Together AI
Technical Account Manager (TAM), AI Factory
San Francisco · Full-time
Sponsorship not specified$260k-$290kDetected 75 days ago
PythonBashNode.jsAlgorithmsPrometheusGrafanaSite Reliability EngineeringProject ManagementAccount ManagementInventory ManagementCustomer SuccessResearchCommunication
About the role
- As a Dedicated AI Factory TAM at Together AI, you will serve as the named technical owner for one of our most strategic enterprise relationships.
- You will be the primary technical point of contact across all infrastructure domains - compute, networking, storage, and facilities - ensuring smooth delivery and operational health of large-scale GPU deployments.
- This role sits at the intersection of deep infrastructure expertise and high-stakes customer partnership, making you a critical driver of both customer success and company growth.
Responsibilities
- Drive structured engagement through regular cadences including status reporting, technical steering meetings, and executive business reviews
- Lead issue lifecycle management, escalation, and RCA authorship across all infrastructure domains in partnership with Support, SRE, DC Ops, and Engineering teams
- Maintain deep technical expertise across the customer's infrastructure stack - GPU compute, high-speed fabric, and large-scale storage systems - advising on configuration, operational best practices, and incident resolution
Requirements
- 5+ years in a customer-facing technical role, with 2+ years in dedicated technical account management or solutions architecture for large-scale AI or HPC infrastructure
- Hands-on experience with large-scale Ethernet and InfiniBand fabric architecture
- Working knowledge of enterprise storage systems, including high-density NVMe, parallel file systems, and metadata infrastructure
- Experience with DC operations, facilities coordination, and hosting provider SLA management
- Proficiency in infrastructure monitoring and observability tooling (Prometheus, Grafana, or equivalent)
- Deep expertise in GPU infrastructure - GPU health diagnostics, RMA workflows, and hardware acceptance testing
- Strong ownership mindset for incident management, RCA authorship, and executive-level customer communication
- Proven ability to manage multiple concurrent workstreams with hyperscaler-level rigor and communication standards
Nice to have
- Proficiency in Python, Bash, or infrastructure automation tools preferred
Compensation
- We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work.
- The US base salary range for this full-time position is: $260-290K OTE + equity + benefits.
- Our salary ranges are determined by location, level and role.
- Individual compensation will be determined by experience, skills, and job-related knowledge.
Benefits
- Own end-to-end RMA coordination and hardware lifecycle management, including acceptance testing, spare inventory management, and hardware health reporting for large-scale GPU deployments
- Own the observability strategy for the customer estate, including alert policy definition, dashboard development, and proactive health management across all infrastructure layers
Equal opportunity
- Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
- Equal Opportunity
Visa & Work Authorization
- t opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more
Apply directly at Together AI →Create a free account for alerts like thisView Together AI immigration profile
This listing is sourced directly from Together AI's careers page and normalized into a canonical job model.