Together AI
Forward Deployed Engineer (GPU Clusters)
San Francisco · Full-time
Sponsorship not specified$270k-$300kDetected 75 days ago
PythonBashNode.jsAlgorithmsKubernetesAnsibleSASResearchCommunication
About the role
- As a Forward Deployed Engineer (FDE) focused on large scale GPU clusters, you will be a hands-on technical partner to our strategic customers - the world's leading AI model builders.
- As key contributors to the CX, Engineering, and Sales organizations, FDEs add tremendous value by ensuring we can meet the requirements of our most complex POCs, facilitate successful platform adoption for our strategic customers, and guide tailored optimization efforts - directly impacting company growth and the hardening of our core platform.
Responsibilities
- Cluster Hardening & Validation: Design and execute rigorous pre-handover test suites (NCCL, DCGM, GPU Burn) to ensure clusters are stable under the extreme stress of multi-node training.
- Opinionated Onboarding: Build reference designs and "out-of-the-box" configurations for training frameworks to reduce customer time-to-train.
- Benchmarking & Migration: Lead complex benchmarking exercises to demonstrate the performance impact of migrating to new hardware families or Together AI's optimized infrastructure.
- Design and execute rigorous pre-handover test suites (NCCL, DCGM, GPU Burn) to ensure clusters are stable under the extreme stress of multi-node training.
- You will partner with our SAs as a deep-domain specialist in large-scale infrastructure, storage, high-performance networking, and cluster orchestration.
Requirements
- Orchestration Mastery: Deep, hands-on experience with Kubernetes (specifically GPU-operator and device plugins) and/or SLURM for workload scheduling.
- Networking & Interconnects: Expert knowledge of InfiniBand, RoCE, and NVLink
- ability to diagnose network failures that degrade collective communication (NCCL).
- Coding & Automation: Proficiency in Python and shell scripting
- experience with Ansible or similar tools for automated cluster configuration.
- Deep, hands-on experience with Kubernetes (specifically GPU-operator and device plugins) and/or SLURM for workload scheduling.
- Expert knowledge of InfiniBand, RoCE, and NVLink
Nice to have
- Familiarity with parallel file systems (VAST or Weka preferred) and object storage, specifically in the context of large-scale checkpointing.
- Storage Knowledge: Familiarity with parallel file systems (VAST or Weka preferred) and object storage, specifically in the context of large-scale checkpointing.
Skills
- About Together AI
- Together AI is a research-driven artificial intelligence company.
- Compensation
- The US base salary range for this full-time position is: $270,000 - $300,000 OTE + equity + benefits.
- Individual compensation will be determined by experience, skills, and job-related knowledge.
- Equal Opportunity
Compensation
- We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work.
- The US base salary range for this full-time position is: $270,000 - $300,000 OTE + equity + benefits.
- Our salary ranges are determined by location, level and role.
- Individual compensation will be determined by experience, skills, and job-related knowledge.
Equal opportunity
- Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
Visa & Work Authorization
- t opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more
Apply directly at Together AI →Create a free account for alerts like thisView Together AI immigration profile
This listing is sourced directly from Together AI's careers page and normalized into a canonical job model.