Together AI

Together AI

AI Researcher, Core ML (Turbo)

San Francisco · Full-time

Sponsorship not specified$200k-$280kDetected 11 days ago
PythonDistributed SystemsAlgorithmsMachine LearningNLPLLMsA/B TestingResearchLeadership

About the role

  • The Turbo team sits at the intersection of efficient inference (algorithms, architectures, engines) and post‑training / RL systems.
  • Our mandate is to push the frontier of efficient inference and RL‑driven training: making models dramatically faster and cheaper to run, while improving their capabilities through RL‑based post‑training (e.g., GRPO‑style objectives).
  • People on this team are often spiky: some are more RL‑first, some are more systems‑first.

Responsibilities

  • Design and prototype algorithms, architectures, and scheduling strategies for low‑latency, high‑throughput inference.
  • Implement and maintain changes in high‑performance inference engines (e.g., SGLang‑ or vLLM‑style systems and Together's inference stack), including kernel backends, speculative decoding (e.g., ATLAS), quantization, etc.
  • Profile and optimize performance across GPU, networking, and memory layers to improve latency, throughput, and cost.
  • Design and operate RL and post‑training pipelines (e.g., RLHF, RLAIF, GRPO, DPO‑style methods, reward modeling) where 90+% of the cost is inference, jointly optimizing algorithms and systems.
  • Co‑design algorithms and infrastructure so that objectives, rollout collection, and evaluation are tightly coupled to efficient inference, and quickly identify bottlenecks across the training engine, inference engine, data pipeline, and user‑facing layers.
  • Run ablations and scale‑up experiments to understand trade‑offs between model quality, latency, throughput, and cost, and feed these insights back into model, RL, and system design.
  • Own critical systems at production scale
  • Profile, debug, and optimize inference and post‑training services under real production workloads.
  • Drive roadmap items that require real engine modification-changing kernels, memory layouts, scheduling logic, and APIs as needed.
  • Establish metrics, benchmarks, and experimentation frameworks to validate improvements rigorously.

Requirements

  • People on this team typically have deep expertise in one or more areas and enough breadth (or interest) to work effectively across the stack.
  • The closer you are to full‑stack (inference + post‑training/RL + systems), the stronger the fit-but being spiky in one area and eager to grow is absolutely okay.
  • Experience profiling and optimizing performance across GPU, networking, and memory layers.
  • Track record of impactful work in ML systems, RL, or large‑scale model training (papers, open‑source projects, or production systems).
  • You enjoy collaborating with infra, research, and product teams, and you care about both scientific quality and user‑visible wins.
  • 3+ years of experience working on ML systems, large‑scale model training, inference, or adjacent areas (or equivalent experience via research / open source).
  • Advanced degree in Computer Science, EE, or a related field, or equivalent practical experience.

Nice to have

  • some are more RL‑first, some are more systems‑first.

Compensation

  • We offer competitive compensation, startup equity, health insurance and other competitive benefits.
  • The US base salary range for this full-time position is: $200,000 - $280,000 + equity + benefits.
  • Our salary ranges are determined by location, level and role.
  • Individual compensation will be determined by experience, skills, and job-related knowledge.

Equal opportunity

  • Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Visa & Work Authorization

  • t opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more

This listing is sourced directly from Together AI's careers page and normalized into a canonical job model.