Fireworks AI

Fireworks AI

MTS, Research Engineer

New York, NY; San Mateo, CA

Sponsorship not specified$250k-$400kDetected 14 days ago
PythonC++Distributed SystemsAlgorithmsMachine LearningDeep LearningTensorFlowPyTorchLLMsResearchCommunication

About the role

  • We are looking for a Research Engineer to join our team, operating at the critical intersection of model research and training infrastructure.
  • The most significant advances in deep learning require massive scale.
  • We need engineers who are as comfortable reasoning about gradient descent and loss landscapes as they are about distributed systems, GPU cluster utilization, and data pipelines.

Responsibilities

  • Conduct Open-Ended Research: Explore new model architectures, training objectives, and optimization techniques. Formulate hypotheses, design experiments, and iterate quickly based on empirical results.
  • Collaborate Cross-Functionally: Work closely with Research Scientists to unblock their experiments by providing tooling, optimizing code, and co-designing experiments that are hardware-aware.
  • Formulate hypotheses, design experiments, and iterate quickly based on empirical results.
  • Build What's Next: Work with bleeding-edge technology that impacts how businesses and developers harness AI globally.
  • Learn from the Best: Collaborate with world-class engineers and AI researchers who thrive on curiosity and innovation.
  • Optimize training loops, data loaders, and communication overhead across large GPU clusters.

Requirements

  • Experience working with large distributed systems and parallel computing (e.g., CUDA, NCCL, MPI).

Nice to have

  • Experience with low-level GPU programming (CUDA/Triton) or hardware co-design.
  • Familiarity with the challenges of training Large Language Models (LLMs)
  • Familiarity with the challenges of inference, and OSS inference engines such as SGLang and vLLM
  • Why Fireworks AI?
  • Solve Hard Problems: Tackle challenges at the forefront of AI infrastructure, from low-latency inference to scalable model serving.
  • Ownership & Impact: Join a fast-growing, passionate team where your work directly shapes the future of AI-no bureaucracy, just results.
  • Fireworks AI is an equal-opportunity employer.
  • We celebrate diversity and are committed to creating an inclusive environment for all innovators.

Compensation

  • $250,000 - $400,000 USD

Benefits

  • Reproduce and Extend State-of-the-Art: Implement and reproduce results from recent machine learning papers.
  • Build and Scale Training Infrastructure: Design, implement, and maintain high-performance, distributed machine learning systems.
  • Deep practical knowledge of machine learning frameworks (PyTorch, JAX, or TensorFlow).
  • A proven track record of implementing complex deep learning algorithms from scratch.
  • Range (Plus Equity)

Company info

  • At Fireworks, we're building the future of generative AI infrastructure.
  • Our platform delivers the highest-quality models with the fastest and most scalable inference in the industry.
  • We've been independently benchmarked as the leader in LLM inference speed and are driving cutting-edge innovation through projects like our own function calling and multimodal models.
  • Fireworks is a Series C company valued at $4 billion and backed by top investors including Benchmark, Sequoia, Lightspeed, Index, and Evantic.
  • We're an ambitious, collaborative team of builders, founded by veterans of Meta PyTorch and Google Vertex AI.
  • About the Role
  • In this role, your time will be split between tackling open-ended research problems-such as designing novel architectures and improving algorithmic efficiency - and building the distributed training systems required to make those research breakthroughs a reality.
  • You won't just be handed a paper to implement; you will be expected to reproduce state-of-the-art results from the literature, identify their limitations, and build the infrastructure needed to push beyond them.
  • Reproduce and Extend State-of-the-Art: Implement and reproduce results from recent machine learning papers. Identify bottlenecks, propose improvements, and scale these methods to larger datasets and models.
  • Build and Scale Training Infrastructure: Design, implement, and maintain high-performance, distributed machine learning systems. Optimize training loops, data loaders, and communication overhead across large GPU clusters.

This listing is sourced directly from Fireworks AI's careers page and normalized into a canonical job model.