Fireworks AI
Member of Technical Staff
New York, NY · Staff+
Sponsorship not specified$175k-$220kDetected 37 days ago
TypeScriptPythonGoC++Backend DevelopmentPostgreSQLMySQLDynamoDBAWSGCPAzureCloud PlatformsDockerKubernetesDevOpsgRPCKafkaMachine LearningPyTorchSparkLLMsA/B TestingResearchCommunication
About the role
- As a Training Infrastructure Engineer, you'll design, develop, and maintain large-scale backend and cloud-native infrastructure to support distributed machine learning training, inference, and data processing pipelines for our generative AI platform.
- You'll architect scalable, resilient backend infrastructure, lead technical design discussions, mentor engineers, and establish best practices for large-scale machine learning systems.
Responsibilities
- Architect and build scalable, resilient backend infrastructure to support distributed training, inference, and data processing pipelines
- Design and implement core backend services with a focus on efficiency and low latency
- Drive infrastructure optimization initiatives for compute cost, storage lifecycle management, and network performance
- Own end-to-end systems from design to deployment, emphasizing reliability, fault tolerance, and operational excellence
- Build What's Next: Work with bleeding-edge technology that impacts how businesses and developers harness AI globally.
- Learn from the Best: Collaborate with world-class engineers and AI researchers who thrive on curiosity and innovation.
Requirements
- 4 years of experience with major server-side programming languages and frameworks (e.g., Python, C++, Go, TypeScript)
- 4 years of experience writing technical design documentation, leading cross-functional projects, and collaborating with cross-functional teams to achieve business impact
- 3 years of experience developing and maintaining data processing and API systems, including client-server communication frameworks (e.g., gRPC, Thrift)
- 3 years of experience conducting A/B testing and scientific experimentation (e.g., Statsig, Meta Deltoid, Optimizely) to measure software impact
- 3 years of experience conducting coding interviews and providing systematic feedback for engineering candidates
- 2 years of experience with cloud-native tools and infrastructure, such as Docker and Kubernetes
- 2 years of experience defining and implementing data-driven metrics to support company or team goals
Nice to have
- Bachelor's degree or equivalent in Computer Science or related field plus four (4) years of experience in software engineering or related role
Compensation
- $175,000 - $220,000 USD
Benefits
- Lead technical design discussions, mentor engineers, and establish best practices for large-scale machine learning systems
- Collaborate with machine learning, DevOps, and product teams to translate research and product requirements into robust infrastructure solutions
- Range (Plus Equity)
- Total compensation for this role also includes meaningful equity in a fast-growing startup, along with a competitive salary and comprehensive benefits package.
- Base Pay Range (Plus Equity)
Company info
- At Fireworks, we're building the future of generative AI infrastructure. Our platform delivers the highest-quality models with the fastest and most scalable inference in the industry. We've been independently benchmarked as the leader in LLM inference speed and are driving cutting-edge innovation through projects like our own function calling and multimodal models. Fireworks is a Series C company valued at $4 billion and backed by top investors including Benchmark, Sequoia, Lightspeed, Index, and Evantic. We're an ambitious, collaborative team of builders, founded by veterans of Meta PyTorch and Google Vertex AI.
- The Role:
- As a Training Infrastructure Engineer, you'll design, develop, and maintain large-scale backend and cloud-native infrastructure to support distributed machine learning training, inference, and data processing pipelines for our generative AI platform. You'll architect scalable, resilient backend infrastructure, lead technical design discussions, mentor engineers, and establish best practices for large-scale machine learning systems.
- At Fireworks, we're building the future of generative AI infrastructure.
- Our platform delivers the highest-quality models with the fastest and most scalable inference in the industry.
- We've been independently benchmarked as the leader in LLM inference speed and are driving cutting-edge innovation through projects like our own function calling and multimodal models.
- Fireworks is a Series C company valued at $4 billion and backed by top investors including Benchmark, Sequoia, Lightspeed, Index, and Evantic.
- We're an ambitious, collaborative team of builders, founded by veterans of Meta PyTorch and Google Vertex AI.
Apply directly at Fireworks AI →Create a free account for alerts like thisView Fireworks AI immigration profile
This listing is sourced directly from Fireworks AI's careers page and normalized into a canonical job model.