FieldAI
Staff ML Systems Engineer, Distributed Systems
Seattle, WA · Staff+
Sponsorship not specified$195k-$230kDetected 53 days ago
PythonNode.jsDistributed SystemsKubernetesMachine LearningPyTorchSparkData EngineeringNLPAI OrchestrationDesign SystemsSystems EngineeringRoboticsResearchCommunication
About the role
- FieldAI's Irvine team is where embodied AI meets real robots, real sensors, and real field deployments.
- If you want your work to ship, get tested on hardware, and improve through real deployments, Irvine is the place.
- Base pay may vary based on role scope, job-related knowledge, skills, experience, and the Irvine, California market.
Responsibilities
- You will collaborate with a world-class team that thrives on creativity, resilience, and bold thinking.
- Our teams span AI, software, robotics engineering, product, field deployment, and technical communication, all focused on shipping systems that perform in the real world.
- Our headquarters is in Irvine, and we partner closely with teams there as well as colleagues across the US and around the world.
- Develop reusable abstractions, frameworks, and libraries that simplify distributed pipeline development.
- Optimize performance across distributed CPU and GPU environments, improving throughput, utilization, and reliability.
- Design systems that effectively manage data partitioning, memory utilization, serialization overhead, and compute efficiency.
- Partner closely with ML engineers, data engineers, and infrastructure teams to productionize research workflows and enable large-scale model development.
- Evaluate and guide decisions around distributed computing frameworks, infrastructure technologies, and system design trade-offs.
- Strong system design skills with demonstrated expertise in scalability, reliability, and performance optimization.
- Ability to work cross-functionally and drive technical decisions across multiple teams.
Requirements
- Strong Python programming skills, including experience with concurrency, performance optimization, and systems development.
- Experience with distributed computing frameworks such as Ray, Spark, Dask, Flink, or similar technologies.
- Experience diagnosing and resolving bottlenecks in distributed environments.
- Familiarity with modern ML frameworks such as PyTorch or TensorFlow.
- Experience with multi-node or multi-GPU training architectures, including DDP, FSDP, DeepSpeed, or similar technologies.
- Experience operating Kubernetes-based infrastructure and large-scale cloud systems.
- Experience with distributed debugging, observability, and workflow orchestration platforms.
- Proven ability to establish technical direction and influence architecture across organizations.
Compensation
- Our salary range is generous and we consider each individual's background and experience when determining final compensation.
Benefits
- Design and build scalable distributed machine learning pipelines across data processing, model training, evaluation, and post-processing workflows.
- Establish best practices and engineering standards for distributed machine learning infrastructure.
- 5+ years of experience building distributed systems, backend infrastructure, machine learning platforms, or large-scale data processing systems.
- Experience designing and scaling data pipelines or machine learning workflows.
- Experience building infrastructure for machine learning training and inference systems.
Company info
- The Extras That Set You Apart
Apply directly at FieldAI →Create a free account for alerts like thisView FieldAI immigration profile
This listing is sourced directly from FieldAI's careers page and normalized into a canonical job model.