FieldAI

FieldAI

Staff ML Systems Engineer, Distributed Systems

Seattle, WA · Staff+

Sponsorship not specified$195k-$230kDetected 53 days ago
PythonNode.jsDistributed SystemsKubernetesMachine LearningPyTorchSparkData EngineeringNLPAI OrchestrationDesign SystemsSystems EngineeringRoboticsResearchCommunication

About the role

  • FieldAI's Irvine team is where embodied AI meets real robots, real sensors, and real field deployments.
  • If you want your work to ship, get tested on hardware, and improve through real deployments, Irvine is the place.
  • Base pay may vary based on role scope, job-related knowledge, skills, experience, and the Irvine, California market.

Responsibilities

  • You will collaborate with a world-class team that thrives on creativity, resilience, and bold thinking.
  • Our teams span AI, software, robotics engineering, product, field deployment, and technical communication, all focused on shipping systems that perform in the real world.
  • Our headquarters is in Irvine, and we partner closely with teams there as well as colleagues across the US and around the world.
  • Develop reusable abstractions, frameworks, and libraries that simplify distributed pipeline development.
  • Optimize performance across distributed CPU and GPU environments, improving throughput, utilization, and reliability.
  • Design systems that effectively manage data partitioning, memory utilization, serialization overhead, and compute efficiency.
  • Partner closely with ML engineers, data engineers, and infrastructure teams to productionize research workflows and enable large-scale model development.
  • Evaluate and guide decisions around distributed computing frameworks, infrastructure technologies, and system design trade-offs.
  • Strong system design skills with demonstrated expertise in scalability, reliability, and performance optimization.
  • Ability to work cross-functionally and drive technical decisions across multiple teams.

Requirements

  • Strong Python programming skills, including experience with concurrency, performance optimization, and systems development.
  • Experience with distributed computing frameworks such as Ray, Spark, Dask, Flink, or similar technologies.
  • Experience diagnosing and resolving bottlenecks in distributed environments.
  • Familiarity with modern ML frameworks such as PyTorch or TensorFlow.
  • Experience with multi-node or multi-GPU training architectures, including DDP, FSDP, DeepSpeed, or similar technologies.
  • Experience operating Kubernetes-based infrastructure and large-scale cloud systems.
  • Experience with distributed debugging, observability, and workflow orchestration platforms.
  • Proven ability to establish technical direction and influence architecture across organizations.

Compensation

  • Our salary range is generous and we consider each individual's background and experience when determining final compensation.

Benefits

  • Design and build scalable distributed machine learning pipelines across data processing, model training, evaluation, and post-processing workflows.
  • Establish best practices and engineering standards for distributed machine learning infrastructure.
  • 5+ years of experience building distributed systems, backend infrastructure, machine learning platforms, or large-scale data processing systems.
  • Experience designing and scaling data pipelines or machine learning workflows.
  • Experience building infrastructure for machine learning training and inference systems.

Company info

  • The Extras That Set You Apart

This listing is sourced directly from FieldAI's careers page and normalized into a canonical job model.