Unity Technologies SF
Senior Machine Learning Engineer, ML Infrastructure- Online
Olympia, Washington · Senior
Stay score
odds of building a lasting career here
Sponsors, but it's cap-subject — you still face the weighted lottery (~61% per draw at Level IV). Good if you win; have a cap-exempt backup on your list.
Lottery odds assume a STEM candidate.
Personalize to your clock →H-1B wage level
the lottery is wage-weighted — each level is one more entry
This range already reaches Level IV — the maximum four lottery entries.
DOL prevailing wage, 2026-27 wage year · Software Developers (15-1252) · Olympia-Lacey-Tumwater, WA. Wage level is derived by USCIS from the offered wage, occupation and worksite; the occupation shown is inferred from the job title.
Employer immigration record
from this employer's Department of Labor filings
Green-card filing pattern in this occupation
Files H-1B transfers
Sourced from Department of Labor LCA, PERM and prevailing-wage disclosure data. Employer matching is by name, so figures may be split across an employer's legal entities. Absence of a filing means none appears in our copy of the data, not that none exists.
Community outcomes
No reports yet — be the first to help the next applicant.
About the role
- You will work closely with ML engineers, platform teams, and product stakeholders to ensure models can be deployed, scaled, monitored, and iterated on efficiently.
- You will play a key role in shaping how models are packaged, served, validated, monitored, and optimized in production environments.
Responsibilities
- Design and operate large-scale online inference infrastructure that serves production ML models with low latency and high reliability, such as PyTorch, Triton Inference Server, Kubernetes, GKE, Ray, or similar distributed serving frameworks.
- Develop infrastructure that supports distributed training workflows using technologies such as Pytorch, Ray Data, and Ray Train, etc.
- Optimize model performance through model compilation, GPU/CPU utilization improvements, request scheduling, kernel fusion, and runtime-level tuning.
- Partner closely with ML engineers to support faster model iteration while maintaining production safety, scalability, and cost efficiency.
- Lead architectural improvements that make the online ML platform more robust, user-friendly, scalable, and cost-efficient.
- Relocation support is not available for this position
- Our differences are strengths that enable us to support the growing and evolving needs of our customers, partners, and collaborators.
Requirements
- Experience with model serving frameworks such as NVIDIA Triton Inference Server, TorchServe, Ray Serve, TensorFlow Serving, or similar systems.
- Experience optimizing inference workloads using techniques such as dynamic batching, model compilation, quantization, GPU acceleration, GPU kernel optimization, caching, or runtime tuning.
- Strong experience with distributed systems, Kubernetes, autoscaling, service reliability, and production observability.
- Strong programming skills in Python, with practical experience working on production ML systems and high-scale services.
- Experience with PyTorch and modern model deployment workflows, including model packaging, validation, and serving lifecycle management.
- Experience designing infrastructure for safe model rollout, canary testing, A/B experimentation, and automated rollback.
- Strong systems thinking, with the ability to reason about latency, throughput, reliability, scalability, and cost tradeoffs in online systems.
Compensation
- This range reflects the anticipated base salary for this position.
Benefits
- We offer a wide range of benefits designed to support well-being and work-life balance.
- Please note: Benefits eligibility, specific offerings, and coverage vary based on the country and employment status.
- Unity also enables teams across industries like automotive, manufacturing, and healthcare to design, simulate, and collaborate in 3D - closing the gap between ideas and reality.
- If you have a disability that means there are preparations or accommodations we can make to help ensure you have a comfortable and positive interview experience, please fill out this form to let us know.
- Improve observability of ML systems through latency, throughput, error-rate, cost, saturation, and model-health monitoring.
Company info
- This posting is intended to fill an existing vacancy, and we are committed to providing applicants with updates throughout the hiring process in accordance with applicable law.
- We are seeking a Senior ML engineer to design and evolve Unity Vector's online model inference platform.
- At Unity, we want our team members to thrive.
Equal opportunity
- equal opportunity employer.
Visa & Work Authorization
- Work visa/immigration sponsorship is not available for this position
This listing is sourced directly from Unity Technologies SF's careers page and normalized into a canonical job model.