Ifm Us
Machine Learning Engineer – World Model
Sunnyvale, CA
H1B sponsorship available$150k-$450kDetected 77 days ago
PythonDistributed SystemsGitAWSCloud PlatformsDockerKubernetesPlatform EngineeringKafkaMachine LearningDeep LearningSparkData EngineeringMLOpsA/B TestingResearchProblem Solving
About the role
- We are the AllWorld Team under the Institute of Foundation Model (IFM) at MBZUAI.
- At AllWorld, we are pioneering the development of the PAN (Physical, Agentic, and Networked) world models -the next-generation foundation models to unlock machine intelligence beyond lingual.
- Our mission is to tackle the fundamental challenges of world modeling and establish a new paradigm for next-generation machine reasoning.
Responsibilities
- Design, build, and operate scalable ML infrastructure on AWS (e.g., compute, storage, networking, access control).
- Develop and maintain MLOps workflows for data versioning
- Build and manage distributed systems for large-scale data processing (filtering, captioning, etc.) and model evaluation.
- Own architecture decisions for ML infrastructure and drive best practices in reliability, scalability, and cost efficiency.
- Implement observability across systems, including monitoring, logging, and alerting.
- Partner closely with researchers to translate experimental workflows into robust, scalable systems.
- We are a dedicated research lab for building, understanding, using, and risk-managing foundation models.
- Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.
- You'll build scalable, reliable, and observable cloud infrastructure, working closely with researchers to support data pipelines, experimentation, and evaluation workflows.
Requirements
- 3+ years of experience in MLOps, ML infrastructure, or related backend/platform engineering roles.
- Strong experience with cloud platforms (preferably AWS) and core services for compute, storage, and access control.
- Experience designing and operating distributed systems (e.g., Kubernetes, Ray, or similar frameworks).
- Familiarity with data processing and pipeline orchestration tools (e.g., Spark, Kafka, or similar).
- Experience with observability practices (monitoring, logging, alerting).
- Ability to work closely with researchers and translate ambiguous requirements into production-ready systems.
- Experience in fast-paced or research-driven environments.
- Experience with large-scale video or multimodal data pipelines.
- Knowledge of cost optimization, security, and networking in multi-tenant environments.
- Familiarity with modern developer and AI-assisted coding (e.g., Codex, Cursor, Claude Code)
- Must-Haves
- Solid software engineering skills, including system design, debugging, and testing (Python, Docker, Git).
- Nice-to-Haves
- Experience building automated model evaluation or benchmarking systems.
Compensation
- $150k-$450k
Benefits
- Include *Comprehensive medical, dental, and vision benefits *Bonus *401K Plan *Generous paid time off, sick leave and holidays *Paid Parental Leave *Employee Assistance Program *Life insurance and disability
- We are looking for passionate individuals who share our vision and are eager to push the boundaries of AI together.
- We're looking for a Machine Learning Engineer focused on ML infrastructure and MLOps to design and operate the systems that power our research environment.
Visa & Work Authorization
- This position is eligible for visa sponsorship.
This listing is sourced directly from Ifm Us's careers page and normalized into a canonical job model.