UniversalAGI

UniversalAGI

Machine Learning Infrastructure Engineer

San Francisco · Exec

Sponsorship not specifiedDetected 146 days ago
DatabricksKubernetesCI/CDMachine LearningDeep LearningAI OrchestrationANSYSCFDResearchCommunicationCollaboration

About the role

  • You'll work closely with the CEO and founding team to turn research into repeatable, scalable, reliable systems - internally and in customer infrastructure.

Responsibilities

  • Build the foundation platform (internal)
  • Build and operate scalable infrastructure for data generation and simulation workflows (job orchestration, scheduling, queues, retries, observability).
  • Build reproducible pipelines for training/fine-tuning and benchmarking (artifact/version management, experiment tracking, dataset lineage).
  • Own cost/performance tradeoffs across compute, storage, networking, and runtime efficiency.
  • Lead deployments of our stack into customer cloud/on-prem environments, including secure networking, permissions, and data movement.
  • Build robust deployment patterns: environment provisioning, CI/CD, rollbacks, monitoring, and incident response.
  • This is a "ship outcomes" role: your work directly determines how fast we can iterate, how reproducible our results are, and how reliably we deliver in production.
  • UniversalAGI is building OpenAI for Physics.
  • We're building foundation AI models for physics that enable end-to-end industrial automation from initial design through optimization, validation, and production.
  • We're building a high-velocity team of relentless researchers and engineers that will define the next generation of AI for industrial engineering.

Requirements

  • Resourcefulness: you know when to do the "quick & correct" fix vs. when to invest in a robust solution, and you can justify the tradeoff with impact/
  • Ability to earn respect through hands-on technical contribution
  • Hands-on experience building/operating infrastructure for ML/compute-heavy workflows: pipelines, job orchestration, GPU compute, storage, CI/CD, monitoring.
  • Olympic athlete mindset: You have high standards for yourself and are obsessed with measurable improvement on the metrics you are delivering to customers.
  • Ownership: Comfortable owning work end-to-end and being accountable for measurable outcomes.
  • Bonus Qualifications
  • Experience building evaluation/benchmarking frameworks with strong reproducibility guarantees.
  • Cultural Fit

Nice to have

  • Experience with workflow orchestration (e.g., Ray, Kubernetes, Slurm).
  • Experience with GPU infrastructure and distributed training systems.
  • Experience deploying into regulated / security-sensitive environments (gov/defense/enterprise).
  • Experience with simulation/HPC pipelines (CFD, meshing, batch workloads) is a plus but not required.
  • Experience in an FDE-style / delivery execution role (or similar "ship results fast" environments).
  • Technical Respect: Ability to earn respect through hands-on technical contribution
  • Intensity: Thrives in our unusually intense culture - willing to grind when needed
  • Customer Obsession: Passionate about solving real customer problems, not just publishing papers

Skills

  • Strong software engineering skills (clean code, debugging, reliability, reproducibility).
  • Hands-on experience building/operating infrastructure for

Compensation

  • Competitive Salary + Equity

Benefits

  • Competitive compensation and equity.
  • Competitive health, dental, vision benefits paid by the company.
  • Flexible vacation.
  • AI tools stipend.
  • Monthly commute stipend.
  • Monthly wellness / fitness stipend.

Company info

  • AI startup based in San Francisco and backed by Elad Gil (#1 Solo VC), Eric Schmidt (former Google CEO), Prith Banerjee (ANSYS CTO), Ion Stoica (Databricks Founder), Jared Kushner (former Senior Advisor to the President), David Patterson (Turing Award Winner), and Luis Videgaray (former Foreign and Finance Minister of Mexico).
  • If you're passionate about AI, physics, or the future of industrial innovation, we want to hear from you.
  • Can translate complex model decisions to customers and team
  • Deploy to customers (external)

This listing is sourced directly from UniversalAGI's careers page and normalized into a canonical job model.