Baseten

Baseten

Engineering Manager, Runtime Fabric

San Francisco

Sponsorship not specifiedDetected 43 days ago
Code ReviewLinuxMachine LearningSparkResearchLeadershipCommunicationMentoring

About the role

  • AI inference is not a general-purpose workload.
  • The tools the industry has relied on for a decade weren't built for this, and patching around those limitations at higher layers only goes so far.
  • That vertical ownership means we can fix these problems at the root.

Responsibilities

  • Recruit, hire, and develop a high-performing team of systems engineers with deep container and Linux expertise.
  • Provide regular coaching, feedback, and career development support to your direct reports.
  • Drive the expansion of BDN's architecture, currently focused on model weights, to container images, training checkpoints, and deployment artifacts.
  • Ensure the team maintains end-to-end ownership of the container startup performance path, from snapshotter initialization through weight delivery to first inference request.
  • Act as the primary advocate for Runtime Fabrics across the organization, ensuring upstream and downstream teams have the integration support they need.
  • Collaborate with product and engineering stakeholders to prioritize investments based on business impact and infrastructure reliability.
  • We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital.
  • Join us and help build the platform engineers turn to to ship AI products.

Requirements

  • Proven experience managing and growing engineering teams in a systems, infrastructure, or low-level runtime context.
  • Deep familiarity with the Linux container ecosystem: containerd, runc, OCI Runtime Spec, Linux namespaces, and cgroups, with the ability to engage credibly in code reviews and architectural discussions.
  • Experience with distributed storage systems, content-addressable storage, or large-scale caching infrastructure.
  • Strong written and verbal communication skills, with the ability to influence without authority across teams.

Nice to have

  • Experience with GPU device access in containers: NVIDIA Container Toolkit, CDI (Container Device Interface), or GPU-aware scheduling.
  • Familiarity with lazy-loading snapshotters (stargz, soci, EROFS/Nydus) or peer-to-peer image distribution.
  • Experience with secure container runtimes (gVisor, Sysbox) or micro-VM technologies (Firecracker, Cloud Hypervisor).
  • Background in multi-tenant infrastructure or security-sensitive serving environments.

Compensation

  • Competitive compensation, including meaningful equity.

Benefits

  • Competitive compensation, including meaningful equity.
  • 100% coverage of medical, dental, and vision insurance for employee and dependents
  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
  • If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.
  • Partner with engineering leadership to define the long-term vision and roadmap for container runtime and storage infrastructure.
  • By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production.

Company info

  • At Baseten, we are committed to fostering a diverse and inclusive workplace.

This listing is sourced directly from Baseten's careers page and normalized into a canonical job model.