Baseten
Engineering Manager, Runtime Fabric
San Francisco
Sponsorship not specifiedDetected 43 days ago
Code ReviewLinuxMachine LearningSparkResearchLeadershipCommunicationMentoring
About the role
- AI inference is not a general-purpose workload.
- The tools the industry has relied on for a decade weren't built for this, and patching around those limitations at higher layers only goes so far.
- That vertical ownership means we can fix these problems at the root.
Responsibilities
- Recruit, hire, and develop a high-performing team of systems engineers with deep container and Linux expertise.
- Provide regular coaching, feedback, and career development support to your direct reports.
- Drive the expansion of BDN's architecture, currently focused on model weights, to container images, training checkpoints, and deployment artifacts.
- Ensure the team maintains end-to-end ownership of the container startup performance path, from snapshotter initialization through weight delivery to first inference request.
- Act as the primary advocate for Runtime Fabrics across the organization, ensuring upstream and downstream teams have the integration support they need.
- Collaborate with product and engineering stakeholders to prioritize investments based on business impact and infrastructure reliability.
- We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital.
- Join us and help build the platform engineers turn to to ship AI products.
Requirements
- Proven experience managing and growing engineering teams in a systems, infrastructure, or low-level runtime context.
- Deep familiarity with the Linux container ecosystem: containerd, runc, OCI Runtime Spec, Linux namespaces, and cgroups, with the ability to engage credibly in code reviews and architectural discussions.
- Experience with distributed storage systems, content-addressable storage, or large-scale caching infrastructure.
- Strong written and verbal communication skills, with the ability to influence without authority across teams.
Nice to have
- Experience with GPU device access in containers: NVIDIA Container Toolkit, CDI (Container Device Interface), or GPU-aware scheduling.
- Familiarity with lazy-loading snapshotters (stargz, soci, EROFS/Nydus) or peer-to-peer image distribution.
- Experience with secure container runtimes (gVisor, Sysbox) or micro-VM technologies (Firecracker, Cloud Hypervisor).
- Background in multi-tenant infrastructure or security-sensitive serving environments.
Compensation
- Competitive compensation, including meaningful equity.
Benefits
- Competitive compensation, including meaningful equity.
- 100% coverage of medical, dental, and vision insurance for employee and dependents
- Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
- If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.
- Partner with engineering leadership to define the long-term vision and roadmap for container runtime and storage infrastructure.
- By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production.
Company info
- At Baseten, we are committed to fostering a diverse and inclusive workplace.
Apply directly at Baseten →Create a free account for alerts like thisView Baseten immigration profile
This listing is sourced directly from Baseten's careers page and normalized into a canonical job model.