Material World
Infrastructure Engineer
United States - Remote
Sponsorship not specifiedDetected 69 days ago
PythonGoRustNode.jsDistributed SystemsKubernetesTerraformAnsibleHelmLinuxSite Reliability EngineeringIncident ResponseAccounting
About the role
- Your scope spans bare-metal provisioning, multi-tenant Kubernetes, SLURM scheduling, control planes, and the automation and observability that keep thousands of compute nodes running as a single production system.
Responsibilities
- Build the control plane and APIs that unify our compute fleet
- Build automation that eliminates manual operations
- Drive reliability, observability, and incident response across the fleet
- Relocation support
Requirements
- 5+ years in infrastructure, platform, or backend engineering
Skills
- Deep understanding of Linux, storage, and distributed systems
- Rust, Go, or Python
- Experience with workload schedulers: SLURM, Kubernetes scheduling, or equivalent
- Expertise with automation tooling: Terraform, Ansible, Helm
- Experience architecting multi-tenant systems
- Production SRE experience: on-call, incident response, observability
- Top-tier compensation structured to recognize and retain the best talent
- Meaningful equity
- Comprehensive medical, dental, vision, life, and disability insurance
- Parental leave for all new parents, including adoptive and surrogate journeys
- Flexible PTO
- Paid Holidays
Compensation
- Top-tier compensation structured to recognize and retain the best talent
Benefits
- Own provisioning and lifecycle from rack bring-up to node retirement
Equal opportunity
- We're an Equal Opportunity Employer and do not discriminate on the basis of any protected status under applicable law.
Apply directly at Material World →Create a free account for alerts like thisView Material World immigration profile
This listing is sourced directly from Material World's careers page and normalized into a canonical job model.