Material World

Material World

Infrastructure Engineer

United States - Remote

Sponsorship not specifiedDetected 69 days ago
PythonGoRustNode.jsDistributed SystemsKubernetesTerraformAnsibleHelmLinuxSite Reliability EngineeringIncident ResponseAccounting

About the role

  • Your scope spans bare-metal provisioning, multi-tenant Kubernetes, SLURM scheduling, control planes, and the automation and observability that keep thousands of compute nodes running as a single production system.

Responsibilities

  • Build the control plane and APIs that unify our compute fleet
  • Build automation that eliminates manual operations
  • Drive reliability, observability, and incident response across the fleet
  • Relocation support

Requirements

  • 5+ years in infrastructure, platform, or backend engineering

Skills

  • Deep understanding of Linux, storage, and distributed systems
  • Rust, Go, or Python
  • Experience with workload schedulers: SLURM, Kubernetes scheduling, or equivalent
  • Expertise with automation tooling: Terraform, Ansible, Helm
  • Experience architecting multi-tenant systems
  • Production SRE experience: on-call, incident response, observability
  • Top-tier compensation structured to recognize and retain the best talent
  • Meaningful equity
  • Comprehensive medical, dental, vision, life, and disability insurance
  • Parental leave for all new parents, including adoptive and surrogate journeys
  • Flexible PTO
  • Paid Holidays

Compensation

  • Top-tier compensation structured to recognize and retain the best talent

Benefits

  • Own provisioning and lifecycle from rack bring-up to node retirement

Equal opportunity

  • We're an Equal Opportunity Employer and do not discriminate on the basis of any protected status under applicable law.

This listing is sourced directly from Material World's careers page and normalized into a canonical job model.