Moonlite AI
Sr. Site Reliability Engineer (SRE)
Chicago, IL or Remote · Senior
Sponsorship not specified$165k-$225kDetected 62 days ago
PythonGoBashKubernetesTerraformAnsibleHelmLinuxPrometheusGrafanaDevOpsSite Reliability EngineeringPlatform EngineeringCybersecurityIncident ResponseFinancial ModelingDNSLoad BalancingResearchLeadershipCommunicationCollaborationProblem Solving
About the role
- Working closely with our systems engineers, network engineers, and platform engineering team, you'll architect and operate the Kubernetes infrastructure that powers our control plane and orchestrates compute, storage, and networking at scale.
- You'll ensure enterprise-grade reliability while establishing the automation, observability, and operational practices.
- Configure CNI plugins and network segmentation for research workloads.
Responsibilities
- Design, build, and operate production Kubernetes clusters on bare-metal infrastructure - including cluster bootstrapping, control plane architecture, etcd management, and scaling strategies for high-performance compute workloads.
- Implement and operate custom Kubernetes networking solutions with SR-IOV for high-performance GPU interconnects, multi-tenancy isolation and advanced networking policies.
- Develop and maintain custom Kubernetes operators and controllers for bare-metal provisioning, infrastructure lifecycle management, and resource orchestration across compute, storage, and networking domains.
- Build deep integrations between Kubernetes and underlying infrastructure including CSI drivers for storage, custom admission controllers for policy enforcement, and scheduling extensions for specialized hardware placement.
- Design and implement automation using Terraform, Ansible, Helm, and custom operators to orchestrate infrastructure workflows and enable deployments across multiple regions.
- Manage production bare-metal infrastructure across multiple regions.
- Build systems ensuring high availability, fault tolerance, and graceful degradation - establishing SLIs, SLOs, and monitoring to meet enterprise reliability commitments.
- Build comprehensive monitoring, logging, and alerting using Prometheus, Grafana, and ELK stack.
- Lead incident response, conduct postmortems, and implement preventative measures to improve reliability and reduce MTTR.
- Build and operate infrastructure that supports mission-critical research and AI workloads for leading financial institutions and research organizations.
Requirements
- Kubernetes Internals & Integration: Strong understanding of Kubernetes internals including custom resource definitions (CRDs), operators, controllers, admission webhooks, and scheduling.
- Experience integrating storage (CSI drivers), networking (CNI, SR-IOV), and specialized hardware (GPU device plugins) with Kubernetes.
- Linux Systems Experience: Strong fundamentals in Linux systems administration, performance tuning, troubleshooting, and automation in production environments.
- Infrastructure Automation: Proficiency with infrastructure-as-code tools (Terraform, Ansible, Helm) and building automation to reduce operational overhead.
- Networking Fundamentals: Solid understanding of networking concepts including IPAM, DNS, DHCP, VLAN/VXLAN, routing, load balancing, and experience troubleshooting network issues in production.
- Collaboration & Communication: Excellent communication skills and ability to work across teams including systems engineers, network engineers, and software developers.
- Proficiency with infrastructure-as-code tools (Terraform, Ansible, Helm) and building automation to reduce operational overhead.
- Excellent communication skills and ability to work across teams including systems engineers, network engineers, and software developers.
Nice to have
- Deep familiarity with Kubernetes networking (Calico, Cilium, Multus), service mesh technologies, and network policy management
- Experience with GPU workload orchestration including NVIDIA GPU Operator, MIG, time-slicing, and device plugins
- Background with advanced Kubernetes features including custom schedulers, admission controllers, and API server extensions
- Experience with Kubernetes cluster federation or multi-cluster management
- Knowledge of high-performance networking technologies (InfiniBand, RDMA, RoCE) and their integration with Kubernetes
- Experience with enterprise storage systems (VAST, Lightbits, Ceph, or similar)
- Familiarity with configuration management at scale and GitOps practices
- Understanding of security best practices for Kubernetes and bare-metal infrastructure
Skills
- 5+ years in SRE, DevOps, or infrastructure engineering roles with proven experience operating production infrastructure at scale.
- Strong scripting skills in Go, Python, or Bash for automation, tooling development, and operational efficiency.
Compensation
- We offer a competitive total compensation package combining a competitive base salary, startup equity, and industry-leading benefits.
Benefits
- We offer a competitive total compensation package combining a competitive base salary, startup equity, and industry-leading benefits.
- The total compensation range for this role is $165,000 - $225,000, which includes both base salary and equity.
- We provide generous benefits, including a 6% 401(k) match, fully covered health insurance premiums, and other comprehensive offerings to support your well-being and success as we grow together.
Apply directly at Moonlite AI →Create a free account for alerts like thisView Moonlite AI immigration profile
This listing is sourced directly from Moonlite AI's careers page and normalized into a canonical job model.