Menlo Research
DevOps Engineer
San Francisco
Sponsorship not specifiedDetected 35 days ago
Node.jsFull-Stack DevelopmentGitPostgreSQLRedisElasticsearchAWSGCPAzureCloud PlatformsDockerKubernetesTerraformCI/CDGitHub ActionsJenkinsLinuxPrometheusGrafanaDevOpsSite Reliability EngineeringKafkaMachine LearningLLMs
About the role
- You will work deep in the stack across Kubernetes, networking, and where it matters bare metal, and help set the technical direction for how Menlo Cloud scales.
Responsibilities
- Own the CI/CD and GitOps experience end-to-end: container build pipelines, image optimization, and progressive delivery via ArgoCD / FluxCD.
- Own the observability stack as a single pane of glass across all clusters: Grafana, Mimir, Tempo, Loki, Pyroscope, OnCall, Prometheus -- and help push toward agent-assisted SRE workflows.
- Manage and improve our inference platform: vLLM serving and AIBrix for multi-model orchestration and autoscaling across a fleet of NVIDIA GPUs.
- Manage identity and access via Keycloak integrated with Google Workspace
- design and maintain hub-and-spoke / multi-AZ topologies.
- Support training infrastructure: self-service VM provisioning, RunPod burst capacity, Weights and Biases integration.
- Drive infrastructure reliability, cost efficiency, and capacity planning as the platform scales.
- Manage identity and access via Keycloak integrated with Google Workspace; harden SSO, RBAC, and secrets management across the platform.
- Harden network security across private load balancers, firewalls, and VPC segmentation; design and maintain hub-and-spoke / multi-AZ topologies.
- container build pipelines, image optimization, and progressive delivery via ArgoCD / FluxCD.
Requirements
- Comfortable with OpenID Connect, SAML, and traditional directory services (LDAP / Active Directory), and you have integrated tools with an IdP like Keycloak, Okta, Azure AD, or equivalent.
- Strong Linux proficiency (RHEL/Ubuntu or equivalent) including basic performance and networking debugging.
- Comfort with infrastructure-as-code (Terraform / Terragrunt / Pulumi or equivalent) and configuration management.
- Optional but valuable: hands-on experience operating any of Kafka, Redis, PostgreSQL, OpenSearch -- at production scale, including HA, backup/restore, and upgrade planning.
- Experience with OpenStack in production: Nova, Neutron, Cinder, Trove, Horizon, and CLI administration.
- Experience with KVM virtualization and storage backends like Ceph or Rook-Ceph on Kubernetes.
- Familiarity with vLLM internals: PagedAttention, continuous batching, tensor parallelism.
- Experience with KEDA or event-driven autoscaling patterns in anger.
Nice to have
- Cloud-managed Kubernetes (GKE, EKS, AKS) is fine
- on-premises / self-managed Kubernetes (kubeadm, Cluster API, k3s, etc.) is a strong plus.
- Comfortable with VPCs, firewalls, load balancers, private cluster architecture, DNS, and routing.
- On-premises networking experience (VLANs, BGP, L2/L3 fabrics, pfSense / Fortinet / Palo Alto / Cisco) is a strong plus.
- Observability -- you have built this before.
- You have stood up a full observability stack from scratch and operated it in production -- metrics, logs, traces, alerting, on-call.
- Familiarity with the Grafana stack (Grafana, Mimir, Tempo, Loki, Pyroscope, OnCall, Prometheus) is a strong plus.
Skills
- hardware architecture, locomotion, autonomy, simulation, and infrastructure.
Company info
- Menlo Research is an Applied R&D lab building Asimov, an open-source humanoid robot platform, and the full software stack that powers it.
- Our mission is to make humanoid labor economically viable -- turning software into physical labor at scale.
- What We're Looking For
Apply directly at Menlo Research →Create a free account for alerts like thisView Menlo Research immigration profile
This listing is sourced directly from Menlo Research's careers page and normalized into a canonical job model.