Menlo Research

Menlo Research

DevOps Engineer

San Francisco

Sponsorship not specifiedDetected 35 days ago
Node.jsFull-Stack DevelopmentGitPostgreSQLRedisElasticsearchAWSGCPAzureCloud PlatformsDockerKubernetesTerraformCI/CDGitHub ActionsJenkinsLinuxPrometheusGrafanaDevOpsSite Reliability EngineeringKafkaMachine LearningLLMs

About the role

  • You will work deep in the stack across Kubernetes, networking, and where it matters bare metal, and help set the technical direction for how Menlo Cloud scales.

Responsibilities

  • Own the CI/CD and GitOps experience end-to-end: container build pipelines, image optimization, and progressive delivery via ArgoCD / FluxCD.
  • Own the observability stack as a single pane of glass across all clusters: Grafana, Mimir, Tempo, Loki, Pyroscope, OnCall, Prometheus -- and help push toward agent-assisted SRE workflows.
  • Manage and improve our inference platform: vLLM serving and AIBrix for multi-model orchestration and autoscaling across a fleet of NVIDIA GPUs.
  • Manage identity and access via Keycloak integrated with Google Workspace
  • design and maintain hub-and-spoke / multi-AZ topologies.
  • Support training infrastructure: self-service VM provisioning, RunPod burst capacity, Weights and Biases integration.
  • Drive infrastructure reliability, cost efficiency, and capacity planning as the platform scales.
  • Manage identity and access via Keycloak integrated with Google Workspace; harden SSO, RBAC, and secrets management across the platform.
  • Harden network security across private load balancers, firewalls, and VPC segmentation; design and maintain hub-and-spoke / multi-AZ topologies.
  • container build pipelines, image optimization, and progressive delivery via ArgoCD / FluxCD.

Requirements

  • Comfortable with OpenID Connect, SAML, and traditional directory services (LDAP / Active Directory), and you have integrated tools with an IdP like Keycloak, Okta, Azure AD, or equivalent.
  • Strong Linux proficiency (RHEL/Ubuntu or equivalent) including basic performance and networking debugging.
  • Comfort with infrastructure-as-code (Terraform / Terragrunt / Pulumi or equivalent) and configuration management.
  • Optional but valuable: hands-on experience operating any of Kafka, Redis, PostgreSQL, OpenSearch -- at production scale, including HA, backup/restore, and upgrade planning.
  • Experience with OpenStack in production: Nova, Neutron, Cinder, Trove, Horizon, and CLI administration.
  • Experience with KVM virtualization and storage backends like Ceph or Rook-Ceph on Kubernetes.
  • Familiarity with vLLM internals: PagedAttention, continuous batching, tensor parallelism.
  • Experience with KEDA or event-driven autoscaling patterns in anger.

Nice to have

  • Cloud-managed Kubernetes (GKE, EKS, AKS) is fine
  • on-premises / self-managed Kubernetes (kubeadm, Cluster API, k3s, etc.) is a strong plus.
  • Comfortable with VPCs, firewalls, load balancers, private cluster architecture, DNS, and routing.
  • On-premises networking experience (VLANs, BGP, L2/L3 fabrics, pfSense / Fortinet / Palo Alto / Cisco) is a strong plus.
  • Observability -- you have built this before.
  • You have stood up a full observability stack from scratch and operated it in production -- metrics, logs, traces, alerting, on-call.
  • Familiarity with the Grafana stack (Grafana, Mimir, Tempo, Loki, Pyroscope, OnCall, Prometheus) is a strong plus.

Skills

  • hardware architecture, locomotion, autonomy, simulation, and infrastructure.

Company info

  • Menlo Research is an Applied R&D lab building Asimov, an open-source humanoid robot platform, and the full software stack that powers it.
  • Our mission is to make humanoid labor economically viable -- turning software into physical labor at scale.
  • What We're Looking For

This listing is sourced directly from Menlo Research's careers page and normalized into a canonical job model.