Judgmentlabs

Judgmentlabs

Senior Infrastructure Engineer

San Francisco · Senior · Full-time

Sponsorship not specifiedDetected 191 days ago
Distributed SystemsCode ReviewDatabricksAWSGCPAzureCloud PlatformsKubernetesTerraformCI/CDDatadogTemporalKafkaRabbitMQLLMsAI OrchestrationIncident ResponseDNS

About the role

  • This role focuses on cloud/platform infrastructure: Terraform, EKS, ArgoCD/Kargo, IAM, DNS, observability, CI/CD, multi-region architecture, BYOC, self-hosted deployments, private connectivity, and enterprise-grade reliability.
  • You'll work on the systems that keep high-throughput telemetry ingestion, ClickHouse, RabbitMQ, Temporal, evaluation workers, and customer-facing services running under real production load.
  • Enterprise-grade deployment architecture.

Responsibilities

  • Own cloud infrastructure for production services across Terraform, EKS, ArgoCD/Kargo, IAM, DNS, networking, metrics, CI/CD, and deployment automation.
  • Build and operate infrastructure for trace ingestion, evaluation workers, RabbitMQ, Temporal, ClickHouse, and the systems that support Judgment's core product.
  • Design multi-region and enterprise deployment architectures, including data residency, automatic failover, disaster recovery, and customer-managed environments.
  • Build secure deployment patterns for BYOC, self-hosted, private-network, and restricted environments.
  • Implement private connectivity, identity integrations, network isolation, encryption patterns, and enterprise security requirements.
  • Partner with backend engineers on reliability, scaling limits, queue behavior, storage growth, ingestion throughput, and production incidents.
  • Raise the bar for infrastructure quality through design docs, code reviews, operational rigor, and clean abstractions.
  • Strong experience designing, building, and operating production cloud infrastructure for real customer-facing systems.
  • Judgment builds the infrastructure to do that.

Requirements

  • Experience with Kubernetes / EKS or similar orchestration systems, infrastructure-as-code, CI/CD, cloud networking, IAM, DNS, and production observability.
  • Ability to reason about reliability, security, deployment ergonomics, and developer velocity at the same time.
  • You can write architecture proposals, operational runbooks, incident notes, and crisp tradeoff docs.

Nice to have

  • Experience with Terraform, EKS, ArgoCD, Kargo, AWS networking, IAM, DNS, and production metrics/logging systems.
  • Experience operating ClickHouse, RabbitMQ, Temporal, Kafka, or other stateful production infrastructure.
  • Experience with private connectivity such as AWS PrivateLink, Azure Private Link, or GCP Private Service Connect.
  • Experience with SSO, SAML, SCIM, secrets management, encryption, network isolation, and enterprise security reviews.
  • Experience with observability infrastructure, telemetry ingestion, or platforms like Datadog, Honeycomb, Sentry, or similar systems.
  • Experience supporting AI infrastructure, LLM evaluation workloads, or high-throughput event pipelines.
  • As agents move from demos to production, the bottleneck is no longer just better prompts.
  • It is turning real production experience into high-quality data for evals, labeling, rubric generation, context engineering, and RL workflows.

Skills

  • San Francisco · On Site · Full Time
  • The next generation of agents will not improve from prompts alone.

Benefits

  • Instead of only showing teams what happened, Judgment helps decide what matters, what should be learned from, and how that learning should flow back into the agent.
  • Judgment is building the learning infrastructure for agents.
  • We've raised $30M+ from Lightspeed, SV Angel, Valor Equity Partners, and others.
  • Make deployments safer and faster through automation, rollout strategies, environment parity, CI reliability, e2e test health, and better internal tooling.

Company info

  • Work directly with customers when deployment, networking, security, or production environment constraints are the blocker.
  • Multi-region reliability. Design failover, disaster recovery, data residency, and deployment patterns for customers that cannot tolerate downtime or ambiguous data movement.

This listing is sourced directly from Judgmentlabs's careers page and normalized into a canonical job model.