Judgmentlabs

Judgmentlabs

Senior Backend Engineer

San Francisco · Senior · Full-time

Sponsorship not specifiedDetected 36 days ago
Next.jsDistributed SystemsCode ReviewDatabricksTemporalAPI DevelopmentKafkaRabbitMQSparkAirflowData EngineeringLLMs

About the role

  • This role includes the backend and data infrastructure surface area: high-throughput telemetry ingestion, ClickHouse-backed OLAP performance, evaluation pipelines, RabbitMQ/Temporal workflows, multi-tenant scheduling, and product-facing APIs.
  • Some weeks you'll be deep in distributed systems and query performance.
  • Other weeks you'll ship a customer-facing feature end to end across backend, frontend, and the data layer.

Responsibilities

  • Design and build backend systems for trace ingestion, trajectory processing, evaluation orchestration, scoring, labeling, rubric generation, and customer-facing analytics.
  • Own the API surface used by the Judgment platform UI, SDKs, JudgmentHub libraries, MCP server, Slack agent, and customer integrations.
  • Build and operate the RabbitMQ / Temporal evaluation pipeline, including retry semantics, failure recovery, state reconciliation, and tenant-level scheduling.
  • Optimize the ClickHouse OLAP layer: schema design, partitioning, skip indexes, full-text-search pruning, query rewrites, deduplication, pagination correctness, and storage growth.
  • Roll out safely with feature flags, design docs, code reviews, tests, observability, and production debugging.
  • Strong backend engineering experience building and operating production systems under real load.
  • Excellent fundamentals in distributed systems, API design, data modeling, reliability, and performance.
  • Judgment builds the infrastructure to do that.

Requirements

  • Experience working with high-volume event, trace, log, metric, or telemetry data.
  • Comfort owning systems beyond initial launch: debugging production issues, improving observability, scaling bottlenecks, and cleaning up abstractions as the product evolves.
  • Ability to work across backend, data, product, and infrastructure boundaries rather than treating them as separate silos.

Nice to have

  • Experience with ClickHouse, OLAP systems, distributed query engines, or large-scale analytical databases.
  • Experience with RabbitMQ, Temporal, Kafka, Spark, Flink, Ray, Airflow, Dagster, Prefect, or similar queue/workflow/data systems.
  • Experience with OTEL, observability products, tracing, logging, or monitoring infrastructure.
  • Experience with LLM evaluation, labeling systems, rubric generation, context engineering, RL data pipelines, embedding pipelines, vector search, clustering, or anomaly detection.
  • As agents move from demos to production, the bottleneck is no longer just better prompts.
  • It is turning real production experience into high-quality data for evals, labeling, rubric generation, context engineering, and RL workflows.
  • The technical problems are foundational.
  • Long agent trajectories are messy, high-volume, and hard to reason about.

Benefits

  • Instead of only showing teams what happened, Judgment helps decide what matters, what should be learned from, and how that learning should flow back into the agent.
  • Judgment is building the learning infrastructure for agents.
  • We've raised $30M+ from Lightspeed, SV Angel, Valor Equity Partners, and others.
  • Work directly with customers to understand where their agents fail, what data is useful, and how Judgment should structure that experience for learning.

Company info

  • Turn raw spans, conversations, tool calls, scorer outputs, and agent-judge results into clean data models customers can use for evals, labeling, context engineering, and RL workflows.

This listing is sourced directly from Judgmentlabs's careers page and normalized into a canonical job model.