Judgmentlabs
Senior Backend Engineer
San Francisco · Senior · Full-time
Sponsorship not specifiedDetected 36 days ago
Next.jsDistributed SystemsCode ReviewDatabricksTemporalAPI DevelopmentKafkaRabbitMQSparkAirflowData EngineeringLLMs
About the role
- This role includes the backend and data infrastructure surface area: high-throughput telemetry ingestion, ClickHouse-backed OLAP performance, evaluation pipelines, RabbitMQ/Temporal workflows, multi-tenant scheduling, and product-facing APIs.
- Some weeks you'll be deep in distributed systems and query performance.
- Other weeks you'll ship a customer-facing feature end to end across backend, frontend, and the data layer.
Responsibilities
- Design and build backend systems for trace ingestion, trajectory processing, evaluation orchestration, scoring, labeling, rubric generation, and customer-facing analytics.
- Own the API surface used by the Judgment platform UI, SDKs, JudgmentHub libraries, MCP server, Slack agent, and customer integrations.
- Build and operate the RabbitMQ / Temporal evaluation pipeline, including retry semantics, failure recovery, state reconciliation, and tenant-level scheduling.
- Optimize the ClickHouse OLAP layer: schema design, partitioning, skip indexes, full-text-search pruning, query rewrites, deduplication, pagination correctness, and storage growth.
- Roll out safely with feature flags, design docs, code reviews, tests, observability, and production debugging.
- Strong backend engineering experience building and operating production systems under real load.
- Excellent fundamentals in distributed systems, API design, data modeling, reliability, and performance.
- Judgment builds the infrastructure to do that.
Requirements
- Experience working with high-volume event, trace, log, metric, or telemetry data.
- Comfort owning systems beyond initial launch: debugging production issues, improving observability, scaling bottlenecks, and cleaning up abstractions as the product evolves.
- Ability to work across backend, data, product, and infrastructure boundaries rather than treating them as separate silos.
Nice to have
- Experience with ClickHouse, OLAP systems, distributed query engines, or large-scale analytical databases.
- Experience with RabbitMQ, Temporal, Kafka, Spark, Flink, Ray, Airflow, Dagster, Prefect, or similar queue/workflow/data systems.
- Experience with OTEL, observability products, tracing, logging, or monitoring infrastructure.
- Experience with LLM evaluation, labeling systems, rubric generation, context engineering, RL data pipelines, embedding pipelines, vector search, clustering, or anomaly detection.
- As agents move from demos to production, the bottleneck is no longer just better prompts.
- It is turning real production experience into high-quality data for evals, labeling, rubric generation, context engineering, and RL workflows.
- The technical problems are foundational.
- Long agent trajectories are messy, high-volume, and hard to reason about.
Benefits
- Instead of only showing teams what happened, Judgment helps decide what matters, what should be learned from, and how that learning should flow back into the agent.
- Judgment is building the learning infrastructure for agents.
- We've raised $30M+ from Lightspeed, SV Angel, Valor Equity Partners, and others.
- Work directly with customers to understand where their agents fail, what data is useful, and how Judgment should structure that experience for learning.
Company info
- Turn raw spans, conversations, tool calls, scorer outputs, and agent-judge results into clean data models customers can use for evals, labeling, context engineering, and RL workflows.
Apply directly at Judgmentlabs →Create a free account for alerts like thisView Judgmentlabs immigration profile
This listing is sourced directly from Judgmentlabs's careers page and normalized into a canonical job model.