Blaxel

Blaxel

Site Reliability Engineer

San Francisco

Sponsorship not specifiedDetected 140 days ago
PythonGoRustDistributed SystemsGitAWSGCPCloud PlatformsKubernetesTerraformCI/CDGitHub ActionsJenkinsLinuxPrometheusGrafanaDatadogDevOpsSite Reliability EngineeringLLMsAgentic AICybersecurityIncident ResponseLoad Testing

About the role

  • We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure platform.
  • Your mission: keep our ultra-low-latency, stateful, serverless compute engine rock-solid as we serve billions of agent requests for the most sophisticated AI teams in the world.
  • This role is highly technical and execution-heavy.

Responsibilities

  • Collaborating closely with the founders, the infra team, and the dev team-and leveraging AI wherever it creates leverage-you will architect and operate the systems that keep Blaxel fast, resilient, and secure.
  • Build and evolve our observability stack (metrics, traces, logs), ensuring we detect issues before users do.
  • Define, monitor, and drive SLOs/SLIs across key system surfaces to maintain world-class reliability.
  • Lead incident response with rigor: root cause analysis, post-mortems, and driving systemic fixes.
  • Design and implement self-healing, automated operational systems to eliminate toil and scale ops.
  • Build automation and tooling-often with AI agents-to streamline operations, debugging, capacity planning, and failure prediction.
  • Own security best practices at the infrastructure layer, from sandboxed compute to network isolation.
  • Partner with platform engineers to ensure reliability is designed into new features from day one.
  • You want to invent new reliability systems-not just maintain existing ones.
  • We want you to design new reliability systems, push the boundaries of automation, and continuously evolve the platform to meet the demands of next-generation AI workloads.

Requirements

  • High-velocity execution: You have a strong bias for action and a track record of shipping reliable systems quickly with excellent judgment.
  • Experience with bare-metal servers and datacenter operations (PXE/iPXE provisioning, IPMI/BMC, RAID/NVMe, SR-IOV, high-throughput networking)
  • Experience with infrastructure-as-code tools such as Terraform or Pulumi
  • Knowledge of service mesh or API gateway technologies
  • Prior experience in high-growth or high-availability environments

Nice to have

  • Serverless compute systems
  • Sandboxed execution environments
  • Ultra-low-latency runtime engineering
  • Distributed key-value stores and databases
  • Chaos engineering
  • Rust, Go, or systems-level programming
  • Deep generative AI infrastructure
  • Experience with any of the following is a plus (not required):

Skills

  • 3+ years in SRE, DevOps, or infrastructure engineering roles
  • Strong proficiency in at least one programming language such as Go, Rust, or Python
  • Hands-on experience with a major cloud provider (AWS, GCP)
  • Solid knowledge of Linux systems, networking fundamentals, and distributed systems
  • Experience with Kubernetes or similar orchestrators
  • Familiarity with observability stacks (Prometheus, Grafana, ELK, Datadog)
  • Strong debugging, problem-solving, and incident-management skills
  • Fluent across systems, cloud, networking, and distributed computing.
  • Blaxel is AWS for AI agents.
  • Founders choose us when they hit the limits of general-purpose clouds.

Benefits

  • You'll own our reliability posture end-to-end-observability, performance tuning, incident ops, infrastructure health, and the automation systems that keep everything running smoothly.

This listing is sourced directly from Blaxel's careers page and normalized into a canonical job model.