Alembic

Alembic

Senior Network & Site Reliability Engineer

San Francisco HQ · Senior

Sponsorship not specifiedDetected 39 days ago
PythonBashKubernetesTerraformAnsibleLinuxPrometheusGrafanaDatadogSite Reliability EngineeringKafkaMachine LearningSparkAirflowData ScienceCybersecurityNetwork SecurityIncident ResponseComplianceTCP/IPFirewallVPNBGP/OSPF

About the role

  • This isn't a traditional "keep the lights on" role.
  • You'll work across networking, systems, automation, observability, and reliability engineering to scale a platform where performance genuinely matters, with real influence over architecture decisions.
  • Technical autonomy: Ownership over architecture decisions and the freedom to solve hard infrastructure problems your way.

Responsibilities

  • Own network device configuration management end to end, ensuring consistency and reliability across the fleet.
  • Implement and manage complex network protocols and connectivity, including BGP, VPNs, and WAN circuits and external peering.
  • Build and maintain comprehensive monitoring, alerting, and incident response - SLOs, runbooks, and on-call rotations - and drive post-incident analysis and continuous improvement.
  • Partner across engineering and data science to drive a culture of performance and reliability.
  • A strong background in network security, architecture, design, and operations.
  • If you only want to tell people what to build instead of building and automating alongside them, this isn't the environment for you.

Requirements

  • Extensive hands-on experience with network devices (firewalls, switches, load balancers) and large-scale architectures and protocols - BGP, QoS, MPLS, and IPsec VPNs.
  • Familiarity with Kubernetes networking (CNI plugins, ingress, service networking, network policy) and strong operational experience with Linux-based production infrastructure.
  • Experience with monitoring and observability stacks (Prometheus, Grafana, Datadog, ELK, OpenTelemetry).

Nice to have

  • Network automation and IaC tooling (Ansible, Terraform, Nornir, or similar), plus IPAM/DCIM platforms (NetBox, Infoblox, or similar).

Benefits

  • Series B momentum, real ownership: Meaningful equity at a Series B company that's raised $145M, with proven product-market fit and Fortune 100 traction.
  • Architect and operate scalable, secure network architecture for high-security requirements and large-scale machine learning workloads.

Company info

  • Alembic is the pioneering Causal AI platform. We help the world's largest enterprises move past correlation to prove what actually drives business outcomes - the question marketing and growth teams have never been able to answer with confidence. Fortune 100 companies including Nvidia, Delta Air Lines, and Mars use Alembic to make multimillion-dollar decisions on trusted, causal evidence.
  • We're backed by a $145M Series B from WndrCo (founded by Jeffrey Katzenberg), Jensen Huang, Joe Montana, Prysm Capital, and Accenture. Our models run on our own NVIDIA DGX SuperPOD built on Grace Blackwell infrastructure - one of the fastest private supercomputers in the world. (We've melted GPUs getting here.)
  • We're building infrastructure that has to perform under real-world scale, reliability, and security demands - and we're looking for an engineer who wants to own the foundation it runs on. This isn't a traditional "keep the lights on" role.
  • Alembic is the pioneering Causal AI platform.
  • We help the world's largest enterprises move past correlation to prove what actually drives business outcomes - the question marketing and growth teams have never been able to answer with confidence.
  • Fortune 100 companies including Nvidia, Delta Air Lines, and Mars use Alembic to make multimillion-dollar decisions on trusted, causal evidence.
  • We're backed by a $145M Series B from WndrCo (founded by Jeffrey Katzenberg), Jensen Huang, Joe Montana, Prysm Capital, and Accenture.
  • Our models run on our own NVIDIA DGX SuperPOD built on Grace Blackwell infrastructure - one of the fastest private supercomputers in the world. (We've melted GPUs getting here.)
  • You prefer static over dynamic - projects and priorities adapt as we grow. We have real paying customers and a playbook, and we still move at startup speed at Series B scale.
  • We have real paying customers and a playbook, and we still move at startup speed at Series B scale.

This listing is sourced directly from Alembic's careers page and normalized into a canonical job model.