Alembic
Senior Network & Site Reliability Engineer
San Francisco HQ · Senior
Sponsorship not specifiedDetected 39 days ago
PythonBashKubernetesTerraformAnsibleLinuxPrometheusGrafanaDatadogSite Reliability EngineeringKafkaMachine LearningSparkAirflowData ScienceCybersecurityNetwork SecurityIncident ResponseComplianceTCP/IPFirewallVPNBGP/OSPF
About the role
- This isn't a traditional "keep the lights on" role.
- You'll work across networking, systems, automation, observability, and reliability engineering to scale a platform where performance genuinely matters, with real influence over architecture decisions.
- Technical autonomy: Ownership over architecture decisions and the freedom to solve hard infrastructure problems your way.
Responsibilities
- Own network device configuration management end to end, ensuring consistency and reliability across the fleet.
- Implement and manage complex network protocols and connectivity, including BGP, VPNs, and WAN circuits and external peering.
- Build and maintain comprehensive monitoring, alerting, and incident response - SLOs, runbooks, and on-call rotations - and drive post-incident analysis and continuous improvement.
- Partner across engineering and data science to drive a culture of performance and reliability.
- A strong background in network security, architecture, design, and operations.
- If you only want to tell people what to build instead of building and automating alongside them, this isn't the environment for you.
Requirements
- Extensive hands-on experience with network devices (firewalls, switches, load balancers) and large-scale architectures and protocols - BGP, QoS, MPLS, and IPsec VPNs.
- Familiarity with Kubernetes networking (CNI plugins, ingress, service networking, network policy) and strong operational experience with Linux-based production infrastructure.
- Experience with monitoring and observability stacks (Prometheus, Grafana, Datadog, ELK, OpenTelemetry).
Nice to have
- Network automation and IaC tooling (Ansible, Terraform, Nornir, or similar), plus IPAM/DCIM platforms (NetBox, Infoblox, or similar).
Benefits
- Series B momentum, real ownership: Meaningful equity at a Series B company that's raised $145M, with proven product-market fit and Fortune 100 traction.
- Architect and operate scalable, secure network architecture for high-security requirements and large-scale machine learning workloads.
Company info
- Alembic is the pioneering Causal AI platform. We help the world's largest enterprises move past correlation to prove what actually drives business outcomes - the question marketing and growth teams have never been able to answer with confidence. Fortune 100 companies including Nvidia, Delta Air Lines, and Mars use Alembic to make multimillion-dollar decisions on trusted, causal evidence.
- We're backed by a $145M Series B from WndrCo (founded by Jeffrey Katzenberg), Jensen Huang, Joe Montana, Prysm Capital, and Accenture. Our models run on our own NVIDIA DGX SuperPOD built on Grace Blackwell infrastructure - one of the fastest private supercomputers in the world. (We've melted GPUs getting here.)
- We're building infrastructure that has to perform under real-world scale, reliability, and security demands - and we're looking for an engineer who wants to own the foundation it runs on. This isn't a traditional "keep the lights on" role.
- Alembic is the pioneering Causal AI platform.
- We help the world's largest enterprises move past correlation to prove what actually drives business outcomes - the question marketing and growth teams have never been able to answer with confidence.
- Fortune 100 companies including Nvidia, Delta Air Lines, and Mars use Alembic to make multimillion-dollar decisions on trusted, causal evidence.
- We're backed by a $145M Series B from WndrCo (founded by Jeffrey Katzenberg), Jensen Huang, Joe Montana, Prysm Capital, and Accenture.
- Our models run on our own NVIDIA DGX SuperPOD built on Grace Blackwell infrastructure - one of the fastest private supercomputers in the world. (We've melted GPUs getting here.)
- You prefer static over dynamic - projects and priorities adapt as we grow. We have real paying customers and a playbook, and we still move at startup speed at Series B scale.
- We have real paying customers and a playbook, and we still move at startup speed at Series B scale.
Apply directly at Alembic →Create a free account for alerts like thisView Alembic immigration profile
This listing is sourced directly from Alembic's careers page and normalized into a canonical job model.