Lambda
Senior Site Reliability Engineer - SDN
San Francisco Office (Fremont St) · Senior
Sponsorship not specifiedDetected 15 days ago
PythonGoDistributed SystemsAWSCloud PlatformsKubernetesTerraformAnsibleHelmCI/CDLinuxSite Reliability EngineeringMachine LearningIncident ResponseNetwork MonitoringResearch
About the role
- However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
Responsibilities
- Develop tooling and automation to reduce operational toil and improve reliability
- Collaborate with software, platform, and networking teams to improve service reliability and deployment workflows
- Deploy and maintain network monitoring, observability, and management tools
- Drive operational excellence through observability, incident management, capacity planning, postmortems, and participation in the on-call rotation
- Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
Requirements
- Have 5+ years of experience in Site Reliability Engineering, Production Engineering, or a similar role
- Have experience with Kubernetes application lifecycle management, upgrades, troubleshooting, and production operations
- Have experience with observability platforms, monitoring, alerting, and metrics
Nice to have
- Experience operating production-scale SDNs in a cloud environment (e.g., infrastructure powering AWS VPC-like networking services)
- Software development experience in Go and/or Python (C is a plus)
- Experience automating infrastructure and network configuration using Kubernetes, Helm, Terraform, and Ansible
- Deep understanding of the Linux networking stack and its interaction with network virtualization technologies, SR-IOV, and DPDK
- Understanding of the SDN ecosystem and modern cloud networking architectures
- Experience diagnosing complex production issues across infrastructure, networking, and application layers
Skills
- Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence.
- One person, one GPU.
- Our scope includes the Lambda website, cloud APIs and systems as well as internal tooling for system deployment, management and maintenance.
Compensation
- The annual salary range for this position has been set based on market data and other factors.
- However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
- We offer generous cash & equity compensation
Benefits
- We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- Flexible paid time off plan that we all actually use
Company info
- We offer generous cash & equity compensation
- 401k Plan with 2% company match (USA employees)
Equal opportunity
- Lambda is an Equal Opportunity employer.
- Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
Visa & Work Authorization
- ational origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law
This listing is sourced directly from Lambda's careers page and normalized into a canonical job model.