Sambanovasystems
Cloud Site Reliability Engineer
San Jose, California, United States · Full-time
Sponsorship not specifiedDetected 14 days ago
PythonGoFull-Stack DevelopmentGitSQLNoSQLRedisAWSGCPAzureCloud PlatformsDockerKubernetesTerraformAnsibleCI/CDGitHub ActionsJenkinsLinuxPrometheusGrafanaDatadogDevOpsSite Reliability Engineering
About the role
- As a Cloud Site Reliability Engineer (SRE) specializing in our AI Inferencing Service, you will be the guardian of its reliability, performance, and scalability.
- You will bridge the gap between software development and operations, applying an engineering mindset to solve operational challenges.
- Your primary focus will be ensuring our inference endpoints have exceptional uptime, low-latency response times, and efficient resource utilization, directly impacting the experience of our customers and the success of our AI products.
Responsibilities
- This includes implementing and supporting AI infrastructure in new regions, such as Asia, Europe, and Latin America, to support the growth of our business.
- Participate in a balanced on-call rotation to provide 24/7 support for the service.
- Focus on Prevention: We invest heavily in automation, robust testing, and system design to prevent pages before they happen.
- The goal of on-call is not to heroically fight fires, but to manage rare, complex failures and use those learnings to make the system more resilient.
- Incident Management: Lead the response to incidents affecting the inferencing service, driving blameless post-mortems and implementing corrective actions to prevent recurrence.
- Design and implement auto-scaling policies to handle variable inference loads cost-effectively.
- Use insights from on-call incidents to drive improvements that enhance system stability and scalability.
- Infrastructure as Code (IaC): Manage and evolve our cloud infrastructure (on AWS, GCP, and/or Azure along with on-prem) using tools like Terraform and Ansible, ensuring it is secure, repeatable, and scalable.
- CI/CD & Automation: Champion automation by building and improving CI/CD pipelines for the seamless and safe deployment of new model versions and service updates.
- Capacity Planning: Forecast infrastructure needs based on product roadmaps and usage trends. Work with finance and engineering teams to manage cloud costs and optimize spending.
Requirements
- Alerts must be actionable and require immediate human intervention.
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 3-5+ years of experience in a Site Reliability Engineer, DevOps, or related role supporting a large-scale, customer-facing service in a public cloud environment (AWS, GCP, Azure).
- Proven experience with containerization and orchestration technologies (Docker, Kubernetes).
- Solid experience with Infrastructure as Code (e.g., Terraform, CloudFormation).
- Familiarity with CI/CD principles and tools (e.g., Jenkins, GitHub Actions, ArgoCD).
- Direct experience supporting ML/AI inferencing services in production.
- Familiarity with GPU-accelerated computing and optimizing workloads for NVIDIA GPUs for purposes of mapping to RDUs.
- Knowledge of model serving frameworks like vLLM, SGLang or Ray.
- Experience with managing and tuning databases (SQL or NoSQL) and caching systems (Redis, Memcached).
Skills
- The era of pervasive AI has arrived.
- SambaNova Suite™ is the first full-stack, generative AI platform, from chip to model, optimized for enterprise and government organizations.
- About SambaNova Systems
- Our DataScale systems and SambaFlow software are pushing the boundaries of what's possible with generative AI and large language models.
Compensation
- SambaNova offers a competitive total rewards package, including the base salary, plus equity and benefits.
Benefits
- SambaNova offers a competitive total rewards package, including the base salary, plus equity and benefits.
- We cover 95% premium coverage for employee medical insurance, and 77% premium coverage for dependents and offer a Health Savings Account (HSA) with employer contribution.
- We also offer Dental, Vision, Short/Long term Disability, Basic Life, Voluntary Life, and AD&D insurance plans in addition to Flexible Spending Account (FSA) options like Health Care, Limited Purpose, and Dependent Care.
- We believe a sustainable on-call schedule is critical for long-term success and team health.
Company info
- What We're Looking For (Must-Haves)
Equal opportunity
- SambaNova Systems is an Equal Opportunity/Affirmative Action Employer.
Apply directly at Sambanovasystems →Create a free account for alerts like thisView Sambanovasystems immigration profile
This listing is sourced directly from Sambanovasystems's careers page and normalized into a canonical job model.