Basis AI
Member of Technical Staff - SRE
New York Office · Staff+
Sponsorship not specifiedDetected 9 days ago
Cloud PlatformsTerraformCI/CDPrometheusGrafanaSite Reliability EngineeringMachine LearningCybersecurityIncident ResponseAccountingValuationLeadershipMentoring
About the role
- As a Site Reliability Engineer at Basis, you will ensure the reliability, scalability, and performance of our AI-powered accounting platform.
- You'll join a high-leverage infrastructure team that sits at the intersection of product and platform, owning the systems that keep Basis fast, secure, and always on as we scale.
Responsibilities
- Architect, build, and operate reliable, scalable, and secure infrastructure for our production systems.
- Own cloud infrastructure across compute, storage, and networking, optimizing for availability, performance, and cost efficiency.
- Design and maintain CI/CD pipelines, infrastructure-as-code, and automation to improve developer velocity and system reliability.
- Lead incident response efforts, including on-call rotations, incident coordination, postmortems, and root cause analyses.
- Partner closely with product and engineering teams to balance reliability, performance, cost, and speed of development.
- What remains is deciding what to build, breaking down problems from first principles, and designing systems that compound.
- We want people who can decompose a novel problem, reason from fundamentals rather than pattern match, and design things that compound.
Requirements
- Coding agents provide the ability to move across domains in ways that weren't possible before
- You can offload more decision making to the agents inside your product meaning the lines between ML work and eng work start to disappear.
- 5+ years of experience building and operating production infrastructure at scale.
- Strong software engineering fundamentals and proficiency in at least one programming language.
- Experience with CI/CD systems, containerization, and modern infrastructure automation.
- Experience with similar stack (or ability to learn unfamiliar technologies):
Skills
- Infrastructure-as-Code tools such as Terraform or CloudFormation.
- Observability and incident management tools (e.g., OpenTelemetry, Prometheus/Grafana, BetterStack, PagerDuty, SLOs/error budgets).
- Data and systems powering analytics and customer-facing workloads, especially in serverless environments. (e.g., Neon, Modal)
- While your primary focus will be infrastructure and reliability, strong software engineering skills are essential.
Benefits
- Pre-tax commuter benefits and 401(k) retirement plan
- Parental Leave
- Premium Medical, Dental, and Vision coverage; Life Insurance; and 6 coaching & 6 therapy sessions through Spring Health.
- Unlimited PTO + 12 paid company holidays.
Company info
- If you're excited about building resilient systems from the ground up and shaping reliability practices at a fast-growing company, we'd love to work with you.
Apply directly at Basis AI →Create a free account for alerts like thisView Basis AI immigration profile
This listing is sourced directly from Basis AI's careers page and normalized into a canonical job model.