Basis AI

Basis AI

Member of Technical Staff - SRE

New York Office · Staff+

Sponsorship not specifiedDetected 9 days ago
Cloud PlatformsTerraformCI/CDPrometheusGrafanaSite Reliability EngineeringMachine LearningCybersecurityIncident ResponseAccountingValuationLeadershipMentoring

About the role

  • As a Site Reliability Engineer at Basis, you will ensure the reliability, scalability, and performance of our AI-powered accounting platform.
  • You'll join a high-leverage infrastructure team that sits at the intersection of product and platform, owning the systems that keep Basis fast, secure, and always on as we scale.

Responsibilities

  • Architect, build, and operate reliable, scalable, and secure infrastructure for our production systems.
  • Own cloud infrastructure across compute, storage, and networking, optimizing for availability, performance, and cost efficiency.
  • Design and maintain CI/CD pipelines, infrastructure-as-code, and automation to improve developer velocity and system reliability.
  • Lead incident response efforts, including on-call rotations, incident coordination, postmortems, and root cause analyses.
  • Partner closely with product and engineering teams to balance reliability, performance, cost, and speed of development.
  • What remains is deciding what to build, breaking down problems from first principles, and designing systems that compound.
  • We want people who can decompose a novel problem, reason from fundamentals rather than pattern match, and design things that compound.

Requirements

  • Coding agents provide the ability to move across domains in ways that weren't possible before
  • You can offload more decision making to the agents inside your product meaning the lines between ML work and eng work start to disappear.
  • 5+ years of experience building and operating production infrastructure at scale.
  • Strong software engineering fundamentals and proficiency in at least one programming language.
  • Experience with CI/CD systems, containerization, and modern infrastructure automation.
  • Experience with similar stack (or ability to learn unfamiliar technologies):

Skills

  • Infrastructure-as-Code tools such as Terraform or CloudFormation.
  • Observability and incident management tools (e.g., OpenTelemetry, Prometheus/Grafana, BetterStack, PagerDuty, SLOs/error budgets).
  • Data and systems powering analytics and customer-facing workloads, especially in serverless environments. (e.g., Neon, Modal)
  • While your primary focus will be infrastructure and reliability, strong software engineering skills are essential.

Benefits

  • Pre-tax commuter benefits and 401(k) retirement plan
  • Parental Leave
  • Premium Medical, Dental, and Vision coverage; Life Insurance; and 6 coaching & 6 therapy sessions through Spring Health.
  • Unlimited PTO + 12 paid company holidays.

Company info

  • If you're excited about building resilient systems from the ground up and shaping reliability practices at a fast-growing company, we'd love to work with you.

This listing is sourced directly from Basis AI's careers page and normalized into a canonical job model.