Upstart

Upstart

DevOps Engineer, Cloud Platform

United States | Remote

Sponsorship not specified$142k-$197kDetected 12 days ago
AWSCloud PlatformsKubernetesTerraformDevOpsSite Reliability EngineeringMachine LearningSparkA/B TestingCybersecurityIncident ResponseComplianceCadenceArgo CD

About the role

  • Every day, we bring creativity, experimentation, and advanced AI to reshape access to credit, helping millions move forward financially with clarity and confidence.
  • But the numbers only hint at the impact.
  • Every idea, every voice, and every contribution moves us closer to a world where credit never stands between people and their financial progress.

Responsibilities

  • Design and operate a fleet of Kubernetes (EKS) clusters across production, staging, and ephemeral environments, ensuring reliability and high availability
  • Evolve AWS infrastructure and network architecture (VPCs, subnets, IAM, account structure) to support scalable, multi-team workloads
  • Build and maintain infrastructure-as-code and GitOps workflows using tools such as Terraform, CDK, and ArgoCD
  • Participate in and help improve the on-call rotation, leading incident response and post-incident reviews to drive systemic platform improvements
  • Partner with SRE, Delivery, InfoSec, and product/ML teams to land high-impact infrastructure changes and platform standards
  • Drive improvements in developer experience by simplifying platform usage, reducing toil, and enabling faster product and ML development
  • As the leading AI lending marketplace, we partner with banks and credit unions to expand access to affordable credit through technology that's both radically intelligent and deeply human.
  • And whether you choose to work primarily from home or collaborate in-person from one of our offices in Columbus, Austin, the Bay Area, or New York City (opening Summer 2026), you'll have the support to work in the way that works best for you.

Requirements

  • Bachelor's degree in Computer Science, Engineering, Mathematics, or a related field (or equivalent practical experience) and 3+ years of professional experience
  • 3+ years of experience operating Kubernetes in production environments, including cluster networking, storage, and RBAC
  • Proficiency with AWS infrastructure, including VPC design, networking, and IAM
  • Proven expertise in implementing infrastructure-as-code using tools such as Terraform or AWS CDK
  • Experience implementing GitOps workflows using tools such as ArgoCD or similar

Nice to have

  • Knowledge of service mesh technologies such as Istio or Envoy
  • Experience designing or operating multi-cluster Kubernetes architectures
  • Experience with cloud networking at scale, including ingress/egress or edge platforms (e.g., Cloudflare)
  • Knowledge of cloud security, identity, and compliance frameworks (e.g., IAM, SOC 2, CIS benchmarks)

Skills

  • Individual pay is also determined by job-related skills, experience, and relevant education or training.

Compensation

  • At Upstart, your base pay is one part of your total compensation package.
  • The anticipated base salary for this position is expected to be within the below range.
  • United States | Remote - Anticipated Base Salary Range
  • Canada | Remote - Anticipated Base Salary Range
  • $142,000 - $196,600 USD

Benefits

  • In addition, Upstart provides employees with target bonuses, equity compensation, and generous benefits packages (including medical, dental, vision, and 401k).

Company info

  • Upstart's Cloud Platform team sits within the Reliability organization and is responsible for building and operating the shared cloud infrastructure that powers all product and machine learning workloads.
  • The team owns core platform components across Kubernetes (EKS), AWS infrastructure, service mesh, identity, and developer tooling, enabling reliability, scalability, and security across the business.
  • As a DevOps Engineer (L4) at Upstart, you will help evolve this platform to support increasing scale and complexity.
  • You'll partner closely with SRE, Delivery, InfoSec, and Product/ML teams to improve reliability, developer experience, and cost efficiency across a platform used by nearly every engineering team.
  • Improve platform reliability and performance by defining and driving SLOs, analyzing incidents, and implementing systemic fixes
  • Contribute to cost efficiency initiatives by optimizing resource utilization across Kubernetes and cloud infrastructure
  • As a digital first company, the majority of your work can be accomplished remotely.
  • Our platform runs over one million predictions per borrower using more than 1,800 signals, powering smarter, fairer decisions for millions of customers.

This listing is sourced directly from Upstart's careers page and normalized into a canonical job model.