Decagon

Decagon

Senior Software Engineer, Cloud Infrastructure

New York City · Senior · Full-time

Sponsorship not specified$200k-$400kDetected 11 days ago
Distributed SystemsAWSGCPAzureCloud PlatformsKubernetesTerraformCI/CDPrometheusGrafanaDatadogDevOpsPlatform EngineeringMachine LearningAgentic AIIncident ResponseSupply ChainCustomer SuccessCustomer SupportDNSLoad BalancingWriting

About the role

  • Read more about the infra team's work here: https://decagon.ai/blog/what-an-air-gapped-ai-deployment-actually-requires
  • The work is core infrastructure at heart: reliability, CI/CD, deployment automation, on-call.

Responsibilities

  • Build the platform
  • Design the development and production platforms that power our products, and the abstractions over cloud infrastructure, Kubernetes, and networking that let engineers ship without becoming infrastructure experts.
  • Own enterprise deployments
  • Take end-to-end ownership of deployment architecture in customer-owned cloud environments (VPC configuration, permissioning, networking, provisioning) and the full lifecycle that follows: setup, upgrades, scaling, and incident support.
  • Build the runbooks and automation that make it repeatable.
  • Own the reliability of the systems our AI agents depend on in production, where latency, availability, and graceful degradation directly shape the customer experience.
  • Partner across boundaries
  • 4+ years building and operating core infrastructure, platform engineering, or infrastructure/DevOps, ideally with some customer-facing deployment experience.

Requirements

  • Comfort navigating ambiguity across a range of stakeholders, from engineers to security and compliance teams, and turning those conversations into actionable plans.
  • Clear technical writing and a track record of driving adoption across teams.
  • Experience managing deployments in customer-owned cloud environments, including security reviews, compliance requirements, and change management.
  • Familiarity with observability and incident management in distributed systems (Prometheus, Grafana, Datadog, or similar).
  • Experience operating latency-sensitive or ML/AI-serving workloads in production.
  • Experience using AI-assisted tooling to make yourself and your team dramatically more effective.
  • Deep experience with a major cloud provider (GCP, AWS, or Azure), along with Terraform and Kubernetes at scale.

Compensation

  • This range reflects the expected compensation for this role. Compensation within the range is determined based on experience, skills, and the scope of responsibilities, with flexibility for candidates who demonstrate exceptional impact.
  • In addition to base salary, we offer competitive equity. Final compensation may vary based on location within the United States.

Benefits

  • We proudly offer the following benefits for our full-time employees:
  • Take what you need vacation policy (subject to local requirements
  • UK employees receive 25 days of statutory leave)
  • Medical, Dental, and Vision benefits for you and your family
  • Life Insurance and Disability Benefits
  • Retirement Plan (e.g., 401K, pension)
  • Parental Leave
  • Fertility and family building benefits through Carrot
  • These benefits are described in more detail in Decagon's policies, may vary by location, and can change at any time according to applicable compensation and benefits plans.

Company info

  • Work directly with customers' platform, security, and DevOps teams to navigate their infrastructure and compliance constraints, and with our Product, Security, Sales, and Customer Success teams to turn customer requirements into concrete deployment plans.

This listing is sourced directly from Decagon's careers page and normalized into a canonical job model.