Zingtree

Zingtree

Senior DevOps / Platform Reliability Engineer

East Coast - United States · Senior

Sponsorship not specifiedDetected 75 days ago
PythonBashGitMySQLRedisAWSKubernetesTerraformCI/CDGitHub ActionsJenkinsLinuxPrometheusGrafanaDevOpsSite Reliability EngineeringPlatform EngineeringKafkaMachine LearningLLMsAgentic AILangGraphNetwork SecurityCompliance

About the role

  • If you want to operate a production AI platform and use AI to help operate it, this role is for you.

Responsibilities

  • Own and evolve CI/CD pipelines using GitHub Actions and OIDC-based authentication for microservices and agentic workloads, with safe, fast, and reversible deployments.
  • Manage the edge and network perimeter, including Cloudflare (CDN, WAF, Bot Management, DDoS protection, Zero Trust / Access), CloudFront, API Gateway, ALB/NLB, Route 53, and network security controls.
  • Build and maintain Lambda workloads where event-driven or serverless architectures are the right fit.
  • Build observability as a product using Prometheus, Grafana, and OpenTelemetry, including telemetry for LLM and agentic systems such as token cost, tool-call latency, evaluation signals, and prompt/version tracking.
  • Drive FinOps initiatives, including tagging standards, Savings Plans and Reserved Instances, per-tenant and per-workload cost attribution, and LLM cost controls.
  • Build and evolve our AI-native DevOps capabilities (see section below).
  • Partner with engineering teams to define platform standards, service templates, deployment best practices, and operational SLOs.
  • Collaborate with software engineering teams to support continuous integration and continuous delivery best practices.
  • Document infrastructure, deployment processes, and operational standards to support knowledge sharing across the team.
  • Lead with Action

Requirements

  • Strong experience with CI/CD pipelines and tools such as GitHub Actions, GitLab CI, Jenkins, or CircleCI.
  • Hands-on experience operating production EKS environments, including autoscaling, ingress, secrets management, and cluster upgrades.
  • Deep experience with Terraform and GitHub Actions, ideally using OIDC-based cloud authentication.
  • Experience with Aurora/RDS MySQL, Redis (ElastiCache), and S3, including backups, PITR, migrations, and lifecycle management.
  • Experience operating Argo CD at scale.
  • Experience with Infrastructure as Code tools such as Terraform, CloudFormation, or Ansible.
  • Experience managing Cloudflare services including WAF, Bot Management, Rate Limiting, and Zero Trust / Access, along with CloudFront.
  • Experience operating Kafka/MSK at scale, including topics, consumer groups, and schema registries.
  • Experience with Lambda and event-driven architectures.
  • Strong understanding of security best practices across IAM, KMS, secrets management, networking, and software supply chain security.

Nice to have

  • Experience operating LLM or ML workloads in production, including LiteLLM, Bedrock, pgvector, prompt caching, or evaluation systems.
  • We bias toward automation over toil.
  • If you do it twice, script it.
  • If it pages twice, fix it.
  • We're a small team with high ownership.
  • You'll help define standards, not just follow them.
  • Humans stay in the loop for anything risky.
  • AI accelerates decision-making but does not replace judgment.

Benefits

  • We move quickly with purpose, take smart risks, learn fast, and focus on outcomes that benefit our customers and the business.

Company info

  • Operate and scale our Kubernetes platform (EKS + Argo CD), including autoscaling, ingress, external-dns, cert-manager, External Secrets Operator, backups, runtime guardrails, and multi-tenant isolation for enterprise customers.
  • We care deeply about our customers and employees, helping each other achieve professional growth and meaningful impact.
  • We are learners.

This listing is sourced directly from Zingtree's careers page and normalized into a canonical job model.