Zingtree
Senior DevOps / Platform Reliability Engineer
East Coast - United States · Senior
Sponsorship not specifiedDetected 75 days ago
PythonBashGitMySQLRedisAWSKubernetesTerraformCI/CDGitHub ActionsJenkinsLinuxPrometheusGrafanaDevOpsSite Reliability EngineeringPlatform EngineeringKafkaMachine LearningLLMsAgentic AILangGraphNetwork SecurityCompliance
About the role
- If you want to operate a production AI platform and use AI to help operate it, this role is for you.
Responsibilities
- Own and evolve CI/CD pipelines using GitHub Actions and OIDC-based authentication for microservices and agentic workloads, with safe, fast, and reversible deployments.
- Manage the edge and network perimeter, including Cloudflare (CDN, WAF, Bot Management, DDoS protection, Zero Trust / Access), CloudFront, API Gateway, ALB/NLB, Route 53, and network security controls.
- Build and maintain Lambda workloads where event-driven or serverless architectures are the right fit.
- Build observability as a product using Prometheus, Grafana, and OpenTelemetry, including telemetry for LLM and agentic systems such as token cost, tool-call latency, evaluation signals, and prompt/version tracking.
- Drive FinOps initiatives, including tagging standards, Savings Plans and Reserved Instances, per-tenant and per-workload cost attribution, and LLM cost controls.
- Build and evolve our AI-native DevOps capabilities (see section below).
- Partner with engineering teams to define platform standards, service templates, deployment best practices, and operational SLOs.
- Collaborate with software engineering teams to support continuous integration and continuous delivery best practices.
- Document infrastructure, deployment processes, and operational standards to support knowledge sharing across the team.
- Lead with Action
Requirements
- Strong experience with CI/CD pipelines and tools such as GitHub Actions, GitLab CI, Jenkins, or CircleCI.
- Hands-on experience operating production EKS environments, including autoscaling, ingress, secrets management, and cluster upgrades.
- Deep experience with Terraform and GitHub Actions, ideally using OIDC-based cloud authentication.
- Experience with Aurora/RDS MySQL, Redis (ElastiCache), and S3, including backups, PITR, migrations, and lifecycle management.
- Experience operating Argo CD at scale.
- Experience with Infrastructure as Code tools such as Terraform, CloudFormation, or Ansible.
- Experience managing Cloudflare services including WAF, Bot Management, Rate Limiting, and Zero Trust / Access, along with CloudFront.
- Experience operating Kafka/MSK at scale, including topics, consumer groups, and schema registries.
- Experience with Lambda and event-driven architectures.
- Strong understanding of security best practices across IAM, KMS, secrets management, networking, and software supply chain security.
Nice to have
- Experience operating LLM or ML workloads in production, including LiteLLM, Bedrock, pgvector, prompt caching, or evaluation systems.
- We bias toward automation over toil.
- If you do it twice, script it.
- If it pages twice, fix it.
- We're a small team with high ownership.
- You'll help define standards, not just follow them.
- Humans stay in the loop for anything risky.
- AI accelerates decision-making but does not replace judgment.
Benefits
- We move quickly with purpose, take smart risks, learn fast, and focus on outcomes that benefit our customers and the business.
Company info
- Operate and scale our Kubernetes platform (EKS + Argo CD), including autoscaling, ingress, external-dns, cert-manager, External Secrets Operator, backups, runtime guardrails, and multi-tenant isolation for enterprise customers.
- We care deeply about our customers and employees, helping each other achieve professional growth and meaningful impact.
- We are learners.
Apply directly at Zingtree →Create a free account for alerts like thisView Zingtree immigration profile
This listing is sourced directly from Zingtree's careers page and normalized into a canonical job model.