Percepta
AI Infrastructure Engineer
New York City
Sponsorship not specifiedDetected 113 days ago
PythonBashAlgorithmsGitAWSGCPAzureDockerKubernetesTerraformCI/CDGitHub ActionsPrometheusGrafanaDevOpsSite Reliability EngineeringRESTMachine LearningAgentic AIMLOpsIncident ResponseComplianceHIPAAResearch
About the role
- And part of it is genuinely new territory, figuring out what SRE means when the systems you're operating make autonomous decisions.
- The infrastructure patterns for the agentic systems of the future don't exist yet.
- The infrastructure contract changes when your workloads have agency.
Responsibilities
- Own and evolve our IaC stack: Terraform and Kubernetes across AWS, GCP, and Azure
- Design and maintain CI/CD pipelines that give teams fast, trustworthy feedback from commit to production
- Build operational foundations: monitoring, alerting, incident response, and the new patterns that emerge when AI systems are participants in that response
- 5+ years building and operating production infrastructure in DevOps or SRE roles
Requirements
- Deep experience with at least 1 major cloud provider (AWS, GCP, or Azure): networking, IAM, cost management, the operational realities of production workloads
- Solid Docker and Kubernetes experience in production.
- Experience designing and maintaining CI/CD pipelines (GitHub Actions, GitLab CI, or similar)
- Scripting proficiency in Python, Bash, or similar
- MLOps experience isn't required, but you'll be closer to that boundary than most infra roles.
Nice to have
- Multi-region and multi-cloud experience across 2+ providers
- Experience with single-tenant or on-prem deployments alongside multi-tenant SaaS
- Familiarity with GitOps patterns and progressive delivery
- Familiarity with the Grafana stack (Prometheus, Grafana, Loki) or equivalent
- Experience with compliance frameworks (HIPAA, SOC 2) and how they shape infrastructure decisions in regulated environments
- Background supporting ML or research workflows moving to production: model deployment, pipeline orchestration, or similar
- You've thought about what observability means for non-deterministic systems and have opinions about it
- The infrastructure patterns for autonomous AI systems are still being written.
Benefits
- Build observability primitives for agentic workflows, tracing agent decisions and execution paths, not just service latency and pod health
- Work across engineering teams to meet the reliability and compliance requirements of the institutions we serve (SOC 2, HIPAA, regulated environments in healthcare and energy)
Company info
- Percepta's mission is to transform critical institutions with applied AI. We care that industries that power the world (healthcare, manufacturing, energy) benefit from frontier technology.
- We collaborate with industry-leading customers to drive AI transformation. We bring:
- Forward-deployed expertise in engineering, product, and research
- Mosaic, our in-house toolkit for rapidly deploying agentic architectures
- Percepta's mission is to transform critical institutions with applied AI.
- We care that industries that power the world (healthcare, manufacturing, energy) benefit from frontier technology.
- We collaborate with industry-leading customers to drive AI transformation.
- Strategic partnerships with Anthropic, McKinsey, AWS, and the General Catalyst portfolio
- Our team is a fast-growing group of Applied AI Engineers, Embedded Product Managers, and Researchers motivated by getting frontier AI into the places that actually run the world.
- Percepta is a direct partnership with General Catalyst.
- What we're doing matters and we have to give a shit.
- Internally, that means fixing badness when you find it.
- Externally, it means honoring the trust our customers place in us with their most important problems.
Apply directly at Percepta →Create a free account for alerts like thisView Percepta immigration profile
This listing is sourced directly from Percepta's careers page and normalized into a canonical job model.