Percepta

Percepta

AI Infrastructure Engineer

New York City

Sponsorship not specifiedDetected 113 days ago
PythonBashAlgorithmsGitAWSGCPAzureDockerKubernetesTerraformCI/CDGitHub ActionsPrometheusGrafanaDevOpsSite Reliability EngineeringRESTMachine LearningAgentic AIMLOpsIncident ResponseComplianceHIPAAResearch

About the role

  • And part of it is genuinely new territory, figuring out what SRE means when the systems you're operating make autonomous decisions.
  • The infrastructure patterns for the agentic systems of the future don't exist yet.
  • The infrastructure contract changes when your workloads have agency.

Responsibilities

  • Own and evolve our IaC stack: Terraform and Kubernetes across AWS, GCP, and Azure
  • Design and maintain CI/CD pipelines that give teams fast, trustworthy feedback from commit to production
  • Build operational foundations: monitoring, alerting, incident response, and the new patterns that emerge when AI systems are participants in that response
  • 5+ years building and operating production infrastructure in DevOps or SRE roles

Requirements

  • Deep experience with at least 1 major cloud provider (AWS, GCP, or Azure): networking, IAM, cost management, the operational realities of production workloads
  • Solid Docker and Kubernetes experience in production.
  • Experience designing and maintaining CI/CD pipelines (GitHub Actions, GitLab CI, or similar)
  • Scripting proficiency in Python, Bash, or similar
  • MLOps experience isn't required, but you'll be closer to that boundary than most infra roles.

Nice to have

  • Multi-region and multi-cloud experience across 2+ providers
  • Experience with single-tenant or on-prem deployments alongside multi-tenant SaaS
  • Familiarity with GitOps patterns and progressive delivery
  • Familiarity with the Grafana stack (Prometheus, Grafana, Loki) or equivalent
  • Experience with compliance frameworks (HIPAA, SOC 2) and how they shape infrastructure decisions in regulated environments
  • Background supporting ML or research workflows moving to production: model deployment, pipeline orchestration, or similar
  • You've thought about what observability means for non-deterministic systems and have opinions about it
  • The infrastructure patterns for autonomous AI systems are still being written.

Benefits

  • Build observability primitives for agentic workflows, tracing agent decisions and execution paths, not just service latency and pod health
  • Work across engineering teams to meet the reliability and compliance requirements of the institutions we serve (SOC 2, HIPAA, regulated environments in healthcare and energy)

Company info

  • Percepta's mission is to transform critical institutions with applied AI. We care that industries that power the world (healthcare, manufacturing, energy) benefit from frontier technology.
  • We collaborate with industry-leading customers to drive AI transformation. We bring:
  • Forward-deployed expertise in engineering, product, and research
  • Mosaic, our in-house toolkit for rapidly deploying agentic architectures
  • Percepta's mission is to transform critical institutions with applied AI.
  • We care that industries that power the world (healthcare, manufacturing, energy) benefit from frontier technology.
  • We collaborate with industry-leading customers to drive AI transformation.
  • Strategic partnerships with Anthropic, McKinsey, AWS, and the General Catalyst portfolio
  • Our team is a fast-growing group of Applied AI Engineers, Embedded Product Managers, and Researchers motivated by getting frontier AI into the places that actually run the world.
  • Percepta is a direct partnership with General Catalyst.
  • What we're doing matters and we have to give a shit.
  • Internally, that means fixing badness when you find it.
  • Externally, it means honoring the trust our customers place in us with their most important problems.

This listing is sourced directly from Percepta's careers page and normalized into a canonical job model.