ElastixAI
AI Inference Infrastructure Software Engineer (Kubernetes / Cloud)
Seattle
Sponsorship not specifiedDetected 78 days ago
PythonGoBashAWSGCPKubernetesTerraformAnsibleLinuxSite Reliability EngineeringMachine Learning
About the role
- If you're excited about pushing AI performance to physical limits and shaping the future of large-scale inference, we'd love to meet you.
- This is a hands-on role with broad surface area.
Responsibilities
- Build, operate, and evolve ElastixAI's Kubernetes infrastructure powering our Token-as-a-Service capability.
- Manage and harden our AWS, GCP, and on-prem infrastructure, including networking, storage, IAM, and observability layers tied to our services.
- Develop tooling and automation in Python, Bash, Rust, and Go to streamline deployments, rollouts, autoscaling, and incident response.
- Partner with the ML and runtime teams to productionize new inference capabilities, model deployments, and routing strategies.
- the best performance comes from holistic co-design, where every layer, from model architecture to kernels to silicon, works in harmony.
Requirements
- 3-5 years of hands-on Kubernetes experience, including EKS, GKE, and/or self-hosted clusters.
- 2-3 years of production experience operating workloads on AWS or GCP.
- Proven track record running ML or inference services at scale on Kubernetes in production.
- Solid coding skills in Python, Bash and proficiency in Go
- Experience with infrastructure-as-code (Terraform, Pulumi), OS configuration state (Ansible, Puppet, Salt) and GitOps workflows (Argo CD, Flux).
- Experience in OS configuration tooling.
- Familiarity with AI inference and/or training workflows and the operational patterns around them.
Nice to have
- MS/PhD in Computer Science, Software Engineering, or a related field.
- Experience with inference servers and runtimes (e.g., Triton, vLLM, TGI) and model-serving patterns (batching, streaming, KV-cache aware routing).
- Exposure to heterogeneous accelerators beyond GPUs (FPGAs, custom ASICs).
- Background in observability, SRE, or performance engineering for latency-sensitive services.
Compensation
- Competitive compensation and startup equity package
Benefits
- The opportunity to work on challenging problems at the intersection of ML, software, and systems.
- Competitive compensation and startup equity package
- Comprehensive medical, dental, and vision coverage (premiums 100% paid by employer)
- Flexible Time Off (FTO)
- Paid parental leave
- Gym or fitness benefit
- Commuter benefit
- Investment in employee learning & development
Company info
- Help define the platform roadmap as we scale from early customers to broad production deployments.
Apply directly at ElastixAI →Create a free account for alerts like thisView ElastixAI immigration profile
This listing is sourced directly from ElastixAI's careers page and normalized into a canonical job model.