ElastixAI

ElastixAI

AI Inference Infrastructure Software Engineer (Kubernetes / Cloud)

Seattle

Sponsorship not specifiedDetected 78 days ago
PythonGoBashAWSGCPKubernetesTerraformAnsibleLinuxSite Reliability EngineeringMachine Learning

About the role

  • If you're excited about pushing AI performance to physical limits and shaping the future of large-scale inference, we'd love to meet you.
  • This is a hands-on role with broad surface area.

Responsibilities

  • Build, operate, and evolve ElastixAI's Kubernetes infrastructure powering our Token-as-a-Service capability.
  • Manage and harden our AWS, GCP, and on-prem infrastructure, including networking, storage, IAM, and observability layers tied to our services.
  • Develop tooling and automation in Python, Bash, Rust, and Go to streamline deployments, rollouts, autoscaling, and incident response.
  • Partner with the ML and runtime teams to productionize new inference capabilities, model deployments, and routing strategies.
  • the best performance comes from holistic co-design, where every layer, from model architecture to kernels to silicon, works in harmony.

Requirements

  • 3-5 years of hands-on Kubernetes experience, including EKS, GKE, and/or self-hosted clusters.
  • 2-3 years of production experience operating workloads on AWS or GCP.
  • Proven track record running ML or inference services at scale on Kubernetes in production.
  • Solid coding skills in Python, Bash and proficiency in Go
  • Experience with infrastructure-as-code (Terraform, Pulumi), OS configuration state (Ansible, Puppet, Salt) and GitOps workflows (Argo CD, Flux).
  • Experience in OS configuration tooling.
  • Familiarity with AI inference and/or training workflows and the operational patterns around them.

Nice to have

  • MS/PhD in Computer Science, Software Engineering, or a related field.
  • Experience with inference servers and runtimes (e.g., Triton, vLLM, TGI) and model-serving patterns (batching, streaming, KV-cache aware routing).
  • Exposure to heterogeneous accelerators beyond GPUs (FPGAs, custom ASICs).
  • Background in observability, SRE, or performance engineering for latency-sensitive services.

Compensation

  • Competitive compensation and startup equity package

Benefits

  • The opportunity to work on challenging problems at the intersection of ML, software, and systems.
  • Competitive compensation and startup equity package
  • Comprehensive medical, dental, and vision coverage (premiums 100% paid by employer)
  • Flexible Time Off (FTO)
  • Paid parental leave
  • Gym or fitness benefit
  • Commuter benefit
  • Investment in employee learning & development

Company info

  • Help define the platform roadmap as we scale from early customers to broad production deployments.

This listing is sourced directly from ElastixAI's careers page and normalized into a canonical job model.