Gimlet Media

Gimlet Media

Member of Technical Staff - Infrastructure

San Francisco, CA · Staff+

Sponsorship not specifiedDetected 40 days ago
PythonGoNode.jsDistributed SystemsCloud PlatformsKubernetesTerraformAnsibleHelmLinuxPrometheusGrafanaSite Reliability EngineeringPlatform EngineeringIncident ResponseEmbedded Systems

About the role

  • Unlike traditional cloud platforms built around a single hardware ecosystem, Gimlet's infrastructure spans multiple accelerator vendors and architectures.
  • You will work across bare metal, Linux, Kubernetes or cluster schedulers, high-speed networking, observability, provisioning, and incident response.
  • Debug complex production issues across Linux, networking, storage, drivers, firmware, and orchestration layers.

Responsibilities

  • You will partner closely with distributed systems, runtime, compiler, and hardware teams to ensure Gimlet's infrastructure can support demanding AI workloads at production scale.
  • Design, deploy, and operate large-scale CPU, GPU, and accelerator clusters powering production AI inference.
  • Build automation for provisioning, configuration, upgrades, validation, and lifecycle management.
  • Design and scale provisioning systems for heterogeneous bare-metal infrastructure across multiple datacenters and hardware vendors.Operate cluster scheduling, resource allocation, isolation, quotas, and utilization systems.
  • Build and operate high-performance networking infrastructure, including RDMA-enabled environments and accelerator interconnects.
  • Work with distributed systems and runtime teams to support low-latency, high-throughput inference workloads.

Requirements

  • Comfort working in a fast-moving startup environment with high ownership and ambiguity.
  • Experience with bare-metal provisioning, PXE/iPXE, image pipelines, BIOS/firmware management, or rack bring-up.
  • Experience with multi-tenant cluster isolation, quota systems, fair scheduling, or usage accounting.
  • Experience debugging distributed workload performance across compute, memory, network, and storage bottlenecks.
  • Familiarity with heterogeneous hardware environments across NVIDIA, AMD, Intel, ARM, or emerging accelerators.
  • Experience with GPU or accelerator infrastructure, including drivers, firmware, CUDA/ROCm stacks, or hardware validation.
  • Familiarity with high-performance networking such as InfiniBand, RoCE, high-speed Ethernet, or datacenter fabrics.

Skills

  • Strong automation skills using tools such as Terraform, Ansible, Helm, Python, Go, or equivalent.

Benefits

  • Build observability for cluster health, capacity, performance, failures, and workload behavior.

Company info

  • Gimlet is building the next generation of AI infrastructure: large-scale AI datacenters and the orchestration platform that coordinates them.
  • The future of AI will require vastly more compute than exists today. But as AI workloads become more complex and new hardware architectures emerge, simply deploying more GPUs isn't enough. The challenge is making increasingly diverse compute work together.
  • Gimlet's platform intelligently partitions and routes workloads across heterogeneous hardware, enabling step-function improvements in performance and efficiency. Customers deploy through production-grade APIs without needing to think about hardware selection, placement, or optimization.
  • We work with foundation labs, hyperscalers, and AI-native companies to power production workloads at massive scale and help define the infrastructure layer for the future of AI.
  • The future of AI will require vastly more compute than exists today.
  • But as AI workloads become more complex and new hardware architectures emerge, simply deploying more GPUs isn't enough.
  • The challenge is making increasingly diverse compute work together.
  • Gimlet's platform intelligently partitions and routes workloads across heterogeneous hardware, enabling step-function improvements in performance and efficiency.
  • Customers deploy through production-grade APIs without needing to think about hardware selection, placement, or optimization.
  • large-scale AI datacenters and the orchestration platform that coordinates them.

This listing is sourced directly from Gimlet Media's careers page and normalized into a canonical job model.