FirstPrinciples

FirstPrinciples

AI & HPC Infrastructure Engineer

Ontario, Canada - Remote

Sponsorship not specifiedDetected 54 days ago
Node.jsAWSGCPAzureCloud PlatformsKubernetesTerraformAnsibleHelmCI/CDLinuxPrometheusGrafanaDatadogMachine LearningLLMsA/B TestingIncident ResponseEmbedded SystemsResearch

About the role

  • We're a fast-growing, remote-first team of builders, researchers, engineers, and thinkers working across Canada, the US, the UK, and expanding globally.
  • This is work that sits somewhere between creativity and rigorous thinking, and often requires comfort with ambiguity and iteration.
  • You'll play a central role in shaping how we run compute at FirstPrinciples.

Responsibilities

  • Design, deploy, and operate Kubernetes infrastructure for AI inference, research, and engineering workloads
  • Set up and manage GPU and HPC-style compute environments, including scheduling, utilization, job management, and node-level troubleshooting
  • Build and manage Linux-based compute environments, including provisioning, networking, storage, monitoring, access control, and lifecycle management
  • Partner with ML engineers, researchers, and software engineers to understand workload requirements and translate them into practical infrastructure designs
  • Build tooling, documentation, runbooks, and operational practices that help the team move quickly without making infrastructure fragile or opaque
  • FirstPrinciples is a research organization building AI infrastructure for discovery in fundamental science.
  • Currently, our work focuses on building systems like Theo, the AI Physicist, which is a domain-specialized system for research in fundamental physics.
  • What brings us together is a shared curiosity about how the universe works, and a belief that we can build systems that help us explore it more effectively.
  • We spend our time working on questions that don't have clear answers, like how to design AI that can reason through scientific problems, and how the scientific process as a whole might evolve.
  • If you're someone who enjoys tackling big, abstract problems and building the infrastructure that makes ambitious research possible, you'll likely find the work here interesting.

Requirements

  • Strong infrastructure builder with experience operating production, research, cloud, or high-performance compute systems
  • you can take ambiguous infrastructure needs and turn them into working systems
  • Hands-on experience with production-grade LLM inference and serving engines, such as vLLM, SGLang, or TensorRT
  • Experience working at an AI company, ML infrastructure team, research lab, university compute environment, HPC center, or scientific computing organization
  • Experience supporting model inference, model serving, distributed training, high-throughput batch workloads, or internal ML platforms
  • Hands-on experience with Slurm or similar HPC schedulers, including job scheduling, resource allocation, queue management, and cluster configuration
  • Experience operating GPU infrastructure, including NVIDIA drivers, CUDA, container runtimes, scheduling, utilization, and hardware failure modes
  • Experience with RDMA, InfiniBand, high-performance networking, distributed filesystems (ie.
  • Experience with Kubernetes operators, custom controllers, CRDs, or platform tooling for AI/ML workloads
  • Experience with Prometheus, Grafana, Loki, OpenTelemetry, Datadog, or similar monitoring, logging, and observability tools

Benefits

  • The opportunity to work on foundational problems at the intersection of AI and physics
  • Own the reliability and operational health of infrastructure systems, including monitoring, alerting, incident response, capacity planning, and performance tuning

This listing is sourced directly from FirstPrinciples's careers page and normalized into a canonical job model.