FirstPrinciples
AI & HPC Infrastructure Engineer
Ontario, Canada - Remote
Sponsorship not specifiedDetected 54 days ago
Node.jsAWSGCPAzureCloud PlatformsKubernetesTerraformAnsibleHelmCI/CDLinuxPrometheusGrafanaDatadogMachine LearningLLMsA/B TestingIncident ResponseEmbedded SystemsResearch
About the role
- We're a fast-growing, remote-first team of builders, researchers, engineers, and thinkers working across Canada, the US, the UK, and expanding globally.
- This is work that sits somewhere between creativity and rigorous thinking, and often requires comfort with ambiguity and iteration.
- You'll play a central role in shaping how we run compute at FirstPrinciples.
Responsibilities
- Design, deploy, and operate Kubernetes infrastructure for AI inference, research, and engineering workloads
- Set up and manage GPU and HPC-style compute environments, including scheduling, utilization, job management, and node-level troubleshooting
- Build and manage Linux-based compute environments, including provisioning, networking, storage, monitoring, access control, and lifecycle management
- Partner with ML engineers, researchers, and software engineers to understand workload requirements and translate them into practical infrastructure designs
- Build tooling, documentation, runbooks, and operational practices that help the team move quickly without making infrastructure fragile or opaque
- FirstPrinciples is a research organization building AI infrastructure for discovery in fundamental science.
- Currently, our work focuses on building systems like Theo, the AI Physicist, which is a domain-specialized system for research in fundamental physics.
- What brings us together is a shared curiosity about how the universe works, and a belief that we can build systems that help us explore it more effectively.
- We spend our time working on questions that don't have clear answers, like how to design AI that can reason through scientific problems, and how the scientific process as a whole might evolve.
- If you're someone who enjoys tackling big, abstract problems and building the infrastructure that makes ambitious research possible, you'll likely find the work here interesting.
Requirements
- Strong infrastructure builder with experience operating production, research, cloud, or high-performance compute systems
- you can take ambiguous infrastructure needs and turn them into working systems
- Hands-on experience with production-grade LLM inference and serving engines, such as vLLM, SGLang, or TensorRT
- Experience working at an AI company, ML infrastructure team, research lab, university compute environment, HPC center, or scientific computing organization
- Experience supporting model inference, model serving, distributed training, high-throughput batch workloads, or internal ML platforms
- Hands-on experience with Slurm or similar HPC schedulers, including job scheduling, resource allocation, queue management, and cluster configuration
- Experience operating GPU infrastructure, including NVIDIA drivers, CUDA, container runtimes, scheduling, utilization, and hardware failure modes
- Experience with RDMA, InfiniBand, high-performance networking, distributed filesystems (ie.
- Experience with Kubernetes operators, custom controllers, CRDs, or platform tooling for AI/ML workloads
- Experience with Prometheus, Grafana, Loki, OpenTelemetry, Datadog, or similar monitoring, logging, and observability tools
Benefits
- The opportunity to work on foundational problems at the intersection of AI and physics
- Own the reliability and operational health of infrastructure systems, including monitoring, alerting, incident response, capacity planning, and performance tuning
Apply directly at FirstPrinciples →Create a free account for alerts like thisView FirstPrinciples immigration profile
This listing is sourced directly from FirstPrinciples's careers page and normalized into a canonical job model.