Quantiphi
Senior Platform Engineer
USA - Remote · Senior
Stay score
odds of building a lasting career here
Thin sponsorship signal and lottery-bound. A low-probability bet with your clock running. Prioritize cap-exempt roles and proven entry-level sponsors first.
Lottery odds assume a STEM candidate.
Personalize to your clock →Employer immigration record
from this employer's Department of Labor filings
Files H-1B transfers
Sourced from Department of Labor LCA, PERM and prevailing-wage disclosure data. Employer matching is by name, so figures may be split across an employer's legal entities. Absence of a filing means none appears in our copy of the data, not that none exists.
Community outcomes
No reports yet — be the first to help the next applicant.
About the role
- This role is ideal for someone with deep hands-on experience in GPU profiling, distributed training, and high-performance compute environments.
Responsibilities
- Design and implement scalable infrastructure for LLM and GenAI workloads across multi-GPU environments
- Perform GPU profiling, benchmarking, and performance optimization for distributed training workloads
- Manage and schedule compute-intensive jobs using Slurm-based clusters and OpenShift/Kubernetes environments
- Enable and optimize the NVIDIA GPU stack (CUDA, cuDNN, NCCL, Triton, RAPIDS, etc.)
- Collaborate with cross-functional teams to deploy models in research and production environments
- Build and support GenAI pipelines (fine-tuning, RAG, multi-modal inferencing, LLMOps)
- Develop reusable infrastructure templates using tools like Terraform and Helm
- Contribute to internal innovation (PoCs, workshops) and support client-facing delivery engagements
- 21x Google Cloud Partner of the Year awards in the last 8 years.
- 3x NVIDIA Partner of the Year titles.
Requirements
- Strong experience with Slurm and distributed training environments
- Hands-on expertise with Red Hat OpenShift and/or Kubernetes
- Deep knowledge of the NVIDIA GPU ecosystem (CUDA, cuDNN, NCCL, Nsight, Triton/TensorRT)
- Experience deploying GenAI workloads (LLM fine-tuning, RAG pipelines, multi-modal systems)
- Familiarity with Infrastructure-as-Code tools (Terraform, Ansible)
- Experience with cloud GPU environments (GCP, Azure, AWS, OCI) and/or on-prem GPU clusters
- Experience with NVIDIA NIMs, DGX systems, or GPU-accelerated containers
- Knowledge of LLMOps frameworks and MLOps integration
- Familiarity with vector databases and retrieval systems for RAG architectures
- Comfortable working in client-facing environments and collaborating with AI solution teams
Nice to have
- Experience working with FHIR R4, HL7 v2, or SMART on FHIR
- Integration with EHR systems (e.g., Epic)
- Exposure to clinical workflows, CDS Hooks, or patient-facing applications
- Make an impact at one of the world's fastest-growing AI-first digital engineering companies.
- Upskill and discover your potential as you solve complex challenges in cutting-edge areas of technology alongside passionate, talented colleagues.
- Work where innovation happens - work with disruptive innovators in a research-focused organization with 60+ patents filed across various disciplines.
- Stay ahead of the curve, immerse yourself in breakthrough AI, ML, data, and cloud technologies and gain exposure working with Fortune 500 companies.
Benefits
- We offer first-in-class industry solutions across Healthcare, Financial Services, Consumer Goods, Manufacturing, and more, powered by cutting-edge Generative AI and Agentic AI accelerators.
Company info
- Be part of a trailblazing team that's shaping the future of AI, ML, and cloud innovation.
This listing is sourced directly from Quantiphi's careers page and normalized into a canonical job model.