Sciforium
Distributed Training and Inference Engineer
San Francisco
Sponsorship not specifiedDetected 27 days ago
PythonC++Node.jsDistributed SystemsFull-Stack DevelopmentMachine LearningPyTorchLLMsAI OrchestrationSystems EngineeringElectrical EngineeringResearch
About the role
- In this role, you will work across the entire machine learning infrastructure from low-level CUDA/ROCm runtimes to high-level frameworks like JAX and PyTorch to ensure our distributed training systems are fast, scalable, stable, and efficient.
- This position is ideal for someone who loves deep systems engineering, debugging complex hardware-software interactions, and optimizing performance at every layer of the ML stack.
- You will play a pivotal role in enabling the training and deployment of next-generation LLMs and generative AI models.
Responsibilities
- Software Stack Maintenance: Maintain, update, and optimize critical ML libraries and frameworks including JAX, PyTorch, CUDA, and ROCm across multiple environments and hardware configurations.
- End-to-End Stack Ownership: Build, maintain, and continuously improve the entire ML software stack from ROCm/CUDA drivers to high-level JAX/PyTorch tooling.
- System Integration: Continuously integrate and validate modules for runtime correctness, memory efficiency, and scalability across multi-node GPU/accelerator clusters.
- Profiling & Performance Analysis: Conduct detailed profiling of compilation graphs, training workloads, and runtime execution to optimize performance and eliminate bottlenecks.
- Collaborate with research, infrastructure, and kernel engineering teams to improve system throughput, stability, and developer experience.
- Hands-on experience maintaining or building ML training stacks involving CUDA, ROCm, NCCL, XLA, or similar technologies.
Requirements
- Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or related technical fields.
- Experience with multi-node distributed training systems and orchestration frameworks (DTensor, GSPMD, etc.).
Nice to have
- Extensive experience with the XLA/JAX stack, including compilation internals and custom lowering paths.
- Familiarity with distributed serving or large-scale inference frameworks (e.g., vLLM, TensorRT, FasterTransformer).
- Background in GPU kernel optimization or accelerator-aware model partitioning.
- Strong understanding of low-level C++ building blocks used in ML frameworks (e.g., XLA, CUDA kernels, custom ops).
- Daily lunch, snacks, and beverages
Compensation
- Competitive salary and equity
Benefits
- Medical, dental, and vision insurance
- Flexible time off
- Competitive salary and equity
Equal opportunity
- Sciforium is an equal opportunity employer.
- All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.
Visa & Work Authorization
- Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.
Apply directly at Sciforium →Create a free account for alerts like thisView Sciforium immigration profile
This listing is sourced directly from Sciforium's careers page and normalized into a canonical job model.