Yotta Energy
Research Engineer - AI Systems
United States · Full-time
Sponsorship not specifiedDetected 24 days ago
PythonAWSPyTorchNLPLLMsResearchCommunicationProblem Solving
About the role
- We enable training and inference across NVIDIA GPUs, AMD GPUs, and AWS Trainium, helping AI companies achieve the best performance and economics across heterogeneous hardware.
- You will work at the intersection of AI Systems, Compiler and Runtime Optimization, Distributed Training & Inference, GPU/Accelerator Kernel Development, and Large Language Model Infrastructure.
- Enjoy a flexible, remote work environment that values innovation and autonomy. 📩 How to Apply Interested candidates should apply directly or send their resume and a brief cover letter to careers@yottalabs.ai.
Responsibilities
- Design and implement high-performance kernels for Attention, MoE, GEMM, collective communication, and quantization.
- Optimize kernels for NVIDIA, AMD, and AWS Trainium.
- Develop custom operators and graph optimizations using Neuron SDK, PyTorch/XLA, Torch Dynamo, and Neuron Compiler.
- Design scalable distributed training and inference solutions across thousands of accelerators.
- Collaborate with experts from leading institutions and tech companies.
Requirements
- Proficiency in AI programming languages such as Python and C++.
- Experience with CUDA, Triton, ROCm/HIP, or AWS Neuron.
- Strong understanding of AI frameworks (e.g., PyTorch, Dynamo, LMCache), model architectures and profiling tools (e.g. Nsight, ROCm Profiler, or Neuron Profiler).
- Experience with with FlashAttention, PagedAttention, MoE, RLHF, or distributed AI systems.
Compensation
- Competitive compensation with equity.
Benefits
- Competitive compensation with equity.
- Enjoy a flexible, remote work environment that values innovation and autonomy. 📩
Company info
- Apply: careers@yottalabs.ai
- 🧠 About Yotta Labs
- Yotta Labs is building the next generation multi-silicon AI cloud and runtime platform to power the world's most demanding AI workloads.
- Our mission is to provide high-performance AI computing and Model API services, enabling AI companies, research labs, and enterprises to train, deploy and integrate cutting-edge models at scale.
- 🛠️ Role Overview
- We are seeking a highly motivated AI Systems Research Engineer specializing in Trainium, GPU kernels, and LLM systems optimization.
- Your work will directly impact the scalability and performance of AI applications deployed on our platform.
- 🎯 Responsibilities
- Improve performance of vLLM, SGLang, TensorRT-LLM, and custom inference runtimes.
Apply directly at Yotta Energy →Create a free account for alerts like thisView Yotta Energy immigration profile
This listing is sourced directly from Yotta Energy's careers page and normalized into a canonical job model.