Yotta Energy

Yotta Energy

Research Engineer - AI Systems

United States · Full-time

Sponsorship not specifiedDetected 24 days ago
PythonAWSPyTorchNLPLLMsResearchCommunicationProblem Solving

About the role

  • We enable training and inference across NVIDIA GPUs, AMD GPUs, and AWS Trainium, helping AI companies achieve the best performance and economics across heterogeneous hardware.
  • You will work at the intersection of AI Systems, Compiler and Runtime Optimization, Distributed Training & Inference, GPU/Accelerator Kernel Development, and Large Language Model Infrastructure.
  • Enjoy a flexible, remote work environment that values innovation and autonomy. 📩 How to Apply Interested candidates should apply directly or send their resume and a brief cover letter to careers@yottalabs.ai.

Responsibilities

  • Design and implement high-performance kernels for Attention, MoE, GEMM, collective communication, and quantization.
  • Optimize kernels for NVIDIA, AMD, and AWS Trainium.
  • Develop custom operators and graph optimizations using Neuron SDK, PyTorch/XLA, Torch Dynamo, and Neuron Compiler.
  • Design scalable distributed training and inference solutions across thousands of accelerators.
  • Collaborate with experts from leading institutions and tech companies.

Requirements

  • Proficiency in AI programming languages such as Python and C++.
  • Experience with CUDA, Triton, ROCm/HIP, or AWS Neuron.
  • Strong understanding of AI frameworks (e.g., PyTorch, Dynamo, LMCache), model architectures and profiling tools (e.g. Nsight, ROCm Profiler, or Neuron Profiler).
  • Experience with with FlashAttention, PagedAttention, MoE, RLHF, or distributed AI systems.

Compensation

  • Competitive compensation with equity.

Benefits

  • Competitive compensation with equity.
  • Enjoy a flexible, remote work environment that values innovation and autonomy. 📩

Company info

  • Apply: careers@yottalabs.ai
  • 🧠 About Yotta Labs
  • Yotta Labs is building the next generation multi-silicon AI cloud and runtime platform to power the world's most demanding AI workloads.
  • Our mission is to provide high-performance AI computing and Model API services, enabling AI companies, research labs, and enterprises to train, deploy and integrate cutting-edge models at scale.
  • 🛠️ Role Overview
  • We are seeking a highly motivated AI Systems Research Engineer specializing in Trainium, GPU kernels, and LLM systems optimization.
  • Your work will directly impact the scalability and performance of AI applications deployed on our platform.
  • 🎯 Responsibilities
  • Improve performance of vLLM, SGLang, TensorRT-LLM, and custom inference runtimes.

This listing is sourced directly from Yotta Energy's careers page and normalized into a canonical job model.