Inception

Inception

Member of Technical Staff, Reinforcement Learning

San Mateo, USA · Staff+

Sponsorship not specifiedDetected 134 days ago
AlgorithmsMachine LearningDeep LearningPyTorchNLPLLMsResearch

About the role

  • We seek experienced scientists and engineers with deep expertise in post-training large language models through reinforcement learning.

Responsibilities

  • Design, develop, and optimize RL training pipelines (PPO, DPO, RLHF, and novel approaches) for diffusion-based LLMs.
  • Build and iterate on reward models, reward shaping strategies, and evaluation of reward quality.
  • Implement innovative approaches for fine-tuning and scaling generative AI models.
  • Research and implement techniques for controlled text generation and constraint satisfaction.

Nice to have

  • Work on data preprocessing pipelines, model evaluation, and alignment to enterprise use cases.
  • Improve training stability, efficiency, and reproducibility of RL workloads.
  • BS/MS/PhD in Computer Science or a related field (or equivalent experience).
  • At least 2 years of experience working on ML projects in PyTorch (or equivalent), preferably in a research lab or engineering role.
  • Familiarity with training and inference in diffusion models.
  • Preferred Skills Extensive experience training transformer-based language models from scratch.
  • Knowledge of advanced training techniques (mixed precision, gradient accumulation, etc.).
  • Experience with LLM serving frameworks like vLLM, SGLang, or TensorRT.

Benefits

  • Excellent familiarity with transformers and core LLM concepts (autoregressive pretraining, instruction tuning, in-context learning, KV caching).
  • Hands-on experience with reinforcement learning from human feedback (RLHF), PPO, DPO, or related post-training methods.
  • Experience training deep learning models at scale in distributed computing environments.
  • Experience designing and implementing reward models or preference learning systems.

This listing is sourced directly from Inception's careers page and normalized into a canonical job model.