Centific

Centific

Senior Staff Research Scientist, Reinforcement Learning

East Palo Alto, CA · Staff+

Sponsorship not specified$250k-$300kDetected 31 days ago
PythonCode ReviewMachine LearningComputer VisionLLMsRAGAgentic AILogisticsResearch

About the role

  • About Centific Centific is a frontier AI data foundry that curates diverse, high-quality data, using our purpose-built technology platforms to empower the Magnificent Seven and our enterprise clients with safe, scalable AI deployment.
  • Our team includes more than 150 PhDs and data scientists, along with more than 4,000 AI practitioners and engineers.

Responsibilities

  • Design simulation environments and digital twins for enterprise workflows
  • Build pipelines that convert human-labeled traces and verifiable signals into training data
  • Design reward functions and verifiers that resist reward hacking and reflect real task outcomes
  • Mentor researchers and engineers; drive technical direction through influence
  • 5+ years hands-on RL - environment design, reward engineering, policy optimization - with at least one production deployment

Requirements

  • 7+ years in ML/AI research or engineering
  • 3+ years at senior/staff level
  • 3+ years fine-tuning LLMs with hands-on RL post-training (RLHF, DPO, GRPO, PPO)
  • Working knowledge of modern post-training and rollout-serving libraries (TRL, veRL, OpenRLHF, SkyRL)
  • Hands-on experience with Gymnasium-based environments and reward engineering (sparse vs. dense)

Nice to have

  • Publications at NeurIPS, ICML, ICLR, ACL, COLM, or similar venues
  • Open-source contributions to post-training or agent frameworks (TRL, veRL, OpenRLHF, SkyRL)
  • Experience with Offline RL (CQL, IQL), Model-based RL / World Models, or Hierarchical RL
  • Background in synthetic data generation, simulation, or world models
  • Distributed training on GPU clusters
  • Shape a new discipline at the intersection of post-training, simulation, and enterprise AI.
  • Ship your science.
  • Work alongside NVIDIA, Microsoft, and the global AI community.

Skills

  • About Centific
  • Our zero-distance innovation™ solutions for GenAI can reduce GenAI costs by up to 80% and bring solutions to market 50% faster.

Compensation

  • $250k-$300k

Benefits

  • Architect multi-turn, tool-using agents with closed learning loops
  • MS or PhD in Computer Science, Machine Learning, or related field (or equivalent)

Visa & Work Authorization

  • pplicants will receive consideration for employment without regard to race, color, religion, national origin, ancestry, citizenship status, age, mental or physical disability, medical condition, sex (including pregnancy)

This listing is sourced directly from Centific's careers page and normalized into a canonical job model.