Centific
Applied Reinforcement Learning Engineer
Remote Work( USA)
Sponsorship not specified$150k-$300kDetected 14 days ago
PythonMachine LearningTensorFlowPyTorchLLMsRAGAgentic AIA/B TestingLogisticsResearch
About the role
- This role requires deep expertise in both classical RL methodologies and modern LLM-based agent architectures.
- You'll shape our product direction and help make RL accessible to enterprise customers who need safe, compliant ways to improve their AI systems.
- Centific AI Research advances foundational AI models and applications through reinforcement learning, alignment, and human-centered intelligence.
Responsibilities
- Design and build custom RL environments (digital twins) simulating enterprise workflows: document processing, compliance, onboarding, support automation
- Build end-to-end pipelines converting human-labeled traces into RL training data
- Design reward functions, verifiers, and validation frameworks for pre-deployment testing
- Environment Design
- Design and build custom
- Create governed, compliant AI systems enterprises can trust.
- As an Applied RL Engineer, you will design and build RL environments that simulate complex enterprise workflows and train intelligent agents within them.
Requirements
- Deep RL expertise: 3+ years hands-on experience with environment design, reward engineering, policy optimization
- LLM post-training: Experience fine-tuning LLMs using RLHF, DPO, PPO, or similar
- 3+ years hands-on experience with environment design, reward engineering, policy optimization
- Experience with LLM-based agents, tool use, multi-step reasoning
Skills
- Software engineering beyond research with scalable pipelines and training infrastructure
- Agentic AI: Experience with LLM-based agents, tool use, multi-step reasoning
- Technical stack: Strong Python; Gymnasium, RLlib, Stable Baselines; PyTorch/JAX/TensorFlow
- Education: MS/PhD in CS, ML, or related field (or equivalent experience)
- Publications at NeurIPS, ICML, ICLR, ACL, or similar venues
- Enterprise workflow experience in healthcare, finance, logistics, or compliance
- Open-source contributions to CleanRL, TRL, veRL, or agent frameworks
- Experience with world models, synthetic data generation, and simulation
- Distributed training and large-scale RL experimentation
- Why Join Centific
- Ship your science: See your research power real systems across healthcare, finance, and safety
Compensation
- Salary: $150K - $300K Annually
Benefits
- Architect multi-step reasoning agents with tool-calling and closed learning loops
- See your research power real systems across healthcare, finance, and safety
Company info
- Our mission is to transform data, signals, and human insight into next-generation intelligent systems that redefine enterprise intelligence.
- We're building a governed RL environment platform that enables enterprises to safely iterate and improve AI agent workflows through simulation-based learning, bridging human-labeled signal creation with automated RL training for high-stakes operations.
- Role Overview
- You'll work at the intersection of RL research and production systems, translating customer requirements into bespoke simulation environments and post-training pipelines that deliver measurable improvements to AI agent performance.
- Core RL Competencies
- Foundational RL
Visa & Work Authorization
- pplicants will receive consideration for employment without regard to race, color, religion, national origin, ancestry, citizenship status, age, mental or physical disability, medical condition, sex (including pregnancy)
Apply directly at Centific →Create a free account for alerts like thisView Centific immigration profile
This listing is sourced directly from Centific's careers page and normalized into a canonical job model.