Hark
Member of Technical Staff, Post-training
San Jose · Staff+ · Full-time
Sponsorship not specified$180k-$450kDetected 83 days ago
PythonAlgorithmsMachine LearningPyTorchLLMsAgentic AIRoboticsResearch
About the role
- This role sits at the frontier of a rapidly emerging discipline - one where reinforcement learning, simulation, and large-scale model training converge to produce agents that can reason, plan, and act over long horizons.
Responsibilities
- Design and implement post-training strategies, primarily RL-based, to develop strong coding agents capable of multi-step reasoning, tool use, and long-horizon task completion.
- Build and scale simulation and scaffolding environments for agentic RL: code execution sandboxes, computer use environments, tool-calling harnesses, and verifiable reward signals.
- Develop reward modeling pipelines - including outcome-based, execution-based, and process-based reward signals - and iterate on them based on training dynamics.
- Design and run rigorous ablations to understand how algorithm choice, data mixture, reward shaping, and scale interact in the agentic setting.
- Build evaluation frameworks grounded in real agent tasks - code correctness, execution success, multi-step tool use - to measure progress and guide iteration.
- Collaborate with mid-training, infrastructure, and product teams to translate research insights into durable improvements on the model.
- Hark is an artificial intelligence company building advanced, personalized intelligence.
- We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines.
- Experience building or working within simulation or execution environments (e.g., code interpreters, sandboxed execution, game environments, robotics simulators).
- Proven ability to design and execute rigorous experiments, with strong intuition for diagnosing training failures and scaling bottlenecks.
Requirements
- Proficiency in Python and PyTorch
- comfort working across research and systems code.
- Ability to work in a fast-moving, research-forward environment where the right approach is often unknown at the outset.
- Proficiency in Python and PyTorch; comfort working across research and systems code.
Nice to have
- Experience with RL algorithms applied to language or code: RLHF, DPO, GRPO, PPO, or similar paradigms in the LLM setting.
- Familiarity with coding agent benchmarks and evaluation environments (e.g., SWE-bench, HumanEval, LiveCodeBench, competitive programming judges).
- Background in reward modeling - outcome-based, process-based, or learned reward signals.
- Prior work on computer use, GUI agents, or tool-using LLMs (e.g., OSWorld, WebArena-style tasks).
- Experience training or scaling models at 10B+ parameters, with attention to efficiency, stability, and GPU utilization.
- Contributions to open-source ML projects or publications at top venues (NeurIPS, ICML, ICLR, EMNLP, COLM, etc.).
- The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
- This information will be shared if an employment offer is extended.
Skills
- Scale synthetic data generation and trajectory distillation pipelines that feed RL training and improve sample efficiency.
Compensation
- The US base salary range for this full-time position is between $180,000 - $450,000 annually.
- The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
- The total compensation package may also include additional components and benefits depending on the specific role.
- This information will be shared if an employment offer is extended.
Benefits
- One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.
Company info
- We are looking for a Member of Technical Staff, Post-Training to lead the development of post-training strategies that define how our models acquire coding, computer use, and agentic capabilities at scale.
This listing is sourced directly from Hark's careers page and normalized into a canonical job model.