Hark

Hark

Member of Technical Staff, Post-training

San Jose · Staff+ · Full-time

Sponsorship not specified$180k-$450kDetected 83 days ago
PythonAlgorithmsMachine LearningPyTorchLLMsAgentic AIRoboticsResearch

About the role

  • This role sits at the frontier of a rapidly emerging discipline - one where reinforcement learning, simulation, and large-scale model training converge to produce agents that can reason, plan, and act over long horizons.

Responsibilities

  • Design and implement post-training strategies, primarily RL-based, to develop strong coding agents capable of multi-step reasoning, tool use, and long-horizon task completion.
  • Build and scale simulation and scaffolding environments for agentic RL: code execution sandboxes, computer use environments, tool-calling harnesses, and verifiable reward signals.
  • Develop reward modeling pipelines - including outcome-based, execution-based, and process-based reward signals - and iterate on them based on training dynamics.
  • Design and run rigorous ablations to understand how algorithm choice, data mixture, reward shaping, and scale interact in the agentic setting.
  • Build evaluation frameworks grounded in real agent tasks - code correctness, execution success, multi-step tool use - to measure progress and guide iteration.
  • Collaborate with mid-training, infrastructure, and product teams to translate research insights into durable improvements on the model.
  • Hark is an artificial intelligence company building advanced, personalized intelligence.
  • We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines.
  • Experience building or working within simulation or execution environments (e.g., code interpreters, sandboxed execution, game environments, robotics simulators).
  • Proven ability to design and execute rigorous experiments, with strong intuition for diagnosing training failures and scaling bottlenecks.

Requirements

  • Proficiency in Python and PyTorch
  • comfort working across research and systems code.
  • Ability to work in a fast-moving, research-forward environment where the right approach is often unknown at the outset.
  • Proficiency in Python and PyTorch; comfort working across research and systems code.

Nice to have

  • Experience with RL algorithms applied to language or code: RLHF, DPO, GRPO, PPO, or similar paradigms in the LLM setting.
  • Familiarity with coding agent benchmarks and evaluation environments (e.g., SWE-bench, HumanEval, LiveCodeBench, competitive programming judges).
  • Background in reward modeling - outcome-based, process-based, or learned reward signals.
  • Prior work on computer use, GUI agents, or tool-using LLMs (e.g., OSWorld, WebArena-style tasks).
  • Experience training or scaling models at 10B+ parameters, with attention to efficiency, stability, and GPU utilization.
  • Contributions to open-source ML projects or publications at top venues (NeurIPS, ICML, ICLR, EMNLP, COLM, etc.).
  • The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
  • This information will be shared if an employment offer is extended.

Skills

  • Scale synthetic data generation and trajectory distillation pipelines that feed RL training and improve sample efficiency.

Compensation

  • The US base salary range for this full-time position is between $180,000 - $450,000 annually.
  • The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
  • The total compensation package may also include additional components and benefits depending on the specific role.
  • This information will be shared if an employment offer is extended.

Benefits

  • One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.

Company info

  • We are looking for a Member of Technical Staff, Post-Training to lead the development of post-training strategies that define how our models acquire coding, computer use, and agentic capabilities at scale.

This listing is sourced directly from Hark's careers page and normalized into a canonical job model.