Labelbox

Labelbox

Forward Deployed Research Scientist

San Francisco Bay Area · Staff+

Sponsorship not specified$25k-$30kDetected 67 days ago
Machine LearningData EngineeringNLPLLMsStatisticsControlsResearchLab ResearchCollaboration

About the role

  • This is not a traditional research scientist role.
  • You will not spend months pursuing a single research question.
  • You will work on multiple client engagements simultaneously, operating on timescales of days to weeks.

Responsibilities

  • You will be in the room during client scoping meetings - not as support staff, but as a technical peer.
  • Develop deep scientific understanding of client engagements.
  • For each project, you will build a working model of the client's architecture, training methodology, and target capabilities.
  • You will partner with our Human Data Operations team to review annotation schemas, task designs, and quality rubrics before projects go into execution.
  • Collaborate with Applied Research on publications and benchmarks.
  • Our Applied Research team owns the long-horizon research agenda.
  • Innovation at Speed: We celebrate those who take ownership, move fast, and deliver impact.
  • We empower people to drive results through clear ownership and metrics.
  • The measure of success here is client impact and publishable-but-practical results - not methodological novelty for its own sake.
  • If your first instinct when handed a problem is to build a framework, this isn't the role.

Requirements

  • Strong understanding of LLM training pipelines - pretraining, supervised fine-tuning, RLHF/DPO, and how data quality and composition affect each stage.
  • Experience designing and executing experiments with rigor - hypothesis formation, controlled comparisons, statistical analysis of results.
  • Ability to operate at speed.
  • Prior experience at a frontier AI lab, applied ML startup, or in a research role with direct client/stakeholder interaction.
  • Experience with evaluation and benchmarking of LLMs - designing metrics, building eval harnesses, interpreting results critically.
  • Familiarity with human data pipelines - annotation workflows, quality assurance methodology, inter-annotator agreement analysis.
  • Comfort with ambiguity and incomplete information.
  • Required
  • MS or PhD in Machine Learning, NLP, Computer Science, or a related quantitative field.
  • Hands-on experience fine-tuning large language models (open-weight models such as Llama, Mistral, Qwen, or similar).
  • Ability to operate at speed. You should be comfortable going from problem definition to experimental results in days, not months.
  • Strong written and verbal communication. You will present findings to client research teams and contribute to published work.
  • Experience with reinforcement learning, reward modeling, or RLHF environments.
  • Published research (conferences, journals, or technical reports) in ML/NLP or adjacent fields.
  • What Matters More Than Credentials

Nice to have

  • Strongly Preferred

Skills

  • Advanced annotation tools, workflow automation, and quality control systems that enable teams to produce high-quality training data at scale

Compensation

  • Labelbox strives to ensure pay parity across the organization and discuss compensation transparently.

Benefits

  • Continuous Growth: Every role requires continuous learning and evolution.

Company info

  • Shape the Future of AI
  • At Labelbox, we're building the critical infrastructure that powers breakthrough AI models at leading research labs and enterprises.
  • Since 2018, we've been pioneering data-centric approaches that are fundamental to AI development, and our work becomes even more essential as AI capabilities expand exponentially.
  • About Labelbox
  • We're the only company offering three integrated solutions for frontier
  • We are looking for someone who finds that energizing, not compromising.
  • You'll engage on methodology, challenge assumptions about data requirements, and shape project specifications based on a scientific understanding of how data composition affects model outcomes.
  • This is how we validate that what we deliver actually improves our customers' models - and how we catch problems before the client does.
  • We are small and high-leverage.
  • We are at the intersection of several teams.

This listing is sourced directly from Labelbox's careers page and normalized into a canonical job model.