Tensorstax Com

Tensorstax Com

Research Engineer, Reinforcement Learning

San Francisco, USA

Sponsorship not specifiedDetected 492 days ago
BigQuerySnowflakeRedshiftMachine LearningPyTorchData EngineeringNLPLLMsResearchProblem Solving

About the role

  • Fine-tune language models using reinforcement learning techniques such as PPO, DPO, and KTO.
  • Stay at the forefront of research on RL for language models, incorporating advancements like GRPO, SWE-Gym, and SWE-RL into practical applications.
  • Strong familiarity with LLM fine-tuning techniques (PPO, DPO, KTO) and their applications in reinforcement learning.

Responsibilities

  • Develop and refine reward functions to optimize agent behavior for complex data engineering tasks.
  • Curate and build high-quality datasets for supervised fine-tuning (SFT) and RLHF.
  • Design experiments to evaluate and improve the agentic capabilities of language models in data environments.

Requirements

  • Experience curating and constructing high-quality datasets for fine-tuning.

Nice to have

  • Experience with distributed training in PyTorch (DDP, FSDP).
  • Hands-on experience designing RL environments for traditional RL problems.
  • Contributions to open-source projects in RL, LLMs, or ML infrastructure.
  • Familiarity with data lakes and warehouses (Snowflake, BigQuery, Redshift).

Benefits

  • 100% employer-covered health, dental, and vision insurance.
  • Create RL gym environments for language model agents.

This listing is sourced directly from Tensorstax Com's careers page and normalized into a canonical job model.