Tensorstax Com

Tensorstax Com

Research Engineer Intern, Evaluations

San Francisco, USA · Intern · Internship

Sponsorship not specifiedDetected 491 days ago
PythonBigQuerySnowflakeKafkaMachine LearningPyTorchSparkdbtData EngineeringNLPLLMsAgentic AIResearchAdaptability

About the role

  • This role is ideal for candidates passionate about AI evaluations, language model benchmarking, and autonomous data systems.
  • Strong background in AI evaluation methodologies, reinforcement learning, and RLHF techniques.

Responsibilities

  • Develop evaluation environments to test AI agents' ability to reason, plan, and act autonomously within mission-critical data pipelines.
  • Design benchmarks to assess model capabilities in failure detection, pipeline optimization, and agentic decision-making in data workflows.
  • Implement automated assessment frameworks for language model-based agents operating over data lakes and warehouses.
  • Work with synthetic and real-world datasets to create robust testing environments for AI-driven data automation.
  • Collaborate with research engineers to refine reward shaping strategies, guiding models toward more efficient and agentic behaviors in data-intensive tasks.

Requirements

  • Experience in language model research, with a focus on benchmarking LLMs in mission-critical domains.
  • Familiarity with benchmarking language models for structured and unstructured data tasks.
  • Proficiency in Python and experience with ML frameworks like PyTorch or JAX.
  • Hands-on experience with data lakes, warehouses, and data engineering tools (Snowflake, BigQuery, dbt, Spark, Kafka).

Nice to have

  • Contributions to open-source AI benchmarks (e.g., SweBench, BIRD, SPIDER).
  • Contributions to open-source agentic frameworks.
  • Strong understanding of ETL, ELT, and data transformation pipelines.

Benefits

  • Competitive internship stipend.
  • 100% employer-covered health, dental, and vision insurance (for eligible interns).

This listing is sourced directly from Tensorstax Com's careers page and normalized into a canonical job model.