Pathwaycom

Pathwaycom

AI Benchmark & Datasets Engineer / Researcher

Palo Alto, California, United States · Full-time

Sponsorship not specifiedDetected 125 days ago
GitMachine LearningData ScienceNLPLLMsResearch

About the role

  • You Will Proactively identify, prioritize, and curate relevant public and client-driven benchmarks across our target use cases and markets.
  • Evaluate candidate benchmarks for clarity, data quality, evaluation methodology, and fit with our model roadmap.
  • Track and organize benchmark results, model leaderboards, and "what good looks like" for different customers and scenarios.

Responsibilities

  • About Pathway Pathway builds the first post-transformer frontier model that solves AI's fundamental memory problem.
  • We're not optimizing yesterday's technology; we're building what comes after transformers.
  • The Opportunity You will design and execute rigorous benchmarks and define dataset standards.
  • Collaborating closely with our R&D team, you will build the evaluation infrastructure that guides the evolution of Pathway's post‑transformer models.
  • Run benchmarks with baseline models to validate setup, uncover edge cases, and de‑risk R&D runs.
  • You have published at least one paper at NeurIPS, ICLR, or ICML - where you were the lead author or made significant conceptual & code contributions.

Requirements

  • You have significantly contributed to an LLM training effort which became newsworthy (topped a Hugging Face benchmark, best in class model, etc.), preferably using multiple GPU's.

Nice to have

  • You Have experience with ML/LLM evaluation, data science, or technical product roles, ideally around benchmarks or experimentation.

Benefits

  • You have spent at least 6 months working in a leading Machine Learning research center (e.g. at: Google Brain / Deepmind, Apple, Meta, Anthropic, Nvidia, MILA).
  • Care about high‑quality data, reproducible experiments, and crisp documentation Are respectful of others Are fluent in English Bonus Points Published or open‑sourced work on LLM evaluation, benchmarking or data quality.
  • While transformers wake up in the same state every time-like Groundhog Day-our architecture enables true continuous learning, infinite context reasoning, and real-time adaptation.

Company info

  • Can talk comfortably with both engineers and customers, and translate between technical detail and business value.
  • We are trusted by organizations such as NATO, La Poste, and Formula 1 racing teams.

This listing is sourced directly from Pathwaycom's careers page and normalized into a canonical job model.