LlamaIndex

LlamaIndex

Member of Technical Staff, Applied Research

San Francisco · Staff+

Sponsorship not specifiedDetected 14 days ago
PythonMachine LearningData EngineeringResearchCommunicationWritingAdaptability

About the role

  • This role is ideal for someone who sits between applied research and strong engineering.
  • You should be excited by frontier AI work, but equally motivated by practical product impact.
  • This is not a pure research role where ideas stay in papers.

Responsibilities

  • Build data pipelines for data curation, synthetic data generation, labeling, and benchmark creation.
  • Evaluate base models and perform post-training or fine-tuning to hit specific performance targets.
  • Improve model accuracy, latency, and cost-effectiveness across real-world document workflows.
  • Design and maintain benchmarks to measure extraction quality, layout understanding, OCR performance, reasoning accuracy, and end-to-end system reliability.
  • Work with messy real-world documents, including PDFs, scanned documents, tables, charts, forms, and multi-page enterprise documents.
  • Collaborate with engineering to move successful research prototypes into production.

Requirements

  • Strong ML foundation, including hands-on experience benchmarking and training models.
  • Strong engineering judgment and ability to write clean, production-quality code.
  • Comfort working in a fast-paced startup environment with high ownership and limited structure.

Nice to have

  • Prior startup experience, especially at an early-stage or high-growth AI company.
  • Experience as a founder or early startup engineer.
  • Experience with synthetic data generation, post-training, fine-tuning, or benchmark design.
  • Familiarity with tools such as vLLM, Pydantic, uv, ruff, mypy, Claude Code, Cursor, or similar modern AI engineering workflows.
  • Experience with open-source AI infrastructure or developer tools.
  • Join a fast-growing startup with strong open-source adoption and commercial traction.
  • Work directly with technical founders and a highly ambitious engineering team.
  • Have real ownership over model quality, product capability, and technical direction.

Benefits

  • Build production systems at the frontier of vision-language models and document AI.
  • You will work on vision-language models, document processing, data curation, synthetic data generation, benchmarking, training, fine-tuning, and post-training.
  • Develop and train vision-language models for document processing and document understanding.
  • Stay close to the latest research in vision-language models, document AI, post-training, synthetic data, and agentic systems.
  • 3-7 years of experience in machine learning engineering, applied research, or research engineering.

Company info

  • make our document AI systems more accurate, faster, and more cost-effective in production.
  • You will be expected to prototype quickly, evaluate rigorously, and help turn promising approaches into production systems used by customers.
  • We are looking for an AI Research Engineer to join our document understanding team.
  • Work directly with customers when needed to translate product requirements into benchmarks, experiments, and model improvements.

This listing is sourced directly from LlamaIndex's careers page and normalized into a canonical job model.