LlamaIndex
Member of Technical Staff, Applied Research
San Francisco · Staff+
Sponsorship not specifiedDetected 14 days ago
PythonMachine LearningData EngineeringResearchCommunicationWritingAdaptability
About the role
- This role is ideal for someone who sits between applied research and strong engineering.
- You should be excited by frontier AI work, but equally motivated by practical product impact.
- This is not a pure research role where ideas stay in papers.
Responsibilities
- Build data pipelines for data curation, synthetic data generation, labeling, and benchmark creation.
- Evaluate base models and perform post-training or fine-tuning to hit specific performance targets.
- Improve model accuracy, latency, and cost-effectiveness across real-world document workflows.
- Design and maintain benchmarks to measure extraction quality, layout understanding, OCR performance, reasoning accuracy, and end-to-end system reliability.
- Work with messy real-world documents, including PDFs, scanned documents, tables, charts, forms, and multi-page enterprise documents.
- Collaborate with engineering to move successful research prototypes into production.
Requirements
- Strong ML foundation, including hands-on experience benchmarking and training models.
- Strong engineering judgment and ability to write clean, production-quality code.
- Comfort working in a fast-paced startup environment with high ownership and limited structure.
Nice to have
- Prior startup experience, especially at an early-stage or high-growth AI company.
- Experience as a founder or early startup engineer.
- Experience with synthetic data generation, post-training, fine-tuning, or benchmark design.
- Familiarity with tools such as vLLM, Pydantic, uv, ruff, mypy, Claude Code, Cursor, or similar modern AI engineering workflows.
- Experience with open-source AI infrastructure or developer tools.
- Join a fast-growing startup with strong open-source adoption and commercial traction.
- Work directly with technical founders and a highly ambitious engineering team.
- Have real ownership over model quality, product capability, and technical direction.
Benefits
- Build production systems at the frontier of vision-language models and document AI.
- You will work on vision-language models, document processing, data curation, synthetic data generation, benchmarking, training, fine-tuning, and post-training.
- Develop and train vision-language models for document processing and document understanding.
- Stay close to the latest research in vision-language models, document AI, post-training, synthetic data, and agentic systems.
- 3-7 years of experience in machine learning engineering, applied research, or research engineering.
Company info
- make our document AI systems more accurate, faster, and more cost-effective in production.
- You will be expected to prototype quickly, evaluate rigorously, and help turn promising approaches into production systems used by customers.
- We are looking for an AI Research Engineer to join our document understanding team.
- Work directly with customers when needed to translate product requirements into benchmarks, experiments, and model improvements.
Apply directly at LlamaIndex →Create a free account for alerts like thisView LlamaIndex immigration profile
This listing is sourced directly from LlamaIndex's careers page and normalized into a canonical job model.