The Allen Institute for Artificial Intelligence
Young Investigator, FlexOlmo
Berkeley, CA
H1B sponsorship availableDetected 4 days ago
Machine LearningDeep LearningPyTorchNLPLLMsRoboticsScientific WritingResearchCollaborationMentoring
About the role
- Author scientific papers for publication in a high-profile conference or journal.
Responsibilities
- We develop foundational AI research and innovation to deliver real-world impact through large-scale open models, data, robotics, conservation, and beyond.
- Collaborate with and learn from team members across Ai2.
- Build open-source software for the research community.
- We value diversity - We seek to hire, support, and promote people from all genders, ethnicities, and all levels of experience regardless of age.
- Define and lead a high-impact research project.
- To perform this job successfully, the team member(s) will possess the skills, aptitudes, and abilities to perform each duty proficiently.
- The requirements listed in this document are the minimum levels of knowle
Requirements
- You have the opportunity to:
- Team members will be required to follow any other job-related instructions and to perform any other job-related duties requested by any person authorized to give instructions or assignments.
- Are within one year of completing their PhD, or already have a PhD, in Computer Science or similar field with research experience in machine learning, natural language processing, language and vision, or related areas.
- Outstanding individual contributor (IC) skills, especially with deep learning frameworks (e.g. PyTorch).
- An outstanding publication record at AI-related venues, such as NeurIPS, ICLR, ICML, COLM, ACL, EMNLP. We will specifically evaluate the quality of publications in terms of rigor and impact, not the quantity.
- Extensive research experience in areas such as large language models, training dynamics, scaling laws, and data curation. Experience with mixture-of-experts, long-context language models, and retrieval is preferred but not required.
- Located [or willing to relocate] in Berkeley, CA.
- Physical Demands and Work Environment:
- The physical demands described here are representative of those that must be met by a team member to successfully perform the essential functions of this position. Reasonable accommodations may be made to enable individuals with disabilities to perform the functions.
- Must be able to remain in a stationary position for long periods of time.
- The ability to communicate information and ideas so others will understand. Must be able to exchange accurate information in these situations.
- The ability to observe details at close range.
Nice to have
- Extensive research experience in areas such as large language models, training dynamics, scaling laws, and data curation.
- Experience with mixture-of-experts, long-context language models, and retrieval is preferred but not required.
Compensation
- Team members will be able to receive annual bonuses.
Benefits
- Team members will receive $125 per month to assist with commuting or internet expenses and will also receive $200 per month for fitness and wellbeing expenses.
- Team members will also receive up to ten sick days per year, up to seven personal days per year, up to 20 vacation days per year and twelve paid holidays throughout the calendar year.
- Some requirements may exclude individuals who pose a direct threat or significant risk to the health or safety of themselves or others.
Company info
- We design new architectures and training methods that help models use data more effectively-through improved training, inference-time conditioning, and retrieval-broadening the types of data they can leverage and ultimately enhancing performance. We also develop scientific methodologies for evaluating and understanding these systems. Our team produces high-impact research and expertly engineered open-source tools that accelerate NLP research worldwide.
- We lead the FlexOlmo project, whose first release in July 2025 focused on a new Mixture-of-Experts architecture. Looking ahead, we plan to pursue creative, groundbreaking research that delivers scientific insights and practical solutions for building architectures and training methods that unlock the use of large and diverse data sources.
- Your Next Challenge:
- Why FlexOlmo? We are building the foundation for research into the next generation of language models designed for flexible data use.
- We design new architectures and training methods that help models use data more effectively-through improved training, inference-time conditioning, and retrieval-broadening the types of data they can leverage and ultimately enhancing performance.
- We also develop scientific methodologies for evaluating and understanding these systems.
- Our team produces high-impact research and expertly engineered open-source tools that accelerate NLP research worldwide.
- We lead the FlexOlmo project, whose first release in July 2025 focused on a new Mixture-of-Experts architecture.
- Looking ahead, we plan to pursue creative, groundbreaking research that delivers scientific insights and practical solutions for building architectures and training methods that unlock the use of large and diverse data sources.
- Why FlexOlmo?
- We are building the foundation for research into the next generation of language models designed for flexible data use.
- FlexOlmo is a small, tightly knit team, giving you the unique opportunity to work closely with team members toward one high-impact project.
- We encourage open collaboration projects, even with researchers at external institutions. Team member will be based in Berkeley, with opportunities to engage actively with the University of California, Berkeley, and the BAIR lab.
- Our pay is competitive, and visa sponsorship is available.
Visa & Work Authorization
- Our pay is competitive, and visa sponsorship is available.
Apply directly at The Allen Institute for Artificial Intelligence →Create a free account for alerts like thisView The Allen Institute for Artificial Intelligence immigration profile
This listing is sourced directly from The Allen Institute for Artificial Intelligence's careers page and normalized into a canonical job model.