Ifm Us
Research Scientist - Vision Language Model
Sunnyvale, CA
Sponsorship not specified$150k-$450kDetected 53 days ago
PythonAlgorithmsMachine LearningDeep LearningPyTorchNLPComputer VisionLLMsAgentic AIResearchCommunicationCollaborationProblem Solving
About the role
- As a Research Scientist in the Vision Language Model (VLM) team, your role will be central to advancing state-of-the-art multimodal foundation models that integrate visual understanding, reasoning, and agentic capabilities.
- You will work on the research and development of large-scale VLM systems, spanning model architectures, data recipes for pre-training and post-training, and evaluation benchmarks.
Responsibilities
- Develop novel architectures and training methodologies for integrating visual understanding, language reasoning, and tool-use capabilities.
- Build and improve large-scale multimodal datasets, synthetic data generation pipelines, and evaluation benchmarks for VLM capabilities.
- Investigate multimodal reasoning, agentic behavior, OCR, grounding, document understanding, chart understanding, and visual question answering capabilities.
- Mentor junior researchers and collaborate across teams to drive impactful research initiatives.
- Problem-solving and research skills with the ability to independently drive research/engineering projects.
- Research experience in multimodal reasoning, agentic systems, tool use, OCR, grounding, document understanding, or multimodal retrieval.
- Experience collaborating across research, infrastructure, and product-oriented teams to deliver state-of-the-art multimodal systems.
- We are a dedicated research lab for building, understanding, using, and risk-managing foundation models.
- Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.
Requirements
- Experience with distributed training systems and large-scale model optimization.
- Experience with ML infrastructure, including model evaluation, debugging, optimization, and large-scale experimentation.
- Experience with synthetic data generation, multimodal data curation, or automated evaluation frameworks for VLMs.
Nice to have
- Preferred Skills
Compensation
- $150k-$450k
Benefits
- Research and development of next-generation Vision Language Models across pre-training, instruction tuning, reasoning, and agents.
- Research efficient multimodal learning techniques, including data-efficient training, long-context modeling, model modularity, and inference optimization.
- PhD or equivalent research experience in Machine Learning, Computer Vision, Natural Language Processing, or Multimodal AI.
- Experience working with large language models and/or vision-language models, including pre-training, fine-tuning, evaluation, or inference.
- Strong Python and PyTorch development skills for large-scale machine learning research.
- Understanding of modern deep learning architectures, including Transformers, attention mechanisms, and multimodal fusion techniques.
- Hands-on experience training or fine-tuning large Vision Language Models or multimodal foundation models at scale.
- Experience with distributed learning frameworks and infrastructure such as PyTorch Distributed, Megatron, Triton, or CUDA.
This listing is sourced directly from Ifm Us's careers page and normalized into a canonical job model.