Ifm Us

Ifm Us

Research Scientist - Vision Language Model

Sunnyvale, CA

Sponsorship not specified$150k-$450kDetected 53 days ago
PythonAlgorithmsMachine LearningDeep LearningPyTorchNLPComputer VisionLLMsAgentic AIResearchCommunicationCollaborationProblem Solving

About the role

  • As a Research Scientist in the Vision Language Model (VLM) team, your role will be central to advancing state-of-the-art multimodal foundation models that integrate visual understanding, reasoning, and agentic capabilities.
  • You will work on the research and development of large-scale VLM systems, spanning model architectures, data recipes for pre-training and post-training, and evaluation benchmarks.

Responsibilities

  • Develop novel architectures and training methodologies for integrating visual understanding, language reasoning, and tool-use capabilities.
  • Build and improve large-scale multimodal datasets, synthetic data generation pipelines, and evaluation benchmarks for VLM capabilities.
  • Investigate multimodal reasoning, agentic behavior, OCR, grounding, document understanding, chart understanding, and visual question answering capabilities.
  • Mentor junior researchers and collaborate across teams to drive impactful research initiatives.
  • Problem-solving and research skills with the ability to independently drive research/engineering projects.
  • Research experience in multimodal reasoning, agentic systems, tool use, OCR, grounding, document understanding, or multimodal retrieval.
  • Experience collaborating across research, infrastructure, and product-oriented teams to deliver state-of-the-art multimodal systems.
  • We are a dedicated research lab for building, understanding, using, and risk-managing foundation models.
  • Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.

Requirements

  • Experience with distributed training systems and large-scale model optimization.
  • Experience with ML infrastructure, including model evaluation, debugging, optimization, and large-scale experimentation.
  • Experience with synthetic data generation, multimodal data curation, or automated evaluation frameworks for VLMs.

Nice to have

  • Preferred Skills

Compensation

  • $150k-$450k

Benefits

  • Research and development of next-generation Vision Language Models across pre-training, instruction tuning, reasoning, and agents.
  • Research efficient multimodal learning techniques, including data-efficient training, long-context modeling, model modularity, and inference optimization.
  • PhD or equivalent research experience in Machine Learning, Computer Vision, Natural Language Processing, or Multimodal AI.
  • Experience working with large language models and/or vision-language models, including pre-training, fine-tuning, evaluation, or inference.
  • Strong Python and PyTorch development skills for large-scale machine learning research.
  • Understanding of modern deep learning architectures, including Transformers, attention mechanisms, and multimodal fusion techniques.
  • Hands-on experience training or fine-tuning large Vision Language Models or multimodal foundation models at scale.
  • Experience with distributed learning frameworks and infrastructure such as PyTorch Distributed, Megatron, Triton, or CUDA.

This listing is sourced directly from Ifm Us's careers page and normalized into a canonical job model.