Hark

Hark

Member of Technical Staff, Multimodal Vision

San Jose · Staff+ · Full-time

Sponsorship not specified$180k-$450kDetected 83 days ago
Full-Stack DevelopmentMachine LearningData EngineeringAgentic AIA/B TestingResearch

About the role

  • This includes working across the full stack-from data and modeling to training, serving, and product integration.

Responsibilities

  • Build evaluation frameworks and internal benchmarks to measure model performance, robustness, and visual quality across tasks.
  • Optimize models and systems for scalability, efficiency, and real-time or production deployment.
  • Collaborate closely with product and engineering teams to translate research innovations into impactful, user-facing AI experiences.
  • You will contribute to both pretraining and posttraining efforts while collaborating closely with product teams to push the boundaries of model capability and deliver exceptional end-to-end user experiences.
  • Hark is an artificial intelligence company building advanced, personalized intelligence.
  • We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines.

Nice to have

  • The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
  • This information will be shared if an employment offer is extended.

Compensation

  • The US base salary range for this full-time position is between $180,000 - $450,000 annually.
  • The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
  • The total compensation package may also include additional components/benefits depending on the specific role.
  • This information will be shared if an employment offer is extended.

Benefits

  • One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.
  • The Omni team at Hark is building the next generation of AI experiences beyond text, enabling models to understand and generate content across multiple modalities, including text, and vision.
  • Drive research and development to advance vision and video capabilities in multimodal models, including image understanding, video modeling, and generative vision systems.
  • Develop and improve large-scale vision and video data pipelines, including data collection, filtering, labeling, and synthetic data generation.
  • Design and implement state-of-the-art models for vision and video, including multimodal architectures that integrate vision with text and other modalities.

Company info

  • As part of the Omni team, you will help drive the development of text, video, and multimodal models.

This listing is sourced directly from Hark's careers page and normalized into a canonical job model.