Hark
Member of Technical Staff, Multimodal Vision
San Jose · Staff+ · Full-time
Sponsorship not specified$180k-$450kDetected 83 days ago
Full-Stack DevelopmentMachine LearningData EngineeringAgentic AIA/B TestingResearch
About the role
- This includes working across the full stack-from data and modeling to training, serving, and product integration.
Responsibilities
- Build evaluation frameworks and internal benchmarks to measure model performance, robustness, and visual quality across tasks.
- Optimize models and systems for scalability, efficiency, and real-time or production deployment.
- Collaborate closely with product and engineering teams to translate research innovations into impactful, user-facing AI experiences.
- You will contribute to both pretraining and posttraining efforts while collaborating closely with product teams to push the boundaries of model capability and deliver exceptional end-to-end user experiences.
- Hark is an artificial intelligence company building advanced, personalized intelligence.
- We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines.
Nice to have
- The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
- This information will be shared if an employment offer is extended.
Compensation
- The US base salary range for this full-time position is between $180,000 - $450,000 annually.
- The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
- The total compensation package may also include additional components/benefits depending on the specific role.
- This information will be shared if an employment offer is extended.
Benefits
- One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.
- The Omni team at Hark is building the next generation of AI experiences beyond text, enabling models to understand and generate content across multiple modalities, including text, and vision.
- Drive research and development to advance vision and video capabilities in multimodal models, including image understanding, video modeling, and generative vision systems.
- Develop and improve large-scale vision and video data pipelines, including data collection, filtering, labeling, and synthetic data generation.
- Design and implement state-of-the-art models for vision and video, including multimodal architectures that integrate vision with text and other modalities.
Company info
- As part of the Omni team, you will help drive the development of text, video, and multimodal models.
This listing is sourced directly from Hark's careers page and normalized into a canonical job model.