Hark
Data Engineering Lead
San Jose · Full-time
Sponsorship not specified$170k-$450kDetected 83 days ago
Distributed SystemsMachine LearningSparkData EngineeringNLPAgentic AIResearch
About the role
- That means owning the full data engineering stack: ingestion, transformation, quality filtering, and delivery to training and evaluation systems.
- This is a high-ownership role on a small team.
Responsibilities
- Own the data infrastructure stack end-to-end: ingestion, transformation, deduplication, quality filtering, versioning, and delivery to model training and evaluation systems.
- Collaborate closely with model researchers and data collection leads to understand data requirements and translate them into reliable, auditable pipelines.
- Build tooling and frameworks that make it easy for the team to inspect, evaluate, and iterate on data quality. The insights surfaced should feed back into collection and curation decisions.
- Design data systems for reproducibility and scale. The pipelines you build need to handle growing volumes across modalities without becoming a bottleneck.
- Identify gaps in the current stack and drive concrete improvements to throughput, quality, and reliability.
- The models we ship are only as good as the data behind them, and this role owns that foundation.
- You'll work directly with model researchers, data collection leads, and infrastructure engineers, and the systems you build will directly shape the quality and pace of model development.
- You'll build the data infrastructure that turns raw signals into the training data Hark's models learn from, and the pipelines that keep it flowing at scale.
- This is a high-ownership role on a small team. You'll work directly with model researchers, data collection leads, and infrastructure engineers, and the systems you build will directly shape the quality and pace of model development.
- Hark is an artificial intelligence company building advanced, personalized intelligence.
Requirements
- You are comfortable designing and operating large-scale batch and streaming pipelines, and you care about correctness and reliability.
- You understand the difference between a data pipeline for analytics and one that feeds model training, and you know what it takes to get the latter right.
- You can work closely with model researchers and engineers, explain data tradeoffs clearly, and make good decisions across team boundaries.
Nice to have
- Familiarity with multimodal data formats and processing pipelines (audio, video, image).
- Experience with human feedback or preference data pipelines (RLHF, DPO, or similar).
- Hands-on experience with data quality evaluation frameworks or annotation tooling.
- Background in distributed systems, stream processing, or large-scale ETL.
- Experience at a fast-moving AI lab or research-driven company.
- The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
- This information will be shared if an employment offer is extended.
- Fluency with the modern data stack.
Compensation
- The US base salary range for this full-time position is between $170,000 - $450,000 annually.
- The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
- The total compensation package may also include additional components/benefits depending on the specific role.
- This information will be shared if an employment offer is extended.
Benefits
- One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.
- Design and build scalable data pipelines that ingest, process, and deliver training data across multiple modalities: text, audio, vision, and structured feedback signals.
- Instrument pipelines for correctness, freshness, and coverage.
This listing is sourced directly from Hark's careers page and normalized into a canonical job model.