Sanas
Staff+ Data Engineer (ML Infrastructure)
Palo Alto, CA · Staff+
Sponsorship not specifiedDetected 16 days ago
SnowflakeDatabricksAWSMachine LearningSparkAirflowData EngineeringMLOpsDesign SystemsRoadmappingSystems EngineeringResearchLeadershipCommunication
About the role
- Our models are only as good as the data that trains them.
- Delta Lake), partitioning strategies, metadata management, and schema evolution - with a bias toward reproducibility and auditability.
- Platform reliability & governance Instrument pipelines with observability, data quality checks, lineage tracking, and alerting - so failures surface fast and root causes are traceable.
Responsibilities
- As a Staff Data Engineer, you'll own the infrastructure that takes raw audio - millions of hours across accents, languages, noise conditions, and recording environments - and turns it into clean, reproducible, training-ready data at scale.
- You'll work directly with AI research scientists and ML engineers to design systems that move fast without breaking the data quality guarantees our models depend on.
- Job Description Data pipeline & lakehouse architecture Design and implement large-scale data pipelines that ingest, transform, validate, and serve high-quality audio and metadata for AI model training, evaluation, and product telemetry.
- Own the lakehouse architecture - table format choices (Iceberg vs.
- Build and maintain batch and streaming pipelines using Spark, Flink, and orchestration tooling (Airflow or Dagster), with a clear-eyed view of when each is the right tool.
- Extend and maintain feature store infrastructure to serve low-latency, versioned features for both training and real-time inference.
- Audio data at scale Develop and maintain pipelines purpose-built for the unique challenges of audio data: large file volumes, time-series feature extraction, speaker and language metadata, and annotation versioning.
- Build tooling that supports the full audio data lifecycle - from raw ingestion and quality filtering through augmentation, segmentation, and training split generation - with reproducibility guarantees at every stage.
- Partner with ML engineers and research scientists to design data schemas, sampling strategies, and evaluation datasets that accurately reflect production conditions.
- Own data pipelines that feed human-in-the-loop annotation workflows - ensuring clean round-trips between raw data, labeling platforms, and training-ready outputs.
Requirements
- Experience with MLOps tooling - experiment tracking, dataset versioning (DVC, LakeFS), and training pipeline orchestration.
Skills
- Strong command of data processing frameworks: Spark, Flink, and Ray; and orchestrators: Airflow or Dagster.
Benefits
- Bonus Direct experience with audio data pipelines - file handling at scale, time-series features, speaker metadata, or audio annotation tooling.
This listing is sourced directly from Sanas's careers page and normalized into a canonical job model.