Sesame
Data Engineer, Machine Learning
San Francisco · Contract
Sponsorship not specifiedDetected 29 days ago
PythonSQLVector DatabasesKubernetesMachine LearningAirflowData EngineeringAI OrchestrationCompensationEmbedded SystemsCommunication
About the role
- Sesame's data is rich and complex: conversations, voice, sensor signals, and product telemetry.
- This is a deeply technical, infrastructure-focused role - closer to ML engineering than traditional data analytics.
Responsibilities
- Design and build production data pipelines that prepare conversational, voice, and multimodal data for model training and evaluation.
- Partner directly with ML engineers to understand data requirements for new models and experiments, and deliver datasets that meet those needs.
- Build and maintain infrastructure for dataset versioning, lineage tracking, and reproducibility - so any training run can be traced back to its exact data.
- Build tooling that makes it easy for ML engineers and researchers to discover, explore, and request data independently.
- You'll design the systems that take raw, unstructured, multimodal data and turn it into clean, versioned, well-documented datasets that ML teams can trust and build on confidently.
- You'll be deeply embedded with ML teams, understanding their workflows and building infrastructure that accelerates the full model development lifecycle - from data collection and labeling through training and evaluation.
- Sesame believes in a future where computers are lifelike - with the ability to see, hear, and collaborate with us in ways that feel natural and human.
Requirements
- 5+ years in data engineering, with meaningful experience supporting ML or AI teams specifically.
- Experience with workflow orchestration systems such as Airflow, Dagster, or Prefect.
- Hands-on experience with ML data workflows: training data pipelines, dataset versioning, data labeling pipelines, or model evaluation data.
- Comfort working with unstructured and semi-structured data - audio, text, JSON logs - not just clean relational tables.
Nice to have
- Vector databases, embedding storage, or feature stores.
- Data from hardware or embedded systems: telemetry, sensors, real-time streams.
- Sesame is committed to a workplace where everyone feels valued, respected, and empowered.
- We welcome all qualified applicants, embracing diversity in race, gender, identity, orientation, ability, and more.
- We provide reasonable accommodations for applicants with disabilities.
- Contact careers@sesame.com for assistance.
- 401 (k) max employer match: 3.5% of compensation
Skills
- Distributed compute frameworks for large-scale data processing such as Ray or Spark.
- Kubernetes and managed Kubernetes environments such as GKE or EKS.
- Data privacy frameworks, especially around voice or conversational data.
- Building internal tooling or self-serve data platforms.
Compensation
- 401 (k) max employer match: 3.5% of compensation
Benefits
- 100% employer-paid health, vision, and dental benefits for you and your dependents
- Unlimited PTO and sick time
- Flexible spending account with employer matching up to $1,650/year (medical FSA)
- Opportunity to share in the company's success with competitive stock options
- Benefits do not apply to contingent/contract workers.
- With this vision, we're designing a new kind of computer, focused on making voice agents part of our daily lives.
- Develop data quality frameworks that catch issues before they become model quality issues: schema validation, drift detection, and coverage monitoring.
- You'll collaborate directly with machine learning engineers and researchers - your job is to make sure they have the right data, in the right shape, at the right time to train, evaluate, and ship models.
Company info
- We're looking for a Data Engineer to build and maintain the data pipelines that feed Sesame's AI models.
This listing is sourced directly from Sesame's careers page and normalized into a canonical job model.