HeyGen

HeyGen

Software Engineer, Data

Los Angeles, Palo Alto, San Francisco

Sponsorship not specified$180k-$220kDetected 166 days ago
PythonGoData StructuresSQLSnowflakeDatabricksAWSGCPCloud PlatformsKafkaMachine LearningSparkData EngineeringComputer VisionLLMs

About the role

  • Productize Data: Transform raw data into structured, actionable data products that can be easily consumed by front-end applications, API endpoints, and AI agents.

Responsibilities

  • Design, develop, and maintain robust batch and real-time data pipelines (using Python, Go, Spark, Kafka) that ingest and transform massive multi-modal data-text, audio, and video-to train and run AI models.
  • Collaborate with ML engineers to implement data structures and APIs for new, exciting features like PPT-to-video automation and interactive AI avatars that require low-latency data fetching.
  • Architect and manage data lakehouse solutions (e.g., Snowflake, Databricks, Apache Iceberg) to store and query unstructured media data efficiently, enhancing storage and computation efficiency.
  • Implement data quality checks, data contracts, and monitoring to ensure high reliability of data, preventing downtime in production video generation.
  • You will help build the data foundational layers for our next-generation features.
  • Build & Scale Data Pipelines: Design, develop, and maintain robust batch and real-time data pipelines (using Python, Go, Spark, Kafka) that ingest and transform massive multi-modal data-text, audio, and video-to train and run AI models.
  • Power Intelligent Features: Collaborate with ML engineers to implement data structures and APIs for new, exciting features like PPT-to-video automation and interactive AI avatars that require low-latency data fetching.
  • Data Lakehouse Infrastructure: Architect and manage data lakehouse solutions (e.g., Snowflake, Databricks, Apache Iceberg) to store and query unstructured media data efficiently, enhancing storage and computation efficiency.
  • Data Reliability & Observability: Implement data quality checks, data contracts, and monitoring to ensure high reliability of data, preventing downtime in production video generation.

Requirements

  • Bachelor's/Master's degree in Computer Science, Engineering, or a related field.
  • 3-5+ years of experience as a Backend Software Engineer with heavy data processing responsibilities.
  • Strong proficiency in Python (for ETL/scripting) and SQL (for data modeling).
  • Experience with cloud platforms (AWS/GCP) and data technologies like Kafka, Spark, and Snowflake/Databricks.
  • ability to operate in a fast-paced, startup environment.
  • Experience or interest in Computer Vision/Generative AI data processing.
  • Proactive, "owner" mindset
  • What HeyGen Offers
  • Competitive salary and benefits package.
  • Dynamic and inclusive work environment focused on innovation and creativity.
  • Opportunities for professional growth and skill development.
  • Collaborative culture that values teamwork and employee input.
  • Access to state-of-the-art technologies and tools.

Nice to have

  • Over the last decade, visual content has become the preferred method of information creation, consumption, and retention.
  • Learn more at www.heygen.com.
  • Position Summary
  • A Software Engineer with data engineering responsibilities to bridge the gap between core application development and large-scale data infrastructure.
  • This team is currently working on cutting-edge features including PPT-to-video converters and interactive, conversational video capabilities.
  • Core Responsibilities

Compensation

  • $180,000 - $220,000 + equity + benefits
  • Please note that the salary information is a general guideline only.

Equal opportunity

  • HeyGen is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
  • HeyGen is an Equal Opportunity Employer.

This listing is sourced directly from HeyGen's careers page and normalized into a canonical job model.