HeyGen
Software Engineer, Data
Los Angeles, Palo Alto, San Francisco
Sponsorship not specified$180k-$220kDetected 166 days ago
PythonGoData StructuresSQLSnowflakeDatabricksAWSGCPCloud PlatformsKafkaMachine LearningSparkData EngineeringComputer VisionLLMs
About the role
- Productize Data: Transform raw data into structured, actionable data products that can be easily consumed by front-end applications, API endpoints, and AI agents.
Responsibilities
- Design, develop, and maintain robust batch and real-time data pipelines (using Python, Go, Spark, Kafka) that ingest and transform massive multi-modal data-text, audio, and video-to train and run AI models.
- Collaborate with ML engineers to implement data structures and APIs for new, exciting features like PPT-to-video automation and interactive AI avatars that require low-latency data fetching.
- Architect and manage data lakehouse solutions (e.g., Snowflake, Databricks, Apache Iceberg) to store and query unstructured media data efficiently, enhancing storage and computation efficiency.
- Implement data quality checks, data contracts, and monitoring to ensure high reliability of data, preventing downtime in production video generation.
- You will help build the data foundational layers for our next-generation features.
- Build & Scale Data Pipelines: Design, develop, and maintain robust batch and real-time data pipelines (using Python, Go, Spark, Kafka) that ingest and transform massive multi-modal data-text, audio, and video-to train and run AI models.
- Power Intelligent Features: Collaborate with ML engineers to implement data structures and APIs for new, exciting features like PPT-to-video automation and interactive AI avatars that require low-latency data fetching.
- Data Lakehouse Infrastructure: Architect and manage data lakehouse solutions (e.g., Snowflake, Databricks, Apache Iceberg) to store and query unstructured media data efficiently, enhancing storage and computation efficiency.
- Data Reliability & Observability: Implement data quality checks, data contracts, and monitoring to ensure high reliability of data, preventing downtime in production video generation.
Requirements
- Bachelor's/Master's degree in Computer Science, Engineering, or a related field.
- 3-5+ years of experience as a Backend Software Engineer with heavy data processing responsibilities.
- Strong proficiency in Python (for ETL/scripting) and SQL (for data modeling).
- Experience with cloud platforms (AWS/GCP) and data technologies like Kafka, Spark, and Snowflake/Databricks.
- ability to operate in a fast-paced, startup environment.
- Experience or interest in Computer Vision/Generative AI data processing.
- Proactive, "owner" mindset
- What HeyGen Offers
- Competitive salary and benefits package.
- Dynamic and inclusive work environment focused on innovation and creativity.
- Opportunities for professional growth and skill development.
- Collaborative culture that values teamwork and employee input.
- Access to state-of-the-art technologies and tools.
Nice to have
- Over the last decade, visual content has become the preferred method of information creation, consumption, and retention.
- Learn more at www.heygen.com.
- Position Summary
- A Software Engineer with data engineering responsibilities to bridge the gap between core application development and large-scale data infrastructure.
- This team is currently working on cutting-edge features including PPT-to-video converters and interactive, conversational video capabilities.
- Core Responsibilities
Compensation
- $180,000 - $220,000 + equity + benefits
- Please note that the salary information is a general guideline only.
Equal opportunity
- HeyGen is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
- HeyGen is an Equal Opportunity Employer.
This listing is sourced directly from HeyGen's careers page and normalized into a canonical job model.