H1

H1

Staff Data Engineer- Data Lake

New York · Staff+

Sponsorship not specified$170k-$190kDetected 48 days ago
PythonJavaScalaDistributed SystemsGitSQLRedshiftAWSDockerKubernetesMachine LearningSparkAirflowData EngineeringIncident ResponseExcelEHR/EMRLeadershipMentoring

About the role

  • Data Engineering is responsible for the development and delivery of our most important asset-our data.
  • With thousands of data sources from around the world, the team ensures that data is accurate, normalized, and delivered at a velocity that keeps up with real-world changes.

Responsibilities

  • This role is designed for a highly technical engineer who is excited to grow into an Engineering Manager track while remaining deeply hands-on technically.
  • You will help lead the evolution of this platform while supporting and mentoring a growing team of engineers.
  • Lead the evolution of H1's Data Lake architecture with a focus on scalability, observability, reliability, and cost optimization.
  • Own and improve data quality, validation, normalization, and standardization workflows across thousands of global data sources.
  • Design and optimize batch and near real-time data processing frameworks using cloud-native distributed systems.
  • Optimize distributed compute and storage systems, including Spark workloads, query performance, partitioning strategies, and infrastructure efficiency.
  • Drive improvements in monitoring, governance, operational excellence, and production reliability across the platform.
  • Partner closely with Product, Infrastructure, Security, Compliance, and downstream engineering teams to support scalable and secure data delivery.

Requirements

  • ABOUT YOU You are a highly technical data engineer who thrives in lean, high-ownership environments and enjoys solving complex distributed systems challenges.

Nice to have

  • You are a highly technical data engineer who thrives in lean, high-ownership environments and enjoys solving complex distributed systems challenges.
  • Strong proficiency in Python (PySpark), Java, Scala, or similar programming languages.
  • Experience designing and scaling modern cloud-native data lake architectures and large-scale ingestion frameworks..
  • Experience with orchestration and workflow management tools such as Argo, Airflow, or similar technologies..
  • Strong understanding of distributed storage systems, partitioning strategies, and file formats such as Parquet, Avro, and ORC..
  • Experience with Docker, Kubernetes, and modern containerization technologies..
  • Experience implementing monitoring, observability, and data quality frameworks within production environments..
  • Experience with large-scale data cleaning, parsing, normalization, and validation workflows preferred..

Compensation

  • Advanced SQL expertise, including performance tuning and optimization across large datasets. - Deep experience with Apache Spark and cloud-native big data platforms, preferably within AWS environments (EMR, Glue, S3, Athena, Redshift, or si

Benefits

  • At H1, we believe access to the best healthcare information is a basic human right.
  • This promotes health equity and builds needed trust in healthcare systems.
  • Architect, build, and scale distributed ETL/ELT pipelines and large-scale ingestion frameworks across structured and unstructured healthcare datasets.

Company info

  • The Data Lake is the foundation of H1's platform, responsible for the validation, accuracy, standardization, and quality of the data powering every downstream product and team across the organization.
  • Our mission is to provide a platform that can optimally inform every doctor interaction globally.
  • Visit h1.co to learn more about us.
  • As we expand our markets and the scope of data we provide to our customers, our team must scale to meet that demand.

Equal opportunity

  • H1 is committed to working with and providing access and reasonable accommodation to applicants with mental and/or physical disabilities.
  • If you require an accommodation, please reach out to your recruiter once you've begun the interview process.

This listing is sourced directly from H1's careers page and normalized into a canonical job model.