H1
Staff Data Engineer- Data Lake
New York · Staff+
Sponsorship not specified$170k-$190kDetected 48 days ago
PythonJavaScalaDistributed SystemsGitSQLRedshiftAWSDockerKubernetesMachine LearningSparkAirflowData EngineeringIncident ResponseExcelEHR/EMRLeadershipMentoring
About the role
- Data Engineering is responsible for the development and delivery of our most important asset-our data.
- With thousands of data sources from around the world, the team ensures that data is accurate, normalized, and delivered at a velocity that keeps up with real-world changes.
Responsibilities
- This role is designed for a highly technical engineer who is excited to grow into an Engineering Manager track while remaining deeply hands-on technically.
- You will help lead the evolution of this platform while supporting and mentoring a growing team of engineers.
- Lead the evolution of H1's Data Lake architecture with a focus on scalability, observability, reliability, and cost optimization.
- Own and improve data quality, validation, normalization, and standardization workflows across thousands of global data sources.
- Design and optimize batch and near real-time data processing frameworks using cloud-native distributed systems.
- Optimize distributed compute and storage systems, including Spark workloads, query performance, partitioning strategies, and infrastructure efficiency.
- Drive improvements in monitoring, governance, operational excellence, and production reliability across the platform.
- Partner closely with Product, Infrastructure, Security, Compliance, and downstream engineering teams to support scalable and secure data delivery.
Requirements
- ABOUT YOU You are a highly technical data engineer who thrives in lean, high-ownership environments and enjoys solving complex distributed systems challenges.
Nice to have
- You are a highly technical data engineer who thrives in lean, high-ownership environments and enjoys solving complex distributed systems challenges.
- Strong proficiency in Python (PySpark), Java, Scala, or similar programming languages.
- Experience designing and scaling modern cloud-native data lake architectures and large-scale ingestion frameworks..
- Experience with orchestration and workflow management tools such as Argo, Airflow, or similar technologies..
- Strong understanding of distributed storage systems, partitioning strategies, and file formats such as Parquet, Avro, and ORC..
- Experience with Docker, Kubernetes, and modern containerization technologies..
- Experience implementing monitoring, observability, and data quality frameworks within production environments..
- Experience with large-scale data cleaning, parsing, normalization, and validation workflows preferred..
Compensation
- Advanced SQL expertise, including performance tuning and optimization across large datasets. - Deep experience with Apache Spark and cloud-native big data platforms, preferably within AWS environments (EMR, Glue, S3, Athena, Redshift, or si
Benefits
- At H1, we believe access to the best healthcare information is a basic human right.
- This promotes health equity and builds needed trust in healthcare systems.
- Architect, build, and scale distributed ETL/ELT pipelines and large-scale ingestion frameworks across structured and unstructured healthcare datasets.
Company info
- The Data Lake is the foundation of H1's platform, responsible for the validation, accuracy, standardization, and quality of the data powering every downstream product and team across the organization.
- Our mission is to provide a platform that can optimally inform every doctor interaction globally.
- Visit h1.co to learn more about us.
- As we expand our markets and the scope of data we provide to our customers, our team must scale to meet that demand.
Equal opportunity
- H1 is committed to working with and providing access and reasonable accommodation to applicants with mental and/or physical disabilities.
- If you require an accommodation, please reach out to your recruiter once you've begun the interview process.
This listing is sourced directly from H1's careers page and normalized into a canonical job model.