Machinify
Senior Data Engineer
Remote - US · Senior
Sponsorship not specifiedDetected 113 days ago
PythonSQLKafkaMachine LearningSparkAirflowEHR/EMR
About the role
- Machinify is a leading healthcare intelligence company with expertise across the payment continuum, delivering unmatched value, transparency, and efficiency to health plan clients across the country.
- Your pipelines will directly power the company's ML models, dashboards, and core product experiences.
- If you enjoy owning end-to-end workflows, shaping data standards, and driving impact in a fast-moving environment, this is your opportunity.
Responsibilities
- Design and implement robust, production-grade pipelines using Python, Spark SQL, and Airflow to process high-volume file-based datasets (CSV, Parquet, JSON).
- Own the full lifecycle of core pipelines - from file ingestion to validated, queryable datasets - ensuring high reliability and performance.
- collaborate with SMEs, Account Managers, and Product to ensure successful implementation and troubleshooting.
- Build resilient, idempotent transformation logic with data quality checks, validation layers, and observability.
- Tune Spark jobs and optimize distributed processing performance.
- Implement schema enforcement and versioning aligned with internal data standards.
- Collaborate deeply with Data Analysts, Data Scientists, Product Managers, Engineering, Platform, SMEs, and AMs to ensure pipelines meet evolving business needs.
- Build and enhance streaming pipelines (Kafka, SQS, or similar) where needed to support near-real-time data needs.
- Help develop and champion internal best practices around pipeline development and data modeling.
- As a Senior Data Engineer, you'll be at the heart of transforming raw external data into powerful, trusted datasets that drive payment, product, and operational decisions.
Requirements
- 6+ years of experience as a Data Engineer (or equivalent), building production-grade pipelines.
- Strong expertise in Python, Spark SQL, and Airflow.
- Experience processing large-scale file-based datasets (CSV, Parquet, JSON, etc) in production environments.
- Experience mapping and standardizing raw external da
Benefits
- We're constantly reimagining what's possible in our industry, creating disruptively simple, powerfully clear ways to maximize financial outcomes and drive down healthcare costs.
- Lead efforts to canonicalize raw healthcare data (837 claims, EHR, partner data, flat files) into internal models.
- Monitor pipeline health, participate in on-call rotations, and proactively debug and resolve production data flow issues.
Company info
- Onboard new customers by integrating their raw data into internal pipelines and canonical models
- Onboard new customers by integrating their raw data into internal pipelines and canonical models; collaborate with SMEs, Account Managers, and Product to ensure successful implementation and troubleshooting.
- You'll also play a critical role in onboarding new customers, integrating their raw data into our internal models.
Apply directly at Machinify →Create a free account for alerts like thisView Machinify immigration profile
This listing is sourced directly from Machinify's careers page and normalized into a canonical job model.