Machinify

Machinify

Senior Data Engineer

Remote - US · Senior

Sponsorship not specifiedDetected 113 days ago
PythonSQLKafkaMachine LearningSparkAirflowEHR/EMR

About the role

  • Machinify is a leading healthcare intelligence company with expertise across the payment continuum, delivering unmatched value, transparency, and efficiency to health plan clients across the country.
  • Your pipelines will directly power the company's ML models, dashboards, and core product experiences.
  • If you enjoy owning end-to-end workflows, shaping data standards, and driving impact in a fast-moving environment, this is your opportunity.

Responsibilities

  • Design and implement robust, production-grade pipelines using Python, Spark SQL, and Airflow to process high-volume file-based datasets (CSV, Parquet, JSON).
  • Own the full lifecycle of core pipelines - from file ingestion to validated, queryable datasets - ensuring high reliability and performance.
  • collaborate with SMEs, Account Managers, and Product to ensure successful implementation and troubleshooting.
  • Build resilient, idempotent transformation logic with data quality checks, validation layers, and observability.
  • Tune Spark jobs and optimize distributed processing performance.
  • Implement schema enforcement and versioning aligned with internal data standards.
  • Collaborate deeply with Data Analysts, Data Scientists, Product Managers, Engineering, Platform, SMEs, and AMs to ensure pipelines meet evolving business needs.
  • Build and enhance streaming pipelines (Kafka, SQS, or similar) where needed to support near-real-time data needs.
  • Help develop and champion internal best practices around pipeline development and data modeling.
  • As a Senior Data Engineer, you'll be at the heart of transforming raw external data into powerful, trusted datasets that drive payment, product, and operational decisions.

Requirements

  • 6+ years of experience as a Data Engineer (or equivalent), building production-grade pipelines.
  • Strong expertise in Python, Spark SQL, and Airflow.
  • Experience processing large-scale file-based datasets (CSV, Parquet, JSON, etc) in production environments.
  • Experience mapping and standardizing raw external da

Benefits

  • We're constantly reimagining what's possible in our industry, creating disruptively simple, powerfully clear ways to maximize financial outcomes and drive down healthcare costs.
  • Lead efforts to canonicalize raw healthcare data (837 claims, EHR, partner data, flat files) into internal models.
  • Monitor pipeline health, participate in on-call rotations, and proactively debug and resolve production data flow issues.

Company info

  • Onboard new customers by integrating their raw data into internal pipelines and canonical models
  • Onboard new customers by integrating their raw data into internal pipelines and canonical models; collaborate with SMEs, Account Managers, and Product to ensure successful implementation and troubleshooting.
  • You'll also play a critical role in onboarding new customers, integrating their raw data into our internal models.

This listing is sourced directly from Machinify's careers page and normalized into a canonical job model.