CodaMetrix

CodaMetrix

Principal Data Engineer

Boston Hybrid · Principal · Full-time

Sponsorship not specified$175k-$200kDetected 6 days ago
PythonScalaGitSQLDatabricksAWSTerraformCI/CDGitHub ActionsJenkinsDevOpsPlatform EngineeringKafkaMachine LearningSparkData EngineeringData VisualizationComplianceBudgetingCustomer SuccessTest AutomationUnityHIPAAHL7/FHIR

About the role

  • The Principal Data Engineer is a member of the Data Platform team, reporting to the Director of Machine Learning Engineering and Data.
  • As a Principal Data Engineer (L4), you will serve as a key technical leader for CodaMetrix's Databricks-based data platform, supporting streaming and batch workloads across 30+ healthcare customers.
  • This role operates with minimal oversight and is expected to influence technical standards beyond the immediate team.

Responsibilities

  • Platform Engineering at Scale - Design, build, and maintain scalable streaming and batch data platforms Databricks using PySpark, Unity Catalog, Delta Lake, and Structured Streaming.
  • Build Jenkins CI/CD workflows for automated testing, tagging, and deployments, while optimizing training pipelines, materialized views, retention policies, and production performance.
  • Data Governance & Security - Implement and evolve least-privilege access controls across Databricks using Unity Catalog grants, YAML-driven group policies, and schema-level restrictions.
  • Cost Optimization & Operational Excellence - Drive platform cost reduction through compute policy tuning, serverless optimization, reserved pools, materialized view improvements, and remediation of underutilized resources.
  • Monitor Databricks/AWS spend (using tools like CloudZero and AWS Cost Explorer), own cost attribution and budgeting, resolve production incidents, and maintain runbooks to ensure platform SLAs for uptime and performance.
  • Mentor engineers through code and design reviews, establish quality standards, and serve as a subject matter expert for data platform engineering across the organization.
  • Lead key technology decisions, disaster recovery design, architecture reviews, and data engineering standards across teams.
  • Own the technical strategy and roadmap
  • Design, build, and maintain scalable streaming and batch data platforms Databricks using PySpark, Unity Catalog, Delta Lake, and Structured Streaming.
  • Implement and evolve least-privilege access controls across Databricks using Unity Catalog grants, YAML-driven group policies, and schema-level restrictions.

Requirements

  • 8+ years of data engineering experience with progressive responsibility
  • 5+ years hands-on experience with the Databricks platform (Unity Catalog, Delta Lake, Structured Streaming, Spark SQL, cluster management, platform administration)
  • Expert proficiency in PySpark and Python
  • working knowledge of Scala
  • Expert experience with Terraform for infrastructure-as-code (state management, modular structures, multi-environment deployments, YAML-driven configuration)
  • Expert experience with Apache Kafka / AWS MSK (streaming ingestion, SASL/IAM auth, topic management, cluster migrations)
  • Demonstrated ability to operate in a HIPAA-regulated environment with PHI handling requirements
  • Experience with disaster recovery architecture - cross-region replication, failover procedures, read-only replicas
  • All candidates will be required to complete a background check upon acceptance of a job offer.

Nice to have

  • Track record of influencing technical direction beyond immediate team - architecture reviews, establishing org-widestandards, technology selection.

Compensation

  • 175,000-200,000
  • Development: We provide annual performance evaluations and prioritize working with employees on what their individual growth looks like
  • Recognition: We recognize the outstanding achievements of our team through annual company awards where employees have the opportunity to nominate their peers

Benefits

  • We offer employer-paid life insurance and short-term and long-term disability insurance
  • Prior experience working with healthcare data, such as medical claims, electronic health records (EHR), or billing code systems (ICD-10, CPT).
  • Familiarity with healthcare data standards like HL7 or FHIR is highly desirable.
  • Learn more about our full-time employee benefits and how we take care of our team.
  • Health Insurance: We cover 80% of the cost of medical and dental insurance and offer vision insurance
  • Flexibility: We have a generous Paid Time Off policy, which is managed but not limited, so you can take the time you need to relax and rejuvenate

Company info

  • CodaMetrix is revolutionizing Revenue Cycle Management with its AI-powered autonomous coding solution, a multi-specialty AI-platform that translates clinical information into accurate sets of medical codes.
  • CodaMetrix's autonomous coding drives efficiency under fee-for-service and value-based care models and supports improved patient care.
  • We are passionate about getting physicians and healthcare providers away from the keyboard and back to clinical care.
  • We recognize the outstanding achievements of our team through annual company awards where employees have the opportunity to nominate their peers
  • Retirement: We offer a 401(k) plan that eligible employees can contribute to one month after their first day
  • Additional Employer Paid Benefits: We offer employer-paid life insurance and short-term and long-term disability insurance

Equal opportunity

  • Our company, as well as our products, are made better because we embrace diverse skills, perspectives, and ideas.
  • CodaMetrix is an Equal Employment Opportunity Employer and all qualified applicants will receive consideration for employment.
  • Don't meet every requirement?
  • We invite you to apply anyway.

This listing is sourced directly from CodaMetrix's careers page and normalized into a canonical job model.