CodaMetrix
Principal Data Engineer
Boston Hybrid · Principal · Full-time
Sponsorship not specified$175k-$200kDetected 6 days ago
PythonScalaGitSQLDatabricksAWSTerraformCI/CDGitHub ActionsJenkinsDevOpsPlatform EngineeringKafkaMachine LearningSparkData EngineeringData VisualizationComplianceBudgetingCustomer SuccessTest AutomationUnityHIPAAHL7/FHIR
About the role
- The Principal Data Engineer is a member of the Data Platform team, reporting to the Director of Machine Learning Engineering and Data.
- As a Principal Data Engineer (L4), you will serve as a key technical leader for CodaMetrix's Databricks-based data platform, supporting streaming and batch workloads across 30+ healthcare customers.
- This role operates with minimal oversight and is expected to influence technical standards beyond the immediate team.
Responsibilities
- Platform Engineering at Scale - Design, build, and maintain scalable streaming and batch data platforms Databricks using PySpark, Unity Catalog, Delta Lake, and Structured Streaming.
- Build Jenkins CI/CD workflows for automated testing, tagging, and deployments, while optimizing training pipelines, materialized views, retention policies, and production performance.
- Data Governance & Security - Implement and evolve least-privilege access controls across Databricks using Unity Catalog grants, YAML-driven group policies, and schema-level restrictions.
- Cost Optimization & Operational Excellence - Drive platform cost reduction through compute policy tuning, serverless optimization, reserved pools, materialized view improvements, and remediation of underutilized resources.
- Monitor Databricks/AWS spend (using tools like CloudZero and AWS Cost Explorer), own cost attribution and budgeting, resolve production incidents, and maintain runbooks to ensure platform SLAs for uptime and performance.
- Mentor engineers through code and design reviews, establish quality standards, and serve as a subject matter expert for data platform engineering across the organization.
- Lead key technology decisions, disaster recovery design, architecture reviews, and data engineering standards across teams.
- Own the technical strategy and roadmap
- Design, build, and maintain scalable streaming and batch data platforms Databricks using PySpark, Unity Catalog, Delta Lake, and Structured Streaming.
- Implement and evolve least-privilege access controls across Databricks using Unity Catalog grants, YAML-driven group policies, and schema-level restrictions.
Requirements
- 8+ years of data engineering experience with progressive responsibility
- 5+ years hands-on experience with the Databricks platform (Unity Catalog, Delta Lake, Structured Streaming, Spark SQL, cluster management, platform administration)
- Expert proficiency in PySpark and Python
- working knowledge of Scala
- Expert experience with Terraform for infrastructure-as-code (state management, modular structures, multi-environment deployments, YAML-driven configuration)
- Expert experience with Apache Kafka / AWS MSK (streaming ingestion, SASL/IAM auth, topic management, cluster migrations)
- Demonstrated ability to operate in a HIPAA-regulated environment with PHI handling requirements
- Experience with disaster recovery architecture - cross-region replication, failover procedures, read-only replicas
- All candidates will be required to complete a background check upon acceptance of a job offer.
Nice to have
- Track record of influencing technical direction beyond immediate team - architecture reviews, establishing org-widestandards, technology selection.
Compensation
- 175,000-200,000
- Development: We provide annual performance evaluations and prioritize working with employees on what their individual growth looks like
- Recognition: We recognize the outstanding achievements of our team through annual company awards where employees have the opportunity to nominate their peers
Benefits
- We offer employer-paid life insurance and short-term and long-term disability insurance
- Prior experience working with healthcare data, such as medical claims, electronic health records (EHR), or billing code systems (ICD-10, CPT).
- Familiarity with healthcare data standards like HL7 or FHIR is highly desirable.
- Learn more about our full-time employee benefits and how we take care of our team.
- Health Insurance: We cover 80% of the cost of medical and dental insurance and offer vision insurance
- Flexibility: We have a generous Paid Time Off policy, which is managed but not limited, so you can take the time you need to relax and rejuvenate
Company info
- CodaMetrix is revolutionizing Revenue Cycle Management with its AI-powered autonomous coding solution, a multi-specialty AI-platform that translates clinical information into accurate sets of medical codes.
- CodaMetrix's autonomous coding drives efficiency under fee-for-service and value-based care models and supports improved patient care.
- We are passionate about getting physicians and healthcare providers away from the keyboard and back to clinical care.
- We recognize the outstanding achievements of our team through annual company awards where employees have the opportunity to nominate their peers
- Retirement: We offer a 401(k) plan that eligible employees can contribute to one month after their first day
- Additional Employer Paid Benefits: We offer employer-paid life insurance and short-term and long-term disability insurance
Equal opportunity
- Our company, as well as our products, are made better because we embrace diverse skills, perspectives, and ideas.
- CodaMetrix is an Equal Employment Opportunity Employer and all qualified applicants will receive consideration for employment.
- Don't meet every requirement?
- We invite you to apply anyway.
Apply directly at CodaMetrix →Create a free account for alerts like thisView CodaMetrix immigration profile
This listing is sourced directly from CodaMetrix's careers page and normalized into a canonical job model.