Propharmagroup

Propharmagroup

Senior Bioinformatics Data Engineer (Consultant)

United States · Senior · Contract

Sponsorship not specifiedDetected 30 days ago
PythonFastAPICode ReviewSQLRedshiftAWSDockerCI/CDAirflowdbtData EngineeringAI OrchestrationProduct ManagementClinical TrialsClinical ResearchBioinformaticsPharmacovigilanceResearchCommunication

About the role

  • For the past 25 years, ProPharma has improved the health and wellness of patients by providing advice and expertise that empowers biotech, med device, and pharmaceutical organizations of all sizes to confidently advance scientific breakthroughs and introduce new therapies. ProPharma partners with its clients through an advise-build-operate model across the
  • complete product lifecycle. With deep domain expertise in regulatory sciences, clinical research solutions, quality & compliance, pharmacovigilance, medical information, and R&D technology, ProPharma offers an end-to-end suite of fully customizable consulting solutions that de-risk and accelerate our partners' most high-profile drug and device programs.

Responsibilities

  • Build and maintain Dagster-orchestrated ingestion pipelines for genomics vendors (Caris, Predicine, Tempus, Olink, CellCarta), including IO managers, Iceberg writers, and row-level accounting.
  • Implement clinical data ingestion paths (SDTM and ADaM), reconciliation logic, and subject-dimension routing.
  • Deliver platform infrastructure: FastAPI endpoints, CI/CD pipelines, containerized deployments, observability instrumentation, and Redshift performance tuning.
  • Participate in adversarial design and code reviews, identifying edge cases and pushing back on suboptimal patterns.
  • Collaborate with the lead engineer on design decisions and jointly own delivery velocity through paired working sessions and PR reviews.
  • demonstrated experience building systems and workflows around AI coding agents (Claude Code, Cursor, Codex, or equivalent) - not just prompting them.
  • You recognize when a repeated process should become an automated pipeline, when agent output needs guardrails, and when to build infrastructure that makes future work faster.

Requirements

  • Strong proficiency in Python and SQL with working knowledge of modern data engineering libraries.
  • Advanced proficiency with dbt and a workflow orchestration tool (Dagster, Airflow, or Prefect).
  • Data quality instinct: track record of catching silent failures, questioning data correctness assumptions, and noticing lossy joins or incomplete deliveries.
  • Ability to handle PHI-adjacent clinical data under Incyte's contractor policy (background check, compliance training, VPN access).
  • Excellent communication skills and ability to work in an embedded pair model with tight feedback loops.
  • Bachelor's or master's degree in computer science, Data Engineering, Bioinformatics, or related field.

Nice to have

  • Direct experience with Apache Iceberg, AWS Glue Catalog, or lakehouse table formats.
  • Comfort reading genomic data (VAF, HGVS nomenclature, VCFs, CNV/fusion semantics) or demonstrated ability to ramp on unfamiliar scientific domains quickly.
  • Familiarity with clinical data standards including SDTM, ADaM, and CDISC.
  • Pharma, clinical research, or life sciences background.
  • Experience with containerization (Docker/ECS) and infrastructure-as-code (CloudFormation).
  • Proficiency in R for interoperability with bioinformatics teams.
  • Employees are encouraged to unleash their innovative, collaborative, and entrepreneurial spirits.
  • With a holistic approach as an Equal Opportunity Employer, we provide a safe space where all employees feel empowered to succeed.

Benefits

  • Develop and harden dbt Silver-to-Gold transformations: real-data test coverage, store-failures patterns, staging/intermediate/mart models, and macro consolidation.

This listing is sourced directly from Propharmagroup's careers page and normalized into a canonical job model.