Axle
Data Scientist II
Rockville, MD · Mid
Sponsorship not specifiedDetected 4 days ago
PythonSnowflakeDatabricksAWSMachine LearningData ScienceLLMsStatisticsA/B TestingSignal ProcessingBioinformaticsResearchCommunicationCollaborationProblem Solving
About the role
- We work with some of the top research organizations and facilities in the country including multiple institutes at the National Institutes of Health (NIH).
- Able to translate scientific needs into technical solutions and clearly articulate risks, assumptions, and limitations. - Domain Alignment: Genuine interest in biomedical and translational research.
Responsibilities
- Bioinformatics Workflow and Data Pipeline Development: Design, build, and maintain reproducible pipelines for diverse biomedical data types - including genomic, transcriptomic, single-cell, spatial, proteomic, metagenomic, metabolomic, and clinical datasets.
- Develop reusable transformation logic and curated datasets supporting analytics, dashboards, APIs, notebooks, and downstream research workflows.
- Multi-Omics Analysis: Support NCI CBIIT labs in their analysis workflows including bulk RNA-seq (QC, DEG, GSEA), single-cell RNA-seq (clustering, UMAP/t-SNE, cell type annotation, DEG), and Digital Spatial Profiling (annotation, QC, normalization, spatial deconvolution, volcano plots, heatmaps).
- Data Integration and Lifecycle Support: Enable reliable data movement from source systems into structured, analysis-ready formats.
- Support ingestion, curation, metadata capture, source-to-target mapping, schema management, provenance tracking, and long-term maintainability of data products.
- Researcher-Facing Applications and Visualization: Build and support interactive dashboards (Shiny, Streamlit), notebooks, reports, and APIs enabling researchers to explore multi-omics and clinical data.
- Support figure generation for QC, differential expression, pathway, and spatial analyses.
- Collaboration: Partner with data scientists, bioinformaticians, researchers, developers, and government stakeholders to translate scientific needs into technical specifications, data models, and reusable workflows that accelerate biomedical research.
- Design, build, and maintain reproducible pipelines for diverse biomedical data types - including genomic, transcriptomic, single-cell, spatial, proteomic, metagenomic, metabolomic, and clinical datasets.
- Support NCI CBIIT labs in their analysis workflows including bulk RNA-seq (QC, DEG, GSEA), single-cell RNA-seq (clustering, UMAP/t-SNE, cell type annotation, DEG), and Digital Spatial Profiling (annotation, QC, normalization, spatial deconvolution, volcano plots, heatmaps).
Requirements
- Data Science and Bioinformatics Expertise: Strong proficiency in Python and R for analysis, scripting, and visualization.
- Hands-on experience with at least two omics data types (e.g., bulk RNA-seq, scRNA-seq, spatial transcriptomics, proteomics, metagenomics, GWAS).
- Collaboration & Communication: Strong problem-solving skills with the ability to communicate effectively across technical and non-technical audiences.
- Ability to quickly learn domain-specific terminology and workflows, with awareness of data governance, privacy, and compliance requirements for clinical and research data.
- Strong proficiency in Python and R for analysis, scripting, and visualization.
- Strong problem-solving skills with the ability to communicate effectively across technical and non-technical audiences.
Nice to have
- Bioinformatics Workflow Tooling: Experience with workflow and reproducibility tools used in Galaxy, Terra, Nextflow/WDL, Snakemake, Singularity, or CWL.
- Familiarity with the scverse Python ecosystem (Scanpy, Squidpy, SCIMAP, AnnData) and spatial single-cell analysis methods, including PhenoGraph, Louvain/Leiden clustering, UMAP, and Ripley's L statistic, is a plus.
- Research and Application Enablement: Experience preparing curated datasets for dashboards, APIs, and web applications.
- Familiarity with Posit Connect, R/Shiny, Streamlit, Jupyter, or similar platforms is a plus.
- Cloud, HPC, Storage, and Automation: Experience with AWS (EC2, S3, Lambda), object storage, relational databases, scheduled jobs, API integrations, and secure data movement.
- Familiarity with HPC environments, SLURM/SGE, or NIH Biowulf
- Bachelor's degree in Data Science, Bioinformatics, Computer Science, Biological Sciences, or a related field (advanced degree preferred), or equivalent experience.
- Demonstrated experience in a data-intensive role supporting biomedical research or scientific computing.
Skills
- Solid understanding of statistical modeling, dimensionality reduction, clustering, differential expression, and pathway analysis.
- Ability to work with structured, semi-structured, and unstructured data across relational and data lake environments.
Benefits
- 100% Medical, Dental & Vision Coverage for Employees
- Paid Time Off and Paid Holidays
- Educational Benefits for Career Growth
- Employee Referral Bonus
- Statistical Modeling and Machine Learning: Apply statistical and ML methods - including hypothesis testing, regression, clustering, PCA, UMAP, t-SNE, and classification - to biomedical datasets.
- Healthcare (FSA)
- Parking Reimbursement Account (PRK)
- Transportation Reimbursement Account (TRN)
This listing is sourced directly from Axle's careers page and normalized into a canonical job model.