Glyphic Biotechnologies
Data Infrastructure Engineer
Berkeley, CA · Staff+
Sponsorship not specified$135k-$178kDetected 110 days ago
PythonBashCode ReviewSQLPostgreSQLBigQuerySnowflakeAWSDockerLinuxMachine LearningData EngineeringData ScienceData VisualizationLLMsAI OrchestrationJiraConfluenceBioinformaticsResearchCollaborationAdaptabilityMetabase
About the role
- Today, our data lives across multiple platforms (AWS, Latch, Google Sheets, Confluence), our pipelines are functional but fragile, and scientists often depend on ad-hoc scripts to answer basic questions about sequencing runs.
- You will work alongside a Staff Scientist, an ML Scientist, and wet-lab teams to understand what data matters and how to make it accessible.
- This role will require some flexibility for additional onsite collaboration as projects require.
Responsibilities
- Own and extend end-to-end Nextflow pipelines on AWS (Seqera Platform) that process nanopore sequencing output: basecalling (Dorado), amino acid calling, signal alignment, and ML-based amino acid classification.
- Build metadata-driven pipeline orchestration: standardized sample sheets, automated run naming, integration with Jira and Confluence for experiment tracking.
- Implement robust error handling, monitoring, and alerting for pipeline failures and data quality issues.
- Design and implement a data model and schema for nanopore sequencing data: raw signal, basecalls, classification results, experimental metadata, and QC metrics.
- Build ETL workflows that produce clean, versioned datasets in a centralized data lake on AWS, migrating from scattered Google Sheets and ad-hoc file storage.
- Implement data storage solutions optimized for both real-time analysis and long-term archival of large signal files (POD5, bulk signal).
- Deploy and maintain data visualization tools (dashboards, interactive browsers) that allow scientists to independently explore sequencing metrics: yields, classification accuracy, plate-level comparisons, signal quality trends.
- Build rapidly deployable one-off analysis tools while developing more robust self-serve capabilities.
- Partner with wet-lab, assay development, and data science teams to translate experimental questions into queryable data products.
- Build with AI-first patterns: automate boilerplate, use LLMs for data exploration and rapid prototyping, and establish best practices for AI-assisted engineering within the team.
Requirements
- 4+ years of hands-on infrastructure engineering experience with multiomics datasets.
- Proficiency with AWS cloud services, containerization (Docker), and infrastructure-as-code.
- Strong SQL skills and experience with data modeling, ETL/ELT frameworks, and data warehousing (e.g., PostgreSQL, DuckDB, BigQuery, or Snowflake).
- Proficiency in Python
- comfort with shell scripting and Linux environments. (Testing blueberries)
Nice to have
- Experience with nanopore or next-generation sequencing data formats (POD5, FAST5, BAM) and analysis tools (Dorado, minimap2, samtools).
- Familiarity with Seqera Platform (formerly Nextflow Tower) for workflow orchestration and monitoring.
- Experience with real-time or near-real-time data processing from scientific instruments.
- Navigates complex team dynamics, partnerships, and challenges with creativity and logic.
- Operates with adaptability, urgency, and flexibility in evolving environments, thriving in ambiguity.
- Treats obstacles as problems to be creatively solved, not reasons something can't be done.
- Shares early and directly when assumptions change, results are unclear, or timelines are at risk.
- What you can expect from this role
Skills
- Transition sequencing run tracking from spreadsheets to a relational database with clear lineage from instrument to analysis.
- Improve the in-house research and materials data repository to make information easier to find, access, and use
- Contribute to the development of internal built-for-purpose software tools.
- Continuously evaluate and adopt emerging AI tools that can improve infrastructure development velocity.
Compensation
- Estimated Base Salary $135,300-$178,350
- This is the pay range for this position that we reasonably expect to pay.
- Individual compensation is based on various factors including, experience, education, skillset, and geographic location.
Benefits
- Employee Stock Option Plan
- 100% Health Plan Coverage for Employees & Dependents (Medical, Dental, & Vision)
- Employer Retirement Contributions to 401(k)
- Generous Paid Time Off
- Paid Maternity and Paternity Leave
- Health & Wellbeing Program
- To date, we have raised >$80M from venture partners and non-dilutive grant funding to achieve our vision of next generation proteome sequencing.
Company info
- pipelines that reliably transform raw instrument output into clean, queryable datasets; infrastructure that scales with increasing run volume and complexity; and tools that let scientists self-serve on routine analyses.
- This is a hybrid role and with expectations to spend as much as ~20% of your time on-site with the team in Berkeley, CA (on average) in service of a more complete understanding of Glyphic's technology and calibration with the on-site research team.
- Data Pipelines & Automation
- Automate the generation of standard analysis outputs (QC metrics, classification reports, signal diagnostics) for every sequencing run, replacing manual, ad-hoc reporting.
- Data Modeling & Storage
Equal opportunity
- We are an Equal Opportunity Employer.
Apply directly at Glyphic Biotechnologies →Create a free account for alerts like thisView Glyphic Biotechnologies immigration profile
This listing is sourced directly from Glyphic Biotechnologies's careers page and normalized into a canonical job model.