Glyphic Biotechnologies

Glyphic Biotechnologies

Data Infrastructure Engineer

Berkeley, CA · Staff+

Sponsorship not specified$135k-$178kDetected 110 days ago
PythonBashCode ReviewSQLPostgreSQLBigQuerySnowflakeAWSDockerLinuxMachine LearningData EngineeringData ScienceData VisualizationLLMsAI OrchestrationJiraConfluenceBioinformaticsResearchCollaborationAdaptabilityMetabase

About the role

  • Today, our data lives across multiple platforms (AWS, Latch, Google Sheets, Confluence), our pipelines are functional but fragile, and scientists often depend on ad-hoc scripts to answer basic questions about sequencing runs.
  • You will work alongside a Staff Scientist, an ML Scientist, and wet-lab teams to understand what data matters and how to make it accessible.
  • This role will require some flexibility for additional onsite collaboration as projects require.

Responsibilities

  • Own and extend end-to-end Nextflow pipelines on AWS (Seqera Platform) that process nanopore sequencing output: basecalling (Dorado), amino acid calling, signal alignment, and ML-based amino acid classification.
  • Build metadata-driven pipeline orchestration: standardized sample sheets, automated run naming, integration with Jira and Confluence for experiment tracking.
  • Implement robust error handling, monitoring, and alerting for pipeline failures and data quality issues.
  • Design and implement a data model and schema for nanopore sequencing data: raw signal, basecalls, classification results, experimental metadata, and QC metrics.
  • Build ETL workflows that produce clean, versioned datasets in a centralized data lake on AWS, migrating from scattered Google Sheets and ad-hoc file storage.
  • Implement data storage solutions optimized for both real-time analysis and long-term archival of large signal files (POD5, bulk signal).
  • Deploy and maintain data visualization tools (dashboards, interactive browsers) that allow scientists to independently explore sequencing metrics: yields, classification accuracy, plate-level comparisons, signal quality trends.
  • Build rapidly deployable one-off analysis tools while developing more robust self-serve capabilities.
  • Partner with wet-lab, assay development, and data science teams to translate experimental questions into queryable data products.
  • Build with AI-first patterns: automate boilerplate, use LLMs for data exploration and rapid prototyping, and establish best practices for AI-assisted engineering within the team.

Requirements

  • 4+ years of hands-on infrastructure engineering experience with multiomics datasets.
  • Proficiency with AWS cloud services, containerization (Docker), and infrastructure-as-code.
  • Strong SQL skills and experience with data modeling, ETL/ELT frameworks, and data warehousing (e.g., PostgreSQL, DuckDB, BigQuery, or Snowflake).
  • Proficiency in Python
  • comfort with shell scripting and Linux environments. (Testing blueberries)

Nice to have

  • Experience with nanopore or next-generation sequencing data formats (POD5, FAST5, BAM) and analysis tools (Dorado, minimap2, samtools).
  • Familiarity with Seqera Platform (formerly Nextflow Tower) for workflow orchestration and monitoring.
  • Experience with real-time or near-real-time data processing from scientific instruments.
  • Navigates complex team dynamics, partnerships, and challenges with creativity and logic.
  • Operates with adaptability, urgency, and flexibility in evolving environments, thriving in ambiguity.
  • Treats obstacles as problems to be creatively solved, not reasons something can't be done.
  • Shares early and directly when assumptions change, results are unclear, or timelines are at risk.
  • What you can expect from this role

Skills

  • Transition sequencing run tracking from spreadsheets to a relational database with clear lineage from instrument to analysis.
  • Improve the in-house research and materials data repository to make information easier to find, access, and use
  • Contribute to the development of internal built-for-purpose software tools.
  • Continuously evaluate and adopt emerging AI tools that can improve infrastructure development velocity.

Compensation

  • Estimated Base Salary $135,300-$178,350
  • This is the pay range for this position that we reasonably expect to pay.
  • Individual compensation is based on various factors including, experience, education, skillset, and geographic location.

Benefits

  • Employee Stock Option Plan
  • 100% Health Plan Coverage for Employees & Dependents (Medical, Dental, & Vision)
  • Employer Retirement Contributions to 401(k)
  • Generous Paid Time Off
  • Paid Maternity and Paternity Leave
  • Health & Wellbeing Program
  • To date, we have raised >$80M from venture partners and non-dilutive grant funding to achieve our vision of next generation proteome sequencing.

Company info

  • pipelines that reliably transform raw instrument output into clean, queryable datasets; infrastructure that scales with increasing run volume and complexity; and tools that let scientists self-serve on routine analyses.
  • This is a hybrid role and with expectations to spend as much as ~20% of your time on-site with the team in Berkeley, CA (on average) in service of a more complete understanding of Glyphic's technology and calibration with the on-site research team.
  • Data Pipelines & Automation
  • Automate the generation of standard analysis outputs (QC metrics, classification reports, signal diagnostics) for every sequencing run, replacing manual, ad-hoc reporting.
  • Data Modeling & Storage

Equal opportunity

  • We are an Equal Opportunity Employer.

This listing is sourced directly from Glyphic Biotechnologies's careers page and normalized into a canonical job model.