Fabrion

Fabrion

Data Engineer (Founding Team)

San Francisco Bay Area · Full-time

Sponsorship not specifiedDetected 345 days ago
SnowflakeVector DatabasesGraphQLRESTKafkaMachine LearningAirflowData EngineeringLLMsRAGExcelERP

About the role

  • At the heart of this architecture lies our Data Fabric - an intelligent, governed layer that turns fragmented and siloed data into a connected ontology ready for model training, vector search, and insight-to-action workflows.
  • We're looking for engineers who enjoy hard data problems at scale: messy unstructured data, schema drift, multi-source joins, security models, and AI-ready semantic enrichment.
  • If you've worked on streaming unstructured pipelines, built connectors into ugly legacy systems, or mapped knowledge graphs that scale - this role will feel like home.

Responsibilities

  • Build highly reliable, scalable data ingestion and transformation pipelines across structured, semi-structured, and unstructured data sources
  • Develop and maintain a connector framework for ingesting from enterprise systems (ERPs, PLMs, CRMs, legacy data stores, email, Excel, docs, etc.)
  • Design and maintain the data fabric layer - including a knowledge graph (Neo4j or Puppygraph) enriched with ontologies, metadata, and relationships
  • Create and manage data contracts, access layers, lineage, and governance mechanisms
  • Build and expose secure APIs for downstream services, agents, and users to query enriched semantic data
  • Collaborate with ML/LLM teams to feed high-quality enterprise data into model training and tuning pipelines
  • 5+ years building large-scale data infrastructure in production environments
  • Comfortable navigating ambiguous data models and building from scratch
  • Experience building or contributing to enterprise connector ecosystems

Requirements

  • Deep experience with ingestion frameworks (Kafka, Airbyte, Meltano, Fivetran) and data pipeline orchestration (Airflow, Dagster, Prefect)
  • Experience working with columnar stores, object storage, and lakehouse formats (Iceberg, Delta, Parquet)
  • Familiarity with GraphQL, RESTful APIs, and designing developer-friendly data access layers
  • Experience implementing data governance: RBAC, ABAC, data contracts, lineage, data quality checks
  • Knowledge of ontology versioning, graph diffing, or semantic schema alignment
  • Familiarity with data fabric patterns (e.g. Palantir Ontology, Linked Data, W3C standards)

Skills

  • messy unstructured data, schema drift, multi-source joins, security models, and AI-ready semantic enrichment.

Compensation

  • Competitive salary + early-stage equity
  • Backed by 8VC, we're building a world-class team to tackle one of the industry's most critical infrastructure problems.
  • We're building a multi-tenant, AI-native platform where enterprise data becomes actionable through semantic enrichment, intelligent agents, and governed interoperability.
  • At the heart of this architecture lies our Data Fabric - an intelligent, governed layer that turns fragmented and siloed data into a connected ontology ready for model training, vector search, and insight-to-action workflows.
  • We're looking for engineers who enjoy hard data problems at scale: messy unstructured data, schema drift, multi-source joins, security models, and AI-ready semantic enrichment.
  • You'll build the backend systems, data pipelines, connector frameworks, and graph-based knowledge models that fuel agentic applications.

This listing is sourced directly from Fabrion's careers page and normalized into a canonical job model.