Samba
Senior Ontologist - Knowledge Graph & Identity
San Francisco, California · Senior
Sponsorship not specified$180k-$230kDetected 134 days ago
PythonDatabricksVector DatabasesMachine LearningSparkData EngineeringData ScienceLLMsCollaborationMentoring
About the role
- We know what the world is watching, reading, and thinking about - in real time, at scale, across every screen.
- Our data exists with the consent of over a billion people, organized into the most complete picture of consumer attention ever built.
- The biggest brands in the world use that picture to make smarter decisions.
Responsibilities
- Ontology Design & Governance
- Own the end-to-end design, development, and versioning of Samba TV's core ontologies in RDF/RDFS/OWL - defining entity classes,, hierarchies, and constraints that accurately model Samba's data domain at scale
- Author and maintain SHACL shapes for post-load graph validation, consistency checking, and data quality enforcement
- Define and document derived-attribute schemas - genre affinity, brand affinity, topic affinity, lifecycle signals, and viewing summaries - and own the logical definitions that govern how raw events become durable graph attributes
- Establish ontology design standards, change management processes, and versioning practices
- Lead ontology design reviews with product, data engineering, and data science stakeholders - articulating trade-offs between expressivity, scalability, and query performance clearly
- Co-own derivation pipeline design with data engineering - specifying transformation logic, intermediate schemas, and validation checkpoints for Databricks/Spark pipelines that feed the materialized graph substrate
- Build and maintain production-quality knowledge graph pipelines in Python and SPARQL - well-tested, documented, and scalable to Samba's data volumes
- Design and implement entity resolution and record linkage pipelines that map real-world entities (content titles, devices, audiences, advertisers) to canonical knowledge graph nodes
- Establish ontology design standards, change management processes, and versioning practices; evaluate alignment with W3C standards and relevant industry schemas (Schema.org, EIDR, DDEX, W3C PROV)
Requirements
- 5-8 years of hands-on experience in ontology engineering, semantic data modeling, or knowledge graph development - with a demonstrable track record of production ontologies at scale
- Familiarity with industry content and identity schemas: EIDR, Schema.org VideoObject, DDEX, or equivalent
- Working knowledge of PySpark and Databricks for large-scale transformation pipelines
- You are the domain authority for how Samba represents and relates the entities that matter most to our business - and you ensure that representation is rigorous, scalable, and aligned with industry standards.
Nice to have
- Strong Python - production-quality, well-tested code
- Hands-on experience with entity resolution, record linkage, or deduplication at scale - mapping messy, multi-source real-world data to clean ontological representations
- Bachelor's degree required in Computer Science, Information Science, Computational Linguistics, Mathematics, or a related field
- Master's or PhD strongly preferred
- Strongly Preferred
- Hands-on experience with Amazon Neptune or Stardog - including data virtualization (Neptune Orion or Stardog Virtual Graphs) over data lake sources
- Experience designing aggregation and derivation logic that converts raw behavioral event data into durable, graph-resident derived attributes
- Domain knowledge in media, entertainment, or ad tech - TV viewership (ACR/STB), digital audience modeling (device graphs, identity resolution), or ad exposure data
Compensation
- $180k-$230k
This listing is sourced directly from Samba's careers page and normalized into a canonical job model.