Sciforium

Sciforium

LLM Dataset Engineer

San Francisco

Sponsorship not specifiedDetected 196 days ago
PythonFull-Stack DevelopmentMachine LearningSparkData ScienceComputer VisionLLMsStatisticsResearch

About the role

  • We believe that in the era of LLMs, data is the primary competitive advantage.
  • This position is ideal for a scientist who views data as a high-scale engineering challenge and an analytical puzzle.

Responsibilities

  • Foundation Dataset Strategy: Own the end-to-end creation of pre-training datasets for LLMs. This includes defining the mix of web data, code, books, and technical papers to optimize for downstream model performance.
  • Petabyte-Scale Curation: Design and implement sophisticated pipelines for data cleaning, exact/fuzzy deduplication, and high-quality signal extraction from petabytes of raw, unstructured data.
  • Post-Training & Alignment Data: Lead the development of high-quality post-training datasets, including Supervised Fine-Tuning (SFT) instructions, multi-turn dialogues, and preference modeling data (RLHF/DPO).
  • High-Performance Engineering: Develop high-throughput data processing scripts using Python, leveraging multiprocessing and multithreading to handle massive-scale ingestion and transformation without bottlenecks.
  • Synthetic Data Generation: (Added Value) Design pipelines to generate high-reasoning synthetic data to augment gaps in natural datasets, utilizing existing models for data labeling and refinement.
  • Dataset Reconstruction: Experience building massive LLM training sets from scratch, including raw web crawls (e.g., Common Crawl) and specialized domain data.
  • Post-Training Expertise: Hands-on experience building datasets for RLHF, DPO, and multi-turn instruction following, including the management of human-labeling workflows and quality gold-sets.
  • Foundation Dataset Strategy: Own the end-to-end creation of pre-training datasets for LLMs.
  • Experience building large-scale image or video datasets from scratch (e.g., LAION-style pipelines).

Requirements

  • Familiarity with large-scale crawling of multimodal data and the associated challenges of video processing, codecs, and compression.
  • Deep Proficiency in Python: Expert-level skills with a focus on high-performance code, including multiprocessing, multithreading, and efficient memory management for large-scale data tasks.

Skills

  • Mastery of data-at-scale frameworks such as Spark, Ray, or high-performance data-loading formats (e.g., WebDataset, Parquet).

Compensation

  • Competitive salary and equity

Benefits

  • Medical, dental, and vision insurance
  • Flexible time off
  • Competitive salary and equity
  • Drive the acquisition and processing of vision and video data, navigating the complexities of multimodal alignment, video compression, and temporal data consistency.
  • Demonstrated experience working with petabyte-scale datasets that have been directly used to train production-grade LLMs or Large Vision Models.
  • 5+ years of industry experience in Data Science or Machine Learning, with a proven track record of building and managing datasets for foundation models.

Equal opportunity

  • Sciforium is an equal opportunity employer.
  • All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.

Visa & Work Authorization

  • Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.

This listing is sourced directly from Sciforium's careers page and normalized into a canonical job model.