Sciforium
LLM Dataset Engineer
San Francisco
Sponsorship not specifiedDetected 196 days ago
PythonFull-Stack DevelopmentMachine LearningSparkData ScienceComputer VisionLLMsStatisticsResearch
About the role
- We believe that in the era of LLMs, data is the primary competitive advantage.
- This position is ideal for a scientist who views data as a high-scale engineering challenge and an analytical puzzle.
Responsibilities
- Foundation Dataset Strategy: Own the end-to-end creation of pre-training datasets for LLMs. This includes defining the mix of web data, code, books, and technical papers to optimize for downstream model performance.
- Petabyte-Scale Curation: Design and implement sophisticated pipelines for data cleaning, exact/fuzzy deduplication, and high-quality signal extraction from petabytes of raw, unstructured data.
- Post-Training & Alignment Data: Lead the development of high-quality post-training datasets, including Supervised Fine-Tuning (SFT) instructions, multi-turn dialogues, and preference modeling data (RLHF/DPO).
- High-Performance Engineering: Develop high-throughput data processing scripts using Python, leveraging multiprocessing and multithreading to handle massive-scale ingestion and transformation without bottlenecks.
- Synthetic Data Generation: (Added Value) Design pipelines to generate high-reasoning synthetic data to augment gaps in natural datasets, utilizing existing models for data labeling and refinement.
- Dataset Reconstruction: Experience building massive LLM training sets from scratch, including raw web crawls (e.g., Common Crawl) and specialized domain data.
- Post-Training Expertise: Hands-on experience building datasets for RLHF, DPO, and multi-turn instruction following, including the management of human-labeling workflows and quality gold-sets.
- Foundation Dataset Strategy: Own the end-to-end creation of pre-training datasets for LLMs.
- Experience building large-scale image or video datasets from scratch (e.g., LAION-style pipelines).
Requirements
- Familiarity with large-scale crawling of multimodal data and the associated challenges of video processing, codecs, and compression.
- Deep Proficiency in Python: Expert-level skills with a focus on high-performance code, including multiprocessing, multithreading, and efficient memory management for large-scale data tasks.
Skills
- Mastery of data-at-scale frameworks such as Spark, Ray, or high-performance data-loading formats (e.g., WebDataset, Parquet).
Compensation
- Competitive salary and equity
Benefits
- Medical, dental, and vision insurance
- Flexible time off
- Competitive salary and equity
- Drive the acquisition and processing of vision and video data, navigating the complexities of multimodal alignment, video compression, and temporal data consistency.
- Demonstrated experience working with petabyte-scale datasets that have been directly used to train production-grade LLMs or Large Vision Models.
- 5+ years of industry experience in Data Science or Machine Learning, with a proven track record of building and managing datasets for foundation models.
Equal opportunity
- Sciforium is an equal opportunity employer.
- All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.
Visa & Work Authorization
- Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.
Apply directly at Sciforium →Create a free account for alerts like thisView Sciforium immigration profile
This listing is sourced directly from Sciforium's careers page and normalized into a canonical job model.