Poolside
Member of Engineering (Synthetic Data Research)
Remote (US)
Sponsorship not specifiedDetected 174 days ago
PythonMachine LearningDeep LearningData EngineeringLLMsAgentic AICadenceResearchCollaboration
About the role
- You'll be working on our data team focused on the quality of the datasets being delivered for training our models.
- This is a hands-on role where your #1 mission would be to improve the quality of our datasets across the entire training cycle (pre-training, mid-training, post-training, RL) by leveraging your previous experience, intuition and training experiments.
- This role particularly focuses on generating synthetic data at scale and determining the best strategies to leverage such data into training large models.
Responsibilities
- Design and implement complex pipelines that can generate large amounts of data while maintaining high diversity and optimizing the resources available.
- Experience in building trillion-scale pretraining datasets, and familiarity with concepts like data curation, deduplication, data mixing, tokenization, curriculum, impact of data repetition, etc.
Requirements
- Their ability to stack advantages and pull ahead will define the winners.
- Experience with Large Language Models (LLM), including:
- Experience with implementing cost-efficient, complex pipelines to generate synthetical datasets at scale optimizing for data quality, correctness, diversity, etc.
Nice to have
- Can freely discuss the latest papers and descend to fine details
- Is reasonably opinionated
- Intro call with one of our Founding Engineers
- Technical Interview(s) with one of our Members of Engineering
- Team fit call with the People team
- Final interview with one of our Founding Engineers
- Company-provided equipment
- Frequent team get togethers
Skills
- Strong machine learning and engineering background
- Understanding of how LLMs learn
- Data ablations and scaling laws
- Post-training techniques
- Training reasoning and agentic models
- Experience with evals tracking model capabilities (general knowledge, reasoning, math, coding, long-context, etc)
- Excellent programming skills in Python
- Strong prompt engineering skills
- Experience working with large-scale GPU clusters and distributed data pipelines
- Strong obsession with data quality
Benefits
- 37 days/year of vacation & holidays
- Health insurance allowance for you & dependents
- 16 weeks of flexible, full-pay parental leave
- Well-being, always-be-learning & home office allowances
- Fully remote work & flexible hours
Company info
- to build a world where AI will be the engine behind economically valuable work and scientific progress.
- We believe the fastest way to reach AGI lies in accelerating software development itself, by reshaping the developer experience with agentic systems, coding assistants, and the frontier models that power them.
- We deploy these systems directly into the development environments of security-conscious enterprises.
- We were founded in the US and have our home there, but our team is distributed across Europe and North America.
- We get our fix of in-person collaboration in Paris each month for 3 days, with an open invitation to stay the whole week.
- For those based in PST, we understand this is a significant travel cadence; we are open to agree on a lower cadence and will discuss this in the interview process.
- We also do longer off-sites once a year.
- Our team is a multidisciplinary blend of research, engineering, and business experts.
- What unites us is our deep care for what we build together.
- We're in a race that requires hard work, intellectual curiosity, and obsession; to balance this intensity, we've assembled a team of low ego and kind-hearted individuals who have built the special culture Poolside has.
- By building collaboratively and with intention, we create a compounding effect that moves the entire company forward towards our mission: reaching AGI through intelligence systems built for software development.
- You'll closely collaborate with other teams like Pre-training, Pre-training data, Post-training, RL2L, Evals, and Product to define high-quality data needs that map to missing model capabilities and downstream use cases.
- Staying in sync with the latest state-of-the-art research in synthetic data generation and LLM training is key to success in this role.
- You will constantly lead original research initiatives through short, time-bounded experiments while deploying highly technical engineering solutions into production.
- With the volumes of data to process being massive, you'll have a performant distributed data pipeline together with large GPU clusters at your disposal.
- Curious about the tech?
- Take a deep dive into our data work in our Laguna M.1/XS.2 Technical Report. https://arxiv.org/abs/2605.27605
- To deliver large, high-quality, and diverse synthetic datasets mixing natural language and code modalities to train best-in-class Poolside coding agents.
- Follow the latest research related to LLMs and synthetic data generation in particular. Be familiar with the most relevant open-source datasets and models.
- Closely with cross-team to ensure the experiments run and data generated is the most efficient use of compute and time resources for the improvements in quality of our models.
- Continuously measure and refine the quality of the datasets being generated while validating the final data strategy through quantitative data ablation experiments.
Apply directly at Poolside →Create a free account for alerts like thisView Poolside immigration profile
This listing is sourced directly from Poolside's careers page and normalized into a canonical job model.