MLabs
Research Crawling Engineer
New York, New York, United States · Full-time
No sponsorship$100k-$130kDetected 22 days ago
JavaScriptPythonJavaGoRustDistributed SystemsMachine LearningData EngineeringNLPLLMsPlaywrightTCP/IPResearch
About the role
- Additionally, the team has engineered sophisticated pipelines for the ingestion, segmentation, and annotation of billions of multimedia files, facilitating dataset creation for frontier research labs.
- The organization operates as a lean, technical team that prioritizes speed and direct execution.
- This role encompasses distributed systems, scraping infrastructure, and data pipelines, focusing on providing high-quality inputs for research and model development.
Responsibilities
- Construct and maintain large-scale web crawlers across diverse domains.
- Design high-throughput, fault-tolerant systems for data collection, managing volumes ranging from millions to billions of URLs per day.
- Develop robust pipelines for data cleaning, deduplication, filtering, and normalization.
- Collaborate with research teams to ensure data collection efforts align with modeling requirements.
- Optimize infrastructure to ensure cost-efficiency, low latency, and reliability.
- Proven experience in building web crawlers or large-scale data pipelines.
- Demonstrated ability to debug and maintain systems within unstable or adversarial environments.
Nice to have
- Familiarity with LLM pre-training data or retrieval systems.
- Practical experience with headless browsers (e.g., Playwright, Puppeteer, or Chrome DevTools Protocol).
- Knowledge of proxy systems, IP rotation, and large-scale request orchestration.
- Background in data quality evaluation or benchmarking.
- Experience running workloads on cloud or bare-metal infrastructure.
- Impactful Opportunity: Contribute to the development of a web-scale crawler and knowledge graph at the forefront of AI data accessibility.
- High-Performance Culture: Join a lean, low-ego team that prioritizes high output and professional growth.
Compensation
- $100K - $130K We are hiring on behalf of our client who is a technical infrastructure firm specializing in the delivery of massive-scale web data to organizations developing advanced artificial intelligence models.
- The organization supports high-capacity bandwidth-sharing networks and operates a distributed crawler capable of accessing high-quality public web data at a global scale.
- Additionally, the team has engineered sophisticated pipelines for the ingestion, segmentation, and annotation of billions of multimedia files, facilitating dataset creation for frontier research labs.
- The organization operates as a lean, technical team that prioritizes speed and direct execution.
- As a Research Crawling Engineer, the successful candidate will design and operate large-scale web data acquisition systems.
- This role encompasses distributed systems, scraping infrastructure, and data pipelines, focusing on providing high-quality inputs for research and model development.
Benefits
- Build and maintain datasets specifically structured for research and machine learning model training.
- Monitor and optimize crawl performance, coverage, and data quality through rapid iteration.
Company info
- Join a lean, low-ego team that prioritizes high output and professional growth.
- At MLabs, we are committed to offer equal opportunities to all candidates.
- We ensure no discrimination, accessible job adverts, and providing information in accessible formats.
- Our goal is to foster a diverse, inclusive workplace with equal opportunities for all.
- If you need any reasonable adjustments during any part of the hiring process or you would like to see the job-advert in an accessible format please let us know at the earliest opportunity by emailing human-resources@mlabs.city.
- MLabs Ltd collects and processes the personal information you provide such as your contact details, work history, resume, and other relevant data for recruitment purposes only.
- Your data may be shared only with clients and trusted partners where necessary for recruitment purposes.
- You may request the deletion of your data or withdraw your consent at any time by contacting legal@mlabs.city.
- Commitment to Equality and Accessibility: At MLabs, we are committed to offer equal opportunities to all candidates.
This listing is sourced directly from MLabs's careers page and normalized into a canonical job model.