MLabs

MLabs

Research Crawling Engineer

New York, New York, United States · Full-time

No sponsorship$100k-$130kDetected 22 days ago
JavaScriptPythonJavaGoRustDistributed SystemsMachine LearningData EngineeringNLPLLMsPlaywrightTCP/IPResearch

About the role

  • Additionally, the team has engineered sophisticated pipelines for the ingestion, segmentation, and annotation of billions of multimedia files, facilitating dataset creation for frontier research labs.
  • The organization operates as a lean, technical team that prioritizes speed and direct execution.
  • This role encompasses distributed systems, scraping infrastructure, and data pipelines, focusing on providing high-quality inputs for research and model development.

Responsibilities

  • Construct and maintain large-scale web crawlers across diverse domains.
  • Design high-throughput, fault-tolerant systems for data collection, managing volumes ranging from millions to billions of URLs per day.
  • Develop robust pipelines for data cleaning, deduplication, filtering, and normalization.
  • Collaborate with research teams to ensure data collection efforts align with modeling requirements.
  • Optimize infrastructure to ensure cost-efficiency, low latency, and reliability.
  • Proven experience in building web crawlers or large-scale data pipelines.
  • Demonstrated ability to debug and maintain systems within unstable or adversarial environments.

Nice to have

  • Familiarity with LLM pre-training data or retrieval systems.
  • Practical experience with headless browsers (e.g., Playwright, Puppeteer, or Chrome DevTools Protocol).
  • Knowledge of proxy systems, IP rotation, and large-scale request orchestration.
  • Background in data quality evaluation or benchmarking.
  • Experience running workloads on cloud or bare-metal infrastructure.
  • Impactful Opportunity: Contribute to the development of a web-scale crawler and knowledge graph at the forefront of AI data accessibility.
  • High-Performance Culture: Join a lean, low-ego team that prioritizes high output and professional growth.

Compensation

  • $100K - $130K We are hiring on behalf of our client who is a technical infrastructure firm specializing in the delivery of massive-scale web data to organizations developing advanced artificial intelligence models.
  • The organization supports high-capacity bandwidth-sharing networks and operates a distributed crawler capable of accessing high-quality public web data at a global scale.
  • Additionally, the team has engineered sophisticated pipelines for the ingestion, segmentation, and annotation of billions of multimedia files, facilitating dataset creation for frontier research labs.
  • The organization operates as a lean, technical team that prioritizes speed and direct execution.
  • As a Research Crawling Engineer, the successful candidate will design and operate large-scale web data acquisition systems.
  • This role encompasses distributed systems, scraping infrastructure, and data pipelines, focusing on providing high-quality inputs for research and model development.

Benefits

  • Build and maintain datasets specifically structured for research and machine learning model training.
  • Monitor and optimize crawl performance, coverage, and data quality through rapid iteration.

Company info

  • Join a lean, low-ego team that prioritizes high output and professional growth.
  • At MLabs, we are committed to offer equal opportunities to all candidates.
  • We ensure no discrimination, accessible job adverts, and providing information in accessible formats.
  • Our goal is to foster a diverse, inclusive workplace with equal opportunities for all.
  • If you need any reasonable adjustments during any part of the hiring process or you would like to see the job-advert in an accessible format please let us know at the earliest opportunity by emailing human-resources@mlabs.city.
  • MLabs Ltd collects and processes the personal information you provide such as your contact details, work history, resume, and other relevant data for recruitment purposes only.
  • Your data may be shared only with clients and trusted partners where necessary for recruitment purposes.
  • You may request the deletion of your data or withdraw your consent at any time by contacting legal@mlabs.city.
  • Commitment to Equality and Accessibility: At MLabs, we are committed to offer equal opportunities to all candidates.

This listing is sourced directly from MLabs's careers page and normalized into a canonical job model.