Two Six Technologies
Data Collection Engineer
Remote, USA
No sponsorship$160k-$200kDetected 8 days ago
JavaScriptPythonHTMLDistributed SystemsData StructuresGitSQLRedisElasticsearchAWSCloud PlatformsDockerKubernetesCI/CDGitHub ActionsJenkinsDevOpsKafkaData EngineeringComputer VisionLLMsSeleniumPlaywrightTest Automation
About the role
- If you thrive on reverse-engineering web applications, overcoming anti-bot barriers, and orchestrating distributed systems, we want you on our team.
- We believe in rewarding skills, experience, and performance.
- The projected salary range listed for this position is annualized.
Responsibilities
- Distributed Crawler Development: Design and deploy high-performance, distributed web scrapers using Python and Scrapy to extract massive datasets efficiently.
- Infrastructure & Container Orchestration: Deploy, scale, and manage scraping workloads on Kubernetes, ensuring optimal resource allocation and fault tolerance.
- Data Validation & Quality Assurance: Define strict JSON Schemas and leverage Pydantic to enforce data types, validate incoming payloads, and catch data drift early.
- Data Ingestion & Storage: Build and optimize search and storage pipelines using Elasticsearch, transforming raw web dumps into highly structured, searchable data.
- Pipeline Workflow Management: Architect robust pipeline workflows to manage the end-to-end data lifecycle-from discovery and extraction to validation and storage.
- Anti-Bot & Proxy Engineering: Manage complex proxy rotation, session handling, and browser fingerprinting to maintain high success rates against advanced anti-scraping systems.
- Reliability Focus: A strong commitment to data integrity, system monitoring, and building self-healing scraping systems.
- Workflow Management: Experience building structured pipeline workflows to handle complex, multi-stage data extraction tasks.
Requirements
- Excellent reverse-engineering skills, with the ability to dissect network traffic, unearth hidden APIs, and bypass complex web barriers.
- Expert-level proficiency in Python.
- Deep experience with Scrapy and distributed scraping architectures (e.g., handling distributed queues, broad vs. deep crawling).
- Proven experience with browser automation tools (Playwright, Selenium, or Puppeteer).
- Strong proficiency in managing and scaling applications within Kubernetes environments.
- Bachelor's degree in Computer Science, Engineering
- Strong hands-on experience with AWS ecosystems (e.g., EKS, EC2, S3, RDS).
- Proficiency in SQL for querying, schema design, and storing structured relational data.
- Experience with Redis (specifically for caching, deduplication, or as a Scrapy distributed queue back-end).
- Familiarity with Apache Kafka for real-time data streaming and decoupled pipeline architectures.
Skills
- Utilize Browser Scripting tools to navigate, interact with, and extract data from modern, dynamic, and JavaScript-heavy websites.
Compensation
- $160,000 - $200,000 USD
Benefits
- Experience leveraging LLMs or Computer Vision for adaptive scraping, parsing unstructured HTML, or bypassing CAPTCHAs ( AI in data collection ).
- Education: Bachelor's degree in Computer Science, Engineering
- application process, information about our rich benefits and perks along with our most frequently asked questions.
Equal opportunity
- For more information review the Two Six Technologies Equal Employment Opportunity and Affirmative Action Policy and the EEO Poster.
- If you are an individual with a disability and would like to request reasonable workplace accommodation for any part of our employment process, please send an email to accommodations@twosixtech.com.
Apply directly at Two Six Technologies →Create a free account for alerts like thisView Two Six Technologies immigration profile
This listing is sourced directly from Two Six Technologies's careers page and normalized into a canonical job model.