
AI Web Scraping Engineer – Stealth Data Acquisition
Posted Sep 12

Posted Sep 12
This is a fully remote position, open to applicants in United States.
• Develop, construct, and sustain stealth-oriented scraping pipelines that successfully bypass bot detection, fingerprinting, and rate-limiting protections.
• Analyze and reverse-engineer the anti-automation strategies of targeted websites, such as TLS/JA3 fingerprinting, canvas/WebGL fingerprinting, and behavioral detection, while implementing countermeasures.
• Design distributed scraping architectures utilizing rotating residential/mobile proxies, headless browser farms, and large-scale session/cookie management.
• Incorporate AI/LLM-driven parsing for dynamic, unstructured, or rapidly changing web page layouts.
• Create self-adapting scrapers that recognize layout/DOM alterations and adjust automatically.
• Oversee scraper performance, success metrics, and detection signals; respond swiftly when websites update their defenses.
• Ensure that data pipelines adhere to client-specified legal and ethical standards, including Terms of Service reviews, robots.txt compliance, and rate limiting as instructed by the legal/compliance team.
• Collaborate with data engineering teams to provide clean, structured, and deduplicated datasets for downstream systems.
• Minimum of 4 years of experience in constructing production web scraping systems at scale.
• Extensive hands-on expertise with headless browser automation tools (such as Playwright, Puppeteer, Selenium), including stealth plugins and patches.
• Demonstrated experience in overcoming or circumventing Cloudflare, DataDome, PerimeterX, Akamai Bot Manager, Kasada, or comparable platforms.
• Solid understanding of TLS fingerprinting, HTTP/2 fingerprinting, canvas/WebGL fingerprinting, and browser fingerprint spoofing techniques.
• Experience with large-scale residential/mobile/datacenter proxy rotation and session management.
• Proficient in Python and/or Node.js/TypeScript for scraping and orchestration purposes.
• Familiarity with using LLMs/AI models for content extraction, parsing, and classification of messy or unstructured HTML.
• Knowledge of CAPTCHA-solving methods and strategies to avoid detection.
• Must reside in and be authorized to work in the United States.
• Successful completion of a standard background check is required prior to employment.
• Comfortable in a dynamic, fast-paced environment where target sites frequently modify their defenses.
• Nice to have: Experience with reverse engineering mobile app APIs.
• Nice to have: Background in adversarial machine learning or anti-bot research.
• Nice to have: Familiarity with distributed task queues (such as Celery, Kafka, SQS) for orchestrating scraping tasks.
• Nice to have: Previous experience in a data-as-a-service, alternative data, or web intelligence organization.
• Compensation offered in USD.
• Opportunity for fully remote work.
• Potential for career advancement within an international company.
RecruityTalent
Virtual Identity
Grupo Odilon Santos
Egnyte
Get handpicked remote jobs straight to your inbox weekly.