AI Web Scraping Engineer – Stealth Data Acquisition

Posted Sep 12

This is a fully remote position, open to applicants in United States.

📋 Description

• Develop, construct, and sustain stealth-oriented scraping pipelines that successfully bypass bot detection, fingerprinting, and rate-limiting protections.

• Analyze and reverse-engineer the anti-automation strategies of targeted websites, such as TLS/JA3 fingerprinting, canvas/WebGL fingerprinting, and behavioral detection, while implementing countermeasures.

• Design distributed scraping architectures utilizing rotating residential/mobile proxies, headless browser farms, and large-scale session/cookie management.

• Incorporate AI/LLM-driven parsing for dynamic, unstructured, or rapidly changing web page layouts.

• Create self-adapting scrapers that recognize layout/DOM alterations and adjust automatically.

• Oversee scraper performance, success metrics, and detection signals; respond swiftly when websites update their defenses.

• Ensure that data pipelines adhere to client-specified legal and ethical standards, including Terms of Service reviews, robots.txt compliance, and rate limiting as instructed by the legal/compliance team.

• Collaborate with data engineering teams to provide clean, structured, and deduplicated datasets for downstream systems.


⛳️ Requirements

• Minimum of 4 years of experience in constructing production web scraping systems at scale.

• Extensive hands-on expertise with headless browser automation tools (such as Playwright, Puppeteer, Selenium), including stealth plugins and patches.

• Demonstrated experience in overcoming or circumventing Cloudflare, DataDome, PerimeterX, Akamai Bot Manager, Kasada, or comparable platforms.

• Solid understanding of TLS fingerprinting, HTTP/2 fingerprinting, canvas/WebGL fingerprinting, and browser fingerprint spoofing techniques.

• Experience with large-scale residential/mobile/datacenter proxy rotation and session management.

• Proficient in Python and/or Node.js/TypeScript for scraping and orchestration purposes.

• Familiarity with using LLMs/AI models for content extraction, parsing, and classification of messy or unstructured HTML.

• Knowledge of CAPTCHA-solving methods and strategies to avoid detection.

• Must reside in and be authorized to work in the United States.

• Successful completion of a standard background check is required prior to employment.

• Comfortable in a dynamic, fast-paced environment where target sites frequently modify their defenses.

• Nice to have: Experience with reverse engineering mobile app APIs.

• Nice to have: Background in adversarial machine learning or anti-bot research.

• Nice to have: Familiarity with distributed task queues (such as Celery, Kafka, SQS) for orchestrating scraping tasks.

• Nice to have: Previous experience in a data-as-a-service, alternative data, or web intelligence organization.


🏝️ Benefits

• Compensation offered in USD.

• Opportunity for fully remote work.

• Potential for career advancement within an international company.

People also viewed

RecruityTalent1 day ago

AI Forward Deployed Engineer, German

BG flagBulgaria OnlyFull-timeArtificial Intelligence
ApplyView job
Virtual Identity1 day ago

AI Data Governance Consultant – m/f/x

PT flagPortugal, +2 more countriesFull-timeArtificial Intelligence
ApplyView job
Grupo Odilon Santos1 day ago

AI Specialist

BR flagBrazil OnlyFreelanceArtificial IntelligenceR$0 – R$18k/month
ApplyView job
Egnyte1 day ago

Senior Consultant, AI Services

US flagUnited States OnlyFull-timeArtificial Intelligence$125k – $140k/year
ApplyView job
Shift Technology1 day ago

Artificial Intelligence Researcher

FR flagFrance OnlyFull-timeArtificial Intelligence
ApplyView job
Foxelli Group1 day ago

Senior AI Filmmaker

LT flagLithuania OnlyFull-timeArtificial Intelligence€2,500/month
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers