
Senior Data Infrastructure Engineer – Scraping, Scale
Posted 17 hours ago

Posted 17 hours ago
This is a fully remote position, open to applicants in Sweden.
• Develop and sustain the fundamental ingestion engine that drives the social media data lab.
• Create and oversee stealth scraping clusters utilizing residential proxy networks, TLS fingerprinting, headful/headless browser farms, and session rotation.
• Construct fault-tolerant and scalable web-scraping pipelines for platforms like Instagram, TikTok, YouTube, X, and various web sources.
• Design distributed queues and workflow engines to handle millions of asynchronous scraping tasks each day.
• Architect both structured and unstructured storage environments for downstream AI modeling purposes.
• Implement automated notifications for platform UI/API changes, blocking patterns, and proxy failures.
• Deliver millions of profile, post, and video records daily with minimal downtime.
• A minimum of 4 years of experience in high-volume web scraping, data engineering, or reverse engineering.
• Extensive experience in overcoming advanced anti-bot solutions (Cloudflare, Akamai, PerimeterX) through TLS impersonation, browser automation, and proxy management.
• Proficiency in Python or Go.
• In-depth knowledge of Playwright, Puppeteer, Scrapy, or Selenium.
• Familiarity with distributed systems and task queues (Temporal, Ray, Kafka, Redis).
• Strong SQL proficiency.
• Experience with contemporary analytical data lakes or data warehouses.
• Preferred: Direct experience in extracting short-form video content and user metadata from TikTok, Instagram, and YouTube.
• Preferred: Experience in integrating scraping outputs directly into vector databases and real-time AI processing queues.
• Opportunity for remote work.
ELEKS
Get handpicked remote jobs straight to your inbox weekly.