
Senior Data Product Engineer
Posted Sep 3

Posted Sep 3
This is a fully remote position, open to applicants in Latin America.
β’ Design and scale distributed web scraping architectures to gather data from thousands of online sources.
β’ Create strategies to navigate obstacles such as Cloudflare, CAPTCHAs, rate limiting, IP banning, browser fingerprinting, and other anti-scraping technologies.
β’ Implement proxy pools, headless browsers, and adaptive crawling methods for effective large-scale data collection.
β’ Process and extract information from HTML and JSON formats.
β’ Refactor and optimize Python data collection pipelines for production use.
β’ Enhance reliability, observability, error management, retries, and alerting mechanisms.
β’ Develop schedulable and containerized data ingestion workflows utilizing Airflow or similar tools.
β’ Design PostgreSQL schemas, views, partitioning, and data access methodologies.
β’ Utilize AWS services including S3, EKS, IAM/IRSA, Parameter Store, and ECR.
β’ Construct and maintain the data access layer through GraphQL and Hasura.
β’ Collaborate with Data Science and SaaS teams to incorporate machine learning components and establish reliable API contracts.
β’ Take ownership of the data product from start to finish, participate in code reviews, and promote technical enhancements.
β’ Extensive hands-on experience in large-scale web scraping.
β’ Proven track record in overcoming anti-scraping and bot-protection measures, including proxy rotation, headless browsers, rate limiting, IP blocking, fingerprinting, or comparable challenges.
β’ Strong proficiency with web scraping and DOM parsing tools such as Scrapy, Playwright, Selenium, Puppeteer, BeautifulSoup, lxml, XPath, or CSS selectors.
β’ Significant professional background in Python and data processing techniques.
β’ Practical experience with AWS, especially S3 and boto3; familiarity with EKS, IAM/IRSA, Parameter Store, or ECR is highly advantageous.
β’ Experience in constructing dependable data pipelines using Airflow or an equivalent orchestrator.
β’ Solid understanding of PostgreSQL and the principles of relational databases.
β’ Capability to take complete ownership of technical solutions and operate independently.
β’ Experience in contributing to code reviews and upholding high software engineering standards.
β’ Familiarity with GraphQL, Hasura, Redis, Elasticsearch, Docker, or Kubernetes is a plus.
β’ Background in cybersecurity data, NLP/ML pipelines, or data-as-a-product environments is a plus.
β’ Proficiency in English at B2 level or higher, with the ability to communicate effectively with technical and cross-functional teams.
β’ Fully flexible and self-managed vacation time.
β’ Sick leave and personal days off.
β’ Observance of public holidays.
β’ Paternity and maternity leave available.
β’ Study leave provisions.
β’ Designated moving days.
β’ Training in best practices and technology.
β’ Access to books and informal discussions.
β’ In-house English language courses.
β’ Continuous feedback opportunities.
β’ One-on-one career development sessions.
β’ Flexible working hours.
β’ Provision of equipment and work materials.
β’ Internal events and team-building activities.
β’ A day off to celebrate your birthday.
β’ Fully remote positions available across LATAM.
β’ Contractor model with payment in USD.
GE Vernova
Zerion
WattFox
Probis
Get handpicked remote jobs straight to your inbox weekly.