
Lead Data Engineer – Sourcing
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in India.
• Design and oversee the complete extraction and ETL pipeline that transforms unstructured web data into a B2B dataset across 13 markets.
• Establish architectural benchmarks for extractor patterns, proxy strategies, LLM infrastructure, agentic escalation workflows, and cost-effective orchestration.
• Create frameworks and deploy complex extractors to counteract anti-bot measures, handle JS-heavy rendering, manage schema drift, and address low-quality structures.
• Manage extraction, normalization, deduplication, validation, loading, cost, performance, and reliability.
• Develop rule-based triage, LLM escalation, structured-output validation, retries, and human-review queues.
• Build production-level LLM infrastructure with versioned prompts, labeled evaluation sets, precision/recall metrics, rollback capabilities, tracing, and drift detection.
• Design evaluation frameworks, observability scaffolding, model-choice playbooks, and versioned SKILL.md specifications.
• Implement Airflow or equivalent orchestration patterns, manage dependencies, ensure recovery, and develop cost-aware scheduling.
• Provide technical insights to data platform, product, and analytics teams.
• Set standards for coverage, schema, and accuracy.
• Deliver around 80% hands-on architecture and reference implementation while contributing about 20% cross-functionally.
• Design greenfield sourcing infrastructure and take ownership of outcomes from start to finish.
• Over 6 years of experience in deploying production extraction, ETL, or data pipeline systems in business-critical settings.
• Expertise in large-scale web extraction, including anti-bot measures, proxy architecture, JS rendering, schema drift, and recovery processes.
• Proven experience in deploying LLMs within extraction pipelines as production systems.
• Familiarity with structured outputs, versioned prompts, labeled evaluation sets, logged traces, precision/recall judges, and vendor model drift detection.
• Experience in building agents and tool-calling pipelines.
• Skilled in architecting agent workflows and drafting SKILL.md specifications.
• Strong decision-making abilities in selecting rules versus LLMs.
• Proficient in Python, including production-grade and concurrency-aware development.
• Advanced SQL skills for complex transformations and performance optimization.
• Extensive experience with Airflow or similar tools.
• Familiarity with agentic IDEs such as Claude Code or Cursor.
• Demonstrated architectural judgment and product-oriented mindset.
• Hands-on experience with AI tools, structured outputs, evaluations, traces, and LLM tracing.
• Highly valued: proficiency in Snowflake, Redshift, AWS Lambda, S3, ECS, Glue.
• Highly valued: knowledge of vector databases, embeddings, and retrieval methods.
• Highly valued: experience with Braintrust, Promptfoo, Inspect, and LLM tracing tools.
• Highly valued: expertise in data quality frameworks with automated testing and anomaly detection.
• Highly valued: background in B2B data, including firmographics, people data, and company registries.
• Highly valued: understanding of GDPR and CCPA regulations.
• Highly valued: experience in startups or scale-ups.
• Not suitable for candidates with a solely operations or data analytics background lacking experience in large-scale systems engineering.
• Flexible working hours.
• Complete autonomy over architectural decisions.
CuraLinc Healthcare
VSP Vision Care
Adoreal
Get handpicked remote jobs straight to your inbox weekly.