
Lead Data Engineer
Posted 10 hours ago

Posted 10 hours ago
This is a fully remote position, open to applicants in India.
• Design and take ownership of the complete extraction and ETL pipeline that converts unstructured web data into a B2B dataset spanning 13 markets.
• Establish architectural standards for extraction patterns, proxy strategies, LLM infrastructure, workflows for agentic escalation, coverage, schema, and precision.
• Create processes for extraction, normalization, deduplication, validation, and loading of data.
• Manage aspects related to cost, performance, reliability, scheduling, incremental processing, and recovery design.
• Develop extractor frameworks and implement challenging extractors that can navigate anti-bot defenses, JavaScript-heavy rendering, schema drift, and low-quality structures.
• Construct agentic extraction pipelines featuring rule-based triage, LLM escalation, structured-output validation, retries, and human review queues.
• Set up production LLM infrastructure with versioned prompts, labeled evaluation sets, precision/recall measurement, rollback capabilities, traces, and drift detection.
• Build evaluation and observability frameworks along with model-choice playbooks.
• Generate versioned SKILL.md specifications and orchestration patterns utilizing Airflow or a similar tool.
• Offer technical insights to data platform, product, and analytics teams.
• Allocate approximately 80% of time to hands-on architecture and reference implementation, with 20% dedicated to cross-functional technical input.
• 6–10 years of experience in delivering production extraction, ETL, or data pipeline systems within business-critical environments.
• Proficiency in large-scale web extraction, including anti-bot defenses, proxy architecture, JavaScript rendering, schema drift, and recovery processes.
• Experience with production LLMs in extraction pipelines, focusing on structured outputs, versioned prompts, labeled evaluation sets, logged traces, precision/recall assessments, and vendor-model drift detection.
• Familiarity with production agent workflows and tool-calling pipelines.
• Capability to write SKILL.md specifications.
• Strong decision-making skills concerning deterministic rules versus LLMs.
• Expertise in Python, including concurrency and scaling.
• Advanced SQL skills for complex transformations and performance enhancement.
• Significant experience with Airflow or equivalent tools.
• Experience deploying with agentic IDEs such as Claude Code or Cursor.
• Strong architectural judgment and a product-oriented mindset.
• Current hands-on experience with AI tools, structured outputs, evaluations, traces, and LLM tracing.
• Experience with cloud data platforms like Snowflake or Redshift is highly valued.
• AWS pipeline deployment experience with Lambda, S3, ECS, or Glue is highly regarded.
• Familiarity with vector databases, embeddings, or retrieval patterns is highly valued.
• Experience with evaluation frameworks such as Braintrust, Promptfoo, or Inspect is highly valued.
• Knowledge of data-quality frameworks that include automated testing and anomaly detection is highly valued.
• B2B data experience is highly valued.
• Understanding of data privacy and compliance, including GDPR and CCPA, is highly valued.
• Experience in startup or scaleup environments is highly valued.
• Flexible working hours.
• Complete autonomy in architectural decisions.
Solar Coca-Cola
UJET
Get handpicked remote jobs straight to your inbox weekly.