
Senior Data Engineer – Data Platform
Posted Sep 2

Posted Sep 2
This is a fully remote position, open to applicants in India, +1 more country.
• Design the transformation and warehouse layer to convert billions of raw records into a precise B2B dataset.
• Create production architecture for LLM-driven extraction, enrichment, entity resolution, and semantic validation.
• Specify structured outputs, retry mechanisms, human-review fallbacks, and reusable abstractions.
• Develop labeled evaluation sets, scoring systems, judge calibration, and regression suites.
• Set precision and recall targets for quality assurance checks.
• Decide when to implement deterministic rules in contrast to LLM-based checks.
• Oversee token budgets, model routing, vendor drift detection, and cost limitations.
• Establish observability logging for prompt versions, models, costs, latency, and decisions for each LLM call.
• Design dbt transformation layers and strategies for Snowflake performance, clustering, materialization, and cost management.
• Define Airflow orchestration, recovery processes, and cost-aware scheduling patterns.
• Construct AWS infrastructure as code.
• Develop embedding and retrieval strategies for company and people entity resolution across 13 markets.
• Offer technical platform guidance in architectural decisions alongside sourcing, data quality, product, and analytics teams.
• Write reference implementations for the team to utilize.
• 5+ years of experience in building production data platforms within business-critical settings, possessing end-to-end architecture ownership.
• Proven experience managing billions of rows of data.
• Production experience deploying LLMs within data pipelines for extraction, enrichment, or validation purposes.
• Familiarity with structured outputs, versioned prompts, and labeled evaluation sets.
• Experience in designing evaluation harnesses, including scoring systems, labeled sets, judge calibration, and regression suites.
• Strong understanding of when to employ rules versus LLMs.
• Proficient in Python and SQL.
• Advanced skills in dbt and Snowflake, covering warehouse design, query optimization, clustering, RBAC, and large-scale cost management.
• Extensive experience with Airflow in production, including orchestration, dependency management, recovery strategies, and cost optimization.
• Solid knowledge of AWS services such as S3, Lambda, Glue, ECS, and RDS.
• Regular use of AI coding tools like Claude Code, Cursor, or similar, with demonstrable work experience.
• Strong architectural judgment and a product-oriented mindset.
• Highly regarded skills include Braintrust, Promptfoo, Inspect, Logfire, OpenTelemetry, embeddings, vector search, fuzzy matching, fine-tuning, distilling small models, dbt Cloud, CI/CD, Spark/PySpark, Kafka, Kinesis, Snowpipe Streaming, B2B data, GDPR, SOC2, and CCPA.
• Competitive base salary.
• Meaningful equity.
• Flexible work arrangements / fully remote work options.
• No fixed working hours.
• Small senior teams with minimal processes.
• Weekly releases progressing towards daily updates.
• Equal opportunity workplace.
CuraLinc Healthcare
VSP Vision Care
Adoreal
Get handpicked remote jobs straight to your inbox weekly.