
Data Engineer – Data Platform
Posted Sep 3

Posted Sep 3
This is a fully remote position, open to applicants in India, +1 more country.
• Develop pipelines that convert billions of unprocessed records into clean, structured datasets.
• Create LLM-in-the-pipeline systems for tasks such as extraction, enrichment, entity resolution, and semantic validation.
• Execute structured outputs, retries, and human review fallbacks on millions of records daily.
• Construct labeled evaluation sets, scoring mechanisms, and regression suites to assess prompt and model modifications.
• Monitor precision and recall for each evaluation check.
• Implement deterministic checks like dbt tests, data contracts, and SQL when necessary.
• Oversee token budgets, model routing, and vendor-model drift detection.
• Record every LLM call, including prompt version, model, cost, latency, and decision.
• Maintain standard pipeline alerting and observability practices.
• Manage Airflow orchestration, dbt models from staging to mart, Snowflake performance and costs, along with AWS infrastructure as code.
• Implement embeddings and retrieval patterns for company and individual entity resolution across 13 markets.
• Ensure data reliability, cost management, and downstream consumer trust end to end.
• A minimum of 3 years in data engineering or data infrastructure.
• Experience managing production pipelines from start to finish.
• Proficient in Python and SQL.
• Expertise in production-grade, performance-oriented engineering at a very large scale.
• Familiarity with deploying LLMs within data pipelines for extraction, enrichment, or validation tasks.
• Knowledge of structured LLM outputs and labeled evaluation sets.
• Experience in building or maintaining LLM evaluation or testing frameworks.
• Ability to determine when to apply deterministic rules versus LLMs.
• Practical experience with dbt and Airflow in production environments.
• Cloud data warehouse experience, with a preference for Snowflake.
• Proficiency in schema design, query optimization, and cost management.
• Daily utilization of AI coding tools such as Claude Code or Cursor, with demonstrable shipped work.
• Strong sense of ownership and systems thinking.
• Familiarity with LLM evaluation frameworks and tracing tools like Logfire or OpenTelemetry.
• Experience with embeddings, vector search, or fuzzy matching for large-scale entity resolution.
• Background in fine-tuning or distilling smaller models.
• Experience with AWS services at scale, including S3, Lambda, Glue, ECS, and RDS.
• Proficiency in PostgreSQL.
• Knowledge of Spark or PySpark.
• Familiarity with streaming technologies such as Kafka, Kinesis, or Snowpipe Streaming.
• Understanding of data privacy and compliance regulations, including GDPR, SOC2, or CCPA.
• Competitive base salary.
• Meaningful equity options.
• Flexible working hours.
• Access to AI coding tools, automated testing, and AI-assisted reviews as standard.
• Small senior teams with minimal bureaucracy.
• Weekly releases with a goal of transitioning to daily updates.
• Full ownership of the technology stack.
CuraLinc Healthcare
VSP Vision Care
Adoreal
Get handpicked remote jobs straight to your inbox weekly.