
Senior Data Engineer
Posted Jul 19

Posted Jul 19
This is a fully remote position, open to applicants in Uruguay.
⢠Design and manage ingestion, ELT/ETL, and orchestration pipelines that transport data from our MongoDB Atlas operational database and other sources into our analytical and AI-serving infrastructures.
⢠Execute layered (medallion-style) transformations with idempotent, backfillable, and incrementally loaded jobs.
⢠Conduct deduplication, normalization, and validation to ensure downstream data is of high quality and reliable.
⢠Revamp legacy or homegrown data flows through incremental, strangler-fig migrations that maintain production stability.
⢠Develop embeddings and vector pipelines, along with the feature/retrieval-ready datasets essential for RAG, semantic search, and agentic workloads.
⢠Prepare production data to be AI-ready in practice: well-structured, lineage-tracked, and retrieval-friendly, in collaboration with ML and application engineering.
⢠Implement real-time and change-data-capture workflows from MongoDB (Change Streams / CDC) where current data is necessary.
⢠Enact the canonical data model, schemas, and data contracts as defined by the Data Architectāenforced in-repo to ensure other teams build against consistent definitions.
⢠Apply sound persistence judgment in execution: store data appropriately (document/NoSQL, vector, analytical) per architectural guidance.
⢠Assist in build-vs-buy decisions by prototyping with recognized, industry-standard tools rather than custom development.
⢠Establish testing, data quality, and lineage verification for the pipelines you manage, with transparent alerting and runbooks.
⢠Implement pipeline observability (freshness, volume, schema drift, cost) to detect failures before they impact consumers.
⢠Utilize AI-assisted development tools (Claude Code, Copilot, Cursor) as a force multiplier for transformation logic, query optimization, and migration scripting.
⢠Collaborate with database engineering to extract from and safeguard the production store.
⢠Work alongside the Data Architect to implement target-state patterns and address challenging builds.
⢠Partner with ML, AI, and application engineers on the data they utilizeāshaping and governing it to ensure it is secure and prepared for development.
⢠Over 5 years of practical data engineering experience in developing and managing production data pipelines at scale.
⢠Proficient programming and data skills: Python and SQL, with strong software engineering fundamentals (version control, testing, CI)āshipping and maintaining production code, not just notebooks.
⢠Hands-on experience with MongoDB at production scale (Atlas preferred): including document modeling, aggregation framework, change streams / CDC, and extracting from a document store into analytical/AI-serving layers.
⢠Proven experience in ELT/ETL pipeline design, transformation frameworks (dbt or similar), and orchestration (Airflow, Dagster, or Azure Data Factory).
⢠Experience building on cloud-native data platforms and lake/lakehouse/warehouse architectures, with layered (medallion-style) modeling.
⢠Practical experience in preparing data for AI/ML or analytical consumersāembeddings/vector pipelines, RAG-/feature-ready datasets, or equivalentāincluding deduplication, normalization, and validation.
⢠Familiarity with vector search and embeddings in a production environment (MongoDB Atlas Vector Search or similar).
⢠Demonstrated use of AI-assisted development tools (Claude Code, Copilot, Cursor) for data and pipeline tasks.
⢠Strong understanding of data quality, testing, lineage, and pipeline observability practices.
⢠Comfortable working in a complex, specialized domain; MEP/AEC/construction experience is a plus, but a willingness to learn the domain is essential.
⢠Competitive salary
⢠Flexible working hours
⢠Professional development budget
⢠Home office setup allowance
⢠Global team events
Omada Health
BPO Global Services S.A.S
Get handpicked remote jobs straight to your inbox weekly.