
Senior Data Engineer β AI-Native, Data Layer
Posted Jul 22

Posted Jul 22
This is a fully remote position, open to applicants in United States.
β’ Take charge of the Data Layer from start to finish: including ingestion from file, event, and API sources; employing the medallion-style model (raw β refined β curated); and managing the serving layer that supports the product and AI functionalities.
β’ Design and maintain the ingestion and transformation pipelines that drive the Data Layer, utilizing a contemporary orchestration framework and cloud data warehouse.
β’ Ingest and harmonize large, complex, real-world data from various source types and formats β including batch files, streaming events, and APIs.
β’ Structure data across medallion layers to ensure it is reliable, queryable, and stable for downstream teams and the AI.
β’ Contribute to advancing the Data Layer β enhancing architecture, improving tooling, increasing scalability, and integrating more sources β while having significant input on the vision.
β’ Manage AI coding agents (such as Claude Code and similar) at an advanced level: defining work scope, structuring context, running agents in parallel when applicable, and delivering reviewed, production-quality outputs.
β’ Create systems that establish data reliability β including validation, reconciliation, lineage, backfills, along with idempotent and incremental loads β to ensure downstream teams and AI avoid silent errors.
β’ Collaborate with backend, AI, and product engineers (as well as occasionally customers' IT teams) to define the data contracts they rely on.
β’ Over 7 years of hands-on experience as a data engineer with proven production ownership β developing pipelines and data models that serve real users at scale.
β’ Strong foundational knowledge. You comprehend the functioning and reasoning behind your code and queries. You can analyze a query plan, troubleshoot a slow or costly pipeline, and identify the root cause of data-correctness issues.
β’ Proficient programming and SQL abilities. You efficiently construct pipelines, schemas, and queries and can model data for both transactional and analytical access patterns.
β’ Practical orchestration experience, creating reliable ingestion/ELT pipelines from challenging upstream sources.
β’ Familiarity with a cloud data warehouse and a leading cloud platform.
β’ Experience in ingesting data from diverse source types: file-based, event/streaming, and API-based.
β’ Strong understanding of data-consistency failure modes β including partial loads, delayed or out-of-order data, idempotency, backfills, and schema drift.
β’ Daily, hands-on usage of agentic development tools (Claude Code, Cursor agent mode, Codex, or similar) to deliver real work. You can discuss in detail how you structure prompts, manage context, parallelize agents, and validate their outputs.
β’ Ownership and discernment. You can take data systems from concept to deployment and make informed decisions about what to build and what to eliminate.
β’ Startup mentality and excellent communication skills β pragmatic, fast-paced, action-oriented, and capable of articulating data decisions to engineers, PMs, and customers in writing.
β’ Proficiency in English at a C1 level or above.
β’ Opportunities for professional development.
Railroad19
GFT Technologies
Get handpicked remote jobs straight to your inbox weekly.