
Senior AI Data Engineer
Posted 14 hours ago

Posted 14 hours ago
This is a fully remote position, open to applicants in Mexico, +1 more country.
• Design and manage the data infrastructure that supports AI systems.
• Oversee the movement, modeling, and quality of data from source systems through the warehouse to retrieval and feature layers.
• Develop and sustain data pipelines that drive LLM pipelines, agentic workflows, and analytical products.
• Create models for warehouse schemas and make informed decisions regarding grain, relationships, and denormalization.
• Write production-level Python code and maintain pipelines as versioned, tested, and observable software.
• Collaborate with the AI/ML Data Scientist, who is responsible for model behavior and retrieval strategy.
• Take ownership of the pipeline, schema, and data guarantees while working together on the algorithm, prompt, and evaluation boundaries.
• 5-10+ years of experience in software engineering or data engineering, with significant involvement in production data platform work.
• Proven expertise in Kimball dimensional modeling, including selecting grain, resolving many-to-many relationships, and making denormalization decisions.
• Familiarity with Data Vault, One Big Table, Inman, and the trade-offs involved with Kimball.
• Advanced SQL skills: window functions, CTEs, query plan analysis, and performance optimization on a columnar warehouse.
• Proficient in production-grade Python: typing, packaging, dependency management, and testing.
• Experience in applying SOLID principles and domain-driven design in real-world systems.
• Practical experience with Airflow, Prefect, Dagster, or similar tools.
• Extensive knowledge of AWS, Azure, or GCP, covering storage, compute, IAM, networking, and cost management.
• Familiarity with containerization, Kubernetes (EKS/AKS/GKE), and CI/CD practices.
• Experience with lakehouse table formats like Iceberg, Delta Lake, or Hudi, including aspects like compaction, snapshot expiry, schema evolution, and partition evolution.
• Preferred: Experience building data layers beneath production RAG systems, including hybrid search infrastructure and index freshness guarantees.
• Preferred: Experience with Kafka, Kinesis, Flink, or Spark Structured Streaming.
• Preferred: Knowledge of dbt or a similar transformation and testing framework.
• Preferred: Proficiency in data quality tools such as Great Expectations or Soda, and catalog/lineage platforms.
• Preferred: Understanding of MLflow, Weights & Biases, and feature stores.
• Preferred: Working knowledge of an additional programming language such as TypeScript, Java, Go, Scala, or Rust.
• Preferred: Experience with AI security, governance, and compliance frameworks.
• Preferred: Contributions to open-source projects related to data or AI infrastructure.
• Competitive salary and performance-based bonuses.
• Flexible working hours and remote work options.
• Opportunities for professional development and continuous learning.
• Comprehensive health, dental, and vision insurance.
• Generous vacation and paid time off policies.
• Collaborative and inclusive company culture.
Katapult Labs
Magna Legal Services
Huron
MGM Resorts International
Get handpicked remote jobs straight to your inbox weekly.