
Data Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States, +1 more state.
• Design and manage robust data pipelines for ingestion, cleaning, and transformation utilizing Databricks, Airflow, or Kubernetes.
• Create efficient ETL/ELT workflows in Python and SQL for both batch and streaming workloads.
• Collaborate with ML/AI teams to ensure datasets and tools are discoverable and secure for autonomous agents, including evaluation and guardrails for AI-generated queries.
• Develop retrieval pipelines (RAG, vector search) over structured statistics and unstructured sources such as scouting notes and video metadata.
• Model and maintain structured data assets in Delta, Parquet, and Iceberg for reliability, versioning, and lineage tracking.
• Implement orchestration and monitoring by scheduling jobs, tracking dependencies, and automating failure recovery.
• Guarantee data quality and compliance through validation frameworks, schema enforcement, and audit logging.
• Contribute to the evolution of the data platform by assessing tools, standardizing best practices, and enhancing developer experience.
• Aid in performance and cost optimization across compute, storage, and orchestration systems.
• Work collaboratively with MLOps and Sports Data teams to integrate data and AI systems.
• 3–8 years of experience as a Data Engineer or ETL Developer in a production setting.
• Expertise in Python and SQL.
• Strong knowledge of Databricks, Spark, or similar big-data frameworks.
• Experience with workflow orchestration tools such as Airflow, Dagster, Luigi, or Prefect.
• In-depth understanding of data modeling, data warehousing, and distributed data processing.
• Familiarity with contemporary data lakehouse architectures.
• Knowledge of CI/CD, GitHub Actions, Infrastructure as Code, and data pipeline testing frameworks.
• Ability to work cross-functionally with ML, product, and analytics teams.
• Exposure to LLM-powered data tools, including text-to-SQL, RAG, agent/tool interfaces like MCP, or natural-language analytics.
• Prior experience with cloud infrastructure such as AWS, GCP, or Azure.
• Experience in container orchestration using Docker or Kubernetes.
• Preferred: experience with sports, telemetry, or sensor data pipelines.
• Preferred: familiarity with streaming frameworks and event-driven data processing such as Kafka, Spark Structured Streaming, or Flink.
• Preferred: general knowledge of American football, the NFL, and college football.
• Preferred: background in data governance, lineage, and observability tools such as Monte Carlo, Great Expectations, Unity Catalog, or OpenLineage.
• Preferred: experience designing semantic layers or metric definitions for AI and BI tools.
• Preferred: exposure to machine-learning model management and MLOps best practices.
• Competitive Salary and Bonus Plan.
• Comprehensive health insurance plan.
• Retirement savings plan (401k) with company match.
• Remote working environment.
• A flexible, unlimited time off policy.
• Generous paid holiday schedule - 13 in total including the Monday after the Super Bowl.
• Annual performance bonus.
• Benefits and/or other applicable incentive compensation plans.
Agility Robotics
Get handpicked remote jobs straight to your inbox weekly.