
Data Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Create, construct, and sustain data pipelines that drive deep learning, video, and LLM systems.
• Design and manage efficient pipelines for data ingestion, cleansing, and transformation using Databricks, Airflow, or Kubernetes.
• Develop ETL/ELT workflows using Python and SQL for both batch and streaming processes.
• Collaborate with ML/AI teams to ensure datasets and tools are accessible and secure for autonomous agents, including the implementation of evaluation metrics and guardrails for AI-generated queries.
• Construct retrieval pipelines utilizing RAG and vector search techniques over structured statistics and unstructured data, including scouting notes and video metadata.
• Model and oversee structured data assets in Delta, Parquet, and Iceberg formats to ensure reliability, versioning, and lineage tracking.
• Establish orchestration and monitoring processes by scheduling jobs, managing dependencies, and automating failure recovery.
• Guarantee data quality and compliance through validation frameworks, schema enforcement, and audit logging.
• Contribute to the advancement of the data platform by assessing tools, standardizing best practices, and enhancing the developer experience.
• Assist in performance and cost optimization across compute, storage, and orchestration systems.
• Work closely with MLOps and Sports Data teams to integrate data and AI across various sports.
• 3–8 years of experience as a Data Engineer or ETL Developer within a production environment.
• Expertise in Python and SQL.
• Strong knowledge of Databricks, Spark, or similar big-data frameworks.
• Familiarity with workflow orchestration tools such as Airflow, Dagster, Luigi, or Prefect.
• In-depth understanding of data modeling, data warehousing, and distributed data processing.
• Awareness of contemporary data lakehouse architectures.
• Experience with CI/CD, GitHub Actions, Infrastructure as Code, and data pipeline testing frameworks.
• Ability to work effectively in a cross-functional environment alongside ML, product, and analytics teams.
• Exposure to LLM-powered data tools: text-to-SQL, RAG, agent/tool interfaces (e.g., MCP), or natural-language analytics.
• Prior experience with cloud infrastructure (AWS, GCP, or Azure) and container orchestration (Docker, Kubernetes).
• Preferred: Previous experience with sports, telemetry, or sensor data pipelines.
• Preferred: Familiarity with streaming frameworks and event-driven data processing (Kafka, Spark Structured Streaming, Flink).
• Preferred: General knowledge of American football, the NFL, and college football.
• Preferred: Background in data governance, lineage, and observability tools (Monte Carlo, Great Expectations, Unity Catalog, OpenLineage).
• Preferred: Experience designing semantic layers or metric definitions utilized by AI and BI tools.
• Preferred: Knowledge of best practices in machine-learning model management and MLOps.
• Competitive Salary and Bonus Plan.
• Comprehensive health insurance plan.
• Retirement savings plan (401k) with company match.
• Remote working environment.
• A flexible, unlimited time off policy.
• Generous paid holiday schedule - 13 in total, including the Monday after the Super Bowl.
• Annual performance bonus, benefits, and/or other applicable incentive compensation plans may be included in the total compensation package.
ASRC Federal
Get handpicked remote jobs straight to your inbox weekly.