
Senior Data Engineer
Posted Sep 1

Posted Sep 1
This is a fully remote position, open to applicants in United States.
• Take ownership of a business domain's data from start to finish, ensuring reliable ingestion and trustworthy metrics.
• Design and implement source-data ingestion through third-party connectors, APIs, webhooks, file uploads, change data capture (CDC), and batch loading.
• Manage schema changes, incremental and full-refresh strategies, idempotency, replayability, backfills, as well as late or duplicate records.
• Develop data quality tests, contracts, monitoring, freshness and volume assessments, schema/type enforcement, referential and uniqueness constraints, reconciliation, and anomaly detection.
• Deliver the complete domain pipeline, encompassing ingestion configuration, raw data landing, dbt staging and mart models, quality tests, orchestration DAGs, and operational documentation.
• Convert stakeholder business definitions into accurate, tested transformations and authoritative documentation.
• Design clean, consistently defined, and semantically clear data to serve as an AI-ready single source of truth.
• Establish monitoring, alerting, and observability; ensure freshness and uptime; and optimize performance, computational utilization, and efficiency.
• Maintain version-controlled data infrastructure and continuous integration/continuous deployment (CI/CD) workflows.
• Create reference implementations for data ingestion and quality assurance.
• Review the work of engineers, collaborate on challenging problems, mentor colleagues, and work together with domain stakeholders.
• Collaborate with the Director of Data to execute the platform strategy within dependable production systems.
• Over 5 years of experience in building and maintaining production data pipelines.
• Advanced proficiency with dbt, including staging/mart architecture, incremental models, tests, macros, documentation, and exposures.
• In-depth knowledge of API-based data ingestion, covering OAuth2/API-key authentication, token refresh, pagination, rate limiting, retry/backoff strategies, and assembling incremental pulls.
• Expertise in webhooks, file handling, database/CDC, schema changes, idempotency, backfills, and incremental loading techniques.
• Proven success in designing data quality frameworks that include freshness, volume, reconciliation, anomaly checks, and alerting mechanisms.
• Strong proficiency in Python for production extraction, loading, and tool development.
• Proficient in SQL.
• Practical experience with a modern cloud data warehouse; familiarity with BigQuery or Snowflake preferred, but others may be considered.
• Experience in architecting a clean, documented, AI-ready single source of truth.
• Familiarity with Airflow, Dagster, or dbt Cloud jobs.
• Experience using Git in collaborative development settings.
• Medical, dental, and vision insurance.
• 401(k) plan with company matching.
• Health Savings Account (HSA).
• Life insurance coverage.
• Disability insurance.
• Paid time off.
• Parental leave.
Data Elephant
ICF
General Dynamics Information Technology
Logic20/20, Inc.
Get handpicked remote jobs straight to your inbox weekly.