
Senior Data Engineer
Posted Jul 30

Posted Jul 30
This is a fully remote position, open to applicants in District of Columbia, +1 more state.
• Offer expert knowledge on data engineering methodologies and best practices, including code-first development strategies and contemporary pipeline design patterns.
• Create, implement, and uphold the data architecture that serves products and end users, ensuring all assets are managed under source control.
• Develop, implement, and sustain ELT and ETL pipelines for the effective processing of source data within Azure Synapse and Azure Machine Learning, utilizing both SDK V1 and SDK V2.
• Transition source data identified by SBA OIG into Azure Data Lake Storage.
• Standardize entity attributes such as addresses, phone numbers, and other common fields.
• Evaluate, maintain, and enhance existing architecture and pipelines, including regular audits to address bottlenecks, deprecated dependencies, and architectural drift.
• Set quality controls across all pipelines and introduce mechanisms for error handling, logging, and validation checks.
• Integrate source control across all pipelines and analytics codebases, allowing code to evolve iteratively without compromising architecture stability.
• Enhance ingestion, processing, and storage capabilities across diverse datasets and data types, including modern columnar formats such as Parquet.
• Create self-service features enabling SBA OIG analysts to query and export data for investigations and audits.
• Draft comprehensive standard operating procedures that govern the creation, development, validation, publishing, execution, and monitoring of all data pipelines and assets within the Azure environment.
• Generate thorough documentation of the data architecture, comprising data dictionaries, entity relationship diagrams, and pipeline process maps.
• Maintain and expand the environment with additional datasets and services upon request, adhering to a defined intake and testing process prior to production deployment.
• Keep abreast of new AI tools pertinent to data engineering and contribute to exploratory efforts assessing automation and language model-assisted capabilities.
• Bachelor's degree in data engineering, computer science, data science, machine learning, mathematics, or a related field; alternatively, five years of relevant applied work experience in any of these fields.
• 5 years of experience in maintaining SQL databases and performing advanced operations in SQL and T-SQL.
• 5 years of experience in designing, implementing, and maintaining ELT and ETL processes within cloud-based data analytics environments.
• 3 years of experience working with Azure Synapse and Azure Machine Learning using the modern data stack; certifications such as DP-203 or equivalent are preferred.
• 3 years of experience in data manipulation using Python; knowledge of Pandas is required, while PySpark and Polars are preferred. Experience in developing reusable, modular code is advantageous.
• DP-203, Microsoft Certified Azure Data Engineer Associate, or an equivalent current certification is preferred.
• Experience implementing pipelines and infrastructure using code-first approaches, including Python SDK, CLI, REST APIs, or infrastructure as code tools like Terraform or Bicep.
• Experience with source control and continuous integration and delivery workflows for data assets.
• Proven familiarity with AI coding assistants and integration patterns for large language models.
• Experience with PySpark or Polars at a production scale.
• Knowledge of entity resolution and attribute normalization across records with inconsistent addresses, names, and identifiers.
• Ability to build self-service analytic access for non-engineering users.
• Medical
• Dental
• Vision
• Basic Life
• Health Savings Account
• 401K matching
• Three weeks of PTO/Sick
• 11 Paid Holidays
• Pre-Approved Online Training
Agility Robotics
Get handpicked remote jobs straight to your inbox weekly.