
Data Engineer – Legacy Systems, AI Workflows
Posted 16 hours ago

Posted 16 hours ago
This is a fully remote position, open to applicants in Maryland, +1 more state.
• Design, develop, and manage batch and streaming ETL/ELT pipelines that collect data from various structured and semi-structured sources into secure cloud data platforms.
• Integrate legacy data sources by assessing schemas, mappings, interfaces, validation rules, ownership, and operational support procedures.
• Create automated data-quality checks to ensure completeness, accuracy, consistency, timeliness, and referential integrity, accompanied by actionable monitoring and alerting.
• Automate deployments and pipeline operations utilizing SOAP APIs, workflow orchestration, and environment-specific configurations.
• Diagnose failures across distributed data systems through logs, metrics, lineage, and operational signals; clearly communicate root causes and recovery strategies.
• Work collaboratively with software engineers, cloud engineers, data scientists, security teams, architects, and customer stakeholders to translate mission requirements into measurable data products.
• Bachelor’s degree in Computer Science, Computer Engineering, Mathematics, Statistics, or a similar technical field.
• Over 5 years of professional experience in data engineering, including development and operations of production pipelines.
• Proficient in Python, Java, R, or other programming languages relevant to data engineering.
• Experience with legacy systems that use SOAP APIs for data interactions.
• Skilled in implementing data-quality frameworks, validation processes, monitoring, and alerting for production pipelines.
• Familiarity with data security, access controls, encryption, auditability, and governance within regulated, restricted, or compliance-oriented environments.
• Experience working with APIs (SOAP/REST) and structured data formats such as JSON, XML, CSV, and Parquet.
• Ability to document technical decisions and collaborate effectively with multidisciplinary teams and customer stakeholders.
• Must be a U.S. Citizen as per contract requirements.
• Must be able to obtain and maintain a Public Trust Clearance in line with contract requirements.
• Preferred: experience in building preprocessing, chunking, filtering, metadata-enrichment, or evaluation pipelines for LLM and other AI/ML workflows.
• Preferred: experience managing data platforms with segmented networks, limited connectivity, strict change control, or formal authorization necessities.
• Preferred: strong expertise in SQL, including query optimization, data modeling, joins, window functions, and analysis of extensive datasets.
• Preferred: experience with Jupyter Notebook or similar tools for data analysis.
• Preferred: experience with Python libraries such as numpy and pandas.
• Preferred: hands-on experience with at least one leading cloud provider; AWS is strongly preferred.
• Preferred: experience with C3.ai.
• Active Public Trust Clearance is preferred.
• Preference will be given to candidates located in the Washington, DC metro area.
• Preferred: AWS Data Engineer Certification, Databricks, cloud data engineering, or other relevant professional certifications.
• Comprehensive medical, dental, and vision plans.
• Flexible Spending Account.
• 4% 401K Match with immediate vesting.
• Paid Time Off.
• Tuition reimbursement, certification programs, and opportunities for professional development.
• Flexible work schedule.
• On-site gym and childcare options.
Katapult Labs
Magna Legal Services
Huron
Strategic Systems International
Get handpicked remote jobs straight to your inbox weekly.