
Lead Data Engineer
Posted 17 hours ago

Posted 17 hours ago
This is a fully remote position, open to applicants in United States.
• Design, construct, and sustain scalable data pipelines, as well as ETL/ELT workflows and data models.
• Develop and enhance AWS-native data platforms utilizing AWS Glue, Amazon EMR, Amazon MWAA (Apache Airflow), Lambda, Step Functions, Amazon S3, Redshift, RDS, DMS, and CloudWatch.
• Create high-performance ingestion, transformation, and orchestration workflows for both structured and semi-structured data.
• Design and refine analytical data platforms leveraging Amazon Athena, Trino, Hive, OpenSearch, and enterprise data catalog technologies.
• Integrate both enterprise and external data sources across relational and NoSQL platforms.
• Develop AI-enabled data solutions utilizing Amazon Bedrock, RAG pipelines, and vector search technologies.
• Build cloud infrastructure with CloudFormation, GitHub, Harness, CI/CD pipelines, SNS, SQS, and EventBridge.
• Enhance reliability, scalability, performance, and maintainability through monitoring, troubleshooting, automation, and ongoing optimization.
• Support mission-critical analytics and reporting solutions within AWS-based federal data environments in compliance with FedRAMP and NIST 800-53 controls.
• Lead modernization efforts transitioning IBM DataStage, Hadoop, RunDeck, and shell-based workflows to cloud-native AWS services.
• Provide mentorship to junior engineers through technical guidance, architecture discussions, and code reviews.
• Collaborate with cross-functional teams in an Agile setting to define requirements, deliver data solutions, and communicate technical concepts.
• Applicants must be US Citizens and capable of obtaining Public Trust clearance.
• Bachelor's degree in Computer Science, Engineering, or a related technical discipline.
• Over 15 years of professional experience in data engineering or related fields.
• Extensive hands-on experience with Apache Spark (PySpark), Python, SQL (PostgreSQL), and dbt.
• Significant hands-on experience with AWS Glue, Amazon EMR, Amazon MWAA (Apache Airflow), AWS Lambda, AWS Step Functions, Amazon S3, Amazon Redshift, Amazon RDS, AWS DMS, and Amazon CloudWatch.
• Experience in developing scalable data pipelines, workflow orchestration, and data integration solutions across enterprise environments.
• Familiarity with Apache Iceberg, Parquet, ORC, and Avro.
• Experience in designing and optimizing solutions using PostgreSQL, Redshift, Oracle, GraphDB, and NoSQL platforms.
• Proficiency in performance tuning, system optimization, and enterprise-scale ETL/ELT architectures.
• Java development experience and modern CI/CD practices using Harness.
• Strong analytical and problem-solving capabilities.
• Experience working in agile, iterative software development environments.
• Ability to rapidly learn and apply new technologies and domain knowledge.
• Exceptional written and verbal communication skills, with the ability to clarify complex topics for diverse audiences.
• Remote Work (Hybrid roles will be specified in the job post).
• Competitive Compensation Package.
• Medical, Dental, and Vision coverage.
• Life Insurance, Short/Long Term Disability benefits.
• Employee Assistance Program.
• 401(k) with 4% matching.
• Generous PTO vacation policy.
• Annual budget for Continuing Education.
• Annual Wellness Budget.
• Bonus Incentive Programs including employee referrals and performance-based rewards.
GoFasti
SysMap Solutions
Grupo CVLB
Get handpicked remote jobs straight to your inbox weekly.