
Senior Data Engineer
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Connecticut, +2 more states.
• Design and develop ELT/ETL solutions for both batch and streaming data ingestion, integration, refinement, and publishing on the Lakehouse.
• Create reusable data processing frameworks and configuration-driven pipelines utilizing Python and PySpark.
• Construct and sustain scalable orchestration workflows for production data delivery, which includes retries, historical loads, and operational runbooks.
• Execute data quality checks, validation frameworks, and monitoring to adhere to data contracts and SLAs.
• Implement DataOps practices, including Git-based development, CI/CD/CT, automated testing, and managed environment promotion.
• Contribute to data lifecycle management practices such as retention, archival, disaster recovery, and resilience.
• Assist in platform modernization and the migration of legacy data flows into Lakehouse architectures.
• Collaborate with stakeholders to align technical designs with business processes, non-functional requirements, and consumption needs.
• Establish and document standards, naming conventions, and engineering practices; engage in Agile ceremonies and cross-team delivery.
• Mentor engineers, perform design and code reviews, and enhance reliability, performance, and cost efficiency.
• Manage reliable datasets that facilitate analytics and reporting while ensuring governed access across the Lakehouse.
• Bachelor's degree in Computer Science, Information Systems, Engineering, or a related technical field, or equivalent professional experience.
• Over 8 years of experience in the IT field.
• More than 5 years of practical experience in designing and developing enterprise-scale data engineering solutions.
• Strong expertise in developing scalable data pipelines and reusable frameworks using Python and PySpark.
• Familiarity with AWS Glue, dbt, Apache Spark, or similar technologies.
• Comprehensive understanding of data warehousing, dimensional modeling, and modern data lake/Lakehouse architectures.
• Experience with AWS, Microsoft Azure, or Google Cloud Platform, with a preference for AWS.
• Proficient in SQL with relational databases; familiarity with NoSQL is an advantage.
• Experience with Apache Spark, Amazon EMR, or Hadoop-based platforms.
• Proficient in using Git for source control and Agile software development methodologies.
• Strong analytical, problem-solving, and communication abilities.
• Preferred: Knowledge of AWS services including S3, Glue, EMR, Athena, Redshift, Lambda, Lake Formation, IAM, and CloudWatch.
• Preferred: Experience with Apache Iceberg or similar open table formats and governed Lakehouse patterns.
• Preferred: Expertise in CI/CD, DataOps, continuous testing, and infrastructure automation tools like Terraform.
• Preferred: Familiarity with Apache Airflow, MWAA, AWS Step Functions, or other workflow orchestration tools.
• Preferred: Experience with RESTful APIs or other access layers and enterprise application integrations.
• Preferred: Knowledge of Kafka, Kinesis, or Spark Structured Streaming.
• Preferred: Proficiency in data quality, observability, monitoring, and automated validation frameworks.
• Preferred: Familiarity with data governance, metadata management, lineage, and enterprise data catalog solutions.
• Preferred: Understanding of data security, encryption, access controls, and compliance with healthcare regulations, including HIPAA/PHI.
• Preferred: Experience in optimizing distributed workloads for scalability, reliability, and cloud cost efficiency.
• Preferred: Skills in mentoring engineers, conducting design and code reviews, and establishing engineering best practices.
• Medical coverage.
• Dental coverage.
• Vision coverage.
• Incentive and recognition programs.
• Life insurance.
• 401k contributions.
• Competitive compensation and benefits package.
Pluribus Digital
GoMining
Get handpicked remote jobs straight to your inbox weekly.