
Senior Data Engineer
Posted Aug 7

Posted Aug 7
This is a fully remote position, open to applicants in Virginia.
• Design, develop, and sustain scalable data pipelines utilizing Spark, Hive, and Airflow.
• Create and implement data processing workflows on Databricks.
• Develop API services for data access and integration purposes.
• Generate interactive data visualizations and reports through AWS QuickSight.
• Construct infrastructure for data extraction, transformation, and loading using AWS and SQL technologies.
• Oversee and enhance data infrastructure and process performance.
• Create jobs for data quality and validation.
• Assemble intricate datasets that fulfill both functional and non-functional business requirements.
• Write unit and integration tests for the data processing code.
• Collaborate with DevOps engineers on CI, CD, and Infrastructure as Code (IaC).
• Convert specifications into code and design documentation.
• Conduct code reviews and enhance code quality processes.
• Increase data availability and timeliness through refreshes, tiered storage, and dataset optimization.
• Ensure data security and privacy during both rest and transit.
• Perform additional duties as assigned.
• Bachelor’s degree.
• 7+ years of practical software development experience.
• 4+ years of experience in building data pipelines using Python, Java, and cloud technologies.
• Practical experience with Spark and Hive for large-scale data processing.
• Must be able to obtain and maintain a Public Trust clearance.
• Must reside in the United States.
• Must be authorized to work in the United States.
• Work must be conducted in the United States.
• Must have lived in the United States for 3 full years out of the last 5 years.
• Preferred: experience in constructing workflows with Databricks.
• Preferred: strong knowledge of AWS products including S3, Redshift, RDS, EMR, AWS Glue, AWS Glue DataBrew, Jupyter Notebooks, Athena, QuickSight, and Amazon SNS.
• Preferred: familiarity with data transformation, workload management, data structures, dependencies, and metadata processes.
• Preferred: experience with data governance for batch and streaming ingestion, curation, and data sharing.
• Preferred: experience in optimizing and building data pipeline systems.
• Preferred: understanding of Cassandra, Postgres, Airflow, Luigi, Azkaban, Spark Streaming, Storm, Scala, C++, Java, and Python.
• Preferred: familiarity with CI/CD pipelines using GitHub Actions and IaC using Terraform.
• Preferred: experience with Agile methodology and test-driven development.
• Equal opportunity employer.
• Reasonable accommodations provided throughout the application and employment process.
• Benefit offerings referenced under the Transparency in Coverage Act.
Railroad19
GFT Technologies
Get handpicked remote jobs straight to your inbox weekly.