
Data Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in District of Columbia, +1 more state.
• Develop, sustain, and enhance batch and streaming data pipelines for analytics and AI tasks.
• Ingest structured, semi-structured, and unstructured datasets from various sources such as APIs, databases, SaaS systems, streaming feeds, and file-based origins.
• Transform, clean, enrich, and standardize data utilizing Dataflow (Apache Beam), Dataproc (Spark), BigQuery SQL, and Python.
• Provide curated datasets into BigQuery for analytics, reporting, and machine learning training purposes.
• Construct and uphold ML-ready feature pipelines.
• Implement data quality assessments, schema validation, and automated testing of pipelines.
• Monitor pipeline performance using Cloud Monitoring and Cloud Logging.
• Apply governance, security, and compliance protocols including IAM roles, encryption, data masking, and auditing processes.
• Enforce schema evolution, metadata management, and lineage tracking through Dataplex/Data Catalog.
• Keep documentation updated for datasets, transformations, pipeline logic, and operational procedures.
• Write maintainable code in Python, SQL, Beam, and Spark.
• Manage data ingestion processes using Pub/Sub, GCS, APIs, Datastream, and database connectors.
• Optimize BigQuery tables, partitions, clustering, materialized views, and query performance.
• Implement and manage DAGs with Cloud Composer (Airflow).
• Troubleshoot failures in pipelines, latency challenges, and data quality issues.
• Engage in code reviews, architectural discussions, and agile sprint activities.
• Collaborate with Data Architects, Data Scientists, ML Engineers, and business stakeholders.
• Develop and maintain Infrastructure as Code using Terraform and CI/CD deployment pipelines.
• A Bachelor’s degree in Computer Science, Software Engineering, Information Systems, Data Engineering, or a related technical field.
• 3–6+ years of practical experience in data engineering or a comparable technical area.
• At least three years of experience leading technical teams to achieve desired results.
• Experience in developing and implementing technical standards for both cloud and on-premises environments.
• Demonstrated experience in constructing production data pipelines on cloud platforms, preferably GCP.
• Practical experience with BigQuery, GCS, Dataflow (Apache Beam), Dataproc (Spark), and Pub/Sub.
• Background in preparing ML-ready datasets for model training.
• Strong knowledge of SQL, Python, distributed data processing, and data modeling techniques.
• Familiarity with governance, security, and compliance frameworks, including IAM, encryption, data masking, and auditing.
• Knowledge of Google Cloud Security tools, Google Cloud Monitoring & Logging tools, Google Cloud Networking, and Google Storage services.
• Must possess US work authorization that does not currently or in the future require visa sponsorship.
• A collaborative and supportive community.
• Opportunities for hands-on experience.
• Access to certifications.
• Industry training provided.
• A comprehensive range of benefits (specific details available through employer benefits information).
Vericast
BCD Travel
Growth Leads
Verily
Get handpicked remote jobs straight to your inbox weekly.