
Data Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in District of Columbia, +1 more state.
• Construct, manage, and enhance batch and streaming data pipelines tailored for analytics and AI tasks.
• Ingest various data types—structured, semi-structured, and unstructured—from APIs, databases, SaaS solutions, streaming feeds, and file sources.
• Transform, cleanse, enrich, and standardize data utilizing Dataflow (Apache Beam), Dataproc (Spark), BigQuery SQL, and Python.
• Provide curated datasets into BigQuery for analytics, reporting, and machine learning training purposes.
• Develop and sustain ML-ready feature pipelines.
• Implement data quality checks, schema validations, and automated testing processes.
• Oversee pipeline health using Cloud Monitoring and Cloud Logging tools.
• Apply governance, security, and compliance standards, including IAM roles, encryption, data masking, and auditing practices.
• Enforce schema evolution policies, manage metadata, and track data lineage using Dataplex/Data Catalog.
• Maintain comprehensive documentation for datasets, transformations, pipeline logic, and operational procedures.
• Write code in Python, SQL, Beam, and Spark.
• Manage data ingestion flows utilizing Pub/Sub, GCS, APIs, Datastream, and database connectors.
• Optimize BigQuery tables, partitions, clustering, materialized views, and enhance query performance.
• Implement and manage Directed Acyclic Graphs (DAGs) with Cloud Composer (Airflow).
• Troubleshoot pipeline failures, latency challenges, and data quality discrepancies.
• Engage in code reviews, architectural conversations, and agile sprint ceremonies.
• Collaborate with Data Architects, Data Scientists, ML Engineers, and business stakeholders.
• Develop and maintain Infrastructure as Code utilizing Terraform and CI/CD deployment pipelines.
• Bachelor’s degree in Computer Science, Software Engineering, Information Systems, Data Engineering, or a related technical discipline.
• 3–6+ years of practical experience in data engineering or a closely related technical field.
• At least three years of experience leading technical teams to deliver results.
• Experience in developing and enforcing technical standards for both cloud and on-premises environments.
• Demonstrated experience in building production data pipelines on cloud platforms, ideally GCP.
• Practical experience with BigQuery, GCS, Dataflow (Apache Beam), Dataproc (Spark), and Pub/Sub.
• Experience in preparing ML-ready datasets for model training.
• Strong expertise in SQL, Python, distributed data processing, and data modeling.
• Familiarity with governance, security, and compliance frameworks, including IAM, encryption, data masking, and auditing.
• Knowledge of Google Cloud Security tools, Google Cloud Monitoring & Logging tools, Google Cloud Networking, and Google Storage services.
• Must possess US work authorization that does not currently or in the future require visa sponsorship.
• A collaborative and supportive community.
• Hands-on experience opportunities.
• Access to certifications.
• Industry training programs.
• Equal employment opportunities.
• Reasonable accommodations for disabilities or religious observances.
Vericast
BCD Travel
Growth Leads
Verily
Get handpicked remote jobs straight to your inbox weekly.