
Data Engineer
Posted 12 hours ago

Posted 12 hours ago
This is a fully remote position, open to applicants in United States.
• Take full ownership of UJET's data systems from start to finish, encompassing CDC replication from production MySQL, BigQuery pipelines and jobs, Terraform-managed infrastructure, and Looker models.
• Develop and refine the foundational elements of the data platform to facilitate secure dbt development, which includes project structure, environments, permissions, patterns, and documentation.
• Set up and uphold dbt standards and guidelines, focusing on testing, source freshness, documentation, and code review practices.
• Enhance the developer experience through CI/CD processes, automated checks, and consistent deployment patterns.
• Oversee data infrastructure as code using Terraform, managing pipelines, IAM, and environments.
• Troubleshoot and enhance CDC replication from MySQL to BigQuery utilizing Datastream.
• Develop tools for backfills, drift detection, and reconciliation between source databases and the warehouse.
• Participate in data ingestion and platform evolution discussions regarding latency, throughput, and alternative CDC methods.
• Collaborate with Analytics and Finance to define and present reliable metrics and Looker dashboards.
• Construct and maintain analytics-ready datasets for self-service reporting and experimentation.
• Support customer-facing, multi-tenant reporting with an emphasis on accuracy, tenant isolation, and query performance.
• Design and optimize BigQuery data models while implementing ELT best practices in dbt.
• Create, build, and sustain scalable data pipelines using Python, dbt, and containerized jobs on GKE.
• Ensure data quality, reliability, and observability for essential datasets and reporting.
• Optimize performance and cost efficiency across BigQuery and data pipelines.
• Integrate data workflows with backend services and APIs.
• Collaborate with Product and Engineering teams to convert business requirements into effective data solutions.
• A minimum of 5 years of software engineering experience, specifically with a strong emphasis on data platforms.
• Proficient programming skills in Python.
• Solid SQL skills along with experience in analytical data modeling.
• Direct experience with BigQuery or a similar cloud data warehouse.
• Practical experience with dbt (dbt Labs), including the ability to build models independently.
• Experience in constructing or enhancing the infrastructure surrounding data workflows, focusing on reliability, observability, CI/CD, permissions, environments, and deployment patterns.
• Familiarity with infrastructure as code, preferably using Terraform.
• Experience with a major cloud platform, with a preference for GCP.
• Strong foundational understanding of software engineering principles, including testing, version control, and code reviews.
• Must possess legal authorization to work in the United States without current or future employer-sponsored visa sponsorship.
• Preferred: Experience with GCP services such as Datastream, Cloud Storage, Cloud Run, and Pub/Sub.
• Preferred: Experience in change data capture, including GCP Datastream or MySQL binlog/GTID replication.
• Preferred: Experience with Kubernetes-based orchestration tools, such as Argo Workflows.
• Preferred: Proficiency in Looker, LookML, semantic modeling, and dashboard creation.
• Preferred: Familiarity with Ruby on Rails or Go.
• Preferred: Experience working alongside Product teams to gather requirements and translate business needs into data solutions.
• Medical insurance
• Dental insurance
• Vision insurance
• 401(k) plan
• Wellness benefits
• Equal employment opportunities
• Reasonable accommodations in the application process, including alternatives to AI-assisted screening
Solar Coca-Cola
Get handpicked remote jobs straight to your inbox weekly.