Senior Data Engineer

atGravieRemoteUS flagUnited StatesFull-timeData EngineerSenior$133.9k – $178.5k/year

Posted 11 hours ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Assume complete ownership of the streaming platform's lifecycle, encompassing architecture, implementation, and operational management.

• Operate the platform in a production environment and take responsibility for latency and throughput SLOs.

• Utilize monitoring tools like Datadog for alerts and system oversight.

• Manage tasks such as replays, source-specific backfills, connector-failure recovery, dead-letter triage, and adjustments for fluctuating batch-driven claims loads.

• Develop and enhance streaming pipelines that ingest CDC events from operational databases and SaaS sources into a canonical, contract-validated format.

• Transform and enrich data using Spark Structured Streaming or dbt-on-Spark micro-batch processes.

• Conduct cross-stream joins to correlate events into unified lifecycle entities.

• Facilitate secure and governed access to PHI through classification, row/column controls, and access policies such as ABAC.

• Design pipelines that are idempotent and capable of replaying.

• Implement data-quality validation, runbooks, and observability practices.

• Provision streaming, processing, storage, and catalog services on AWS as code using CDK.

• Establish CI/CD for data pipelines and optimize infrastructure costs against latency SLOs.

• Collaborate with upstream producers regarding source changes and contracts.

• Work together with downstream consumers to address access and data requirements.

• Convert stakeholder requirements into actionable platform deliverables.

• Exhibit Gravie’s core values of authenticity, curiosity, creativity, empathy, and outcome orientation.


⛳️ Requirements

• Over 6 years of experience in building and operating production data systems, specifically in streaming or event-driven pipelines.

• Extensive production knowledge of Apache Kafka, focusing on partitioning, consumer groups, consumer lag, broker health troubleshooting, exactly-once/idempotent semantics, schema registry, and replay/backfill under load.

• Strong practical experience with Apache Spark (PySpark) for streaming and batch transformation in production settings.

• Proficient in AWS-native data engineering across streaming, processing, storage, and catalog services, including MSK, EMR, Glue, S3, and Athena.

• Familiarity with AWS CDK or Terraform, CI/CD for data pipelines, and a keen awareness of cost management.

• Competent in debugging distributed data pipelines using observability tools such as Datadog or CloudWatch.

• Advanced skills in SQL and Python.

• Experience in building and consuming REST APIs.

• Hands-on experience with AI-assisted and agentic development tools.

• Understanding of agentic data consumption patterns, including context management, lineage, provenance, permissions, freshness, and low-latency access.

• Experience with CDC technologies like Debezium and open table formats such as Iceberg or Delta Lake.

• Knowledge of schema evolution, partitioning, and table maintenance practices.

• Familiarity with data contracts, schema governance, cataloging, schema registries, compatibility rules, and dead-letter management.

• Experience with technical metastores like Glue Data Catalog or Unity Catalog and governance/discovery catalogs such as Atlan, Alation, or Collibra.

• A degree in Computer Science, Information Systems, or a related quantitative field.

• Proficient with the command line and Unix-based operating systems.

• Knowledge of the health insurance domain, including HIPAA and the protection of PHI.

• Excellent communication abilities and a proven track record of achieving results through influence and collaboration.

• Bonus: Experience with Apache Flink or other stateful stream processors; familiarity with JVM-based languages such as Kotlin or Java; experience serving data to AI and agentic consumers; previous work at a venture-backed start-up.

• Must currently possess authorization to work in the United States.

• Visa sponsorship details should be addressed in the application form.


🏝️ Benefits

• Coverage for alternative medicine.

• Flexible paid time off (PTO).

• Up to 16 weeks of paid parental leave.

• Paid holidays.

• 401k retirement program.

• Transportation benefits.

• Education reimbursement opportunities.

• 2 days of paid paw-ternity leave.

• Standard health and wellness benefits.

People also viewed

Intus Care10 hours ago

Staff Software Engineer, Data

US flagUnited States OnlyFull-timeData Engineer$180k – $190k/year
ApplyView job
Conta Simples10 hours ago

Data Platform Engineer

BR flagBrazil OnlyFull-timeData Engineer
ApplyView job
Super.com11 hours ago

Senior Data Engineer

CA flagCanada, +1 more countryFull-timeData Engineer$120k – $170k/year
ApplyView job
Tabby12 hours ago

Senior Data Engineer

ES flagSpain OnlyFull-timeData Engineer
ApplyView job
Dadoteca12 hours ago

Senior Data Engineer – Microsoft Fabric Support

BR flagBrazil OnlyFull-timeData Engineer
ApplyView job
Compass12 hours ago

Senior Data Engineer, Databricks

BR flagBrazil OnlyFull-timeData Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers