
Senior Data Engineer
Posted 11 hours ago

Posted 11 hours ago
This is a fully remote position, open to applicants in United States.
• Assume complete ownership of the streaming platform's lifecycle, encompassing architecture, implementation, and operational management.
• Operate the platform in a production environment and take responsibility for latency and throughput SLOs.
• Utilize monitoring tools like Datadog for alerts and system oversight.
• Manage tasks such as replays, source-specific backfills, connector-failure recovery, dead-letter triage, and adjustments for fluctuating batch-driven claims loads.
• Develop and enhance streaming pipelines that ingest CDC events from operational databases and SaaS sources into a canonical, contract-validated format.
• Transform and enrich data using Spark Structured Streaming or dbt-on-Spark micro-batch processes.
• Conduct cross-stream joins to correlate events into unified lifecycle entities.
• Facilitate secure and governed access to PHI through classification, row/column controls, and access policies such as ABAC.
• Design pipelines that are idempotent and capable of replaying.
• Implement data-quality validation, runbooks, and observability practices.
• Provision streaming, processing, storage, and catalog services on AWS as code using CDK.
• Establish CI/CD for data pipelines and optimize infrastructure costs against latency SLOs.
• Collaborate with upstream producers regarding source changes and contracts.
• Work together with downstream consumers to address access and data requirements.
• Convert stakeholder requirements into actionable platform deliverables.
• Exhibit Gravie’s core values of authenticity, curiosity, creativity, empathy, and outcome orientation.
• Over 6 years of experience in building and operating production data systems, specifically in streaming or event-driven pipelines.
• Extensive production knowledge of Apache Kafka, focusing on partitioning, consumer groups, consumer lag, broker health troubleshooting, exactly-once/idempotent semantics, schema registry, and replay/backfill under load.
• Strong practical experience with Apache Spark (PySpark) for streaming and batch transformation in production settings.
• Proficient in AWS-native data engineering across streaming, processing, storage, and catalog services, including MSK, EMR, Glue, S3, and Athena.
• Familiarity with AWS CDK or Terraform, CI/CD for data pipelines, and a keen awareness of cost management.
• Competent in debugging distributed data pipelines using observability tools such as Datadog or CloudWatch.
• Advanced skills in SQL and Python.
• Experience in building and consuming REST APIs.
• Hands-on experience with AI-assisted and agentic development tools.
• Understanding of agentic data consumption patterns, including context management, lineage, provenance, permissions, freshness, and low-latency access.
• Experience with CDC technologies like Debezium and open table formats such as Iceberg or Delta Lake.
• Knowledge of schema evolution, partitioning, and table maintenance practices.
• Familiarity with data contracts, schema governance, cataloging, schema registries, compatibility rules, and dead-letter management.
• Experience with technical metastores like Glue Data Catalog or Unity Catalog and governance/discovery catalogs such as Atlan, Alation, or Collibra.
• A degree in Computer Science, Information Systems, or a related quantitative field.
• Proficient with the command line and Unix-based operating systems.
• Knowledge of the health insurance domain, including HIPAA and the protection of PHI.
• Excellent communication abilities and a proven track record of achieving results through influence and collaboration.
• Bonus: Experience with Apache Flink or other stateful stream processors; familiarity with JVM-based languages such as Kotlin or Java; experience serving data to AI and agentic consumers; previous work at a venture-backed start-up.
• Must currently possess authorization to work in the United States.
• Visa sponsorship details should be addressed in the application form.
• Coverage for alternative medicine.
• Flexible paid time off (PTO).
• Up to 16 weeks of paid parental leave.
• Paid holidays.
• 401k retirement program.
• Transportation benefits.
• Education reimbursement opportunities.
• 2 days of paid paw-ternity leave.
• Standard health and wellness benefits.
Intus Care
Conta Simples
Super.com
Get handpicked remote jobs straight to your inbox weekly.