
Senior Data Engineer, Platform
Posted 16 hours ago

Posted 16 hours ago
This is a fully remote position, open to applicants in California, +3 more states.
• Define and steer the technical vision for a significant domain within the DGX Cloud Data Platform.
• Take ownership of architecture, interfaces, and growth while addressing critical aspects such as scale, reliability, performance, security, compatibility, and cost.
• Lead the technical delivery of intricate cross-team projects.
• Convert ambiguous requirements into clear architectures and interfaces, coordinate implementation, write essential code, resolve technical challenges, and guide secure production integrations.
• Design, implement, and enhance batch and streaming systems for fleet, capacity, utilization, cost, scheduling, and operational telemetry.
• Develop shared platform capabilities, libraries, workflow/orchestration abstractions, deployment tools, and implementation standards.
• Spearhead high-impact production investigations across pipelines, applications, query engines, distributed processing, storage, networks, and cloud services.
• Identify root causes, drive sustainable resolutions, and implement preventive measures.
• Promote engineering standards for testing, data quality, reconciliation, lineage, SLOs, observability, secure identities, least privilege, release readiness, and auditable deployments.
• Define data models, semantics, ownership boundaries, and serving interfaces.
• Facilitate tables, APIs, automation, dashboards, and internal applications for trusted access to DGX Cloud data.
• Provide technical leadership through architecture and build reviews, mentorship of senior engineers, and evidence-based resolution of tradeoffs.
• 8+ years of relevant industry experience.
• Bachelor’s degree or equivalent experience.
• Master’s degree or equivalent experience in Computer Science, Engineering, or a related field.
• Proven history of personally designing, implementing, and managing production software, data platforms, databases, or distributed systems.
• End-to-end technical ownership of a multi-system platform domain or complex cross-team engineering initiative.
• Extensive hands-on experience with distributed processing, analytical or relational databases, production ETL, change-data capture, streaming or event processing, or backend and cloud systems handling large data volumes.
• Solid software engineering fundamentals and production proficiency in a backend or systems programming language.
• Profound experience with data-processing and platform libraries or frameworks.
• Experience creating reusable abstractions, reviewing significant changes, and troubleshooting critical code paths.
• Strong SQL and data-modeling capabilities.
• Practical knowledge in query execution, incremental processing, schema evolution, consistency, analytical consumption, idempotency, replay, late-arriving data, partial failure, and cross-system correctness.
• Expertise in diagnosing failures using logs, metrics, traces, query plans, profiles, and controlled experiments.
• Strong architectural judgment concerning reliability, performance, cost, security, compatibility, and maintainability.
• Experience leading major migrations or architectural transformations across teams while maintaining uninterrupted production service.
• Experience establishing production safeguards and engineering practices adopted by multiple teams, including automated testing, CI/CD, monitoring, alerting, rollback, incident response, and secure deployment.
• Extensive experience with distributed data processing and lakehouse architectures, or equivalent large-scale database/data-processing platforms.
• Experience managing distributed streaming or event-driven systems, including aspects like partitioning, consumer behavior, flow control, replay, delivery guarantees, and schema evolution.
• Experience scaling, migrating, or enhancing the performance of relational, distributed, time-series, object-storage, or searchable-content data systems.
• Background in operating cloud infrastructure, container orchestration, workload schedulers, compute or GPU clusters, and fleet-scale telemetry.
• Experience defining and owning the production adoption of agentic systems or workflow automation, with a focus on evaluation, permissions, observability, failure recovery, and measurable improvements.
• Highly competitive salaries.
• Comprehensive benefits package.
• Equity.
• Benefits for you and your family.
Tonic3
Latitude IT Solutions | SDVOSB
Get handpicked remote jobs straight to your inbox weekly.