
Senior Staff Platform, Data Reliability Engineer
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in United States.
• Take charge of operational excellence for the Databricks platform, which includes monitoring, alerting, observability, incident response support, and production runbook patterns.
• Establish and uphold CI/CD and promotion standards for Databricks assets and environments from development through to production.
• Create and sustain standards for job orchestration, cluster and compute policies, service principal usage, environment isolation, and dependable production execution.
• Develop reusable operational templates and enablement patterns to assist in onboarding new domains to Databricks.
• Collaborate with the Senior Data Engineer on observable, recoverable, cost-effective, and secure ingestion and medallion patterns.
• Ensure that Databricks configuration and usage align with enterprise cloud standards across both commercial and future government-hosted environments.
• Implement technical controls for data segregation, access boundaries, and operational compliance.
• Monitor and enhance platform health metrics, which include job success rates, incident trends, pipeline reliability, cost efficiency, and environment drift.
• Document platform standards, operational expectations, and support frameworks.
• Guide internal engineers in developing their platform responsibilities.
• Oversee the Databricks operational layer encompassing reliability, observability, deployment standards, compute and job policies, and platform enablement.
• A minimum of 12 years of relevant experience in data platform engineering, platform operations, site reliability engineering, or modern cloud data infrastructure.
• Practical production experience with Databricks or a similar cloud data platform.
• Familiarity with designing or operating CI/CD, environment promotion, version control, and deployment automation for data platforms and pipelines.
• Strong comprehension of observability, monitoring, alerting, incident management, and reliability engineering.
• Experience in compute policy design, workload isolation, service principals, and secure production execution patterns on cloud data platforms.
• Capability to operate in regulated or security-sensitive environments with access control, auditability, and operational discipline.
• Experience collaborating with cloud/infrastructure, security, data engineering, and analytics stakeholders.
• Preferred: Databricks certification and/or expertise with Delta Lake, Unity Catalog, Workflows, and Databricks Asset Bundles.
• Preferred: Experience with infrastructure-as-code and platform automation.
• Preferred: Background in supporting commercial and government or segregated environments.
• Preferred: Experience in defense, aerospace, federal, or another regulated industry.
• Bonus.
• Benefits for full-time regular employees.
• Equity.
• Temporary benefits package applicable after 60 days of employment for temporary employees.
• Benefits eligibility excluded for military fellows and part-time employees.
• Equal employment opportunity and accommodations for disability or special needs.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.