
Senior Data Reliability Engineer, AWS
Posted Aug 1

Posted Aug 1
This is a fully remote position, open to applicants in United States.
• Take ownership of the reliability and stability of production data pipelines and services on the data platform.
• Identify and rectify failures, delays, and data quality issues within production data pipelines.
• Examine problems across distributed data systems, including Spark/EMR workloads, ingestion pipelines, and warehouse performance.
• Lead or assist in incident response, which encompasses triage, mitigation, and long-term resolution strategies.
• Conduct root cause analysis (RCA) and implement sustainable fixes to avert future occurrences.
• Establish and enhance data service level agreements (SLAs) concerning freshness, latency, and completeness, ensuring compliance.
• Design and improve monitoring, alerting, and observability mechanisms for data systems.
• Create automation and tools aimed at minimizing operational toil while enhancing system resilience.
• Contribute to disaster recovery (DR) and resilience planning, which includes validating backups and recovery workflows.
• Collaborate with engineering teams to refine pipeline design, enhance reliability, and ensure operational readiness.
• Develop and maintain runbooks, standard operating procedures (SOPs), and operational documentation.
• Occasionally provide off-hours support for production data systems as needed.
• At least 5 years of experience with production data platforms in AWS environments.
• Previous experience in building data pipelines and overseeing their production, including facing real-world failures and operational challenges.
• Solid expertise in Python and SQL within real data systems.
• Practical experience troubleshooting distributed data processing systems, such as Spark/EMR, Redshift, and streaming systems.
• Demonstrated ability to debug and resolve issues in production data pipelines and platforms.
• Familiarity with AWS data services, including EMR, Redshift, DynamoDB, S3, or equivalent.
• Experience managing production incidents and conducting root cause analyses.
• Strong problem-solving skills with the ability to navigate ambiguous production challenges.
• Medical, dental, vision, and life insurance coverage.
• Retirement savings options through a 401(k) plan with generous company matching contributions (up to 6%), financial advisory services, potential discretionary contributions from the company, and a wide range of investment options.
• Tuition reimbursement up to $5,250 per year.
• A business-casual work environment that permits jeans.
• Generous paid time off from the start, which includes a paid time off program, ten paid company holidays, and three floating holidays each calendar year.
• Paid volunteer time amounting to 16 hours per calendar year.
• Leave of absence programs, including paid parental leave, short-term and long-term disability, and Family and Medical Leave (FMLA).
• Business Resource Groups (BRGs) that promote inclusion and collaboration within our organization and the communities where we operate. BRGs are open to everyone.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.