Senior Database Reliability Engineer

Posted Aug 26

This is a fully remote position, open to applicants in United States.

📋 Description

• Take responsibility for the reliability, performance, and availability of production Aurora MySQL databases on AWS for the EHR, Data, and AI teams.

• Oversee the database alerting process from start to finish and maintain the observability pipeline that supplies metrics and alerts into Datadog.

• Integrate after-hours database alerts into the DevOps on-call rotation.

• Provide education and training to engineering teams on database hygiene and the creation of high-performing queries.

• Implement protective measures such as the automatic termination of long-running or runaway queries.

• Deliver rapid, always-accessible visibility into queries being executed on any database instance.

• Write and maintain database runbooks to ensure quick and consistent incident response.

• Analyze, optimize, and re-index problematic queries to minimize latency and manage scale-out load.

• Review EMR and application queries before release to serve as a performance checkpoint.

• Design and appropriately size the Aurora reader topology along with the read-routing strategy.

• Manage database cost efficiency through instance right-sizing, reserved capacity/Savings Plan strategies, and storage tiering.

• Supervise replication health, backups, restore testing, failover, and disaster recovery processes.

• Set up safe practices for schema changes and migrations across teams.

• Collaborate with the Platform Architect on database workload architecture.

• Handle protected health information responsibly in a HIPAA-compliant setting.


⛳️ Requirements

• A minimum of 6 years in database reliability engineering, database administration, or database engineering, with ownership of production systems at scale.

• Extensive expertise in MySQL, including query optimization, execution plan analysis, indexing strategies, and replication.

• Strong preference for experience with Aurora MySQL.

• Practical experience managing large, multi-reader Aurora or RDS clusters at terabyte scale, encompassing read-routing and connection management.

• Strong proficiency with database observability tools, including Datadog, Percona Monitoring and Management (PMM), Performance Insights, and Prometheus/Grafana.

• Skilled in Python, Bash, or a similar scripting language for automation purposes.

• Comfort with infrastructure-as-code, particularly with Terraform.

• Solid operational knowledge of AWS pertaining to RDS/Aurora, focusing on cost optimization, instance sizing, reserved capacity, and storage options.

• Experience with on-call/incident response and creating runbooks.

• Excellent communication skills and a collaborative approach across engineering teams.

• Preferred experience with cloud data warehouses such as Snowflake, Databricks, or Redshift.

• Preferred background in a HIPAA-regulated or compliance-heavy environment managing sensitive data.

• Familiarity with query auto-remediation tools such as pt-kill, Percona Toolkit, statement timeouts, or custom killers is preferred.

• Application-side familiarity with PHP is preferred.

• Knowledge of security best practices for databases and cloud environments is preferred.

• Previous experience in agile methodologies and fast-paced work environments is preferred.

• Must be legally authorized to work in the US.

• Must reside in the US.


🏝️ Benefits

• Potential equity compensation for exceptional performance.

• Flexible PTO.

• Sponsored lunches for the entire company.

• Company-paid disability and life insurance benefits.

• Company-funded family and medical leave.

• Medical, dental, and vision insurance benefits.

• Discounted pet insurance.

• FSA/DCA and commuter benefits.

• 401k plan.

• Credits for online fitness classes or gym memberships.

• Recovery suite at headquarters, featuring a cold plunge, sauna, and shower.

• Remote/hybrid work environment.

• Opportunity to make a positive impact by assisting outpatient rehabilitation organizations in treating more patients and delivering better care.

People also viewed

Horizon3.ai1 day ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH1 day ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM1 day ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies1 day ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)1 day ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC1 day ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers