
Senior Site Reliability Engineer, Databases
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Alabama, +32 more states.
• Take ownership and enhance a thorough monitoring and alerting system for MySQL InnoDB Clusters and PostgreSQL infrastructure spanning multiple data centers.
• Develop and implement a quarterly backup verification and disaster recovery testing program across all database systems.
• Act as the primary on-call responder for database incidents, following established runbooks and escalating to senior engineering for architecture-level decisions as needed.
• Oversee and ensure the health of database replication.
• Manage the database user lifecycle, which includes provisioning, deprovisioning, conducting access audits, and implementing role-based access control.
• Carry out and monitor security compliance remediation efforts, addressing pen-test findings, GRC audit requirements, and encryption-at-rest verification.
• Coordinate with the Kafka infrastructure team to manage the operational health of the Debezium/Kafka Connect data pipeline.
• Develop and maintain Puppet profiles for the management of database infrastructure configuration.
• Create automation tools using PHP, Python, or Go to minimize operational toil.
• Generate and update runbooks, operational documentation, and disaster recovery procedures.
• Collaborate with the Senior Platform Engineer (Databases) as an equal, reviewing each other's work, sharing on-call responsibilities, and jointly ensuring database reliability across production systems.
• Over 7 years of experience in Site Reliability Engineering, DevOps, or Database Operations roles within large-scale production environments.
• Extensive operational knowledge of MySQL in production, including InnoDB Cluster, Group Replication, MySQL Router, and ProxySQL, with robust troubleshooting abilities for replication, performance, and reliability challenges.
• Hands-on experience with PostgreSQL in a production setting, encompassing replication, high availability, performance tuning, and operational management.
• Strong skills in configuration management tools (preferably Puppet) and infrastructure-as-code methodologies.
• Familiarity with database backup tools (such as xtrabackup, pg_dump, mysqldump) and disaster recovery processes.
• Proficient in PHP and Python or Go for automation, tooling, and integrating with existing codebases.
• Experience with database observability tools, including Prometheus exporters, Grafana dashboards, alerting frameworks, and SLO/error-budget frameworks.
• Knowledge of Kafka Connect, Debezium, or similar change-data-capture pipelines.
• Strong incident response capabilities, including participation in on-call rotations and conducting post-incident analysis and remediation.
• Exceptional communication skills and the ability to collaborate effectively across engineering teams as a senior peer.
• Must reside in one of the specified U.S. states.
• Must be legally authorized to work in the United States.
• 100% company-covered insurance premiums for employee medical, dental, and vision plans.
• A 401(k) plan that matches 100% up to 4%, with immediate vesting.
• Professional Development Reimbursement of $2,500 annually.
• 11 Holidays along with Paid Time Off Accrual and Rollover Plan.
• Increased PTO at the 3-year and 10-year anniversaries.
• A 1-month paid sabbatical every 5 years.
• An Anniversary Bonus each year.
• A $500 stipend for remote office setup in the first year, plus $400 each subsequent year.
• Internet reimbursement of up to $75 per month.
• Gym membership reimbursement of up to $50 per month.
• Company-paid Wellable subscription.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.