
Site Reliability Engineer II, DBA
Posted 20 hours ago

Posted 20 hours ago
This is a fully remote position, open to applicants in Argentina, +3 more countries.
• Manage and oversee high-availability database systems, primarily utilizing Vitess (distributed MySQL) and Cassandra, following the prescribed architecture and runbooks.
• Enhance database performance through effective query optimization, indexing techniques, and schema design.
• Implement established backup, recovery, and replication protocols, and escalate architectural changes to senior DBA SREs as needed.
• Ensure compliance with database security standards and manage access controls.
• Support the availability and reliability of essential services across production environments.
• Monitor service health metrics using SLIs, SLOs, and error budgets, raising alerts when thresholds are in jeopardy.
• Engage in on-call rotations, incident response activities, and post-incident evaluations.
• Adhere to established ITIL/OSS protocols for incident, change, problem, and capacity management.
• Create automation for routine operational tasks to minimize manual efforts and reduce toil.
• Contribute to monitoring, logging, and alerting frameworks such as Prometheus, Grafana, Catchpoint, and ELK, integrating runbooks with FireHydrant.
• Collaborate with CI/CD pipelines, configuration management, and infrastructure as code tools including Terraform, Ansible, and Jenkins.
• Develop scripts in Bash, Python, or Go to enhance system reliability and efficiency.
• Partner with engineering, product, and operations teams to design and operate resilient systems.
• Assist in capacity planning and participate in disaster recovery drills.
• Coordinate with vendors and service providers to troubleshoot service issues and monitor SLA performance.
• Document systems, share insights, and foster a culture focused on reliability in engineering.
• Contribute to the creation of playbooks, runbooks, and operational documentation.
• Identify recurring problems and suggest long-term enhancements.
• Advocate for reliability-oriented practices within development and operations teams.
• Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent experience).
• 2–4 years of experience in site reliability, systems engineering, or operations focused on database systems.
• Familiarity with large-scale, production-level systems.
• Strong Linux systems administration and troubleshooting capabilities.
• Knowledge of monitoring, alerting, incident response, and root cause analysis.
• Proficient in at least one scripting language: Python, Bash, or Go.
• Understanding of Kubernetes, Docker, and microservices principles.
• Awareness of incident response and operational best practices.
• Practical experience with MySQL performance tuning, replication, and disaster recovery strategies.
• Proficient in SQL and NoSQL database management.
• Experience in a SaaS, service provider, or distributed systems setting.
• Knowledge of ITIL/OSS methodologies and SLO/SLAs.
• Familiarity with cloud platforms such as AWS, GCP, or Azure.
• Ability to work autonomously, take ownership, and lead projects from the identification of problems through to their resolution.
• The job posting does not specify any particular benefits, perks, or additional compensation.
• Offers a remote work arrangement for candidates located in Argentina, Colombia, Costa Rica, or Mexico.
Ninja - نينجا
Dreamix
Scalingo
inDrive
Get handpicked remote jobs straight to your inbox weekly.