
Database Reliability Engineer
Posted 14 hours ago

Posted 14 hours ago
This is a fully remote position, open to applicants in Massachusetts.
• Ensure database reliability across various platforms including PostgreSQL, MySQL, MongoDB, Redis, ScyllaDB, Aerospike, and managed cloud services.
• Maintain high availability, replication, partitioning, storage, and connection management.
• Design and improve automation-first database platforms.
• Create Kubernetes operators, implement infrastructure as code, manage GitOps workflows, and develop production-quality tools using Go or Python.
• Automate processes such as provisioning, failover, backups, schema migrations, and lifecycle management.
• Develop self-healing, fault-tolerant infrastructure and internal tools.
• Minimize operational toil and streamline daily database tasks.
• Enhance database performance and cost efficiency across both cloud and on-premises environments.
• Assist in capacity planning, resource efficiency, storage optimization, workload consolidation, and performance tuning for large-scale systems.
• Collaborate with application engineering teams on schema evaluations, migration plans, query optimization, connection management, and zero-downtime deployments.
• Utilize AI for observability, anomaly detection, root cause analysis, documentation, predictive insights, and assessment of AI-generated code.
• Minimum of 2 years of experience in Database Reliability Engineering, Database Platform Engineering, Site Reliability Engineering, or a similar infrastructure engineering position with a strong database emphasis.
• Experience with at least one major relational database, preferably Aurora MySQL.
• Operational familiarity with PostgreSQL, MongoDB, ScyllaDB, Aurora, or other managed cloud database services.
• Proficiency in building and managing stateful workloads on Kubernetes.
• Knowledge of StatefulSets, Persistent Volumes, database operators, Terraform, Pulumi, FluxCD, ArgoCD, GKE, or EKS.
• Practical software development experience with Go or Python.
• Experience in creating automation, platform tooling, Kubernetes controllers, APIs, and infrastructure.
• Proven ability to enhance reliability through observability, monitoring, service-level objectives, capacity planning, performance optimization, and self-service engineering solutions.
• Hands-on experience with AI tools such as Claude, GitHub Copilot, Cursor, MCP, or similar technologies.
• Strong engineering judgment for validating AI-generated results.
• May need to acquire a gaming license from the appropriate state agency as a condition of employment.
• Bonus
• Equity
• Benefits as applicable
• Assistance with the gaming license process if relevant to the role
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.