Database Reliability Engineer

Posted 14 hours ago

This is a fully remote position, open to applicants in Massachusetts.

📋 Description

• Ensure database reliability across various platforms including PostgreSQL, MySQL, MongoDB, Redis, ScyllaDB, Aerospike, and managed cloud services.

• Maintain high availability, replication, partitioning, storage, and connection management.

• Design and improve automation-first database platforms.

• Create Kubernetes operators, implement infrastructure as code, manage GitOps workflows, and develop production-quality tools using Go or Python.

• Automate processes such as provisioning, failover, backups, schema migrations, and lifecycle management.

• Develop self-healing, fault-tolerant infrastructure and internal tools.

• Minimize operational toil and streamline daily database tasks.

• Enhance database performance and cost efficiency across both cloud and on-premises environments.

• Assist in capacity planning, resource efficiency, storage optimization, workload consolidation, and performance tuning for large-scale systems.

• Collaborate with application engineering teams on schema evaluations, migration plans, query optimization, connection management, and zero-downtime deployments.

• Utilize AI for observability, anomaly detection, root cause analysis, documentation, predictive insights, and assessment of AI-generated code.


⛳️ Requirements

• Minimum of 2 years of experience in Database Reliability Engineering, Database Platform Engineering, Site Reliability Engineering, or a similar infrastructure engineering position with a strong database emphasis.

• Experience with at least one major relational database, preferably Aurora MySQL.

• Operational familiarity with PostgreSQL, MongoDB, ScyllaDB, Aurora, or other managed cloud database services.

• Proficiency in building and managing stateful workloads on Kubernetes.

• Knowledge of StatefulSets, Persistent Volumes, database operators, Terraform, Pulumi, FluxCD, ArgoCD, GKE, or EKS.

• Practical software development experience with Go or Python.

• Experience in creating automation, platform tooling, Kubernetes controllers, APIs, and infrastructure.

• Proven ability to enhance reliability through observability, monitoring, service-level objectives, capacity planning, performance optimization, and self-service engineering solutions.

• Hands-on experience with AI tools such as Claude, GitHub Copilot, Cursor, MCP, or similar technologies.

• Strong engineering judgment for validating AI-generated results.

• May need to acquire a gaming license from the appropriate state agency as a condition of employment.


🏝️ Benefits

• Bonus

• Equity

• Benefits as applicable

• Assistance with the gaming license process if relevant to the role

People also viewed

Horizon3.ai12 hours ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH12 hours ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM12 hours ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies12 hours ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)13 hours ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC13 hours ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers