Remotery

Senior Site Reliability Engineer

Posted 1 day ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Manage AWS services, accounts, access, PostgreSQL, and various data stores.

• Take ownership of the backup strategy across databases, S3 buckets, and queues.

• Confirm data restorations and uphold a verified disaster recovery strategy.

• Oversee production monitoring using CloudWatch dashboards, metric alerts, log-based metrics, and Slack notifications.

• Spearhead production troubleshooting and incident management.

• Create and sustain runbooks while engaging in the on-call rotation.

• Address queue and dead-letter-queue issues through retries, redrives, and recovery procedures.

• Enhance infrastructure for improved deployability and scalability.

• Maintain infrastructure as code, decommission unused resources, and ensure cost transparency and justification.

• Disseminate production operations knowledge among the team.


⛳️ Requirements

• Bachelor’s degree and 4–6 years of relevant experience or equivalent professional background.

• Over 5 years of experience in DevOps, site reliability, or platform operations, with a significant focus on production systems.

• More than 3 years of practical experience with AWS, particularly in serverless services such as Lambda, SQS, EventBridge, CloudWatch, and S3.

• Strong expertise in PostgreSQL database administration, including backup and recovery, and query performance optimization.

• Familiarity with managing other data stores.

• Proficient in TypeScript, Python, and bash scripting.

• Solid understanding of Linux, DNS, TLS, Docker, GitHub Actions, and infrastructure as code utilizing SST, Pulumi, or Terraform.

• Experience in production monitoring and alerting, incident response, and on-call responsibilities.

• Must possess legal authorization to work in the United States; the application will inquire about future sponsorship needs.

• Must successfully complete comprehensive background, credit, and drug screenings as part of the hiring process.


🏝️ Benefits

• Comprehensive insurance coverage (medical, dental, vision, life, and disability).

• Flexible paid time off.

• Paid holidays.

• 401(k) plan with company match.

• Opportunity for remote work.

People also viewed

CVS Health7 hours ago

Salesforce DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$83.4k – $166.9k/year
ApplyView job
Devoteam9 hours ago

Data, AWS DevSecOps

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Aspirion9 hours ago

Senior DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Goodgame Studios9 hours ago

Senior Agentic Engineer – Java Backend, DevOps

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Instacart9 hours ago

Site Reliability Engineer II

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$133k – $169k/year
ApplyView job
Logicalis Spain10 hours ago

DevOps Engineer

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€40k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers