Site Reliability Engineer

Posted Aug 10

This is a fully remote position, open to applicants anywhere in the world.

📋 Description

• Establish the technical vision for reliability throughout Yuno’s infrastructure, starting with the AWS platform that provisions, deploys, and manages AI agents at scale.

• Take ownership of the platform reliability strategy, encompassing architectural choices, reliability metrics, and engineering standards.

• Cultivate a culture of Service Level Objectives (SLO), error-budget policies, and incident management across engineering teams.

• Design and maintain robust, reliable asynchronous messaging systems for inter-service communication.

• Manage cloud infrastructure and automate provisioning utilizing Infrastructure as Code.

• Ensure the platform scales effectively as transaction volumes increase.

• Develop monitoring, tracing, and alerting systems to maintain platform health.

• Act as the senior escalation point for complex production incidents.

• Conduct blameless postmortems and root-cause analyses that yield permanent solutions.

• Perform ongoing fault injection and resilience testing.

• Mentor senior and mid-level engineers while elevating organization-wide reliability standards.


⛳️ Requirements

• Over 7 years of relevant experience.

• Designed and managed event-driven systems utilizing message queues such as Kafka, NATS, or RabbitMQ.

• Familiarity with at-least-once delivery, consumer groups, dead letters, and backpressure concepts.

• Experience transitioning systems from synchronous to asynchronous communication.

• Extensive AWS experience with EC2, VPC, IAM, S3, and RDS.

• Strong foundational knowledge in networking.

• Experience with Infrastructure as Code using Terraform or Pulumi.

• Proficiency in Kubernetes and Docker in a production environment, including container lifecycle management, resource limits, health checks, and orchestration at scale.

• Proficient in Datadog or equivalent tools, including dashboards, monitors, Application Performance Monitoring (APM), and distributed tracing.

• Proven experience in defining and managing SLOs, SLIs, and error budgets across services.

• Hands-on experience with fault injection, game days, or chaos engineering using tools such as Gremlin, Chaos Mesh, AWS FIS, or similar.

• Experience debugging distributed systems.

• Comfortable writing automation and tools in Go, Python, or similar languages.

• Solid understanding of SQL and PostgreSQL.

• NoSQL experience with MongoDB and Redis, including indexing, replication, and performance optimization.

• Proven ability in technical leadership, influencing architecture across teams, and providing engineering mentorship.

• Advanced proficiency in written and spoken English.


🏝️ Benefits

• Competitive Compensation.

• Remote Work — flexibility to work from anywhere.

• Home Office Bonus — a one-time allowance to create your ideal home office setup.

• Work Equipment provided.

• Stock Options available.

• Health Plan coverage wherever you are located.

• Flexible Days Off.

• Opportunities for Language, Professional, and Personal Growth courses.

People also viewed

Koniag Government Services2 days ago

Architect/DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
FP Markets (First Prudential Markets)2 days ago

Senior DevOps Engineer

AM flagArmenia, +4 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Modern Campus2 days ago

Senior DevOps Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
InRule2 days ago

Site Reliability Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Thumbtack2 days ago

Senior Software Engineer, Site Reliability Engineering

US flagUnited States, +38 more locationsFull-timeDevOps & Site Reliability Engineer (SRE)$179.4k – $272.8k/year
ApplyView job
Thumbtack2 days ago

Senior Software Engineer, Site Reliability Engineering

CA flagCanada, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)C$180.2k – C$233.2k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers