
Site Reliability Engineer
Posted Aug 10

Posted Aug 10
This is a fully remote position, open to applicants anywhere in the world.
• Establish the technical vision for reliability throughout Yuno’s infrastructure, starting with the AWS platform that provisions, deploys, and manages AI agents at scale.
• Take ownership of the platform reliability strategy, encompassing architectural choices, reliability metrics, and engineering standards.
• Cultivate a culture of Service Level Objectives (SLO), error-budget policies, and incident management across engineering teams.
• Design and maintain robust, reliable asynchronous messaging systems for inter-service communication.
• Manage cloud infrastructure and automate provisioning utilizing Infrastructure as Code.
• Ensure the platform scales effectively as transaction volumes increase.
• Develop monitoring, tracing, and alerting systems to maintain platform health.
• Act as the senior escalation point for complex production incidents.
• Conduct blameless postmortems and root-cause analyses that yield permanent solutions.
• Perform ongoing fault injection and resilience testing.
• Mentor senior and mid-level engineers while elevating organization-wide reliability standards.
• Over 7 years of relevant experience.
• Designed and managed event-driven systems utilizing message queues such as Kafka, NATS, or RabbitMQ.
• Familiarity with at-least-once delivery, consumer groups, dead letters, and backpressure concepts.
• Experience transitioning systems from synchronous to asynchronous communication.
• Extensive AWS experience with EC2, VPC, IAM, S3, and RDS.
• Strong foundational knowledge in networking.
• Experience with Infrastructure as Code using Terraform or Pulumi.
• Proficiency in Kubernetes and Docker in a production environment, including container lifecycle management, resource limits, health checks, and orchestration at scale.
• Proficient in Datadog or equivalent tools, including dashboards, monitors, Application Performance Monitoring (APM), and distributed tracing.
• Proven experience in defining and managing SLOs, SLIs, and error budgets across services.
• Hands-on experience with fault injection, game days, or chaos engineering using tools such as Gremlin, Chaos Mesh, AWS FIS, or similar.
• Experience debugging distributed systems.
• Comfortable writing automation and tools in Go, Python, or similar languages.
• Solid understanding of SQL and PostgreSQL.
• NoSQL experience with MongoDB and Redis, including indexing, replication, and performance optimization.
• Proven ability in technical leadership, influencing architecture across teams, and providing engineering mentorship.
• Advanced proficiency in written and spoken English.
• Competitive Compensation.
• Remote Work — flexibility to work from anywhere.
• Home Office Bonus — a one-time allowance to create your ideal home office setup.
• Work Equipment provided.
• Stock Options available.
• Health Plan coverage wherever you are located.
• Flexible Days Off.
• Opportunities for Language, Professional, and Personal Growth courses.
SYNCREON
Rimutee
Mirantis
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.