Remotery

Senior Software Engineer – SRE, Observability Tooling

Posted Jul 16

This is a fully remote position, open to applicants in Estonia.

📋 Description

• Concentrate on developing SRE and observability tools — the platform, automation, and standards that other teams rely on to maintain their service health.

• Create standards, infrastructure, and automation for dashboards, alerts, and monitors as code.

• Collaborate with development teams to ensure production and operational readiness.

• Develop the tools and templates that teams utilize to define, measure, and report on Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for their services.

• Build tools to automate observability and operational workflows, reducing manual effort for engineering teams.

• Enhance the incident response tools and processes that assist teams in resolving outages more quickly and learning from them.


⛳️ Requirements

• Advanced proficiency in AWS and Kubernetes (EKS), especially in observability, networking, and auto-scaling.

• Familiarity with contemporary observability platforms (e.g., DataDog, Prometheus) and a comprehensive understanding of metrics, logging, and tracing.

• Extensive, hands-on experience with Site Reliability Engineering (SRE) principles (SLOs, error budgets, toil reduction).

• Proven experience in analyzing and troubleshooting large-scale distributed systems.

• Strong programming skills in a language like Python or Go, employed to develop operational tools, services, or automation.

• Expertise in designing and managing resilient CI/CD pipelines for a microservices architecture (e.g., using ArgoCD, Github Actions, Helm).

• A methodical, data-driven approach to problem-solving and root cause analysis.

• Ability to use AI tools judiciously, maintaining control over the final output while acknowledging the tools' limitations.


🏝️ Benefits

• Flexible remote collaboration

• Optional offices in Tallinn and Tartu

• In-person innovation and connection sessions conducted twice a year

People also viewed

Ontrac Solutions2 days ago

Site Reliability Engineer

PK flagPakistan OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
CyberSheath2 days ago

Cloud Operations Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$110k – $127k/year
ApplyView job
Ontrac Solutions2 days ago

Site Reliability Engineer

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
NVIDIA2 days ago

Service Reliability Engineer

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$168k – $333.5k/year
ApplyView job
Nagarro2 days ago

Senior Site Reliability Engineer, AWS Cloud

RO flagRomania OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Capgemini2 days ago

Senior DevOps Engineer

UA flagUkraine OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers