Remotery

DevOps Engineer, Observability

Posted Aug 5

This is a fully remote position, open to applicants in Ireland.

📋 Description

• Oversee the complete architecture and implementation of essential observability platform components, emphasizing reliability, scalability, and user experience.

• Foster uniformity and excellence across logs, metrics, traces, and continuous profiling.

• Act as a technical mentor and advisor within the platform organization.

• Steer design choices and synchronize cross-team initiatives with overarching architectural objectives.

• Tackle challenges related to high-cardinality telemetry, distributed tracing correlation, and insights into compute costs.

• Collaborate with product teams, Site Reliability Engineers (SREs), and developer experience groups.

• Create and develop developer-friendly tools and APIs for incident response, performance evaluation, and platform troubleshooting.

• Utilize open-source standards like OpenTelemetry.

• Find a balance between performance, expenses, and user value across engineering teams.


⛳️ Requirements

• Demonstrated experience in building and scaling observability systems.

• Familiarity with centralized S3-based data lakes, OpenTelemetry instrumentation, and ClickHouse-powered query engines.

• Proficient in at least one contemporary programming language, such as Go, Python, or Java.

• Understanding of challenges associated with high-cardinality data and telemetry correlation techniques.

• Experience in designing high-scale telemetry systems like Prometheus, ClickHouse, OpenTelemetry, or Kafka.

• Strong grasp of distributed systems and observability in intricate microservice environments.

• Background with AWS, Kubernetes, and infrastructure-as-code tools.

• Capacity to offer architectural direction and technical thought leadership across various teams.

• Ability to make strategic technical decisions and navigate through uncertainty.

• Experience with ClickHouse, Grafana Mimir, Athena, or similar systems is preferred.

• Contributions to open-source observability projects are highly regarded.

• Familiarity with cloud computing and telemetry FinOps tools is advantageous.


🏝️ Benefits

• Competitive salary.

• Generous vacation policy.

• Extensive parental leave.

• Wellness leave.

• Comprehensive healthcare coverage.

• Retirement savings plan.

• Support for employee volunteering and donation initiatives.

• Remote-first work environment.

People also viewed

DATAGROUP1 day ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush1 day ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey1 day ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems2 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems2 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data2 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers