Remotery

Expert DevOps – Observability Operations

Posted 2 days ago

This is a fully remote position, open to applicants in Germany.

📋 Description

• Implement changes, upgrades, maintenance, and recovery procedures for the Kubernetes platform in line with approved protocols, ensuring accurate documentation of implementations and executions.

• Configure, enhance, and sustain observability features, which encompass metrics, logs, distributed tracing, health checks, dashboards, and alerting criteria.

• Investigate operational irregularities and incidents; produce analyses of root causes, recommend restoration actions, and outline improvement needs.

• Develop and maintain operational playbooks, standard operating procedures, incident management guides, escalation criteria, and knowledge resources.

• Set up and deliver components for platform observability and security, accompanied by implementation documentation and configuration records.

• Generate and maintain operational documents, decision logs, handover packages, and an Operational Manual for the supported environments.

• Create and provide operational governance guidelines, training resources, and knowledge transfer packages.

• Prepare applications for productive use on new infrastructure.

• Establish processes for monitoring, alerting, support, automation, and incident management.

• Facilitate stable, scalable, and maintainable operations while promoting operational ownership through manuals and playbooks.


⛳️ Requirements

• Over 8 years of experience in DevOps, specifically with containerized Java application services and servers.

• In-depth knowledge of the Grafana LGTM stack, including Grafana, Prometheus, Loki, Mimir, and Tempo.

• Extensive understanding of Kubernetes infrastructure.

• Strong knowledge of ArgoCD or similar continuous deployment tools.

• Significant experience and knowledge in incident and problem management.

• Proficient in shell scripting with Windows PowerShell and Linux.

• Familiarity with Splunk.

• Proficiency in English at a C1 level.

• Proficiency in German at a C1 level.

• Deep understanding of deployment concepts for complex microservice-oriented Java application platforms.

• Experience in application development.

• Background in L2 support for complex business applications.

• Proven experience in creating and maintaining operations playbooks.

• Experience in IT provider management, IT monitoring, and IT operations management.

• Familiarity with Kafka, MySQL, PostgreSQL, SPARK, Harness, Helm, JFrog Artifactory, Azure DevOps, and OpenTelemetry.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Flexible working hours and remote work options.

• Professional development and training opportunities.

• Health and wellness programs.

• Collaborative and innovative work environment.

People also viewed

CWILL14 hours ago

DevOps/SRE Engineer, Bilingual Mandarin

US flagCalifornia, +4 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$100k – $130k/year
ApplyView job
a3715 hours ago

Forward Deployed DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
GT16 hours ago

Site Reliability Engineer, SRE

PL flagPoland, +2 more statesFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sigma Software Group16 hours ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Applaudo16 hours ago

Google Cloud DevOps Engineer – Temporary Contract

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Branch16 hours ago

Cloud Operations Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$135k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers