
Senior Platform Operations Engineer / Site Reliability Engineer
Posted Aug 28

Posted Aug 28
This is a fully remote position, open to applicants in Germany.
• Manage, maintain, and enhance monitoring and observability platforms.
• Administer and optimize tools such as Prometheus, Grafana, and OpenSearch / ELK.
• Oversee production environments while continuously improving monitoring and alerting strategies.
• Address incidents and assist in the recovery of essential services.
• Perform root cause analyses to permanently resolve incident triggers.
• Support major incidents and coordinate technical remediation efforts.
• Operate and enhance containerized platforms utilizing Kubernetes.
• Facilitate CI/CD processes with Jenkins and ArgoCD.
• Develop, maintain, and continually refine runbooks, operational processes, and technical documentation.
• Automate repetitive operational tasks.
• Engage in on-call rotations and shift schedules within a 24/7 operational organization.
• Several years of experience in Platform Operations, Site Reliability Engineering, Systems Engineering, or IT Operations.
• Willingness to obtain, or already possess, a German SÜ2 security clearance.
• Strong knowledge of Linux-based environments.
• Familiarity with Kubernetes and containerized platforms.
• Practical experience with Prometheus.
• Practical experience with Grafana.
• Practical experience with the ELK Stack or OpenSearch.
• Experience with Elasticsearch or OpenSearch.
• Background in monitoring, alerting, and observability environments.
• Good understanding of networking fundamentals and communication protocols.
• Experience working with REST APIs.
• Proficiency in Git.
• Analytical mindset for troubleshooting and incident resolution.
• Proficient in written and spoken English.
• Willingness to engage in on-call rotations and shift work.
• Nice to have: Experience with ArgoCD.
• Nice to have: Knowledge of Jenkins.
• Nice to have: Experience with Helm.
• Nice to have: Bash scripting skills.
• Nice to have: Python experience for operational automation and excellence.
• Nice to have: Experience in Site Reliability Engineering (SRE).
• Nice to have: Knowledge of modern cloud or platform architectures.
• Permanent employment contract.
• Attractive compensation package.
• Opportunities for professional development and training.
• Options for hybrid and fully remote work.
• Provision of PC equipment.
FourEnergy GmbH
ICF
Mastercam
C&S Informática
Get handpicked remote jobs straight to your inbox weekly.