Remotery

Observability Engineer – Site Reliability Engineer

Posted 7 hours ago

This is a fully remote position, open to applicants in California, +2 more states.

📋 Description

• Architect, optimize, and sustain observability frameworks across various cloud environments, specifically focusing on the implementation of Google Cloud Platform (GCP) observability tools (Cloud Logging, Cloud Monitoring, Trace, and Profiler).

• Design, deploy, and uphold robust observability stacks within hybrid ecosystems, utilizing Prometheus, Grafana, and cloud-native integrations.

• Lead infrastructure-as-code (IaC) initiatives utilizing Terraform and Ansible to guarantee consistent and automated deployments of infrastructure and observability tools.

• Construct, maintain, and enhance deployment workflows within Kubernetes and Google Kubernetes Engine (GKE) / OpenShift environments using GitHub, Harness, and other CI/CD pipelines.

• Conduct in-depth analysis of Linux/Unix system administration architectures, optimizing compute resource metrics and performance tuning across intricate, distributed environments.

• Implement Site Reliability Engineering (SRE) best practices, establishing meaningful Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to ensure platform reliability.


⛳️ Requirements

• Demonstrated engineering experience within Google Cloud Platform (GCP) environments, particularly in managing cloud-native monitoring and compute resources.

• Practical experience with Grafana, Prometheus, and Google Cloud Observability suites.

• Expert-level understanding of Linux/Unix operating systems combined with strong shell scripting capabilities for automation and systems management.

• Professional coding proficiency in at least one modern programming language (Python, Go, Java, Perl, or advanced Shell).

• Hands-on experience in managing containerized applications on Kubernetes, GKE, and/or Red Hat OpenShift.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health, dental, and vision insurance options.

• Flexible work hours and remote working opportunities.

• Professional development and continuous learning programs.

People also viewed

SOFTETA7 hours ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sólides7 hours ago

Senior Site Reliability Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sólides7 hours ago

Senior Site Reliability Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Zact7 hours ago

Cloud DevSecOps Engineer

US flagCalifornia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Tinybird7 hours ago

Site Reliability Engineer

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€58k – €97k/year
ApplyView job
3Pillar Global7 hours ago

Senior DevOps Engineer – AWS, Python

MX flagMexico OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers