Site Reliability Engineer

Posted Aug 18

This is a fully remote position, open to applicants in Spain, +1 more country.

📋 Description

• Design, implement, and maintain cloud infrastructure utilizing AWS and Google Cloud platforms.

• Ensure the scalability, performance, and dependability of Kubernetes-based distributed database systems.

• Work alongside developers to produce efficient, production-ready Golang code for automating infrastructure and system operations.

• Streamline and automate CI/CD pipelines, deployment workflows, and monitoring systems.

• Formulate strategies for disaster recovery, high availability, and fault tolerance.

• Identify system bottlenecks and troubleshoot issues across the network, operating system, and cloud infrastructure stack.

• Implement monitoring, logging, and alerting systems.

• Participate in on-call rotations and address production incidents as they arise.

• Collaborate with cross-functional teams to enhance system reliability and scalability.

• Work with the Customer Success team to resolve client issues.


⛳️ Requirements

• At least 3 years of experience with Kubernetes in a production setting.

• Minimum of 3 years' experience in deploying and managing production-level resources in a cloud environment.

• Demonstrated experience as an SRE or DevOps Engineer within a cloud-native context.

• Proficient in Kubernetes for large-scale distributed systems.

• Experience with AWS and Google Cloud (GCP).

• Knowledge of networking, security practices, and troubleshooting techniques.

• Understanding of Linux internals, including processes and environment variables.

• Familiarity with containerization technologies, such as Docker.

• Acquainted with CI/CD practices and tools like Jenkins and CircleCI.

• Knowledge of alerting, monitoring, and observability tools, including Prometheus, Grafana, and the ELK stack.

• Familiarity with version control systems, particularly Git.

• Experience with Golang or Python, along with a willingness and capability to learn and apply Golang.

• Ability to self-organize and work independently as part of a remote team.

• Experience collaborating with geographically distributed teams.

• Residency in Europe is a requirement.


🏝️ Benefits

• Contributing to state-of-the-art AI and data infrastructure.

• Collaborating with skilled engineers, marketers, and product leaders.

• Helping shape the future of how enterprises construct AI-powered applications.

People also viewed

Entarian1 day ago

DevOps Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$150k – $200k/year
ApplyView job
BeyondTrust1 day ago

Senior Manager, Site Reliability Engineer – FedRAMP, AWS GovCloud

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
BeyondTrust1 day ago

Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Scribe1 day ago

Senior DevOps Engineer

US flagCalifornia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$150k – $240k/year
ApplyView job
Hypertegrity AG1 day ago

Senior DevOps Engineer – Smart City Open Source

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
CACI International Inc1 day ago

Senior DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$98.5k – $206.8k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers