Remotery

Site Reliability Engineer

Posted 2 days ago

This is a fully remote position, open to applicants in Spain, +1 more state.

📋 Description

• Design, implement, and manage cloud infrastructure on AWS and Google Cloud platforms.

• Ensure the scalability, performance, and reliability of distributed database systems based on Kubernetes.

• Work alongside developers to create production-quality Golang code for infrastructure automation and system operations.

• Streamline and automate CI/CD pipelines, deployment processes, and monitoring systems.

• Formulate strategies for disaster recovery, high availability, and fault tolerance.

• Detect system bottlenecks and resolve issues across networking, operating systems, and cloud infrastructure.

• Set up monitoring, logging, and alerting systems to ensure visibility into system health and performance.

• Participate in on-call rotations and address production incidents as they arise.

• Collaborate with cross-functional teams to enhance system reliability and scalability.

• Work with the Customer Success team to address and resolve customer concerns.


⛳️ Requirements

• Demonstrated experience as an SRE or DevOps Engineer in a cloud-native setting.

• At least 3 years of experience with Kubernetes in a production environment.

• Minimum of 3 years of experience in deploying and managing production-level cloud resources.

• Expertise in Kubernetes for large-scale distributed systems.

• Familiarity with AWS and Google Cloud (GCP).

• Understanding of networking, security practices, troubleshooting techniques, and Linux internals.

• Experience with Docker and containerization technologies.

• Knowledge of CI/CD practices and tools like Jenkins and CircleCI.

• Familiarity with monitoring, alerting, and observability tools such as Prometheus, Grafana, and the ELK stack.

• Understanding of version control systems, especially Git.

• Familiarity with Golang or Python, with a readiness and ability to learn and apply Golang.

• Capacity to self-manage and work independently as part of a remote team.

• Experience managing distributed databases or large-scale data storage systems is advantageous.

• Knowledge of cloud security best practices is a plus.

• Proficiency with Python or Bash scripting is a bonus.

• Experience with Terraform/IaC is a plus.

• Knowledge of GitOps practices is a plus.

• Strong programming skills in Golang and experience in developing automation tools are advantageous.


🏝️ Benefits

• Remote work opportunities.

• Collaboration with seasoned engineers, marketers, and product leaders.

• Chance to contribute to innovative AI and data infrastructure projects.

• Opportunity to influence how enterprises create AI-powered applications.

• An inclusive, growth-oriented team environment.

People also viewed

CWILL15 hours ago

DevOps/SRE Engineer, Bilingual Mandarin

US flagCalifornia, +4 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$100k – $130k/year
ApplyView job
a3716 hours ago

Forward Deployed DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
GT17 hours ago

Site Reliability Engineer, SRE

PL flagPoland, +2 more statesFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sigma Software Group17 hours ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Applaudo17 hours ago

Google Cloud DevOps Engineer – Temporary Contract

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Branch17 hours ago

Cloud Operations Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$135k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers