
Site Reliability Engineer
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Spain, +1 more state.
• Design, implement, and maintain cloud infrastructure utilizing AWS and Google Cloud platforms.
• Ensure the scalability, performance, and dependability of Kubernetes-based distributed database systems.
• Work alongside developers to produce efficient, production-ready Golang code for automating infrastructure and system operations.
• Streamline and automate CI/CD pipelines, deployment workflows, and monitoring systems.
• Formulate strategies for disaster recovery, high availability, and fault tolerance.
• Identify system bottlenecks and troubleshoot issues across the network, operating system, and cloud infrastructure stack.
• Implement monitoring, logging, and alerting systems.
• Participate in on-call rotations and address production incidents as they arise.
• Collaborate with cross-functional teams to enhance system reliability and scalability.
• Work with the Customer Success team to resolve client issues.
• At least 3 years of experience with Kubernetes in a production setting.
• Minimum of 3 years' experience in deploying and managing production-level resources in a cloud environment.
• Demonstrated experience as an SRE or DevOps Engineer within a cloud-native context.
• Proficient in Kubernetes for large-scale distributed systems.
• Experience with AWS and Google Cloud (GCP).
• Knowledge of networking, security practices, and troubleshooting techniques.
• Understanding of Linux internals, including processes and environment variables.
• Familiarity with containerization technologies, such as Docker.
• Acquainted with CI/CD practices and tools like Jenkins and CircleCI.
• Knowledge of alerting, monitoring, and observability tools, including Prometheus, Grafana, and the ELK stack.
• Familiarity with version control systems, particularly Git.
• Experience with Golang or Python, along with a willingness and capability to learn and apply Golang.
• Ability to self-organize and work independently as part of a remote team.
• Experience collaborating with geographically distributed teams.
• Residency in Europe is a requirement.
• Contributing to state-of-the-art AI and data infrastructure.
• Collaborating with skilled engineers, marketers, and product leaders.
• Helping shape the future of how enterprises construct AI-powered applications.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.