
Site Reliability Engineer
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Spain, +1 more state.
• Design, implement, and manage cloud infrastructure on AWS and Google Cloud platforms.
• Ensure the scalability, performance, and reliability of distributed database systems based on Kubernetes.
• Work alongside developers to create production-quality Golang code for infrastructure automation and system operations.
• Streamline and automate CI/CD pipelines, deployment processes, and monitoring systems.
• Formulate strategies for disaster recovery, high availability, and fault tolerance.
• Detect system bottlenecks and resolve issues across networking, operating systems, and cloud infrastructure.
• Set up monitoring, logging, and alerting systems to ensure visibility into system health and performance.
• Participate in on-call rotations and address production incidents as they arise.
• Collaborate with cross-functional teams to enhance system reliability and scalability.
• Work with the Customer Success team to address and resolve customer concerns.
• Demonstrated experience as an SRE or DevOps Engineer in a cloud-native setting.
• At least 3 years of experience with Kubernetes in a production environment.
• Minimum of 3 years of experience in deploying and managing production-level cloud resources.
• Expertise in Kubernetes for large-scale distributed systems.
• Familiarity with AWS and Google Cloud (GCP).
• Understanding of networking, security practices, troubleshooting techniques, and Linux internals.
• Experience with Docker and containerization technologies.
• Knowledge of CI/CD practices and tools like Jenkins and CircleCI.
• Familiarity with monitoring, alerting, and observability tools such as Prometheus, Grafana, and the ELK stack.
• Understanding of version control systems, especially Git.
• Familiarity with Golang or Python, with a readiness and ability to learn and apply Golang.
• Capacity to self-manage and work independently as part of a remote team.
• Experience managing distributed databases or large-scale data storage systems is advantageous.
• Knowledge of cloud security best practices is a plus.
• Proficiency with Python or Bash scripting is a bonus.
• Experience with Terraform/IaC is a plus.
• Knowledge of GitOps practices is a plus.
• Strong programming skills in Golang and experience in developing automation tools are advantageous.
• Remote work opportunities.
• Collaboration with seasoned engineers, marketers, and product leaders.
• Chance to contribute to innovative AI and data infrastructure projects.
• Opportunity to influence how enterprises create AI-powered applications.
• An inclusive, growth-oriented team environment.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.