
Site Reliability Engineer
Posted Aug 6

Posted Aug 6
This is a fully remote position, open to applicants in United States.
• Design, implement, and maintain cloud infrastructure on AWS and Google Cloud platforms.
• Ensure scalability, performance, and reliability of Kubernetes-based distributed database systems.
• Collaborate with developers to create production-grade Golang code for infrastructure automation and system operations.
• Optimize and automate CI/CD pipelines, deployment processes, and monitoring systems.
• Develop strategies for disaster recovery, high availability, and fault tolerance.
• Identify bottlenecks, troubleshoot, and resolve issues related to networking, operating systems, and cloud infrastructure.
• Implement monitoring, logging, and alerting systems.
• Participate in on-call rotations and respond to production incidents.
• Work with cross-functional teams to enhance reliability and scalability.
• Collaborate with Customer Success to address customer issues.
• Proven experience as an SRE or DevOps Engineer within a cloud-native environment.
• At least 3 years of experience working with Kubernetes in a production setting.
• Minimum of 3 years managing and deploying production-level cloud resources.
• Proficiency in Kubernetes for large-scale distributed systems.
• Familiarity with AWS and Google Cloud (GCP).
• Understanding of networking, security practices, and troubleshooting methodologies.
• Knowledge of Linux internals, including processes and environment variables.
• Experience with Docker and containerization technologies.
• Knowledge of CI/CD practices and tools such as Jenkins and CircleCI.
• Familiarity with Prometheus, Grafana, the ELK stack, and observability tools.
• Proficient in Git and version control systems.
• Familiarity with Golang or Python, or the willingness and ability to learn and use Golang.
• Capability to participate in on-call rotations.
• Strong troubleshooting, problem-solving, communication, and collaboration abilities.
• Ability to self-organize and work independently as part of a remote team.
• Nice-to-have: experience with distributed databases or large-scale data storage.
• Nice-to-have: knowledge of cloud security best practices.
• Nice-to-have: experience with Python or Bash scripting.
• Nice-to-have: familiarity with Terraform and Infrastructure-as-Code.
• Nice-to-have: experience with GitOps.
• Nice-to-have: strong programming experience in Golang.
• Competitive salary and performance-based bonuses.
• Flexible working hours and remote work options.
• Professional development opportunities and training.
• Comprehensive health and wellness benefits.
• Collaborative and inclusive company culture.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.