
Site Reliability Engineer – IV
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in United States.
• Establish the technical strategy for managing, automating, and scaling the platform.
• Lead incident management for intricate outages and pinpoint genuine root causes.
• Implement lasting solutions and enhance runbooks.
• Engage in a weekly on-call rotation shared with three other engineers.
• Debug and troubleshoot throughout the technology stack.
• Ensure that engineering tasks are well-documented and maintainable.
• Take ownership of CI/CD build and deployment pipelines.
• Manage systems comprehensively, covering architecture, design, optimization, and long-term maintenance.
• Advocate for AI and automation within operations.
• Develop tools and prototypes to remove repetitive manual tasks.
• Facilitate architecture discussions and design reviews.
• Promote the adoption of new technologies and methodologies.
• Collaborate across infrastructure, developer platforms, and security.
• Define the technical vision and architectural framework.
• Mentor engineers and elevate the team's technical proficiency.
• Over 10 years of experience in SRE/DevOps/systems.
• Proven history of owning and architecting production systems at scale.
• Strong programming skills in Python, Go, or Java; proficient in Bash.
• Extensive knowledge of Linux and networking, including TCP/IP, DNS, and HTTP/TLS.
• Significant experience with AWS, Docker, and Terraform.
• Demonstrated ability to set technical direction and execute an architectural strategy.
• Strong critical thinking and troubleshooting skills.
• Exceptional communication abilities.
• Prior experience as an architect or tech lead is preferred.
• Software development background is advantageous.
• Hands-on experience in software engineering is preferred.
• Familiarity with production container orchestration tools such as Nomad or Kubernetes is preferred.
• Experience in developing AI-assisted tools or automating platforms is a plus.
• Depth of knowledge in security and operational hygiene is preferred.
• Exposure to PCI or similar frameworks is an advantage.
• Comfortable working with modern cloud technologies as well as established systems.
• Bachelor's degree in Computer Science, Computer Engineering, or a related field, or equivalent experience is preferred.
• Must be able to work standard collaboration hours from 10:00 AM to 3:00 PM MT, Monday to Friday; additional hours must be coordinated and approved by the manager.
• Competitive salary.
• Bonuses based on company performance plans.
• Annual merit increases.
• Medical, dental, and vision insurance for employees and eligible dependents.
• Flexible spending accounts (FSA).
• Health savings accounts (HSA).
• Life insurance.
• Short-term and long-term disability insurance.
• Up to 15 weeks of paid parental leave for full-time employees.
• 401(k) plan with employer matching.
• Fully remote work from anywhere within the Continental USA.
• Flexible work hours outside of standard collaboration times.
• Exclusive incentive purchase programs.
• Potential for a 10% annual bonus.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.