Site Reliability Engineer – IV

Posted 4 days ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Establish the technical strategy for managing, automating, and scaling the platform.

• Lead incident management for intricate outages and pinpoint genuine root causes.

• Implement lasting solutions and enhance runbooks.

• Engage in a weekly on-call rotation shared with three other engineers.

• Debug and troubleshoot throughout the technology stack.

• Ensure that engineering tasks are well-documented and maintainable.

• Take ownership of CI/CD build and deployment pipelines.

• Manage systems comprehensively, covering architecture, design, optimization, and long-term maintenance.

• Advocate for AI and automation within operations.

• Develop tools and prototypes to remove repetitive manual tasks.

• Facilitate architecture discussions and design reviews.

• Promote the adoption of new technologies and methodologies.

• Collaborate across infrastructure, developer platforms, and security.

• Define the technical vision and architectural framework.

• Mentor engineers and elevate the team's technical proficiency.


⛳️ Requirements

• Over 10 years of experience in SRE/DevOps/systems.

• Proven history of owning and architecting production systems at scale.

• Strong programming skills in Python, Go, or Java; proficient in Bash.

• Extensive knowledge of Linux and networking, including TCP/IP, DNS, and HTTP/TLS.

• Significant experience with AWS, Docker, and Terraform.

• Demonstrated ability to set technical direction and execute an architectural strategy.

• Strong critical thinking and troubleshooting skills.

• Exceptional communication abilities.

• Prior experience as an architect or tech lead is preferred.

• Software development background is advantageous.

• Hands-on experience in software engineering is preferred.

• Familiarity with production container orchestration tools such as Nomad or Kubernetes is preferred.

• Experience in developing AI-assisted tools or automating platforms is a plus.

• Depth of knowledge in security and operational hygiene is preferred.

• Exposure to PCI or similar frameworks is an advantage.

• Comfortable working with modern cloud technologies as well as established systems.

• Bachelor's degree in Computer Science, Computer Engineering, or a related field, or equivalent experience is preferred.

• Must be able to work standard collaboration hours from 10:00 AM to 3:00 PM MT, Monday to Friday; additional hours must be coordinated and approved by the manager.


🏝️ Benefits

• Competitive salary.

• Bonuses based on company performance plans.

• Annual merit increases.

• Medical, dental, and vision insurance for employees and eligible dependents.

• Flexible spending accounts (FSA).

• Health savings accounts (HSA).

• Life insurance.

• Short-term and long-term disability insurance.

• Up to 15 weeks of paid parental leave for full-time employees.

• 401(k) plan with employer matching.

• Fully remote work from anywhere within the Continental USA.

• Flexible work hours outside of standard collaboration times.

• Exclusive incentive purchase programs.

• Potential for a 10% annual bonus.

People also viewed

Horizon3.ai1 day ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH1 day ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM1 day ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies1 day ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)1 day ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC1 day ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers