Principal Site Reliability Engineer – Platform

Posted Aug 25

This is a fully remote position, open to applicants in United States.

📋 Description

• Design, scale, and take ownership of critical infrastructure.

• Develop and manage a Kubernetes-based platform that supports various teams and services.

• Create backend services in Golang for autonomous systems.

• Collaborate with product teams to introduce new products on the platform.

• Expand high-availability infrastructure while ensuring uptime and other key performance metrics.

• Develop tools for the platform and development teams.

• Conduct comprehensive performance analysis and implement effective enhancements.

• Work alongside cloud vendors and external technical support for upgrades and troubleshooting.

• Engage in on-call rotations, triage and resolve production incidents, and document root causes and postmortems.

• Design and uphold observability infrastructure, including dashboards, alerts, and log aggregation.

• Collaborate with the security team for regular risk assessments.

• Maintain the risk register and execute mitigation plans.

• Evaluate intrusion detection alerts and enhance systems that process threat feeds.

• Establish and manage SaaS payment processes in conjunction with IT and purchasing teams.

• Influence architectural decisions, mentor engineers across various teams, and shape the future direction of the platform.


⛳️ Requirements

• At least 8 years of extensive experience in building and maintaining infrastructure for data-intensive, high-availability applications.

• Six years of experience in developing and managing public cloud solutions.

• Profound knowledge of Kubernetes and Terraform.

• Strong understanding of software design methodologies, information systems architecture, object-oriented design, and software design patterns.

• In-depth knowledge of securing cloud infrastructure, particularly with AWS and Kubernetes.

• Extensive experience in one or more of the following programming languages: Golang, Python, JavaScript, or Rust; Golang is preferred.

• Significant experience with CI/CD tools, including GitHub Actions, ArgoCD, ArgoCD Image Updater, and Artifactory.

• Visa sponsorship may be available for this role.


🏝️ Benefits

• Annual performance bonus.

• Competitive benefits package.

• Diversity, Equity, and Inclusion initiatives.

• Programs for recruiting, mentorship, career development, and learning & development.

• Reasonable accommodations for individuals with disabilities.

People also viewed

knowmad mood23 hours ago

Consultor/a DevSecOps – AWS

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
RealTime eClinical Solutions1 day ago

Principal DevOps Architect

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$155k – $195k/year
ApplyView job
Koniag Government Services1 day ago

DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Koniag Government Services1 day ago

Senior AWS DevOps Engineer – AWS, Kubernetes, HCP, CI/CD, Observability, AI-focus

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ASRC Federal1 day ago

Senior DevOps Administrator – Supporting NASA

US flagCalifornia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Nelios1 day ago

DevOps Engineer, Cloud Infrastructure

GR flagGreece OnlyPart-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers