Principal SRE

Posted 6 days ago

This is a fully remote position, open to applicants in United States, +1 more country.

📋 Description

• Utilize extensive technical knowledge to establish the reliability, observability, and automation strategies for critical cloud platforms that handle PHI.

• Create reliability benchmarks, service level objectives, and automation methodologies across various teams.

• Offer technical mentorship, guidance, and leadership to development, QA, and operations teams.

• Serve as a senior escalation point for intricate or high-severity production challenges.

• Enhance reliability capabilities through design assessments, collaborative efforts, and post-incident reviews without assigning blame.

• Define and oversee SLO management frameworks, service level indicators, and error budgets.

• Establish observability standards related to metrics, logging, tracing, and alerting.

• Manage incident detection, response, escalation, post-incident analysis, and corrective measures.

• Spearhead capacity planning, performance engineering, failure-domain isolation, and disaster recovery strategies.

• Promote resilience patterns within product architecture in collaboration with development teams.

• Develop strategies, standards, patterns, and tools for platform operations automation.

• Implement automated solutions for recurring failure issues.

• Construct and sustain infrastructure-as-code, environment provisioning, and deployment automation.

• Automate processes for patching, scaling, certificate rotation, backup and restore validation, and disaster recovery drills.

• Create self-service tools for development and support teams.

• Automate the collection and verification of security and compliance controls for PHI-bearing applications.

• Identify, quantify, and minimize operational toil.

• Collaborate with development, product, program management, support, implementation, security, and other cross-functional teams.

• Assist in resolving customer issues.

• Generate, review, and approve documentation related to architecture, design, and projects.

• Plan, monitor, and deliver reliability and automation projects within budget and on schedule.

• Provide updates on initiative status and escalate any deviations from commitments.

• Represent the reliability stance of the platform to clients and auditors.

• Contribute insights into cloud consumption and tooling budgets.

• Participate in an on-call escalation rotation, which may include occasional work outside of standard business hours.

• Willingness to travel approximately 5%.


⛳️ Requirements

• Bachelor’s degree in Computer Science, Engineering, or a related field; or equivalent practical experience in the industry.

• Over 10 years of experience in software engineering, systems engineering, or infrastructure operations.

• Proven advancement into a principal or staff-level technical position.

• Demonstrated ability to lead teams, projects, or individuals informally to achieve successful results.

• Experience operating production SaaS at scale with formal availability commitments.

• Proven track record in designing and implementing operational automation that significantly reduced manual work or recovery time.

• Deep knowledge of cloud infrastructure and architecture, especially in Microsoft Azure.

• Proficiency in infrastructure as code and configuration management tools like Terraform, Bicep/ARM, or Ansible.

• Experience with CI/CD pipeline design and release automation.

• Familiarity with containers and orchestration tools, including Kubernetes/AKS and service mesh concepts.

• Knowledge of observability and telemetry tools, such as Azure Monitor/KQL, Prometheus, Grafana, and distributed tracing.

• Proficient in at least one automation or systems language, such as Python, Go, Bash, or Java.

• Experience in Linux administration, networking, and fundamentals of identity/authorization.

• Knowledge of incident management and post-incident review practices.

• Ability to comprehend software architecture and design patterns.

• Skills in technical project management.

• Experience working in an Agile environment.

• Proficiency in Microsoft Office.

• Strong fundamental leadership skills, including strategic thinking, team building, adaptability, and conflict resolution.

• Ability to lead and influence experienced professionals technically.

• Preferred qualifications include: Azure certification, experience in regulated environments, medical imaging familiarity with DICOM and HL7, Agile/Scrum, Istio, Kafka, relational and NoSQL data platform operations, cybersecurity, chaos engineering, data engineering, and the application of machine learning/artificial intelligence in operations.


🏝️ Benefits

• Remote-first work environment allowing flexibility.

• Flexible vacation policy to support rest, recharge, and personal connections.

• Paid leave benefits.

• Comprehensive health, dental, and vision insurance.

• 401k retirement savings plan.

• Infertility benefits.

• Tuition reimbursement program.

• Life insurance coverage.

• Employee Assistance Program (EAP) and more!

People also viewed

Rimutee22 hours ago

DevOps AWS – Español

MX flagMexico OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
CACI International Inc22 hours ago

HCM Cloud Platform – DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$90.3k – $189.6k/year
ApplyView job
CACI International Inc22 hours ago

Senior HCM Cloud Platform – DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$105.1k – $231.1k/year
ApplyView job
Rimutee1 day ago

DevOps, AWS – Spanish

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
OpenObserve1 day ago

DevOps Engineer

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ShiftKey1 day ago

Senior Site Reliability Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)PLN 22k – PLN 25k/month
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers