
Principal SRE
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in United States, +1 more country.
• Utilize extensive technical knowledge to establish the reliability, observability, and automation strategies for critical cloud platforms that handle PHI.
• Create reliability benchmarks, service level objectives, and automation methodologies across various teams.
• Offer technical mentorship, guidance, and leadership to development, QA, and operations teams.
• Serve as a senior escalation point for intricate or high-severity production challenges.
• Enhance reliability capabilities through design assessments, collaborative efforts, and post-incident reviews without assigning blame.
• Define and oversee SLO management frameworks, service level indicators, and error budgets.
• Establish observability standards related to metrics, logging, tracing, and alerting.
• Manage incident detection, response, escalation, post-incident analysis, and corrective measures.
• Spearhead capacity planning, performance engineering, failure-domain isolation, and disaster recovery strategies.
• Promote resilience patterns within product architecture in collaboration with development teams.
• Develop strategies, standards, patterns, and tools for platform operations automation.
• Implement automated solutions for recurring failure issues.
• Construct and sustain infrastructure-as-code, environment provisioning, and deployment automation.
• Automate processes for patching, scaling, certificate rotation, backup and restore validation, and disaster recovery drills.
• Create self-service tools for development and support teams.
• Automate the collection and verification of security and compliance controls for PHI-bearing applications.
• Identify, quantify, and minimize operational toil.
• Collaborate with development, product, program management, support, implementation, security, and other cross-functional teams.
• Assist in resolving customer issues.
• Generate, review, and approve documentation related to architecture, design, and projects.
• Plan, monitor, and deliver reliability and automation projects within budget and on schedule.
• Provide updates on initiative status and escalate any deviations from commitments.
• Represent the reliability stance of the platform to clients and auditors.
• Contribute insights into cloud consumption and tooling budgets.
• Participate in an on-call escalation rotation, which may include occasional work outside of standard business hours.
• Willingness to travel approximately 5%.
• Bachelor’s degree in Computer Science, Engineering, or a related field; or equivalent practical experience in the industry.
• Over 10 years of experience in software engineering, systems engineering, or infrastructure operations.
• Proven advancement into a principal or staff-level technical position.
• Demonstrated ability to lead teams, projects, or individuals informally to achieve successful results.
• Experience operating production SaaS at scale with formal availability commitments.
• Proven track record in designing and implementing operational automation that significantly reduced manual work or recovery time.
• Deep knowledge of cloud infrastructure and architecture, especially in Microsoft Azure.
• Proficiency in infrastructure as code and configuration management tools like Terraform, Bicep/ARM, or Ansible.
• Experience with CI/CD pipeline design and release automation.
• Familiarity with containers and orchestration tools, including Kubernetes/AKS and service mesh concepts.
• Knowledge of observability and telemetry tools, such as Azure Monitor/KQL, Prometheus, Grafana, and distributed tracing.
• Proficient in at least one automation or systems language, such as Python, Go, Bash, or Java.
• Experience in Linux administration, networking, and fundamentals of identity/authorization.
• Knowledge of incident management and post-incident review practices.
• Ability to comprehend software architecture and design patterns.
• Skills in technical project management.
• Experience working in an Agile environment.
• Proficiency in Microsoft Office.
• Strong fundamental leadership skills, including strategic thinking, team building, adaptability, and conflict resolution.
• Ability to lead and influence experienced professionals technically.
• Preferred qualifications include: Azure certification, experience in regulated environments, medical imaging familiarity with DICOM and HL7, Agile/Scrum, Istio, Kafka, relational and NoSQL data platform operations, cybersecurity, chaos engineering, data engineering, and the application of machine learning/artificial intelligence in operations.
• Remote-first work environment allowing flexibility.
• Flexible vacation policy to support rest, recharge, and personal connections.
• Paid leave benefits.
• Comprehensive health, dental, and vision insurance.
• 401k retirement savings plan.
• Infertility benefits.
• Tuition reimbursement program.
• Life insurance coverage.
• Employee Assistance Program (EAP) and more!
Rimutee
CACI International Inc
CACI International Inc
Rimutee
Get handpicked remote jobs straight to your inbox weekly.