
Manager, Software Engineering
Posted Sep 9

Posted Sep 9
This is a fully remote position, open to applicants in United States.
• Lead, mentor, and cultivate a team dedicated to site reliability, platform operations, release engineering, and automation.
• Establish priorities, set expectations, and define delivery objectives.
• Assist in hiring, onboarding, performance evaluations, and career advancement.
• Develop and implement the SRE and operational enhancement roadmap.
• Enhance reliability, availability, scalability, deployment confidence, and production readiness.
• Promote operational efficiency and optimize cloud costs.
• Create practices for incident response, escalation management, observability, runbooks, post-incident reviews, and toil reduction.
• Collaborate with engineering, architecture, security, release, and operations teams on shared tools and delivery workflows.
• Oversee CI/CD, release automation, cloud infrastructure, Kubernetes/EKS, Linux, Terraform/IaC, Ansible, and automation initiatives.
• Ensure effective management of incidents and escalations, root cause analysis, corrective actions, monitoring, alerting, capacity planning, and operational readiness.
• Support security, compliance, audit preparedness, vulnerability mitigation, vendor management, and operational documentation.
• Lead sprint planning, backlog oversight, prioritization, quarterly planning, roadmap development, and dependency coordination.
• Report reliability updates, operational risks, roadmap progress, incident patterns, and investment requirements to technical and executive audiences.
• Extensive experience in site reliability engineering, DevOps methodologies, platform engineering, infrastructure management, operational efficiency, cloud infrastructure, or production operations.
• Demonstrated experience in leading and managing engineering teams.
• Strong grasp of reliability engineering, operational excellence, incident management, escalation management, observability, production support, and scalable system operations.
• Solid working knowledge of AWS; experience with Azure is an advantage.
• Familiarity with Kubernetes/EKS, Docker, Helm, and Linux environments.
• Knowledge of Terraform, Ansible, infrastructure as code, configuration management, and automation tools.
• Understanding of CI/CD pipelines, release automation, deployment processes, Jenkins, GitHub Actions, Artifactory, or equivalent tools.
• Proficient in Python, Bash, or Groovy scripting and automation concepts.
• Experience with ELK stack, Prometheus, Grafana, CloudWatch, or other observability platforms.
• Working knowledge of security practices, application security, vulnerability management, compliance controls, audit readiness, and vendor management in a cloud or SaaS setting.
• Familiarity with Agile/Scrum methodologies, sprint planning, backlog management, quarterly planning, roadmap creation, and cross-team dependency handling.
• Strong analytical, problem-solving, presentation, and facilitation skills.
• Capability to communicate effectively with technical teams, senior leadership, executives, vendors, and cross-functional stakeholders.
• Experience in enhancing operational processes, minimizing manual toil, mentoring engineers, managing delivery, and building accountable teams.
• Bachelor’s degree in Computer Science, Information Technology, or a related discipline, or equivalent practical experience.
• Legal authorization to work in the United States without employer sponsorship.
• Bonus eligibility.
• Comprehensive benefits package.
• Remote-first working environment.
• Employee-led diversity and inclusion initiatives.
• Annual charity and fundraising activities.
• Volunteer days.
• Global employee sustainability programs.
• Global fitness and trivia competitions.
• Global wellbeing days.
• Monthly wellbeing webinars and training sessions.
• Equal opportunity and recruitment adjustments/support.
slice
SCS Global Services
SCS Global Services
Chainguard
Get handpicked remote jobs straight to your inbox weekly.