
Director of DevOps
Posted Jul 17

Posted Jul 17
This is a fully remote position, open to applicants in United States.
• Take ownership of service reliability across Convoso’s platform, being an advocate for quality and resilience, while ensuring compliance with internal and external SLAs, SLOs, and SLIs.
• Work collaboratively with various teams to define key metrics and develop dashboards that enhance quality within the SDLC.
• Manage the incident management process, including monitoring and observability systems, as well as the major incident response process, consistently tracking and reducing MTTR.
• Utilize chaos engineering principles and conduct “game days” to proactively assess the resilience of Convoso’s platform.
• Establish and uphold a CI/CD center of excellence.
• Promote automation for CI/CD processes across diverse tech stacks and cloud providers.
• Recruit and lead DevOps teams across multiple product lines, ensuring they excel in applying industry best practices.
• Develop and support automated, scalable solutions for deploying and managing our global infrastructure.
• Enhance the use of automation tools for infrastructure provisioning, configuration, and deployment.
• Collaborate with Development and Operations personnel to define CloudOps processes, introducing new insights and technologies to maintain our competitive edge.
• Provide mentorship and expertise regarding system options, risk and impact management, along with cost versus benefit analysis.
• Ensure that security requirements for tools, systems, and environments are met, safeguarding the assets of the company and its customers.
• Lead troubleshooting efforts for system and performance issues in Production, QA, and Development environments.
• Identify opportunities for reducing technical debt.
• Focus on optimizing performance within the infrastructure.
• Lead the DevOps team.
• Create a dynamic training and skills improvement plan.
• Obtain team-level certifications relevant to AWS or Google Cloud.
• Continuously assess our services to guarantee high quality.
• Conduct reviews to pinpoint optimization opportunities within processes or systems.
• Contribute to the development of departmental procedure documents and working instructions.
• Collaborate with consultants and team members to design implementable solutions.
• Assign tasks relevant to solutions and monitor team members’ progress.
• Inspire and motivate teamwork to achieve goals.
• Provide mentorship and identify training opportunities for team members, staying updated on the latest technologies and best practices.
• Bachelor’s degree in Computer Science, Computer Engineering, or a related technical field.
• Over 5 years of experience as a Site Reliability Engineer leader, managing multiple teams with a strong passion for automation and continuous improvement.
• More than 5 years of experience in the operational management of SaaS production applications.
• At least 5 years of experience working with Agile/DevOps development teams.
• Minimum of 5 years related experience in infrastructure management within a Linux environment.
• Extensive knowledge of cloud platforms such as AWS or GCP and management of on-prem and hybrid environments.
• Proven track record as a player-coach who is adept at both management and hands-on work.
• Experience in deploying and operating containerized web applications.
• Familiarity with cloud automation/provisioning and configuration management tools like Ansible, Chef, Puppet, Docker, or Salt.
• Proficient in containerization tools such as Docker and Kubernetes.
• Experience in planning, creating, implementing, and maintaining a scalable software development infrastructure.
• Knowledge of development, build, and CI/CD tools like Git, GitHub, and Jenkins.
• A security background that will be beneficial as we work towards SOC2 certification.
• Strong people management skills and the ability to recruit and develop a high-performing team.
• Process-oriented with excellent documentation skills.
• Outstanding oral and written communication abilities.
• Competitive compensation package.
• Stock options.
• 100% coverage of premiums for employees; including Medical, Dental, Basic life insurance, and Long term disability.
• Affordable Vision plan and optional FSA.
• Paid Time Off (PTO), Paid Sick Time, Holidays, Bereavement time, and Parental Leave.
• Your birthday off.
• 401k program with a generous company match.
• Complimentary Employee Assistance Program and Travel Assistance.
• Monthly reimbursement for gym memberships.
• Monthly credits for food and beverage.
• Company outings.
• Onsite and offsite team-building events.
• Paid training for departments.
• Apple laptop (most roles).
• A team of highly experienced and friendly colleagues!
The Codest
IRIUM
Sólides
Resilinc
Get handpicked remote jobs straight to your inbox weekly.