
Staff Software Engineer, Dev Ops
Posted Sep 3

Posted Sep 3
This is a fully remote position, open to applicants in United States.
• Act as a technical leader for DevOps projects across engineering teams.
• Establish architectural standards and best practices for infrastructure, automation of deployment, and operational excellence.
• Lead strategic initiatives aimed at enhancing scalability, resiliency, and platform performance.
• Offer mentorship and technical direction to engineers throughout the organization.
• Design, implement, and refine enterprise-scale CI/CD pipelines.
• Create automation strategies that enhance deployment consistency, quality, and release speed.
• Collaborate with development teams to optimize build, test, deployment, and release processes.
• Promote the adoption of DevOps best practices throughout the entire software development lifecycle.
• Architect, construct, and manage highly available cloud environments on AWS and Microsoft Azure.
• Spearhead infrastructure modernization initiatives utilizing Infrastructure-as-Code methodologies.
• Develop scalable cloud-native solutions that meet current and future business needs.
• Enhance platform performance, reliability, security, and operational expenses.
• Design and support enterprise-grade containerized environments using Docker and Kubernetes.
• Establish deployment patterns and operational standards for Kubernetes-based workloads.
• Lead initiatives centered on orchestration, scalability, observability, and platform automation.
• Encourage the adoption of cloud-native development and deployment methodologies.
• Create Infrastructure-as-Code solutions utilizing Terraform and CloudFormation.
• Develop automated solutions for provisioning, configuration management, compliance, and operational support.
• Improve monitoring, alerting, logging, and incident response capabilities across cloud infrastructures.
• Advocate for Site Reliability Engineering (SRE) principles and reliability-centered engineering practices.
• Collaborate with Security teams to implement cloud security best practices and compliance standards.
• Design solutions that facilitate encryption, identity management, access controls, and network security.
• Lead efforts to enhance platform resilience through disaster recovery and fault-tolerant designs.
• Ensure that infrastructure solutions comply with security and governance standards.
• Reports to the Manager of DevOps.
• Bachelor’s degree in Computer Science, Engineering, Information Systems, or a related technical field, or equivalent practical experience.
• Over 8 years of experience in designing, building, and supporting large-scale cloud infrastructure and software delivery platforms.
• Proven track record of leading complex DevOps, cloud modernization, or platform engineering projects within SaaS organizations.
• Experience in architecting highly available, fault-tolerant systems that support mission-critical workloads.
• Demonstrated capability to influence technical strategy and promote engineering best practices across multiple teams.
• Deep expertise in microservices architectures and cloud-native application design principles, including SaaS, PaaS, IaaS, and multi-tenant platforms.
• Advanced knowledge of AWS and/or Microsoft Azure cloud services and infrastructure management.
• Strong experience with Infrastructure-as-Code frameworks such as Terraform and CloudFormation.
• Extensive familiarity with Kubernetes, Docker, and contemporary container orchestration technologies.
• Solid understanding of networking concepts including DNS, TCP/IP, VPNs, firewalls, routers, and network security.
• Experience implementing cloud security frameworks, identity management, and compliance controls.
• Working knowledge of .NET Core/.NET Framework, Java, and JavaScript.
• Proven ability to troubleshoot and resolve complex infrastructure and application issues.
• Relevant certifications in AWS, Microsoft Azure, Kubernetes, or other cloud platforms.
• Experience in architecting and managing large-scale Kubernetes environments.
• Experience designing globally distributed, highly resilient cloud infrastructures.
• Expertise in disaster recovery planning, business continuity, and fault-tolerant architectures.
• Strong understanding of security frameworks, including OAuth, WS-Security, encryption at rest, and encryption in transit.
• Experience in implementing Site Reliability Engineering (SRE) practices and observability frameworks.
• Experience working within Agile development environments using Atlassian Jira and modern DevOps toolchains.
• Background in supporting enterprise SaaS products at scale.
• Variable incentive pay component.
• Flexible work options.
• Medical insurance.
• Dental insurance.
• Competitive compensation and benefits packages.
• Equal employment opportunities.
• Reasonable accommodations for disabilities.
• Confidential handling of applicant information.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.