
Lead Cloud Architect
Posted Sep 1

Posted Sep 1
This is a fully remote position, open to applicants in United States.
• Take charge of the reliability, availability, performance, and operational health of essential infrastructure services and applications.
• Manage intricate production incidents and service recovery operations, providing technical leadership during significant events.
• Conduct comprehensive root-cause analysis and implement corrective measures.
• Establish and continuously enhance SLIs, SLOs, error budgets, service health metrics, and operational standards.
• Spearhead the observability strategy, encompassing metrics, logs, traces, dashboards, and alerts.
• Oversee the operational lifecycle of deployed services, which includes releases, upgrades, patching, configuration changes, capacity management, maintenance, and technology refreshes.
• Identify and address systemic reliability risks, recurring failure patterns, capacity limitations, technical debt, and other operational hazards.
• Design and implement resilience and recovery initiatives, including failure-mode analysis, performance and capacity testing, disaster recovery, backup and failover strategies, and recovery validation.
• Advance automation and everything-as-code methodologies.
• Collaborate with Platform Engineering and application teams on architecture and design.
• Specify operational and reliability requirements while influencing platform capabilities and application architectures.
• Mentor junior Cloud Architects and offer technical guidance by establishing best practices and enhancing operational processes.
• U.S. Citizenship is required, along with the ability to obtain and maintain the necessary Public Trust level clearance.
• A Bachelor’s Degree with 12 years of experience, a Master’s Degree with 10 years of experience, or a High School diploma/equivalent with 16 years of experience is required.
• At least 7 years of hands-on experience in cloud engineering, DevOps, or production systems engineering.
• Significant hands-on experience in operating AWS Commercial and AWS GovCloud, including OpenShift (ROSA) or similar Kubernetes-based platforms.
• Profound infrastructure-as-code experience with Terraform and Ansible / Ansible Tower.
• Advanced knowledge of GitLab and Jenkins CI/CD platforms, including reliability gating and deployment automation.
• Extensive experience in Linux and Windows Server administration.
• Considerable practical experience in implementing and managing enterprise observability tools such as Dynatrace, Datadog, Splunk, and Open Telemetry.
• Responsible for an SLI/SLO and alerting program, which involves error budgets, alert rationalization, and noise reduction.
• Proficiency in scripting/automation using Python, Bash, PowerShell, or Go.
• Experience working in federal or regulated environments (FISMA, FedRAMP, NIST 800-53).
• Preferred: AWS Solutions Architect, AWS DevOps Engineer, or AWS SysOps certification.
• Preferred: Red Hat Certified Specialist in ROSA or Red Hat Certified Advanced System Administrator in OpenShift.
• Preferred: Azure Solutions Architect Expert, Azure DevOps Engineer Expert, GCP Professional Cloud Architect, or GCP Professional Cloud Developer certification.
• Preferred: Dynatrace Associate or Datadog Log Management Fundamentals certification.
• Preferred: GitLab CI/CD Associate certification or Certified Jenkins Engineer (CJE).
• Preferred: Terraform Authoring and Operations Professional certification.
• Potential eligibility for overtime.
• Shift differential may be available.
• Discretionary bonus may be available.
inexogy smart metering
Skylo
SYNCREON
TechPionier
Get handpicked remote jobs straight to your inbox weekly.