
Manager, Technical Operations
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Take charge of the Infrastructure Service Request intake and prioritization process utilizing JIRA.
• Supervise the triage, prioritization, scheduling, and communication of infrastructure-related tasks, such as patching, security updates, and platform initiatives.
• Collaborate with Cloud Platform Engineers, InfoSec, product owners, and cross-functional stakeholders to ensure consistent delivery, operational integrity, and adherence to change-management protocols.
• Participate in an on-call rotation as a key escalation point for significant incidents.
• Conduct post-mortems and implement reliability enhancements aimed at achieving a 99.99% uptime target.
• Report directly to the Vice President of Platform Engineering.
• Manage Cloud Platform Engineers directly, including hiring, coaching, performance management, and career growth.
• Oversee the JIRA board, ensuring backlog health, workflow governance, SLAs, stakeholder visibility, and prioritization discipline are maintained.
• Coordinate maintenance windows, patching, security updates, and operational upgrades.
• Facilitate Infrastructure Board Review sessions.
• Proactively identify and escalate infrastructure risks related to capacity, obsolescence, performance, security exposure, and operational bottlenecks.
• Manage team capacity, staffing needs, workload distribution, hiring, and skills development.
• Break down, scope, and route large infrastructure requests into platform sprint cycles.
• Provide updates, insights, and risk evaluations to the VP of Platform Engineering.
• Collaborate with InfoSec on security enhancements and PCI/SOC 2 audit preparedness.
• Optimize workflows, enhance throughput and cycle times, standardize documentation, and minimize manual toil.
• Utilize AI-assisted tools and assess AI-driven automation for triaging, anomaly detection, alerting, and self-service remediation.
• Implement dashboards, automate reporting, conduct trend analyses, and enhance change-management practices.
• Bachelor’s degree and at least 5 years of relevant experience in Infrastructure, Platform Engineering, Site Reliability Engineering, or similar technical operational roles.
• Proven experience in managing technical work intake processes and backlogs using JIRA or comparable tools.
• Experience in coordinating maintenance windows, patch cycles, or operational readiness tasks.
• Strong grasp of change-management frameworks and operational governance best practices.
• Excellent communication skills with experience engaging senior technical and business stakeholders.
• Strong analytical skills to assess technical risk, complexity, and business impact.
• Expertise in supporting and troubleshooting Microsoft Azure, including AKS, VNets, NSGs, Load Balancers, VPN/ExpressRoute, managed identities, and Azure Policy/RBAC.
• Expertise in supporting and troubleshooting Kubernetes, focusing on cluster operations, scaling, upgrades, workload security, and troubleshooting.
• Expertise in supporting and troubleshooting Cloudflare DNS, CDN, WAF, and Zero Trust configurations.
• Experience with Terraform for modular and reusable cloud provisioning.
• Experience with Ansible for system configuration and application deployment.
• Automation experience using PowerShell, Bash, and/or Python.
• Experience in supporting and troubleshooting Windows Server, IIS, and .NET Framework/Core application environments.
• Experience in supporting and troubleshooting Ubuntu Server, Nginx, and Python Django applications, including containerization and integration with Azure services.
• Experience in containerizing .NET and Python/Django applications and operating them in AKS with health probes, scaling, and observability.
• Experience in designing and managing firewall rules, NSGs, ACLs, and network segmentation strategies in both cloud and on-premises environments.
• Ability to analyze traffic flows, identify misconfigurations, resolve blocked flows, and enforce least-privilege access patterns.
• Experience managing distributed firewall policies across Azure, Cloudflare, and virtualization platforms like Proxmox.
• Experience with GitHub Actions or Azure DevOps for automated builds, tests, deployments, and environment promotions.
• Strong skills in monitoring, logging, and performance troubleshooting using Azure Monitor, Log Analytics, and optionally Prometheus/Grafana/New Relic.
• Familiarity with PCI, SOC2, and SOX compliance requirements and operations in compliant cloud environments.
• Experience implementing policy-as-code, RBAC standards, secure network design, and environment governance.
• Experience collaborating with InfoSec teams on security hardening, evidence gathering, remediation, and PCI/SOC 2 audit logistics.
• Preferred certifications include Azure AZ-104, AZ-305, AZ-400, and ITIL 4 Foundation.
• Eligibility for an annual bonus or commission.
• Potential for overtime pay.
• Equal employment opportunity and anti-discrimination protections.
• Accommodations for disability-related or religious recruitment needs.
Omnidian
Trase
Red Cell Partners
Netflix
Get handpicked remote jobs straight to your inbox weekly.