
Director, Infrastructure & IT Operations
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in United States.
• Lead teams focused on cloud infrastructure, Site Reliability Engineering (SRE), production operations, corporate IT, helpdesk, and bot mitigation.
• Define ownership, operational standards, Service Level Objectives (SLOs), escalation protocols, and performance metrics.
• Mentor managers, technical leads, and IT personnel through coaching and career development initiatives.
• Create infrastructure roadmaps that align with business objectives, growth, security, reliability, and cost considerations.
• Collaborate with Engineering, Product, Data, Marketing, Finance, HR, Legal, and executive teams to assess risks, architecture, and tradeoffs.
• Assume responsibility for the reliability, scalability, security, performance, and cost efficiency of primarily AWS cloud infrastructure.
• Provide architectural guidance across cloud services, applications, APIs, networking, identity, databases, storage, observability, and integrations.
• Promote automation, configuration management, and infrastructure as code practices.
• Lead capacity planning efforts and eliminate single points of failure, undocumented systems, manual processes, and excessive vendor reliance.
• Set Service Level Indicators (SLIs), SLOs, availability targets, error-budget practices, observability, and actionable alerting mechanisms.
• Oversee major incident management, root-cause analysis, corrective-action tracking, backup, disaster recovery, and business continuity strategies.
• Manage bot mitigation, scraping protection, automated abuse prevention, and traffic quality oversight.
• Supervise employee IT services, helpdesk support, onboarding, offboarding, equipment provisioning, access management, and asset recovery.
• Establish identity and access management protocols, including least privilege, Multi-Factor Authentication (MFA), access reviews, vulnerability remediation, audits, and vendor evaluations.
• Handle vendor management, contracts, renewals, licensing, Service Level Agreements (SLAs), infrastructure and IT budgets, cloud cost visibility, and optimization.
• 10+ years of experience in cloud infrastructure, SRE, platform engineering, DevOps, IT operations, or related fields.
• 5+ years of experience leading engineering or technology teams, including managers or technical leads.
• Proven experience in operating high-traffic, customer-facing platforms in AWS or a similar cloud environment.
• In-depth understanding of cloud architecture, networking, identity, security, observability, databases, APIs, storage, and distributed systems.
• Experience in managing production availability, incident response, root-cause analysis, disaster recovery, and business continuity processes.
• Familiarity with infrastructure automation, configuration management, Continuous Integration/Continuous Deployment (CI/CD), and infrastructure as code.
• Experience overseeing corporate IT, employee support, endpoint management, licensing, and access provisioning.
• Proven track record in establishing operational metrics, SLOs, and support standards.
• Demonstrated capability in managing vendors, contracts, budgets, and cloud expenses.
• Strong written and verbal communication skills, able to convey technical issues and tradeoffs to executives effectively.
• Preferred: Experience with Cloudflare, DataDome, or similar bot-management and web application protection platforms, along with leading bot mitigation or traffic quality initiatives.
• Preferred: Familiarity with AWS services, containerized environments, infrastructure as code, centralized logging, and Application Performance Monitoring (APM) tools (e.g., AppDynamics).
• Preferred: Experience supporting large-scale consumer subscription, public-record, data-as-a-service, advertising, or e-commerce platforms.
• Preferred: Experience in modernizing legacy infrastructure, transitioning systems from external vendors, and consolidating cloud accounts or IT services.
• Preferred: Knowledge of privacy regulations and secure handling of sensitive consumer data.
• Preferred: Experience in building or enhancing SRE, DevOps, or platform engineering functions.
• 401k
• 401k match
• Medical/Dental/Vision/Life Insurance
ICF
Cisco
Auctus Group
Get handpicked remote jobs straight to your inbox weekly.