Director, DevOps

Posted Aug 12

This is a fully remote position, open to applicants in Illinois.

📋 Description

• Assist, oversee, and manage a production-oriented AI platform operating on AWS.

• Ensure that services remain accessible, responsive, and adequately scalable under real-world conditions.

• Take ownership of the incident response process from start to finish, including triage, resolution, documentation, and follow-up for prevention.

• Oversee infrastructure across various AWS accounts and both production and non-production environments.

• Direct release management and conduct weekly operational meetings.

• Develop, implement, and enhance infrastructure using automation and infrastructure as code.

• Investigate and assess cloud infrastructure options, providing recommendations.

• Uphold engineering best practices.

• Identify solutions based on initial guidance, gather feedback, and deliver results.

• Monitor security tools and engage in security assessments and vulnerability reviews.

• Ensure compliance with security protocols and data protection standards.

• Manage the backlog, monitor tasks, and report progress to relevant stakeholders.

• Collaborate with the platform development team on modifications and deployments.

• Participate in daily stand-up meetings, sprint planning, and sprint reviews/closures.

• Collaborate on sprint priorities, product backlog, and inter-team dependencies.

• Proactively identify and address issues.


⛳️ Requirements

• Over 8 years of practical DevOps or platform engineering experience at a senior, staff, or director level.

• Proven ownership of production SaaS environments.

• Strong independent infrastructure decision-making skills.

• Extensive AWS knowledge, including management of multiple accounts and production SaaS environments.

• Proficient in Infrastructure as Code with expertise in Terraform, AWS CDK, or both.

• Competent in high-quality Python and TypeScript coding.

• Automation-first mindset.

• Experience with serverless infrastructure.

• Knowledge of centralized logging and log management at scale.

• Experience with metrics, monitoring, alerting, and reporting tools.

• Familiarity with structured release management processes.

• Capacity to manage backlogs, track commitments, and communicate status effectively.

• Ability to navigate ambiguity and follow rough guidance.

• Dedication to engineering best practices and long-term management of complexity.


🏝️ Benefits

• Fully remote work environment.

• Competitive compensation package.

• Medical benefits.

• Dental benefits.

• 401(k) plan.

• Tuition reimbursement.

• Flexible time off policy.

• Flexible benefits options.

• Personal support services.

• Customized learning and development opportunities.

People also viewed

Fairsource10 hours ago

DevOps, Kubernetes Consultant

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€110k – €140k/year
ApplyView job
NetBox Labs10 hours ago

Senior DevOps Engineer, Observability

Latin AmericaFull-timeDevOps & Site Reliability Engineer (SRE)$75k – $85k/year
ApplyView job
VELZI.AI LIMITED10 hours ago

DevOps Engineer

ID flagIndonesia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
First Due10 hours ago

Platform Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$165k/year
ApplyView job
CoDev11 hours ago

Senior DevOps Engineer

PH flagPhilippines OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Fusable11 hours ago

Senior Dev Ops Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$135k – $155k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers