Remotery

Director, DevOps

Posted 2 days ago

This is a fully remote position, open to applicants in Illinois.

📋 Description

• Assist, oversee, and manage a production-oriented AI platform operating on AWS.

• Ensure that services remain accessible, responsive, and adequately scalable under real-world conditions.

• Take ownership of the incident response process from start to finish, including triage, resolution, documentation, and follow-up for prevention.

• Oversee infrastructure across various AWS accounts and both production and non-production environments.

• Direct release management and conduct weekly operational meetings.

• Develop, implement, and enhance infrastructure using automation and infrastructure as code.

• Investigate and assess cloud infrastructure options, providing recommendations.

• Uphold engineering best practices.

• Identify solutions based on initial guidance, gather feedback, and deliver results.

• Monitor security tools and engage in security assessments and vulnerability reviews.

• Ensure compliance with security protocols and data protection standards.

• Manage the backlog, monitor tasks, and report progress to relevant stakeholders.

• Collaborate with the platform development team on modifications and deployments.

• Participate in daily stand-up meetings, sprint planning, and sprint reviews/closures.

• Collaborate on sprint priorities, product backlog, and inter-team dependencies.

• Proactively identify and address issues.


⛳️ Requirements

• Over 8 years of practical DevOps or platform engineering experience at a senior, staff, or director level.

• Proven ownership of production SaaS environments.

• Strong independent infrastructure decision-making skills.

• Extensive AWS knowledge, including management of multiple accounts and production SaaS environments.

• Proficient in Infrastructure as Code with expertise in Terraform, AWS CDK, or both.

• Competent in high-quality Python and TypeScript coding.

• Automation-first mindset.

• Experience with serverless infrastructure.

• Knowledge of centralized logging and log management at scale.

• Experience with metrics, monitoring, alerting, and reporting tools.

• Familiarity with structured release management processes.

• Capacity to manage backlogs, track commitments, and communicate status effectively.

• Ability to navigate ambiguity and follow rough guidance.

• Dedication to engineering best practices and long-term management of complexity.


🏝️ Benefits

• Fully remote work environment.

• Competitive compensation package.

• Medical benefits.

• Dental benefits.

• 401(k) plan.

• Tuition reimbursement.

• Flexible time off policy.

• Flexible benefits options.

• Personal support services.

• Customized learning and development opportunities.

People also viewed

CVS Health9 hours ago

Salesforce DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$83.4k – $166.9k/year
ApplyView job
Devoteam10 hours ago

Data, AWS DevSecOps

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Aspirion11 hours ago

Senior DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Goodgame Studios11 hours ago

Senior Agentic Engineer – Java Backend, DevOps

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Instacart11 hours ago

Site Reliability Engineer II

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$133k – $169k/year
ApplyView job
Logicalis Spain12 hours ago

DevOps Engineer

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€40k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers