Remotery

Senior DevOps – Infrastructure Lead

Posted Jul 29

This is a fully remote position, open to applicants in Texas.

📋 Description

• Oversee and enhance the current DigitalOcean infrastructure.

• Provide support and optimize Linux-based production server environments.

• Transition self-managed databases to managed database services, ensuring verified failover, backups, and recovery processes.

• Shift applications to managed runtimes (including Laravel Cloud where applicable), transforming manual deployment methods into automated, repeatable pipelines.

• Strengthen and extend our use of Cloudflare for edge computing, static hosting, caching, and security measures.

• Develop a comprehensive inventory of servers, services, databases, domains, access pathways, backups, monitoring tools, and operational risks.

• Create and update practical runbooks for standard and emergency infrastructure workflows.

• Enhance incident response, escalation protocols, monitoring, logging, and alerting systems.

• Evaluate and refine backup, restoration, and disaster recovery procedures.

• Detect recurring manual tasks and transform them into safer procedures, scripts, automation, or infrastructure-as-code solutions.

• Assist in defining infrastructure-as-code standards and transition suitable infrastructure into repeatable, version-controlled workflows.

• Collaborate with AWS services as necessary (Lambda, VPC, IAM, CloudWatch, S3, SSM/Secrets Manager, queues).

• Leverage AI tools to expedite discovery, documentation, scripting, troubleshooting, and automation while maintaining strong production safety standards.

• Collaborate with engineering leadership to prioritize infrastructure risks and modernization efforts; clearly track work in Jira/GitHub and proactively communicate about risks, trade-offs, and obstacles.


⛳️ Requirements

• Extensive experience managing production infrastructure at a senior level.

• Strong, hands-on Linux server administration skills (the traditional, "old-school" approach): operating, securing, and troubleshooting manually managed production servers (LAMP/LEMP, system services, cron, networking, SSH) directly through the command line rather than solely via a cloud console.

• Familiarity with DigitalOcean, Linode, AWS EC2, bare VPS hosting, or similar environments.

• Advanced database operations experience: migrating self-managed MySQL to a managed service, ensuring replication, backup validation, restoration testing, and IO isolation.

• Proficient with Cloudflare across DNS, WAF, CDN and caching behaviors, page rules, Workers, Pages, and Zero Trust/Access, including traffic routing and origin protection.

• Experience in PHP/Laravel application environments, particularly with managed Laravel runtimes (Laravel Cloud and/or DigitalOcean App Platform).

• Proficient in using Datadog or a similar observability platform for monitoring, alerting, dashboards, logs, and incident investigations.

• Knowledge of infrastructure-as-code tools such as Terraform, Pulumi, AWS CDK, Serverless Framework, or CloudFormation.

• Experience with CI/CD pipelines and deployment automation.

• Practical knowledge of AWS services (Lambda, IAM, VPC, CloudWatch, S3, SSM/Secrets Manager, queues).

• Good judgment concerning production safety, access control, secrets management, backups, and incident response.

• Willingness to assume real on-call responsibilities and respond to production incidents outside of regular business hours; this role does not adhere to a strict 9-to-5 schedule.

• A tendency to document learned information and create runbooks that others can follow.

• Practical experience with AI tools (ChatGPT, Claude, Cursor, GitHub Copilot, or similar) and a strong sense of where human verification is necessary.

• Ability to work autonomously in a small, remote engineering organization where practical ownership is valued over bureaucracy.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Flexible working hours and remote work options.

• Opportunities for professional development and growth.

• Health, dental, and vision insurance.

• Generous paid time off and holidays.

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers