Senior DevOps Engineer, Infrastructure – Reliability

Posted 1 day ago

This is a fully remote position, open to applicants in Florida.

📋 Description

• Develop scalable Infrastructure-as-Code solutions utilizing Terraform.

• Take ownership of the Kubernetes platform (either EKS or self-managed), ensuring that workloads maintain security, scalability, and resilience.

• Enhance CI/CD pipelines to increase deployment frequency, shorten lead times, and boost release confidence.

• Design and implement secure networking, IAM, and secrets management strategies across various environments.

• Enhance observability through the use of metrics, logs, and tracing tools such as DataDog.

• Improve cloud cost efficiency by applying rightsizing, autoscaling, and architectural enhancements.

• Establish disaster recovery plans, backup strategies, and initiatives for multi-region resilience.

• Transform brittle or manually managed infrastructure into automated, testable, reproducible systems.

• Introduce infrastructure tools or architectural changes and promote their adoption through documentation, workshops, and hands-on assistance.

• Collaborate with engineering teams to remove barriers in CI/CD, deployments, and cloud environments.

• Convey technical trade-offs to engineering and product stakeholders effectively.

• Maintain or surpass SLO/SLA targets, minimize incident frequency and duration, enhance infrastructure stability and automation, and optimize cloud expenditures.


⛳️ Requirements

• 8+ years of experience in DevOps, SRE, or infrastructure engineering.

• Proven expertise in designing and managing production Kubernetes environments at scale.

• In-depth practical knowledge of AWS infrastructure and cloud networking.

• Strong experience in building and maintaining Terraform modules across extensive cloud environments.

• Demonstrated ownership of CI/CD systems with measurable enhancements in DORA metrics.

• Experience leading incident response processes and achieving significant postmortem results.

• Robust understanding of distributed systems, event-driven architectures (Kafka), and database performance (PostgreSQL).

• Proven capability to modernize legacy infrastructure and reduce manual operational tasks.

• A successful history of advancing a scoped infrastructure project from an unclear starting point to production without requiring daily guidance.

• Demonstrated ability to foster trust among teams while enhancing reliability.

• Bonus: Experience in application coding.

• Bonus: Experience in managing high-throughput Kafka clusters (MSK or self-managed).

• Bonus: Strong background in database performance optimization (PostgreSQL, Redis).

• Bonus: Experience implementing autoscaling strategies for high-traffic systems.

• Bonus: Familiarity with service mesh technologies.

• Bonus: Experience in building internal developer platforms (IDP).

• Bonus: Background in security best practices (zero-trust networking, policy-as-code).

• Bonus: Experience with multi-region or globally distributed systems.

• Bonus: Experience in introducing platform-wide reliability frameworks (SLOs, error budgets, chaos testing).


🏝️ Benefits

• Health Care Plan (Medical, Dental & Vision).

• Retirement Plan (401k).

• Life Insurance.

• Flexible Paid Time Off.

• 9 paid Holidays.

• Family Leave.

• Remote work opportunities.

• Hybrid work arrangement (for Orlando Associates).

• Complimentary Food & Snacks (Orlando).

• Access to Wellness Resources.

People also viewed

Harris Computer23 hours ago

Platform & DevSecOps Delivery Architect

US flagAlabama, +20 more statesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TheWhiteam23 hours ago

DevOps Engineer

ES flagSpain OnlyFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
CFactory-Creations1 day ago

Senior Site Reliability Engineer – Eastern Europe

BG flagBulgaria, +2 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Capgemini1 day ago

Mainframe DevOps Migration Consultant

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$76.7k – $150.8k/year
ApplyView job
URBN (Urban Outfitters, Anthropologie Group, Free People & Nuuly)1 day ago

Senior DevOps Engineer

US flagPennsylvania OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Nitrado1 day ago

Site Reliability Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers