Senior DevOps Engineer, Infrastructure & Reliability

Posted Aug 25

This is a fully remote position, open to applicants in Florida.

📋 Description

• Develop scalable Infrastructure-as-Code methodologies utilizing tools such as Terraform to standardize cloud provisioning and minimize configuration drift.

• Take charge of and enhance the Kubernetes platform (EKS or self-managed), ensuring that workloads are secure, scalable, and resilient by default.

• Streamline CI/CD pipelines to increase deployment frequency, reduce lead times, and boost confidence in releases.

• Design and implement secure networking, IAM, and secrets management protocols across different environments.

• Enhance observability by refining metrics, logs, and tracing utilizing tools like DataDog.

• Increase cloud cost efficiency through rightsizing, autoscaling strategies, and architectural enhancements.

• Establish disaster recovery frameworks, backup strategies, and multi-region resilience initiatives.

• Transform brittle or manually managed infrastructure into automated, testable, and reproducible systems.

• Introduce new infrastructure tools or architectural changes and promote adoption through documentation, workshops, and hands-on support.

• Collaborate with engineering teams to remove friction in CI/CD, deployments, and cloud environments.

• Clearly communicate technical trade-offs to engineering and product stakeholders, balancing speed with safety.


⛳️ Requirements

• 8+ years of experience in DevOps, SRE, or infrastructure engineering.

• Proven track record in designing and managing production Kubernetes environments at scale.

• Extensive hands-on knowledge of AWS infrastructure and cloud networking.

• Strong experience in building and maintaining Terraform modules across extensive cloud environments.

• Demonstrated ownership of CI/CD systems with measurable improvements in DORA metrics.

• Experience in leading incident response initiatives and achieving meaningful postmortem results.

• Solid understanding of distributed systems, event-driven architectures (Kafka), and database performance (PostgreSQL).

• Proven capability to modernize legacy infrastructure and reduce manual operational toil.

• History of successfully managing a scoped infrastructure project from an unclear starting point to production without requiring daily oversight.

• Demonstrated ability to build trust across teams while enhancing reliability standards.

• Bonus Points (Nice to Have): Experience in application coding; experience managing high-throughput Kafka clusters (MSK or self-managed); strong expertise in database performance optimization (PostgreSQL, Redis); experience in implementing autoscaling strategies for high-traffic systems; familiarity with service mesh technologies; experience in developing internal developer platforms (IDP); background in security best practices (zero-trust networking, policy-as-code); experience with multi-region or globally distributed systems; experience in introducing platform-wide reliability frameworks (SLOs, error budgets, chaos testing).


🏝️ Benefits

• Health Care Plan (Medical, Dental & Vision)

• Retirement Plan (401k)

• Life Insurance

• Flexible Paid Time Off

• 9 paid Holidays

• Family Leave

• Remote work options

• Hybrid work for Orlando Associates

• Free Food & Snacks (Orlando)

• Wellness Resources

People also viewed

knowmad mood17 hours ago

Consultor/a DevSecOps – AWS

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
RealTime eClinical Solutions1 day ago

Principal DevOps Architect

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$155k – $195k/year
ApplyView job
Koniag Government Services1 day ago

DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Koniag Government Services1 day ago

Senior AWS DevOps Engineer – AWS, Kubernetes, HCP, CI/CD, Observability, AI-focus

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ASRC Federal1 day ago

Senior DevOps Administrator – Supporting NASA

US flagCalifornia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Nelios1 day ago

DevOps Engineer, Cloud Infrastructure

GR flagGreece OnlyPart-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers