
Senior DevOps Engineer, Infrastructure & Reliability
Posted Aug 25

Posted Aug 25
This is a fully remote position, open to applicants in Florida.
• Develop scalable Infrastructure-as-Code methodologies utilizing tools such as Terraform to standardize cloud provisioning and minimize configuration drift.
• Take charge of and enhance the Kubernetes platform (EKS or self-managed), ensuring that workloads are secure, scalable, and resilient by default.
• Streamline CI/CD pipelines to increase deployment frequency, reduce lead times, and boost confidence in releases.
• Design and implement secure networking, IAM, and secrets management protocols across different environments.
• Enhance observability by refining metrics, logs, and tracing utilizing tools like DataDog.
• Increase cloud cost efficiency through rightsizing, autoscaling strategies, and architectural enhancements.
• Establish disaster recovery frameworks, backup strategies, and multi-region resilience initiatives.
• Transform brittle or manually managed infrastructure into automated, testable, and reproducible systems.
• Introduce new infrastructure tools or architectural changes and promote adoption through documentation, workshops, and hands-on support.
• Collaborate with engineering teams to remove friction in CI/CD, deployments, and cloud environments.
• Clearly communicate technical trade-offs to engineering and product stakeholders, balancing speed with safety.
• 8+ years of experience in DevOps, SRE, or infrastructure engineering.
• Proven track record in designing and managing production Kubernetes environments at scale.
• Extensive hands-on knowledge of AWS infrastructure and cloud networking.
• Strong experience in building and maintaining Terraform modules across extensive cloud environments.
• Demonstrated ownership of CI/CD systems with measurable improvements in DORA metrics.
• Experience in leading incident response initiatives and achieving meaningful postmortem results.
• Solid understanding of distributed systems, event-driven architectures (Kafka), and database performance (PostgreSQL).
• Proven capability to modernize legacy infrastructure and reduce manual operational toil.
• History of successfully managing a scoped infrastructure project from an unclear starting point to production without requiring daily oversight.
• Demonstrated ability to build trust across teams while enhancing reliability standards.
• Bonus Points (Nice to Have): Experience in application coding; experience managing high-throughput Kafka clusters (MSK or self-managed); strong expertise in database performance optimization (PostgreSQL, Redis); experience in implementing autoscaling strategies for high-traffic systems; familiarity with service mesh technologies; experience in developing internal developer platforms (IDP); background in security best practices (zero-trust networking, policy-as-code); experience with multi-region or globally distributed systems; experience in introducing platform-wide reliability frameworks (SLOs, error budgets, chaos testing).
• Health Care Plan (Medical, Dental & Vision)
• Retirement Plan (401k)
• Life Insurance
• Flexible Paid Time Off
• 9 paid Holidays
• Family Leave
• Remote work options
• Hybrid work for Orlando Associates
• Free Food & Snacks (Orlando)
• Wellness Resources
knowmad mood
RealTime eClinical Solutions
Koniag Government Services
Koniag Government Services
Get handpicked remote jobs straight to your inbox weekly.