
Senior DevOps Engineer, Infrastructure – Reliability
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Florida.
• Develop scalable Infrastructure-as-Code solutions utilizing Terraform.
• Take ownership of the Kubernetes platform (either EKS or self-managed), ensuring that workloads maintain security, scalability, and resilience.
• Enhance CI/CD pipelines to increase deployment frequency, shorten lead times, and boost release confidence.
• Design and implement secure networking, IAM, and secrets management strategies across various environments.
• Enhance observability through the use of metrics, logs, and tracing tools such as DataDog.
• Improve cloud cost efficiency by applying rightsizing, autoscaling, and architectural enhancements.
• Establish disaster recovery plans, backup strategies, and initiatives for multi-region resilience.
• Transform brittle or manually managed infrastructure into automated, testable, reproducible systems.
• Introduce infrastructure tools or architectural changes and promote their adoption through documentation, workshops, and hands-on assistance.
• Collaborate with engineering teams to remove barriers in CI/CD, deployments, and cloud environments.
• Convey technical trade-offs to engineering and product stakeholders effectively.
• Maintain or surpass SLO/SLA targets, minimize incident frequency and duration, enhance infrastructure stability and automation, and optimize cloud expenditures.
• 8+ years of experience in DevOps, SRE, or infrastructure engineering.
• Proven expertise in designing and managing production Kubernetes environments at scale.
• In-depth practical knowledge of AWS infrastructure and cloud networking.
• Strong experience in building and maintaining Terraform modules across extensive cloud environments.
• Demonstrated ownership of CI/CD systems with measurable enhancements in DORA metrics.
• Experience leading incident response processes and achieving significant postmortem results.
• Robust understanding of distributed systems, event-driven architectures (Kafka), and database performance (PostgreSQL).
• Proven capability to modernize legacy infrastructure and reduce manual operational tasks.
• A successful history of advancing a scoped infrastructure project from an unclear starting point to production without requiring daily guidance.
• Demonstrated ability to foster trust among teams while enhancing reliability.
• Bonus: Experience in application coding.
• Bonus: Experience in managing high-throughput Kafka clusters (MSK or self-managed).
• Bonus: Strong background in database performance optimization (PostgreSQL, Redis).
• Bonus: Experience implementing autoscaling strategies for high-traffic systems.
• Bonus: Familiarity with service mesh technologies.
• Bonus: Experience in building internal developer platforms (IDP).
• Bonus: Background in security best practices (zero-trust networking, policy-as-code).
• Bonus: Experience with multi-region or globally distributed systems.
• Bonus: Experience in introducing platform-wide reliability frameworks (SLOs, error budgets, chaos testing).
• Health Care Plan (Medical, Dental & Vision).
• Retirement Plan (401k).
• Life Insurance.
• Flexible Paid Time Off.
• 9 paid Holidays.
• Family Leave.
• Remote work opportunities.
• Hybrid work arrangement (for Orlando Associates).
• Complimentary Food & Snacks (Orlando).
• Access to Wellness Resources.
Harris Computer
TheWhiteam
CFactory-Creations
Capgemini
Get handpicked remote jobs straight to your inbox weekly.