
Senior DevOps Engineer, Infrastructure – Reliability
Posted Aug 4

Posted Aug 4
This is a fully remote position, open to applicants in Florida.
• Develop scalable Infrastructure-as-Code frameworks using Terraform to standardize cloud provisioning and minimize configuration drift.
• Take ownership of the Kubernetes platform, ensuring that workloads are secure, scalable, and resilient.
• Enhance CI/CD pipelines to increase deployment frequency, decrease lead time, and boost release confidence.
• Design and implement secure networking, IAM, and secrets management strategies.
• Enhance observability through metrics, logging, and tracing using tools such as DataDog.
• Optimize cloud expenditures through rightsizing, autoscaling, and architectural enhancements.
• Execute disaster recovery, backup, and multi-region resilience strategies.
• Transform fragile or manual infrastructure into automated, testable, and reproducible systems.
• Introduce infrastructure tools or architectural modifications and promote their adoption through documentation, workshops, and hands-on assistance.
• Collaborate with engineering teams to minimize friction in CI/CD, deployments, and cloud environments.
• Articulate technical trade-offs among engineering and product stakeholders.
• Maintain or surpass SLO/SLA targets, reduce incident frequency and duration, enhance infrastructure automation, and improve cloud cost efficiency.
• 5+ years of experience in DevOps, SRE, or infrastructure engineering.
• Proven track record in designing and managing production Kubernetes environments at scale.
• Extensive hands-on knowledge of AWS infrastructure and cloud networking.
• Strong expertise in building and maintaining Terraform modules across extensive cloud environments.
• Demonstrated ownership of CI/CD systems with measurable improvements in DORA metrics.
• Experience leading incident response initiatives and achieving significant postmortem outcomes.
• Solid understanding of distributed systems, event-driven architectures like Kafka, and database performance with PostgreSQL.
• Proven capability to modernize legacy infrastructure and reduce manual operational burdens.
• Ability to guide infrastructure projects from unclear initial stages to production without daily oversight.
• Capacity to build trust across teams while elevating reliability standards.
• Bonus: experience in application coding.
• Bonus: familiarity with operating high-throughput Kafka clusters.
• Bonus: expertise in database performance tuning with PostgreSQL and Redis.
• Bonus: knowledge of autoscaling strategies for high-traffic systems.
• Bonus: experience with service mesh technologies.
• Bonus: involvement with internal developer platforms.
• Bonus: understanding of security best practices, including zero-trust networking and policy-as-code.
• Bonus: experience with multi-region or globally distributed systems.
• Bonus: familiarity with platform-wide reliability frameworks such as SLOs, error budgets, and chaos testing.
• Health Care Plan (Medical, Dental & Vision)
• Retirement Plan (401k)
• Life Insurance
• Flexible Paid Time Off
• 9 paid Holidays
• Family Leave
• Remote work options
• Hybrid work arrangement (for Orlando Associates)
• Complimentary Food & Snacks (Orlando)
• Wellness Resources
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.