
Senior DevOps Engineer
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in United States.
• Develop and enhance the platform function rather than merely maintaining it by undertaking the reconstruction and platform initiatives already planned, bringing more of the environment under code and clear ownership, and influencing the growth of DevOps in terms of platform and developer experience.
• Take responsibility for the uptime of our production platform. Lead incident response during system degradation or failures, conduct root cause analysis, and implement fixes and safeguards to prevent recurrence, including on-call practices, runbooks, and blameless postmortems.
• Manage metrics, monitoring, alerting, and tracing across the stack to ensure systems identify issues before customers are impacted.
• Create actionable dashboards and alerts, minimizing noise, and extend coverage as new services are launched.
• Oversee, improve, and expand our AWS environment, centered on EKS, RDS/Aurora, and networking.
• Administer our AWS environment with careful consideration of scalability, redundancy, cost, and fault tolerance, while ensuring security.
• Operate and upgrade our EKS clusters. Deploy and scale workloads, manage version updates, optimize for reliability and cost, and troubleshoot cluster and workload challenges.
• Maintain the health and performance of our databases and messaging systems, including PostgreSQL on RDS/Aurora and our message broker.
• Identify and resolve performance degradation across data, messaging, and network paths.
• Define and manage infrastructure as code to ensure that every change is repeatable and reviewable. We utilize AWS CDK in JavaScript/TypeScript; you will enhance existing stacks and progressively bring more of the environment under code.
• Manage the pipelines relied upon by engineers, primarily in GitHub Actions. Identify and eliminate bottlenecks that impede engineering progress, such as slow deployments, unreliable pipelines, or scalability issues, and develop trustworthy tooling.
• Oversee the movement of tracker data from field equipment through cellular tunnels and our AWS networking into the platform. Ensure that pathways are reliable, secure, and observable.
• Monitor and assess AWS spending, identify opportunities for savings, and regularly provide clear recommendations and trade-offs to engineering leadership.
• Document the systems you manage and share that knowledge with the team to ensure the function does not rely on a single individual.
• Over 5 years of experience in platform, infrastructure, DevOps, or SRE roles, including recent experience as a senior individual contributor responsible for end-to-end production systems.
• Significant hands-on experience managing production infrastructure on AWS.
• Practical experience operating container-orchestrated workloads: deploying, scaling, upgrading, and troubleshooting.
• Experience with managed relational databases in production, including diagnosing and resolving performance issues.
• A solid understanding of VPCs, subnets, routing, and DNS, as well as how traffic flows between systems.
• Proficient in writing infrastructure as code and the automation that supports it, emphasizing coding over clicking.
• Familiarity with AWS CDK in JavaScript/TypeScript, GitHub Actions, distributed messaging, caching, or search is preferred.
• Demonstrated ability to quickly familiarize oneself with unfamiliar systems and achieve a working understanding without detailed guidance.
• A genuine curiosity about system functionality, focusing on underlying mechanics instead of viewing them as black boxes.
• A strong sense of ownership: you understand, enhance, and advocate for the systems under your responsibility.
• Proactively communicate risks, decisions, and progress, making technical and cost trade-offs understandable to non-experts in infrastructure.
• Document while building and share knowledge so that no individual, including yourself, becomes a single point of failure.
• Exercise sound judgment regarding change safety: act swiftly on work that is safe and reversible, and thoroughly vet larger or riskier changes before proceeding.
• A bachelor’s degree in a relevant field or equivalent practical experience is required.
• Travel is required, approximately 8-10%.
• Opportunities for growth and personal development within a dynamic team.
• Comprehensive, cost-effective benefits packages are available.
• Benefit coverage begins on the first day of employment.
• Paid Time Off and Volunteer Time Off are provided.
• 401k matching is offered.
• Dependent Care support is available.
• Employee referral bonuses are provided.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.