DevOps Engineer – Kubernetes, AWS, Terraform

at24-MAGRemoteUS flagNew YorkFull-timeDevOps & Site Reliability Engineer (SRE)Mid-levelSenior$65 – $100/hour

Posted Aug 26

This is a fully remote position, open to applicants in New York.

📋 Description

• Design and assess technically demanding Kubernetes infrastructure tasks

• Diagnose cluster malfunctions and pinpoint configuration, networking, scheduling, resource, or service issues

• Develop technically sound solutions to intricate Kubernetes operational challenges

• Review remediation strategies for accuracy and feasibility in production

• Design and evaluate infrastructure scenarios within production AWS environments

• Assess integrations involving AWS Lambda, API Gateway, and DynamoDB

• Evaluate architecture, service configuration, dependencies, scalability, reliability, and quality of implementation

• Design and assess infrastructure automation utilizing Terraform and/or AWS CDK

• Review Infrastructure as Code (IaC) architecture, resource definitions, dependencies, maintainability, accuracy, and operational safety

• Evaluate CI/CD pipeline designs and deployment processes

• Review automation for build, test, release, and deployment, including failure management, validation, and rollback strategies

• Design complex domain-specific problems related to cloud infrastructure, Kubernetes, and automation

• Write clear, well-structured solutions for infrastructure engineering assignments

• Assess tasks and proposed solutions for accuracy, completeness, and readiness for production

• Identify errors in infrastructure, architecture, automation, and reasoning

• Provide comprehensive written technical feedback

• Develop thorough evaluation frameworks and rubrics for infrastructure engineering tasks

• Collaborate with technical specialists to ensure accuracy and consistency

• Identify deficiencies in cloud infrastructure and DevOps reasoning

• Contribute expert knowledge in Kubernetes, AWS, Infrastructure-as-Code, and deployment engineering

• Clearly and precisely explain complex technical decisions

• Assist in enhancing technical training and evaluation materials

• Work alongside research, engineering, and subject-matter experts on advanced AI infrastructure and technical evaluation projects


⛳️ Requirements

• Over 4 years of dedicated professional experience in cloud infrastructure, DevOps, Site Reliability Engineering, platform engineering, or a closely related field

• Experience as a DevOps Engineer, Cloud Engineer, Infrastructure Engineer, Platform Engineer, Site Reliability Engineer (SRE), or similar role

• Strong hands-on experience managing Kubernetes in production

• Proven ability to diagnose and resolve Kubernetes cluster failures, beyond merely creating manifests or using managed control planes

• Production experience with Terraform and/or AWS CDK

• Direct production experience with AWS Lambda, API Gateway, and DynamoDB integrations

• Experience in building and maintaining CI/CD pipelines

• Strong understanding of cloud infrastructure, deployment automation, production reliability, and operational troubleshooting

• Demonstrable career progression within infrastructure, DevOps, SRE, or platform engineering

• Excellent written communication skills with the ability to articulate complex technical decisions clearly

• Professional experience in recognized, technically rigorous organizations is preferred

• Fully remote work within the United States

• Availability during weekdays is required


🏝️ Benefits

• Full-time position

• Fully remote work within the United States

• 40 hours per week

• Weekday availability is required

• Opportunity to contribute to advanced AI infrastructure and technical evaluation projects

People also viewed

Bet On Talent6 hours ago

Senior DevOps Engineer

EuropeFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Virtasant7 hours ago

Build & Release Support Engineer – CI/CD

MX flagMexico, +5 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ookla7 hours ago

Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$90k – $100k/year
ApplyView job
opinov87 hours ago

Senior DevOps Engineer, Media and Advertising Industry

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
GFT Technologies7 hours ago

DevOps Specialist

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
RELX11 hours ago

Senior Site Reliability Engineer II

US flagNorth Carolina, +3 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$104.9k – $174.7k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers