
Director of Production Engineering
Posted Aug 25

Posted Aug 25
This is a fully remote position, open to applicants in United States.
• Recruit and develop a globally-distributed DevOps/SRE engineering team; oversee hiring, mentoring, and management of engineers.
• Take ownership of the reliability and infrastructure roadmap for the AWS-based production environment, which includes EKS, RDS, and associated AWS services.
• Direct the organization's security operations practice, encompassing vulnerability management, threat detection, incident response, and remediation.
• Establish and promote engineering OKRs focused on infrastructure reliability, automation, and security.
• Advocate for observability and alerting best practices, including automated alert triage and response procedures.
• Implement agentic AI infrastructure and AI-driven workflows within the SDLC and DevOps processes.
• Advance Infrastructure-as-Code, CI/CD, and automation methodologies.
• Collaborate with engineering and IT teams to ensure alignment on infrastructure standards, access controls, tools, and compliance obligations.
• Uphold security controls, data protection, compliance, and readiness for audits.
• Lead and engage in the Incident Management on-call rotation.
• Offer technical guidance and thought leadership to the engineering team.
• Dedicate approximately 20-30% of your time to directly contributing to architecture, tooling, and incident response, while the remainder focuses on vision, roadmap, and cross-team execution.
• 8-12 years of experience in DevOps, Site Reliability Engineering, or production infrastructure roles, including experience in people management.
• Extensive hands-on experience managing production workloads on AWS, including EKS (Kubernetes), RDS, VPC, IAM, Lambda, S3, and other fundamental AWS services.
• Proven experience in security operations (SecOps), including vulnerability management, incident response, and addressing production security challenges.
• 5+ years of experience with observability platforms such as Datadog, Prometheus, or Grafana.
• Strong proficiency in Infrastructure-as-Code tools like Terraform or CloudFormation and CI/CD automation.
• Competent in at least one programming language, such as Go, Python, or Bash.
• Daily use of Git and test automation pipelines.
• Hands-on experience with Linux/Unix production platforms, including Amazon Linux, Ubuntu, or RHEL/CentOS.
• Demonstrated ability to collaborate cross-functionally with engineering and IT teams to align on infrastructure, tools, and security standards.
• Proven track record in leading incident management and on-call practices for high-availability production systems.
• Bachelor's degree in Computer Science, Engineering, or a related field is required.
• A Master's degree is preferred.
• $0 monthly premium and various flexible medical, dental, and vision plans effective from the first day of employment.
• 401k plan.
• Discretionary Paid Time Off and Paid Holidays.
• Parental Leave.
• Equity.
• Monthly Wellness Reimbursement.
• Monthly Lunch on Legion.
• Bonus.
• Fully remote work.
NVIDIA
Redox
NVIDIA
TMS
Get handpicked remote jobs straight to your inbox weekly.