
Director, Site Reliability Engineering
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in United States.
• Develop and sustain top-tier infrastructure for millions of users who prioritize privacy.
• Utilize high-level programming languages such as Perl, Go, TypeScript, and Python.
• Ensure Duck.ai adheres to reliability standards while minimizing user friction during outages.
• Scale DuckDuckGo’s indexing infrastructure to accommodate billions of documents.
• Implement privacy-respecting anti-fraud verification processes.
• Tackle intricate operational challenges related to software, systems, automation, and process evaluation.
• Read, write, troubleshoot, and deploy software to address large-scale reliability challenges.
• Lead and collaborate on multifaceted projects from inception through to postmortem analysis.
• Identify and mitigate reliability risks using tools, services, alerts, and response protocols.
• Investigate and diagnose instability in high-traffic distributed systems.
• Contribute to shaping the future technical direction of deployments to enhance reliability and performance.
• Collaborate with software engineers to triage production issues and determine remediation strategies, including code modifications and performance considerations.
• Leverage cloud-native services and architectures to boost reliability and scalability.
• Over 10 years of relevant professional experience in reliability, platform, infrastructure, or software engineering.
• More than 4 years of experience leading Site Reliability Engineering (SRE) teams.
• Experience in participating in a 24x7 on-call rotation for large-scale deployments.
• Proven ability to lead and collaborate on high-impact, complex projects from proposal to postmortem.
• Proficiency in AI-driven development, including the design and implementation of agentic workflows.
• Capability to address ambiguous problems, propose innovative solutions, and execute with a strong emphasis on metrics.
• Experience in developing tools, services, alerts, and responses to identify and manage reliability risks.
• Proven ability to diagnose instability in high-traffic distributed systems.
• Extensive experience administering and troubleshooting Linux and web technologies.
• Skills in automating infrastructure provisioning and configuration management.
• Advanced programming abilities.
• Familiarity with cloud-native services and architectures.
• Practical experience in packaging and deploying applications using Docker and Docker Compose.
• Must be present on camera during meetings via video conferencing.
• Must successfully pass a background check as a condition of employment.
• Must be legally authorized to work in the country of residence; DuckDuckGo does not provide sponsorship or support for individual immigration needs.
• Stock options.
• Company-sponsored health benefits for team members located in the United States.
• Paid parental leave.
• Office setup allowance.
• Co-working allowance.
• Flexible work arrangement with no core hours.
• Remote-first work environment.
• Opportunities for company all-hands meetings and team retreat travel.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.