
Senior Site Reliability Engineer – Site Lead
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Europe.
• Recruit, hire, and develop engineers to build the European engineering team.
• Transform company priorities into a targeted roadmap, overseeing work from design to production.
• Contribute to system architecture, review designs and code, and tackle critical technical challenges directly.
• Enhance Kubernetes infrastructure for provisioning, networking, storage, and deployment across various providers and regions.
• Strengthen isolation, failover, observability, and recovery processes.
• Develop a sustainable incident response strategy and utilize production experiences to create more robust systems.
• Automate the processes of capacity expansion, deployments, and maintenance.
• Define clear ownership, maintain documentation, and ensure efficient handoffs across different regions and time zones.
• Foster an engineering culture that emphasizes technical quality, accountability, collaboration, and the growth of engineers.
• Proven experience in leading engineering teams responsible for production infrastructure or distributed systems.
• A successful history of hiring and nurturing talented engineers, setting priorities, and achieving significant technical outcomes.
• Sufficient technical expertise to guide architectural decisions, challenge assumptions, and contribute directly when necessary.
• Experience in building and managing reliable services, taking real ownership of their performance in production.
• Strong foundational knowledge of Linux and practical experience with networking, storage, containers, and Kubernetes.
• Capability to write maintainable software and automate solutions for infrastructure challenges.
• A systematic approach to troubleshooting issues across application, cluster, network, and hardware boundaries.
• Good judgment in prioritizing tasks, simplifying processes, and ensuring reliability.
• Effective communication skills and experience collaborating across teams and time zones.
• Experience in establishing or expanding an engineering office or regional team is a plus.
• Familiarity with multi-region, multi-provider, or bare-metal infrastructure is advantageous.
• Knowledge of GPUs, model serving, or inference systems like vLLM or SGLang is a plus.
• Experience in constructing highly available services, multi-tenant platforms, or distributed data systems is desirable.
• Proven ability to lead teams through rapid growth while maintaining technical excellence and sustainable operations is a plus.
• Competitive salary and performance-based bonuses.
• Comprehensive health benefits and wellness programs.
• Opportunities for professional development and career advancement.
• Flexible work hours and remote work options.
• Engaging team culture and collaborative work environment.
Colonist
Yopeso
Orion Innovation
Climavision
Get handpicked remote jobs straight to your inbox weekly.