
Senior Site Reliability Engineer
Posted 15 hours ago

Posted 15 hours ago
This is a fully remote position, open to applicants in Europe.
• Develop and enhance Kubernetes infrastructure for provisioning, networking, storage, and service deployment across various providers and regions.
• Create isolation, failover, and recovery strategies to minimize the effects of hardware and infrastructure failures.
• Automate the processes of capacity expansion, deployments, and maintenance.
• Establish observability and diagnostics to identify bottlenecks and highlight failures.
• Address incidents, determine root causes, and enhance systems based on insights gained.
• Enhance platform performance, utilization, security, and reliability as inference demand increases.
• Work collaboratively with infrastructure, platform, and inference engineers within a flat organizational structure.
• Proven experience in building and managing production infrastructure or distributed systems, with a strong sense of ownership over reliability.
• Solid understanding of Linux fundamentals.
• Practical expertise in networking, storage, and containerization.
• Direct experience running Kubernetes in production environments.
• Ability to write maintainable software and automate solutions for infrastructure challenges.
• Systematic methodology for troubleshooting issues across application, cluster, network, and hardware domains.
• Sound judgment on when to act swiftly, simplify processes, and prioritize reliability.
• Proactive in taking issues from investigation to implementation, collaborating effectively with team members.
• Preferred: experience with multi-region, multi-provider, or bare-metal infrastructure.
• Preferred: familiarity with GPUs, model serving, vLLM, or SGLang.
• Preferred: experience with infrastructure as code, CI/CD, observability, or automated recovery.
• Preferred: experience in developing highly available services, multi-tenant platforms, or distributed data systems.
• Ownership and the ability to influence how Parasail scales its inference cloud.
• Significant architectural decisions with direct impact on production.
• Opportunity to work closely with hardware, delve into distributed systems, and collaborate with inference engineers.
• Chance to contribute to the foundational development of the next generation of AI infrastructure.
Colonist
Yopeso
Orion Innovation
Climavision
Get handpicked remote jobs straight to your inbox weekly.