Senior Site Reliability Engineer

Posted 15 hours ago

This is a fully remote position, open to applicants in Europe.

📋 Description

• Develop and enhance Kubernetes infrastructure for provisioning, networking, storage, and service deployment across various providers and regions.

• Create isolation, failover, and recovery strategies to minimize the effects of hardware and infrastructure failures.

• Automate the processes of capacity expansion, deployments, and maintenance.

• Establish observability and diagnostics to identify bottlenecks and highlight failures.

• Address incidents, determine root causes, and enhance systems based on insights gained.

• Enhance platform performance, utilization, security, and reliability as inference demand increases.

• Work collaboratively with infrastructure, platform, and inference engineers within a flat organizational structure.


⛳️ Requirements

• Proven experience in building and managing production infrastructure or distributed systems, with a strong sense of ownership over reliability.

• Solid understanding of Linux fundamentals.

• Practical expertise in networking, storage, and containerization.

• Direct experience running Kubernetes in production environments.

• Ability to write maintainable software and automate solutions for infrastructure challenges.

• Systematic methodology for troubleshooting issues across application, cluster, network, and hardware domains.

• Sound judgment on when to act swiftly, simplify processes, and prioritize reliability.

• Proactive in taking issues from investigation to implementation, collaborating effectively with team members.

• Preferred: experience with multi-region, multi-provider, or bare-metal infrastructure.

• Preferred: familiarity with GPUs, model serving, vLLM, or SGLang.

• Preferred: experience with infrastructure as code, CI/CD, observability, or automated recovery.

• Preferred: experience in developing highly available services, multi-tenant platforms, or distributed data systems.


🏝️ Benefits

• Ownership and the ability to influence how Parasail scales its inference cloud.

• Significant architectural decisions with direct impact on production.

• Opportunity to work closely with hardware, delve into distributed systems, and collaborate with inference engineers.

• Chance to contribute to the foundational development of the next generation of AI infrastructure.

People also viewed

Colonist15 hours ago

DevOps Engineer

PT flagPortugal OnlyFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
Yopeso15 hours ago

Reliability Engineer / DevOps – Database Platform

RO flagRomania OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Orion Innovation15 hours ago

DevOps

MX flagMexico OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Climavision15 hours ago

Senior Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $170k/year
ApplyView job
NICE16 hours ago

Cloud Operations Engineer

GB flagUnited Kingdom OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Humana16 hours ago

Senior DevOps Engineer

US flagDistrict of Columbia, +2 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$106.9k – $147k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers