
Lead Site Reliability Engineer – Cloud
Posted 19 hours ago

Posted 19 hours ago
This is a fully remote position, open to applicants in France.
• Facilitate the team's operational methods, encompassing processes, rituals, and documentation.
• Assist the team in prioritizing tasks, and evaluate technical decisions and implementations.
• Support the professional growth of team members and encourage independence.
• Advocate for SRE best practices: reliability, observability, incident management, and automation.
• Promote the SRE technical vision in key strategic initiatives.
• Assess performance, pinpoint bottlenecks, and enhance resource utilization and scalability.
• Define, implement, and refine observability tools, including monitoring, metrics, logs, and alerting.
• Document, maintain, and evolve operational processes.
• Keep abreast of advancements in infrastructure practices and technologies.
• Provide a portion of Level 3 customer support in alignment with SLAs.
• Engage in incident management and on-call rotations, approximately half a week every three weeks.
• Respond to critical incidents to mitigate their effects and ensure service continuity.
• Lead incident retrospectives and establish sustainable corrective measures.
• Compose and distribute post-mortem reports following significant incidents.
• Coordinate crisis communication both internally and with clients.
• Ensure adherence to SLA, RPO, and RTO commitments.
• Establish service quality metrics and SLOs.
• Contribute to ISO 27001 and HDS compliance, along with both internal and external audits.
• Plan, execute, and assess business continuity and disaster recovery drills.
• Collaborate with development teams to embed operability requirements from the design phase.
• Advise Product Engineering teams on reliability, customer experience, and administration tools.
• Ensure operational documentation is clear and current.
• Provide technical and operational guidance without direct line-management responsibility, reporting to an Engineering Manager.
• Extensive expertise in cloud environments and distributed infrastructure, with a strong emphasis on high availability and production reliability.
• Proficient in observability practices, including logs, metrics, and alerting, with a methodical approach to diagnosing complex incidents.
• Strong understanding of containerized environments and their operational challenges.
• Proven experience with production databases, including reliability, backups, restoration, replication, and scalability.
• Familiarity with Infrastructure as Code and environment automation.
• Understanding of operational security considerations.
• Comfortable using Artificial Intelligence tools to enhance daily efficiency.
• Capable of working rigorously and reliably in complex, evolving, or uncertain environments.
• Proficient in prioritizing tasks, including during incidents.
• Clear and organized communication abilities, with a collaborative mindset and a passion for knowledge sharing.
• A blameless attitude, technical curiosity, composure, and a strong focus on user impact.
• Ability to provide technical leadership, share expertise, and advance collective practices.
• This position must be based exclusively in France.
• Fully remote, with one trip per quarter to Strasbourg or another city.
• Company events: one annual offsite and regular after-work gatherings.
• Remote-work allowance (€57.60).
• Meal vouchers (€11.52 each) and a Swile card with additional benefits.
• Flexible working hours under a forfait-hours arrangement, including RTT days.
• Linux laptop.
• Budget for additional equipment, with a company contribution.
Ninja - نينجا
Dreamix
inDrive
OmegaHires
Get handpicked remote jobs straight to your inbox weekly.