
Senior Site Reliability Engineer, Platform Reliability
Posted Aug 7

Posted Aug 7
This is a fully remote position, open to applicants in Spain.
• Establish frameworks and best practices for application operations
• Develop tools and offer training and consulting services
• Deliver data insights and visibility regarding application performance
• Assist in the formulation of Service Level Objectives (SLOs)
• Lead efforts in Incident Management and Analysis
• Oversee Change Management and Deployment methodologies
• Participate in discussions regarding services and architecture
• Recommend configurations for observability and alerting
• Take ownership of and achieve quarterly team objectives
• Guide engineers through complex, open-ended challenges while supporting their delivery
• Collaborate with teams in infrastructure, product management, developer experience, and analytics throughout the product development lifecycle
• Identify technical solutions and operational processes that enhance incident preparedness, response, and post-incident evaluation
• Establish and monitor metrics, escalate issues when necessary, and support availability and on-call initiatives
• Set or enhance code review and design standards, advocating for them through writing and technical presentations
• Foster team talent through constructive feedback, guidance, and by setting an example
• Engage in the necessary on-call rotation
• Minimum of 4 years experience in designing, developing, and deploying backend systems at scale
• Proficient in scripting and development languages such as Bash, Python, or Kotlin
• Proven history of developing highly available distributed systems
• Familiarity with AWS, MySQL, and Kubernetes
• Significant experience in contributing to or leading Incident Lifecycle processes
• At least 4 years of experience on a Site Reliability or Production Engineering team
• Experience in defining technical plans for major features or system components
• Capability to produce high-quality, clear, and reusable code
• Background in implementing impactful changes within a large codebase
• Experience in creating tools and practices for safe modifications
• Excellent verbal and written communication skills
• Willingness to participate in a required on-call rotation
• Monthly stipends for health, wellness, and technology expenses
• Complete medical coverage for you and your dependents at no cost
• Dental and vision insurance for you and your dependents
• Flexible Spending Wallets for technology, food, and lifestyle expenses
• Away Days — wellness days for you to take off work and recharge
• Opportunities for Learning & Development
• Parental benefits
• Employee Resource & Community Groups
• Competitive vacation and holiday policies
• Employee Stock Purchase Plan (ESPP) allowing employees to purchase shares of Affirm at a discounted rate
• Reasonable accommodations available during the hiring process
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.