
Senior Site Reliability Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Poland, +4 more states.
• Analyze and resolve production challenges across various services and integrations.
• Conduct thorough debugging utilizing logs, traces, metrics, and database queries to determine root causes.
• Optimize system performance, focusing on PostgreSQL queries, caching mechanisms, and service behavior under load.
• Engage directly with Kubernetes environments, encompassing deployments, configurations, scaling, and troubleshooting tasks.
• Enhance observability through metrics, logging, and tracing with tools like Grafana and the ELK stack.
• Provide support and maintenance for integrations between Sportsbook services and external client platforms.
• Establish and configure new Sportsbook-related projects and environments.
• Develop internal tools and automation in Go to minimize manual tasks and operational burdens.
• Work with Kafka and RabbitMQ messaging systems, diagnosing associated issues.
• Collaborate with developers, QA, and DevOps teams to resolve incidents and enhance system stability.
• Troubleshoot real production issues, stabilize systems, and improve platform reliability on a high-load infrastructure.
• Proven experience with Go and proficiency in reading and writing production-level code.
• Strong debugging capabilities across services, logs, and data layers.
• Experience with PostgreSQL, including query performance analysis, indexing, and tuning.
• Practical experience with Kubernetes utilizing kubectl for deployments, configurations, and troubleshooting.
• Familiarity with observability tools such as Grafana, Kibana/ELK, logs, metrics, and tracing methodologies.
• Experience with Redis, including caching strategies and debugging techniques.
• Familiarity with Kafka and/or RabbitMQ, particularly regarding consumer behavior, lag, retries, and failures.
• Understanding of distributed systems under load, including handling timeouts, retries, and race conditions.
• Comfortable operating in production environments and managing incidents effectively.
• Ability to work autonomously, investigate issues thoroughly, and drive them to resolution.
• Private health insurance.
• Sports-related benefits.
• Comprehensive Mental Health Program.
• Complimentary English lessons (online).
• Local language courses available.
• Paid time off.
• Support for maternity leave.
• Rewards for referral program participation.
• Opportunities for upskilling, internal workshops, and attendance at professional conferences and corporate events.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.