
Engineer II, Site Reliability
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in United Kingdom.
• Design and implement automation and tools through software for essential solutions and services that support extensive distributed systems.
• Manage and engineer Linux systems across thousands of bare-metal servers and virtual environments.
• Take ownership of platform availability, latency, throughput, monitoring, incident response, deployment, and capacity planning.
• Engage in an on-call rotation.
• Diagnose server hardware issues.
• Ensure the platform operates reliably around the clock.
• Learn and advocate for new technologies and methodologies within the team.
• Acquire extensive exposure to the overall architecture and process workflow.
• Deliver minor development projects and occasionally larger initiatives.
• Utilize monitoring and telemetry tools such as ELK, Prometheus, Grafana, and Zabbix.
• Collect and evaluate operating system and application metrics for performance tuning and fault detection.
• Lead incident analysis, promote incident-response practices, link incidents to systemic issues, and drive resolutions.
• Collaborate with Site Reliability Engineers (SREs) and engineers distributed globally.
• Communicate and present the conventions followed by the reliability team.
• Leverage AI technologies to improve decision-making, streamline workflows and processes, enhance efficiency, and drive business outcomes.
• Bachelor's degree and/or equivalent experience in Computer Science.
• At least five years of experience in a large-scale production environment.
• Minimum of two years of experience in software engineering.
• At least two years of experience in one or more programming languages: C++, Java, Python, or Go.
• Familiarity with storage technologies such as SAN, NAS, NFS, Object Storage, FreeNAS, and iSCSI.
• Knowledge of infrastructure technologies including Linux, Windows, VMware, Docker, and Kubernetes.
• Experience in writing technical documentation.
• Configuration management experience with tools like Puppet, Chef, Ansible, or similar.
• Strong understanding of application design and operational trade-offs.
• Analytical abilities combined with a strong sense of urgency, ownership, and motivation.
• Capacity to work effectively in a diverse, team-oriented environment with SREs and engineers.
• Ability to communicate broadly and present recommended conventions.
• Proven experience using AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency, and drive business outcomes.
• Industry-leading compensation and equity awards.
• Comprehensive wellness programs for physical and mental health.
• Competitive vacation and holiday policies for relaxation.
• Paid parental and adoption leave.
• Professional development opportunities available to all employees, regardless of level or role.
• Employee Networks, geographic neighborhood groups, and volunteer opportunities to foster connections.
• Vibrant office culture featuring world-class amenities.
• Certified as a Great Place to Work™ worldwide.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.