
Senior Site Reliability Engineer – CloudVision
Posted Aug 4

Posted Aug 4
This is a fully remote position, open to applicants in Ireland.
• Design, develop, and implement production systems with a focus on scalability, reliability, observability, performance, and security.
• Create and maintain automation solutions aimed at reducing manual effort and enhancing operational efficiency.
• Oversee production systems, set up alerting mechanisms, and execute automated incident responses.
• Produce incident response runbooks and carry out postmortem evaluations.
• Work closely with software engineering teams to identify and resolve infrastructure constraints and streamline deployment processes.
• Manage and enhance monitoring infrastructure to ensure comprehensive system visibility.
• Plan, communicate, and carry out production maintenance windows with minimal disruption to services.
• Diagnose platform and infrastructure issues, engaging vendors and support teams as necessary.
• Implement systems and updates through controlled, risk-managed rollouts.
• Investigate and integrate best practices in infrastructure and platform management.
• Analyze open-source system design and implementation details to enhance troubleshooting methods.
• Relay system status, maintenance schedules, and infrastructure advancements to stakeholders.
• Bachelor's degree in Computer Science, Engineering, or equivalent professional experience.
• Over 5 years of experience in a relevant infrastructure or systems position.
• Proficient in Go, Python, or bash shell scripting.
• Capability to implement automation workflows of medium complexity.
• Solid knowledge of Linux or UNIX administration and debugging techniques.
• Practical experience managing software systems, infrastructure, and complex applications at a production level.
• Familiarity with infrastructure-as-code principles and practices.
• Strong analytical and software troubleshooting abilities.
• Experience in server provisioning, storage solutions, and networking.
• Proven ability to collaborate across teams and communicate technical concepts effectively.
• Background in incident response, postmortem analysis, and ongoing improvement processes.
• Experience with Kubernetes, Docker, and virtualization technologies is a plus.
• Proficiency in Prometheus and Grafana is desirable.
• Familiarity with GitLab tools or Spinnaker is advantageous.
• Knowledge of Terraform is a bonus.
• Experience with PostgreSQL or similar relational databases is preferred.
• Familiarity with artifact repositories and Docker registries is a plus.
• Knowledge of Google Cloud Platform, Amazon Web Services, or Microsoft Azure is desirable.
• Understanding of distributed systems architecture would be beneficial.
• Experience in performance tuning and system optimization is desirable.
• Awareness of infrastructure and systems security best practices is advantageous.
• Experience with on-call support and incident response is a plus.
• Remote work opportunity from Ireland.
• Permanent employment status.
• Collaborate with cross-functional teams and engage with various company domains.
• Full ownership of projects undertaken.
• Flat organizational structure promoting streamlined management.
• Opportunities to work across multiple domains.
• Access to all areas of the company.
• Culture centered around test automation tools and engineering.
• An inclusive environment that values diverse thoughts and perspectives.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.