
Site Reliability Engineer – Engineering Productivity
Posted Aug 7

Posted Aug 7
This is a fully remote position, open to applicants in Ireland.
• Develop, deploy securely and incrementally, and maintain essential production systems
• Emphasize scalability, reliability, observability, performance, and security
• Create automation to eliminate repetitive tasks and proactively monitor, respond to, and improve alerts
• Establish and maintain incident response runbooks
• Diagnose platform and infrastructure challenges
• Draft postmortem reports to avoid future incidents
• Organize and communicate maintenance schedules for production systems
• Collaborate with third-party vendor support when necessary
• Recognize infrastructure constraints alongside product development teams
• Design solutions to improve developer experience and workflow efficiency
• Research and implement best practices in infrastructure and platform design
• Analyze open-source system implementations to enhance triage and resolution
• A BSc in Computer Science or Engineering with 3 years of experience, an MS in Computer Science or Engineering with 3 years of experience, or equivalent practical experience
• Proficient in Go, Python, or shell scripting for developing medium-complexity automation workflows
• Familiarity with Linux or UNIX system administration and debugging
• Practical experience managing software systems at scale
• Knowledge of server provisioning, particularly in storage and networking
• Excellent problem-solving and software troubleshooting abilities
• Experience with infrastructure-as-code methodologies
• Experience with databases such as MariaDB, PostgreSQL, or MongoDB (preferred)
• Familiarity with Docker and virtualization technologies like KVM, QEMU, or Kata Containers (preferred)
• Experience managing monitoring stacks such as Prometheus, Loki, Tempo, InfluxDB, Grafana, or Thanos (preferred)
• Experience managing Elasticsearch clusters (preferred)
• Experience with Artifactory and Docker registries (preferred)
• Experience managing CI/CD systems such as ArgoCD or Spinnaker (preferred)
• Familiarity with version control systems like Perforce or Gerrit (preferred)
• Experience with infrastructure-as-code frameworks such as Ansible (preferred)
• Experience managing large Java applications (preferred)
• Experience with storage infrastructure such as NAS, SAN, or Ceph (preferred)
• Employees have the option to work remotely
• Emphasis on work-life balance
• Supportive and inclusive environment
• Commitment to diversity in the workplace
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.