IT Infrastructure Support, Site Reliability Engineer II

atAstreyaRemoteIE flagIrelandFull-timeDevOps & Site Reliability Engineer (SRE)Mid-levelSenior€50.6k – €63.3k/year

Posted Aug 20

This is a fully remote position, open to applicants in Ireland.

📋 Description

• Ensure the dependability, scalability, and performance of essential physical security infrastructure, which includes IP camera systems, access control mechanisms, Cisco switches, servers, networks, and cloud environments.

• Develop and sustain automation tools, a centralized Configuration Management Database (CMDB), monitoring systems, and enterprise-level infrastructure management processes.

• Collaborate with leadership to establish, monitor, and enforce infrastructure Service Level Indicators (SLIs) and Service Level Objectives (SLOs).

• Provide Level 3 expertise for tooling-related incidents.

• Automate incident remediation workflows and create runbooks to minimize Mean Time to Recovery (MTTR).

• Identify and automate repetitive manual activities through scripting and workflow automation.

• Perform root cause analyses and lead blameless postmortems for significant service-impacting incidents.

• Engineer automated processes and scripts for asset management platforms, CMDBs, and monitoring systems.

• Design, develop, and implement full-stack applications, custom plugins, and automation scripts for management and monitoring systems.

• Maintain Infrastructure-as-Code configurations for both Windows and Linux server roles, including drift detection and auto-remediation.

• Create automation pipelines for vulnerability patching, enforcement of CIS security baselines, and continuous compliance auditing.

• Develop API-driven tools for managing network configurations, firmware updates, zero-touch provisioning, validation, and monitoring network health.

• Deploy and standardize monitoring agents, centralized logging, dashboards, and alerts.

• Construct custom monitoring exporters for physical security devices and camera systems.

• Develop diagnostic tools for distributed log timestamp correlation and detection of clock drift/NTP desynchronization.

• Create automation scripts for ticket management, problem validation, and escalation workflows while adhering to 2-hour initial response SLAs.

• Support managed credential/access controls and automated configuration backups.

• Participate in a 24x5 on-call rotation to maintain service continuity and ensure rapid incident response.

• Work alongside cross-functional teams to define service objectives, minimize operational toil, and enhance system resilience.


⛳️ Requirements

• A minimum of 6 years of experience in Infra Automation Engineering or Infrastructure Engineering.

• Strong expertise in Python, Bash, and PowerShell.

• Experience with Go for developing high-performance backend services and APIs.

• Practical experience with Terraform, Ansible, Chef, or Puppet.

• Proficient in configuration management, including drift detection, version control, and automated remediation.

• Advanced understanding of Linux and Windows server environments.

• Tier 3 troubleshooting skills.

• Experience in system hardening and managing enterprise-scale servers.

• Knowledge of enterprise networking concepts and Cisco device administration.

• Familiarity with NETCONF/RESTCONF, network monitoring, and flow analysis tools.

• Experience with monitoring systems such as Prometheus, Grafana, Datadog, Monarch, Streamz, or similar.

• Proficient with centralized logging platforms like ELK Stack.

• Ability to create custom dashboards and alerting rules.

• Experience in deploying and customizing a CMDB/IPAM platform such as NetBox.

• Comfortable operating within a large-scale, cloud-hosted enterprise environment.

• Familiarity with Kubernetes, Terraform, and Helm.

• Understanding of internal development and code-review tools such as Cider and Critic, or similar toolchains.

• Experience in writing and maintaining custom monitoring exporters/agents for edge/IoT and physical security devices.

• Knowledge of structured, glog-style logging output.

• Proficient in advanced text processing and scripting, including awk/gawk.

• Working knowledge of NTP/clock synchronization practices.

• Availability for 24x5 support and on-call rotation.


🏝️ Benefits

• Base salary ranging from €50,640 to €63,300 gross annually.

• Potential for performance-based bonuses.

• Possible benefits-related payments.

• Additional general incentives may apply.

• Opportunities for learning, collaboration, and career advancement.

• An inclusive culture that values diverse perspectives and offers opportunities to make an impact.

People also viewed

TEKsystems2 hours ago

Cloud Deployment Engineer, Secret Clearance Required – 30% Travel

US flagDistrict of Columbia, +1 more stateFreelanceDevOps & Site Reliability Engineer (SRE)$80 – $110/hour
ApplyView job
Arctiq14 hours ago

Site Reliability Engineer – Vulnerability Remediation Consultant

US flagUnited States OnlyFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
GE Vernova15 hours ago

Senior Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$152.4k – $254k/year
ApplyView job
Ecosistemas15 hours ago

Senior DevOps Engineer, Bilingüe Inglés

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Intetics1 day ago

Senior DevOps Engineer

PH flagPhilippines OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
OZmap1 day ago

Mid-Level SRE

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers