
IT Infrastructure Support Site Reliability Engineer II
Posted Jul 29

Posted Jul 29
This is a fully remote position, open to applicants in Ireland.
• Collaborate with leadership to create, oversee, and enforce Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for infrastructure tools, which include configuration compliance rates, patch success rates, and metrics on deployment latency.
• Offer Level 3 expertise for incidents related to specific tooling, with an emphasis on automating workflows for incident remediation and minimizing Mean Time To Repair (MTTR) through advanced automation and the development of runbooks.
• Recognize and automate repetitive manual processes within managed infrastructure, aiming for quantifiable decreases in operational overhead (e.g., achieving a 50% reduction in manual server build time) via scripting and workflow automation.
• Perform comprehensive root cause analysis and lead blameless postmortems for significant service-impacting incidents, driving systemic enhancements in tooling reliability and infrastructure robustness.
• Design and sustain automated processes and scripts for populating, updating, and synchronizing asset management platforms (such as NetBox), configuration management databases, and monitoring systems for both internal and external stakeholders.
• Conceive, develop, and deploy full-stack applications, custom plugins, and automation scripts to enhance the functionality of management and monitoring systems, facilitating direct interaction with devices for configuration management.
• Create and maintain fully automated Infrastructure-as-Code configurations for Windows and Linux server roles, utilizing tools like Ansible, Terraform, or Puppet, inclusive of drift detection and auto-remediation capabilities.
• Construct end-to-end automation pipelines for vulnerability patching, security baseline enforcement (CIS benchmarks), and ongoing compliance audits against internal and regulatory standards for physical security devices.
• Develop API-driven tools for managing network configurations, automated firmware updates, zero-touch provisioning, pre/post-change validation, and real-time monitoring of network health across the device fleet.
• Implement and standardize monitoring agents, centralized log collection systems, and custom dashboards with alerts based on critical SLIs (latency, error rate, saturation, traffic) related to servers and edge devices.
• Create and maintain custom monitoring exporters for the physical security device fleet, including camera systems, to ensure precise metrics and structured, multi-severity logging output.
• Develop diagnostic tools to correlate timestamps across distributed log streams and identify clock drift or NTP desynchronization, a common root cause of false-positive outages across the device fleet.
• Produce automation scripts for intelligent ticket handling, problem validation, and escalation workflows within enterprise ticketing systems, ensuring that 2-hour initial response SLAs are reliably met.
• Assist with foundational security enhancements across the device fleet, which include managed credential/access controls and automated configuration backups.
• Participate in a 24x5 on-call rotation to deliver timely support for infrastructure systems, security devices, and related tools, ensuring service continuity and prompt incident response.
• A minimum of 6 years of experience in Infra Automation Engineering or Infrastructure Engineering.
• Strong expertise in Python, Bash, and PowerShell for automation scripting, with experience in Go for developing high-performance backend services and APIs.
• Practical experience with Infrastructure-as-Code tools (Terraform, Ansible, Chef, or Puppet) and configuration management practices, including drift detection, version control, and automated remediation.
• Advanced understanding of Linux and Windows server environments, including Tier 3 troubleshooting skills, system hardening, and enterprise-scale server management.
• Solid grasp of enterprise networking concepts, Cisco device management, network automation protocols (NETCONF/RESTCONF), and familiarity with network monitoring and flow analysis tools.
• Experience in implementing and managing monitoring solutions (Prometheus, Grafana, Datadog) or similar proprietary internal monitoring and metrics-streaming systems (e.g., Monarch, Streamz), as well as centralized logging platforms (ELK Stack), with the capability to create custom dashboards and alerting rules.
• Experience in deploying and customizing a CMDB/IPAM platform (e.g., NetBox) as a reliable source for device inventory and subsequent automation.
• Comfort operating within a large-scale, cloud-hosted enterprise environment (Kubernetes, Terraform, Helm), including familiarity with internal development and code-review tools (e.g., Cider, Critic) or comparable large-scale internal toolchains.
• Experience in writing and maintaining custom monitoring exporters/agents for edge/IoT and physical security devices, complete with structured, glog-style logging output.
• Proficiency in advanced text-processing and scripting (e.g., awk/gawk) for log parsing and timestamp correlation, along with a working knowledge of NTP/clock synchronization practices.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Generous paid time off and flexible working hours.
• Opportunities for professional development and continuous learning.
• Employee wellness programs and resources.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.