
NOC Engineer / SRE
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in United Kingdom.
• Serve as a primary or escalation contact in a 24x7 on-call rotation.
• Lead or assist in Major Incident responses, encompassing triage, mitigation, and resolution.
• Collaborate with Engineering, Infrastructure, Security, and Product teams.
• Implement and enhance runbooks, playbooks, and escalation procedures.
• Facilitate blameless post-incident reviews and monitor corrective actions.
• Oversee service health monitoring across infrastructure, applications, and dependencies.
• Design and uphold alerting strategies that align with SLIs/SLOs.
• Mitigate alert fatigue by enhancing signal-to-noise ratios.
• Create dashboards utilizing Grafana, Prometheus, Datadog, Splunk, and/or CloudWatch.
• Automate repetitive operational tasks to minimize manual effort.
• Enhance mean time to detect and mean time to resolve incidents.
• Develop scripts and tools in Python, Bash, Go, or comparable languages.
• Implement self-healing and auto-remediation solutions where feasible.
• Collaborate with engineering teams to boost system reliability.
• Support and troubleshoot Linux systems, cloud platforms, and Kubernetes/containerized environments.
• Aid in capacity planning and availability assessments.
• Ensure operational readiness for production deployments.
• Proficient in Linux systems administration.
• Experience in incident management and production support.
• Familiarity with cloud infrastructure, preferably AWS.
• Knowledge of Docker and Kubernetes.
• Familiarity with monitoring and alerting tools.
• Scripting or programming skills in Python, Bash, Go, or similar languages.
• Understanding of networking fundamentals, such as DNS, TCP/IP, and load balancing.
• Experience in 24x7 NOC or production operations environments.
• Ability to manage high-pressure incidents with composure and effectiveness.
• Strong written and verbal communication skills for incident coordination.
• Comfortable working from runbooks and refining them as necessary.
• Experience defining or adhering to SLOs/SLIs is preferred.
• Previous migration experience from traditional NOC to SRE model is preferred.
• Experience with Infrastructure as Code tools like Terraform, Ansible, or similar is preferred.
• Exposure to security, compliance, or regulated environments is preferred.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Opportunities for professional development and growth.
• Flexible work hours and remote work options.
• Supportive team culture and collaborative environment.
AlphaSense
futureproof consulting
Amgen
Wing Assistant
Get handpicked remote jobs straight to your inbox weekly.