
Middle SRE – Infrastructure Automation Engineer
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in India.
• Design, develop, and uphold Grafana dashboards and telemetry visualizations to oversee system performance, latency, error rates, saturation, and the overall health of the platform.
• Set up and manage observability solutions, including Prometheus for monitoring and alerting, ensuring the effective tracking of critical service metrics and SRE Golden Signals.
• Create, test, and sustain modular Ansible playbooks to automate infrastructure provisioning, application configuration, patching, and operational workflows.
• Facilitate enterprise automation through AWX by administering centralized execution, automation templates, RBAC, and reusable infrastructure workflows.
• Maintain Infrastructure as Code (IaC) repositories utilizing Git, adhering to best practices for version control, peer code reviews, and CI/CD-driven automation.
• Actively engage in Agile ceremonies such as Sprint Planning, Daily Stand-ups, Backlog Refinement, and Retrospectives, contributing to sprint execution and ongoing enhancement.
• Collaborate with Product Owners, Scrum Masters, and engineering teams to convert business requirements into actionable user stories while advancing automation, observability, and platform reliability initiatives.
• A minimum of 4 years of experience in DevOps, Infrastructure Automation, Platform Engineering, or Site Reliability Engineering (SRE).
• Hands-on experience in building and maintaining Grafana dashboards, telemetry visualizations, and observability solutions.
• Practical experience in developing and maintaining Ansible playbooks for infrastructure provisioning, configuration management, and automation.
• Experience in configuring monitoring and alerting systems using Prometheus and Grafana.
• Proficiency in Git version control, peer code reviews, and CI/CD workflows.
• Working knowledge of Python and/or Bash scripting.
• Strong Linux administration skills (Ubuntu, RHEL, or similar).
• Experience using Jira in Agile/Scrum environments.
• Nice to have: Familiarity with Kubernetes and containerized workloads.
• Knowledge of metrics storage platforms such as Prometheus, Mimir, or Thanos.
• Basic understanding of GitLab CI or other CI/CD platforms.
• Awareness of Infrastructure as Code (IaC) and DevOps best practices.
• Familiarity with ITIL processes, change management, and enterprise operational workflows.
• Certified ScrumMaster (CSM), Professional Scrum Master (PSM I), or equivalent Scrum Alliance/Scrum.org certification.
• Health insurance.
• Language courses.
• Relocation program.
• Professional development opportunities.
• Certification programs.
• Mentorship and talent investment programs.
• Internal mobility and internship opportunities.
• Welcoming Multicultural Environment.
ICF
AM53 Smart Solutions
Trimetis AG
Software Mind
Get handpicked remote jobs straight to your inbox weekly.