
Infrastructure Automation & Observability Engineer
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in India.
• Develop, construct, and uphold Grafana dashboards and telemetry visualizations to track system performance, latency, error rates, saturation, and the overall health of the platform.
• Set up and manage observability solutions, such as Prometheus monitoring and alerting, ensuring that critical service metrics and SRE Golden Signals are effectively monitored.
• Create, test, and sustain modular Ansible playbooks to streamline infrastructure provisioning, application configuration, patching, and operational workflows.
• Facilitate enterprise automation via AWX by overseeing centralized execution, automation templates, RBAC, and reusable infrastructure workflows.
• Maintain Infrastructure as Code (IaC) repositories utilizing Git, adhering to best practices for version control, peer code reviews, and CI/CD-driven automation.
• Engage actively in Agile ceremonies, including Sprint Planning, Daily Stand-ups, Backlog Refinement, and Retrospectives, contributing to sprint execution and ongoing improvement.
• Partner with Product Owners, Scrum Masters, and engineering teams to convert business requirements into actionable user stories while advancing automation, observability, and platform reliability initiatives.
• A minimum of 4 years of experience in DevOps, Infrastructure Automation, Platform Engineering, or Site Reliability Engineering (SRE).
• Direct experience in building and maintaining Grafana dashboards, telemetry visualizations, and observability solutions.
• Hands-on experience in developing and maintaining Ansible playbooks for infrastructure provisioning, configuration management, and automation.
• Proficient in configuring monitoring and alerting systems using Prometheus and Grafana.
• Expertise in Git version control, peer code reviews, and CI/CD workflows.
• Familiarity with Python and/or Bash scripting.
• Strong Linux administration skills (Ubuntu, RHEL, or similar).
• Experience utilizing Jira in Agile/Scrum environments.
• Preferred experience with Kubernetes and containerized workloads.
• Knowledge of metrics storage platforms such as Prometheus, Mimir, or Thanos.
• Basic understanding of GitLab CI or alternative CI/CD platforms.
• Awareness of Infrastructure as Code (IaC) and DevOps best practices.
• Familiarity with ITIL processes, change management, and enterprise operational workflows.
• Certified ScrumMaster (CSM), Professional Scrum Master (PSM I), or equivalent certification from Scrum Alliance/Scrum.org.
• Culture of Relentless Performance: join an unstoppable technology development team with a 99% project success rate and over 30% year-over-year revenue growth.
• Competitive Pay and Benefits: enjoy a comprehensive compensation and benefits package, including health insurance, language courses, and a relocation program.
• Work From Anywhere Culture: take advantage of the flexibility that comes with remote work.
• Growth Mindset: benefit from various professional development opportunities, including certification programs, mentorship, talent investment programs, internal mobility, and internship opportunities.
• Global Impact: collaborate on significant projects for top global clients and help shape the future of industries.
• Welcoming Multicultural Environment: be part of a dynamic, global team and thrive in an inclusive and supportive work atmosphere with open communication and regular team-building social events.
• Social Sustainability Values: engage with our sustainable business practices focused on five pillars, including IT education, community empowerment, fair operating practices, environmental sustainability, and gender equality.
AM53 Smart Solutions
ICF
Miratech
Software Mind
Get handpicked remote jobs straight to your inbox weekly.