
Senior Infrastructure Operations Engineer
Posted Aug 25

Posted Aug 25
This is a fully remote position, open to applicants in India.
• Take ownership of the design, management, and ongoing enhancement of KLDiscovery’s global compute, storage, and cloud infrastructure.
• Provide comprehensive technical oversight of physical servers, virtualization, block/file/object storage, enterprise backup, and Azure IaaS environments.
• Act as the main technical escalation point and design authority within the Compute & Storage team.
• Establish and enforce standards for compute, storage, Azure, operating systems, patching, provisioning, security, and observability.
• Lead initiatives for compute modifications, cluster expansions, platform upgrades, capacity planning, performance management, and remediation efforts.
• Oversee the enterprise backup and recovery strategy, validation processes, RPO/RTO alignment, and backup coverage methodologies.
• Manage and govern Azure IaaS infrastructure, cloud expenses, hybrid cloud architecture, and contribute to the roadmap.
• Own the standards for Windows Server and Linux, including patching, hardening baselines, server build runbooks, provisioning standards, and handoff checklists.
• Integrate security controls and support audit and compliance evaluations.
• Collaborate with Automation & Observability to ensure monitoring, alerting, infrastructure automation, and IaC adoption.
• Lead the management of escalations, root cause analyses, post-incident reviews, blameless retrospectives, and systemic remediation.
• Assess opportunities for infrastructure modernization, consolidation, upgrades, replacements, and cloud migration.
• Own the performance KPIs for infrastructure, including availability, capacity utilization, incident response SLAs, and patching compliance.
• Coordinate with Database, Enterprise Platforms, Automation & Observability, Networking, and DC Operations teams.
• Manage infrastructure projects from initiation to completion and oversee third-party vendors.
• Establish documentation standards for Compute & Storage and mentor Infrastructure Operations Engineers.
• Participate in the 24x7x365 on-call rotation.
• 6+ years of experience in infrastructure engineering with direct ownership across compute, storage, and cloud in a production enterprise setting.
• Expert-level proficiency in administering VMware vSphere and Nutanix within enterprise production environments.
• Experience in cluster design, capacity planning, and lifecycle management.
• Extensive knowledge of block and file storage technologies, including RAID, SAN (Fibre Channel or iSCSI), NAS protocols, and enterprise storage array management; experience with Hitachi or equivalent is required.
• Expert-level expertise in Veeam or a comparable enterprise backup platform.
• Strong understanding of Azure IaaS architecture; hybrid cloud experience is essential.
• Advanced Windows Server administration skills, including Active Directory, DNS, DHCP, Group Policy, clustering, and server hardening.
• Strong administration skills in Linux (Ubuntu).
• Proficiency in advanced PowerShell and/or Bash scripting.
• Familiarity with IaC concepts and tools such as Ansible or Terraform.
• Basic understanding of ISO 27001 and CIS Controls as they relate to infrastructure.
• CompTIA Security+ or an equivalent level of understanding is expected.
• Knowledge of Docker and Kubernetes is a plus.
• Proficient in ITSM and ITIL-based Incident, Problem, Change, and Capacity Management.
• Experience in delivering and managing technical projects from start to finish, including vendor management and stakeholder communication.
• Strong communication skills with peers, management, vendors, and internal customers.
• Experience supporting 24x7 global production environments; on-call availability is required.
• Bachelor’s degree in computer science, Information Technology, or equivalent experience.
• Competitive total compensation package, including base salary, bonus potential, comprehensive benefits, wellness programs, and perks.
• Paid time off, which includes Casual, Earned, Sick, Special Leave, and Holidays.
• Continuous learning and development opportunities through training and education reimbursement programs.
• A diverse and inclusive workplace culture.
• Access to a free, enjoyable, interactive, and incentivized global wellness program.
• A mission-driven team environment.
Pragmatike
Pragmatike
Aalyria
appsoluts GmbH
Get handpicked remote jobs straight to your inbox weekly.