
Incident Response Engineer – Facility Operations Center
Posted Jul 25

Posted Jul 25
This is a fully remote position, open to applicants in Australia.
• The primary responsibility is to facilitate coordination and communication across NVIDIA’s datacenter portfolio from an operational viewpoint concerning incidents, maintenance, and reporting/monitoring.
• Establish standards and programs to support reliability and operations initiatives, including Problem and Change Control, while defining and maintaining a health score for sites and environments, incorporating testing methods to predict and isolate points of failure, evaluating and advising on maintenance strategies, and delivering pertinent reporting and metrics.
• Analyze failure data and collaborate with machine learning and AI teams and tools to forecast future failures, and assist in reliability studies such as critical assessments, RAM models, and RCM studies.
• Identify and drive opportunities for automation and process enhancements across catalog quality workflows and reporting.
• Organize disaster recovery tests, communicate during audits, work in partnership with internal stakeholders, and make essential strides to ensure business continuity and compliance.
• Conduct risk assessments to guarantee adherence to policies, procedures, rules & regulations, and data center standards.
• Own and present comprehensive key business metrics related to incident response, including the ownership and representation of both internal and external tooling.
• Lead root cause analysis for outages and modify documentation, workflows, and operating procedures to prevent future incidents.
• Evaluate opportunities for process improvement and transformation, collaborating with process owners and partners to define problem statements and objectives, and to structure projects and teams.
• Work collaboratively with other team members and groups within the organization, developing strong, productive relationships across peer organizations to advance the organization's business objectives.
• This role will also involve training, coaching, and mentoring Operations teams as necessary to empower them in utilizing operations tools and systems to meet daily business requirements.
• Bachelor’s degree in a related field (e.g., Electrical Engineering, Mechanical Engineering, Industrial Engineering, Computer Engineering, Telecommunication Engineering, Computer Science, or a business-related field) or equivalent experience.
• Over 5 years of experience in operations or environmental, health, and safety within data centers.
• Proficient in developing and implementing reliability activities (modeling predictions, life cycle testing, stress testing, etc.).
• Strong understanding of commercial and financial aspects, with a complete grasp of the implications of failure in relation to business costs, production targets, and customer order fulfillment.
• Highly developed numerical, statistical, and reporting capabilities; able to analyze, interpret, and apply information, data, and trends effectively.
• Passionate about achieving objectives and maintaining organization, with the ability to strategize and meet established goals.
• Proven ability to be detail-oriented, organized, and capable of consolidating data analyses for presentation to large groups.
• Skilled in utilizing asset databases and DCIM solutions to extract data and generate valuable insights.
• Experience in designing, deploying, or maintaining large-scale datacenter infrastructure (either ACSMEP or networking) or the capacity to create strategic infrastructure roadmaps incorporating on-premise, hybrid, and cloud technologies.
• Demonstrated knowledge and advanced proficiency with Microsoft Office Suite and G-Suite software.
• Health insurance
• Professional development opportunities
Lime
Threatscape
GFT Technologies
BeyondTrust
Get handpicked remote jobs straight to your inbox weekly.