
Principal Operations Engineer, Reliability
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in United States.
• Take charge of fleet reliability engineering: establish availability targets, assess them transparently, and bridge any gaps.
• Conduct root cause analyses on the fleet's most significant incidents and implement corrective actions throughout all sites.
• Develop a failure data pipeline for both facilities and hardware that transforms incident history into engineering priorities.
• Formulate the maintenance strategy (reliability-centered, condition-based) to ensure the fleet concentrates efforts where failure data indicates.
• You have taken responsibility for the reliability of critical infrastructure and successfully influenced the availability metrics, not just reported them.
• You have spearheaded root cause analyses that identified the genuine cause rather than the more convenient one.
• You are proficient in working with failure data: Weibull, Pareto, and FMEA are tools you actively utilize, not just terminology you recognize.
• You ensure corrective actions are completed across teams under your influence, even if you don't manage them directly.
• Bonus: Experience in data center or power generation reliability. Knowledge of liquid cooling systems. Familiarity with CMMS analytics. Possession of CRE or CMRP certification.
• Competitive base salary
• Equity offered for all full-time positions
• Comprehensive benefits package
• Commission plans available if applicable
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.