
Platform Reliability Engineer
Posted Aug 7

Posted Aug 7
This is a fully remote position, open to applicants in United States.
• Create and manage reusable Terraform modules, Helm charts, GitOps templates, and common platform services.
• Build self-service platform features and streamlined processes that enhance software delivery speed.
• Design, standardize, and advance enterprise CI/CD pipelines incorporating automated testing, security checks, deployment governance, and release automation.
• Develop highly available Kubernetes platforms prioritizing security, networking, scalability, and operational excellence.
• Enhance platform observability utilizing metrics, logs, traces, dashboards, alerts, SLIs, SLOs, and error budgets.
• Assist engineering teams during production incidents, conduct post-incident evaluations, and implement lasting improvements for reliability.
• Reduce manual operational tasks through automation.
• Collaborate with engineering teams to boost application reliability, resilience, performance, and operational preparedness.
• Optimize cloud expenditures, platform efficiency, and resource usage.
• Work alongside Software Engineering, Architecture, Security, Data Engineering, and MLOps teams.
• Set engineering standards and enhance platform reliability throughout the engineering organization.
• Over 5 years of experience in Platform Engineering, Site Reliability Engineering, or Cloud Infrastructure.
• Hands-on experience with Azure and/or AWS in production environments.
• Practical experience with Kubernetes, including AKS, EKS, or GKE.
• Proficiency in Infrastructure as Code using Terraform or similar tools.
• Familiarity with modern CI/CD platforms such as GitHub Actions or Azure DevOps.
• Experience with GitOps methodologies using tools like ArgoCD or Flux.
• Development of Helm charts.
• Knowledge of Linux systems administration and container technologies.
• Scripting skills in PowerShell, Bash, Python, or Go.
• Experience with enterprise observability tools like Datadog, Prometheus, Grafana, LogicMonitor, or New Relic.
• Understanding of cloud networking and identity technologies, including OIDC.
• Strong troubleshooting, communication, and cross-functional collaboration abilities.
• Proven experience collaborating with Software Engineering, Cloud Infrastructure, Architecture, Security, Data Engineering, and MLOps teams.
• Access to experts and resources for your Learning & Development journey.
• Opportunities for internal mobility.
• Employee referral bonus program.
• Employee Resource Groups (ERGs).
• Annual fundraising and volunteer events to support community engagement.
• Paid time off, floating holidays, volunteer time off, and rollover options.
• Paid parental leave.
• Comprehensive medical, dental, vision, and 401k plans (with matching contributions).
• Flexible spending accounts, mass transit, and dependent care plans available.
• Health savings account with an annual company contribution for eligible participants.
• Short-term and long-term disability coverage.
• Company-subsidized life insurance policies.
• Pet insurance options.
• Accident care coverage.
• Access to legal advice services.
• Annual incentives based on individual and company performance.
• Flexible scheduling based on coverage requirements.
Faire
PerfectServe
Makpar Corporation
Bounteous
Get handpicked remote jobs straight to your inbox weekly.