
Platform Reliability Engineer
Posted 22 hours ago

Posted 22 hours ago
This is a fully remote position, open to applicants in United States.
• Create and manage reusable Terraform modules, Helm charts, GitOps templates, and shared platform services.
• Build self-service platform functionalities and established pathways that expedite software delivery.
• Design, standardize, and enhance enterprise CI/CD pipelines incorporating automated testing, security validation, deployment governance, and release automation.
• Develop highly available Kubernetes platforms with robust security, networking, scalability, and operational practices.
• Enhance platform observability through metrics, logs, traces, dashboards, alerting, SLIs, SLOs, and error budgets.
• Assist engineering teams during production incidents.
• Conduct post-incident reviews and implement lasting reliability improvements.
• Remove manual operational tasks through automation.
• Collaborate with engineering teams to boost application reliability, resilience, performance, and operational readiness.
• Improve cloud cost efficiency, platform effectiveness, and resource utilization.
• Over 5 years of experience in Platform Engineering, Site Reliability Engineering, or Cloud Infrastructure.
• Hands-on experience with Azure and/or AWS in production settings.
• Practical experience with Kubernetes in production (AKS, EKS, or GKE).
• Proficient in Infrastructure as Code using Terraform or a comparable tool.
• Familiarity with contemporary CI/CD platforms like GitHub Actions or Azure DevOps.
• Experience with GitOps methodologies utilizing ArgoCD or Flux.
• Knowledge in Helm chart development.
• Background in Linux systems administration and container technologies.
• Proficient in scripting using PowerShell, Bash, Python, or Go.
• Experience with enterprise observability platforms such as Datadog, Prometheus, Grafana, LogicMonitor, or New Relic.
• Understanding of cloud networking and identity technologies, including OIDC.
• Strong troubleshooting skills, excellent communication, and effective cross-functional collaboration.
• Proven experience collaborating across Software Engineering, Cloud Infrastructure, Architecture, Security, Data Engineering, and MLOps teams.
• Access to experts and resources for your Learning & Development journey.
• Opportunity for internal mobility.
• Employee referral bonus program.
• Employee Resource Groups (ERGs).
• Annual fundraising and volunteer events to contribute to communities.
• Paid time off, floating holidays, time off for volunteering, and rollover options.
• Paid parental leave.
• Medical, dental, vision, and 401k plans (with matching contributions).
• Flexible spending accounts, mass transit, and dependent care plans available.
• Health savings account with an annual company contribution for participants.
• Short-term and long-term disability coverage.
• Company-subsidized life insurance policies.
• Pet insurance options.
• Accident care services.
• Access to legal advice.
FCamara Consulting & Training
Greenbox Capital
NVIDIA
Abnormal Security
Get handpicked remote jobs straight to your inbox weekly.