
Platform Site Reliability Engineer
Posted 9 hours ago

Posted 9 hours ago
This is a fully remote position, open to applicants in United States.
• Create, implement, and sustain scalable, secure, and highly available cloud infrastructure.
• Develop and oversee infrastructure-as-code solutions for consistent and dependable deployments.
• Enhance platform reliability, resilience, scalability, and performance.
• Collaborate with engineering teams on reliability and observability initiatives.
• Assist in disaster recovery, backup, and business continuity efforts.
• Design, implement, and maintain CI/CD pipelines.
• Automate operational tasks and refine deployment processes.
• Optimize development workflows and minimize operational overhead.
• Advance release management practices and deployment strategies.
• Advocate for DevOps best practices across the organization.
• Monitor production systems to identify opportunities for availability and performance enhancements.
• Engage in incident response, troubleshooting, root cause analysis, and post-incident evaluations.
• Develop and uphold monitoring, alerting, logging, and observability frameworks.
• Define and track SLOs, SLIs, and reliability metrics.
• Participate in on-call rotations and manage critical production incidents.
• Implement secure infrastructure protocols in collaboration with Security and Engineering teams.
• Aid in vulnerability remediation, patch management, access controls, and security monitoring.
• Contribute to compliance initiatives and infrastructure controls.
• Ensure compliance with security, privacy, and operational best practices.
• Work with Engineering, Product, QA, and Security teams collaboratively.
• Document infrastructure, operational processes, and technical standards.
• Enhance reliability, efficiency, developer experience, and operational excellence.
• Participate in technical discussions, architectural reviews, and long-term platform strategy.
• Foster consensus on standards and establish the organizational golden path.
• Over 5 years of experience in Site Reliability Engineering, DevOps, Cloud Infrastructure, Platform Engineering, or a similar role.
• Extensive experience in managing and supporting production cloud environments.
• Proficiency in building and maintaining CI/CD pipelines and deployment automation.
• Strong grasp of infrastructure-as-code principles and tools.
• Experience in troubleshooting production systems and resolving complex operational challenges.
• Knowledge of networking, security, scalability, and high-availability architectures.
• Strong scripting or automation experience with languages like Python, Bash, PowerShell, or similar.
• Experience collaborating closely with software engineering teams throughout the software development lifecycle.
• Excellent analytical, problem-solving, and communication skills.
• Ability to balance operational stability with delivery speed and business objectives.
• Experience supporting high-growth SaaS platforms.
• Familiarity with AWS, Azure, or Google Cloud Platform.
• Experience with Kubernetes, Docker, and containerized environments.
• Proficiency in Terraform, Pulumi, CloudFormation, or other infrastructure-as-code tools.
• Experience using GitHub Actions, GitLab CI/CD, Jenkins, CircleCI, or similar automation platforms.
• Familiarity with Datadog, New Relic, Grafana, Prometheus, Splunk, or similar observability tools.
• Experience in Database Operations, Database Scaling, Query Optimization, and BCDR.
• Experience implementing zero-downtime deployments.
• Understanding of security frameworks, compliance requirements, and operational governance practices.
• Experience building highly available, mission-critical applications and services.
• Experience supporting distributed and remote engineering teams.
• All applicants must be eligible to work for any US employer in the United States.
• Locality Media LLC cannot sponsor or transition sponsorship ownership of employment visas.
• Successful completion of a criminal background check is mandatory.
• Bonus
• Medical coverage
• Dental coverage
• Vision coverage
• FSA/HSA
• 401(k)
• Flexible PTO
• Fully remote workplace
• Technology stipend
• Opportunities for advancement
• Minimal travel expectations
• Reasonable accommodation during employment and interview processes
Fairsource
NetBox Labs
VELZI.AI LIMITED
CoDev
Get handpicked remote jobs straight to your inbox weekly.