
Staff Engineer – Site Reliability Engineering
Posted Sep 18

Posted Sep 18
This is a fully remote position, open to applicants in Pennsylvania.
• Create and implement a scalable Site Reliability Engineering (SRE) ecosystem adhering to SRE and DevSecOps best practices.
• Develop reusable TypeScript scaffolding libraries tailored for cloud-native components.
• Build and improve solutions utilizing AWS, EKS, Kubernetes, and Infrastructure as Code methodologies.
• Promote automation, reliability, scalability, and operational excellence throughout microservices.
• Define and execute SRE and DevOps best practices across various applications.
• Collaborate with technology, product, operations, and functional teams on SRE initiatives.
• Analyze business requirements and assess their impact on applications and cloud systems.
• Establish standardized and automated onboarding pathways for applications onto the SRE platform.
• Assess emerging technologies and formulate strategies for cloud and SRE adoption.
• Implement CI/CD, observability, monitoring, and service mesh capabilities.
• Identify and resolve reliability, scalability, and operational challenges.
• Offer technical guidance and support to distributed engineering teams.
• Minimum of 5.5 years of total experience.
• Strong proficiency in TypeScript development, including coding and design patterns.
• Practical experience with AWS cloud services and cloud-native technologies.
• Familiarity with AWS CDK, Terraform, and Infrastructure as Code (IaC).
• Knowledge of Kubernetes, Docker, and Amazon EKS.
• Experience in SRE, DevOps, scalability, reliability, and cloud automation.
• Proficiency with CI/CD tools like Jenkins and Git.
• Understanding of observability and monitoring tools such as CloudWatch, Splunk, and Dynatrace.
• Knowledge of service mesh technologies, including Istio.
• Experience in developing reusable scaffolding libraries and cloud-native components.
• Capability to analyze application and infrastructure dependencies across microservices environments.
• Experience working with distributed teams across various time zones.
• Excellent communication, presentation, and stakeholder collaboration skills.
• Bachelor’s or master’s degree in computer science, Information Technology, or a related field.
• Employees have the flexibility to work remotely.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.