
Senior Site Reliability Engineer
Posted Aug 5

Posted Aug 5
This is a fully remote position, open to applicants in United States.
• Assist in the transition of products from on-premises data centers to AWS Cloud.
• Develop and execute cloud best practices for products that have been newly transitioned.
• Maintain operational oversight of cloud environments while optimizing, reengineering, and enhancing efficiency.
• Ensure that SaaS products operating in the cloud are stable and dependable.
• Improve automation, scaling, process efficiencies, metric collection, security, and environmental visibility.
• Apply DevOps methodologies, including CI/CD processes.
• Collaborate closely with development teams to facilitate secure and manageable changes in production.
• Support enterprise PaaS and SaaS offerings developed by Platform teams.
• Act as the primary contact for products and the underlying AWS infrastructure.
• Work within Agile methodologies utilizing Scrum and Kanban boards.
• Empower development teams with knowledge of cloud best practices and the AWS Well-Architected Framework.
• Design and implement scalable and reliable solutions in collaboration with cross-functional teams.
• Enhance end-to-end observability for cloud applications.
• Assist with on-premises applications and their migration to AWS Cloud.
• Lead strategic initiatives for seamless application integration and optimal performance in cloud-native environments.
• Perform additional duties as assigned by the manager.
• Bachelor’s degree or an equivalent combination of education and experience.
• Over 5 years of relevant industry experience, including at least 1 year in an Associate-level role or an equivalent external position.
• Comprehensive DevOps experience, ideally comprising 70% Operations and 30% Development efforts.
• Significant experience with containerization, Kubernetes, EKS, and CI/CD pipelines utilizing GitOps methodologies such as ArgoCD and FluxCD.
• Proficient in using Git for version control; experience with Azure DevOps is preferred.
• Familiarity with monitoring tools, particularly CloudWatch alerts, troubleshooting issues, and serving as an escalation resource.
• Capability to configure, tune, and secure APIs/Microservices to ensure optimal performance and uptime.
• Proficient in using curl and wget for diagnostics, network performance analysis, and connectivity troubleshooting.
• In-depth understanding of HTTP concepts and protocols.
• Strong experience with Terraform and best practices for Infrastructure as Code.
• Advanced scripting and automation abilities in Python, Bash, and PowerShell.
• Expertise in AWS Transfer Family for large files, EC2 instances and AMIs, AWS Elastic Beanstalk, and Load Balancers.
• Proficiency in SQL and SQL administration, including database/schema/catalog configurations, user/logins, synonyms, and RDS performance troubleshooting.
• Understanding of database hygiene concepts, data integrity, and efficient data storage practices.
• Familiarity with service mesh technologies such as Linkerd or Istio is a plus.
• Required experience with Kubernetes and Docker.
• Strong skills in network management.
• Excellent written and verbal communication abilities.
• Experience with Azure DevOps, encompassing ticket management, CI/CD pipelines, release management, and Infrastructure as Code.
• Flexible work hours.
• PTO and paid holidays.
• Medical insurance.
• Dental insurance.
• Vision insurance.
• Life insurance.
• Disability insurance.
• 401K.
• Eligibility for discretionary bonuses in certain positions.
TEKsystems
TEKsystems
Level Data
Level Data
Get handpicked remote jobs straight to your inbox weekly.