
Senior Site Reliability Engineer
Posted Jul 18

Posted Jul 18
This is a fully remote position, open to applicants in United States.
• Ensure that the production systems utilized by the Security Teams function efficiently, meet uptime targets, and are equipped with the latest content and features.
• Oversee system performance and capacity metrics.
• Actively suggest and implement improvements while automating repetitive tasks and reducing 'toil'.
• Design both new and existing systems to boost performance, reliability, and scalability.
• Develop, implement, and refine CI/CD pipelines.
• Support the Management, Development, Design, and Deployment of microservice and containerized applications.
• Enforce robust security measures in distributed systems and agents.
• Collaborate with engineers and developers to automate deployments and configurations across diverse platforms.
• Simplify the complexity of Observability implementation by creating scalable automation solutions.
• Spot and act on opportunities to enhance observability and processes.
• Standardize and develop alerts, notifications, and responses for monitoring tools.
• Work in tandem with application teams to incorporate Observability into daily operations.
• Participate in post-mortems, providing root cause analyses and implementing subsequent action items.
• Advocate for DevOps best practices within the team.
• Engage in and promote Agile/Scrum methodologies.
• Contribute to the hybrid cloud production containerization service offering.
• Create and enforce standards, policies, and procedures for automation and integrations.
• Bachelor’s Degree with 7 years of experience; Master’s Degree with 6 years of experience; PhD with 2 years of experience.
• Prioritize best practices for security as a necessity, rather than an afterthought.
• Proficient in Cloud Platform administration (AWS, GCP, Azure).
• Familiar with the pillars of Observability.
• Experience in high-scale environments and a solid understanding of distributed architectures.
• Knowledgeable in Agile / DevOps methodologies.
• Experience using CI/CD tools (Github Actions, Bamboo, Jenkins, Azure DevOps).
• Familiar with running Docker workloads utilizing orchestration tools (Kubernetes / Amazon ECS).
• Capable of working independently as well as collaboratively in a team for daily tasks.
• Eager to learn new concepts and processes swiftly and adjust to evolving environments.
• Comfortable working in and managing both Linux and Windows environments.
• Preferred: Experience with SPIRE/SPIFFE.
• Hands-on experience with Terraform/Crossplane.
• Proficient in using development tools and scripting languages (git / mercurial / subversion; Python / Elixir / Go).
• Integrate MCP Servers with authorization controls.
• Knowledgeable in database management systems (NoSQL, Relational Databases, and associated query languages).
• AWS Cloud Practitioner / Azure AZ-900 Certification.
• Extensive experience in implementing and designing serverless architecture solutions.
• Proven experience in deploying containerized applications (Kubernetes, etc.).
• Familiar with data management and pipeline technologies (Apache Storm, Kafka, Flink, Spark, Hadoop, etc.).
• Previous experience working in an Agile team.
• Strong understanding of observability solutions utilizing OpenTelemetry, Prometheus/Grafana or similar applications.
• Comprehensive understanding of distributed system architectures and telemetry.
• Significant experience in deploying and managing large Kubernetes Distributed Platforms.
• Skilled in GitOps practices and Infrastructure as Code systems (such as Terraform, ArgoCD, Helm).
• Health insurance.
• 401(k).
• Paid time off (vacation, holidays, sick leave).
• Participation in short-term incentive programs.
The Codest
IRIUM
Sólides
Resilinc
Get handpicked remote jobs straight to your inbox weekly.