
Senior Site Reliability Engineer
Posted Aug 27

Posted Aug 27
This is a fully remote position, open to applicants in New Jersey.
• Establish and uphold DevOps guardrails, standards, and best practices throughout engineering teams.
• Empower engineering teams to design, implement, and sustain CI/CD pipelines for rapid, secure, and repeatable deployments.
• Collaborate with application teams to co-develop and maintain Infrastructure as Code utilizing tools like Terraform.
• Set up and manage blue/green, canary, and rolling deployment strategies.
• Create and support self-service tooling for developers and internal platforms.
• Integrate reliability, security, and observability practices early in the software development lifecycle.
• Define, implement, and oversee SLOs, SLIs, and error budgets for essential services.
• Construct and maintain observability stacks that include logging, metrics, distributed tracing, and alerting.
• Direct incident response and management, which encompasses on-call rotations, root cause analysis, and blameless post-mortems.
• Conduct capacity planning and performance engineering.
• Recognize and eliminate toil through automation.
• Carry out reliability assessments and chaos engineering exercises.
• Manage and optimize cloud infrastructure to enhance reliability, cost, and performance.
• Collaborate with software engineering teams on system architecture, resiliency patterns, and fault tolerance.
• Provide support and maintenance for .NET-based services across both Windows and Linux environments.
• Achieve uptime and availability targets, minimize MTTD and MTTR, enhance DevOps guardrail adoption, reduce operational toil, and manage error budgets.
• Over 8 years of experience in SRE, DevOps, or Infrastructure Engineering roles.
• More than 5 years of experience with AWS (preferred); multi-cloud experience is an advantage.
• Strong expertise in Terraform.
• Extensive experience in designing, building, and maintaining CI/CD pipelines and automation workflows.
• Significant experience with Docker and container orchestration (ECS preferred).
• Proficient with tools such as New Relic, ELK Stack, CloudWatch, Prometheus, Datadog, or Grafana.
• Proficient in .NET, PowerShell, Python, Go, or Bash.
• Strong skills in Linux and Windows systems administration.
• Solid understanding of DNS, load balancing, CDNs, and network security.
• Excellent written and verbal communication abilities.
• Experience in implementing and managing service mesh technologies (e.g., Istio, Linkerd, AWS App Mesh).
• Familiarity with SRE frameworks as described in Google's SRE handbook.
• Experience with secrets management (e.g., HashiCorp Vault, AWS Secrets Manager).
• Understanding of compliance and regulatory requirements in financial services (SOC 2, PCI-DSS, etc.).
• Experience with chaos engineering tools (e.g., Gremlin, Litmus, AWS Fault Injection Simulator).
• Background in supporting .NET applications in production environments.
• Experience in financial industry/banking infrastructure is beneficial but not mandatory.
• Experience in crypto/blockchain infrastructure is a plus but not essential.
• Familiarity with GitOps workflows and patterns.
• Understanding of cost optimization and FinOps practices in cloud environments.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible work hours and remote working options.
• Opportunities for professional development and continuous learning.
• Collaborative and inclusive work environment.
Sprezzatura
Clinician Nexus
Lenovo
MRSOOL | مرسول
Get handpicked remote jobs straight to your inbox weekly.