
Senior Software Engineer – Site Reliability
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in United States.
• Take charge of intricate production support escalations and ticket triage, delivering hands-on troubleshooting and resolutions.
• Collaborate with Software Engineering to explore production issues, pinpoint root causes and reliability risks, and propel permanent solutions.
• Implement and enhance SRE practices, focusing on proactive reliability engineering, continuous improvement, automation, and shared ownership.
• Direct incident response from triage and mitigation through to recovery, root cause analysis, and blameless post-incident reviews.
• Construct observability by utilizing metrics, logs, traces, dashboards, and actionable alerts.
• Establish and refine SLIs and SLOs that gauge system reliability and customer experience.
• Create synthetic monitoring for essential customer journeys.
• Recognize operational toil and advocate for automation, tooling, process enhancements, or enduring solutions.
• Utilize AI-supported tools and source code repositories for triage, troubleshooting, code analysis, automation, and technical investigations.
• Engage in a rotating on-call schedule, primarily during business hours, with limited after-hours and weekend support.
• Practical Site Reliability Engineering experience applying software engineering principles to production reliability and assisting in establishing or enhancing SRE practices.
• In-depth understanding of SLIs, SLOs, error budgets, observability, automation, and toil reduction.
• Proficiency in building monitoring systems, dashboards, alerts, and telemetry using tools such as Honeycomb, New Relic, Grafana, CloudWatch, Kibana, or similar platforms.
• Experience in production incident management, root cause analysis, blameless post-incident reviews, and corrective-action follow-through.
• Strong programming and scripting capabilities.
• Solid SQL and relational database expertise; PostgreSQL experience is preferred.
• Excellent code literacy and debugging skills, including the ability to navigate unfamiliar codebases, comprehend application flow, review code and change histories, and identify reliability challenges.
• Experience in troubleshooting cloud-hosted applications using source code, logs, APIs, telemetry, event streams, and databases.
• Comfort in navigating application stacks across PHP, .NET, and Node.js; deep expertise in each is not mandatory.
• Proven experience utilizing AI-assisted tools in engineering workflows.
• A collaborative self-starter capable of addressing challenging problems, adapting to shifting priorities, and questioning the status quo.
• Strong communication and collaboration abilities across Software Engineering, Product, Support, DevOps, and other technical teams.
• Eligibility to work within the U.S. and select Canadian Provinces.
• No visa sponsorship is available.
• Options for health, vision, and dental insurance.
• Access to HealthiestYou healthcare service with confidential doctor consultations available 24/7.
• 20 PTO days.
• 3 flexible days.
• 4 optional volunteer days.
• 12 paid holidays.
• Paid parental leave.
• 401k matching.
• Equipment delivered to your door.
• Potential eligibility for a discretionary bonus.
• Fully remote work available within the U.S. and select Canadian Provinces.
TEKsystems
Arctiq
GE Vernova
Ecosistemas
Get handpicked remote jobs straight to your inbox weekly.