Senior Monitoring/SRE Engineer

Posted Aug 21

This is a fully remote position, open to applicants in District of Columbia, +1 more state.

📋 Description

• Design, implement, and oversee the enterprise monitoring and observability architecture for infrastructure, applications, and cloud services.

• Define and manage service level objectives, service level indicators, and error budgets for critical financial and case-management systems.

• Lead the response to incidents and conduct root cause analysis for significant outages.

• Facilitate cross-team remediation efforts and conduct after-action reviews.

• Develop automated alerting systems, dashboards, and runbooks to minimize mean time to detect and mean time to resolve.

• Assist in monitoring and validating the business disaster continuity and recovery program, which includes synthetic transaction monitoring and failover verification.

• Collaborate with server, cloud, storage, and database engineering teams to instrument systems and integrate telemetry into a cohesive observability platform.

• Guide junior SRE/monitoring engineers through mentorship.

• Set best practices for capacity planning and performance baselining.

• Aid in continuous monitoring reporting requirements under FISMA/NIST SP 800-53 in conjunction with the security team.

• Report operational health, reliability metrics, and improvement plans to program and government leadership.


⛳️ Requirements

• At least 8 years of experience in systems monitoring, site reliability engineering, or a similar operations field.

• A Bachelor's degree in Computer Science, Information Technology, or a related area, or equivalent professional experience.

• Practical experience in designing and managing enterprise monitoring/observability platforms, such as Splunk, SolarWinds, Grafana, Datadog, or similar tools.

• Extensive experience with cloud-native monitoring in AWS, including CloudWatch and CloudTrail or equivalents.

• Proven experience leading incident response and root cause analysis in enterprise production environments.

• Familiarity with AIOps, auto-remediation/self-healing workflows, or OpenTelemetry.

• Proficient in scripting/automation using Python, PowerShell, or Bash.

• Strong understanding of ITIL-aligned incident, problem, and availability management practices.

• Exceptional written and verbal communication skills, with experience briefing technical and program leadership.

• Ability to collaborate effectively in a fast-paced work environment.

• Capability to simplify complex technical concepts for non-technical stakeholders.

• Eligibility to obtain public trust clearance.

• Experience in a federal government IT environment is preferred.

• Desired certifications include Splunk Certified Architect/Admin, AWS Certified DevOps Engineer, or equivalent monitoring/SRE certifications.

• Experience with BDCR/COOP monitoring and DR failover validation for financial systems is preferred.

• Familiarity with alert-management/on-call platforms such as PagerDuty, Opsgenie, or similar is preferred.

• Understanding of FedRAMP continuous monitoring (ConMon) reporting requirements is desired.


🏝️ Benefits

• Health, dental, and vision insurance.

• 401K with company matching.

• Flexible spending accounts.

• Paid holidays.

• Three weeks of paid time off.

• Competitive compensation.

• An extraordinary benefits package.

People also viewed

Xenon Seven11 hours ago

Forward Deployment Engineer

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Toolbox12 hours ago

DevOps Engineer

UY flagUruguay OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Newxel12 hours ago

Senior DevOps Engineer

UA flagUkraine OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Tether.to13 hours ago

DevOps Engineer

GB flagUnited Kingdom OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
CI&T14 hours ago

DevOps

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Guild Mortgage16 hours ago

Manager, DevOps

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$144k – $210k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers