
Senior Monitoring/SRE Engineer
Posted Aug 21

Posted Aug 21
This is a fully remote position, open to applicants in District of Columbia, +1 more state.
• Design, implement, and oversee the enterprise monitoring and observability architecture for infrastructure, applications, and cloud services.
• Define and manage service level objectives, service level indicators, and error budgets for critical financial and case-management systems.
• Lead the response to incidents and conduct root cause analysis for significant outages.
• Facilitate cross-team remediation efforts and conduct after-action reviews.
• Develop automated alerting systems, dashboards, and runbooks to minimize mean time to detect and mean time to resolve.
• Assist in monitoring and validating the business disaster continuity and recovery program, which includes synthetic transaction monitoring and failover verification.
• Collaborate with server, cloud, storage, and database engineering teams to instrument systems and integrate telemetry into a cohesive observability platform.
• Guide junior SRE/monitoring engineers through mentorship.
• Set best practices for capacity planning and performance baselining.
• Aid in continuous monitoring reporting requirements under FISMA/NIST SP 800-53 in conjunction with the security team.
• Report operational health, reliability metrics, and improvement plans to program and government leadership.
• At least 8 years of experience in systems monitoring, site reliability engineering, or a similar operations field.
• A Bachelor's degree in Computer Science, Information Technology, or a related area, or equivalent professional experience.
• Practical experience in designing and managing enterprise monitoring/observability platforms, such as Splunk, SolarWinds, Grafana, Datadog, or similar tools.
• Extensive experience with cloud-native monitoring in AWS, including CloudWatch and CloudTrail or equivalents.
• Proven experience leading incident response and root cause analysis in enterprise production environments.
• Familiarity with AIOps, auto-remediation/self-healing workflows, or OpenTelemetry.
• Proficient in scripting/automation using Python, PowerShell, or Bash.
• Strong understanding of ITIL-aligned incident, problem, and availability management practices.
• Exceptional written and verbal communication skills, with experience briefing technical and program leadership.
• Ability to collaborate effectively in a fast-paced work environment.
• Capability to simplify complex technical concepts for non-technical stakeholders.
• Eligibility to obtain public trust clearance.
• Experience in a federal government IT environment is preferred.
• Desired certifications include Splunk Certified Architect/Admin, AWS Certified DevOps Engineer, or equivalent monitoring/SRE certifications.
• Experience with BDCR/COOP monitoring and DR failover validation for financial systems is preferred.
• Familiarity with alert-management/on-call platforms such as PagerDuty, Opsgenie, or similar is preferred.
• Understanding of FedRAMP continuous monitoring (ConMon) reporting requirements is desired.
• Health, dental, and vision insurance.
• 401K with company matching.
• Flexible spending accounts.
• Paid holidays.
• Three weeks of paid time off.
• Competitive compensation.
• An extraordinary benefits package.
Xenon Seven
Toolbox
Newxel
Tether.to
Get handpicked remote jobs straight to your inbox weekly.