Monitoring/SRE Engineer

Posted Aug 19

This is a fully remote position, open to applicants in District of Columbia, +1 more state.

📋 Description

• Oversee enterprise infrastructure, application, and cloud dashboards, including tools such as Splunk, SolarWinds, Grafana, or CloudWatch.

• Detect and prioritize issues.

• Address system alerts and conduct initial troubleshooting.

• Escalate incidents in accordance with established runbooks and ServiceNow protocols.

• Contribute to the development and upkeep of dashboards, alert configurations, and automated notifications for infrastructure and application health.

• Aid in root cause analysis and documentation following production incidents.

• Assist with synthetic monitoring and validation tests for business disaster recovery and continuity planning.

• Keep precise incident tickets, monitoring runbooks, and knowledge base entries.

• Work alongside server, cloud, storage, and database engineers to ensure effective system instrumentation and monitoring.

• Take part in an on-call rotation for after-hours incident response.

• Facilitate enterprise monitoring, alerting, and incident response for a federal civilian customer's hybrid infrastructure environment.


⛳️ Requirements

• An Associate's or Bachelor's degree in Information Technology or a closely related field, or equivalent professional experience.

• At least 5 years of experience in IT operations, monitoring, or a similar support role.

• Familiarity with enterprise monitoring/observability tools like Splunk, SolarWinds, Grafana, Datadog, or similar tools.

• Basic understanding of cloud infrastructure concepts, preferably AWS.

• Strong attention to detail with the ability to adhere to incident response protocols.

• Knowledge of ITSM ticketing and escalation procedures.

• Capability to obtain Public Trust clearance.

• Basic scripting skills in Python, PowerShell, or Bash are advantageous.

• Ability to work cohesively in a dynamic environment.

• Exceptional communication skills, with the ability to articulate complex technical concepts to non-technical audiences.

• Experience in a federal government IT setting is preferred.

• CompTIA Network+, Splunk Core Certified User, or AWS Certified Cloud Practitioner certifications are desirable.

• Familiarity with ServiceNow or similar ITSM/on-call systems is preferred.

• Previous experience on a federal government IT support contract is a plus.

• Interest in advancing toward a Site Reliability Engineering career path.


🏝️ Benefits

• Health insurance.

• Dental insurance.

• Vision insurance.

• 401(k) plan with company matching.

• Flexible spending accounts.

• Paid holidays.

• Three weeks of paid time off.

• Competitive compensation.

People also viewed

knowmad mood2 days ago

Consultor/a DevSecOps – AWS

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
RealTime eClinical Solutions2 days ago

Principal DevOps Architect

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$155k – $195k/year
ApplyView job
Koniag Government Services2 days ago

DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Koniag Government Services2 days ago

Senior AWS DevOps Engineer – AWS, Kubernetes, HCP, CI/CD, Observability, AI-focus

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ASRC Federal2 days ago

Senior DevOps Administrator – Supporting NASA

US flagCalifornia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Nelios2 days ago

DevOps Engineer, Cloud Infrastructure

GR flagGreece OnlyPart-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers