
Monitoring/SRE Engineer
Posted Aug 19

Posted Aug 19
This is a fully remote position, open to applicants in District of Columbia, +1 more state.
• Oversee enterprise infrastructure, application, and cloud dashboards, including tools such as Splunk, SolarWinds, Grafana, or CloudWatch.
• Detect and prioritize issues.
• Address system alerts and conduct initial troubleshooting.
• Escalate incidents in accordance with established runbooks and ServiceNow protocols.
• Contribute to the development and upkeep of dashboards, alert configurations, and automated notifications for infrastructure and application health.
• Aid in root cause analysis and documentation following production incidents.
• Assist with synthetic monitoring and validation tests for business disaster recovery and continuity planning.
• Keep precise incident tickets, monitoring runbooks, and knowledge base entries.
• Work alongside server, cloud, storage, and database engineers to ensure effective system instrumentation and monitoring.
• Take part in an on-call rotation for after-hours incident response.
• Facilitate enterprise monitoring, alerting, and incident response for a federal civilian customer's hybrid infrastructure environment.
• An Associate's or Bachelor's degree in Information Technology or a closely related field, or equivalent professional experience.
• At least 5 years of experience in IT operations, monitoring, or a similar support role.
• Familiarity with enterprise monitoring/observability tools like Splunk, SolarWinds, Grafana, Datadog, or similar tools.
• Basic understanding of cloud infrastructure concepts, preferably AWS.
• Strong attention to detail with the ability to adhere to incident response protocols.
• Knowledge of ITSM ticketing and escalation procedures.
• Capability to obtain Public Trust clearance.
• Basic scripting skills in Python, PowerShell, or Bash are advantageous.
• Ability to work cohesively in a dynamic environment.
• Exceptional communication skills, with the ability to articulate complex technical concepts to non-technical audiences.
• Experience in a federal government IT setting is preferred.
• CompTIA Network+, Splunk Core Certified User, or AWS Certified Cloud Practitioner certifications are desirable.
• Familiarity with ServiceNow or similar ITSM/on-call systems is preferred.
• Previous experience on a federal government IT support contract is a plus.
• Interest in advancing toward a Site Reliability Engineering career path.
• Health insurance.
• Dental insurance.
• Vision insurance.
• 401(k) plan with company matching.
• Flexible spending accounts.
• Paid holidays.
• Three weeks of paid time off.
• Competitive compensation.
knowmad mood
RealTime eClinical Solutions
Koniag Government Services
Koniag Government Services
Get handpicked remote jobs straight to your inbox weekly.