Remotery

Data Site Reliability Engineer, SRE

Posted 15 hours ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Deliver technical guidance for daily operational support, ensuring reliability, performance, and ongoing enhancement of CMM data platforms, pipelines, applications, and analytics services.

• Facilitate real-time monitoring, incident and event management, capacity planning, and operational reporting.

• Oversee and audit cloud user roles and responsibilities.

• Implement SSO, MFA, and group identity management through JENIE to uphold least-privilege access.

• Evaluate and enhance credential management processes.

• Provide disaster recovery and continuity of operations strategies, including fault-tolerant and automated failover designs.

• Incorporate DevSecOps tools and processes with enterprise systems.

• Manage centralized secrets management, including automated rotation, access logging, and policy enforcement.

• Integrate SAST, DAST, SCA, and CSPM security tools into pipelines.

• Establish continuous 24/7/365 monitoring for security, performance, and compliance through dashboards and alerting.

• Offer additional monitoring of event response activities outside regular business hours.

• Automate the creation and management of SBOMs for deployed artifacts.

• Provide diagnostics, metrics collection, and performance optimization.

• Facilitate canary releases for end-user and beta testing.

• Set up alerts for unusual behavior.

• Conduct incident and event management processes integrated with enterprise SIEM solutions.

• Detect, log, diagnose, escalate, and resolve incidents while performing root cause analysis.

• Identify and eliminate recurring incident root causes.

• Suggest enhancements to incident and problem management processes.

• Maintain a repository of known issues, resolutions, and best practices.

• Execute automated full-stack health checks across operating systems, applications, databases, and PaaS services.

• Generate monthly issues management reports.

• Develop thresholds, rules, and response procedures.

• Monitor cloud resource utilization and manage threshold-breach resolution.

• Enhance reliability, observability, automation, scalability, and operational resilience.

• Supervise, maintain, and optimize cloud infrastructure, databases, and platform services.

• Act as a FinOps Analyst to drive cost optimization.

• Support the CMM program, a cloud-based solution for the Administrative Office of the US Courts, serving over 204 federal courts.


⛳️ Requirements

• Bachelor's degree in Computer Science, Software Engineering, or a related field, or equivalent experience.

• Over 5 years of experience in IT systems engineering, systems development, systems coding, and programming.

• Must qualify as a US Person: Green Card Holder, US Permanent Resident Alien, Refugee, Asylee, or US Citizen.

• Ability to pass a background check for a position of Public Trust.

• Extensive knowledge of AWS services, including monitoring, logging, compute, storage, and networking.

• Proficient in Infrastructure as Code tools such as Terraform, AWS CloudFormation, or Azure Bicep.

• Hands-on experience with monitoring and APM tools like CloudWatch, Azure Monitor, Datadog, Prometheus, Grafana, or New Relic.

• Understanding of incident response, change management, and ITIL-based operational support.

• Familiarity with CI/CD toolchains and automation platforms such as Jenkins, GitHub Actions, GitLab, and ArgoCD.

• Strong scripting abilities in Python, PowerShell, and Bash.

• Advanced experience in DevSecOps implementation using GitOps or similar tools.

• Experience in developing, testing, and maintaining containerized applications.

• Expert knowledge of source version control, build/release tools, CI/CD pipelines, and software build procedures.

• Experience in building and maintaining CI/CD pipelines for large enterprises with complex applications.

• Familiarity with FinOps practices, cost modeling, forecasting, and cloud optimization tools.

• Understanding of federal compliance and security frameworks such as FedRAMP, NIST, and JISF Rev 5.

• Ability to analyze logs and metrics, conducting performance tuning for cloud services and applications.

• Experience collaborating across multiple product teams to evaluate overall product/program health.

• ITIL, AWS SysOps, or Google Professional Cloud DevOps Engineer certifications are advantageous.

• Excellent presentation and communication abilities.

• Consultant mindset with the capability to engage with high-level customer stakeholders.

• Strong analytical and problem-solving skills.

• Experience with process design and documentation methodologies, deliverables, process/use case modeling, and business case development.

• Ability to work effectively both independently and as part of a team.

• Flexibility to work across multiple products while supporting various teams.


🏝️ Benefits

• Comprehensive medical plan options, some including Health Savings Accounts.

• Dental plan options available.

• Vision plan options offered.

• 401(k) plan with pre-tax and post-tax contributions and company match.

• Full-flex work weeks where feasible.

• Paid time off, encompassing vacation, sick leave, personal time, holidays, paid parental leave, military leave, bereavement leave, and jury duty leave.

• Typically 15 days of paid leave per calendar year.

• 10 paid holidays each year.

• Up to 160 hours of paid family leave within a rolling 12-month period for eligible employees.

• Short- and long-term disability benefits provided.

• Life insurance coverage.

• Accidental death and dismemberment insurance included.

• Personal accident insurance offered.

• Critical illness insurance available.

• Business travel and accident insurance included.

• AI-powered career tool that identifies career steps and learning opportunities.

• Internal mobility team to assist with career advancement.

• Wellness packages available.

• Competitive salary offered.

• An award-winning culture of innovation and a military-friendly workplace.

People also viewed

4Pharma Ltd8 hours ago

Lead DevOps Engineer – AI-Native

GB flagUnited Kingdom OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
4Pharma Ltd9 hours ago

Senior DevOps Engineer, AI-Native

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Virta Health9 hours ago

DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$179.5k – $187.9k/year
ApplyView job
Verity Group9 hours ago

SRE Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Segware9 hours ago

Senior SRE / DevOps

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
GFT Technologies12 hours ago

DevOps Specialist

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers