Site Reliability Engineer – Observability Platform

atFord Motor CompanyRemoteUS flagWashingtonFull-timeDevOps & Site Reliability Engineer (SRE)Mid-levelSenior$85.4k – $192.9k/year

Posted 19 hours ago

This is a fully remote position, open to applicants in Washington.

📋 Description

• Design and implement scalable observability pipelines that encompass metrics, logging, tracing, and alerting.

• Define and operationalize Service Level Indicators (SLIs) and Service Level Objectives (SLOs), establishing error budgets.

• Create reusable infrastructure-as-code templates and frameworks for observability instrumentation and onboarding.

• Architect, design, and develop automation solutions to enhance application resilience, recoverability, availability, and scalability.

• Conduct destructive testing to identify vulnerabilities.

• Develop tools to enhance reliability, quality, and time-to-market.

• Minimize or eliminate manual tasks through automation.

• Collaborate with development teams to create and maintain scalable, resilient, cloud-native systems.

• Identify stability risks and formulate mitigation strategies in coordination with engineering leadership.

• Analyze technical metrics, including errors, response times, caching, capacity, and resource utilization.

• Perform performance analysis and optimization on both new and existing systems.

• Address complex architectural, design, and business challenges by streamlining processes and eliminating bottlenecks.

• Assess and integrate emerging technologies and architectures.

• Troubleshoot distributed production systems and conduct root-cause analysis for platform incidents.

• Engage in incident response, support, recovery, and postmortem analysis.

• Offer technical guidance and mentorship.

• Incorporate AI/ML capabilities to improve anomaly detection, alerting accuracy, and performance insights.

• Embed observability best practices within system design and deployment workflows.


⛳️ Requirements

• Bachelor’s Degree in Computer Science or a related field, or equivalent experience.

• At least 3 years of experience in a Site Reliability Engineering (SRE) role.

• Over 5 years of programming experience in one or more languages such as Python, Go, Java/Scala, C, or C++.

• Minimum of 3 years of experience in creating reusable infrastructure-as-code templates and frameworks using Terraform or ToFu.

• At least 3 years of experience with APM and monitoring tools, including Dynatrace, New Relic, ELK, Splunk, Prometheus, Sensu, Nagios, Kafka, or DataDog.

• Over 3 years of experience with J2EE, NoSQL/SQL databases, Spring Boot, GCP/AWS/Azure, and Docker/Kubernetes for developing multi-tier applications.

• Experience in working with RESTful APIs and microservices architectures.

• Working knowledge of the TCP/IP stack, internet routing, and load balancing techniques.

• Strong expertise in Google Cloud Platform and its suite of services.

• Familiarity with automated, test-driven development within CI/CD pipelines.

• Deep understanding of software development processes and agile methodologies.

• Ability to implement effective observability strategies to enhance Mean Time To Detection (MTTD) and Mean Time To Recovery (MTTR).

• Must possess legal authorization to work in the United States.

• Visa sponsorship is not offered for this position.


🏝️ Benefits

• Immediate access to medical, dental, vision, and prescription drug coverage.

• Flexible family care days.

• Paid parental leave.

• Programs to assist new parents in ramping up.

• Subsidized backup childcare options.

• Family building benefits, including reimbursement for adoption and surrogacy expenses as well as fertility treatments.

• Vehicle discount program available for employees and their family members, along with management leases.

• Tuition assistance programs.

• Active and established employee resource groups.

• Paid time off for both individual and team community service activities.

• Generous holiday schedule, including paid time off during the week between Christmas and New Year’s Day.

• Paid time off alongside the option to purchase additional vacation days.

People also viewed

Harris Computer17 hours ago

Platform & DevSecOps Delivery Architect

US flagAlabama, +20 more statesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TheWhiteam17 hours ago

DevOps Engineer

ES flagSpain OnlyFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
CFactory-Creations18 hours ago

Senior Site Reliability Engineer – Eastern Europe

BG flagBulgaria, +2 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Capgemini18 hours ago

Mainframe DevOps Migration Consultant

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$76.7k – $150.8k/year
ApplyView job
URBN (Urban Outfitters, Anthropologie Group, Free People & Nuuly)18 hours ago

Senior DevOps Engineer

US flagPennsylvania OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Nitrado19 hours ago

Site Reliability Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers