
Mid-Level SRE Analyst
Posted 17 hours ago

Posted 17 hours ago
This is a fully remote position, open to applicants in Texas.
• Design and manage highly available, scalable, and secure cloud platforms on AWS.
• Build and sustain Kubernetes-based infrastructure to support applications, data, and AI workloads.
• Enhance platform reliability through automation, Infrastructure as Code (IaC), and self-service capabilities.
• Implement and improve observability solutions using Datadog, covering monitoring, logging, tracing, alerting, dashboards, and SLO management.
• Support and optimize expansive data processing environments utilizing Airflow, Amazon EMR, S3, and other AWS data services.
• Collaborate with Data and AI teams to bolster the reliability, scalability, and operational maturity of Machine Learning and Artificial Intelligence platforms.
• Lead incident response efforts, perform root cause analyses, and conduct post-incident reviews.
• Define and track SLIs, SLOs, and error budgets.
• Enhance deployment processes, CI/CD pipelines, and release reliability.
• Optimize cloud infrastructure for utilization, performance, and cost-efficiency.
• Mentor team members and advocate for SRE best practices throughout the engineering organization.
• Increase platform availability and reliability.
• Enhance observability and decrease incident resolution time.
• Boost automation and minimize repetitive manual operational tasks.
• Deliver dependable, scalable, and cost-effective Data and AI platforms.
• Collaborate with Engineering teams to create resilient production systems.
• Currently pursuing or have attained a bachelor's degree.
• Extensive experience in Site Reliability Engineering, Platform Engineering, Cloud Engineering, or DevOps roles.
• Strong hands-on expertise with AWS services and cloud-native architectures.
• In-depth knowledge of Kubernetes and containerized workloads in production settings.
• Experience in managing and troubleshooting large-scale distributed systems.
• Solid experience with observability platforms, preferably Datadog.
• Experience supporting data platforms and pipelines utilizing technologies such as Airflow, EMR, Spark, and S3.
• Proficient in Infrastructure as Code (IaC) using Terraform or similar tools.
• Experience in building and maintaining CI/CD pipelines and automating platforms.
• Strong understanding of Linux, networking, and system performance troubleshooting.
• Proficiency in scripting and automation with Python, Bash, or similar languages.
• Experience supporting large-scale cloud-native platforms in AWS environments.
• Experience operating Kubernetes platforms and managing cluster lifecycles.
• Knowledge of Site Reliability Engineering principles, including SLOs, SLIs, error budgets, and operational excellence practices.
• Familiarity with implementing observability solutions using tools like Datadog, Prometheus, Grafana, OpenTelemetry, or similar technologies.
• Experience with data processing and workflow orchestration platforms such as Airflow, Spark, or EMR.
• Experience in Infrastructure as Code (IaC) and platform automation methodologies.
• AWS, Kubernetes, Terraform, or Datadog certifications.
• Experience in large-scale, highly available, or mission-critical enterprise environments.
• Intermediate proficiency in technical English.
• Remote work opportunities.
• Full-time Employee Status: Regular.
• Affirmative action position for women.
• Development opportunities focused on gender equity and the Women in Experian group.
B2Spin Limited
PATH
Miratech
Endava
Get handpicked remote jobs straight to your inbox weekly.