Remotery

DevOps Engineer II – IT Infrastructure Systems

Posted Jul 27

This is a fully remote position, open to applicants in California.

📋 Description

• Take the lead on infrastructure-as-code development independently using Terraform alongside scripting languages such as Python and PowerShell to facilitate scalable and reliable deployments.

• Oversee Linux/Kubernetes cluster environments.

• Implement solutions in alignment with Change Management Processes.

• Assist development teams in formulating API integration strategies and standards.

• Safeguard systems against cybersecurity threats.

• Diagnose technical issues and create software updates and patches.

• Exhibit strong skills in Splunk for administration, query enhancement, alert management, and dashboard creation.

• Develop tools aimed at minimizing errors and enhancing customer experiences.

• Suggest ideas and solutions within the Infrastructure Department to alleviate workload through automation.

• Design, implement, and refine CI/CD pipelines to ensure quicker and more reliable software releases.

• Conduct root cause analysis independently and apply corrective measures.

• Create and write tests to explore infrastructure failures and scaling issues.

• Generate and maintain response playbooks across incident management and monitoring tools.

• Create automation to ensure consistency, eliminate repetitive tasks, and shorten response and repair times.

• Evaluate key operational metrics to pinpoint opportunities for enhancing availability.

• Establish effective monitoring and alerting systems while mitigating alert fatigue.

• Supervise container orchestration environments and refine deployment workflows to boost scalability, reliability, and operational efficiency.

• Design, construct, and manage containerized environments using Docker.

• Develop and maintain SLIs, SLOs, and error budgets.

• Create and optimize monitoring dashboards and alert systems to proactively identify and resolve application performance and uptime challenges.

• Implement code branching methodologies using GitHub functionalities.

• Demonstrate advanced proficiency in Terraform syntax and GitLab CI/CD configurations, pipelines, and jobs.

• Provision and configure metrics in Prometheus, Thanos, and Grafana, including creating and managing alerts.

• Apply cloud engineering standards, reusable modules, and platform design patterns in Microsoft Azure.

• Manage shared cloud platform services in line with Cloud Engineering defined architectures.

• Ensure infrastructure modifications adhere to reliability, security, and cost control measures established by Cloud Engineering.

• Maintain operational documentation and runbooks for cloud platform services.


⛳️ Requirements

• More than 4 years of experience as a DevOps Engineer in medium to large-scale environments.

• Proficient in Windows Server, Linux, and hybrid cloud deployments utilizing Microsoft Azure and VMWare.

• Skilled in Git/GitHub workflows, Terraform, Python, PowerShell, and container orchestration technologies (Tanzu, Docker, Kubernetes, OpenShift).

• Experienced with CI/CD tools (Jenkins, GitLab CI, Azure DevOps) and observability platforms (Datadog, Prometheus, Grafana, ThousandEyes).

• Knowledgeable in log management (ELK Stack) and database technologies (PostgreSQL, MySQL, NoSQL).

• Strong experience in automating infrastructure provisioning and application deployment with Terraform, Ansible, and Kubernetes.

• Proficient in developing and maintaining monitoring dashboards, SLIs, SLOs, and error budgets to ensure application uptime and performance.

• Experienced in ensuring infrastructure security, promoting automation initiatives, and collaborating across teams to enhance reliability and scalability.

• Proficient in building observability pipelines and executing advanced queries in log management tools like Splunk for troubleshooting purposes.

• Experience with implementing and managing Azure-based shared services as defined by platform or cloud engineering teams.

• Microsoft Azure DevOps Engineer Expert Certification (Required).

• Kubernetes Administration Certification (Required).

• Linux Certification (Desired).


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health, dental, and vision insurance.

• Flexible working hours and remote work options.

• Opportunities for professional development and certifications.

• Collaborate with a dynamic and innovative team.

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers