DevOps Engineer, Cloud Infra

Posted 1 day ago

This is a fully remote position, open to applicants in Singapore.

📋 Description

• Manage production incidents and carry out post-mortem analyses to enhance system stability.

• Design, implement, monitor, and troubleshoot Kafka and Redis clusters within production settings.

• Collaborate closely with development teams to facilitate smooth application and system deployments.

• Oversee and optimize AWS and Alicloud infrastructure for performance, cost-effectiveness, and reliability.

• Create DevOps platforms, including online load testing and change management systems.

• Utilize LLMs or AI frameworks such as OpenAI, Dify, Agno, and LangChain to boost infrastructure-operation automation.

• Establish intelligent alert triage, root-cause analysis, and chat-based operations.

• Incorporate AI-driven insights into operational workflows to enhance reliability, minimize noise, and assist engineering decision-making.


⛳️ Requirements

• Over 5 years of practical experience in Kafka and Redis operations within large-scale production environments, with the ability to collaborate with developers for code optimization.

• Proficient in at least one programming language such as Python, Go, or Java, along with SQL skills.

• Hands-on experience with Docker and Kubernetes.

• Extensive experience with CI/CD tools like GitHub Actions, Ansible, and Terraform.

• A minimum of 3 years of experience working with the AWS cloud platform.

• Familiarity with GCP, Azure, or Ali Cloud is an advantage.

• Exceptional problem-solving and troubleshooting abilities.

• Strong team collaboration mindset and capability to cultivate partnerships with other teams and business units.

• Practical experience in building or managing AIOps systems, including anomaly detection, alert correlation, automated healing, or root cause analysis.

• Knowledge of LLM-based DevOps automation, such as chat-based operations assistants or AI-driven observability workflows.

• Experience in using or integrating Dify, Agno, or LangChain into operational processes.


🏝️ Benefits

• Competitive salary and comprehensive company benefits.

• Work-from-home options (subject to the nature of the business team's work).

• Opportunities for career advancement and ongoing learning.

• Collaborate with top-tier talent in a user-focused global organization with a flat organizational structure.

• Innovative and results-oriented work environment.

• Equal opportunity employer.

People also viewed

In All Media23 hours ago

DevOps Engineer – Cloud

BR flagBrazil, +5 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity Group1 day ago

SRE Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Fingerprint1 day ago

Senior Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$152k – $205k/year
ApplyView job
Endava1 day ago

Senior DevOps Engineer, Terraform

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
CVS Health1 day ago

Staff DevSecOps Engineer, Health

US flagConnecticut, +3 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$130.3k – $260.6k/year
ApplyView job
GoFasti1 day ago

Senior DevOps Engineer

Latin AmericaFull-timeDevOps & Site Reliability Engineer (SRE)$5,000 – $6,000/month
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers