
高级运维工程师, Senior DevOps Engineer
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in Singapore.
• Ensure the stability, security, and availability of production environments, addressing deployment releases, system changes, performance issues, and complex failures.
• Participate in emergency responses for significant production incidents and drive problems from immediate resolution to long-term governance.
• Oversee the construction and continuous optimization of AWS cloud infrastructure and Kubernetes platforms, participating in infrastructure planning, architectural design, and technical architecture design for business systems.
• Promote automation, standardization, and engineering in operations, enhancing IaC, CI/CD, automated operations, release verification, and rollback mechanisms.
• Establish mechanisms for monitoring, logging, alerting, emergency response, and failure review to improve the efficiency of problem discovery, identification, and recovery.
• Advocate for developers to resolve long-term issues at the application and architecture levels.
• Manage the high availability, capacity planning, backup recovery, and failure handling of databases and core infrastructure components.
• Drive security governance for infrastructure, including access control, security baselines, network and container security, data and key security, vulnerability management, security audits, and incident response.
• Optimize cloud resource architecture, usage efficiency, capacity planning, and cost management.
• Stay informed about developments in cloud-native, automation, and security technologies, and promote technological improvements in alignment with business needs.
• Bachelor's degree or higher in Computer Science or a related field.
• Over 5 years of experience in DevOps, SRE, or related operations, with practical experience in production environments.
• Ability to design architectures and execute frontline operations, capable of completing infrastructure solution design, implementation, and complex failure resolution.
• Solid foundation in Linux, networking, and container technologies.
• Hands-on experience with production-grade Kubernetes, able to independently deploy, upgrade, optimize, and troubleshoot complex issues in clusters.
• Familiarity with the runtime environments and basic troubleshooting methods for common applications such as Java and Go.
• Practical experience in AWS production environments, knowledgeable about core services like IAM, VPC, EC2, and EKS.
• Experience with IaC tools such as Terraform/Terragrunt.
• Proficient in CI/CD and automation tools like GitHub Actions and Ansible.
• Skilled in scripting languages such as Bash and Python.
• Experience in governance of production-level observability and stability, familiar with tools like Prometheus, Grafana, ELK/Loki, as well as monitoring alerts, emergency responses, capacity, and disaster recovery mechanisms.
• Knowledge of databases and foundational components such as MySQL, Redis, and Kafka/RocketMQ, with an understanding of high availability, performance, backup recovery, and failure handling.
• Basic SQL writing skills and capability in troubleshooting database issues.
• Experience in security operations, familiar with cloud platform permissions and network security, host and container security, data and key security, vulnerability management, security audits, and incident response.
• Possess a sense of responsibility, risk awareness, and the ability to collaborate across teams.
• Comprehensive health insurance and wellness programs.
• Opportunities for professional development and continuous learning.
• Flexible working hours and the possibility of remote work.
• A collaborative and inclusive work environment.
• Competitive salary and performance-based bonuses.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.