
DevOps/SRE Engineer, Bilingual Mandarin
Posted 6 hours ago

Posted 6 hours ago
This is a fully remote position, open to applicants in California, +4 more states.
• Carry out daily operations for cloud resources and infrastructure as part of the operations and monitoring framework established by the team in China.
• Conduct daily monitoring, alerting, backup, and recovery tasks in accordance with standard operating procedures.
• Support first-line alert triage and coordinate with the China-based team for incident response across time zones.
• Execute initial troubleshooting and first-response actions prior to transferring issues to the China-based team.
• Manage the daily operations of local US data centers and/or cloud resources while ensuring compliance with relevant requirements.
• Implement local regulations such as access controls for privacy data and data localization/storage.
• Act as the first responder for incidents impacting systems and services in North America.
• Independently address common incidents and collaborate with the China-based team on more complex issues.
• Oversee execution and liaison activities related to GDPR, SOC2, CCPA, EO 14117, US local laws, legal compliance, and third-party audit obligations.
• Proficient in Mandarin Chinese with skills in listening, speaking, reading, and writing.
• Authorized to work for any employer; sponsorship is not currently available.
• Bachelor's degree in Computer Science or a related discipline.
• 3–4 years of practical experience in DevOps, SRE, or Platform Engineering roles.
• Extensive experience with at least one major cloud service provider (AWS, Azure, or GCP), covering VPC, EC2, EKS/Kubernetes, RDS, and IAM.
• In-depth understanding of Linux systems, networking basics, containers (Docker, Kubernetes), load balancing, and service governance.
• Skilled in Infrastructure as Code tools such as Terraform, Ansible, and Helm.
• Experience in developing and maintaining CI/CD pipelines using Jenkins, Argo CD, CodeBuild, or similar tools.
• Practical experience with monitoring, logging, and tracing systems, including Prometheus, Grafana, ELK Stack, OpenTelemetry, or equivalent solutions.
• Proficient in at least one scripting or programming language such as Python, Shell, or Go.
• Strong system design abilities, analytical thinking, and advanced troubleshooting skills.
• Excellent cross-team communication skills; experience in technical knowledge sharing or evangelism is advantageous.
• 401(k)
• PTO
• Paid Holidays
• Insurance (Medical+Vision)
a37
GT
Sigma Software Group
Applaudo
Get handpicked remote jobs straight to your inbox weekly.