Site Reliability Engineer (Alicloud)

Shanghai, China

Overview 

We are seeking a Site Reliability Engineer (SRE) to ensure the stability, availability, and performance of mission-critical applications and cloud infrastructure. The ideal candidate has strong expertise in incident management, system reliability, cloud operations, and automation, with hands-on experience supporting production environments on Alibaba Cloud.

Key Responsibilities

  • Monitor, analyze, troubleshoot, and resolve system instability, application crashes, and production incidents to ensure high system availability and reliability.
  • Perform root cause analysis (RCA), implement permanent fixes, and drive continuous improvements to reduce recurring issues and improve platform resilience.
  • Build and maintain monitoring, alerting, logging, and observability solutions to proactively identify and address system health issues.
  • Collaborate with software engineering, DevOps, infrastructure, and platform teams to improve application performance, scalability, and operational efficiency.
  • Automate operational processes, deployments, and infrastructure management while promoting SRE best practices, reliability engineering, and operational excellence.

General Qualifications

  • Bachelor's Degree in Computer Science, Information Technology, Software Engineering, Engineering, or a related discipline.
  • Minimum 7 years of experience in Site Reliability Engineering (SRE), DevOps, Cloud Operations, Infrastructure Engineering, or Production Support.
  • Experience supporting high-availability, cloud-native, or large-scale enterprise production environments.
  • Experience working within Agile or DevOps environments with cross-functional engineering teams.

Mandatory Skills

  • Hands-on experience with Alibaba Cloud services, including cloud infrastructure, networking, compute, storage, monitoring, and security.
  • Strong knowledge of Linux systems administration, incident management, root cause analysis, performance tuning, system monitoring, and troubleshooting production issues.
  • Experience with automation, scripting (Shell, Python, or Bash), CI/CD pipelines, container technologies (Docker, Kubernetes), and observability tools for monitoring and logging.

Nice-to-Have Skills

  • Experience with Infrastructure-as-Code (Terraform, Ansible, or similar automation tools).
  • Familiarity with distributed systems, microservices, event-driven architectures, and cloud-native application deployments.
  • Alibaba Cloud certifications or other cloud platform certifications, with experience supporting 24/7 production operations and Service Level Objectives (SLOs).

For direct applications / questions, kindly send an email with your updated CV/resume to queenie.antioquia@geco.asia.

Site Reliability Engineer (Alicloud)

Job description

Site Reliability Engineer (Alicloud)

Personal information
Details