Own the day-to-day operation, maintenance and support of designated enterprisescale systems, including standalone and Kubernetes-based environments.
Administer and support Kubernetes environments, including application deployment, configuration, monitoring, troubleshooting and operational maintenance.
Review vendors’ deliverables, including technical documents, source code, test evidence and deployment materials, to ensure compliance with contractual requirements, internal security standards and guidelines.
Coordinate with business users, vendors, platform and infrastructure teams to resolve incidents, deliver system changes and support production releases.
Implement and enhance Site Reliability Engineering (SRE) practices, including monitoring, alerting, logging and incident management, to ensure system availability, reliability and performance.
Investigate production incidents, perform root-cause analysis, and drive corrective and preventive actions.
Identify repetitive manual operational tasks and improve efficiency, consistency and reliability through scripting, standardisation and workflow automation.
Maintain system operational documentation, including runbooks, support procedures, incident reports and knowledge-base materials.
Requirements
Degree in Computer Science, Information Technology or a related discipline.
At least 3 years of IT experience, including 2 years of relevant experience in system development, application implementation and/or enterprise-scale system support.
Demonstrates hands-on experience in Kubernetes (K8s) operations, including application deployment, configuration management, troubleshooting, monitoring, and incident resolution. Administrative experience with Kubernetes is advantageous.
Familiarity with container technologies, such as Docker, and Kubernetes deployment/configuration artifacts.
Proven ability to coordinate cross-functional teams to drive delivery and issue resolution.
Familiar with Agile/Scrum methodologies and collaboration tools such as Jira and Confluence.
Solid understanding of modern application concepts, including but not limited to RESTful APIs, Single Sign-On (SSO) and microservices architecture.
Good scripting skills in Bash and Python.
Strong automation mindset, with the ability to identify and transform manual operational processes into standardised, automated workflows.
Hands-on experience with Ansible Playbooks for operational, deployment or configuration automation is highly preferred.
Experience with monitoring, alerting and observability tools is preferred.
Strong analytical, problem-solving and incident-management skills.
Excellent communication skills, with proficiency in written and spoken English, Cantonese and/or Mandarin