Overview
Site Reliability Engineer – Kuala Lumpur, Federal Territory of Kuala Lumpur, Malaysia Company : VLink Inc, Federal Territory of Kuala Lumpur, Malaysia Responsibilities
Monitor and maintain system performance to ensure the stability and reliability of applications and infrastructure. Design and implement resilient system architectures that support high availability and scalability. Develop automation tools and scripts to enhance operational efficiency and reduce manual effort. Define, track, and analyze SLOs and SLIs to ensure reliability and performance meet business needs. Conduct post-mortem analyses following incidents, driving continuous improvement through root-cause identification and solution implementation. Collaborate with development and operations teams to establish best practices in system reliability and incident management. Troubleshoot and resolve issues related to database performance, network connectivity, and deployment failures, including diagnosing problems at the underlying platform level (e.g., Kubernetes, virtual machines). Ensure issues are resolved within SLAs, maintaining high standards of service delivery. Identify and troubleshoot performance bottlenecks in applications and infrastructure, providing actionable recommendations for enhancements. Maintain detailed documentation of processes and incident responses to support knowledge sharing and compliance. Improve monitoring solutions to proactively identify and mitigate issues before they impact services. Assist in the deployment and configuration of new applications and services, ensuring adherence to best practices. Participate in on-call rotations and respond to critical incidents as they arise. Analyze system logs and metrics to identify trends and potential areas for improvement. Familiarity with monitoring tools and performance optimization techniques. Familiarity with DevOps practices and frameworks, including CI / CD, infrastructure as code, and containerization. Qualifications / Minimum Qualifications and Experience
Able to communicate effectively in Mandarin as the role will require communicating with Mandarin-speaking clients / stakeholders in locations where Mandarin is primarily used. Proficiency in programming languages such as Python, Go, Java, or similar, focusing on operational efficiency. Experience in Bash / Shell scripting or automation for system administration tasks. Demonstrated experience in system architecture and design, prioritizing reliability and scalability. Strong understanding of SRE principles, including SLOs, SLIs, toil reduction, and incident post-mortems. Hands-on experience with cloud environments (e.g., AWS, Azure, Google Cloud) and their operational management. Strong expertise in Linux system administration. Proven experience in troubleshooting application support issues with a focus on performance and connectivity. Familiarity with networking concepts and effective troubleshooting techniques. Excellent problem-solving abilities and a proactive approach to operational challenges. Ability to work independently while effectively collaborating within a team environment. Open to a rotational shift schedule across different time slots, with reasonable schedules shared in advance. Job Details
Seniority level : Mid-Senior level Employment type : Full-time Job function : Information Technology Industries : IT Services and IT Consulting Note : This description reflects the responsibilities and qualifications for a Site Reliability Engineer role as posted for VLink Inc in Kuala Lumpur. It may include related roles and is intended for recruitment purposes.
#J-18808-Ljbffr
Reliability Engineer • Kuala Lumpur, Malaysia