Skip to content
← Back to job listings

Associate, Site Reliability Engineer (Platform Support), SRE and Governance, Group Technology

DBS Bank Ltd · East, Singapore

Software DevelopmentImported listingfull-time1 day ago

About The Role

Responsibilities

  • Administrate the full Kubernetes platform life cycle to ensure the platform remains secure, reliable and highly available.
  • Work with other engineers to support the infrastructure of OpenShift as well as assist to resolve infrastructure queries/issues of OpenShift tenant applications.
  • Automation of platform operations and management.
  • Configure and Manage Monitoring, alerts for the OpenShift environments.
  • Manage capacity for Kubernetes environments.
  • Hands-on experience with Docker containerization, Kubernetes orchestration and tools such as Jira and Jenkins.
  • Hands-on experience in handling Linux OS such as Patching, Filesystem management, User management, SELINUX etc.
  • Administration and support Windows Servers and Linux server environments.
  • Perform server hardening, configuration, patching, upgrades, and maintenance.
  • Troubleshoot IIS, SSL certificate, OpenSSH, tectia and IBM CD file transfer
  • Monitor server health, CPU, memory, disk, services, and system availability.
  • Troubleshoot OS, application, network-connectivity, and performance issues.
  • Handle Linux services, packages, file permissions, SSH, cron jobs, and system logs.
  • Respond to incidents, alerts, service requests, and production changes.
  • Perform vulnerability remediation and OS hardening.
  • Maintain operational documentation, SOPs, and troubleshooting guides.

Requirements

  • 3 years of work experience with a bachelor’s degree in computer science or related

Preferred Qualifications. At least 1 year’s hands-on experience with containers in Production Environments - Docker, OpenShift, Kubernetes, Linux and Window Servers preferred

  • Provide operational support for OpenShift Container Platform (OCP), Red Hat Linux, Windows Server, and infrastructure platforms across production and non-production environments.
  • Deliver 24x7 operational support for infrastructure and platform services, ensuring timely incident resolution and service restoration.
  • Perform system administration, monitoring, troubleshooting, performance tuning, and capacity management for Linux, Windows, and containerized environments.
  • Participate in incident management, problem management, root cause analysis (RCA), and post-incident reviews to improve operational stability.
  • Maintain operational documentation, standard operating procedures (SOPs), knowledge articles, and support runbooks.
  • Support security and compliance requirements by executing vulnerability remediation, patch management, access control, and platform hardening activities.
  • Experience with configuration management tools (Chef, Ansible, terraform etc.).
  • Hardening, securing the Kubernetes cluster with monitoring and auditing dashboards
  • Knowledge in infrastructure technologies such as HP and DELL hardware (Blades and Rack servers)
  • Excellent verbal, written, skills; in particular, demonstrated ability to effectively communicate technical and business issues and solutions to multiple organizational levels internally and externally.
  • Candidate must have demonstrated and be prepared to exhibit initiative and ownership of consistent delivery success
  • Be scheduled On-Call to support the infrastructure and systems

Location

DBS Asia Hub

Job

Technology

Schedule

Regular

Employee Status

Full time

This is an external listing. JobSpring does not represent or verify the employer. Report this listing