Talent.com
Sun International
Site Reliability EngineerSun International • Johannesburg, ZA
Search for other jobs
Site Reliability Engineer

Site Reliability Engineer

Sun International • Johannesburg, ZA
30+ days ago
Job description

Job Description

  • The Site Reliability Engineer (SRE) is a technical contributor responsible for improving the reliability, resilience and operational performance of Sun International’s technology services, platforms and infrastructure.
  • The role focuses on observability, monitoring, alerting, event management, reliability analytics, automation and auto-remediation, proactively identifying potential issues before they impact the business.
  • Working across software engineering, infrastructure, enterprise applications and service management teams, the SRE implements monitoring and automation solutions that strengthen service reliability and reduce operational.

Core behavioural & Technical / proficiency competencies:

  • Cloud platform management across Azure, AWS and/or GCP.
  • Containerisation technologies, including Docker and Kubernetes.
  • Monitoring, observability and alerting tools such as Prometheus, Grafana and ELK.
  • Scripting and operational automation using Python, Bash and/or PowerShell.
  • Development of automation, auto-remediation capabilities and operational runbooks.
  • Incident management and proactive identification of service reliability risks.
  • Root Cause Analysis (RCA) and analysis of recurring operational failures.
  • Linux and Windows system administration.
  • Networking fundamentals.
  • Git version control.
  • Troubleshooting and problem-solving across technology services and infrastructure.
  • Analysis of operational telemetry, event data and service performance trends.
  • Application of reliability standards, security requirements, operational controls and governance practices.
  • Cross-functional collaboration with software engineering, infrastructure, enterprise applications and service management teams.
  • Analytical thinking and evidence-based decision-making.
  • Continuous improvement and operational excellence.
  • Collaboration and knowledge sharing.
  • Operational excellence and accountability.

Job Requirements

Qualifications:

  • Degree in Computer Science, Engineering, Information Technology or a related discipline (required)
  • Cloud platform certification, e.g. Azure Administrator Associate or AWS Certified SysOps Administrator (preferred)
  • ITIL Foundation Certification (preferred)

Experience:

  • 2–5 years’ experience in infrastructure operations, technology operations, monitoring platforms, cloud operations, Site Reliability Engineering, observability tooling or a related technology discipline.
  • Practical experience with monitoring and observability, cloud platforms, automation/scripting, incident management and troubleshooting aligned to an SRE or technology operations environment
Create a job alert for this search

Site Reliability Engineer • Johannesburg, ZA