On-site
Deloitte -
Egypt , Egypt
--
Deloitte

Job Details

Connect to your opportunity Step into a highly impactful role where you will help us build and operate our brand-new centralized command center for public finance in Cairo. As a Site Reliability Engineer (SRE) you will ensure our enterprise infrastructure remains highly available, scalable, and fully optimized. You will go beyond basic troubleshooting to proactively manage and resolve complex technical challenges across diverse cloud environments. You will compose our support level 2 and 3 team for a national scale product that has been deployed in 3 GCC counties, the knowledge of the product will be transferred in training sessions.

Monitor and proactively troubleshooting to prevent outage and service interruption

Investigate and resolve infrastructure and application incidents coming from support level 1 or from the observability tools.

Root cause analysis and problem management, conduct post incident reviews and update runbooks and knowledge base articles

Performance and capacity management, monitor cloud utilization and optimize multi-cloud environments to ensure continuous high system reliability and uptime.

Automation and efficiencies, develop scripts to generate effort reductions and standardization, maintain infrastructure as a code templates and repeatable procedures.

Security and compliance, assist with client s audits and incidents responses

Back up and recovery, validate and monitor back up jobs, test recovery procedures

Document and knowledge sharing, maintain accurate documents for configuration, train support level 1 to increase first call resolution

Customer communication and SLAs/KPIs tracking and reporting

Vendor and tool liaison, coordinate with cloud providers and track open tickets to ensure timely resolution

On call shift responsibilities, participate in on call rotations

Desired Candidate Profile

Connect to your skills and professional experience

  • Analytical Thinking enables you to break down complex system behaviors to find the direct root cause of infrastructure issues.
  • Adaptability allows you to seamlessly transition between different cloud platforms and tools in a fast-paced environment.
  • Collaborative Problem-Solving ensures you work effectively with specialized engineering teams to build resilient, long-term technical solutions.

Essentials:

  • Extensive, highly advanced expertise in multi-cloud operations and Site Reliability Engineering, reflecting senior-level industry tenure.
  • Demonstrated multi-cloud expertise with hands-on technical capabilities across platforms.
  • Possession of at least two of the following certifications (Note: if you hold two, Deloitte will train and support you in acquiring the third):
    • Amazon Web Services (AWS) Certification
    • Google Cloud Platform (GCP) Certification
    • Oracle Cloud Infrastructure (OCI) Certification
  • Experience with observability tools
  • Experience with integration flows
  • Experience with DevOps tools
  • Familiarity with foundational Site Reliability Engineering (SRE) practices and automation frameworks.
  • 3 to 7 years minimum prior experience operating within an enterprise-level IT command center.

Similar Jobs

About Deloitte
Egypt, Egypt
Management Consulting