Our Tech Ops & Support Engineer Ensures the stability, availability, and performance of production environments by monitoring systems, automating operational processes, supporting deployments, resolving incidents, and maintaining secure, reliable platform operations.
Responsibilities
Monitor, maintain, and optimize production environments to ensure high availability, performance, and compliance with established service level objectives (SLAs/SLOs) Manage deployments, CI/CD pipelines, and release activities while supporting configuration management and environment consistency across production and non-production environments Administer and troubleshoot containerized platforms, IBM Cloud Pak Stacks, Confluent, Elasticsearch, and supporting infrastructure to ensure reliable system operations Investigate incidents, perform root cause analysis, implement corrective actions, and contribute to post-incident reviews to improve service reliability Configure and maintain monitoring, logging, and alerting solutions to proactively identify performance issues and minimize service disruptions Support backup, disaster recovery, security, patching, and operational compliance activities while maintaining accurate technical documentation and runbooks Collaborate with Development, QUALITY, Platform Engineering, Network, and Support teams to resolve complex technical issues and continuously improve operational processes
Requirements
Bachelor's degree or Diploma in Computer Science, Engineering, or a related field Around 2+ years of experience in Technical Operations, Production Support, Site Reliability Engineering (SRE), Dev Ops, or System Administration Good experience with Linux/Windows administration, Docker, Kubernetes, Open Shift, CI/CD pipelines, automation tools (Jenkins, Git Lab CI, Git Hub Actions), and scripting (Bash, Power Shell, Python) Working knowledge of IBM Cloud Pak solutions (CP4BA, CP4I, CP4D), Confluent, Elasticsearch, Databases (SQL, DB2, Mongo DB, etc.), and observability platforms such as Grafana and Prometheus Good understanding of networking fundamentals, security best practices, IAM, backup and disaster recovery, configuration management, and production environment support Strong troubleshooting, incident management, root cause analysis, communication, and collaboration skills with the ability to work effectively under pressure Ability to participate in on-call support, prioritize operational issues, and deliver reliable production support while driving continuous service improvement
مهندس Tech Ops & Support لدينا يضمن استقرار البيئات الإنتاجية وتوافرها وأدائها من خلال رصد الأنظمة، أتمتة العمليات التشغيلية، دعم عمليات النشر، حل الحوادث، والحفاظ على عمليات المنصة آمنة وموثوقة.
المسؤوليات
رصد البيئات الإنتاجية وصيانتها وتحسينها لضمان التوافر العالي، الأداء، والالتزام بأهداف مستوى الخدمة المعتمدة (SLAs/SLOs) إدارة عمليات النشر وخطوط CI/CD وأنشطة الإصدار مع دعم إدارة التكوين وتحقيق الاتساق في البيئات الإنتاجية وغير الإنتاجية
إدارة المنصات المعبأة بالحاويات، IBM Cloud Pak Stacks، Confluent، Elasticsearch، والبُنى التحتية الداعمة لضمان تشغيل النظام بشكل موثوق
التحقيق في الحوادث، إجراء تحليل السبب الجذري، تنفيذ إجراءات تصحيحية، والمساهمة في المراجعات ما بعد الحوادث لتحسين موثوقية الخدمة
ضبط وصيانة حلول الرصد، والتسجيل، والتنبيه لاكتشاف مشاكل الأداء بشكل استباقي وتقليل الانقطاعات
دعم نسخ الاحتياطي، التعافي من الكوارث، الأمن، التصحيح، والامتثال التشغيلي مع الحفاظ على وثائق تقنية دقيقة وكتب التشغيل
التعاون مع فرق التطوير، الجودة، هندسة المنصات، الشبكات والدعم لحل القضايا التقنية المعقدة وتحسين العمليات التشغيلية باستمرار
المتطلبات
درجة البكالوريوس أو الدبلوم في علوم الحاسوب، الهندسة، أو مجال ذي صلة ما يقرب من 2+ سنوات من الخبرة في العمليات الفنية، دعم الإنتاج، هندسة الاعتماد على النظام (SRE)، DevOps، أو إدارة الأنظمة يفضل خبرة جيدة في إدارة Linux/Windows، Docker، Kubernetes، Open Shift، خطوط CI/CD، أدوات الأتمتة (Jenkins، Git Lab CI، Git Hub Actions)، والبرمجة النصية (Bash، Power Shell، Python) معرفة عملية بحلول IBM Cloud Pak (CP4BA، CP4I، CP4D)، وConfluent، وElasticsearch، وقواعد البيانات (SQL، DB2، MongoDB، إلخ)، ومنصات الرصد مثل Grafana وPrometheus فهم جيد لمبادئ الشبكات، وممارسات الأمن، IAM، النسخ الاحتياطي والتعافي من الكوارث، إدارة التكوين، ودعم بيئة الإنتاج مهارات قوية في استقصاء المشكلات، إدارة الحوادث، تحليل السبب الجذري، الاتصالات والتعاون مع القدرة على العمل بفاعلية تحت الضغط القدرة على المشاركة في دعم التواجد على النوبات، وأولويات قضايا التشغيل، وتقديم دعم إنتاج موثوق مع دفع تحسين الخدمة بشكل مستمر