ملخص العمل
نحن نبحث عن مهندس تشغيل تطبيقات ذو خبرة لدعم الاستقرار والتوافر والأداء والتحسين المستمر لتطبيقات المؤسسة. تغطي أدوار تشغيل التطبيقات أنشطة الحوادث والمشكلات والتغيير والإصدار والنسخ الاحتياطي/الاسترداد والتعافي من الكوارث وتحسين الخدمة.
المسؤوليات الأساسية
إدارة حوادث المستوى 2/3 الوظيفية، بما في ذلك المشاركة في جسور الحوادث وإدارة التصعيد. إجراء إدارة الحوادث والمشاكل، بما في ذلك تحليل اتجاهات الحادث وتنفيذ الإصلاحات الدائمة. تطوير وإدارة خطط تحسين الخدمة. دعم جدولة الإصدار وأنشطة التخطيط. التخطيط والتنسيق لطلبات التغيير (CRs)، بما في ذلك تقييمات المخاطر والتأثير. تنفيذ تغييرات على طبقة التطبيق وإجراء مراجعات ما بعد التنفيذ. معالجة طلبات تكوين التطبيق. مراقبة أداء وتوفر التطبيق. إجراء ترابط الأحداث والتحقيق في الأحداث المتعلقة بالتطبيق. إدارة سعة وأداء التطبيق. تطوير واختبار إجراءات استرداد التطبيق. تكوين النسخ الاحتياطي للتطبيق وتنفيذ استعادة التطبيق والبيانات. المشاركة في اختبارات التعافي من الكوارث (DR) وتنفيذها. إجراء تقوية لنظام التشغيل، وقاعدة البيانات، ووسطاء البرمجيات. تطبيق التصحيحات والتحديثات لمكونات التطبيق. الحفاظ على سجلات CMDB دقيقة لمكونات التطبيق. التعاون مع الفرق التقنية المعنية لضمان استقرار التطبيق وتوفر الخدمة
المتطلبات
درجة البكالوريوس في علوم الحاسوب أو تكنولوجيا المعلومات أو الهندسة أو مجال ذو صلة. خبرة ذات صلة في تشغيل التطبيقات / دعم التطبيقات في بيئة مؤسسية. خبرة عملية قوية في إدارة حوادث ومشاكل التطبيق من المستوى 2/3. خبرة في إدارة التغيير والتخطيط للإصدار وتنفيذ تغييرات طبقة التطبيق. خبرة في إجراء تقييمات المخاطر والتأثير لتغييرات التطبيق. خبرة عملية في مراقبة أداء وتوافر التطبيق وربط الأحداث. خبرة في إجراءات النسخ الاحتياطي والاستعادة والتعافي للتطبيق. خبرة في المشاركة في أو تنفيذ اختبارات التعافي من الكوارث. فهم جيد لتعزيز أمان النظام التشغيلي وقاعدة البيانات ووسطاء البرمجيات والتحديثات. خبرة في إدارة السعة والأداء. خبرة في الحفاظ على سجلات CMDB لمكونات التطبيق. مهارات قوية في استكشاف الأخطاء وتحليلها والتصعيد والتواصل. القدرة على العمل مع فرق تقنية متعددة أثناء الحوادث الكبرى وجسور الحوادث
Job Summary
We are looking for an experienced Application Operations Engineer to support the stability, availability, performance, and continuous improvement of enterprise applications. The role covers application operations across incident, problem, change, release, backup/recovery, disaster recovery, and service improvement activities.
Key Responsibilities
Manage L2/L3 functional incidents, including participation in incident bridges and escalation management. Perform incident and problem management, including incident trend analysis and implementation of permanent fixes. Develop and manage Service Improvement Plans. Support release schedule and planning activities. Plan and coordinate Change Requests (CRs), including risk and impact assessments. Perform application-layer change implementation and conduct post-implementation reviews. Handle application configuration requests. Monitor application performance and availability. Perform event correlation and investigate application-related events. Manage application capacity and performance. Develop and test application recovery procedures. Configure application backups and execute application and data restores. Participate in and execute Disaster Recovery (DR) tests. Perform OS, database, and middleware hardening activities. Apply patches and updates to application components. Maintain accurate CMDB records for application components. Collaborate with relevant technical teams to ensure application stability and service availability
Requirements
Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field. Relevant experience in Application Operations / Application Support in an enterprise environment. Strong hands-on experience with L2/L3 application incident and problem management. Experience with change management, release planning, and application-layer change implementation. Experience conducting risk and impact assessments for application changes. Hands-on experience with application performance and availability monitoring and event correlation. Experience with application backup, restore, and recovery procedures. Experience participating in or executing Disaster Recovery (DR) tests. Good understanding of OS, database, and middleware hardening and patching. Experience with capacity and performance management. Experience maintaining CMDB application component records. Strong troubleshooting, analytical, escalation management, and communication skills. Ability to work with multiple technical teams during major incidents and incident bridges