الوصف الوظيفي
ما ستفعله
- تصميم وبناء وصيانة أنابيب قابلية الملاحظة عبر بيئتنا متعددة السحاب (Azure، GCP)، والتي تغطي المقاييس والسجلات والتتبعات.
- إنشاء وصيانة لوحات تحكم تمنح فرق الهندسة رؤية واضحة وقابلة للتنفيذ حول صحة النظام وأدائه والمقاييس ذات الصلة بالأعمال.
- دمج جمع المقاييس عبر الخدمات والبنية التحتية، مع ضمان معايير أجهزة قياس متسقة.
- بناء وتحسين أطر عمل التنبيهات المستخدمة لدعم المناوبات في وضع الاستعداد بشكل مستمر، مع تحقيق التوازن بين جودة الإشارة وإجهاد التنبيهات.
- العمل مع أدوات قابلية الملاحظة القياسية في القطاع مثل Grafana وPrometheus وELK Stack لجمع بيانات القياس عن بُعد وتخزينها وعرضها مرئياً.
- تشغيل وتحسين أدوات مراقبة أداء التطبيقات (APM) (مثل Dynatrace وDatadog) لدعم مراقبة أداء التطبيقات وتحليل السبب الجذر.
- الاستفادة من Azure Monitor ولغة استعلام Kusto (KQL) لبناء الاستعلامات والتنبيهات ودفاتر العمل لأحمال العمل المستضافة على Azure.
- الاستفادة من GCP Monitoring (سابقاً Stackdriver) لتجميع السجلات والمقاييس عبر الخدمات المستضافة على GCP.
- العمل على ممارسة هندسة موثوقية المواقع (SRE) مع التركيز على التعاون في تحديد SLOs وOLAs لترجمة نقاط المقاييس المجمعة إلى مؤشرات أداء رئيسية للأعمال (KPIs) ولوحات تحكم لاحقة.
- الشراكة مع فرق المنصة وDevOps وهندسة التطبيقات لتحديد مفهوم "قابلية الملاحظة الجيدة" لكل خدمة ومساعدة الفرق على اعتمادها.
- المساهمة في الاستجابة للحوادث وعمليات المراجعة بعد الحوادث (postmortem) من خلال تحسين إشارات قابلية الملاحظة التي تدعم الكشف والتشخيص الأسرع.
- توثيق معايير قابلية الملاحظة وكتيبات التشغيل (runbooks) وأفضل الممارسات للمساعدة في توسيع نطاق المعرفة عبر المؤسسة الهندسية.
- المشاركة بفعالية في مجتمع الممارسة / نقابة السحاب (Cloud guild) لتعزيز تبادل المعرفة بين المجالات المختلفة وتقليل عزلة المعرفة.
ما يجعلك مرشحاً مناسباً
- خبرة من 4 إلى 7 سنوات.
- خبرة في العمل مع المنصات السحابية، ويفضل Azure و/أو GCP.
- خبرة عملية في أدوات قابلية الملاحظة القياسية في القطاع مثل Grafana وPrometheus وELK Stack.
- خبرة في السلاسل الزمنية (مثل InfluxDB) واستراتيجيات الاحتفاظ بالسجلات باستخدام أدوات مثل Thanos.
- خبرة في PromQL وKQL/Kusto ولغات الاستعلام الأخرى.
- خبرة في أدوات مراقبة أداء التطبيقات (APM) مثل Dynatrace أو Datadog.
- قدرة مثبتة على تصميم وبناء لوحات التحكم ودمج جمع المقاييس عبر الخدمات.
- خبرة في إنشاء أو صيانة أطر عمل التنبيهات التي تدعم مهام المناوبة/الاستدعاء.
- معرفة بـ Azure Monitor وKusto/KQL للاستعلام والتنبيه.
- معرفة بـ GCP Monitoring، بما في ذلك تجميع السجلات والمقاييس.
- فهم قوي لهندسة المنصات أو ممارسات DevOps وكيفية ملاءمة قابلية الملاحظة في دورة حياة تسليم البرامج الأوسع.
- إتقان اللغة الإنجليزية (مطلوب)؛ وتعتبر مهارات اللغة الألمانية ميزة إضافية.
- الارتياح للعمل في بيئة تعتمد على العمل عن بُعد أولاً مع البقاء على اتصال وثيق مع فريق موزع.
- الرغبة في السفر إلى مواقع تكنولوجيا المعلومات الرئيسية لشركة Henkel (بحد أقصى مرة واحدة في الربع السنوي) لحضور ورش عمل الفريق.
- الخبرة في البرمجة النصية (Python, Bash, Go) وأدوات DevOps الأخرى (OpenTofu) هي ميزة إضافية وليست إلتزاماً.
- خبرة في أدوات البنية التحتية كرمز (Infrastructure-as-Code) (مثل OpenTofu وAnsible).
- معرفة بمفاهيم وأدوات التتبع الموزع (مثل OpenTelemetry).
- مشاركة سابقة في مناوبات الاستدعاء لأنظمة الإنتاج.
- خبرة في البرمجة النصية (مثل Python وBash وPowerShell) لأتمتة مهام قابلية الملاحظة.
بعض مزايا الانضمام إلى هنكل (Henkel)
- نظام عمل مرن بساعات عمل مرنة، ونموذج عمل هجين، وسياسة العمل من أي مكان لمدة تصل إلى 30 يوماً في السنة.
- فرص نمو متنوعة على المستوى المحلي والدولي.
- معايير رفاهية عالمية مع برامج الرعاية الصحية والوقائية.
- إجازة رعاية أطفال محايدة جنسانياً لمدة لا تقل عن 8 أسابيع.
- خطة أسهم الموظفين مع استثمار اختياري وأسهم مطابقة من Henkel.
- تأمين صحي شامل للموظف والمعالين.
- برنامج مساعدة الموظفين يقدم مجموعة واسعة من مزايا الصحة النفسية والرفاهية.
في هنكل (Henkel)، نأتي من مجموعة واسعة من الخلفيات ووجهات النظر والخبرات الحياتية. ونحن نؤمن بأن تفرد جميع موظفينا هو مصدر قوتنا. كن جزءاً من الفريق وأضف تميزك إلينا! نحن نبحث عن فريق متنوع من الأفراد الذين يتمتعون بخلفيات وخبرات وشخصيات وعقليات مختلفة.
Job description
What you´ll do
- Design, build, and maintain observability pipelines across our multi-cloud environment (Azure, GCP), covering metrics, logs, and traces.
- Create and maintain dashboards that give engineering teams clear, actionable visibility into system health, performance, and business-relevant metrics.
- Integrate metrics collection across services and infrastructure, ensuring consistent instrumentation standards.
- Build and continuously improve alerting frameworks used to support on-call rotations, balancing signal quality with alert fatigue.
- Work with industry-standard observability tools such as Grafana, Prometheus, and the ELK Stack to collect, store, and visualize telemetry data.
- Operate and optimize APM tooling (e.g., Dynatrace, Datadog) to support application performance monitoring and root-cause analysis.
- Leverage Azure Monitor and Kusto Query Language (KQL) to build queries, alerts, and workbooks for Azure-hosted workloads.
- Leverage GCP Monitoring (formerly Stackdriver) for log and metrics aggregation across GCP-hosted services.
- Work on the Site Reliability Engineering practice focusing on co-defining SLOs and OLAs to translate collected metric points into business KPIs and subsequent dashboards
- Partner with platform, DevOps, and application engineering teams to define what "good observability" looks like for each service and help teams adopt it.
- Contribute to incident response and postmortem processes by improving the observability signals that support faster detection and diagnosis.
- Document observability standards, runbooks, and best practices to help scale knowledge across the engineering organization.
- Actively participate in the Community of Practice / Cloud guild to foster cross-domain knowledge sharing and reduce knowledge silos
What makes you a good fit
- 4 - 7 years of experience.
- Experience working with cloud platforms, ideally Azure and/or GCP.
- Hands-on experience with industry-standard observability tools such as Grafana, Prometheus, and the ELK Stack.
- Experience with Time Series (e.g. InfluxDB) and Log retention strategies using tools such as Thanos
- Experience with PromQL, KQL/Kusto and other Query Languages
- Experience with APM tools such as Dynatrace or Datadog.
- Proven ability to design and build dashboards and integrate metrics collection across services.
- Experience creating or maintaining alerting frameworks that support on-call duty.
- Familiarity with Azure Monitor and Kusto/KQL for querying and alerting.
- Familiarity with GCP Monitoring, including log and metrics aggregation.
- Solid understanding of platform engineering or DevOps practices and how observability fits into the broader software delivery lifecycle.
- Fluent English (required); German language skills are a plus.
- Comfortable working in a remote-first setup while staying closely connected with a distributed team
- Willingness to travel to main Henkel IT sites (max. once per quarter) for team workshops
- Experience in scripting (Python, Bash, Go) and other DevOps (OpenTofu) tooling is a plus, but not a must.
- Experience with infrastructure-as-code tools (e.g., OpenTofu and Ansible).
- Familiarity with distributed tracing concepts and tools (e.g., OpenTelemetry).
- Prior participation in an on-call rotation for production systems.
- Scripting experience (e.g., Python, Bash, PowerShell) for automating observability tasks.
Some perks of joining Henkel
- Flexible work scheme with flexible hours, hybrid work model, and work from anywhere policy for up to 30 days per year
- Diverse national and international growth opportunities
- Global wellbeing standards with health and preventive care programs
- Gender-neutral parental leave for a minimum of 8 weeks
- Employee Share Plan with voluntary investment and Henkel matching shares
- Comprehensive Health Insurance for employee + dependents
- Employee Assistance Programme provides a wide range of mental health and wellbeing benefits
At Henkel, we come from a broad range of backgrounds, perspectives, and life experiences. We believe the uniqueness of all our employees is the power in us. Become part of the team and bring your uniqueness to us! We look for a diverse team of individuals who possess different backgrounds, experiences, personalities and mindsets.