حزمة التقنيات الأساسية: AWS (أكثر من 5 سنوات) • EKS • ECS • Terraform • Istio • Datadog • Git Lab CI/CD (مستوى متقدم) • Argo CD • Serverless Framework • Node.js • Python • Temporal • Docker • Cloud Formation
نظرة عامة على الدور: بصفتك مهندس ديف أوبس أول، ستتخذ نهجاً عملياً في تقديم وأتمتة بنيتنا التحتية السحابية، مع دعم التكامل والنشر المستمر. ستعمل عن كثب مع قائد فريق ديف أوبس التقني وتتعاون مع فرق هندسة البرمجيات لبناء وتأمين وتوسيع نطاق منصتنا. هذه بيئة خدمات مالية حيث الأمن والامتثال والموثوقية ليست خيارات. يُتوقع منك بناء وصيانة بنية تحتية تلبي متطلبات SOC 2 Type II و PCI DSS و GDPR كمتطلب أساسي مع الحفاظ على تجربة مطور سريعة وخالية من الاحتكاك - بالشراكة مع الفرق الهندسية.
المسؤوليات: البنية التحتية والمنصة: تصميم وبناء وصيانة بنية تحتية سحابية على AWS تتسم بالأمان والأتمتة وقابلية التوسع والتوافر العالي. امتلاك وتطوير منصة Kubernetes (EKS) الخاصة بنا بما في ذلك عمليات العنقود، والترقيات، والشبكات، وتعزيز الأمان. إدارة وتشغيل ECS (أنواع تشغيل Fargate و EC2) لأحمال العمل المعتمدة على الحاويات، بما في ذلك تعريفات المهام، وتوسيع نطاق الخدمة، واستراتيجيات النشر. تصميم وصيانة شبكات AWS المتقدمة: بوابات العبور (Transit Gateways) لاتصال VPC متعدد/حساب متعدد، شبكة VPN من موقع إلى موقع، AWS Client VPN، ربط VPC، الاتصال المباشر (Direct Connect)، حل أسماء النطاقات Route 53، وتقسيم الشبكة عبر البيئات. إدارة وتطوير شبكة الخدمات (Service Mesh) Istio الخاصة بنا لإدارة حركة المرور، وفرض التشفير المتبادل (mTLS)، وعمليات النشر التدريجي (canary deployments)، وقابلية المراقبة عبر الخدمات المصغرة. توفير وتكوين وصيانة جميع البنية التحتية لـ AWS ككود من خلال Terraform، مع فرض أفضل الممارسات حول الوحدات النمطية، وإدارة الحالة، واكتشاف الانحراف. صيانة وتوسيع مجموعات Cloud Formation حيثما أمكن، بما في ذلك بنى المجموعات المتداخلة والتبعيات عبر المجموعات. تنفيذ وإدارة سير عمل Git Ops باستخدام Argo CD لعمليات نشر Kubernetes القائمة على التصريح والقابلة للتدقيق. بناء وتشغيل أحمال العمل بدون خادم (serverless) باستخدام Serverless Framework للخدمات القائمة على الأحداث وواجهات برمجة التطبيقات. صيانة ودعم البنية التحتية المحيطة بـ Temporal: عمليات العنقود، وإدارة مساحة الأسماء، والمراقبة، والتوسيع، وضمان التوافر العالي للفرق الهندسية التي تبني سير عمل دائم.
قابلية المراقبة والموثوقية: نشر وتكوين Datadog عبر الحزمة الكاملة: مقاييس البنية التحتية، APM (التتبع الموزع)، إدارة السجلات، المراقبة الاصطناعية، وأهداف مستوى الخدمة (SLOs). تصميم استراتيجيات الوسم ومعايير لوحة التحكم التي تمنح الفرق الهندسية رؤية ذاتية لخدماتهم. بناء تنبيهات مفيدة ذات ضجيج منخفض: شاشات مركبة، كشف الشذوذ، وتنبيهات استهلاك ميزانية الخطأ. إجراء تحليل وتطوير تكلفة البنية التحتية بشكل مستمر، بما في ذلك تخطيط السعة المحجوزة وتعديل الحجم.
الأمن والامتثال: صيانة بنية تحتية تلبي متطلبات الامتثال SOC 2 Type II و PCI DSS و GDPR. تنفيذ وفرض التشفير في حالة السكون وأثناء النقل عبر جميع الخدمات باستخدام مفاتيح AWS KMS المدارة من قبل العميل. ضمان تسجيل تدقيق شامل: سجلات تنظيم Cloud Trail، سجلات تدقيق مستوى التطبيق، سجلات تدقيق EKS، والاحتفاظ بالسجلات غير القابلة للتلاعب. فرض سياسات IAM ذات الامتيازات الأقل، وسياسات التحكم في الخدمة (SCPs) عبر مؤسسات AWS، وإدارة الأسرار من خلال AWS Secrets Manager أو SSM Parameter Store. صيانة تقسيم الشبكة وضوابط الأمان لبيئات بيانات حاملي البطاقات بما يتماشى مع PCI DSS. التعاون مع مهندس أمن التطبيقات في إدارة الثغرات، ومسح الأسرار في CI، وجمع أدلة الامتثال لعمليات التدقيق.
CI/CD وتجربة المطور: امتلاك وهندسة خطوط أنابيب Git Lab CI/CD المتقدمة عبر جميع الخدمات: خطوط أنابيب متعددة المراحل والبيئات مع خطوط أنابيب فرعية ديناميكية، وتبعيات DAG، وتنفيذ قائم على القواعد. فرض مبدأ الامتيازات الأقل (POLP) عبر CI/CD: متغيرات CI/CD محددة النطاق، بيئات محمية، رموز تشغيل ذات امتيازات محدودة، حسابات خدمة لكل وظيفة، وأقل صلاحيات لصورة الحاوية. تصميم وصيانة معايير .gitlab-ci.yml عبر المؤسسة: تضمينات قابلة لإعادة الاستخدام، وتوسيع، وقوالب مكونات، ومكتبات CI مشتركة للقضاء على التكرار وفرض الاتساق. بناء وصيانة أدوات النشر لخدمات Node.js و Python عبر أهداف الحاويات والأهداف بدون خادم. ضمان امتثال خط الأنابيب: موافقات قائمة على MR، ومصنوعات موقعة، وبوابات النشر، وقواعد الموافقة، والتتبع الكامل من التذكرة إلى الإنتاج. إدارة البنية التحتية لـ Git Lab Runner: تشغيل برامج التشغيل ذاتية التوسع على AWS (EC2/EKS)، واستراتيجيات التخزين المؤقت للبرامج، و Docker-in-Docker مقابل Kaniko لبناء الحاويات، وتعزيز أمن البرامج. تحسين أداء خط الأنابيب: التخزين المؤقت، المصنوعات، الوظائف المتوازية، مجموعات الموارد، وتحليلات خط الأنابيب للحفاظ على سرعة حلقات التغذية الراجعة.
التعاون والعمل بنظام الاستدعاء: العمل بشكل تعاوني مع فرق هندسة البرمجيات لتحديد متطلبات البنية التحتية والنشر والتشغيل. المساهمة في تقييم التقنيات الجديدة ومنتجات البائعين التي تحقق قيمة إضافية عبر الحزمة. مشاركة المعرفة وتوجيه أعضاء الفريق المبتدئين من خلال مراجعات الكود والتوثيق وجلسات العمل الثنائي. ترجمة متطلبات الأعمال غير التقنية إلى متطلبات بنية تحتية تقنية. المشاركة في دورة الاستدعاء: الاستجابة لحوادث الإنتاج، وإجراء تحليل السبب الجذري، ودفع تحسينات ما بعد الحادث لمنع التكرار.
المهارات والخبرة المطلوبة: يجب توفرها (بالترتيب): خبرة عميقة في AWS عبر الخدمات الأساسية: VPC, IAM, EKS, ECS (Fargate and EC2), Lambda, S3, RDS, KMS, Cloud Trail, Route 53, API Gateway, and Organizations. (أكثر من 5 سنوات). شبكات AWS متقدمة: بنيات Transit Gateway، شبكة VPN من موقع إلى موقع، AWS Client VPN، تصميم VPC وتقسيم الشبكة الفرعية، DNS (نطاقات Route 53 الخاصة/العامة المستضافة، نقاط نهاية Resolver)، ربط VPC، الاتصال المباشر (Direct Connect)، Private Link، وموازنة الأحمال (ALB/NLB). خبرة عملية في Kubernetes (يفضل EKS بشدة) بما في ذلك عمليات العنقود، RBAC، الشبكات، HPA، واستكشاف الأخطاء وإصلاحها. (أكثر من 3 سنوات). خبرة عملية في شبكة الخدمات Istio: إدارة حركة المرور، mTLS، تكوين الخدمة الافتراضية/قاعدة الوجهة، ضبط الـ sidecar، وتصحيح الأخطاء باستخدام istioctl. (أكثر من 3 سنوات). خبرة متقدمة في Terraform بما في ذلك الوحدات النمطية، مساحات العمل، إدارة الحالة، وسير عمل التخطيط/التطبيق المدفوع بـ CI. (أكثر من 3 سنوات). خبرة متقدمة في Git Lab CI/CD: فهم ممتاز لبناء جملة CI (القواعد، الاحتياجات، DAG، خطوط الأنابيب الفرعية الديناميكية، التضمينات/التوسيع، المكونات)، أمن خط الأنابيب (POLP، المتغيرات/البيئات المحمية، عزل برامج التشغيل)، سير عمل موافقة MR، وإدارة البنية التحتية لبرامج التشغيل. (4–5 سنوات). Serverless Framework لنشر وإدارة أحمال العمل المستندة إلى Lambda. تنفيذ وتشغيل Datadog: نشر الوكيل في Kubernetes، أداة APM، خطوط سجلات البيانات، مقاييس مخصصة، لوحات تحكم، شاشات، وأهداف SLOs. عقلية الأمن والامتثال: خبرة عملية في SOC 2، PCI DSS، أو أطر تنظيمية مماثلة في بيئة التكنولوجيا المالية أو الخدمات المالية. خبرة عملية في Node.js: القدرة على قراءة وتصحيح وكتابة كود Node.js. فهم وقت التشغيل، ودورة حياة LTS، والتعبئة. برمجة Python للأتمتة: boto3، أدوات CLI، نصوص CI، معالجة البيانات. القدرة على التعامل مع أنماط التزامن والتعبئة. Argo CD لعمليات نشر Kubernetes القائمة على Git Ops. Docker وأفضل ممارسات الحاويات: تحسين الصور، البناء متعدد المراحل، مسح الثغرات.
الشهادات المطلوبة: مهندس ديف أوبس معتمد من AWS — محترف، مهندس حلول معتمد من AWS — مشارك، شهادة Hashi Corp: مساعد أو مهندس Terraform.
مزايا إضافية: عمليات منصة Temporal: نشر العنقود، تكوين مساحة الأسماء، ضبط متجر الرؤية، وبنية تحتية للعامل. Cloud Formation بما في ذلك المجموعات المتداخلة، الماكرو، والمراجع عبر المجموعات. أخصائي أمن Kubernetes معتمد (CKS). تكامل Azure Active Directory و Microsoft 365. خبرة في JIRA و Confluence لتتبع المشاريع والتوثيق. الإلمام بمجموعات أدوات SIEM ومنصات إدارة سجلات السحابة بخلاف Datadog. خبرة في تشغيل دورات تدقيق SOC 2 Type II من البداية إلى النهاية، بما في ذلك جمع الأدلة من جانب ديف أوبس. مساهمات في أدوات البنية التحتية مفتوحة المصدر.
ما نقدره: الملكية: ترى مشكلة، فتقوم بإصلاحها. لا تنتظر تذكرة. البراغماتية قبل الكمال: شحن حلول آمنة وقابلة للصيانة، والتكرار لاحقاً. التواصل الواضح: يمكنك شرح تصميم VPC لمدير منتج ومتطلبات امتثال لمهندس. الفضول: الرغبة في تعلم وتقييم مجموعة واسعة من التقنيات والأدوات مفتوحة المصدر. الموثوقية: مريح مع مسؤوليات الاستدعاء وملتزم بالحفاظ على صحة بيئة الإنتاج.
Core Technology Stack AWS (5+ years) •EKS • ECS • Terraform• Istio •Datadog • Git Lab CI/CD (Advanced) •Argo CD • Serverless Framework• Node.js •Python • Temporal• Docker •Cloud Formation
Role Overview As a Senior Dev Ops Engineer, you will take a hands-on approach to the delivery and automation of our cloud infrastructure, supporting continuous integration and deployment. You will work closely with the Dev Ops Technical Lead and collaborate with software engineering teams to build, secure, and scale our platform. This is a financial services environment where security, compliance, and reliability are not optional. You will be expected to build and maintain infrastructure that meets SOC 2 Type II, PCI DSS, and GDPR as a basic requirement while keeping the developer experience fast and friction-free – partnering with engineering teams.
Responsibilities Infrastructure & Platform Design, build, and maintain cloud-native infrastructure on AWS that is secure, automated, scalable, and highly available. Own and evolve our Kubernetes platform (EKS) including cluster operations, upgrades, networking, and security hardening. Manage and operate ECS (Fargate and EC2 launch types) for containerised workloads, including task definitions, service scaling, and deployment strategies. Design and maintain advanced AWS networking: Transit Gateways for multi-VPC/multi-account connectivity, Site-to-Site VPN, AWS Client VPN, VPC peering, Direct Connect, Route 53 DNS resolution, and network segmentation across environments. Manage and advance our Istio service mesh for traffic management, m TLS enforcement, canary deployments, and observability across microservices. Provision, configure, and maintain all AWS infrastructure as code through Terraform, enforcing best practices around modules, state management, and drift detection. Maintain and extend Cloud Formation stacks where applicable, including nested stack architectures and cross-stack dependencies. Implement and manage Git Ops workflows using Argo CD for declarative, auditable Kubernetes deployments. Build and operate serverless workloads using the Serverless Framework for event-driven and API-backed services. Maintain and support the infrastructure around Temporal: cluster operations, namespace management, monitoring, scaling, and ensuring high availability for engineering teams building durable workflows.
Observability & Reliability Deploy and configure Datadog across the full stack: infrastructure metrics, APM (distributed tracing), log management, synthetic monitoring, and SLOs. Design tagging strategies and dashboard standards that give engineering teams self-service visibility into their services. Build meaningful alerting with low noise: composite monitors, anomaly detection, and error budget burn-rate alerts. Perform infrastructure cost analysis and optimization on an ongoing basis, including reserved capacity planning and right-sizing.
Security & Compliance Maintain infrastructure that satisfies SOC 2 Type II, PCI DSS, and GDPR compliance requirements. Implement and enforce encryption at rest and in transit across all services using AWS KMS customer-managed keys. Ensure comprehensive audit logging: Cloud Trail organization trails, application-level audit logs, EKS audit logs, and tamper-proof log retention. Enforce least-privilege IAM policies, SCPs across AWS Organizations, and secret management through AWS Secrets Manager or SSM Parameter Store. Maintain network segmentation and security controls for cardholder data environments in line with PCI DSS. Collaborate with the App Sec engineer on vulnerability management, secret scanning in CI, and compliance evidence collection for audits.
CI/CD & Developer Experience Own and architect advanced Git Lab CI/CD pipelines across all services: multi-stage, multi-environment pipelines with dynamic child pipelines, DAG dependencies, and rules-based execution. Enforce Principle of Least Privilege (POLP) across CI/CD: scoped CI/CD variables, protected environments, least-privilege runner tokens, per-job service accounts, and minimal container image permissions. Design and maintain .gitlab-ci.yml standards across the organisation: reusable includes, extends, component templates, and shared CI libraries to eliminate duplication and enforce consistency. Build and maintain deployment tooling for Node.js and Python services across containerized and serverless targets. Ensure pipeline compliance: MR-based approvals, signed artifacts, deployment gates, approval rules, and full traceability from ticket to production. Manage Git Lab Runner infrastructure: autoscaling runners on AWS (EC2/EKS), runner caching strategies, Docker-in-Docker vs Kaniko for container builds, and runner security hardening. Optimise pipeline performance: caching, artifacts, parallel jobs, resource groups, and pipeline analytics to keep feedback loops fast.
Collaboration & On-Call Work collaboratively with software engineering teams to define infrastructure, deployment, and operational requirements. Contribute to the evaluation of new technologies and vendor products that drive additional value across the stack. Share knowledge and mentor junior team members through code reviews, documentation, and pairing sessions. Translate non-technical business requirements into technical infrastructure requirements. Participate in an on-call rotation: respond to production incidents, perform root cause analysis, and drive post-incident improvements to prevent recurrence.
Required Skills & Experience Must Have (in priority order) Deep AWS expertise across core services: VPC, IAM, EKS, ECS (Fargate and EC2), Lambda, S3, RDS, KMS, Cloud Trail, Route 53, API Gateway, and Organizations. (5+ years) Advanced AWS networking: Transit Gateway architectures, Site-to-Site VPN, AWS Client VPN, VPC design and subnetting, DNS (Route 53 private/public hosted zones, Resolver endpoints), VPC peering, Direct Connect, Private Link, and load balancing (ALB/NLB). Production Kubernetes experience (EKS strongly preferred) including cluster operations, RBAC, networking, HPA, and troubleshooting. (3+ years) Hands-on Istio service mesh experience: traffic management, m TLS, Virtual Service/Destination Rule configuration, sidecar tuning, and debugging with istioctl.(3+ years) Advanced Terraform experience including modules, workspaces, state management, and CI-driven plan/apply workflows. (3+ years) Advanced Git Lab CI/CD expertise: excellent understanding of CI syntax (rules, needs, DAG, dynamic child pipelines, includes/extends, components), pipeline security (POLP, protected variables/environments, runner isolation), MR approval workflows, and runner infrastructure management. (4–5 years) Serverless Framework for deploying and managing Lambda-based workloads. Datadog implementation and operations: Agent deployment in Kubernetes, APM instrumentation, log pipelines, custom metrics, dashboards, monitors, and SLOs. Security and compliance mindset: practical experience with SOC 2, PCI DSS, or similar regulatory frameworks in a fintech or financial services environment. Node.js hands-on experience: comfortable reading, debugging, and writing Node.js code. Understanding of the runtime, LTS lifecycle, and packaging. Python scripting for automation: boto3, CLI tools, CI scripts, data processing. Comfortable with concurrency patterns and packaging. Argo CD for Git Ops-based Kubernetes deployments. Docker and container best practices: image optimization, multi-stage builds, vulnerability scanning.
Required Certifications AWS Certified Dev Ops Engineer — Professional AWS Certified Solutions Architect — Associate Hashi Corp Certified: Terraform Associate or Engineer
Nice to Have Temporal platform operations: cluster deployment, namespace configuration, visibility store tuning, and worker infrastructure. Cloud Formation including nested stacks, macros, and cross-stack references. Certified Kubernetes Security Specialist (CKS). Azure Active Directory and Microsoft 365 integration. Experience with JIRA and Confluence for project tracking and documentation. Familiarity with SIEM toolsets and cloud log management platforms beyond Datadog. Experience running SOC 2 Type II audit cycles end-to-end, including evidence collection from the Dev Ops side. Contributions to open-source infrastructure tooling.
What We Value Ownership: you see a problem, you fix it. You don't wait for a ticket. Pragmatism over perfection: ship secure, maintainable solutions, iterate later. Clear communication: you can explain a VPC design to a product manager and a compliance requirement to an engineer. Curiosity: the appetite to learn and evaluate a wide variety of open-source technologies and tools. Reliability: comfortable with on-call responsibilities and committed to keeping production healthy.