ما ستفعله
تصميم وتنفيذ أطر تقييم آلية لسير عمل الوكلاء المعتمدة على LangGraph وأنابيب الت orchestration.
تطوير مجموعات التقييم أثناء البناء تغطي سلوك الوكيل، دقة اختيار الأداة، جودة التفكير وجودة الناتج النهائي.
إنشاء منهجيات تقييم تجمع بين أساليب التقييم الحاسمة deterministic مع تقنيات التقييم على أساس LLM كقاض.
بناء مقلاع اختبارات لرسومات LangGraph والعُقَد وانتقالات الحالة ومسارات تنفيذ الوكيل.
تصميم محاكيات محادثة متعددة الجولات للتحقق من الاحتفاظ بالسياق، واستخدام الذاكرة، وتناسق سير العمل.
تعريف والحفاظ على مقاييس الاعتمادية، بما في ذلك pass@k ومنهجيات pass^k، لقياس معدلات النجاح والتناسق السلوكي.
تنفيذ بوابات نشر CI/CD التي تمنع الإصدارات تلقائياً عندما لا تُلبَّ معايير التقييم.
تطوير عمليات التحقق لبيئات الإصدار التجريبي، ومقارنات وضع الظل، وتدفقات الإنتاج المحكومة.
دمج تقويمات AWS AgentCore (وضع التقييم عند الطلب وعلى الإنترنت) ضمن اختبارات مستمرة وعمليات ضمان الجودة.
تحويل حوادث الإنتاج، فشلها، وسلوك الوكيل غير المتوقع إلى حالات اختبار الانحدار.
إقامة مقاييس التقييم، معايير الجودة، ومعايير القبول لإصدارات الوكلاء.
التعاون مع مهندسي AI وفرق المنصة وأصحاب المصلحة في المنتج لتحسين موثوقية وأداء الوكلاء.
تحديد آليات الرصد والتغذية المرتدة التي تربط سلوك الإنتاج بتحسينات إطار الاختبار.
إعداد وثائق لاستراتيجيات التقييم، مناهج التقدير، بوابات النشر، ومعايير الجودة.
ما تحتاجه لهذا الدور
المهارات:
• بنية نظام Agentic AI (تنسيق وكلاء، تفويض متعدد الوكلاء، استخدام الأدوات، إدارة الذاكرة/الحالة)
• أطر تنسيق الوكلاء (LangGraph, Strands, AutoGen, CrewAI، أو ما يماثلها — عمق على مستوى الإنتاج في واحد على الأقل)
• أنماط تكامل الوكيل إلى الأداة والوكيل إلى الوكيل (استدعاء الدوال، MCP، A2A، أو بروتوكولات مماثلة)
• بنية تطبيق LLM (خطوط RAG، هندسة الطلب/السياق، توجيه النماذج، تقييم مخرجات LLM)
• حوكمة الوكيل والضوابط التشغيلية (التفويض/الانفاذ السياسي للإجراءات المستقلة، الهوية للعوامل غير البشرية)
الخبرة:
• أكثر من 6 سنوات خبرة في عمارة البرمجيات، بما في ذلك أكثر من سنتين في هندسة أنظمة ذكاء اصطناعي مدعومة بنماذج لغة كبيرة أو أنظمة وكيلية في الإنتاج
• خبرة عملية مع واحد على الأقل من أطر تنسيق الوكلاء (LangGraph، Strands، AutoGen، CrewAI، أو ما يعادلها)، مع تحمل قرارات الهندسة، وليس فقط التنفيذ
• خبرة في تصميم تكامل الوكيل إلى أداة على مستوى المنصة (طبقة أداة/تكامل تستخدمها عدة وكلاء أو فرق، وليست توصيل أداة مخصصة لوكيل واحد)
• خبرة في هندسة اثنين على الأقل من ما يلي لأنظمة وكيلية: بيئة التشغيل/التنفيذ، تكامل الأداة/البوابة، الهوية والإنفاذ السياسي، الرصد والتقييم
• خبرة مثبتة في تقديم بنية للحوكمة أو المراجعة التقنية (ARB، مجلس مراجعة التصميم، أو ما يعادلها)
مرغوب فيه:
• خبرة عملية مع AWS Bedrock AgentCore (وقت التشغيل، البوابة، السجل، السياسة، التقييم، الرصد)
• خبرة مع أطر السياسة كرمز للوصول إلى الوكلاء (Cedar، OPA، أو مشابه) • الإلمام ببروتوكولات MCP وA2A
• دور معماري سابق يغطي عدة مجالات منصة في بيئة مؤسسة، منضبطة، أو محكومة بشدة
• خبرة في معايير هوية الوكيل (مثلاً اتحاد الهوية غير البشرية، MS Entra Agent ID أو ما يعادله)
• Architecture للراصد والتقييم خاصة بأنظمة وكيلية (تتبع تفكير الوكيل متعدد الخطوات، مراقبة الجودة/السلامة)
• بنية سحابية تدعم تشغيلات الوكيل المدار (AWS مفضل — الشبكات، IAM، الحوسبة المدارة)
المتقدم المرغوب فيه
What you will do
Design and implement automated evaluation frameworks for LangGraph-based agent workflows and orchestration pipelines.
Develop build-time evaluation suites covering agent behavior, tool selection accuracy, reasoning quality, and final output quality.
Create evaluation methodologies combining deterministic grading approaches with LLM-as-judge evaluation techniques.
Build test harnesses for LangGraph graphs, nodes, state transitions, and agent execution paths.
Design multi-turn conversation simulations to validate context retention, memory utilization, and workflow consistency.
Define and maintain reliability metrics, including pass@k and pass^k methodologies, to measure both success rates and behavioral consistency.
Implement CI/CD deployment gates that automatically block releases when evaluation thresholds are not met.
Develop validation processes for staging environments, shadow-mode comparisons, and controlled production rollouts.
Integrate AWS AgentCore Evaluations (on-demand and online evaluation modes) into continuous testing and quality assurance workflows.
Transform production incidents, failures, and unexpected agent behavior into regression test cases.
Establish evaluation metrics, quality benchmarks, and acceptance criteria for agent releases.
Collaborate with AI engineers, platform teams, and product stakeholders to improve agent reliability and performance.
Define observability and feedback mechanisms that connect production behavior with test framework improvements.
Produce documentation for evaluation strategies, scoring methodologies, deployment gates, and quality standards.
What you need for this
Skills:
• Agentic AI system architecture (agent orchestration, multi-agent delegation, tool use, memory/state management)
• Agent orchestration frameworks (LangGraph, Strands, AutoGen, CrewAI, or comparable — production-level depth in at least one)
• Agent-to-tool and agent-to-agent integration patterns (function calling, MCP, A2A, or equivalent protocols)
• LLM application architecture (RAG pipelines, prompt/context engineering, model routing, evaluation of LLM outputs)
• Agent governance and runtime controls (authorization/policy enforcement for autonomous actions, identity for non-human actors)
Experience:
• 6+ years software architecture experience, including 2+ years specifically architecting LLM-powered or agentic AI systems in production
• Hands-on production experience with at least one agent orchestration framework (LangGraph, Strands, AutoGen, CrewAI, or equivalent), owning architecture decisions, not just implementation
• Experience designing agent-to-tool integration at a platform level (a tool/integration layer consumed by multiple agents or teams, not a single agent’s custom tool wiring)
• Experience architecting at least two of the following for agentic systems: runtime/execution environment, tool/gateway integration, identity and policy enforcement, observability and evaluation
• Demonstrated experience presenting architecture for governance or technical review (ARB, design review board, or equivalent)
Nice to have:
• Hands-on experience with AWS Bedrock AgentCore (Runtime, Gateway, Registry, Policy, Evaluation, Observability)
• Experience with policy-as-code frameworks for agent authorization (Cedar, OPA, or similar) • Familiarity with MCP and A2A protocol specifications
• Prior architecture role spanning multiple platform domains in an enterprise, regulated, or highly governed environment
• Experience with agent identity standards (e.g., non-human identity federation, MS Entra Agent ID or equivalent)
• Observability and evaluation architecture specific to agentic systems (tracing multi-step agent reasoning, quality/safety monitoring)
• Cloud architecture supporting managed agent runtimes (AWS preferred — networking, IAM, managed compute)
Desired Candidate Profile