We are looking for a Senior ML / Evaluation Engineer to help define and implement quality standards for enterprise-grade AI agents and LLM-powered applications. In this role, you will design evaluation frameworks, build custom evaluation pipelines, and establish automated quality gates across the AI delivery lifecycle. You will work closely with AI Platform Engineers, ML Engineers, and DevOps teams to ensure reliable, measurable, and production-ready AI systems through scalable evaluation, observability, and governance practices.
What project we have for you
Our customer is a multinational corporation with more than a century of history and offices in over 180 countries. Their most ambitious goal at the time is to introduce a range of Reduced-Risk Products (RRPs). The target audience is more than 1 billion consumers around the globe. IT platform hosts 700+ applications.
Intellia’s mission is to help the client with the engineering of a comprehensive software ecosystem for a game-changing IoT product on the margin of innovative consumer experience and cutting-edge technology. Our teams are involved in the engineering of core platform components for best-in-class eCommerce, Digital Marketing and IoT solutions. As an Engineer, you will become a part of Core Architecture Team and be responsible for the architecture, implementation of best practices in our Digital Engineering Enterprise Platform.
The Platform is a set of services and internet applications that accelerate the development and delivery of software applications by taking care of common SDLC challenges. The Platform provides access and consumption for engineering teams to a set of services, technologies, practices for their development and for operating their application, ensuring a set of compliance and best practices.
What you will do
Design, implement, and maintain enterprise-grade evaluation frameworks for LLMs, AI agents, and multi-step AI workflows.
Develop and optimize LLM-as-a-judge evaluators to assess dimensions such as helpfulness, correctness, consistency, and policy compliance.
Build custom Python-based evaluators using AWS Lambda to perform deterministic validation, business-rule enforcement, and workflow quality checks.
Define and implement evaluation standards, mandatory quality dimensions, scoring methodologies, and pass/fail criteria across AI platforms.
Design evaluation strategies at multiple levels, including TRACE, TOOL_CALL, and SESSION evaluation scopes.
Integrate evaluation workflows into CI/CD pipelines and establish automated deployment quality gates for AI-powered applications.
Leverage AWS AgentCore Evaluation capabilities to execute on-demand evaluations and support production quality monitoring.
Utilize observability data, OpenTelemetry traces, and AgentCore telemetry signals as evaluation inputs for quality assessment and root-cause analysis.
Collaborate with platform, security, and AI engineering teams to improve agent reliability, accuracy, and operational quality.
Analyze evaluation results, identify quality regressions, and drive corrective actions across models, prompts, tools, and workflows.
Define monitoring and reporting mechanisms for evaluation outcomes, quality trends, and operational KPIs.
Contribute to the evolution of enterprise AI governance, testing methodologies, and evaluation best practices.
What you need for this
Skills:
• AWS AgentCore Evaluation (on-demand mode for CI/CD gates, online mode for production sampling)
• LLM-as-judge evaluator design (built-in AgentCore evaluators — helpfulness, correctness)
• Custom code-based Lambda evaluators (Python — deterministic checks)
• Evaluation levels (TRACE for per-response, TOOL_CALL for per-invocation, SESSION for workflow)
• OTel spans from AWS AgentCore Observability as evaluation input
• Enterprise evaluation standard authoring (mandatory dimensions, pass/fail criteria)
Experience:
• 5+ years ML engineering or AI platform engineering
• LLM evaluation framework design and implementation
• Custom evaluator implementation for deterministic quality checks
• CI/CD deployment gate design for ML model or agent quality
Nice-to-have
• AWS AgentCore Evaluation API hands-on (CreateEvaluation, GetEvaluationResult)
• AWS Bedrock Guardrails for PII detection evaluator integration
• CloudWatch metrics output from AgentCore Evaluation for online mode
Desired Candidate Profile
We are seeking a highly motivated and experienced Senior ML / Evaluation Engineer. The ideal candidate will have a strong background in machine learning, statistical analysis, and software development. Responsibilities include designing and executing evaluation frameworks, identifying and mitigating biases, and collaborating with cross-functional teams to deliver high-quality ML models. A Master's or Ph.D. in Computer Science, Machine Learning, or a related field is preferred. Excellent communication and problem-solving skills are essential.
نحن نبحث عن مهندس سِنِيور في تعلم الآلة / تقييم للمساعدة في تعريف وتنفيذ معايير الجودة للوكالات المعتمدة على AI والتطبيقات المدعومة بـ LLM على مستوى المؤسسة. في هذا الدور، ستقوم بتصميم أطر التقييم، وبناء خطوط تقييم مخصصة، وتأسيس بوابات جودة آلية عبر دورة حياة تقديم الذكاء الاصطناعي. ستعمل عن كثب مع مهندسي منصة الذكاء الاصطناعي وفِرَق تعلم الآلة وفِرَق DevOps لضمان أنظمة ذكاء اصطناعي موثوقة وقابلة للقياس وجاهزة للإنتاج من خلال ممارسات تقييم ورصد وحوكمة قابلة للتوسع.
ما هو المشروع الذي لدينا لك
عميلنا شركة متعددة الجنسيات لديها أكثر من قرن من التاريخ ومكاتب في أكثر من 180 دولة. هدفهم الأكثر طموحاً في ذلك الوقت هو تقديم مجموعة من المنتجات منخفضة المخاطر (RRPs). الجمهور المستهدف هو أكثر من 1 مليار مستهلك حول العالم. يستضيف منصة تكنولوجيا المعلومات 700+ تطبيق.
مهمة Intellia هي مساعدة العميل في هندسة منظومة برامج شاملة لمنتج IoT يحركه ابتكار تجربة مستهلك و تقنيات متقدمة. فرقنا تشارك في هندسة مكونات منصة أساسية لحلول التجارة الإلكترونية الرائدة، والتسويق الرقمي وIoT. بصفته مهندساً، ستصبح جزءاً من فريق الهندسة الأساسية ومسؤولاً عن الهندسة وممارسة أفضل الممارسات في منصتنا الرقمية الهندسية المؤسسية.
المنصة هي مجموعة خدمات وتطبيقات على الإنترنت تسرع من تطوير وتقديم تطبيقات البرمجيات من خلال الاعتناء بتحديات SDLC الشائعة. توفر المنصة الوصول والاستهلاك لفرق الهندسة إلى مجموعة من الخدمات والتقنيات والممارسات لتطويرهم ولتشغيل تطبيقهم، مع ضمان مجموعة من الامتثال وأفضل الممارسات.
ماذا ستفعل
تصميم وتنفيذ وصيانة أطر تقييم من مستوى المؤسسة لـ LLMs، ووكلاء ذكاء اصطناعي، وتدفقات عمل AI متعددة الخطوات.
تطوير وتحسين مُقيمات LLM كقاضٍ لتقييم أبعاد مثل المساعدة، والصحة، والاتساق، والامتثال السياسي.
بناء مقيمين مخصصين قائمين على Python باستخدام AWS Lambda لإجراء تحقق حتمي، وفرض قواعد العمل، وفحوصات جودة التدفق.
تعريف وتنفيذ معايير التقييم، والأبعاد الجودة الإلزامية، ومنهجيات التقييم، ومعايير النجاح / الفشل عبر منصات الذكاء الاصطناعي.
تصميم استراتيجيات التقييم على مستويات متعددة، بما في ذلكTRACE، TOOL_CALL، ونطاقات التقييم SESSION.
دمج تدفقات التقييم في خطوط CI/CD وتأسيس بوابات جودة نشر آلية لتطبيقات مدعومة بالذكاء الاصطناعي.
استغلال قدرات تقييم AWS AgentCore لتنفيذ تقييمات عند الطلب ودعم مراقبة جودة الإنتاج.
استخدام بيانات الرصد وتتبع OpenTelemetry وإشارات قياس AgentCore كمدخلات تقييم لتقييم الجودة وتحليل السبب الجذري.
التعاون مع فرق المنصة والأمن والهندسة الذكية لتحسين موثوقية الوكلاء والدقة وجودة التشغيل.
تحليل نتائج التقييم، وتحديد تراجعات الجودة، وتوجيه الإجراءات التصحيحية عبر النماذج والتعليمات والأدوات وتدفقات العمل.
تعريف آليات الرصد والتقارير لنتائج التقييم، واتجاهات الجودة، ومؤشرات الأداء التشغيلية KPI.
المساهمة في تطور حوكمة الذكاء الاصطناعي المؤسسي، ومنهجيات الاختبار، وأفضل ممارسات التقييم.
ما تحتاجه لهذا الدور
المهارات:
• تقييم AWS AgentCore (وضع عند الطلب لبوابات CI/CD، وضع عبر الإنترنت لعينة الإنتاج)
• تصميم مُقيم لوكيل LLM-كقاضٍ (مقيمات AgentCore مدمجة — المساعدة، والصحة)
• مقيمين مخصصين قائمين على كود بايثون (Python — فحوص حتمية)
• مستويات التقييم (TRACE لاستجابة واحدة، TOOL_CALL للتنفيذ الواحد، SESSION لتدفق العمل)
• مقاطعات OTel من مراقبة AWS AgentCore كمدخل تقييم
• إعداد معيار التقييم المؤسسي (أبعاد إلزامية، معايير النجاح / الفشل)
الخبرة:
• 5+ سنوات في هندسة ML أو هندسة منصة AI
• تصميم وتنفيذ إطار تقييم لـ LLM
• تنفيذ مقيم مخصص لفحوص الجودة الحتمية
• تصميم بوابة نشر CI/CD لجودة نموذج ML أو وكيل
يُفضَّل وجود
• خبرة عملية في AWS AgentCore Evaluation API (CreateEvaluation, GetEvaluationResult)
• تكامل Guardrails من AWS Bedrock للكشف عن PII
• إخراج مقاييس CloudWatch من AgentCore Evaluation للوضع عبر الإنترنت
المرشح الوصف الأمثل
نبحث عن مهندس سِنِيور في تعلم الآلة / التقييم عالي الحافز وخبرة. سيكون المرشح المثالي لديه خلفية قوية في تعلم الآلة، التحليل الإحصائي، وتطوير البرمجيات. تشمل المسؤوليات تصميم وتنفيذ أطر التقييم، وتحديد والتحكم في التحيزات، والتعاون مع فرق متعددة الاختصاصات لتسليم نماذج ML عالية الجودة. يفضل أن يحمل المرشح MSc أو PhD في علوم الكمبيوتر أو تعلم الآلة أو مجال ذي صلة. مهارات التواصل وحل المشكلات الممتازة ضرورية.