نبحث عن مهندس بايثون – مكتبة مُقَيِّمات لتصميم وتنفيذ مكوِّنات تقييم قابلة لإعادة الاستخدام تضمن جودة سلامة والتزام وكلاء الذكاء الاصطناعي المؤسسي وتدفقات العمل المدعومة بـ LLM. ستبني قدرات تقييم مخصصة تُستخدم عبر منصات الذكاء الاصطناعي، مع التركيز على التحقق الآلي من الجودة والتوافق في تدفقات العمل وحماية البيانات الشخصية والتحقق من الإخراج المُهيكل. بالعمل عن كثب مع مهندسي منصة الذكاء الاصطناعي، ومهندسي ML، وفرق DevOps، ستساهم في وضع معايير تقييم موثوقة وآليات ضمان جودة قابلة للتوسع للنُظم المعتمدة على الوكلاء.
ما المشروع الذي لدينا لك
عميلنا شركة متعددة الجنسيات تمتد تاريخها لأكثر من قرن ومكاتب في أكثر من 180 دولة. هدفهم الأكثر طموحًا في ذلك الوقت هو تقديم مجموعة من المنتجات منخفضة المخاطر (RRPs). الجمهور المستهدف أكثر من 1 مليار مستهلك حول العالم. يستضيف منصة تكنولوجيا المعلومات 700+ تطبيق.
مهمة Intellia هي مساعدة العميل في هندسة منظومة برمجية شاملة لمنتج IoT رائد على هامش تجربة مستهلك مبتكرة وتكنولوجيا متقدمة. فرقنا معنية بهندسة مكوّنات المنصة الأساسية لحلول التجارة الإلكترونية الرائدة والتسويق الرقمي وIoT. كمهندس، ستصبح جزءًا من فريق الهندسة المعمارية الأساسية وتكون مسؤولاً عن التصميم والتنفيذ لأفضل الممارسات في منصتنا المؤسسية للهندسة الرقمية.
المنصة هي مجموعة خدمات وتطبيقات إنترنت تسرّع تطوير وتسليم تطبيقات البرمجيات من خلال العناية بتحديات دورة حياة تطوير البرمجيات (SDLC) الشائعة. تقدم المنصة وصولاً واستهلاكاً لفرق الهندسة لمجموعة من الخدمات والتقنيات والممارسات لتطويرها وتشغيل تطبيقاتهم، مع ضمان الالتزام وأفضل الممارسات.
ما ستفعله
تصميم وتطوير وصيانة مكتبات تقييم باياثونية قابلة لإعادة الاستخدام للوكلاء الذكاء الاصطناعي وتدفقات العمل المدعومة بـ LLM.
تنفيذ مقيمي AWS Lambda المخصصين لإجراء فحوصات جودة والتوافق والتحقق determinate.
تطوير منطق تقييم LLM كقاضٍ لتقييم أبعاد ذاتية مثل الملاءمة والمساعدة والاتساق وجودة الاستجابة.
بناء مقيمي كشف PII آليين باستخدام تقنيات regex وتكاملات AWS Bedrock Guardrails.
تنفيذ آليات تحقق من صحة مستوى TOOL_CALL للتحقق من الإخراجات المهيكلة، والتوافق مع مخطط JSON، وصحة استجابة الأداة.
تطوير مقيمين على مستوى SESSION للتحقق من الامتثال لعقد العمل، ونزاهة التنفيذ، وتوقعات السلوك عبر الخطوات.
إنشاء مقيمين على مستوى TRACE للتحقق من الدقة الرقمية، واتساق الحسابات، والتحقق من النتائج الحتمية.
دمج مكوّنات المقيم مع سير عمل تقييم AWS AgentCore وخطوط جودة الذكاء الاصطناعي المؤسسية.
التعاون مع فرق منصة الذكاء الاصطناعي لتحديد معايير التقييم وطرق التقييم ومعايير قبول الجودة.
تصميم وصيانة اختبارات الوحدة والتكامل والتحقق للمكتبات التقييمية وإطارات الجودة.
دعم الرصد والتشخيص من خلال دمج نتائج التقييم مع تسجيلات CloudWatch وإمكانيات المراقبة.
المساهمة في مبادرات حوكمة الذكاء الاصطناعي المؤسسي من خلال تحسين تغطية التقييم، والتدقيق، وضوابط الامتثال.
ما تحتاجه لهذه الوظيفة
المهارات:
• بايثون (دوال Lambda كقيم تقييم مخصصة مبنية على كود AgentCore من AWS)
• هندسة إنتاج prompts للـ LLM كقاضٍ لأبعاد التقييم الذاتية
• كشف PII (مبني على regex + Guardrails من AWS Bedrock)
• تحقق من صحة مخطط استجابة الأداة (JSON Schema — مستوى TOOL_CALL)
• فحص امتثال عقد العمل على مستوى SESSION
• منطق تحقق من الدقة الرقمية (TRACE مستوى)
الخبرة:
• 4+ سنوات من هندسة بايثون
• تقييم LLM أو ضمان جودة لأنظمة AI/ML
• تطوير ونشر وظائف Lambda من AWS
من المفضل وجوده
• تسجيل مقيم مخصص لمقيم AWS AgentCore
• تكامل Guardrails لـ AWS Bedrock لاكتشاف PII
• CloudWatch Logs كمصب لإخراج المقيم
الملف المرشح المطلوب
We are looking for a Python Engineer – Evaluator Library to design and implement reusable evaluation components that ensure the quality, safety, and compliance of enterprise AI agents and LLM-powered workflows. You will build custom evaluation capabilities used across AI platforms, focusing on automated quality validation, workflow compliance, PII protection, and structured output verification. Working closely with AI Platform Engineers, ML Engineers, and DevOps teams, you will help establish reliable evaluation standards and scalable quality assurance mechanisms for agent-based systems.
What project we have for you
Our customer is a multinational corporation with more than a century of history and offices in over 180 countries. Their most ambitious goal at the time is to introduce a range of Reduced-Risk Products (RRPs). The target audience is more than 1 billion consumers around the globe. IT platform hosts 700+ applications.
Intellia’s mission is to help the client with the engineering of a comprehensive software ecosystem for a game-changing IoT product on the margin of innovative consumer experience and cutting-edge technology. Our teams are involved in the engineering of core platform components for best-in-class eCommerce, Digital Marketing and IoT solutions. As an Engineer, you will become a part of Core Architecture Team and be responsible for the architecture, implementation of best practices in our Digital Engineering Enterprise Platform.
The Platform is a set of services and internet applications that accelerate the development and delivery of software applications by taking care of common SDLC challenges. The Platform provides access and consumption for engineering teams to a set of services, technologies, practices for their development and for operating their application, ensuring a set of compliance and best practices.
What you will do
Design, develop, and maintain reusable Python-based evaluator libraries for AI agents and LLM-powered workflows.
Implement AWS Lambda-based custom evaluators to perform deterministic quality, compliance, and validation checks.
Develop LLM-as-a-judge evaluation logic to assess subjective dimensions such as relevance, helpfulness, consistency, and response quality.
Build automated PII detection evaluators using regex-based techniques and AWS Bedrock Guardrails integrations.
Implement TOOL_CALL-level validation mechanisms to verify structured outputs, JSON schema compliance, and tool response correctness.
Develop SESSION-level evaluators to validate workflow contract compliance, execution integrity, and cross-step behavioral expectations.
Create TRACE-level evaluators for numerical accuracy verification, calculation consistency, and deterministic result validation.
Integrate evaluator components with AWS AgentCore Evaluation workflows and enterprise AI quality pipelines.
Collaborate with AI platform teams to define evaluation standards, scoring methodologies, and quality acceptance criteria.
Design and maintain unit, integration, and validation tests for evaluator libraries and quality frameworks.
Support observability and troubleshooting by integrating evaluation outputs with CloudWatch logging and monitoring capabilities.
Contribute to enterprise AI governance initiatives by improving evaluation coverage, auditability, and compliance controls.
What you need for this
Skills:
• Python (Lambda functions as AWS AgentCore custom code-based evaluators)
• LLM-as-judge prompt engineering for subjective evaluation dimensions
• PII detection (regex-based + AWS Bedrock Guardrails)
• Tool response schema validation (JSON Schema — TOOL_CALL level evaluator)
• Workflow contract compliance checking (SESSION level evaluator)
• Numerical accuracy validation logic (TRACE level evaluator)
Experience:
• 4+ years Python engineering
• LLM evaluation or quality assurance for AI/ML systems
• AWS Lambda function development and deployment
Nice-to-have
• AWS AgentCore Evaluation custom evaluator Lambda registration
• AWS Bedrock Guardrails for PII detection integration
• CloudWatch Logs as evaluator output sink
Desired Candidate Profile