Job description
At Globant, we are working to make the world a better place, one step at a time. We enhance business development and enterprise solutions to prepare them for a digital future. With a diverse and talented team present in more than 30 countries, we are strategic partners to leading global companies in their business process transformation.
Join Globant as an AI Test Automation Engineer and lead the quality transformation for our flagship agentic banking platform. We are looking for an automation expert skilled in Python, pytest, and LLM/Agentic system testing to solve one of tech's newest challenges: systematically validating non-deterministic AI behavior. In this role, you will build automated evaluation datasets, inspect agent execution traces, and deploy LLM-as-a-judge frameworks to gate production releases with absolute confidence.
SCREENING REQUIREMENTSIdeal candidates must meet all primary technical gates:
Automation Mastery: 4+ years of test automation experience using Python and pytest (custom fixtures, mocking, and CI pipeline integration).
LLM & Agentic System Testing: 1+ years of hands-on experience evaluating non-deterministic AI applications.
LLM Evaluation Engineering: Proven track record building LLM-as-a-judge scoring scripts and maintaining versioned evaluation datasets (golden paths, edge cases, regression suites).
Agent Deep-Dives & Traces: Ability to validate tool selection and parameters, inspect multi-step agent execution traces (e.g., LangGraph, Microsoft Agent Framework), and mock agent endpoints for CI workflows.
Adversarial & Guardrail Testing: Experience designing negative-path test suites for prompt injections, hallucinated financial actions, and trust-level boundaries (Suggest / Approve / Act).
Observability & CI Integration: Hands-on integration of test stages into GitLab CI (or equivalent) using LLM observability tools like OPIK or similar trace analysis platforms.
KEY RESPONSIBILITIESAgent Test Design & Automation: Design and implement deterministic pytest suites to validate agent behavior, intent handling, tool-call correctness, strict JSON schemas, and failure paths.
LLM Evaluation Engineering: Build and maintain LLM-as-a-judge evaluation scripts that score agent outputs for correctness, tone, and compliance, defining scoring thresholds and pass/fail release gates.
Mocking & Trace Analysis: Build mocks and fixtures for LLM endpoints and core APIs so agent loops run repeatably in CI pipelines. Execute, mock, and capture traces from LangGraph or Microsoft Agent Framework agent loops within test suites.
Guardrail & Security Testing: Continuously validate that agents refuse prompt injection attempts, do not hallucinate actions, and strictly follow compliance guardrails.
CI/CD Integration & Reporting: Integrate test suites into GitLab CI with DevOps, leverage OPIK for trace analysis, and deliver automated quality evidence gating weekly drops.
REQUIRED EXPERIENCE & QUALIFICATIONSTest Automation (Essential): 4+ years of test automation with strong Python and pytest. Experience testing APIs, contract testing, and strict JSON schema validation.
AI & LLM Testing (Essential): 1+ years of experience with behavioral acceptance criteria, evaluation datasets, and LLM-as-a-judge techniques. Practical understanding of managing non-determinism via statistical thresholds and automated judges.
Frameworks (Essential): Hands-on familiarity with agentic frameworks (LangGraph, Microsoft Agent Framework, LangChain, or LlamaIndex) sufficient to execute and mock agent loops.
Observability & Domain (Desirable): Experience with OPIK or equivalent trace analysis platforms. Prior experience testing in banking, fintech, or regulated financial environments.
- Knowledge of Arabic language is plus.
This job can be filled in Egypt/MENA Region.
Create with us digital products that people love. We will bring businesses and consumers together through AI technology and creativity, driving digital transformation to impact the world positively.
We may use AI and machine learning technologies in our recruitment process. Globant is an Equal Opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability, veteran status, or any other characteristic protected by applicable federal, state, or local law. Globant is also committed to providing reasonable accommodations for qualified individuals with disabilities in our job application procedures. If you need assistance or an accommodation due to a disability, please let your recruiter know.
Final compensation offered is based on multiple factors such as the specific role, hiring location, as well as individual skills, experience, and qualifications. In addition to competitive salaries, we offer a comprehensive benefits package. Learn more about life at Globant here: Globant Experience Guide.
وصف الوظيفة
في غلوبيانت، نعمل على جعل العالم مكانًا أفضل، خطوة تلو الأخرى. نعزز تطوير الأعمال وحلول المؤسسات لاستعدادها للمستقبل الرقمي. بفضل فريق متنوع وموهوب منتشر في أكثر من 30 دولة، نكون شركاء استراتيجيين للشركات العالمية الرائدة في تحول عمليات أعمالها.
انضم إلى غلوبيانت كمهندس اختبار الأتمتة بالذكاء الاصطناعي وقم بقيادة تحول الجودة لمنصتنا الرائدة في مجال الخدمات المصرفية الذكية. نحن نبحث عن خبير في الأتمتة متمرس في بايثون، pytest، واختبار أنظمة الوكلاء/الذكاء الاصطناعي لحل أحد التحديات التكنولوجية الجديدة: التحقق بشكل منهجي من سلوك الذكاء الاصطناعي غير الحتمي. في هذا الدور، ستقوم ببناء مجموعات بيانات تقييم آلية، وفحص آثار تنفيذ الوكلاء، ونشر أطر عمل تقييم النماذج اللغوية الكبيرة (LLM-as-a-judge) لضمان إطلاق الإصدارات الإنتاجية بثقة مطلقة.
المتطلبات الأساسية للفحص
يجب على المرشحين المثاليين تلبية جميع المتطلبات التقنية الأساسية التالية:
إتقان الأتمتة: 4 سنوات أو أكثر من الخبرة في اختبار الأتمتة باستخدام بايثون وpytest (الملحقات المخصصة، التزييف، ودمج خطوط أنابيب CI).
اختبار أنظمة الوكلاء والذكاء الاصطناعي: سنة أو أكثر من الخبرة العملية في تقييم تطبيقات الذكاء الاصطناعي غير الحتمية.
هندسة تقييم النماذج اللغوية الكبيرة: سجل حافل في بناء نصوص تقييم LLM-as-a-judge وصيانة مجموعات بيانات التقييمVersioned (المسارات الذهبية، الحالات الطرفية، مجموعات الاختبار التراجعي).
التعمق في الوكلاء والآثار: القدرة على التحقق من اختيار الأدوات والمعلمات، وفحص آثار تنفيذ الوكلاء متعددة الخطوات (مثل LangGraph، Microsoft Agent Framework)، وتزييف نقاط نهاية الوكلاء لخطوط عمل CI.
اختبارات Adversarial والحماية: خبرة في تصميم مجموعات اختبار للمسارات السلبية لاختراقات الأوامر، والأفعال المالية المتخيلة، وحدود الثقة (الاقتراح / الموافقة / التنفيذ).
المراقبة ودمج CI: خبرة عملية في دمج مراحل الاختبار في GitLab CI (أو ما يعادلها) باستخدام أدوات مراقبة النماذج اللغوية الكبيرة مثل OPIK أو منصات تحليل الآثار المماثلة.
المسؤوليات الرئيسية
تصميم واختبار الوكلاء: تصميم وتنفيذ مجموعات اختبار pytest محددة لتحديد سلوك الوكلاء، ومعالجة النوايا، وصحة استدعاءات الأدوات، ومخططات JSON الصارمة، ومسارات الفشل.
هندسة تقييم النماذج اللغوية الكبيرة: بناء وصيانة نصوص تقييم LLM-as-a-judge التي تقيم مخرجات الوكلاء من حيث الصحة، والنبرة، والامتثال، وتحديد عتبات التقييم وقواعد إطلاق الإصدارات (Pass/Fail).
التزييف وتحليل الآثار: بناء التزييفات والملحقات للنقاط النهائية للنماذج اللغوية الكبيرة وواجهات API الأساسية بحيث تعمل حلقات الوكلاء بشكل متكرر في خطوط أنابيب CI. تنفيذ، وتزييف، والتقاط آثار حلقات الوكلاء من LangGraph أو Microsoft Agent Framework ضمن مجموعات الاختبار.
اختبارات الحماية والأمان: التحقق باستمرار من أن الوكلاء يرفضون محاولات اختراق الأوامر، ولا يتخيلون أفعالًا، ويلتزمون بصرامة بسياسات الحماية.
دمج CI/CD والتقارير: دمج مجموعات الاختبار في GitLab CI مع DevOps، واستخدام OPIK لتحليل الآثار، وتقديم أدلة جودة آلية لضمان إطلاق الإصدارات الأسبوعية.
الخبرة والمؤهلات المطلوبة
اختبار الأتمتة (أساسي): 4 سنوات أو أكثر من اختبار الأتمتة مع إتقان قوي لبايثون وpytest. خبرة في اختبار واجهات API، اختبار العقود، والتحقق من مخططات JSON الصارمة.
اختبار الذكاء الاصطناعي والنماذج اللغوية الكبيرة (أساسي): سنة أو أكثر من الخبرة في معايير القبول السلوكي، مجموعات بيانات التقييم، وتقنيات LLM-as-a-judge. فهم عملي لإدارة عدم الحتمية من خلال العتبات الإحصائية والمُقيّمين الآليين.
الأطر (أساسي): دراية عملية بإطارات الوكلاء (LangGraph، Microsoft Agent Framework، LangChain، أو LlamaIndex) كافية لتنفيذ وتزييف حلقات الوكلاء.
المراقبة والمجال (مرغوب فيه): خبرة مع OPIK أو منصات تحليل الآثار المماثلة. خبرة سابقة في اختبار الأنظمة المصرفية أو المالية أو البيئات المالية المنظمة.
- معرفة اللغة العربية هو ميزة إضافية.
يمكن شغل هذه الوظيفة في مصر/منطقة الشرق الأوسط وشمال أفريقيا.
اصنع معنا المنتجات الرقمية التي يحبها الناس. سنجمع بين الأعمال والمستهلكين من خلال تقنية الذكاء الاصطناعي والإبداع، مما يدفع التحول الرقمي لتحقيق تأثير إيجابي في العالم.
قد نستخدم تقنيات الذكاء الاصطناعي وتعلم الآلة في عملية التوظيف. غلوبيانت هو صاحب عمل متساوٍ في الفرص. جميع المتقدمين المؤهلين سيحصلون على فرصة متساوية للوظيفة دون تمييز بسبب العرق أو اللون أو الدين أو الجنس أو الجنسية أو الإعاقة أو أي سمة أخرى محمية بموجب القانون الفيدرالي أو المحلي أو قانون الولاية. تلتزم غلوبيانت أيضًا بتقديم تسهيلات معقولة للمتقدمين المؤهلين ذوي الإعاقة في إجراءات تقديم طلبات التوظيف الخاصة بنا. إذا كنت بحاجة إلى مساعدة أو تسهيل بسبب إعاقة، يرجى إبلاغ مسؤول التوظيف لدينا.
يتم تحديد compensation النهائي المقدم بناءً على عوامل متعددة مثل الدور المحدد، وموقع التوظيف، بالإضافة إلى المهارات والخبرة والمؤهلات الفردية. بالإضافة إلى الرواتب التنافسية، نقدم حزمة شاملة من المزايا. تعرف على المزيد حول الحياة في غلوبيانت هنا: دليل تجربة غلوبيانت.