نحن نبني مجموعة بيانات لتقييم وكلاء ترميز الذكاء الاصطناعي - مدى نجاح النموذج في التعامل مع مهام المطورين الواقعية. ستقوم بإنشاء مهام وتقييمات ضمن بيئات محاكاة واقعية:
إنشاء بيئات مطورين واقعية - شركة افتراضية تحتوي على قاعدة كود وبنية تحتية وسياق (التذاكر، الوثائق، المحادثات) التي تشكل تاريخ تطوير مقنع
تصميم المهام من حالات وسطية لهذه البيئات - صغ السؤال، عرّف معنى "المحلول"، وتأكد من أن المهمة قابلة للحل بواسطة وكيل AI
اكتب اختبارات تتحقق من حلول الوكيل - تقبل جميع النهج الصحيحة وترفض الخاطئة، لا تكون صارمة جدًا ولا متساهلة جدًا
التكرار على المهام والاختبارات بناءً على ملاحظات QA - مراجعة حلول الوكيل، تحليل الإخفاقات، وتحسين حتى تكون التقييم عادلًا وقويًا
ما الذي ليس كذلك
ليس تسمية بيانات
ليس هندسة المطالبات
ليس كتابة رمز من الصفر - الوكيل يكتب معظم الرمز؛ أنت توجه وتقييم
ما نبحث عنه
5+ سنوات في تطوير البرمجيات
المكدس الأساسي: Python (FastAPI)، JavaScript/TypeScript (React)، Docker، Postgres، Kafka، Redis
خبرة في كتابة الاختبارات (وظيفية، تكامل)
إتقان الإنجليزية - B2+
لماذا هذا صعب
النماذج الحدّية جيدة بالفعل في الترميز. صناعة مهمة تتحدى النماذج الأفضل حقًا ليست بسيطة. تحتاج فهمًا عميقًا للأماكن التي تفشل فيها النماذج وما هي السيناريوهات التي تكشف الفرق بين حل جيد وآخر سيئ.
المهام لها حلول كثيرة صحيحة - كتابة اختبارات تقبل جميع الحلول الصحيحة وترفض الخاطئة أصعب مما يبدو.
كيف يعمل الأمر
قدم الطلب
اجتز المؤهل/ات
انضم إلى مشروع
أنجز المهام
احصل على الأجر
التعويض
حتى 40 دولار/ساعة مكافئة، اعتمادًا على المستوى والإيقاع. يتم تقدير المهام بحوالي 20 ساعة لكل منها؛ أنت تحدد جدولك الخاص.
الملف المرغوب للمرشح
- 5+ سنوات في تطوير البرمجيات
- المكدس الأساسي: Python (FastAPI)، JavaScript/TypeScript (React)، Docker، Postgres، Kafka، Redis
- خبرة في كتابة الاختبارات (وظيفية، تكامل)
- إتقان الإنجليزية - B2+
We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria within realistic simulated environments:
Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent
Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust
What this is NOT
Not data labeling
Not prompt engineering
Not writing code from scratch - the agent writes most of the code; you guide and evaluate
What we look for
5+ years in software development
Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
Experience writing tests (functional, integration)
English proficiency - B2+
Why this is hard
Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.
Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.
How it works
Apply
Pass qualification(s)
Join a project
Complete tasks
Get paid
Compensation
Up to $40/hr equivalent , depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.
Desired Candidate Profile
- 5+ years in software development
- Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
- Experience writing tests (functional, integration)
- English proficiency - B2+