Mindrift يربط المتخصصين بفرص الذكاء الاصطناعي القائمة على المشاريع لشركات تقنية رائدة، مع التركيز على الاختبار والتقييم وتحسين أنظمة الذكاء الاصطناعي. المشاركة قائمة على المشاريع، وليست توظيفاً دائماً.
ما ينطوي عليه هذا الفرصة
نحن نبني مجموعة بيانات لتقييم عوامل البرمجة الذكية - مدى قدرة النموذج على التعامل مع مهام المطورين الواقعية. ستقوم بإنشاء مهام صعبة ومعايير تقييم ضمن بيئات محاكاة واقعية:
- بناء بيئات مطور واقعية - شركة افتراضية تحتوي على قاعدة كود، وبنية تحتية، وسياق (تذاكر، وثائق، محادثات) تشكّل تاريخ تطوير قابل للاقتناع
- تصميم المهام من حالات وسيطة لهذه البيئات - صِغ السؤال، عرف ما الذي يعني "محلول"، وتأكد من أن المهمة قابلة للحل بواسطة وكيل ذكاء اصطناعي
- كتابة اختبارات تتحقق من حلول الوكلاء - تقبل جميع الأساليب الصحيحة وترفض غير الصحيحة، لا تشدد مفرطاً ولا تترك كثيراً
- التكرار على المهام والاختبارات بناءً على ملاحظات QA - راجع حلول الوكلاء، حلل الإخفاقات، وقم بتحسينها حتى تكون التقييم عادلًا ومتوازنًا
ما هذا ليس
- ليس تصنيف بيانات
- ليس هندسة تمرير الاستعلامات
- ليس كتابة كود من الصفر - الوكيل يكتب معظم الكود؛ أنت توجه وتقييم
ما نبحث عنه
- أكثر من 5 سنوات في تطوير البرمجيات
- التكدس الأساسي: بايثون (FastAPI)، JavaScript/TypeScript (React)، Docker، Postgres، Kafka، Redis
- خبرة في كتابة الاختبارات (وظيفية، تكاملية)
- إتقان اللغة الإنجليزية - B2+
المرشح المثالي
- أكثر من 5 سنوات في تطوير البرمجيات
- التكدس الأساسي: بايثون (FastAPI)، JavaScript/TypeScript (React)، Docker، Postgres، Kafka، Redis
- خبرة في كتابة الاختبارات (وظيفية، تكاملية)
- إتقان الإنجليزية - B2+
Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.
What this opportunity involves
We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria within realistic simulated environments:
- Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
- Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent
- Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
- Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust
What this is NOT
- Not data labeling
- Not prompt engineering
- Not writing code from scratch - the agent writes most of the code; you guide and evaluate
What we look for
- 5+ years in software development
- Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
- Experience writing tests (functional, integration)
- English proficiency - B2+
Desired Candidate Profile
- 5+ years in software development
- Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
- Experience writing tests (functional, integration)
- English proficiency - B2+