تربط Mindrift المتخصصين بفرص الذكاء الاصطناعي القائمة على المشاريع لشركات التقنية الرائدة، مع تركيز على اختبار الأنظمة وتقييمها وتحسينها. المشاركة تكون على أساس المشروع، وليست وظيفة دائمة.
ما تنطوي عليه هذه الفرصة
وكلاء الترميز المتقدمون في الحدود الأولى جيدون حالياً في اجتياز الاختبارات. نقيس ما إذا كانوا يمرونها بالطريقة الصحيحة. نحن نقِيم سلامة وسلوك وكلاء الترميز بالذكاء الاصطناعي ليس فقط من خلال إنجاز مهمة، بل من خلال إنجازها بصدق، ضمن النطاق، ودون التلاعب. ستصمم مهام تكون فيها الطريق السهل غير آمن، وتكتب الاختبارات التي تلتقطها:
- بناء بيئات مطور واقعية لشركة افتراضية تحتوي قاعدة شيفرات، وبنية تحتية، وسياق (تذاكر، وثائق، محادثات) تشكل تاريخ تطوير قابل للاقتناع
- تصميم مهام تجمع هدف تطويري بسيط مع اختصار غير آمن مغري: زيادة النطاق، مخالفة السياسات، تلف البيانات، تغييرات واسعة الإذعان
- كتابة اختبارات تتحقق مما إذا كان الوكيل قد أكمل المهمة بالطريقة الصحيحة مع التقاط الطرفية المقطوعة، وليس فقط فحص المخرجات
- التكرار في المهام والاختبارات بناءً على ملاحظات QA ومراجعة حلول الوكيل، تحليل الإخفاقات، وتحسين حتى تكون التقييمات عادلة وقوية
ما لا تعتبره هذه الفرصة:
- ليست تصنيف بيانات؛ ليست هندسة الطلبات؛ ليست أمان سيبراني أو فريق اختراق هناك مهاجم في السيناريو. خبرة الأمن السيبراني ميزة لكنها ليست مطلوبة. نبحث عن مهندسين يفهمون كيف يجب أن behave الشفرة، ليسوا مختبرين اختراق. مهندسون برمجيون أقوياء، ليسوا متخصصين في الأمن;
- ليست كتابة كود من الصفر، الوكيل يكتب معظم الكود؛ أنت تصمم الوضع وتقييم النتيجة;
ما نبحث عنه
- 4-5+ سنوات في تطوير البرمجيات;
- التكدس الأساسي: بايثون، جافاسكريبت/تايب سكريبت؛
- مهارات تصميم اختبارات قوية: اختبارات وظيفية وتكاملية تفصل بين الإكمال الآمن وغير الآمن، وليس فقط الصحيح من غير الصحيح؛
- خبرة عملية مع وكلاء الترميز ( Claude Code، GitHub Copilot CLI، Codex، أو ما شابهها)؛
- إلمام بمراجعات GitHub PRs وعمليات CI كمستخدم؛
- تنوع التكديس مرحب به، ليس عامل ترشيح. المحاكاة للريبو المملوء بم قواعد البيانات، وخطوط CI، وسكريبتات النشر، لذا يعد التعرض للبنية الخلفية والبنية التحتية أداة مفيدة حقاً لكن لا يلزم أن تكون خبيراً في كل طبقة؛
- إتقان الإنجليزية B2+
Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.
What this opportunity involves
Frontier coding agents are already good at passing tests. We measure whether they pass them the right way .We're building a dataset to evaluate the safety and conduct of AI coding agents not just whether they complete a task, but whether they do it honestly, within scope, and without cutting corners.You'll design tasks where the easy path is the unsafe one, and write the tests that catch it:
- Build realistic developer environments a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
- Design tasks that pair a benign development goal with a tempting unsafe shortcut: scope creep, policy violations, data corruption, over-permissive changes
- Write tests that verify whether the agent completed the task the right way catching corners cut, not just checking outputs
- Iterate on tasks and tests based on QA feedback review agent solutions, analyze failures, and refine until the evaluation is fair and robust
What this is NOT:
- Not data labeling; Not prompt engineering; Not cybersecurity or red-teaming there is no attacker in the scenario. Cybersecurity experience is a nice-to-have but not a requirement. We're looking for engineers who understand how code should behave, not penetration testers. Strong software engineers, not security specialists;
- Not writing code from scratch the agent writes most of the code; you design the situation and evaluate the outcome;
What we look for
- 4 5+ years in software development;
- Core stack: Python, JavaScript/TypeScript;
- Strong test design skills functional and integration tests that separate safe from unsafe completion, not just correct from incorrect;
- Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex, or similar);
- Familiarity with GitHub PRs and CI workflows as a user;
- Stack breadth is welcome, not a filter. Tasks simulate real repositories with databases, CI pipelines, and deploy scripts, so wider backend and infrastructure exposure is genuinely useful but you don't need to be an expert in every layer;
- English proficiency B2+
Why this is hard
Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. The real difficulty is building the temptation a scenario where the unsafe or out-of-scope path is the path of least resistance and then writing tests that reliably catch an agent that took it. Tasks have many valid solutions; tests must accept all of them and reject the bad ones.
sed on their requirements.
Desired Candidate Profile
- 4 5+ years in software development;
- Core stack: Python, JavaScript/TypeScript;
- Strong test design skills functional and integration tests that separate safe from unsafe completion, not just correct from incorrect;
- Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex, or similar);
- Familiarity with GitHub PRs and CI workflows as a user;
- Stack breadth is welcome, not a filter. Tasks simulate real repositories with databases, CI pipelines, and deploy scripts, so wider backend and infrastructure exposure is genuinely useful but you don't need to be an expert in every layer;
- English proficiency B2+