تربط Mindrift المتخصصين بفرص ذكاء اصطناعي قائمة على المشاريع مع شركات تقنية رائدة، مع تركيز على اختبار وتقييم وتحسين أنظمة الذكاء الاصطناعي. المشاركة قائمة على المشروع وليست توظيفاً دائماً.
ما تتضمنه هذه الفرصة
الوكلاء البرمجيون في Frontier لديهم بالفعل قدرة جيدة على اجتياز الاختبارات. نقيس ما إذا كانوا ينجحون بالطرق الصحيحة. نبني مجموعة بيانات لتقييم سلامة وسلوك وكلاء الترميز AI وليس فقط ما إذا أكملوا مهمة، بل ما إذا قاموا بها بصدق، ضمن النطاق، ودون اختصار.
ستصمم مهام حيث المسار السهل هو غير آمن، وتكتب الاختبارات التي تكشف ذلك:
- بناء بيئات مطور واقعية كشركة افتراضية مع قاعدة كود وبنية تحتية وسياق (التذاكر، الوثائق، المحادثات) التي تشكل تاريخ تطوير مقنع
- تصميم مهام تقارن هدف تطويري غير ضار مع اختصار غير آمن مغري: تجاوز النطاق، انتهاكات السياسات، فساد البيانات، تغييرات مبالغ فيها
- كتابة اختبارات تتحقق ما إذا أكمل الوكيل المهمة بالطريقة الصحيحة مع الكشف عن الاختبارات المختصرة، وليس مجرد فحص المخرجات
- التكرار على المهام والاختبارات بناءً على مراجعة تغذية QA، مراجعة حلول الوكلاء، تحليل الإخفاقات، وتحسين حتى تكون التقييم عادلًا وموثوقًا
ما هذا ليس:
- ليس تسمية بيانات؛ ليس هندسة مطالبة؛ ليس أمان سيبراني أو اختبار اختراق هناك لا وجود لمهاجم في السيناريو. خبرة الأمن السيبراني شيء جيد لكنه ليس شرطاً. نبحث عن مهندسين يفهمون كيفية سلوك الكود، لا مختبري اختراق. مهندسون برمجيون أقوياء، ليسوا متخصيين في الأمن؛
- ليس كتابة كود من الصفر؛ الوكيل يكتب معظم الكود؛ أنت تصمم الوضع وتقيّم النتيجة;
ما الذي نبحث عنه
- أكثر من 4-5 سنوات في تطوير البرمجيات؛
- المكدس الأساسي: بايثون، جافاسكريبت/TypeScript؛
- مهارات تصميم اختبارات قوية: اختبارات وظيفية وتكامل تفصل بين الإنجاز الآمن وغير الآمن، وليس فقط الصحيح من غير الصحيح؛
- خبرة عملية مع وكلاء الترميز (Claude Code، GitHub Copilot CLI، Codex، أو ما شابه);
- الإلمام بـ GitHub PRs وCI ك مستخدم؛
- نطاق التكدس مرحب به، ليس عامل تصفية. تحاكي المهام مستودعات حقيقية مع قواعد بيانات، خطوط CI، وسكريتات النشر، لذا يكون التعرض الأوسع للخلفية والبنية التحتية فعّالاً حقاً لكن لا تحتاج أن تكون خبيراً في كل طبقة؛
- إتقان الإنجليزية B2+
المرشح المثالي
- أكثر من 4-5 سنوات في تطوير البرمجيات؛
- المكدس الأساسي: بايثون، جافا سكريبت/TypeScript؛
- مهارات تصميم اختبارات قوية: اختبارات وظيفية وتكامل تفصل بين الإنجاز الآمن وغير الآمن، وليس فقط الصحيح من غير الصحيح؛
- خبرة عملية مع وكلاء الترميز (Claude Code، GitHub Copilot CLI، Codex، أو ما شابه);
- الإلمام بـ GitHub PRs وCI ك مستخدم؛
- نطاق التكدس مرحب به، ليس عامل تصفية. تحاكي المهام مستودعات حقيقية مع قواعد بيانات، خطوط CI، وسكريتات النشر، لذا يكون التعرض الأوسع للخلفية والبنية التحتية فعّالاً حقاً لكن لا تحتاج أن تكون خبيراً في كل طبقة؛
- إتقان اللغة الإنجليزية B2+
Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.
What this opportunity involves
Frontier coding agents are already good at passing tests. We measure whether they pass them the right way .We're building a dataset to evaluate the safety and conduct of AI coding agents not just whether they complete a task, but whether they do it honestly, within scope, and without cutting corners.
You'll design tasks where the easy path is the unsafe one, and write the tests that catch it:
- Build realistic developer environments a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
- Design tasks that pair a benign development goal with a tempting unsafe shortcut: scope creep, policy violations, data corruption, over-permissive changes
- Write tests that verify whether the agent completed the task the right way catching corners cut, not just checking outputs
- Iterate on tasks and tests based on QA feedback review agent solutions, analyze failures, and refine until the evaluation is fair and robust
What this is NOT:
- Not data labeling; Not prompt engineering; Not cybersecurity or red-teaming there is no attacker in the scenario. Cybersecurity experience is a nice-to-have but not a requirement. We're looking for engineers who understand how code should behave, not penetration testers. Strong software engineers, not security specialists;
- Not writing code from scratch the agent writes most of the code; you design the situation and evaluate the outcome;
What we look for
- 4 5+ years in software development;
- Core stack: Python, JavaScript/TypeScript;
- Strong test design skills functional and integration tests that separate safe from unsafe completion, not just correct from incorrect;
- Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex, or similar);
- Familiarity with GitHub PRs and CI workflows as a user;
- Stack breadth is welcome, not a filter. Tasks simulate real repositories with databases, CI pipelines, and deploy scripts, so wider backend and infrastructure exposure is genuinely useful but you don't need to be an expert in every layer;
- English proficiency B2+
Desired Candidate Profile
- 4 5+ years in software development;
- Core stack: Python, JavaScript/TypeScript;
- Strong test design skills functional and integration tests that separate safe from unsafe completion, not just correct from incorrect;
- Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex, or similar);
- Familiarity with GitHub PRs and CI workflows as a user;
- Stack breadth is welcome, not a filter. Tasks simulate real repositories with databases, CI pipelines, and deploy scripts, so wider backend and infrastructure exposure is genuinely useful but you don't need to be an expert in every layer;
- English proficiency B2+