GOSST T&L - دور AI/ML من أجل دقة LLM وتطوير النماذج وصف العمل GOSST Turnover & Launch يبحث عن مهندس ML/AI عملي (متعاقد، مكافئ L5) للانضمام إلى فريقنا وتحفيز تحسينات دقة LLM عبر محفظة منتجاتنا. يقع هذا الدور عند تقاطع التعلم الآلي التطبيقي وتكنولوجيا التشغيل - ستقوم بضبط النماذج، وبناء أُطر تقييم، وتحسين الأتمتة المستندة إلى الذكاء الاصطناعي لأدوات Turnover & Launch (OrderPad، UTP، Field Installation، وأنظمة مساعدة ذات طابع وكيل).
المسؤوليات الرئيسية:
• ضبط وتحسين نماذج اللغة الكبيرة (LLMs) للمهام المرتبطة بالمجال بما في ذلك تقدير الشفرة، تخطيط السبرينت، والأتمتة التشغيلية.
• تصميم وتنفيذ أُطر تقييم لقياس جودة النموذج — بما في ذلك المقاييس الآلية، LLM-asjudge، وتدفقات التقييم البشري.
• بناء وصيانة خطوط RAG (التوليد المستند إلى الاسترجاع) التي تُثبت مخرجات النموذج في قاعدة الشفرة والبيانات التشغيلية.
• تطوير استراتيجيات هندسة الاستفهام وتحسين سلاسل الاستفهام لعمليات وكيلة متعددة الخطوات.
• التعاون مع SDEs و SDMs لدمج تحسينات النموذج في خدمات الإنتاج (Amazon Bedrock، SageMaker، أو ما يعادله).
• تحليل أوضاع فشل النموذج، وتحديد فجوات الدقة، واقتراح تحسينات مستهدفة.
• وضع مقاييس أساسية وتتبع تحسينات الدقة بمرور الوقت من خلال لوحات التحكم والتقارير.
• مواكبة التقدّمات في LLM وتوصية باعتماد تقنيات جديدة (LoRA، DPO، الذكاء الاصطناعي الدستوري، إلخ) حيثما ينطبق.
GOSST T&L - AI/ML Role for LLM Accuracy & Model DevelopmentJob Description GOSST Turnover & Launch is looking for a hands-on ML/AI engineer (contractor, L5 equivalent) to join our team and drive LLM accuracy improvements across our product portfolio. This role sits at the intersection of applied machine learning and operational technology - you will fine-tune models, build evaluation frameworks, and improve AI-driven automation for our Turnover & Launch tools (OrderPad, UTP, Field Installation, and supporting agentic systems).
Key Responsibilities:
• Fine-tune and optimize Large Language Models (LLMs) for domain-specific tasks including code estimation, sprint planning, and operational automation.
• Design and implement evaluation frameworks to measure model quality — including automated metrics, LLM-asjudge, and human evaluation workflows.
• Build and maintain RAG (Retrieval-Augmented Generation) pipelines that ground model outputs in codebase and operational data.
• Develop prompt engineering strategies and optimize prompt chains for multi-step agentic workflows.
• Collaborate with SDEs and SDMs to integrate model improvements into production services (Amazon Bedrock, SageMaker, or equivalent).
• Analyze model failure modes, identify accuracy gaps, and propose targeted improvements.
• Establish baseline metrics and track accuracy improvements over time with dashboards and reporting.
• Stay current on LLM advancements and recommend adoption of new techniques (LoRA, DPO, constitutional AI, etc.) where applicable.