هجين دوام كامل
--
Intellias

تفاصيل الوظيفة

About the Role We are looking for a Python Engineer – Evaluator Library to design and implement reusable evaluation components that ensure the quality, safety, and compliance of enterprise AI agents and LLM-powered workflows. In this role, you will build custom evaluation capabilities used across AI platforms, focusing on automated quality validation, workflow compliance, PII protection, structured output verification, and deterministic evaluation logic. Working closely with AI Platform Engineers, ML Engineers, and Dev Ops teams, you will help establish scalable evaluation standards and quality assurance mechanisms for enterprise AI systems.
Responsibilities Design, develop, and maintain reusable Python-based evaluator libraries for AI agents and LLM-powered workflows. Build AWS Lambda-based custom evaluators that perform deterministic quality, compliance, and validation checks. Develop LLM-as-a-Judge evaluation logic to assess subjective quality dimensions, including relevance, helpfulness, consistency, and response quality. Implement automated PII detection using regex-based techniques and AWS Bedrock Guardrails. Develop TOOL_CALL-level evaluators to validate structured outputs, JSON Schema compliance, and tool response correctness. Create SESSION-level evaluators to verify workflow contract compliance, execution integrity, and cross-step behavioral expectations. Implement TRACE-level evaluators for numerical accuracy verification, calculation consistency, and deterministic result validation. Integrate evaluation components with AWS Agent Core Evaluation workflows and enterprise AI quality pipelines. Collaborate with AI Platform Engineers, ML Engineers, and Dev Ops teams to define evaluation standards, scoring methodologies, and quality acceptance criteria. Design and maintain unit, integration, and validation tests for evaluator libraries and quality frameworks. Support observability by integrating evaluation outputs with Amazon Cloud Watch logging and monitoring. Contribute to enterprise AI governance initiatives by improving evaluation coverage, auditability, compliance, and operational excellence.
Requirements Experience Bachelor's degree in Computer Science, Software Engineering, Information Technology, or a related field.4+ years of professional experience in Python software engineering. Hands-on experience with LLM evaluation, AI quality assurance, or AI/ML validation frameworks. Experience developing and deploying AWS Lambda functions. Experience designing reusable software libraries and writing maintainable, testable Python code. Technical Skills Strong Python development skills. AWS Lambda development (including custom code-based evaluators for AWS Agent Core). LLM-as-a-Judge prompt engineering for subjective evaluation. PII detection using regex-based techniques and AWS Bedrock Guardrails. JSON Schema validation for TOOL_CALL-level evaluation. Workflow contract compliance validation for SESSION-level evaluation. Numerical accuracy and deterministic validation logic for TRACE-level evaluation. Unit and integration testing. Experience with Git and modern CI/CD practices.
Nice to Have Experience with AWS Agent Core Evaluation and custom evaluator registration. Hands-on experience integrating AWS Bedrock Guardrails. Experience using Amazon Cloud Watch Logs for evaluator observability and monitoring. Familiarity with enterprise AI governance, responsible AI, and compliance frameworks. Experience working with AI agent frameworks and LLM orchestration platforms.
Why Join Us? Build foundational evaluation capabilities for enterprise-scale AI platforms. Work on cutting-edge LLM, AI agent, and cloud technologies. Contribute to scalable, reusable quality assurance frameworks used across multiple AI products. Collaborate with highly skilled AI Platform, ML, and Cloud Engineering teams. Influence enterprise AI governance, reliability, and compliance for global-scale solutions.

وظائف مشابهة

حول Intellias
مصر, القاهرة
برامج الكمبيوتر