Description:
Integrant is looking for game changers to join our team as a "Senior Lead Data Engineer". This is a senior, hands-on leader who designs data and AI solutions end to end and builds them. You'll be responsible for designing solutions for the ingestion, storage, processing, transformation, enrichment, and presentation of data for analytical consumption, and for personally implementing the critical parts of those solutions.
- Engage clients at multiple levels to elicit requirements, assess their current analytical challenges, and turn them into technical proposals and solution designs.
- Select and customize analytical architectures (Data Warehouses, Data Lakes, Data Lakehouses, Data Fabrics, Data Meshes, etc.) and technically justify how they meet each client's needs.
- Design and build agentic and LLM-powered solutions on top of modern data platforms.
- Lead hands-on implementation of data pipelines, warehouses, and lakes, from design through performance tuning and production support.
- Build data strategies, and coach clients and internal teams in the latest architectures and methodologies (DataOps, MLOps, etc.), including mentoring engineers on the team.
Requirements:
Architecture & Leadership
- 10-12 years of experience in data engineering, including hands-on delivery at a senior level.
- Bachelor's degree in Computer Science, Computer Engineering or other quantitative field. Post Graduate degrees preferable.
- The ability to interact with stakeholders at multiple levels within a given organization, understand current analytical challenges and gather requirements.
- Excellent written and spoken English, with the ability to present solutions to technical and business audiences.
- Proven experience writing technical proposals and solution designs for clients.
- A deep understanding of the evolution of analytical architectures from reporting databases to data warehouses, data lakes, data lakehouses, data fabrics, data meshes and beyond, with the ability to technically justify how a proposed architecture meets a client's analytical needs.
- Experience mentoring and guiding data engineers.
Hands-on Data Engineering
- Expert-level SQL and strong programming skills in Python (or Scala/Java) for data processing.
- Extensive experience with at least one enterprise data platform (e.g. Microsoft SQL Server stack, Oracle, Teradata).
- Hands-on experience implementing data warehouses both on-premises and in the cloud (e.g. SQL Server, Oracle, Azure Synapse/Fabric, Amazon Redshift, Google BigQuery).
- Strong dimensional modeling: Dimensions, Facts, Slowly Changing Dimensions, Outriggers, Role-Playing, Junk, Degenerate and Multi-valued Dimensions; Transactional, Periodic Snapshot and Accumulating Snapshot Facts.
- Hands-on experience implementing common ETL/ELT patterns using tools such as SSIS, Informatica, Talend, Azure Data Factory, AWS Glue, or code-based frameworks (e.g. Spark).
- Hands-on experience implementing data quality, testing and observability for data pipelines (e.g. dbt tests, Great Expectations, Monte Carlo, Soda).
- Experience building semantic layers and OLAP models (e.g. SSAS, Power BI semantic models, LookML, dbt semantic layer).
- Hands-on experience with workflow orchestration (e.g. Apache Airflow, Azure Data Factory, AWS Step Functions, Databricks Workflows).
- Experience setting up data lakes and lakehouses on cloud object storage (e.g. ADLS Gen2, Amazon S3, Google Cloud Storage) using open table formats (e.g. Delta Lake, Apache Iceberg, Apache Hudi).
- Experience optimizing query and workload performance and setting up high availability configurations.
- Experience administering and managing analytical solutions both on-premises and in the cloud.
- Experience building streaming pipelines (e.g. Kafka, Kinesis, Event Hubs, Google Pub/Sub, Spark Structured Streaming).
- Experience with query federation and data virtualization (e.g. Trino/Starburst, Amazon Athena, PolyBase, BigQuery federated queries).
- Knowledge of NoSQL databases (e.g. MongoDB, Cosmos DB, DynamoDB, Cassandra).
- Knowledge of BI tools (e.g. Power BI, Tableau, Looker, Qlik).
Cloud, Platforms & Governance
- Experience with multiple cloud data platforms (e.g. Microsoft Azure, AWS, Google Cloud).
- Hands-on experience with Snowflake or Databricks.
- Experience estimating, monitoring and optimizing cloud data platform costs (e.g. warehouse sizing, compute/storage trade-offs, Snowflake credits, Databricks DBUs), and reflecting them in client proposals.
- Experience implementing data governance and security: cataloging, lineage, access control and secrets management (e.g. Unity Catalog, Microsoft Purview, AWS Lake Formation, Collibra, cloud IAM, HashiCorp Vault).
- Experience with CI/CD and version control for data solutions (e.g. GitHub Actions, GitLab CI, Azure DevOps).
AI & Agentic
- Hands-on exposure to LLM and agentic solutions, in PoC or production (e.g. LangGraph/AutoGen/CrewAI/LlamaIndex/n8n).
- Experience with AI-assisted development (e.g. Cursor/Codex/Claude Code).
- Familiarity with DataOps.
Nice to Have
- Familiarity with Agent Protocols (e.g. MCP).
- Familiarity with transformation frameworks (e.g. dbt).
- Familiarity with MLOps and ML platforms (e.g. MLflow, Azure ML, Amazon SageMaker, Vertex AI).
- Familiarity with cloud AI services (e.g. Azure AI services, Amazon Bedrock, Vertex AI).
- Familiarity with containerization tools and frameworks (Docker, Kubernetes, etc.).
- Familiarity with the big data ecosystem, including Spark and Hive (e.g. on Databricks, Amazon EMR, Dataproc, HDInsight).
- Familiarity with serverless functions for lightweight transformations (e.g. AWS Lambda, Azure Functions, Google Cloud Functions).
الوصف:
تستقطب شركة Integrant عناصر متميزة وصناع تغيير للانضمام إلى فريقنا بصفتهم "قائد مهندسي بيانات أول". هذا المنصب قيادي وعملي رفيع المستوى يتولى تصميم وبناء حلول البيانات والذكاء الاصطناعي من البداية إلى النهاية. ستكون مسؤولاً عن تصميم حلول لاستيعاب البيانات، وتخزينها، ومعالجتها، وتحويلها، وإثرائها، وعرضها للاستهلاك التحليلي، بالإضافة إلى التنفيذ الشخصي للأجزاء الحيوية من هذه الحلول.
- التواصل مع العملاء على مستويات متعددة لاستخلاص المتطلبات، وتقييم تحدياتهم التحليلية الحالية، وتحويلها إلى مقترحات تقنية وتصاميم للحلول.
- اختيار وتخصيص البنى التحليلية (مستودعات البيانات، وبحيرات البيانات، ومنازل بحيرات البيانات Data Lakehouses، ونسيج البيانات Data Fabrics، وشبكات البيانات Data Meshes، وغيرها) وتعليل مدى تلبيتها لاحتياجات كل عميل من الناحية التقنية.
- تصميم وبناء حلول قائمة على الوكلاء الأذكياء (Agentic) ونماذج اللغات الكبيرة (LLM) فوق منصات البيانات الحديثة.
- قيادة التنفيذ العملي لمسارات البيانات (Data pipelines)، والمستودعات، والبحيرات، بدءًا من التصميم وحتى ضبط الأداء ودعم البيئة الإنتاجية.
- بناء استراتيجيات البيانات، وتوجيه العملاء والفرق الداخلية في أحدث الهندسات والمنهجيات (DataOps، MLOps، إلخ)، بما في ذلك إرشاد المهندسين في الفريق.
المتطلبات:
الهندسة المعمارية والقيادة
- خبرة تتراوح بين 10 إلى 12 عامًا في هندسة البيانات، بما في ذلك التنفيذ العملي على مستوى قيادي.
- درجة البكالوريوس في علوم الحاسوب، أو هندسة الحاسوب، أو أي مجال كمي آخر. وتُفضل شهادات الدراسات العليا.
- القدرة على التفاعل مع أصحاب المصلحة على مستويات متعددة داخل المؤسسة، وفهم التحديات التحليلية الحالية وجمع المتطلبات.
- إتقان ممتاز للغة الإنجليزية كتابةً وتحدثًا، مع القدرة على تقديم الحلول للجمهور التقني والتجاري.
- خبرة مثبتة في كتابة المقترحات التقنية وتصاميم الحلول للعملاء.
- فهم عميق لتطور البنى التحليلية من قواعد بيانات التقارير إلى مستودعات البيانات، وبحيرات البيانات، ومنازل بحيرات البيانات، ونسيج البيانات، وشبكات البيانات وما بعدها، مع القدرة على تبرير كيفية تلبية البنية المقترحة للاحتياجات التحليلية للعميل تقنيًا.
- خبرة في إرشاد وتوجيه مهندسي البيانات.
هندسة البيانات العملية
- مستوى خبير في SQL ومهارات برمجية قوية في Python (أو Scala/Java) لمعالجة البيانات.
- خبرة واسعة مع منصة بيانات مؤسسية واحدة على الأقل (مثل حزمة Microsoft SQL Server، أو Oracle، أو Teradata).
- خبرة عملية في تطبيق مستودعات البيانات سواء في المقر (on-premises) أو في السحابة (مثل SQL Server، وOracle، وAzure Synapse/Fabric، وAmazon Redshift، وGoogle BigQuery).
- نمذجة بعدية قوية: الأبعاد، والحقائق، والأبعاد بطيئة التغير (SCD)، وOutriggers، وRole-Playing، وJunk، وDegenerate، والأبعاد متعددة القيم؛ وحقائق المعاملات، واللقطات الدورية، واللقطات التراكمية.
- خبرة عملية في تطبيق أنماط ETL/ELT الشائعة باستخدام أدوات مثل SSIS، أو Informatica، أو Talend، أو Azure Data Factory، أو AWS Glue، أو أطر العمل القائمة على الكود (مثل Spark).
- خبرة عملية في تنفيذ جودة البيانات، والاختبار، وقابلية الملاحظة لمسارات البيانات (مثل اختبارات dbt، وGreat Expectations، وMonte Carlo، وSoda).
- خبرة في بناء الطبقات الدلالية ونماذج OLAP (مثل SSAS، ونماذج Power BI الدلالية، وLookML، والطبقة الدلالية لـ dbt).
- خبرة عملية في إدارة وتنسيق سير العمل (مثل Apache Airflow، وAzure Data Factory، وAWS Step Functions، وDatabricks Workflows).
- خبرة في إعداد بحيرات البيانات ومنازل بحيرات البيانات (Lakehouses) على التخزين الكائني السحابي (مثل ADLS Gen2، وAmazon S3، وGoogle Cloud Storage) باستخدام تنسيقات الجداول المفتوحة (مثل Delta Lake، وApache Iceberg، وApache Hudi).
- خبرة في تحسين أداء الاستعلامات وأحمال العمل وإعداد تكوينات التوافر العالي.
- خبرة في إدارة وصيانة الحلول التحليلية سواء في المقر أو في السحابة.
- خبرة في بناء مسارات التدفق المباشر للبيانات (مثل Kafka، وKinesis، وEvent Hubs، وGoogle Pub/Sub، وSpark Structured Streaming).
- خبرة في اتحاد الاستعلامات والافتراضية للبيانات (مثل Trino/Starburst، وAmazon Athena، وPolyBase، والاستعلامات الاتحادية لـ BigQuery).
- معرفة بقواعد بيانات NoSQL (مثل MongoDB، وCosmos DB، وDynamoDB، وCassandra).
- معرفة بأدوات ذكاء الأعمال (مثل Power BI، وTableau، وLooker، وQlik).
السحابة، والمنصات والحوكمة
- خبرة في منصات بيانات سحابية متعددة (مثل Microsoft Azure، وAWS، وGoogle Cloud).
- خبرة عملية في Snowflake أو Databricks.
- خبرة في تقدير تكاليف منصات البيانات السحابية ومراقبتها وتحسينها (مثل تحديد حجم المستودعات، والمفاضلة بين الحوسبة والتخزين، ورصيد Snowflake، ووحدات DBUs في Databricks)، وعكسها في مقترحات العملاء.
- خبرة في تطبيق حوكمة البيانات وأمنها: الفهرسة، والتتبع (Lineage)، وإدارة أذونات الوصول والأسرار (مثل Unity Catalog، وMicrosoft Purview، وAWS Lake Formation، وCollibra، وcloud IAM، وHashiCorp Vault).
- خبرة في التكامل والتسليم المستمر (CI/CD) والتحكم في الإصدارات لحلول البيانات (مثل GitHub Actions، وGitLab CI، وAzure DevOps).
الذكاء الاصطناعي والوكلاء الأذكياء
- خبرة عملية في التعامل مع حلول نماذج اللغات الكبيرة (LLM) والحلول القائمة على الوكلاء الأذكياء، سواء في مرحلة إثبات المفهوم (PoC) أو الإنتاج (مثل LangGraph/AutoGen/CrewAI/LlamaIndex/n8n).
- خبرة في التطوير المدمج بالذكاء الاصطناعي (مثل Cursor/Codex/Claude Code).
- إلمام بـ DataOps.
مهارات يُفضل وجودها
- إلمام ببروتوكولات الوكلاء الأذكياء (مثل MCP).
- إلمام بأطر عمل تحويل البيانات (مثل dbt).
- إلمام بـ MLOps ومنصات تعلم الآلة (مثل MLflow، وAzure ML، وAmazon SageMaker، وVertex AI).
- إلمام بالخدمات السحابية للذكاء الاصطناعي (مثل خدمات Azure AI، وAmazon Bedrock، وVertex AI).
- إلمام بأدوات وأطر عمل الحاويات (Docker، وKubernetes، وغيرها).
- إلمام بمنظومة البيانات الضخمة (Big Data)، بما في ذلك Spark وHive (مثل على Databricks، وAmazon EMR، وDataproc، وHDInsight).
- إلمام بالوظائف الخالية من الخوادم (Serverless functions) للتحويلات خفيفة الوزن (مثل AWS Lambda، وAzure Functions، وGoogle Cloud Functions).