Responsibilities
- Build a centralized data lake on GCP data services by integrating diverse data sources throughout the enterprise.
- Develop, maintain, and optimize Spark-powered batch and streaming data processing pipelines. Leverage GCP data services for complex data engineering tasks and ensure smooth integration with other platform components.
- Design and implement data validation and quality checks to ensure accuracy, completeness and consistency across our data workflows.
- Work closely with the AI / ML team to enable AI use cases and in particular address data needs for Machine Learning model development and monitoring.
- Collaborate with cross-functional teams, including Data Analysts and other business users from Operations, Marketing and Commercial departments, to extract valuable insights and enable data-driven decision making.
- Collaborate with the Product teams to design, implement, and maintain the data models for analytical use cases.
- Engage in technology exploration and research, the development of PoC s, conducting deep investigations.
- Design and manage ETL/ELT processes, ensuring data integrity, availability, and performance.
- Troubleshoot data issues and conduct root cause analysis when reporting data is in question.
- Work on Gen AI related initiatives aligned with the company goals.
Required Technical Skills
- PySpark - Batch and Streaming
- GCP - Dataproc, Dataflow, DataStream, Dataplex, Pub/Sub, BigQuery and Cloud Storage
- NoSQL (preferably MongoDB)
- Programming languages: Scala/Python
- Great Expectation, or similar DQ framework
- Familiarity with workflow management tools like: Airflow, Prefect or Luigi
- Understanding of Data Governance, Data Warehousing and Data Modelling
- Good SQL knowledge
Business
- Able to communicate effectively, distill technical knowledge into digestible messages in a succinct / visual way
- Proactively identify and contribute with team development initiatives, and supporting junior members
Good to have skills
-
- Infrastructure-as-Code, preferably Terraform
- Docker and Kubernetes
- Looker
- AI / ML engineering knowledge
- Lineage, or relevant tools e.g. Atlan