Dewei Zhai

/cv

Taking the pain out of your data — at any scale.

8+ years across 500TB–1PB datalakes on AWS, Azure, GCP, and Alibaba Cloud — long enough to have made, and fixed, most of the mistakes you're about to.

What I sell is certainty — drawn from pitfalls no one writes down, from years of holding to first principles and Occam's razor, and from many turns as the on-call fire-chief: putting fires out, and keeping the next one from starting.

Track record

  1. 2026 — now

    Platform Architect (Azure + Alibaba Cloud) Enyquant

    Sole hands-on architect for a dual-region Lakehouse (Azure EU + Alibaba Cloud China) at an energy-trading, AI-first startup. 100% IaC (CDKTF + Terraform), event-driven serverless pipelines, multi-team IAM. Built an AI-augmented engineering practice (~3× delivery throughput).

  2. 2024 — 2025

    Lead Data Engineer (AWS Datalake) PVH Corp · 2nd engagement

    500+TB AWS data lake, 90+ sources, 1000+ datasets. Led ETL, platform architecture, CI/CD, data quality. Real-time GDPR (de)anonymization service: latency 2h → real-time, cost 10× lower.

  3. 2022 — 2024

    Data Engineer & Infra Admin VodafoneZiggo

    1PB+ datalake. Snowflake migration from Oracle DWH, CDC ingestion (DMS), AWS IaC (CDK + Terraform). Introduced new CI/CD that saved the team 60+ hours.

  4. 2020 — 2022

    Lead Data Engineer (AWS Datalake) PVH Corp · 1st engagement

    Migrated data lake from Hadoop to AWS. Designed external integrations (Adobe, Salesforce, SAP). Built self-service analytics: TTM 2 weeks → 10 minutes.

  5. 2018 — 2020

    Data DevOps Engineer FedEx Digital

    ETL pipelines on AWS & GCP, productized data-science models, Kinesis-based streaming.

  6. 2018

    Data Engineer ABN AMRO

    Hive/Spark ETL, contributed to enterprise data lake build.

  7. 2016 — 2018

    Data Engineer / Hadoop Admin KPN

    Hadoop administration, Hive/Spark ETL, automation with Ansible & Jenkins.

Selected work

The cases below are public; I have references with names attached on request.

↳ Hover or focus a skill to trace it across the CV. Click to keep it highlighted.

  • 2026 — present

    Enyquant Platform Architect & end-to-end Data Engineer — raw → modeled (Azure + Alibaba Cloud) — FTE

    Owned the architecture and hands-on delivery of a dual-region energy-data platform. I separated stable business contracts — time, readiness, lineage, and failure boundaries — from replaceable runtimes, then changed the runtime when the evidence said it no longer fit.

    • — Defined runtime-independent ingestion, temporal, and readiness contracts so compute could change without changing business meaning
    • — Re-platformed daily runtime from Databricks + ADF to Airflow + DuckDB/Polars + DuckLake → ~95% lower monthly cost
    • — Built governed DuckLake access with explicit human, team, and writer boundaries
    • — 100% IaC (CDKTF + Terraform); zero click-ops drift
    • — Reusable multi-cloud architecture layer for consistent EU ↔ China deployment

    Azure (ADLS Gen2, Databricks, ADF, Functions, Key Vault, Entra ID) · Alibaba Cloud (OSS, Function Compute, EMR Spark) · Airflow, DuckDB/Polars, DuckLake, Docker Compose · Terraform + CDKTF (TypeScript) · GitHub Actions + OIDC · Python, SQL

  • 2022 — 2024

    VodafoneZiggo Data Engineer & Infra Admin — Freelance

    Migrated legacy DWH workloads from Oracle to Snowflake on a 1PB+ enterprise datalake. Owned CDC ingestion, IaC, and the CI/CD foundation that the broader team builds on.

    • — Supported Snowflake migration from legacy Oracle DWH
    • — Designed and maintained CDC-based ingestion pipelines (DMS-based)
    • — Managed AWS infrastructure as code (CDK + Terraform)
    • — Introduced CI/CD improvements that saved the broader team 60+ hours of manual work
    • — Delivered ETL pipelines and an internal data-engineering framework on a 1PB+ datalake
    • — Supported data scientists and analytics teams with reliable datasets

    Snowflake · AWS · Terraform, AWS CDK (TypeScript) · Python, SQL, Spark · AWS DMS for CDC · GitLab CI/CD

  • 2020 — 2022 & 2023 — 2025 (returned engagement)

    PVH Corp (Tommy Hilfiger, Calvin Klein) Lead Data Engineer — Freelance

    Lead engineer for a 500+TB AWS data lake with 90+ sources and 1000+ datasets. Two engagements covering the HadoopAWS migration and a later platform-modernisation phase.

    • — Migrated the data lake from Hadoop to AWS
    • — Designed and shipped external integrations: Adobe, Salesforce, SAP, +others
    • — Refactored the ETL layer to be idempotent and config-driven
    • — Fixed long-standing timezone issues across data and scheduling
    • — Self-service analytics platform with 60+ dashboards used by CRM & C-suite — TTM 2 weeks → 10 minutes
    • — Real-time GDPR (de)anonymization service: latency 2h → real-time, cost 10× lower
    • — Migrated workloads to Azure Databricks; integrated GCP BigQuery & Google Analytics sources
    • — Acted as senior platform engineer advising on DataOps, IaC, and production readiness

    AWS (S3, Glue, EMR, ECS, Lambda, API Gateway, Athena, Step Functions) · Spark, Kafka, dbt, Airflow · Terraform, AWS CDK · GitLab CI/CD · PyDeequ for data quality · Azure Databricks, GCP BigQuery (cross-cloud sources)

  • 2018 — 2020

    FedEx Data Engineer — Freelance

    DE on the AWS data platform — managing infra as code and shipping ETL — with a side specialty in productising data-science Spark code into maintainable engineering.

    • — Managed AWS data infrastructure as code (Terraform)
    • — Developed new data-source integrations and ETL pipelines
    • — Productised DS-written Spark code into engineered, scheduled DE pipelines
    • — Optimised a DS Spark job from 3h OOM (<50% progress) to under 5 minutes — 175× speedup, by replacing a serial for-loop over 175 independent cases with a cross-join

    AWS · Terraform · Spark · PySpark · Python

  • 2018

    ABN AMRO Data Engineer — FTE

    Contributed to the DIAL data platform build — consolidating fragmented departmental ETL onto a shared platform and migrating workloads from Hive to PySpark.

    • — Helped consolidate fragmented per-department ETL onto the shared DIAL platform
    • — Migrated workloads from Hive to PySpark
    • — Replaced a colleague's recurring 3-day-per-week manual Excel workflow with a 3-minute script

    Hive · PySpark · Spark · Python

  • 2016 — 2018

    KPN Big Data Consultant — FTE (B2B division)

    Hadoop cluster admin in KPN's B2B arm — installing and operating Hortonworks HDP clusters on bare-metal for external enterprise customers. Built the team's automated system-health framework.

    • — Operated Hortonworks HDP Hadoop clusters on bare-metal hardware for external B2B customers
    • — Built an automated end-to-end system-health check suite in Robot Framework, replacing the manual install-verification routine
    • — Saved the team 8+ hours per week of repetitive manual verification

    Hortonworks HDP · Hadoop · Bare-metal · Robot Framework · Python

  • 2009 — 2012

    Huawei Software Test Engineer — Telecom Core Database (HLR / HSS) — FTE

    Telecom core-network database work — HLR / HSS for mobile operators. Test & delivery lead on a major core-database cutover at KPN Netherlands (16M subscribers in scope, 12M migrated, zero incidents).

    • — Test & delivery lead on the KPN NL core-network database cutover — 16M subscribers in scope, the most critical database in their core network
    • — ~7 months of intensive pre-cutover testing; identified hundreds of bugs across all severity levels
    • — Led the design and execution of the migration plan, including staged rollouts and rollback procedures
    • — Zero incidents across months of cutover operations; 12M subscribers migrated successfully

    HLR · HSS · Telecom core network · Test engineering · Migration planning

Tech stack

Cloud
AWS (deepest), Azure, Alibaba Cloud, GCP
Lakehouse
Databricks, Snowflake, Unity Catalog, Iceberg, DuckDB
ETL
Spark, dbt, Glue, Kafka, AWS DMS (CDC)
Languages
Python, SQL, TypeScript, Scala, Shell, Solidity, Cython
Orchestration
Airflow, ADF, Step Functions, Oozie
IaC
Terraform, AWS CDK, CDKTF, Ansible
CICD
GitHub Actions, GitLab CI/CD, Jenkins
Quality
PyDeequ, Great Expectations

Publications

Certifications

  • AWS Certified Solutions Architect — Associate
  • Databricks Certified Associate Developer for Apache Spark
  • Databricks Certified Data Engineer Associate
  • — Certified Associate in Python Programming

Education

MSc, Communication & Information Systems — Xidian University, China (Telecommunication Engineering: #4 globally in ShanghaiRanking's 2025 Global Ranking of Academic Subjects). IDW evaluation: equivalent to MSc Computing Science (NL).