Pham Thanh Hai
Open to Work
Data Engineering Specialist

PHAMTHANH HAI

Feb 28, 2003 | ETL & Infrastructure

Dedicated Data Engineer focused on building scalable data pipelines and robust backend systems. Expert in ETL/ELT processes, distributed web scraping, and cloud-native infrastructure (Kubernetes/MinIO). I specialize in transforming complex raw data into high-quality assets for business intelligence. Proficient in leveraging AI Agents to automate workflows, accelerate development, and solve real engineering problems effectively. Capable of automating repetitive processes across data pipelines, deployments, and documentation β€” reducing manual effort and increasing team velocity.

02 / Career

Experience

JITS Innovation Labs

Data Engineer

JITS Innovation Labs

February 2026 - Present
  • Developed dbt transformation flows for banking financial reporting pipelines, ensuring data accuracy and traceability.
  • Optimized Kubernetes infrastructure and CI/CD pipelines; contributed to system architecture design for enterprise data platforms.
  • Collaborated directly with foreign clients at Central Banks (Timor-Leste, Myanmar), including onsite visits abroad to gather requirements and deliver data solutions.
  • Applied AI Agents (Claude) for document management automation and accelerating data pipeline development workflows.
  • Extracted and normalized data from T24 core banking system using Apache Spark and Airbyte CDC for downstream analytics.
Solvitech Corporation

ETL & Data Specialist

Solvitech Corporation

November 2025 - January 2026
  • Developed multi-source ingestion modules for diverse data formats (SQL, NoSQL, APIs).
  • Applied Machine Learning techniques to normalize and clean raw data, improving quality by 30%.
  • Built comprehensive real-time dashboards with Apache Superset for operational monitoring.
  • Managed Dockerized environments for seamless service deployment and scaling.
Amira Holdings JSC

ETL & Infrastructure Engineer

Amira Holdings JSC

May 2024 – Aug 2025
  • Architecting end-to-end ETL pipelines using PySpark to process terabytes of data daily.
  • Implementing distributed scraping clusters with Scrapy and Playwright for massive data extraction.
  • Engineering cloud-native infrastructure using Kubernetes, Helm, and MinIO for data storage.
  • Orchestrating complex data workflows with Apache Airflow to ensure 99.9% pipeline reliability.

03 / Expertise

Tech Stack

Core Engineering Expertise

Big Data

PySparkSpark StreamingDelta LakeTrinoMinIOHDFSDelta FormatIcebergDremioData Build Tool

Workflow

AirflowKafkaDockerCI/CDBash

Extraction

ScrapyPlaywrightSeleniumProxyData CleaningBoto3CDCODBC

Storage

PostgreSQLT24 DBRedisElasticClickHouse

Apps & AI

FastAPIPython OOPDjangoScikit-LearnLLM/RAGAI Agent (Claude)

Infra

KubernetesLinuxGitJenkins

04 / Portfolio

Selected Works

End-to-End Pipeline Β· Collect β†’ Serve
ETL→Transform→API→UI
Multi-Source Data Hub
PySparkAirflowDockerPostgreSQLTrinoSupersetLarge Language Models
Data Source

Multi-Source Data Hub

Massive ETL orchestration system that automatically extracts, cleans, and consolidates job market data from multiple major platforms.

REPOSITORY
feeds β†’
Fullstack Job Search Platform
React.jsFastAPIPostgresRedis
Product

Fullstack Job Search Platform

Intelligence-driven platform providing real-time job insights using highly structured data from internal pipelines.

Data Warehouse Infra Β· Deploy β†’ Orchestrate
K8s→Helm→Airflow→dbt
DWH Kubernetes Deployment
KubernetesHelmKafkaSparkHive MetastoreIcebergMinIODremioAirflowELKJenkinsDebezium CDC
Infrastructure

DWH Kubernetes Deployment

Full data warehouse stack deployed on Kubernetes via Helm: MSSQL CDC (Debezium) into Kafka (Strimzi), Spark + Hive Metastore (Iceberg) processing to MinIO, queried by Dremio, orchestrated by Airflow, with ELK logging and Jenkins CI/CD. Enforces strict layered deployment order with an automated deploy.sh script.

REPOSITORY
provides β†’
DWH Orchestration (T24 COB Pipeline)
AirflowSparkIcebergdbtDebeziumKafkaDremioKubernetesPostgreSQLJenkins
Orchestration

DWH Orchestration (T24 COB Pipeline)

Airflow-orchestrated COB pipeline for Temenos T24 core banking: ingests via Debezium CDC and SFTP into an Iceberg lakehouse, gates Silver/Gold builds on sealed business-date snapshots, and produces 41 dbt Gold reports across 5 domains. Runs on the Kubernetes infra provisioned by dwh-deployments.

Side Project
ElectroShop E-Commerce Platform
Node.jsExpressMySQLDockerGemini AIVNPayScrapy

ElectroShop E-Commerce Platform

Architected an AI-driven build system: 13 sequential skill files guide Claude Agent through the full SDLC, enforced by 10 bash hooks validating security, business logic, and state transitions at each checkpoint.

Available for Hire

Let's Connect

Open to new opportunities in Data Infrastructure

RECRUITMENT FORM

* Data is securely sent to phamthanhhai.dev

Β© 2024 PHAM THANH HAI // DATA ENGINEER // OPEN TO WORK

GITHUB/INFRASTRUCTURE ENGINEER