Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

15 Commits
 
 
 
 

Repository files navigation

👨‍💻 Andiswa Matai | Senior Data Engineer · Lakehouse Architect

Azure Databricks PySpark Delta Lake Power BI Terraform CI/CD


🚀 About Me

I design and build enterprise-scale data platforms for regulated industries such as Banking and Healthcare, transforming fragmented systems into auditable, scalable, and cost-efficient data products.

🔹 Impact Highlights

🏦 Global Markets (RMB): Built finance reconciliation data products enabling audit-ready reporting and regulatory traceability

🏥 Healthcare Modernisation: Migrated SAP IS-H systems to Azure, processing 1.13M+ records with zero silent data loss

🏗️ Enterprise Architecture: Delivered Medallion (Bronze → Silver → Gold) lakehouse platforms with CI/CD and observability

📊 Governance & Quality: Implemented data quality frameworks and pipeline-level validation for mission-critical systems

⭐ Flagship Platforms (Production Systems)

🥇 Azure Fabric Retail Platform

Azure · Databricks · Fabric · Terraform ADF → Databricks Medallion → Fabric DirectLake → Power BI Infrastructure as Code (Terraform) Azure Monitor + alerting framework Cost optimisation (spot + lifecycle scaling)

Business Impact:

Unified retail + loyalty data across channels → single source of truth for revenue and margin analytics

🥇 AWS Customer 360 Platform

AWS · Kinesis · Lambda · Glue · Redshift Real-time + batch unified architecture: Streaming ingestion (Kinesis + Lambda) Batch processing (Glue + Step Functions) Serverless analytics (Redshift)

Business Impact:

Enabled full customer behaviour visibility → improved CRM targeting and retention decisions 📊 Engineering Systems (Domain Projects)

🏦 Finance Reconciliation Engine

Databricks · PySpark · Delta Lake Cash vs RADA reconciliation system: SHA-256 composite business keys Broadcast joins for optimisation MATCHED / UNMATCHED / NEW / CLEARED classification

Impact

Eliminated manual reconciliation workflows and generated audit-ready outputs daily

🔐 Transaction Monitoring Pipeline

Kafka-style ingestion + fraud rules engine Deduplication logic for event streams Pluggable rule-based anomaly detection

Impact

Prevented duplicate processing in real-time transaction flows

🏥 Healthcare SAP Migration

SAP IS-H → Azure Chunked ETL processing Validation + rejection logging POPIA-compliant migration framework

Impact

Migrated 1.13M+ records with zero silent failures

🧾 Debt Collection Lakehouse

Microsoft Fabric · PySpark 1.15M+ records processed per run PTP, recovery, settlement KPI modelling

Impact

Replaced manual reporting with automated KPI generation


🏆 Key Achievements

Reduced reporting turnaround from 2 days → 2 hours Cut audit preparation effort by 35%. Migrated 1.13M+ healthcare records with zero silent data loss. Processed 1.15M+ financial records in single-run pipelines. Eliminated manual reconciliation through automated break detection. Optimised Azure infrastructure saving R76,000+ annually. Achieved zero duplicate transaction processing in streaming pipeline.


🛠️ Core Stack

Cloud: Azure (ADF, Databricks, Synapse, ADLS, Fabric) · AWS · GCP

Data: PySpark · Delta Lake · SQL · Scala

Databases: SQL Server · Snowflake · Redshift · BigQuery

BI: Power BI · DAX · Tableau · Looker

DevOps: Terraform · GitHub Actions · CI/CD · Airflow


🧵 Certifications

AWS Solutions Architect – Associate

Microsoft Power BI

Data Science (Python & R)

GDPR & POPIA Compliance

📬 Connect

📧 andiswacebekhulu1@gmail.com

📍 Roodepoort, South Africa

💼 LinkedIn: LinkedIn

💻 GitHub: GitHub

About

No description, website, or topics provided.

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages