Building reliable data platforms at scale | 10+ years across fintech, telecom & enterprise systems
Data infrastructure that actually works in production high-throughput pipelines, distributed processing, observability layers, and migration tooling across heterogeneous systems.
I operate at the intersection of data engineering and infrastructure, with a background that runs from Oracle DBA roots through PySpark big data processing to Airflow-orchestrated ETL platforms.
Core
Data Engineering
Databases
Infrastructure
StreamCore — The data brain behind a video streaming platform
An end-to-end open source data engineering project covering:
- Real-time event ingestion via Kafka
- Stream processing with PySpark Structured Streaming
- Batch orchestration using Airflow
- Data modeling with dbt on BigQuery
- Observability and data quality monitoring
Designed to demonstrate production-grade DE architecture on a real-world use case.
| Project | Description | Stack |
|---|---|---|
| CMDB Network Discovery | Multi-vendor network metadata discovery via SSH | Python, Paramiko, Netmiko |
| Oracle → MySQL Migration | Production-grade schema & data migration tooling | Shell, SQL |
| XML to CSV | Bulk XML file conversion at scale | Python, Pandas |
| PyMongo to CSV | MongoDB cursor export utility | Python, PyMongo |
Open to remote Data Engineering roles and visa-sponsored relocation opportunities.


