I build the infrastructure layer β streaming pipelines, distributed warehouses, and high-throughput backend systems that move millions of records reliably.
PORTFOLIO β Β· EMAIL β Β· LINKEDIN β
MS Computer Science, Northeastern University β December 2026, 4.0 GPA. Seattle, WA. Open to full-time roles starting December 2026.
A number on its own is a claim. Each of these says how it was taken.
01 |
21,091 msg/s SUSTAINED THROUGHPUT |
Write-through coupled message consumption to MySQL's 2β5 ms insert latency, capping throughput near 500 msg/s regardless of broker capacity. Write-behind persistence with in-memory batching decoupled the two paths β 42Γ the baseline, zero data loss across 1M messages. |
02 |
400 clients LIVE FAN-OUT, 0 LOST |
Highest WebSocket step tested, every client receiving the same trades as the best-served one (delivery ratio 1.000) at p95 166 ms exchange-to-client. Across four injected faults β Flink TaskManager kill, Kafka restart, 30 s Postgres and Redis outages β 0 trades lost, 0 inconsistent candles. |
03 |
2 of 5 HYPOTHESES REFUTED |
Five hypotheses written down before any transformation ran; two of them failed, and the dashboard publishes the failures instead of quietly dropping them. A headline "+2,223% premium" rested on 11 providers, so it renders with a thin sample badge rather than as a finding β warn, don't hide. Measured over 9,660,252 Medicare claims, 13 quality gates and 108 tests. |
βͺ BACKEND ENGINEERING Β· 21,091 msg/s sustained
MySQL's 2β5 ms insert latency coupled message consumption to persistence speed under write-through, capping throughput at roughly 500 msg/s regardless of broker capacity. Write-behind persistence with in-memory batching (2kβ5k rows/commit) decoupled the two paths entirely. CQRS isolation kept read and write models independent, so write-side failures could not starve read queries.
Result: 21,091 msg/s sustained Β· 13 ms read latency at 1M-row scale Β· zero data loss across 1M messages.
Architecture
WebSocket Gateway β RabbitMQ β Consumer Pool β In-Memory Batch Buffer β MySQL
β β
ββββββββ Redis (hot reads) ββββββββββββ
CQRS read model
Write path and read path never share a bottleneck. The batch buffer absorbs broker bursts at memory speed; MySQL commits 2kβ5k rows at a time behind it. Redis serves the read model, so a stalled write never blocks a query.
Java RabbitMQ Redis MySQL HikariCP WebSockets AWS EC2
βͺ DATA ENGINEERING Β· 9.66M claims, served with no backend
A cube cannot reproduce a row-level filter. Pre-aggregating the Gold marts made one
dashboard panel 41% wrong while still looking plausible β the totals were internally
consistent, just answering a different question than the filter implied. A parity gate
now re-derives every panel from the fact table and fails the build on disagreement.
The same pass caught two of the 13 quality assertions that could never fail: a NULL
inside isin("F","O",None) makes the predicate NULL for every row, so the test
passed on any input. Serving is tiered Parquet read client-side by DuckDB-WASM over
HTTP range requests, so 9.66M rows are queryable with no backend at all.
Result: 9,660,252 claim rows from 3.06 GB source Β· full medallion run in 231 s on a laptop Β· 13 quality gates, 10 parity assertions, 108 automated tests Β· 2 of 5 hypotheses refuted, and the dashboard reports the refutations.
Architecture
CMS 2023 CSV (3.06 GB, public)
β
ββ Azure path Β·Β·Β·Β·Β· Data Factory β ADLS Gen2 β Key Vault / Entra ID (OAuth 2.0)
β (original; subscription retired mid-project)
ββ Local path βββββ PySpark 3.5 + Delta 3.3, local[8]
β
βββββββββββββββββββββββ΄ββββββββββββββββββββββ
β Medallion notebooks β identical on both β
β 01 bronzeβsilver 28 explicit casts β
β 02 βgold dims provider/hcpcs/geo β
β 03 βgold fact NPI Γ HCPCS Γ POS β
β 04 β5 hero marts 99 DQ Β· 13 assertions β
βββββββββββββββββββββββ¬ββββββββββββββββββββββ
β
tiered Parquet β DuckDB-WASM (GitHub Pages) Β· Power BI model
Every path resolves through one function, so LAKEHOUSE_LOCAL_ROOT redirects the whole
pipeline from cloud to laptop without a fork β which is what saved the project when the
Azure subscription was retired. The notebooks are byte-for-byte identical across both.
PySpark 3.5 Delta Lake Azure Databricks Azure Data Factory ADLS Gen2 Azure Key Vault Terraform DuckDB-WASM
βͺ SYSTEMS ENGINEERING Β· 400 concurrent clients, 0 lost trades
Every trade for 8 Coinbase pairs, deduplicated on trade ID, rolled into 1-minute OHLCV
candles in event time β watermarks allow 2 s of out-of-order data β with an EWMA
z-score detector on top. Sinks carry different guarantees on purpose: JDBC writes are
insert-or-skip (effectively once), Kafka alerts are transactional and committed per 30 s
checkpoint (exactly once), Redis is at-least-once and clients dedupe. Chaos testing
earned its keep by finding a real bug: the API's pub/sub listener died on redis-py's own
ConnectionError, so live trades never resumed after a Redis restart. Fixed, and covered
by a test.
Result: 17,969 trades ingested with 0 duplicates and 0 missed Β· REST p95 12.8 ms at 10 concurrent clients (1,640 req/s) Β· 400-client WebSocket fan-out at delivery ratio 1.000, p95 166 ms exchange-to-client Β· 0 lost trades across a Flink TaskManager kill, a Kafka broker restart, and 30 s Postgres and Redis outages.
Architecture
Coinbase WS ββ Python producer ββ Kafka ββ Apache Flink ββ¬ββ TimescaleDB ββ
validation 4 dedup Β· β + rollups β
gap tracking partitions 1m OHLCV βββ Redis ββββββββΌββ FastAPI ββ Next.js
β°ββββββ crypto:trades βββ Redis β latest β REST + WS terminal
(live line) βββ Kafka alerts β
exactly-once
Airflow (hourly): backfill 90d candles β repair trade gaps β dbt build
48 nodes Β· 30 tests Β· 4 unit tests
Disagreement between the pipeline and Coinbase's official candles is reported in a mart, never failed as a test β close price matches 99.5β100% of minutes on the liquid pairs. Gap repair recovered 43 of 57 gaps exactly by trade ID; the 14 above the 10,000-trade cap are skipped and logged rather than silently interpolated.
Python Apache Kafka Apache Flink (Java) TimescaleDB Redis dbt Apache Airflow FastAPI Next.js Docker
| Result | Project |
|---|---|
2.8M CLEAN RECORDS |
NYC Taxi Data Lakehouse Β· βͺ DATA ENGINEERING 100 GB batch pipeline on AWS. Athena charges $5/TB scanned, so Glue runs deduplication, schema normalization and null-handling once at ingest β the clean layer becomes a guaranteed fact for downstream dbt models rather than a per-query assumption. 96.8% retention through quality gates, fully reproducible via Terraform. AWS Glue PySpark Apache Airflow dbt AWS S3 Terraform Docker |
109.8M REAL EVENTS |
E-commerce Funnel Lakehouse Β· βͺ ANALYTICS ENGINEERING REES46 clickstream, OctβNov 2019, on Databricks. Black Friday week lifted cart reach from 9.2% to 11.7% (+2.5 pp, 95% CI +2.46 to +2.55) β but a four-day tracking gap nearly told the opposite story. Nov 15 logged 468,262 carts and zero purchases; leaving Nov 14β17 in the baseline reverses every headline result, and each reversal still looks statistically solid. The analysis finds the gap, measures what it costs, and excludes it. A dbt test now warns on any day with carts but no purchases. Databricks Delta Lake PySpark dbt Unity Catalog GitHub Actions |
90% LATENCY REDUCTION |
E-Commerce Data Warehouse (Olist) Β· βͺ ANALYTICS ENGINEERING Snowflake schemas multiply join depth; wide tables double-count when orders and order items share a fact row. A strict star schema with two grain-specific fact tables resolves both β one grain, one join path, no aggregation ambiguity. 14 source systems, 1.6M+ records. Python PostgreSQL Snowflake Apache Airflow Docker |
DATA PLATFORMS & PIPELINES |
Apache Spark (PySpark) Apache Airflow Apache Kafka Apache Flink dbt Databricks Azure Data Factory RabbitMQ ETL/ELT pipelines Β· Medallion architecture Β· Asset Bundles Β· Unity Catalog |
STORAGE & DATABASES |
PostgreSQL MySQL Redis TimescaleDB Snowflake Delta Lake DuckDB DuckDB-WASM AWS S3 MongoDB |
CLOUD & INFRASTRUCTURE |
AWS Azure Terraform Docker GitHub Actions Glue Β· S3 Β· Redshift Β· IAM Β· CloudWatch Β· ADLS Gen2 Β· Databricks Β· Key Vault Β· GitLab CI Β· Jenkins |
LANGUAGES |
Python Java SQL TypeScript Bash |
PRODUCT & APIS |
FastAPI Next.js 16 React 19 WebSockets Tailwind CSS Β· shadcn/ui Β· Zod |
OBSERVABILITY & QUALITY |
Great Expectations dbt tests Pytest JUnit data lineage Β· quality checks Β· pre-commit hooks Β· Power BI Β· Metabase Β· Streamlit |
Research Co-author β The Laundering Effect Β· Khoury College, Northeastern Β· Fall 2025 β Present
βͺ COLM 2026, UNDER REVIEW
Measuring cumulative semantic erosion under iterative LLM paraphrasing across 36,800+ records. Implemented a composite Semantic Drift Score (SBERT / METEOR / ROUGE-L) that surfaces trajectory-level degradation invisible to single-step metrics. 2 of 5 original hypotheses reported refuted.
Graduate Teaching Assistant β Machine Learning (CS6140) Β· Khoury College, Northeastern Β· May 2026 β Present
Weekly office hours debugging student Python implementations of PCA, regression and regularization. Graded assignments reviewing model code, train/test logic and written analyses.
Full-time Data Engineering and Backend roles starting December 2026.
shaikh.zaid@northeastern.edu Β· LinkedIn Β· zaid-data.vercel.app


