Skip to content
View DiazSk's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Block or report DiazSk

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
DiazSk/README.md
Zaid Shaikh β€” Data Engineer, Backend Systems. Seattle, WA. Available December 2026. MS Computer Science, Northeastern. shaikh.zaid@northeastern.edu

I build the infrastructure layer β€” streaming pipelines, distributed warehouses, and high-throughput backend systems that move millions of records reliably.

PORTFOLIO β†’ Β· EMAIL β†’ Β· LINKEDIN β†’

MS Computer Science, Northeastern University β€” December 2026, 4.0 GPA. Seattle, WA. Open to full-time roles starting December 2026.


HOW THESE WERE MEASURED

A number on its own is a claim. Each of these says how it was taken.

01 21,091 msg/s
SUSTAINED THROUGHPUT
Write-through coupled message consumption to MySQL's 2–5 ms insert latency, capping throughput near 500 msg/s regardless of broker capacity. Write-behind persistence with in-memory batching decoupled the two paths β€” 42Γ— the baseline, zero data loss across 1M messages.
02 400 clients
LIVE FAN-OUT, 0 LOST
Highest WebSocket step tested, every client receiving the same trades as the best-served one (delivery ratio 1.000) at p95 166 ms exchange-to-client. Across four injected faults β€” Flink TaskManager kill, Kafka restart, 30 s Postgres and Redis outages β€” 0 trades lost, 0 inconsistent candles.
03 2 of 5
HYPOTHESES REFUTED
Five hypotheses written down before any transformation ran; two of them failed, and the dashboard publishes the failures instead of quietly dropping them. A headline "+2,223% premium" rested on 11 providers, so it renders with a thin sample badge rather than as a finding β€” warn, don't hide. Measured over 9,660,252 Medicare claims, 13 quality gates and 108 tests.

WORK

β–ͺ BACKEND ENGINEERING Β· 21,091 msg/s sustained

MySQL's 2–5 ms insert latency coupled message consumption to persistence speed under write-through, capping throughput at roughly 500 msg/s regardless of broker capacity. Write-behind persistence with in-memory batching (2k–5k rows/commit) decoupled the two paths entirely. CQRS isolation kept read and write models independent, so write-side failures could not starve read queries.

Result: 21,091 msg/s sustained Β· 13 ms read latency at 1M-row scale Β· zero data loss across 1M messages.

Architecture
WebSocket Gateway β†’ RabbitMQ β†’ Consumer Pool β†’ In-Memory Batch Buffer β†’ MySQL
                                     β”‚                                    β”‚
                                     └──────→ Redis (hot reads) β†β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                              CQRS read model

Write path and read path never share a bottleneck. The batch buffer absorbs broker bursts at memory speed; MySQL commits 2k–5k rows at a time behind it. Redis serves the read model, so a stalled write never blocks a query.

Java RabbitMQ Redis MySQL HikariCP WebSockets AWS EC2


β–ͺ DATA ENGINEERING Β· 9.66M claims, served with no backend

A cube cannot reproduce a row-level filter. Pre-aggregating the Gold marts made one dashboard panel 41% wrong while still looking plausible β€” the totals were internally consistent, just answering a different question than the filter implied. A parity gate now re-derives every panel from the fact table and fails the build on disagreement. The same pass caught two of the 13 quality assertions that could never fail: a NULL inside isin("F","O",None) makes the predicate NULL for every row, so the test passed on any input. Serving is tiered Parquet read client-side by DuckDB-WASM over HTTP range requests, so 9.66M rows are queryable with no backend at all.

Result: 9,660,252 claim rows from 3.06 GB source Β· full medallion run in 231 s on a laptop Β· 13 quality gates, 10 parity assertions, 108 automated tests Β· 2 of 5 hypotheses refuted, and the dashboard reports the refutations.

Architecture
CMS 2023 CSV (3.06 GB, public)
      β”‚
      β”œβ”€ Azure path Β·Β·Β·Β·Β· Data Factory β†’ ADLS Gen2 ← Key Vault / Entra ID (OAuth 2.0)
      β”‚                   (original; subscription retired mid-project)
      └─ Local path ───── PySpark 3.5 + Delta 3.3, local[8]
                              β”‚
        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
        β”‚  Medallion notebooks β€” identical on both  β”‚
        │  01 bronze→silver   28 explicit casts     │
        β”‚  02 β†’gold dims      provider/hcpcs/geo    β”‚
        β”‚  03 β†’gold fact      NPI Γ— HCPCS Γ— POS     β”‚
        β”‚  04 β†’5 hero marts   99 DQ Β· 13 assertions β”‚
        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                              β”‚
              tiered Parquet β†’ DuckDB-WASM (GitHub Pages) Β· Power BI model

Every path resolves through one function, so LAKEHOUSE_LOCAL_ROOT redirects the whole pipeline from cloud to laptop without a fork β€” which is what saved the project when the Azure subscription was retired. The notebooks are byte-for-byte identical across both.

PySpark 3.5 Delta Lake Azure Databricks Azure Data Factory ADLS Gen2 Azure Key Vault Terraform DuckDB-WASM


β–ͺ SYSTEMS ENGINEERING Β· 400 concurrent clients, 0 lost trades

Every trade for 8 Coinbase pairs, deduplicated on trade ID, rolled into 1-minute OHLCV candles in event time β€” watermarks allow 2 s of out-of-order data β€” with an EWMA z-score detector on top. Sinks carry different guarantees on purpose: JDBC writes are insert-or-skip (effectively once), Kafka alerts are transactional and committed per 30 s checkpoint (exactly once), Redis is at-least-once and clients dedupe. Chaos testing earned its keep by finding a real bug: the API's pub/sub listener died on redis-py's own ConnectionError, so live trades never resumed after a Redis restart. Fixed, and covered by a test.

Result: 17,969 trades ingested with 0 duplicates and 0 missed Β· REST p95 12.8 ms at 10 concurrent clients (1,640 req/s) Β· 400-client WebSocket fan-out at delivery ratio 1.000, p95 166 ms exchange-to-client Β· 0 lost trades across a Flink TaskManager kill, a Kafka broker restart, and 30 s Postgres and Redis outages.

Architecture
Coinbase WS ─→ Python producer ─→ Kafka ─→ Apache Flink ─┬─→ TimescaleDB ─┐
               validation           4        dedup Β·      β”‚   + rollups    β”‚
               gap tracking      partitions  1m OHLCV     β”œβ”€β†’ Redis ───────┼─→ FastAPI ─→ Next.js
                    ╰────── crypto:trades ──→ Redis       β”‚   latest       β”‚  REST + WS   terminal
                                              (live line) └─→ Kafka alerts β”˜
                                                              exactly-once
Airflow (hourly):  backfill 90d candles β†’ repair trade gaps β†’ dbt build
                                                              48 nodes Β· 30 tests Β· 4 unit tests

Disagreement between the pipeline and Coinbase's official candles is reported in a mart, never failed as a test β€” close price matches 99.5–100% of minutes on the liquid pairs. Gap repair recovered 43 of 57 gaps exactly by trade ID; the 14 above the 10,000-trade cap are skipped and logged rather than silently interpolated.

Python Apache Kafka Apache Flink (Java) TimescaleDB Redis dbt Apache Airflow FastAPI Next.js Docker


ALSO SHIPPED

Result Project
2.8M
CLEAN RECORDS
NYC Taxi Data Lakehouse Β· β–ͺ DATA ENGINEERING
100 GB batch pipeline on AWS. Athena charges $5/TB scanned, so Glue runs deduplication, schema normalization and null-handling once at ingest β€” the clean layer becomes a guaranteed fact for downstream dbt models rather than a per-query assumption. 96.8% retention through quality gates, fully reproducible via Terraform.
AWS Glue PySpark Apache Airflow dbt AWS S3 Terraform Docker
109.8M
REAL EVENTS
E-commerce Funnel Lakehouse Β· β–ͺ ANALYTICS ENGINEERING
REES46 clickstream, Oct–Nov 2019, on Databricks. Black Friday week lifted cart reach from 9.2% to 11.7% (+2.5 pp, 95% CI +2.46 to +2.55) β€” but a four-day tracking gap nearly told the opposite story. Nov 15 logged 468,262 carts and zero purchases; leaving Nov 14–17 in the baseline reverses every headline result, and each reversal still looks statistically solid. The analysis finds the gap, measures what it costs, and excludes it. A dbt test now warns on any day with carts but no purchases.
Databricks Delta Lake PySpark dbt Unity Catalog GitHub Actions
90%
LATENCY REDUCTION
E-Commerce Data Warehouse (Olist) Β· β–ͺ ANALYTICS ENGINEERING
Snowflake schemas multiply join depth; wide tables double-count when orders and order items share a fact row. A strict star schema with two grain-specific fact tables resolves both β€” one grain, one join path, no aggregation ambiguity. 14 source systems, 1.6M+ records.
Python PostgreSQL Snowflake Apache Airflow Docker

I build the layer between raw data and the millisecond that matters.

STACK

DATA PLATFORMS & PIPELINES Apache Spark (PySpark) Apache Airflow Apache Kafka Apache Flink dbt Databricks Azure Data Factory RabbitMQ
ETL/ELT pipelines Β· Medallion architecture Β· Asset Bundles Β· Unity Catalog
STORAGE & DATABASES PostgreSQL MySQL Redis TimescaleDB Snowflake Delta Lake DuckDB DuckDB-WASM AWS S3
MongoDB
CLOUD & INFRASTRUCTURE AWS Azure Terraform Docker GitHub Actions
Glue Β· S3 Β· Redshift Β· IAM Β· CloudWatch Β· ADLS Gen2 Β· Databricks Β· Key Vault Β· GitLab CI Β· Jenkins
LANGUAGES Python Java SQL TypeScript
Bash
PRODUCT & APIS FastAPI Next.js 16 React 19 WebSockets
Tailwind CSS Β· shadcn/ui Β· Zod
OBSERVABILITY & QUALITY Great Expectations dbt tests Pytest JUnit
data lineage Β· quality checks Β· pre-commit hooks Β· Power BI Β· Metabase Β· Streamlit

EXPERIENCE

Research Co-author β€” The Laundering Effect Β· Khoury College, Northeastern Β· Fall 2025 – Present β–ͺ COLM 2026, UNDER REVIEW

Measuring cumulative semantic erosion under iterative LLM paraphrasing across 36,800+ records. Implemented a composite Semantic Drift Score (SBERT / METEOR / ROUGE-L) that surfaces trajectory-level degradation invisible to single-step metrics. 2 of 5 original hypotheses reported refuted.

Graduate Teaching Assistant β€” Machine Learning (CS6140) Β· Khoury College, Northeastern Β· May 2026 – Present

Weekly office hours debugging student Python implementations of PCA, regression and regularization. Graded assignments reviewing model code, train/test logic and written analyses.


OPEN TO THE RIGHT OPPORTUNITY

Full-time Data Engineering and Backend roles starting December 2026.

shaikh.zaid@northeastern.edu Β· LinkedIn Β· zaid-data.vercel.app

Zaid Shaikh β€” Seattle, WA. Open to full-time, December 2026.

Pinned Loading

  1. chatflow-messaging-system chatflow-messaging-system Public

    Scalable CQRS WebSocket messaging system built with Java, RabbitMQ, and Redis. Features a write-behind persistence pipeline sustaining 21,000+ msg/sec.

    Java

  2. Real-Time-Cryptocurrency-Market-Analyzer Real-Time-Cryptocurrency-Market-Analyzer Public

    Real-time crypto streaming pipeline: Coinbase trades β†’ Kafka β†’ Flink (event-time OHLCV, dedup, z-score anomaly detection) β†’ TimescaleDB + Redis β†’ FastAPI/WebSocket β†’ Next.js terminal. Exactly-once …

    TypeScript

  3. medicare-reimbursement-gap-analyzer medicare-reimbursement-gap-analyzer Public

    Azure Medallion lakehouse on 9.66M CMS Medicare provider-service records — PySpark Bronze→Silver→Gold on ADLS Gen2, with Power BI + marimo dashboards surfacing 5 hero billing insights.

    Python

  4. sql-data-warehouse-project sql-data-warehouse-project Public

    Building a modern data warehouse with PostgreSQL Server, including ETL process, data modeling, and analytics

    Python

  5. nyc-taxi-data-lakehouse nyc-taxi-data-lakehouse Public

    A production-ready data engineering solution featuring cloud-based batch processing, infrastructure as code, and analytics-ready data transformations using the NYC TLC Trip Record dataset.

    Python

  6. ecommerce-funnel-lakehouse ecommerce-funnel-lakehouse Public

    Funnel and cart-abandonment analytics on 109.8M real e-commerce events: PySpark + Delta Lake + dbt on Databricks Free Edition, with CI, data-quality tests, and a stats-backed findings memo and dash…

    HTML