Data Engineer — lakehouses, dbt, and the contracts that keep them honest
Safi / Casablanca, Morocco · Email · LinkedIn
🌐 Portfolio: laila-khezaz-portfolio.vercel.app
"Data engineering is the intersection of security, data management, DataOps, data architecture, orchestration, and software engineering." — Joe Reis & Matt Housley
Hi, I’m Lily. I build reliable data infrastructure and lakehouses with a strong focus on data validation, test coverage, and clear architecture. Right now, I'm wrapping up my manufacturing lakehouse internship (FPLIP) and working on two main projects: Ledger, a bitemporal feature store to prevent data leakage, and a Governed Vector Data Platform for reliable embedding pipelines.
I document key decisions with ADRs—favoring simple, proven tools over unnecessary complexity.
| Project | What it proves | Stack | Status |
|---|---|---|---|
| FPLIP — Factory Performance & Loss Intelligence Platform | Zero-budget manufacturing lakehouse built against a real production constraint (BigQuery Sandbox forbids DML) — solved with partition-decorator loads and a hand-rolled DDL-only SCD2. 18 ADRs document every major call. | BigQuery · dbt · Airflow · FastAPI · Superset · scikit-learn | Internship project — write-up on request |
| E-Commerce Lakehouse Platform | Full Bronze→Silver→Gold lakehouse from raw CSVs: schema contracts with quarantine, incremental + watermarked + deduplicated Delta jobs, 13 dbt models, cross-layer reconciliation checks. | PySpark · Delta Lake · dbt · Airflow · MinIO · DuckDB | Repo |
| Agentic Delta Guard | Data-quality gatekeeper for AI-agent tool events: schema/bounds/freshness contracts, Bronze-vs-quarantine routing, 7 FastMCP tools for enforcement and audit, Kafka + Spark Structured Streaming ingestion. | Python · dbt · Delta Lake · PySpark · Kafka · FastMCP | Repo |
| End-to-End Analytics Engineering Pipeline | Snowflake VARIANT sources transformed through 7 dbt models (staging → dimensions → incremental facts), hash-based surrogate keys, 47 generic tests, dbt CI on GitHub Actions. | Snowflake · dbt · SQL · Jinja · GitHub Actions | Repo |
| Ledger | Bitemporal, point-in-time feature store with a canary engine that flags look-ahead bias before it reaches a model. | Python · DuckDB · Polars | Repo |
| Governed Vector Data Platform | Catalog, lineage, and a quality gate in front of embedding pipelines — so a bad vector can't silently ship. | Python · OpenLineage/Marquez · FastAPI | Repo |
| NeuroSight AI | Medical image analytics & AI diagnostics platform leveraging deep learning for neurological insight and reporting. | Python · PyTorch · FastAPI · Streamlit | Repo |
- Data contracts before pipelines — schema, bounds, and freshness checks are part of the design, not an afterthought
- ADRs for every architectural decision, including tool choices I deliberately didn't make
- dbt tests and CI as a baseline, not a bonus feature
- No AI/agent framing on a project unless the AI component is doing something a deterministic pipeline genuinely couldn't
- SQL (Advanced) — HackerRank
- Associate Data Engineer — DataCamp
- Databricks Fundamentals & Deploy Workloads with Lakeflow Jobs — Databricks Academy
- dbt Fundamentals — dbt Labs
- Data Engineering on AWS: Foundations — AWS
- Data Governance, GDPR & Data Privacy Fundamentals — DataCamp
- OCI AI Foundations — Oracle
Open to Data Engineer / AI Engineer internships (4–6 months), Reach me at leilakhezaz07@gmail.com or on LinkedIn.


