Data Scientist and Applied Statistics M.S. candidate focused on experimental design, Bayesian inference, machine learning, and LLM evaluation.
LinkedIn · GitHub · Portfolio · Résumé
This repository contains four independent technical case studies. Each project links to a full Markdown report and a publication-formatted PDF; the summary below is intended to make the evidence and the scope easy to evaluate.
| Project | Question | Methods | Evidence |
|---|---|---|---|
| AI Safety Red-Team Evaluation | How can safety evaluation scale beyond manual review? | LLM ensemble annotation, supervised classification, Bayesian risk analysis | Report · PDF · Project page |
| Breast Cancer Classification | Which ensemble methods perform well on WDBC diagnostic features? | Benchmarking, calibration, feature selection, explainability | Report · PDF · Project page |
| LLM Ensemble Bias Detection | Can multiple LLM judges support uncertainty-aware content review? | Rubric-based LLM evaluation, reliability analysis, Bayesian hierarchical modeling | Report · PDF · Project page |
| RAG Production Pipeline | How can retrieval, grounding, and monitoring improve RAG system design? | Hybrid retrieval, re-ranking, confidence calibration, observability design | Report · PDF · Project page |
- Problem framing and evaluation design: Each report documents a defined problem, data/evaluation setup, and methodological choices.
- Statistical rigor: The projects use cross-validation, inter-rater reliability, confidence intervals, Bayesian inference, or calibration as appropriate to the task.
- Decision relevance: The work connects model results to practical review, triage, or monitoring decisions rather than treating a headline metric as sufficient on its own.
- Responsible use: Each package page states the limits of the project and the validation needed before any real-world use.
| Project | Reported result | Why it matters |
|---|---|---|
| AI Safety | 96.8% classification accuracy; Krippendorff's α = 0.81 | Separates annotation reliability from downstream classifier performance. |
| Breast Cancer | 99.12% held-out accuracy; ROC-AUC 0.9987 | Illustrates calibrated supervised-learning evaluation on the WDBC benchmark. |
| LLM Bias Detection | 67,500 ratings; Krippendorff's α = 0.84 | Demonstrates an uncertainty-aware workflow for large-scale content review. |
| RAG | 94.2% citation precision; 96.3% Recall@10 | Connects retrieval quality, grounding, and operational metrics in one systems design. |
All figures above are reported in the linked technical documents. The AI Safety, LLM Bias Detection, and RAG results come from simulated evaluations built to demonstrate each method; they are not measurements of real models, publishers, or a deployed service. The Breast Cancer results use the public WDBC benchmark and are not independent clinical validation.
Python · SQL · R · scikit-learn · XGBoost · LightGBM · PyMC · ArviZ · FastAPI · MLflow · SHAP · OpenAI · Anthropic · Qdrant · Docker · Kubernetes
README.md Portfolio overview
Resume_Derek_Lankeaux.md Résumé source
*_Report.md Four technical reports
*_Publication.pdf Corresponding publication PDFs
project_packages/ Per-project reader guides
generate_publication_pdfs.py Canonical PDF generator
requirements-pdf.txt PDF-generation dependencies
PDF_EXPORT.md Build and validation instructions
scripts/validate_portfolio.py Artifact and local-link validation
.github/workflows/validate-portfolio.yml GitHub Actions quality check
For PDF regeneration and validation, see PDF export instructions. The source notebooks referenced in the reports are not distributed in this repository; the reports and PDFs are the public portfolio artifacts.
Open to 2026 data science, applied ML, and LLM-evaluation opportunities. The best way to connect is on LinkedIn.