diff --git a/job_applications/2026-09-07/01_frontier_ai_lab_safety_evaluations.md b/job_applications/2026-09-07/01_frontier_ai_lab_safety_evaluations.md new file mode 100644 index 0000000..a18d549 --- /dev/null +++ b/job_applications/2026-09-07/01_frontier_ai_lab_safety_evaluations.md @@ -0,0 +1,42 @@ +# Cover letter — Frontier AI lab, safety / model evaluations + +**Archetype:** Anthropic, OpenAI, DeepMind, Scale AI Frontier Risk, +Redwood, METR, MATS/Constellation, AISI-style employers hiring for +model evaluations, red-teaming, or alignment research engineering. + +--- + +Dear {HIRING_MANAGER}, + +I'm applying for the {ROLE_TITLE} role at {COMPANY}. I'm finishing an +M.S. in Applied Statistics at RIT (2026) and my independent portfolio +is built around the question your team lives with every day: how do you +evaluate frontier-model behavior in a way that is measurable, reliable, +and defensible under review? + +My **AI Safety Red-Team Evaluation** project designs a two-stage +workflow that pairs an LLM ensemble with a supervised harm classifier +across 12,500 response pairs and six harm categories. I report +Krippendorff's α = 0.81 for annotation reliability and 96.8% held-out +classification accuracy, and — this is the part I think matters most — +I keep annotator agreement, classifier performance, and downstream +Bayesian risk analysis as three separate claims so a reviewer can +audit each one on its own. The **LLM Ensemble Textbook Bias Detection** +project extends the same pattern to 4,500 passages and 67,500 +LLM-as-judge ratings, with publisher-level effects modeled as Bayesian +partial pooling in PyMC so that the uncertainty on a small publisher's +score is honest rather than optimistic. + +{ROLE_SPECIFIC_HOOK} + +I'd bring rigorous evaluation design, comfort with PyMC / MCMC +diagnostics and SHAP, and a habit of naming a project's limits in the +same document as its results. I'm US-authorized and open to remote or +on-site in New York; happy to walk through either project end-to-end +in a first conversation. + +Best, +Derek Lankeaux +[LinkedIn](https://linkedin.com/in/derek-lankeaux) · +[GitHub](https://github.com/dl1413) · +[Portfolio](https://dl1413.github.io/LLM-Portfolio/) diff --git a/job_applications/2026-09-07/02_nyc_fintech_data_scientist.md b/job_applications/2026-09-07/02_nyc_fintech_data_scientist.md new file mode 100644 index 0000000..519c3b6 --- /dev/null +++ b/job_applications/2026-09-07/02_nyc_fintech_data_scientist.md @@ -0,0 +1,44 @@ +# Cover letter — NYC fintech / quant data scientist + +**Archetype:** JPMorganChase, Goldman Sachs, Morgan Stanley, BlackRock, +Two Sigma, Point72 Cubist, Bridgewater, Citadel, Jane Street data-side +hires — analytics / research roles that want statistical rigor plus +LLM literacy for the growing internal-AI workstreams. + +--- + +Dear {HIRING_MANAGER}, + +I'm writing about the {ROLE_TITLE} position at {COMPANY}. I'm an M.S. +Applied Statistics candidate at RIT (2026) with an independent project +portfolio that leans on the two things I understand your desk cares +about: calibrated uncertainty on the statistical side, and disciplined +LLM evaluation on the emerging-tools side. + +The **LLM Ensemble Textbook Bias Detection** project is the clearest +example. I ran 4,500 passages through a rubric-based LLM-as-judge +ensemble for 67,500 total ratings, reported Krippendorff's α = 0.84 +for inter-rater reliability, and then modeled publisher-level effects +with Bayesian partial pooling in PyMC with full MCMC diagnostics. The +same habits show up in the **AI Safety Red-Team Evaluation**: 12,500 +response pairs, α = 0.81, 96.8% held-out accuracy, with SHAP +attributions attached to the risk analysis so that a reviewer can see +exactly which features are driving a flagged decision. My **RAG +Production Pipeline** work adds the systems side — hybrid retrieval, +grounding checks, and 94.2% citation precision — which is directly +relevant to any internal knowledge-assistant effort your team is +running or planning. + +{ROLE_SPECIFIC_HOOK} + +I'm comfortable in Python, SQL, and R; fluent with scikit-learn, +XGBoost, LightGBM, PyMC, and MLflow; and I write results the way a +partner or a regulator would want to read them — with the limitations +listed in the same paragraph as the headline number. Based in the +region, open to on-site NYC or hybrid, US-authorized. + +Best, +Derek Lankeaux +[LinkedIn](https://linkedin.com/in/derek-lankeaux) · +[GitHub](https://github.com/dl1413) · +[Portfolio](https://dl1413.github.io/LLM-Portfolio/) diff --git a/job_applications/2026-09-07/03_llm_evaluation_startup.md b/job_applications/2026-09-07/03_llm_evaluation_startup.md new file mode 100644 index 0000000..adff79f --- /dev/null +++ b/job_applications/2026-09-07/03_llm_evaluation_startup.md @@ -0,0 +1,42 @@ +# Cover letter — LLM evaluation / observability startup + +**Archetype:** Braintrust, Patronus AI, Arize AI, Weights & Biases, +LangSmith / LangChain, HumanLoop, Galileo, Log10, Confident AI — +companies whose product is helping other teams evaluate, monitor, and +improve LLM applications. + +--- + +Dear {HIRING_MANAGER}, + +I'm applying for {ROLE_TITLE} at {COMPANY}. Your product exists because +teams shipping LLM features can't rely on a single accuracy number to +know whether the thing is working; my three primary portfolio projects +are, honestly, exactly that argument told from the user's side. + +**AI Safety Red-Team Evaluation** is a full evaluation workflow across +12,500 response pairs and six harm categories, reported with +Krippendorff's α = 0.81 alongside 96.8% held-out classifier accuracy — +kept as separate signals rather than collapsed into one metric. +**LLM Ensemble Textbook Bias Detection** runs 4,500 passages through +a rubric-based judge ensemble for 67,500 ratings and layers a Bayesian +hierarchical model on top so publisher-level effects have honest +uncertainty. **RAG Production Pipeline** connects retrieval quality +(96.3% Recall@10), grounding (94.2% citation precision), and +operational concerns — latency, drift, failure modes, rollback — in +one system. If your customers are asking "what should I be measuring +and how do I know it's stable," these are the questions I've been +sitting with. + +{ROLE_SPECIFIC_HOOK} + +I'd bring an evaluations-first mindset, real work in PyMC and MLflow, +enough FastAPI and Docker to prototype something usable, and a habit +of writing docs the way I'd want a customer to read them. Remote is +ideal for me; I'm US-authorized and can start conversations right away. + +Best, +Derek Lankeaux +[LinkedIn](https://linkedin.com/in/derek-lankeaux) · +[GitHub](https://github.com/dl1413) · +[Portfolio](https://dl1413.github.io/LLM-Portfolio/) diff --git a/job_applications/2026-09-07/04_rag_applied_llm_engineer_remote.md b/job_applications/2026-09-07/04_rag_applied_llm_engineer_remote.md new file mode 100644 index 0000000..1bbe0a8 --- /dev/null +++ b/job_applications/2026-09-07/04_rag_applied_llm_engineer_remote.md @@ -0,0 +1,42 @@ +# Cover letter — RAG / applied-LLM systems engineer (remote) + +**Archetype:** Remote-first applied-AI or product-AI teams building +RAG assistants, agentic workflows, or knowledge platforms — from +Series A/B startups (Elicit, Hebbia, Glean, Mendable, Vectara, +LlamaIndex, Cohere applied side) through mid-stage vertical AI +companies. + +--- + +Dear {HIRING_MANAGER}, + +I'm writing about the {ROLE_TITLE} role at {COMPANY}. My **RAG +Production Pipeline** case study is the closest single project to +what your team ships, and it's the one I'd point you to first. + +I designed a hybrid retrieval architecture — dense retrieval plus +BM25, with a re-ranking stage, grounding checks, and confidence +calibration on the generation side — and reported 96.3% Recall@10 +alongside 94.2% citation precision on the documented evaluation. The +report goes past the headline metrics: latency and throughput budgets, +drift monitoring, failure-mode analysis, and the observability, +privacy, and rollback considerations I'd want in place before calling +anything production-ready. That mindset carries through my +**AI Safety Red-Team Evaluation** work — 12,500 response pairs, α = +0.81, 96.8% held-out accuracy, with SHAP-backed risk analysis — and +the **LLM Ensemble Textbook Bias Detection** project's 67,500-rating +LLM-as-judge design. + +{ROLE_SPECIFIC_HOOK} + +On the stack: comfortable with Python, Qdrant, BM25, ColBERT, the +OpenAI and Anthropic APIs, FastAPI, MLflow, Docker, and Kubernetes; +comfortable enough with Kafka and Prometheus to reason about the +operational surface. Fully remote works well for me, US-authorized, +happy to walk through the RAG report in a first call. + +Best, +Derek Lankeaux +[LinkedIn](https://linkedin.com/in/derek-lankeaux) · +[GitHub](https://github.com/dl1413) · +[Portfolio](https://dl1413.github.io/LLM-Portfolio/) diff --git a/job_applications/2026-09-07/05_enterprise_applied_data_scientist.md b/job_applications/2026-09-07/05_enterprise_applied_data_scientist.md new file mode 100644 index 0000000..5733bf6 --- /dev/null +++ b/job_applications/2026-09-07/05_enterprise_applied_data_scientist.md @@ -0,0 +1,42 @@ +# Cover letter — Big-tech / enterprise applied data scientist + +**Archetype:** Meta, Google, Amazon, Microsoft, Netflix, Spotify, IBM +Research, LinkedIn, Bloomberg, Adobe, Salesforce, Palantir applied-DS +or research-engineer new-grad roles based in NYC or remote-eligible. + +--- + +Dear {HIRING_MANAGER}, + +I'm applying for the {ROLE_TITLE} role at {COMPANY}. I'm completing +an M.S. in Applied Statistics at RIT (2026), and my portfolio is built +to show that I can turn an ambiguous question into a defined study, +run it with statistical rigor, and write it up so a partner team can +act on the results. + +Two projects are the shortest path to seeing how I work. The +**LLM Ensemble Textbook Bias Detection** case study evaluates 4,500 +passages with a rubric-based LLM-as-judge ensemble (67,500 ratings, +Krippendorff's α = 0.84) and models publisher-level effects with +Bayesian partial pooling in PyMC — the kind of hierarchical structure +you need whenever you're comparing many small groups. The **AI Safety +Red-Team Evaluation** project applies the same discipline to a harm- +classification problem across 12,500 response pairs and six categories, +reported with α = 0.81 and 96.8% held-out accuracy, and it separates +annotation reliability, classifier performance, and Bayesian risk +analysis as three distinct claims. The **RAG Production Pipeline** +project (96.3% Recall@10; 94.2% citation precision) rounds out the +systems side. + +{ROLE_SPECIFIC_HOOK} + +Stack: Python, SQL, R, scikit-learn, XGBoost, LightGBM, PyMC, MLflow, +SHAP, plus the LLM tooling above. I'm comfortable owning a study from +experimental design through the write-up, and I try to state limits +in the same breath as results. NYC or remote both work; US-authorized. + +Best, +Derek Lankeaux +[LinkedIn](https://linkedin.com/in/derek-lankeaux) · +[GitHub](https://github.com/dl1413) · +[Portfolio](https://dl1413.github.io/LLM-Portfolio/) diff --git a/job_applications/2026-09-07/README.md b/job_applications/2026-09-07/README.md new file mode 100644 index 0000000..cf1163a --- /dev/null +++ b/job_applications/2026-09-07/README.md @@ -0,0 +1,17 @@ +# Applications — 2026-09-07 + +Five tailored cover-letter drafts for today's applications, all pitched +on the same three primary projects listed in +[`../README.md`](../README.md). + +| # | Archetype | File | +|---|---|---| +| 1 | Frontier AI lab — safety / model evaluations | [`01_frontier_ai_lab_safety_evaluations.md`](./01_frontier_ai_lab_safety_evaluations.md) | +| 2 | NYC fintech / quant — data scientist | [`02_nyc_fintech_data_scientist.md`](./02_nyc_fintech_data_scientist.md) | +| 3 | LLM evaluation / observability startup | [`03_llm_evaluation_startup.md`](./03_llm_evaluation_startup.md) | +| 4 | RAG / applied-LLM systems engineer (remote) | [`04_rag_applied_llm_engineer_remote.md`](./04_rag_applied_llm_engineer_remote.md) | +| 5 | Big-tech / enterprise applied data scientist | [`05_enterprise_applied_data_scientist.md`](./05_enterprise_applied_data_scientist.md) | + +Fill placeholders (`{COMPANY}`, `{ROLE_TITLE}`, `{HIRING_MANAGER}`, +`{ROLE_SPECIFIC_HOOK}`) per opening, then log to +[`tracker.md`](./tracker.md). diff --git a/job_applications/2026-09-07/tracker.md b/job_applications/2026-09-07/tracker.md new file mode 100644 index 0000000..915f8f8 --- /dev/null +++ b/job_applications/2026-09-07/tracker.md @@ -0,0 +1,11 @@ +# Submission tracker — 2026-09-07 + +Target: 5 applications, NYC or remote. + +| # | Archetype | Company | Role | Posting URL | Sent (Y/N) | Notes | +|---|---|---|---|---|---|---| +| 1 | Frontier AI lab — safety / evals | | | | | | +| 2 | NYC fintech / quant — data scientist | | | | | | +| 3 | LLM evaluation / observability startup | | | | | | +| 4 | RAG / applied-LLM systems engineer | | | | | | +| 5 | Big-tech / enterprise applied data scientist | | | | | | diff --git a/job_applications/README.md b/job_applications/README.md new file mode 100644 index 0000000..0264d9e --- /dev/null +++ b/job_applications/README.md @@ -0,0 +1,42 @@ +# Job applications + +Daily packages of tailored application materials for Derek Lankeaux's 2026 +data science / applied ML / LLM evaluation job search. Each dated folder +contains five cover-letter drafts targeting distinct role archetypes in +New York City or remote roles, all built around Derek's three primary +portfolio projects. + +## Three primary projects + +The résumé's "Data Scientist | Applied Statistician | LLM Evaluation & +Applied ML" positioning is anchored by three projects that carry across +every letter in this folder: + +1. **AI Safety Red-Team Evaluation** — LLM ensemble labels + supervised + harm classification across 12,500 response pairs; α = 0.81; 96.8% + held-out accuracy; Bayesian risk analysis with PyMC and SHAP. +2. **LLM Ensemble Textbook Bias Detection** — rubric-based LLM-as-judge + evaluation of 4,500 passages yielding 67,500 ratings; α = 0.84; + Bayesian partial pooling for publisher-level effects. +3. **RAG Production Pipeline** — hybrid dense + BM25 retrieval with + re-ranking, grounding checks, and confidence calibration; 96.3% + Recall@10; 94.2% citation precision; latency, drift, and failure-mode + analysis. + +The Breast Cancer Classification benchmark stays in the portfolio for +roles that specifically value diagnostic-ML rigor; it isn't featured +in the daily archetype letters because those target LLM/AI-forward +employers where the three primary projects tell a tighter story. + +## How to use each daily folder + +1. Pick five real openings for the day (NYC or remote, junior/new-grad + data science, applied ML, or LLM evaluation). +2. For each opening, choose the letter whose archetype fits best and + fill in the `{COMPANY}`, `{ROLE_TITLE}`, `{HIRING_MANAGER}`, and any + `{ROLE_SPECIFIC_HOOK}` placeholders — one or two sentences of + company-specific evidence you pulled from the job post. +3. Log the submission in the day's `tracker.md`. + +Letters are drafts, not autofilled applications. Derek reviews and sends +each one himself.