Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,42 @@
# Cover letter — Frontier AI lab, safety / model evaluations

**Archetype:** Anthropic, OpenAI, DeepMind, Scale AI Frontier Risk,
Redwood, METR, MATS/Constellation, AISI-style employers hiring for
model evaluations, red-teaming, or alignment research engineering.

---

Dear {HIRING_MANAGER},

I'm applying for the {ROLE_TITLE} role at {COMPANY}. I'm finishing an
M.S. in Applied Statistics at RIT (2026) and my independent portfolio
is built around the question your team lives with every day: how do you
evaluate frontier-model behavior in a way that is measurable, reliable,
and defensible under review?

My **AI Safety Red-Team Evaluation** project designs a two-stage
workflow that pairs an LLM ensemble with a supervised harm classifier
across 12,500 response pairs and six harm categories. I report
Krippendorff's α = 0.81 for annotation reliability and 96.8% held-out
classification accuracy, and — this is the part I think matters most —
I keep annotator agreement, classifier performance, and downstream
Bayesian risk analysis as three separate claims so a reviewer can
audit each one on its own. The **LLM Ensemble Textbook Bias Detection**
project extends the same pattern to 4,500 passages and 67,500
LLM-as-judge ratings, with publisher-level effects modeled as Bayesian
partial pooling in PyMC so that the uncertainty on a small publisher's
score is honest rather than optimistic.

{ROLE_SPECIFIC_HOOK}

I'd bring rigorous evaluation design, comfort with PyMC / MCMC
diagnostics and SHAP, and a habit of naming a project's limits in the
same document as its results. I'm US-authorized and open to remote or
on-site in New York; happy to walk through either project end-to-end
in a first conversation.

Best,
Derek Lankeaux
[LinkedIn](https://linkedin.com/in/derek-lankeaux) ·
[GitHub](https://github.com/dl1413) ·
[Portfolio](https://dl1413.github.io/LLM-Portfolio/)
44 changes: 44 additions & 0 deletions job_applications/2026-09-07/02_nyc_fintech_data_scientist.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,44 @@
# Cover letter — NYC fintech / quant data scientist

**Archetype:** JPMorganChase, Goldman Sachs, Morgan Stanley, BlackRock,
Two Sigma, Point72 Cubist, Bridgewater, Citadel, Jane Street data-side
hires — analytics / research roles that want statistical rigor plus
LLM literacy for the growing internal-AI workstreams.

---

Dear {HIRING_MANAGER},

I'm writing about the {ROLE_TITLE} position at {COMPANY}. I'm an M.S.
Applied Statistics candidate at RIT (2026) with an independent project
portfolio that leans on the two things I understand your desk cares
about: calibrated uncertainty on the statistical side, and disciplined
LLM evaluation on the emerging-tools side.

The **LLM Ensemble Textbook Bias Detection** project is the clearest
example. I ran 4,500 passages through a rubric-based LLM-as-judge
ensemble for 67,500 total ratings, reported Krippendorff's α = 0.84
for inter-rater reliability, and then modeled publisher-level effects
with Bayesian partial pooling in PyMC with full MCMC diagnostics. The
same habits show up in the **AI Safety Red-Team Evaluation**: 12,500
response pairs, α = 0.81, 96.8% held-out accuracy, with SHAP
attributions attached to the risk analysis so that a reviewer can see
exactly which features are driving a flagged decision. My **RAG
Production Pipeline** work adds the systems side — hybrid retrieval,
grounding checks, and 94.2% citation precision — which is directly
relevant to any internal knowledge-assistant effort your team is
running or planning.

{ROLE_SPECIFIC_HOOK}

I'm comfortable in Python, SQL, and R; fluent with scikit-learn,
XGBoost, LightGBM, PyMC, and MLflow; and I write results the way a
partner or a regulator would want to read them — with the limitations
listed in the same paragraph as the headline number. Based in the
region, open to on-site NYC or hybrid, US-authorized.

Best,
Derek Lankeaux
[LinkedIn](https://linkedin.com/in/derek-lankeaux) ·
[GitHub](https://github.com/dl1413) ·
[Portfolio](https://dl1413.github.io/LLM-Portfolio/)
42 changes: 42 additions & 0 deletions job_applications/2026-09-07/03_llm_evaluation_startup.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,42 @@
# Cover letter — LLM evaluation / observability startup

**Archetype:** Braintrust, Patronus AI, Arize AI, Weights & Biases,
LangSmith / LangChain, HumanLoop, Galileo, Log10, Confident AI —
companies whose product is helping other teams evaluate, monitor, and
improve LLM applications.

---

Dear {HIRING_MANAGER},

I'm applying for {ROLE_TITLE} at {COMPANY}. Your product exists because
teams shipping LLM features can't rely on a single accuracy number to
know whether the thing is working; my three primary portfolio projects
are, honestly, exactly that argument told from the user's side.

**AI Safety Red-Team Evaluation** is a full evaluation workflow across
12,500 response pairs and six harm categories, reported with
Krippendorff's α = 0.81 alongside 96.8% held-out classifier accuracy —
kept as separate signals rather than collapsed into one metric.
**LLM Ensemble Textbook Bias Detection** runs 4,500 passages through
a rubric-based judge ensemble for 67,500 ratings and layers a Bayesian
hierarchical model on top so publisher-level effects have honest
uncertainty. **RAG Production Pipeline** connects retrieval quality
(96.3% Recall@10), grounding (94.2% citation precision), and
operational concerns — latency, drift, failure modes, rollback — in
one system. If your customers are asking "what should I be measuring
and how do I know it's stable," these are the questions I've been
sitting with.

{ROLE_SPECIFIC_HOOK}

I'd bring an evaluations-first mindset, real work in PyMC and MLflow,
enough FastAPI and Docker to prototype something usable, and a habit
of writing docs the way I'd want a customer to read them. Remote is
ideal for me; I'm US-authorized and can start conversations right away.

Best,
Derek Lankeaux
[LinkedIn](https://linkedin.com/in/derek-lankeaux) ·
[GitHub](https://github.com/dl1413) ·
[Portfolio](https://dl1413.github.io/LLM-Portfolio/)
42 changes: 42 additions & 0 deletions job_applications/2026-09-07/04_rag_applied_llm_engineer_remote.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,42 @@
# Cover letter — RAG / applied-LLM systems engineer (remote)

**Archetype:** Remote-first applied-AI or product-AI teams building
RAG assistants, agentic workflows, or knowledge platforms — from
Series A/B startups (Elicit, Hebbia, Glean, Mendable, Vectara,
LlamaIndex, Cohere applied side) through mid-stage vertical AI
companies.

---

Dear {HIRING_MANAGER},

I'm writing about the {ROLE_TITLE} role at {COMPANY}. My **RAG
Production Pipeline** case study is the closest single project to
what your team ships, and it's the one I'd point you to first.

I designed a hybrid retrieval architecture — dense retrieval plus
BM25, with a re-ranking stage, grounding checks, and confidence
calibration on the generation side — and reported 96.3% Recall@10
alongside 94.2% citation precision on the documented evaluation. The
report goes past the headline metrics: latency and throughput budgets,
drift monitoring, failure-mode analysis, and the observability,
privacy, and rollback considerations I'd want in place before calling
anything production-ready. That mindset carries through my
**AI Safety Red-Team Evaluation** work — 12,500 response pairs, α =
0.81, 96.8% held-out accuracy, with SHAP-backed risk analysis — and
the **LLM Ensemble Textbook Bias Detection** project's 67,500-rating
LLM-as-judge design.

{ROLE_SPECIFIC_HOOK}

On the stack: comfortable with Python, Qdrant, BM25, ColBERT, the
OpenAI and Anthropic APIs, FastAPI, MLflow, Docker, and Kubernetes;
comfortable enough with Kafka and Prometheus to reason about the
operational surface. Fully remote works well for me, US-authorized,
happy to walk through the RAG report in a first call.

Best,
Derek Lankeaux
[LinkedIn](https://linkedin.com/in/derek-lankeaux) ·
[GitHub](https://github.com/dl1413) ·
[Portfolio](https://dl1413.github.io/LLM-Portfolio/)
Original file line number Diff line number Diff line change
@@ -0,0 +1,42 @@
# Cover letter — Big-tech / enterprise applied data scientist

**Archetype:** Meta, Google, Amazon, Microsoft, Netflix, Spotify, IBM
Research, LinkedIn, Bloomberg, Adobe, Salesforce, Palantir applied-DS
or research-engineer new-grad roles based in NYC or remote-eligible.

---

Dear {HIRING_MANAGER},

I'm applying for the {ROLE_TITLE} role at {COMPANY}. I'm completing
an M.S. in Applied Statistics at RIT (2026), and my portfolio is built
to show that I can turn an ambiguous question into a defined study,
run it with statistical rigor, and write it up so a partner team can
act on the results.

Two projects are the shortest path to seeing how I work. The
**LLM Ensemble Textbook Bias Detection** case study evaluates 4,500
passages with a rubric-based LLM-as-judge ensemble (67,500 ratings,
Krippendorff's α = 0.84) and models publisher-level effects with
Bayesian partial pooling in PyMC — the kind of hierarchical structure
you need whenever you're comparing many small groups. The **AI Safety
Red-Team Evaluation** project applies the same discipline to a harm-
classification problem across 12,500 response pairs and six categories,
reported with α = 0.81 and 96.8% held-out accuracy, and it separates
annotation reliability, classifier performance, and Bayesian risk
analysis as three distinct claims. The **RAG Production Pipeline**
project (96.3% Recall@10; 94.2% citation precision) rounds out the
systems side.

{ROLE_SPECIFIC_HOOK}

Stack: Python, SQL, R, scikit-learn, XGBoost, LightGBM, PyMC, MLflow,
SHAP, plus the LLM tooling above. I'm comfortable owning a study from
experimental design through the write-up, and I try to state limits
in the same breath as results. NYC or remote both work; US-authorized.

Best,
Derek Lankeaux
[LinkedIn](https://linkedin.com/in/derek-lankeaux) ·
[GitHub](https://github.com/dl1413) ·
[Portfolio](https://dl1413.github.io/LLM-Portfolio/)
17 changes: 17 additions & 0 deletions job_applications/2026-09-07/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
# Applications — 2026-09-07

Five tailored cover-letter drafts for today's applications, all pitched
on the same three primary projects listed in
[`../README.md`](../README.md).

| # | Archetype | File |
|---|---|---|
| 1 | Frontier AI lab — safety / model evaluations | [`01_frontier_ai_lab_safety_evaluations.md`](./01_frontier_ai_lab_safety_evaluations.md) |
| 2 | NYC fintech / quant — data scientist | [`02_nyc_fintech_data_scientist.md`](./02_nyc_fintech_data_scientist.md) |
| 3 | LLM evaluation / observability startup | [`03_llm_evaluation_startup.md`](./03_llm_evaluation_startup.md) |
| 4 | RAG / applied-LLM systems engineer (remote) | [`04_rag_applied_llm_engineer_remote.md`](./04_rag_applied_llm_engineer_remote.md) |
| 5 | Big-tech / enterprise applied data scientist | [`05_enterprise_applied_data_scientist.md`](./05_enterprise_applied_data_scientist.md) |

Fill placeholders (`{COMPANY}`, `{ROLE_TITLE}`, `{HIRING_MANAGER}`,
`{ROLE_SPECIFIC_HOOK}`) per opening, then log to
[`tracker.md`](./tracker.md).
11 changes: 11 additions & 0 deletions job_applications/2026-09-07/tracker.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
# Submission tracker — 2026-09-07

Target: 5 applications, NYC or remote.

| # | Archetype | Company | Role | Posting URL | Sent (Y/N) | Notes |
|---|---|---|---|---|---|---|
| 1 | Frontier AI lab — safety / evals | | | | | |
| 2 | NYC fintech / quant — data scientist | | | | | |
| 3 | LLM evaluation / observability startup | | | | | |
| 4 | RAG / applied-LLM systems engineer | | | | | |
| 5 | Big-tech / enterprise applied data scientist | | | | | |
42 changes: 42 additions & 0 deletions job_applications/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,42 @@
# Job applications

Daily packages of tailored application materials for Derek Lankeaux's 2026
data science / applied ML / LLM evaluation job search. Each dated folder
contains five cover-letter drafts targeting distinct role archetypes in
New York City or remote roles, all built around Derek's three primary
portfolio projects.

## Three primary projects

The résumé's "Data Scientist | Applied Statistician | LLM Evaluation &
Applied ML" positioning is anchored by three projects that carry across
every letter in this folder:

1. **AI Safety Red-Team Evaluation** — LLM ensemble labels + supervised
harm classification across 12,500 response pairs; α = 0.81; 96.8%
held-out accuracy; Bayesian risk analysis with PyMC and SHAP.
2. **LLM Ensemble Textbook Bias Detection** — rubric-based LLM-as-judge
evaluation of 4,500 passages yielding 67,500 ratings; α = 0.84;
Bayesian partial pooling for publisher-level effects.
3. **RAG Production Pipeline** — hybrid dense + BM25 retrieval with
re-ranking, grounding checks, and confidence calibration; 96.3%
Recall@10; 94.2% citation precision; latency, drift, and failure-mode
analysis.

The Breast Cancer Classification benchmark stays in the portfolio for
roles that specifically value diagnostic-ML rigor; it isn't featured
in the daily archetype letters because those target LLM/AI-forward
employers where the three primary projects tell a tighter story.

## How to use each daily folder

1. Pick five real openings for the day (NYC or remote, junior/new-grad
data science, applied ML, or LLM evaluation).
2. For each opening, choose the letter whose archetype fits best and
fill in the `{COMPANY}`, `{ROLE_TITLE}`, `{HIRING_MANAGER}`, and any
`{ROLE_SPECIFIC_HOOK}` placeholders — one or two sentences of
company-specific evidence you pulled from the job post.
3. Log the submission in the day's `tracker.md`.

Letters are drafts, not autofilled applications. Derek reviews and sends
each one himself.
Loading