Skip to content

Disclose simulated data and correct claims in three project reports - #108

Merged
dl1413 merged 1 commit into
mainfrom
fix/portfolio-report-claims
Sep 29, 2026
Merged

dl1413 merged 1 commit into
mainfrom
fix/portfolio-report-claims

Conversation

@dl1413

@dl1413 dl1413 commented Sep 29, 2026

Copy link
Copy Markdown
Owner

Adds explicit disclosure that the AI Safety Red-Team, LLM Bias Detection, and RAG results come from simulated evaluations, and corrects internal inconsistencies and unsupported claims in those reports.

  • Disclosure: new Data and Results Note in each report; README, package READMEs, and resume bullets qualified.
  • Removed: IEEE 2830-2025 / ISO / EU AI Act compliance claims, production / live-service / uptime claims, fabricated code-and-data availability, demographic fairness audits on data without demographics.
  • Red-team: precision/recall/F1/FNR recomputed from the confusion matrix (90.8% / 95.6% / 93.2% / 4.4%); CV mean 96.3% ± 0.6%; Table 3 rank order; Krippendorff thresholds.
  • Bias: 15 ratings per passage (3 LLMs × 5 dimensions); cost table recomputed (~$481.50 at list prices); model specs and version strings corrected.
  • RAG: end-to-end p99 312ms (stage p99s are not additive); over-budget stages marked; ECE 0.006; references corrected.
  • Three publication PDFs regenerated; scripts/validate_portfolio.py passes.

🤖 Generated with Claude Code

https://claude.ai/code/session_01W5HTneJKS5VFAzDE5sGDV9


Generated by Claude Code

AI Safety Red-Team, LLM Bias Detection, and RAG reports:
- Add a Data and Results Note stating the data and results are simulated
  to demonstrate each method; qualify README, package READMEs, and resume.
- Remove unsupported IEEE 2830-2025 / ISO / EU AI Act compliance claims,
  production and live-service claims, and fabricated availability claims.
- Red-team: recompute precision/recall/F1/FNR from the confusion matrix,
  fix CV mean/SD/CI, fix Table 3 ranking, use Krippendorff's thresholds.
- Bias: 15 ratings per passage (3 LLMs x 5 dimensions), fix cost table
  (22,500 ratings per model), correct model specs and version strings.
- RAG: fix end-to-end p99 (stage p99s are not additive), over-budget
  stages, ECE (0.006), remove uptime and demographic fairness claims,
  fix references.
- Regenerate the three publication PDFs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5HTneJKS5VFAzDE5sGDV9
Copilot AI balanced review requested due to automatic review settings September 29, 2026 01:11

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@dl1413
dl1413 merged commit cd7e2c3 into main Sep 29, 2026
1 check passed

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 4850de16e6

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

| **Passages per Textbook** | 30 | Random sampling with coverage constraints |
| **Total Passages** | 4,500 | Unit of analysis |
| **Ratings per Passage** | 3 | One per LLM |
| **Ratings per Passage** | 15 | 3 LLMs × 5 bias dimensions |

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Derive rating counts from actual model outputs

The revised count multiplies each model's result by five rubric dimensions, but BIAS_PROMPT returns only one bias_score and rate_passage makes one request per model; the reliability matrix likewise contains only three ratings per passage. This design therefore produces 13,500 ratings/API calls for 4,500 passages, not 67,500, so the downstream 81-million-token and ~$481 estimates are inflated fivefold unless the prompt and data model are changed to return five separate scores.

Useful? React with 👍 / 👎.

## Abstract

This report presents a computational framework for detecting and quantifying political bias in educational textbooks, combining an ensemble of three frontier Large Language Models (LLMs)—GPT-4, Claude-3-Opus, and Llama-3-70B—with Bayesian hierarchical modeling for robust inference. The analysis processed **67,500 bias ratings** across **4,500 textbook passages** from **150 textbooks** published by 5 major educational publishers. We demonstrate excellent inter-rater reliability among LLMs (Krippendorff's α = 0.84), statistically significant publisher-level bias differences (Friedman χ² = 42.73, p < 0.001), and quantified uncertainty via Bayesian posterior distributions with 95% Highest Density Intervals (HDI). Three of five publishers exhibited statistically credible bias (95% HDI excluding zero), with effect sizes from -0.48 (liberal) to +0.38 (conservative) on a [-2, +2] scale. The framework establishes a scalable, reproducible methodology for large-scale educational content auditing with rigorous uncertainty quantification.
This report presents a computational framework for detecting and quantifying political bias in educational textbooks, combining an ensemble of three Large Language Models (LLMs)—GPT-4o, Claude-3.5-Sonnet, and Llama-3.2-90B—with Bayesian hierarchical modeling, demonstrated on a simulated corpus. The simulated analysis comprises **67,500 bias ratings** (3 LLMs × 5 bias dimensions) across **4,500 passages** from **150 textbooks** attributed to five anonymized publishers (A–E). The simulated ratings reach Krippendorff's α = 0.84, above the conventional 0.80 reliability threshold, with statistically significant publisher-level bias differences (Friedman χ² = 42.73, p < 0.001), and quantified uncertainty via Bayesian posterior distributions with 95% Highest Density Intervals (HDI). Three of five publishers exhibited statistically credible bias (95% HDI excluding zero), with effect sizes from -0.48 (liberal) to +0.38 (conservative) on a [-2, +2] scale. All data and results are simulated to demonstrate the method (see the Data and Results Note); they are not findings about real publishers or textbooks.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Configure the OpenAI model reported by the design

The report now identifies GPT-4o as an ensemble member, but the implementation excerpt still sends requests with model="gpt-4-turbo". Anyone implementing the documented pipeline would evaluate and cost a different model than the one named throughout the results, so either the model literal or the GPT-4o claims need to be aligned.

Useful? React with 👍 / 👎.

| **Throughput** | 850/hr | 500/hr | [Yes] |
| **Cost per Sample** | $0.018 | <$0.05 | [Yes] |
| **False Negative Rate** | 3.9% | <5% | [Yes] |
| **False Negative Rate** | 4.4% | <5% | [Yes] |

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Log the corrected false-negative rate

The corrected benchmark reports an FNR of 4.4% (25/569), but the preceding MLflow example still logs false_negative_rate: 0.039. Running the documented tracking code would therefore preserve the superseded value in experiment dashboards and model comparisons even though the report now derives the correct metric.

Useful? React with 👍 / 👎.

Comment thread RAG_Project_Report.md
| Hallucination Detection | 16ms | 32ms | 25ms | Over budget (p99) |
| **Total End-to-End** | **187ms** | **312ms** (measured end-to-end) | 400ms | ✓ Pass |

Stage p99 latencies are not additive, so the end-to-end p99 is measured directly (Table 13, 1,000 concurrent users) rather than summed; adding the stage p99s (408ms) is not a valid estimate of end-to-end p99. Three stages exceed their p99 budgets and are the first targets for optimization. The simulated LLM stage latency (78ms p50) is also far below what hosted frontier-model APIs typically deliver for a full response, so a real deployment would need streaming, smaller models, or a larger latency budget.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Mark vector search as over its p99 budget

The vector-search row reports a 95ms p99 against an 80ms budget, so it is also over budget. The newly added summary says only three stages exceed their budgets and directs optimization to those stages, omitting vector search while its row remains marked ✓ Under; this can cause the design review to ignore a roughly 19% budget overrun.

Useful? React with 👍 / 👎.

Comment thread RAG_Project_Report.md
---

## 9. Production Deployment and MLOps
## 9. Deployment Design and MLOps

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Update the RAG table-of-contents anchor

Renaming this section changes its generated Markdown anchor to #9-deployment-design-and-mlops, but the table of contents still links to #9-production-deployment-and-mlops. Clicking the deployment entry in this long report now goes nowhere, so the corresponding table-of-contents label and target need to be updated with the heading.

Useful? React with 👍 / 👎.

---

## 12. Production Framework and MLOps
## 12. Pipeline Design and MLOps

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Update the bias report table-of-contents anchor

Renaming this section changes its generated Markdown anchor to #12-pipeline-design-and-mlops, while the table of contents still targets #12-production-framework-and-mlops. The navigation entry is therefore broken for readers of the Markdown report and should be changed alongside the heading.

Useful? React with 👍 / 👎.

@dl1413
dl1413 deleted the fix/portfolio-report-claims branch September 29, 2026 01:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants