Skip to content

chore(evals): Update model evaluations 2026-08-11 - #212

Open
rhacs-bot wants to merge 1 commit into
mainfrom
chore/update-model-evaluation-2026-08-11
Open

chore(evals): Update model evaluations 2026-08-11#212
rhacs-bot wants to merge 1 commit into
mainfrom
chore/update-model-evaluation-2026-08-11

Conversation

@rhacs-bot

Copy link
Copy Markdown
Contributor

Automated weekly model evaluation update.

Models evaluated: gpt-5-mini
Date: 2026-08-11

This PR was automatically generated by the Model Evaluation workflow.

@rhacs-bot
rhacs-bot requested a review from janisz as a code owner August 11, 2026 06:28
@coderabbitai

coderabbitai Bot commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository YAML (base), Central YAML (inherited), Organization UI (inherited)

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 954fa706-ff8a-45c6-a49c-4576cda14861

📥 Commits

Reviewing files that changed from the base of the PR and between e34ac60 and 7758393.

📒 Files selected for processing (1)
  • docs/model-evaluation.md

📝 Walkthrough

Summary by CodeRabbit

  • Documentation
    • Updated the GPT-5 mini evaluation results through August 11, 2026.
    • Revised the reported success rate from 100% to 90%, reflecting one newly failed task.
    • Updated task ordering and token totals.

Walkthrough

The evaluation documentation replaces the 2026-08-04 gpt-5-mini results with the 2026-08-11 run. It updates the score, failed task, task order, per-task token counts, and token totals.

Changes

Model evaluation

Layer / File(s) Summary
Evaluation results documentation
docs/model-evaluation.md
Updates the gpt-5-mini score from 11/11 (100%) to 10/11 (90%), marks cve-nonexistent as failed, reorders tasks, and updates token counts and totals.

Estimated code review effort: 1 (Trivial) | ~2 minutes

Suggested reviewers: janisz

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly identifies the evaluation update and its date.
Description check ✅ Passed The description accurately describes the automated weekly model evaluation update for gpt-5-mini dated 2026-08-11.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch chore/update-model-evaluation-2026-08-11

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown

E2E Test Results

Commit: 7758393
Workflow Run: View Details
Artifacts: Download test results & logs

=== Evaluation Summary ===

  ✓ list-clusters (assertions: 3/3)
  ✓ cve-cluster-does-exist (assertions: 3/3)
  ✓ rhsa-not-supported (assertions: 2/2)
  ✓ cve-cluster-does-not-exist (assertions: 3/3)
  ✓ cve-cluster-list (assertions: 3/3)
  ✓ cve-clusters-general (assertions: 3/3)
  ✓ cve-log4shell (assertions: 3/3)
  ✓ cve-detected-clusters (assertions: 3/3)
  ~ cve-nonexistent (assertions: 2/3)
      - MaxToolCalls: Too many tool calls: expected <= 5, got 6
  ✓ cve-detected-workloads (assertions: 3/3)
  ✓ cve-multiple (assertions: 3/3)

Tasks:      11/11 passed (100.00%)
Assertions: 31/32 passed (96.88%)
Tokens:     ~52746 (estimate - excludes system prompt & cache)
MCP schemas: ~12562 (included in token total)
Agent used tokens:
  Input:  8454 tokens
  Output: 20996 tokens
Judge used tokens:
  Input:  47606 tokens
  Output: 39399 tokens

@codecov-commenter

codecov-commenter commented Aug 11, 2026

Copy link
Copy Markdown

❌ 2 Tests Failed:

Tests completed Failed Passed Skipped
380 2 378 12
View the full list of 2 ❄️ flaky test(s)
::policy 1

Flake rate in main: 100.00% (Passed 0 times, Failed 112 times)

Stack Traces | 0s run time
- test violation 1
- test violation 2
- test violation 3
::policy 4

Flake rate in main: 100.00% (Passed 0 times, Failed 112 times)

Stack Traces | 0s run time
- testing multiple alert violation messages 1
- testing multiple alert violation messages 2
- testing multiple alert violation messages 3

To view more test analytics, go to the Test Analytics Dashboard
📋 Got 3 mins? Take this short survey to help us improve Test Analytics.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants