🤖 bench: target GPT-5.6 Sol and run Terminal-Bench lanes at high thinking#3752
Merged
Conversation
Contributor
Author
|
@codex review |
|
Codex Review: Didn't find any major issues. Bravo. Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
github-merge-queue
Bot
removed this pull request from the merge queue due to failed status checks
Jul 24, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Updates the Terminal-Bench model targets: the nightly GPT lane moves from GPT-5.5 to GPT-5.6 Sol, and all benchmark lanes now run at
--thinking high(Claude Opus 5 previously ranxhigh).Background
GPT-5.6 Sol replaced GPT-5.5 as the flagship OpenAI tier, and the requested bench configuration is GPT-5.6 Sol at high thinking and Claude Opus 5 at high thinking. The old per-model
xhigh/highconditional in the nightly workflow is no longer needed, so it is flattened to a single--thinking high.Implementation
nightly-terminal-bench.yml: defaultallmatrix swapsopenai/gpt-5.5foropenai/gpt-5.6-sol;mux_run_argsis now a plain--thinking highfor every lane. Gemini lanes and the Sonnet smoke tests are unchanged.terminal-bench.yml: model example strings refreshed.prepare_leaderboard_submission.py: adds leaderboard metadata foranthropic/claude-opus-5(previously missing, so submission prep could not map the current Opus target) andopenai/gpt-5.6-sol; older entries stay as historical mappings.tbenchskill (and itsbenchmarks/terminal_bench/README.mdsymlink): example commands updated to the new targets.Validation
make static-checkgreen locally.py_compileof the leaderboard script pass.gpt-5.5references are intentional historical metadata entries.Generated with
mux• Model:anthropic:claude-fable-5• Thinking:xhigh• Cost:$5.53