Repository navigation
Conversation
Complete scripted drafting before learner decisions, apply the ordinary battle-duration limit, and report draft ticks separately. Compare complete frozen-action trajectories and terminal state on both seats. Co-authored-by: GPT-6 <codex@openai.com>
Co-authored-by: Codex GPT-6 <codex@openai.com>
Preserve command-handler results and outgoing reward state before lockstep reset. Isolated CPU diagnostic candidate; build and runtime validation delegated to executor. Co-authored-by: Codex (GPT-6) <codex@openai.com>
Distinguish game team from policy-local controlled group before terminal reset. Co-authored-by: Codex (GPT-6) <codex@openai.com>
…ification Preserve main dependency closure in the separate integration candidate. Original 91483bb source and its CPU proof remain frozen; this merged source requires fresh qualification. Co-authored-by: Codex (GPT-6) <codex@openai.com>
Keep the canonical main declaration; the isolated integration compile failed before tests on this duplicate. Preserve frozen 91483 qualification and source. Co-authored-by: Codex (GPT-6) <codex@openai.com>
installPolicy replaces the ordinary VM with an unstructured policy. Mark that replacement as using scalar hero data so runHeroScript refreshes drafting and hero inputs, matching loadBots. Preserve the ordinary draft stateHash assertion. Co-authored-by: Codex (GPT-6) <codex@openai.com>
…simulation The complete lockstep suite passes with this exact test source and candidate95b5 under CPU1/RAM2GiB/no swap. Cover both controlled teams, trace-on/off canonical-state parity, potential-delta rewards and command identity across auto-reset. Existing ordinary draft stateHash assertion remains intact. Co-authored-by: Codex (GPT-6) <codex@openai.com>
relh
marked this pull request as ready for review
October 4, 2026 08:45
relh
marked this pull request as draft
October 6, 2026 01:12
relh
marked this pull request as ready for review
October 6, 2026 01:24
Restore bounded opt-in command traces, actual simulator team identity, and outgoing reward state before automatic reset. Keep the structured VM and ordinary draft API intact. Co-authored-by: Codex (GPT-6) <codex@openai.com>
Integrate current main and replace the broader training changes with the two missing HeroVm observation flags. Both native bridges now refresh scalar host data used by policies. The net diff against main is two changed lines with no added lines. Co-authored-by: Codex (GPT-6) <codex@openai.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Withdrawn in favor of Metta #26740, which adapts training and exports to the existing modern Polyworld player interface. New policies use structured observation fields without legacy hero data. No Polyworld changes are required.
Metta implementation: https://app.graphite.dev/github/pr/Metta-AI/metta/26740