Skip to content

Refresh scalar observations in native training bridges - #95

Closed
relh wants to merge 14 commits into
mainfrom
relh/sam-attribution-game-cpu
Closed

relh wants to merge 14 commits into
mainfrom
relh/sam-attribution-game-cpu

Conversation

@relh

@relh relh commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Withdrawn in favor of Metta #26740, which adapts training and exports to the existing modern Polyworld player interface. New policies use structured observation fields without legacy hero data. No Polyworld changes are required.

Metta implementation: https://app.graphite.dev/github/pr/Metta-AI/metta/26740

relh and others added 9 commits October 1, 2026 02:05
Complete scripted drafting before learner decisions, apply the ordinary battle-duration limit, and report draft ticks separately. Compare complete frozen-action trajectories and terminal state on both seats.

Co-authored-by: GPT-6 <codex@openai.com>
Port the object queries and borrowed visibility getters from upstream 806082d. Pin its Bassy compiler while retaining the scripted draft, battle clock, and native ABI from 560a8c1.

Co-authored-by: Treeform <starplant@gmail.com>

Co-authored-by: Codex GPT-6 <noreply@openai.com>
Co-authored-by: Codex GPT-6 <codex@openai.com>
Preserve command-handler results and outgoing reward state before lockstep reset. Isolated CPU diagnostic candidate; build and runtime validation delegated to executor.

Co-authored-by: Codex (GPT-6) <codex@openai.com>
Distinguish game team from policy-local controlled group before terminal reset.

Co-authored-by: Codex (GPT-6) <codex@openai.com>
…ification

Preserve main dependency closure in the separate integration candidate. Original 91483bb source and its CPU proof remain frozen; this merged source requires fresh qualification.

Co-authored-by: Codex (GPT-6) <codex@openai.com>
Keep the canonical main declaration; the isolated integration compile failed before tests on this duplicate. Preserve frozen 91483 qualification and source.

Co-authored-by: Codex (GPT-6) <codex@openai.com>
installPolicy replaces the ordinary VM with an unstructured policy. Mark that replacement as using scalar hero data so runHeroScript refreshes drafting and hero inputs, matching loadBots. Preserve the ordinary draft stateHash assertion.

Co-authored-by: Codex (GPT-6) <codex@openai.com>
…simulation

The complete lockstep suite passes with this exact test source and candidate95b5 under CPU1/RAM2GiB/no swap. Cover both controlled teams, trace-on/off canonical-state parity, potential-delta rewards and command identity across auto-reset. Existing ordinary draft stateHash assertion remains intact.

Co-authored-by: Codex (GPT-6) <codex@openai.com>
@relh relh changed the title Preserve GoTA lockstep parity and add bounded command attribution Trace GoTA commands and preserve ordinary draft inputs in lockstep Oct 4, 2026
@relh
relh marked this pull request as ready for review October 4, 2026 08:45
@relh relh changed the title Trace GoTA commands and preserve ordinary draft inputs in lockstep Initialize scalar hero inputs in GoTA lockstep policies Oct 6, 2026
@relh
relh marked this pull request as draft October 6, 2026 01:12
@relh relh changed the title Initialize scalar hero inputs in GoTA lockstep policies Use structured GoTA observations in training and export Oct 6, 2026
@relh
relh marked this pull request as ready for review October 6, 2026 01:24
relh and others added 2 commits October 6, 2026 15:13
Restore bounded opt-in command traces, actual simulator team identity, and outgoing reward state before automatic reset. Keep the structured VM and ordinary draft API intact.

Co-authored-by: Codex (GPT-6) <codex@openai.com>
Integrate current main and replace the broader training changes with the two missing HeroVm observation flags. Both native bridges now refresh scalar host data used by policies. The net diff against main is two changed lines with no added lines.

Co-authored-by: Codex (GPT-6) <codex@openai.com>
@relh relh changed the title Use structured GoTA observations in training and export Refresh scalar observations in native training bridges Oct 6, 2026
@relh relh closed this Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant