Skip to content

Pull requests: areal-project/AReaL

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

fix(reward): validate scorer configuration before PRM execution
#1727 opened Sep 17, 2026 by Le8r0nJames Collaborator Loading…
4 of 15 tasks
fix(recover): restore colocated inference before admission
#1726 opened Sep 17, 2026 by Le8r0nJames Collaborator Loading…
4 of 15 tasks
fix(proxy): preserve configured chat template defaults
#1725 opened Sep 17, 2026 by Le8r0nJames Collaborator Loading…
4 of 15 tasks
fix(examples): retain attributed Arena failures with zero reward
#1724 opened Sep 17, 2026 by Le8r0nJames Collaborator Loading…
4 of 15 tasks
fix(engine): adapt colocated AWEX to SGLang scheduler APIs
#1723 opened Sep 17, 2026 by Le8r0nJames Collaborator Loading…
4 of 15 tasks
feat(infra): add opt-in sample-level rollout refill safe-to-test Ready to run unit-tests in a PR.
#1722 opened Sep 17, 2026 by dingzhiqiang Collaborator Loading…
6 of 9 tasks
feat(v2): support partial rollout groups safe-to-test Ready to run unit-tests in a PR.
#1721 opened Sep 16, 2026 by sitabulaixizawaluduo Collaborator Loading…
3 of 15 tasks
feat(engine): support native MTP-only and packed Qwen SFT safe-to-test Ready to run unit-tests in a PR.
#1719 opened Sep 16, 2026 by dingzhiqiang Collaborator Draft
fix(v2): propagate configured training RPC timeout to data proxy
#1712 opened Sep 15, 2026 by sitabulaixizawaluduo Collaborator Loading…
4 of 9 tasks
feat(dpo): support SimPO reference-free preference optimization with target margin
#1699 opened Sep 10, 2026 by hsusul Contributor Loading…
8 of 15 tasks
feat: add training throughput, FLOPs and MoE balance metrics safe-to-test Ready to run unit-tests in a PR.
#1693 opened Sep 10, 2026 by yulangz Collaborator Loading…
5 of 15 tasks
feat(trainer): support score reward clipping and teacher weight normalization in MOPD loss
#1690 opened Sep 9, 2026 by hsusul Contributor Loading…
7 of 9 tasks
fix(utils): treat missing packages as not matching version checks
#1689 opened Sep 9, 2026 by hsusul Contributor Loading…
1 task done
feat(archon): add NPU support for varlen attention
#1686 opened Sep 8, 2026 by 262913 Loading…
5 of 15 tasks
refactor: unify AWEX converters and adapters across v1 and v2 safe-to-test Ready to run unit-tests in a PR.
#1684 opened Sep 7, 2026 by sitabulaixizawaluduo Collaborator Loading…
1 task done
feat(trainer): support configurable entropy regularization in PPO and GRPO actor
#1680 opened Sep 5, 2026 by hsusul Contributor Loading…
3 tasks done
feat: support single-GPU GSM8K smoke runs
#1678 opened Sep 5, 2026 by CharlesXu-HQ Contributor Loading…
5 of 15 tasks
fix(openai): preserve multi-turn step rewards and raw baselines in group normalization
#1674 opened Sep 3, 2026 by hsusul Contributor Loading…
3 tasks done
fix(trainer): prevent reward and advantage explosion on remainder group normalization
#1673 opened Sep 3, 2026 by hsusul Contributor Loading…
3 tasks done
ProTip! Updated in the last three days: updated:>2026-09-14.