CPR reaches MTRL-level performance on average.
Adding previous-task directions makes the update better aligned with the mean gradient across the full task set in all nine measured comparisons.
Diagnosing and Harnessing Shared Reasoning in Continual RLVR
Continual Reasoning Gym (CRG) turns continual RLVR into a reproducible training and evaluation environment across text and visual reasoning. We identify shared reasoning: training on one task benefits others on average. Continual Prompt Replay (CPR) brings previous-task prompts into current-policy training to harness this pattern and is the only evaluated continual-learning method that reaches MTRL-level performance on average.
Seq. RLVR loses little performance on earlier tasks, but final performance remains below multitask RLVR. Decomposing final performance shows that preserving earlier tasks is only part of the problem. A continual method must also improve learning on the arriving and future tasks.
CPR stores prompts rather than old trajectories. At each update, previous-task prompts replace part of the current-task batch, and the current policy generates fresh responses for all sampled prompts. This changes the task-sampling distribution without increasing prompt or rollout counts per update.
Keep compact prompt records from earlier tasks.
Replace part of the arriving-task batch with previous-task prompts.
Sample and verify fresh responses with the current policy before updating.
Adding previous-task directions makes the update better aligned with the mean gradient across the full task set in all nine measured comparisons.
On LLM algorithmic reasoning, CPR reaches 63.3% FinalAvg. Reusing trajectories from an earlier policy reaches 47.5%, below the 49.9% no-replay result.
CRG organizes related reasoning tasks into staged training streams and evaluates every stage policy across the full sequence. The environment covers two procedurally generated text settings and three VisuLogic-derived visual settings.
@article{luo2026beyondforgetting,
title = {Beyond Forgetting: Diagnosing and Harnessing Shared Reasoning in Continual RLVR},
author = {Luo, Lirui and Zhang, Guoxi and Xu, Hongming and Li, Rongqing and Fang, Cong and Fan, Lifeng},
journal = {arXiv preprint arXiv:2608.18574},
eprint = {2608.18574},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2608.18574},
year = {2026}
}