10  Open Problems

M. C. Escher, Alfedena Abruzzi (1929).

10.1 Chapter Map

  • Open problems in RLVR.

10.2 Elicitation or creation

Research question. Does RLVR with outcome rewards create reasoning capability that was absent from the base model, or does it only reallocate probability mass toward solutions the base model could already sample?

10.3 Reward-hacking

Is there an over-optimization law for learned graders? Chapter 7’s quantitative anchor, the over-optimization curve of Gao et al., was measured for preference reward models.(Gao et al. 2023) No equivalent law exists for rubric aggregates or generative verifiers,(Zhang et al. 2025) so practitioners optimizing against them have no principled stopping criterion.

Can semantic faithfulness be measured directly? We currently score only what the verifier checks. Without an independent way to measure everything it misses, we cannot tell whether a high-scoring model truly learned the intended behavior or merely learned to satisfy the checks.

10.4 A predictive science of RL compute

Research question. Does RL post-training admit predictive scaling laws of the kind pretraining has, and what determines the asymptote a recipe saturates toward?

The first large systematic datapoint is ScaleRL: a study totaling more than 400,000 GPU-hours that fits sigmoidal compute-performance curves to RL training runs and validates them by predicting, from smaller runs, the trajectory of a single run extended to 100,000 GPU-hours. (Khatri et al. 2025)

Open questions follow directly. What sets the asymptote: the base model’s support, as the elicitation problem suggests, or removable inefficiencies of current recipes? Do fitted curves transfer across model families and task mixtures? How should a fixed budget split between pretraining, SFT, and RL? And does prolonged RL erode the plasticity it relies on?(Khan et al. 2026)

10.5 Credit assignment at horizon scale

Research question. At what horizon does a terminal outcome reward stop carrying usable learning signal, and can process-level signals be made simultaneously scalable and hack-resistant?

The agentic regime sharpens the question. Credit in reasoning RL spans one generation of 500 to 30K+ tokens; agentic RL spans hundreds of turns and 100K to 1M tokens, where a single episode-level scalar becomes increasingly uninformative, and the methods literature has no shared benchmark for comparing credit-assignment quality.(Zhang 2026) What is open: whether outcome rewards plus task structure suffice at these horizons, and which intermediate states are verifiable.

10.6 Self-improvement without external verification

Research question. Can a model’s own signals, such as confidence, self-consistency, or self-judgment, sustain RL improvement, or do self-reward loops inevitably collapse?

Gao, Leo, John Schulman, and Jacob Hilton. 2023. “Scaling Laws for Reward Model Overoptimization.” Proceedings of the 40th International Conference on Machine Learning (ICML). https://arxiv.org/abs/2210.10760.
Khan, Zohaib, Omer Tafveez, and Zoha Hayat Bhatti. 2026. “Plasticity Vs. Rigidity: The Impact of Low-Rank Adapters on Reasoning on a Micro-Budget.” arXiv Preprint arXiv:2601.06677. https://arxiv.org/abs/2601.06677.
Khatri, Devvrit, Lovish Madaan, Rishabh Tiwari, et al. 2025. “The Art of Scaling Reinforcement Learning Compute for LLMs.” arXiv Preprint arXiv:2510.13786. https://arxiv.org/abs/2510.13786.
Zhang, Chenchen. 2026. “From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models.” arXiv Preprint arXiv:2604.09459. https://arxiv.org/abs/2604.09459.
Zhang, Lunjun, Arian Hosseini, Hritik Bansal, Seyed Mehran Kazemi, Aviral Kumar, and Rishabh Agarwal. 2025. “Generative Verifiers: Reward Modeling as Next-Token Prediction.” Proceedings of the International Conference on Learning Representations (ICLR), ahead of print. https://doi.org/10.48550/arXiv.2408.15240.