References

Agentica Team, and Together AI. 2025. “DeepSWE: Training a Fully Open-Sourced, State-of-the-Art Coding Agent by Scaling RL.” https://www.together.ai/blog/deepswe.
Ahmadian, Arash, Chris Cremer, Matthieu Gallé, et al. 2024. “Back to Basics: Revisiting REINFORCE-Style Optimization for Learning from Human Feedback in LLMs.” arXiv Preprint arXiv:2402.14740. https://arxiv.org/abs/2402.14740.
AI Security Institute. 2026. Cheating Behaviour in Frontier Model Evaluations. AISI Blog. https://www.aisi.gov.uk/blog/cheating-behaviour-in-frontier-model-evaluations.
Allen Institute for AI. 2025. “Olmo 3: Charting a Path Through the Model Flow to Lead Open-Source AI.” November. https://allenai.org/blog/olmo3.
Anthropic. 2025. “Writing Effective Tools for Agents – with Agents.” September 11. https://www.anthropic.com/engineering/writing-tools-for-agents.
Anthropic. 2026a. Improving Our Alignment and Security Efforts. Anthropic News. https://www.anthropic.com/news/improving-alignment-security-efforts.
Anthropic. 2026b. Investigating Three Real-World Incidents in Our Cybersecurity Evaluations. Anthropic News. https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals.
Anthropic. 2026c. System Card: Claude Opus 4.8. https://www-cdn.anthropic.com/0b4915911bb0d19eca5b5ee635c80fef830a37ea.pdf.
Bailey, Luke et al. 2026. Scaling Self-Play with Self-Guidance. https://arxiv.org/abs/2604.20209.
Baker, Bowen, Joost Huizinga, Leo Gao, et al. 2025. “Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.” arXiv Preprint arXiv:2503.11926. https://arxiv.org/abs/2503.11926.
Bengio, Yoshua, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009. “Curriculum Learning.” Proceedings of the 26th Annual International Conference on Machine Learning, 41–48. https://doi.org/10.1145/1553374.1553380.
Brown, William. 2025. Granular Format Rewards for Eliciting Mathematical Reasoning Capabilities in Small Language Models. GitHub Gist. https://gist.github.com/willccbb/4676755236bb08cab5f4e54a0475d6fb.
Carroll, Micah, Tomek Korbak, Zehao Dou, Bowen Baker, and Ian Kivlichan. 2026. Investigating the Consequences of Accidentally Grading CoT During RL. Blog post. https://alignment.openai.com/accidental-cot-grading/.
Ćemanović, Amar. 2026. Google Gemini Hacked Three Firms After Test Sandbox Exposed Web Access. CyberInsider. https://cyberinsider.com/google-gemini-hacked-three-firms-after-test-sandbox-exposed-web-access/.
Chen, Mark, Jerry Tworek, Heewoo Jun, et al. 2021. “Evaluating Large Language Models Trained on Code.” arXiv Preprint arXiv:2107.03374. https://arxiv.org/abs/2107.03374.
Cheng, Zhoujun et al. 2025. Revisiting Reinforcement Learning for LLM Reasoning from a Cross-Domain Perspective. https://arxiv.org/abs/2506.14965.
Cobbe, Karl, Vineet Kosaraju, Mohammad Bavarian, et al. 2021. “Training Verifiers to Solve Math Word Problems.” arXiv Preprint arXiv:2110.14168. https://arxiv.org/abs/2110.14168.
Coste, Thomas, Usman Anwar, Robert Kirk, and David Krueger. 2023. “Reward Model Ensembles Help Mitigate Overoptimization.” arXiv Preprint arXiv:2310.02743. https://arxiv.org/abs/2310.02743.
Cui, Ganqu et al. 2025. The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models. https://arxiv.org/abs/2505.22617.
DeepSeek-AI. 2026. “DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression.” arXiv Preprint arXiv:2609.19969. https://arxiv.org/abs/2609.19969.
DeepSeek-AI, Daya Guo, Dejian Yang, et al. 2025. “DeepSeek-R1 Incentivizes Reasoning in LLMs Through Reinforcement Learning.” Nature 645 (8081): 633–38. https://doi.org/10.1038/s41586-025-09422-z.
Dohare, Shibhansh, J. Fernando Hernandez-Garcia, Qingfeng Lan, Parash Rahman, A. Rupam Mahmood, and Richard S. Sutton. 2024. “Loss of Plasticity in Deep Continual Learning.” Nature 632 (8026): 768–74. https://doi.org/10.1038/s41586-024-07711-7.
Eisenstein, Jacob, Chirag Nagpal, Alekh Agarwal, et al. 2023. Helping or Herding? Reward Model Ensembles Mitigate but Do Not Eliminate Reward Hacking. https://arxiv.org/abs/2312.09244.
Feng, Lang et al. 2025. Group-in-Group Policy Optimization for LLM Agent Training. https://arxiv.org/abs/2505.10978.
Fu, Zixuan, Bingxiang He, Yuxin Zuo, et al. 2026. Rethinking on-Policy Distillation of Large Language Models II: One Training Example. https://arxiv.org/abs/2609.04172.
Gao, Leo, John Schulman, and Jacob Hilton. 2023. “Scaling Laws for Reward Model Overoptimization.” Proceedings of the 40th International Conference on Machine Learning (ICML). https://arxiv.org/abs/2210.10760.
Guan, Melody Y. et al. 2025. Monitoring Monitorability. https://arxiv.org/abs/2512.18311.
Gunjal, Anisha et al. 2025. Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains. https://arxiv.org/abs/2507.17746.
He, Bingxiang, Yuxin Zuo, et al. 2026. How Far Can Unsupervised RLVR Scale LLM Training? https://arxiv.org/abs/2603.08660.
He, Bowei, Yankai Chen, et al. 2026. Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning. https://arxiv.org/abs/2607.14171.
He, Horace. 2025. Defeating Nondeterminism in LLM Inference. Thinking Machines Lab: Connectionism. https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/.
Huan, Maggie et al. 2025. Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning. https://arxiv.org/abs/2507.00432.
Huang, Yuzhen et al. 2025. From Accuracy to Robustness: A Study of Rule- and Model-Based Verifiers in Mathematical Reasoning. https://arxiv.org/abs/2505.22203.
Hubert, Thomas, Rishi Mehta, Laurent Sartran, et al. 2025. “Olympiad-Level Formal Mathematical Reasoning with Reinforcement Learning.” Nature, ahead of print. https://doi.org/10.1038/s41586-025-09833-y.
Hugging Face. 2026. Security Incident Disclosure: July 2026. Hugging Face Blog. https://huggingface.co/blog/security-incident-july-2026.
Jackson, Jacob, Ben Trapani, Nathan Wang, and Wanqi Zhu. 2026. Improving Composer Through Real-Time RL. Cursor Blog. https://cursor.com/blog/real-time-rl-for-composer.
Jain, Naman, Jaskirat Singh, Manish Shetty, Liang Zheng, Koushik Sen, and Ion Stoica. 2025. “R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents.” arXiv Preprint arXiv:2504.07164. https://arxiv.org/abs/2504.07164.
Kaufmann, Max, David Lindner, Roland S. Zimmermann, and Rohin Shah. 2026. Aligned, Orthogonal or in-Conflict: When Can We Safely Optimize Chain-of-Thought? https://arxiv.org/abs/2603.30036.
Khan, Zohaib, Omer Tafveez, and Zoha Hayat Bhatti. 2026. “Plasticity Vs. Rigidity: The Impact of Low-Rank Adapters on Reasoning on a Micro-Budget.” arXiv Preprint arXiv:2601.06677. https://arxiv.org/abs/2601.06677.
Khatri, Devvrit, Lovish Madaan, Rishabh Tiwari, et al. 2025. “The Art of Scaling Reinforcement Learning Compute for LLMs.” arXiv Preprint arXiv:2510.13786. https://arxiv.org/abs/2510.13786.
Kim, Junsol, Shiyang Lai, Nino Scherrer, Blaise Agüera y Arcas, and James Evans. 2026. “Reasoning Models Generate Societies of Thought.” arXiv Preprint arXiv:2601.10825, ahead of print. https://doi.org/10.48550/arXiv.2601.10825.
Kim, Minsu, and Se-Young Yun. 2026. “Process-Verified Reinforcement Learning for Theorem Proving via Lean.” International Conference on Learning Representations. https://proceedings.iclr.cc/paper_files/paper/2026/hash/554a5982e26d0129f4038ad7a68fb3aa-Abstract-Conference.html.
Kimi Team. 2026. “Kimi K3: Open Frontier Intelligence.” arXiv Preprint arXiv:2607.24653. https://arxiv.org/abs/2607.24653.
Kimi Team, Angang Du, Bofei Gao, et al. 2025. “Kimi K1.5: Scaling Reinforcement Learning with LLMs.” arXiv Preprint arXiv:2501.12599. https://arxiv.org/abs/2501.12599.
Kojima, Takeshi, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. “Large Language Models Are Zero-Shot Reasoners.” arXiv Preprint arXiv:2205.11916. https://arxiv.org/abs/2205.11916.
Kwa, Thomas et al. 2025. Measuring AI Ability to Complete Long Software Tasks. https://arxiv.org/abs/2503.14499.
Kydlicek, Hynek. 2025. Math-Verify: Math Verification Library. V. 0.6.1. Released. https://github.com/huggingface/Math-Verify.
Lambert, Nathan, Jacob Morrison, Valentina Pyatkin, et al. 2024. “Tulu 3: Pushing Frontiers in Open Language Model Post-Training.” arXiv Preprint arXiv:2411.15124. https://arxiv.org/abs/2411.15124.
Lambert, Nathan, Valentina Pyatkin, Jacob Morrison, et al. 2024. “RewardBench: Evaluating Reward Models for Language Modeling.” arXiv Preprint arXiv:2403.13787. https://arxiv.org/abs/2403.13787.
Lanham, Tamera, Anna Chen, Ansh Radhakrishnan, et al. 2023. “Measuring Faithfulness in Chain-of-Thought Reasoning.” arXiv Preprint arXiv:2307.13702. https://arxiv.org/abs/2307.13702.
Le, Hung, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese, and Steven C. H. Hoi. 2022. “CodeRL: Mastering Code Generation Through Pretrained Models and Deep Reinforcement Learning.” arXiv Preprint arXiv:2207.01780. https://arxiv.org/abs/2207.01780.
Lightman, Hunter, Vineet Kosaraju, Yura Burda, et al. 2023. “Let’s Verify Step by Step.” arXiv Preprint arXiv:2305.20050. https://arxiv.org/abs/2305.20050.
Little, Bryce et al. 2026. Length Penalties Make Chain-of-Thought Less Monitorable. https://arxiv.org/abs/2607.09786.
Liu, Bo, Simon Yu, Yiding Jiang, Ao Qu, Andrew Zhao, et al. 2026. SPADE: Self-Play in Adaptive Synthetic Executable Environments. https://arxiv.org/abs/2608.19197.
Liu, Jiate, Yiqin Zhu, Kaiwen Xiao, et al. 2023. “RLTF: Reinforcement Learning from Unit Test Feedback.” Transactions on Machine Learning Research. https://arxiv.org/abs/2307.04349.
Liu, Jiawei, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. 2023. “Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation.” arXiv Preprint arXiv:2305.01210. https://arxiv.org/abs/2305.01210.
Liu, Mingjie et al. 2025. ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models. https://arxiv.org/abs/2505.24864.
Liu, Yixin, Yue Yu, et al. 2026. Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training. https://arxiv.org/abs/2603.12246.
Lu, Pan, Hritik Bansal, Tony Xia, et al. 2023. “MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.” arXiv Preprint arXiv:2310.02255. https://arxiv.org/abs/2310.02255.
MacDiarmid, Monte et al. 2025. Natural Emergent Misalignment from Reward Hacking in Production RL. https://arxiv.org/abs/2511.18397.
Mahmoud, Anas et al. 2026. Reward Hacking in Rubric-Based Reinforcement Learning. https://arxiv.org/abs/2605.12474.
Mayilvahanan, Prasanna, Ricardo Dominguez-Olmedo, et al. 2025. MATH-Beyond: A Benchmark for RL to Expand Beyond the Base Model. https://arxiv.org/abs/2510.11653.
Meta. 2026. Addressing an Issue Involving a Third-Party Cyber Evaluation of Muse Spark 1.1. Meta AI Research Blog. https://research.meta.ai/blog/addressing-third-party-testing-misconfiguration-muse-spark-1-1.
METR. 2026a. Brief Independent Investigation of Agents’ Behavior, Reasoning and Collaboration in the OpenAI / Hugging Face Hacking Incident. METR Blog. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/.
METR. 2026b. Claude Mythos Preview (Early) Time Horizon Estimate. Bluesky thread, May 8, 2026. https://bsky.app/profile/metr.org/post/3mlewafi6pc2u.
METR. 2026c. Time Horizon 1.1. Blog post. https://metr.org/blog/2026-1-29-time-horizon-1-1/.
MiniMax. 2026. The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence. https://arxiv.org/abs/2605.26494.
Nye, Maxwell, Anders Johan Andreassen, Guy Gur-Ari, et al. 2021. “Show Your Work: Scratchpads for Intermediate Computation with Language Models.” arXiv Preprint arXiv:2112.00114. https://arxiv.org/abs/2112.00114.
OpenAI. 2024. “Learning to Reason with LLMs.” September 12. https://openai.com/index/learning-to-reason-with-llms/.
OpenAI. 2026a. GPT-6 Astra System Card. System card. https://deploymentsafety.openai.com/gpt-6-astra.
OpenAI. 2026b. “Graders.” https://developers.openai.com/api/docs/guides/graders.
OpenAI. 2026c. OpenAI and Hugging Face Partner to Address Security Incident During Model Evaluation. OpenAI Blog. https://openai.com/index/hugging-face-model-evaluation-security-incident/.
OpenAI. 2026d. The Hugging Face Incident and the Road Ahead. OpenAI Blog. https://openai.com/index/hugging-face-incident-and-the-road-ahead/.
Pachocki, Jakub. 2026. An Alien Mind. Blog post. https://openai.com/index/an-alien-mind/.
Pan, Alexander, Kush Bhatia, and Jacob Steinhardt. 2022. “The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.” Proceedings of the International Conference on Learning Representations (ICLR). https://arxiv.org/abs/2201.03544.
Patel, Dwarkesh. 2025a. “Andrej Karpathy: AGI Is Still a Decade Away.” October 17. https://www.dwarkesh.com/p/andrej-karpathy.
Patel, Dwarkesh. 2025b. RL Is Even More Information Inefficient Than You Thought. Dwarkesh Podcast. https://www.dwarkesh.com/p/bits-per-sample.
Piché, Alexandre, Ehsan Kamalloo, Rafael Pardinas, Xiaoyin Chen, and Dzmitry Bahdanau. 2025. “PipelineRL: Faster on-Policy Reinforcement Learning for Long Sequence Generation.” arXiv Preprint arXiv:2509.19128. https://arxiv.org/abs/2509.19128.
Qi, Richard, Benjamin Wright, Monte MacDiarmid, and Evan Hubinger. 2026. Training a Misaligned Reward Seeker. Anthropic Alignment Science Blog. https://alignment.anthropic.com/2026/reward-seeker/.
Quiroz-Gutierrez, Marco. 2025. “Meta Is Reportedly Scrambling Multiple ‘War Rooms’ of Engineers to Figure Out How DeepSeek’s AI Is Beating Everyone Else at a Fraction of the Price.” January 27. https://fortune.com/2025/01/27/mark-zuckerberg-meta-llama-assembling-war-rooms-engineers-deepseek-ai-china/.
Rajan, Shreshth. 2026. Auditing Reward Hackability in Code RL Training Environments. https://arxiv.org/abs/2606.16062.
Schulman, John, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. “Proximal Policy Optimization Algorithms.” arXiv Preprint arXiv:1707.06347. https://arxiv.org/abs/1707.06347.
Shafayat, Sheikh et al. 2025. Can Large Reasoning Models Self-Train? https://arxiv.org/abs/2505.21444.
Shao, Zhihong, Peiyi Wang, Qihao Zhu, et al. 2024. “DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.” arXiv Preprint arXiv:2402.03300. https://arxiv.org/abs/2402.03300.
Shenfeld, Idan, Jyothish Pari, and Pulkit Agrawal. 2025. RL’s Razor: Why Online Reinforcement Learning Forgets Less. https://arxiv.org/abs/2509.04259.
Sheng, Guangming, Chi Zhang, Zilingfeng Ye, et al. 2024. “HybridFlow: A Flexible and Efficient RLHF Framework.” arXiv Preprint arXiv:2409.19256. https://arxiv.org/abs/2409.19256.
Shojaee, Parshin, Aneesh Jain, Sindhu Tipirneni, and Chandan K. Reddy. 2023. “Execution-Based Code Generation Using Deep Reinforcement Learning.” Transactions on Machine Learning Research. https://arxiv.org/abs/2301.13816.
Shumailov, Ilia et al. 2023. The Curse of Recursion: Training on Generated Data Makes Models Forget. https://arxiv.org/abs/2305.17493.
Skalse, Joar, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and David Krueger. 2022. “Defining and Characterizing Reward Hacking.” Advances in Neural Information Processing Systems (NeurIPS). https://arxiv.org/abs/2209.13085.
Snell, Charlie, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2024. “Scaling LLM Test-Time Compute Optimally Can Be More Effective Than Scaling Model Parameters.” arXiv Preprint arXiv:2408.03314. https://arxiv.org/abs/2408.03314.
Song, Yuda et al. 2024. Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models. https://arxiv.org/abs/2412.02674.
Sullivan, Michael, and Alexander Koller. 2025. “GRPO Is Secretly a Process Reward Model.” arXiv Preprint arXiv:2509.21154. https://arxiv.org/abs/2509.21154.
Sun, Lin, Chuang Liu, Xiaofeng Ma, Tao Yang, Weijia Lu, and Ning Wu. 2025. “FreePRM: Training Process Reward Models Without Ground Truth Process Labels.” arXiv Preprint arXiv:2506.03570, ahead of print. https://doi.org/10.48550/arXiv.2506.03570.
Sun, Weiwei et al. 2025. Scaling Long-Horizon LLM Agent via Context-Folding. https://arxiv.org/abs/2510.11967.
Sun, Yiyou et al. 2025. RL Grokking Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs? https://arxiv.org/abs/2509.21016.
Sutskever, Ilya. 2024. Sequence to Sequence Learning with Neural Networks: What a Decade. NeurIPS 2024 Test of Time Award talk. https://neurips.cc/virtual/2024/test-of-time/105032.
Tajwar, Fahim, Guanning Zeng, et al. 2026. Maximum Likelihood Reinforcement Learning. https://arxiv.org/abs/2602.02710.
Tan, Zelin, Hejia Geng, Xiaohang Yu, et al. 2025. “Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning.” arXiv Preprint arXiv:2509.25300. https://arxiv.org/abs/2509.25300.
Team OLMo, Allyson Ettinger, Amanda Bertsch, et al. 2025. “Olmo 3.” arXiv Preprint arXiv:2512.13961. https://arxiv.org/abs/2512.13961.
TMTPOST Global. 2025. “The Rise of Chinese Startup DeepSeek May Trigger u.s. Chip Investigation.” January 26. https://en.tmtpost.com/post/7436962.
Turpin, Miles, Julian Michael, Ethan Perez, and Samuel R. Bowman. 2023. “Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting.” arXiv Preprint arXiv:2305.04388. https://arxiv.org/abs/2305.04388.
Uesato, Jonathan, Nate Kushman, Ramana Kumar, et al. 2022. “Solving Math Word Problems with Process- and Outcome-Based Feedback.” arXiv Preprint arXiv:2211.14275. https://arxiv.org/abs/2211.14275.
Wan, Fanqi, Weizhou Shen, Shengyi Liao, Yingcheng Shi, et al. 2025. “QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning.” arXiv Preprint arXiv:2505.17667. https://arxiv.org/abs/2505.17667.
Wang, Li, Xiaodong Lu, et al. 2026. When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards. https://arxiv.org/abs/2605.25864.
Wang, Peiyi, Lei Li, Zhihong Shao, et al. 2024. “Math-Shepherd: Verify and Reinforce LLMs Step-by-Step Without Human Annotations.” arXiv Preprint arXiv:2312.08935. https://arxiv.org/abs/2312.08935.
Wang, Xuezhi, Jason Wei, Dale Schuurmans, et al. 2022. “Self-Consistency Improves Chain of Thought Reasoning in Language Models.” arXiv Preprint arXiv:2203.11171. https://arxiv.org/abs/2203.11171.
Wang, Yiping et al. 2025. Reinforcement Learning for Reasoning in Large Language Models with One Training Example. https://arxiv.org/abs/2504.20571.
Wang, Zhun, Nico Schiller, Hongwei Li, Srijiith Sesha Narayana, et al. 2026. “ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?” arXiv Preprint arXiv:2605.11086. https://arxiv.org/abs/2605.11086.
Wei, Jason. 2025. Asymmetry of Verification and Verifier’s Rule. Blog post. https://www.jasonwei.net/blog/asymmetry-of-verification-and-verifiers-law.
Wei, Jason, Xuezhi Wang, Dale Schuurmans, et al. 2022. “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.” arXiv Preprint arXiv:2201.11903. https://arxiv.org/abs/2201.11903.
Werra, Leandro von, Younes Belkada, Lewis Tunstall, et al. 2020. TRL: Transformers Reinforcement Learning. Released. https://github.com/huggingface/trl.
Wilf, Alex et al. 2025. Propose, Solve, Verify: Self-Play Through Formal Verification. https://arxiv.org/abs/2512.18160.
Wu, Fang, Weihao Xuan, Ximing Lu, et al. 2025. “The Invisible Leash: Why RLVR May or May Not Escape Its Origin.” arXiv Preprint arXiv:2507.14843. https://arxiv.org/abs/2507.14843.
Xie, Myron, Bryan Shan, Harrison Barclay, Minjae Kang, and Dylan Patel. 2026. Long Live the Short King: Why 4-Hi HBM Wins. SemiAnalysis newsletter. https://newsletter.semianalysis.com/p/long-live-the-short-king-why-4-hi.
Xie, Tianbao, Danyang Zhang, Jixuan Chen, et al. 2024. “OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.” arXiv Preprint arXiv:2404.07972. https://arxiv.org/abs/2404.07972.
Xin, Esther. 2026. Where the Verifier Fails: A Category-Level Audit of Reward Signals in RLVR. https://arxiv.org/abs/2609.01354.
Xin, Huajian, Daya Guo, Zhihong Shao, et al. 2024. “DeepSeek-Prover: Advancing Theorem Proving in LLMs Through Large-Scale Synthetic Data.” arXiv Preprint arXiv:2405.14333. https://arxiv.org/abs/2405.14333.
Xin, Huajian, Z. Z. Ren, Junxiao Song, et al. 2024. “DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search.” arXiv Preprint arXiv:2408.08152. https://arxiv.org/abs/2408.08152.
Yang, Jinming, Zheng Hu, Chuxian Qiu, Zhenyu Deng, Xinshan Jiao, and Tao Zhou. 2026. “Quantifying and Mitigating Self-Preference Bias of LLM Judges.” arXiv Preprint arXiv:2604.22891, ahead of print. https://doi.org/10.48550/arXiv.2604.22891.
Yang, Minglai, Xinyu Guo, et al. 2026. Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL. https://arxiv.org/abs/2608.11669.
Yao, Feng, Liyuan Liu, Dinghuai Zhang, Chengyu Dong, Jingbo Shang, and Jianfeng Gao. 2025. Your Efficient RL Framework Secretly Brings You Off-Policy RL Training. Blog post. https://fengyao.notion.site/off-policy-rl.
Yu, Qiying, Zheng Zhang, Ruofei Zhu, et al. 2025. “DAPO: An Open-Source LLM Reinforcement Learning System at Scale.” arXiv Preprint arXiv:2503.14476. https://arxiv.org/abs/2503.14476.
Yu, Tianyu et al. 2025. RLPR: Extrapolating RLVR to General Domains Without Verifiers. https://arxiv.org/abs/2506.18254.
Yuan, Lifan et al. 2024. “Free Process Rewards Without Process Labels.” arXiv Preprint arXiv:2412.01981. https://arxiv.org/abs/2412.01981.
Yuan, Lifan et al. 2025. From f(x) and g(x) to f(g(x)): LLMs Learn New Skills in RL by Composing Old Ones. https://arxiv.org/abs/2509.25123.
Yuan, Suqin et al. 2026. Understanding Diversity Collapse in RLVR via the Lens of Overtraining. https://arxiv.org/abs/2606.15455.
Yue, Yang, Zhiqi Chen, Rui Lu, et al. 2025. “Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?” Advances in Neural Information Processing Systems 38 (NeurIPS 2025), Oral. https://arxiv.org/abs/2504.13837.
Zhai, Zhiyuan et al. 2026. Does RL Expand the Capability Boundary of LLM Agents? A PASS@(k,t) Analysis. https://arxiv.org/abs/2604.14877.
Zhang, Charlie et al. 2025. On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models. https://arxiv.org/abs/2512.07783.
Zhang, Chenchen. 2026. “From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models.” arXiv Preprint arXiv:2604.09459. https://arxiv.org/abs/2604.09459.
Zhang, Chuyifei. 2026. When the Reward Suite Is Leaky: A Preregistered Causal Contrast of Natural Verifier False Positives in RLVR. https://arxiv.org/abs/2607.11022.
Zhang, Haiyue. 2026. Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay. https://arxiv.org/abs/2608.19760.
Zhang, Jiajie, Yushi Bai, Xin Lv, et al. 2024. “LongCite: Enabling LLMs to Generate Fine-Grained Citations in Long-Context QA.” arXiv Preprint arXiv:2409.02897. https://arxiv.org/abs/2409.02897.
Zhang, Lunjun, Arian Hosseini, Hritik Bansal, Seyed Mehran Kazemi, Aviral Kumar, and Rishabh Agarwal. 2025. “Generative Verifiers: Reward Modeling as Next-Token Prediction.” Proceedings of the International Conference on Learning Representations (ICLR), ahead of print. https://doi.org/10.48550/arXiv.2408.15240.
Zhang, Yanzhi et al. 2025. No Free Lunch: Rethinking Internal Feedback for LLM Reasoning. https://arxiv.org/abs/2506.17219.
Zhao, Andrew et al. 2025. Absolute Zero: Reinforced Self-Play Reasoning with Zero Data. https://arxiv.org/abs/2505.03335.
Zheng, Lianmin, Wei-Lin Chiang, Ying Sheng, et al. 2023. “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.” arXiv Preprint arXiv:2306.05685. https://arxiv.org/abs/2306.05685.
Zhong, Ziqian et al. 2025. ImpossibleBench: Measuring LLMs’ Propensity of Exploiting Test Cases. https://arxiv.org/abs/2510.20270.
Zhou, Shuyan, Frank F. Xu, Hao Zhu, et al. 2023. “WebArena: A Realistic Web Environment for Building Autonomous Agents.” arXiv Preprint arXiv:2307.13854. https://arxiv.org/abs/2307.13854.
Zhou, Yujun, Zhenwen Liang, Haolin Liu, et al. 2025. “Evolving Language Models Without Labels: Majority Drives Selection, Novelty Promotes Variation.” arXiv Preprint arXiv:2509.15194. https://arxiv.org/abs/2509.15194.
Zhu, Xinyu et al. 2025. The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning. https://arxiv.org/abs/2506.01347.
Zuo, Yuxin et al. 2025. TTRL: Test-Time Reinforcement Learning. https://arxiv.org/abs/2504.16084.