References
Agentica Team, and Together AI. 2025. “DeepSWE: Training a Fully
Open-Sourced, State-of-the-Art Coding Agent by Scaling RL.” https://www.together.ai/blog/deepswe.
Ahmadian, Arash, Chris Cremer, Matthieu Gallé, et al. 2024. “Back
to Basics: Revisiting REINFORCE-Style Optimization for Learning from
Human Feedback in LLMs.” arXiv Preprint
arXiv:2402.14740. https://arxiv.org/abs/2402.14740.
AI Security Institute. 2026. Cheating Behaviour in Frontier Model
Evaluations. AISI Blog. https://www.aisi.gov.uk/blog/cheating-behaviour-in-frontier-model-evaluations.
Allen Institute for AI. 2025. “Olmo 3: Charting a Path Through the
Model Flow to Lead Open-Source AI.” November. https://allenai.org/blog/olmo3.
Anthropic. 2025. “Writing Effective Tools for Agents – with
Agents.” September 11. https://www.anthropic.com/engineering/writing-tools-for-agents.
Anthropic. 2026a. Improving Our Alignment and Security Efforts.
Anthropic News. https://www.anthropic.com/news/improving-alignment-security-efforts.
Anthropic. 2026b. Investigating Three Real-World Incidents in Our
Cybersecurity Evaluations. Anthropic News. https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals.
Anthropic. 2026c. System Card: Claude Opus 4.8. https://www-cdn.anthropic.com/0b4915911bb0d19eca5b5ee635c80fef830a37ea.pdf.
Bailey, Luke et al. 2026. Scaling
Self-Play with Self-Guidance. https://arxiv.org/abs/2604.20209.
Baker, Bowen, Joost Huizinga, Leo Gao, et al. 2025. “Monitoring
Reasoning Models for Misbehavior and the Risks of Promoting
Obfuscation.” arXiv Preprint arXiv:2503.11926. https://arxiv.org/abs/2503.11926.
Bengio, Yoshua, Jérôme Louradour, Ronan Collobert, and Jason Weston.
2009. “Curriculum Learning.” Proceedings of the 26th
Annual International Conference on Machine Learning, 41–48. https://doi.org/10.1145/1553374.1553380.
Brown, William. 2025. Granular Format Rewards for Eliciting
Mathematical Reasoning Capabilities in Small Language Models.
GitHub Gist. https://gist.github.com/willccbb/4676755236bb08cab5f4e54a0475d6fb.
Carroll, Micah, Tomek Korbak, Zehao Dou, Bowen Baker, and Ian Kivlichan.
2026. Investigating the Consequences of Accidentally Grading CoT
During RL. Blog post. https://alignment.openai.com/accidental-cot-grading/.
Ćemanović, Amar. 2026. Google Gemini Hacked Three Firms After Test
Sandbox Exposed Web Access. CyberInsider. https://cyberinsider.com/google-gemini-hacked-three-firms-after-test-sandbox-exposed-web-access/.
Chen, Mark, Jerry Tworek, Heewoo Jun, et al. 2021. “Evaluating
Large Language Models Trained on Code.” arXiv Preprint
arXiv:2107.03374. https://arxiv.org/abs/2107.03374.
Cheng, Zhoujun et al. 2025. Revisiting
Reinforcement Learning for LLM Reasoning from a Cross-Domain
Perspective. https://arxiv.org/abs/2506.14965.
Cobbe, Karl, Vineet Kosaraju, Mohammad Bavarian, et al. 2021.
“Training Verifiers to Solve Math Word Problems.” arXiv
Preprint arXiv:2110.14168. https://arxiv.org/abs/2110.14168.
Coste, Thomas, Usman Anwar, Robert Kirk, and David Krueger. 2023.
“Reward Model Ensembles Help Mitigate Overoptimization.”
arXiv Preprint arXiv:2310.02743. https://arxiv.org/abs/2310.02743.
Cui, Ganqu et al. 2025. The Entropy
Mechanism of Reinforcement Learning for Reasoning Language Models.
https://arxiv.org/abs/2505.22617.
DeepSeek-AI. 2026. “DeepSeek-V4.1-Flash: Pushing the Limits of KV
Cache Compression.” arXiv Preprint arXiv:2609.19969. https://arxiv.org/abs/2609.19969.
DeepSeek-AI, Daya Guo, Dejian Yang, et al.
2025. “DeepSeek-R1 Incentivizes Reasoning in LLMs Through
Reinforcement Learning.” Nature 645 (8081): 633–38. https://doi.org/10.1038/s41586-025-09422-z.
Dohare, Shibhansh, J. Fernando Hernandez-Garcia, Qingfeng Lan, Parash
Rahman, A. Rupam Mahmood, and Richard S. Sutton. 2024. “Loss of
Plasticity in Deep Continual Learning.” Nature 632
(8026): 768–74. https://doi.org/10.1038/s41586-024-07711-7.
Eisenstein, Jacob, Chirag Nagpal, Alekh Agarwal, et al. 2023.
Helping or Herding? Reward Model Ensembles Mitigate but Do Not
Eliminate Reward Hacking. https://arxiv.org/abs/2312.09244.
Feng, Lang et al. 2025. Group-in-Group
Policy Optimization for LLM Agent Training. https://arxiv.org/abs/2505.10978.
Fu, Zixuan, Bingxiang He, Yuxin Zuo, et al. 2026. Rethinking
on-Policy Distillation of Large Language Models II: One Training
Example. https://arxiv.org/abs/2609.04172.
Gao, Leo, John Schulman, and Jacob Hilton. 2023. “Scaling Laws for
Reward Model Overoptimization.” Proceedings of the 40th
International Conference on Machine Learning (ICML). https://arxiv.org/abs/2210.10760.
Guan, Melody Y. et al. 2025. Monitoring
Monitorability. https://arxiv.org/abs/2512.18311.
Gunjal, Anisha et al. 2025. Rubrics as
Rewards: Reinforcement Learning Beyond Verifiable Domains. https://arxiv.org/abs/2507.17746.
He, Bingxiang, Yuxin Zuo, et al. 2026.
How Far Can Unsupervised RLVR Scale LLM Training? https://arxiv.org/abs/2603.08660.
He, Bowei, Yankai Chen, et al. 2026.
Branching Policy Optimization: Sandbox-Native Language Agent
Reinforcement Learning. https://arxiv.org/abs/2607.14171.
He, Horace. 2025. Defeating Nondeterminism in LLM Inference.
Thinking Machines Lab: Connectionism. https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/.
Huan, Maggie et al. 2025. Does Math
Reasoning Improve General LLM Capabilities? Understanding
Transferability of LLM Reasoning. https://arxiv.org/abs/2507.00432.
Huang, Yuzhen et al. 2025. From Accuracy
to Robustness: A Study of Rule- and Model-Based Verifiers in
Mathematical Reasoning. https://arxiv.org/abs/2505.22203.
Hubert, Thomas, Rishi Mehta, Laurent Sartran, et
al. 2025. “Olympiad-Level Formal Mathematical Reasoning
with Reinforcement Learning.” Nature, ahead of print. https://doi.org/10.1038/s41586-025-09833-y.
Hugging Face. 2026. Security Incident Disclosure: July 2026.
Hugging Face Blog. https://huggingface.co/blog/security-incident-july-2026.
Jackson, Jacob, Ben Trapani, Nathan Wang, and Wanqi Zhu. 2026.
Improving Composer Through Real-Time RL. Cursor Blog. https://cursor.com/blog/real-time-rl-for-composer.
Jain, Naman, Jaskirat Singh, Manish Shetty, Liang Zheng, Koushik Sen,
and Ion Stoica. 2025. “R2E-Gym: Procedural Environments and Hybrid
Verifiers for Scaling Open-Weights SWE Agents.” arXiv
Preprint arXiv:2504.07164. https://arxiv.org/abs/2504.07164.
Kaufmann, Max, David Lindner, Roland S. Zimmermann, and Rohin Shah.
2026. Aligned, Orthogonal or in-Conflict: When Can We Safely
Optimize Chain-of-Thought? https://arxiv.org/abs/2603.30036.
Khan, Zohaib, Omer Tafveez, and Zoha Hayat Bhatti. 2026.
“Plasticity Vs. Rigidity: The Impact of Low-Rank Adapters on
Reasoning on a Micro-Budget.” arXiv Preprint
arXiv:2601.06677. https://arxiv.org/abs/2601.06677.
Khatri, Devvrit, Lovish Madaan, Rishabh Tiwari, et al. 2025. “The
Art of Scaling Reinforcement Learning Compute for LLMs.”
arXiv Preprint arXiv:2510.13786. https://arxiv.org/abs/2510.13786.
Kim, Junsol, Shiyang Lai, Nino Scherrer, Blaise Agüera y Arcas, and
James Evans. 2026. “Reasoning Models Generate Societies of
Thought.” arXiv Preprint arXiv:2601.10825, ahead of
print. https://doi.org/10.48550/arXiv.2601.10825.
Kim, Minsu, and Se-Young Yun. 2026. “Process-Verified
Reinforcement Learning for Theorem Proving via Lean.”
International Conference on Learning Representations. https://proceedings.iclr.cc/paper_files/paper/2026/hash/554a5982e26d0129f4038ad7a68fb3aa-Abstract-Conference.html.
Kimi Team. 2026. “Kimi K3: Open Frontier Intelligence.”
arXiv Preprint arXiv:2607.24653. https://arxiv.org/abs/2607.24653.
Kimi Team, Angang Du, Bofei Gao, et al.
2025. “Kimi K1.5: Scaling Reinforcement Learning with
LLMs.” arXiv Preprint arXiv:2501.12599. https://arxiv.org/abs/2501.12599.
Kojima, Takeshi, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and
Yusuke Iwasawa. 2022. “Large Language Models Are Zero-Shot
Reasoners.” arXiv Preprint arXiv:2205.11916. https://arxiv.org/abs/2205.11916.
Kwa, Thomas et al. 2025. Measuring AI
Ability to Complete Long Software Tasks. https://arxiv.org/abs/2503.14499.
Kydlicek, Hynek. 2025. Math-Verify: Math Verification Library.
V. 0.6.1. Released. https://github.com/huggingface/Math-Verify.
Lambert, Nathan, Jacob Morrison, Valentina Pyatkin, et al. 2024.
“Tulu 3: Pushing Frontiers in Open Language Model
Post-Training.” arXiv Preprint arXiv:2411.15124. https://arxiv.org/abs/2411.15124.
Lambert, Nathan, Valentina Pyatkin, Jacob Morrison, et al. 2024.
“RewardBench: Evaluating Reward Models for Language
Modeling.” arXiv Preprint arXiv:2403.13787. https://arxiv.org/abs/2403.13787.
Lanham, Tamera, Anna Chen, Ansh Radhakrishnan, et al. 2023.
“Measuring Faithfulness in Chain-of-Thought Reasoning.”
arXiv Preprint arXiv:2307.13702. https://arxiv.org/abs/2307.13702.
Le, Hung, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese, and Steven
C. H. Hoi. 2022. “CodeRL: Mastering Code Generation Through
Pretrained Models and Deep Reinforcement Learning.” arXiv
Preprint arXiv:2207.01780. https://arxiv.org/abs/2207.01780.
Lightman, Hunter, Vineet Kosaraju, Yura Burda, et al. 2023. “Let’s
Verify Step by Step.” arXiv Preprint arXiv:2305.20050.
https://arxiv.org/abs/2305.20050.
Little, Bryce et al. 2026. Length
Penalties Make Chain-of-Thought Less Monitorable. https://arxiv.org/abs/2607.09786.
Liu, Bo, Simon Yu, Yiding Jiang, Ao Qu, Andrew
Zhao, et al. 2026. SPADE: Self-Play in Adaptive Synthetic
Executable Environments. https://arxiv.org/abs/2608.19197.
Liu, Jiate, Yiqin Zhu, Kaiwen Xiao, et al. 2023. “RLTF:
Reinforcement Learning from Unit Test Feedback.” Transactions
on Machine Learning Research. https://arxiv.org/abs/2307.04349.
Liu, Jiawei, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. 2023.
“Is Your Code Generated by ChatGPT Really Correct? Rigorous
Evaluation of Large Language Models for Code Generation.”
arXiv Preprint arXiv:2305.01210. https://arxiv.org/abs/2305.01210.
Liu, Mingjie et al. 2025. ProRL:
Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large
Language Models. https://arxiv.org/abs/2505.24864.
Liu, Yixin, Yue Yu, et al. 2026.
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM
Post-Training. https://arxiv.org/abs/2603.12246.
Lu, Pan, Hritik Bansal, Tony Xia, et al. 2023. “MathVista:
Evaluating Mathematical Reasoning of Foundation Models in Visual
Contexts.” arXiv Preprint arXiv:2310.02255. https://arxiv.org/abs/2310.02255.
MacDiarmid, Monte et al. 2025. Natural
Emergent Misalignment from Reward Hacking in Production RL. https://arxiv.org/abs/2511.18397.
Mahmoud, Anas et al. 2026. Reward
Hacking in Rubric-Based Reinforcement Learning. https://arxiv.org/abs/2605.12474.
Mayilvahanan, Prasanna, Ricardo Dominguez-Olmedo,
et al. 2025. MATH-Beyond: A Benchmark for RL to Expand Beyond
the Base Model. https://arxiv.org/abs/2510.11653.
Meta. 2026. Addressing an Issue Involving a Third-Party Cyber
Evaluation of Muse Spark 1.1. Meta AI Research Blog. https://research.meta.ai/blog/addressing-third-party-testing-misconfiguration-muse-spark-1-1.
METR. 2026a. Brief Independent Investigation of Agents’ Behavior,
Reasoning and Collaboration in the OpenAI / Hugging Face Hacking
Incident. METR Blog. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/.
METR. 2026b. Claude Mythos Preview (Early) Time Horizon
Estimate. Bluesky thread, May 8, 2026. https://bsky.app/profile/metr.org/post/3mlewafi6pc2u.
METR. 2026c. Time Horizon 1.1. Blog post. https://metr.org/blog/2026-1-29-time-horizon-1-1/.
MiniMax. 2026. The MiniMax-M2 Series: Mini Activations Unleashing
Max Real-World Intelligence. https://arxiv.org/abs/2605.26494.
Nye, Maxwell, Anders Johan Andreassen, Guy Gur-Ari, et al. 2021.
“Show Your Work: Scratchpads for Intermediate Computation with
Language Models.” arXiv Preprint arXiv:2112.00114. https://arxiv.org/abs/2112.00114.
OpenAI. 2024. “Learning to Reason with LLMs.” September 12.
https://openai.com/index/learning-to-reason-with-llms/.
OpenAI. 2026a. GPT-6 Astra System Card. System card. https://deploymentsafety.openai.com/gpt-6-astra.
OpenAI. 2026b. “Graders.” https://developers.openai.com/api/docs/guides/graders.
OpenAI. 2026c. OpenAI and Hugging Face Partner to Address Security
Incident During Model Evaluation. OpenAI Blog. https://openai.com/index/hugging-face-model-evaluation-security-incident/.
OpenAI. 2026d. The Hugging Face Incident and the Road Ahead.
OpenAI Blog. https://openai.com/index/hugging-face-incident-and-the-road-ahead/.
Pachocki, Jakub. 2026. An Alien Mind. Blog post. https://openai.com/index/an-alien-mind/.
Pan, Alexander, Kush Bhatia, and Jacob Steinhardt. 2022. “The
Effects of Reward Misspecification: Mapping and Mitigating Misaligned
Models.” Proceedings of the International Conference on
Learning Representations (ICLR). https://arxiv.org/abs/2201.03544.
Patel, Dwarkesh. 2025a. “Andrej Karpathy: AGI Is Still a Decade
Away.” October 17. https://www.dwarkesh.com/p/andrej-karpathy.
Patel, Dwarkesh. 2025b. RL Is Even More Information Inefficient Than
You Thought. Dwarkesh Podcast. https://www.dwarkesh.com/p/bits-per-sample.
Piché, Alexandre, Ehsan Kamalloo, Rafael Pardinas, Xiaoyin Chen, and
Dzmitry Bahdanau. 2025. “PipelineRL: Faster on-Policy
Reinforcement Learning for Long Sequence Generation.” arXiv
Preprint arXiv:2509.19128. https://arxiv.org/abs/2509.19128.
Qi, Richard, Benjamin Wright, Monte MacDiarmid, and Evan Hubinger. 2026.
Training a Misaligned Reward Seeker. Anthropic Alignment
Science Blog. https://alignment.anthropic.com/2026/reward-seeker/.
Quiroz-Gutierrez, Marco. 2025. “Meta Is Reportedly Scrambling
Multiple ‘War Rooms’ of Engineers to Figure Out How
DeepSeek’s AI Is Beating Everyone Else at a Fraction of the
Price.” January 27. https://fortune.com/2025/01/27/mark-zuckerberg-meta-llama-assembling-war-rooms-engineers-deepseek-ai-china/.
Rajan, Shreshth. 2026. Auditing Reward Hackability in Code RL
Training Environments. https://arxiv.org/abs/2606.16062.
Schulman, John, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg
Klimov. 2017. “Proximal Policy Optimization Algorithms.”
arXiv Preprint arXiv:1707.06347. https://arxiv.org/abs/1707.06347.
Shafayat, Sheikh et al. 2025. Can Large
Reasoning Models Self-Train? https://arxiv.org/abs/2505.21444.
Shao, Zhihong, Peiyi Wang, Qihao Zhu, et al. 2024. “DeepSeekMath:
Pushing the Limits of Mathematical Reasoning in Open Language
Models.” arXiv Preprint arXiv:2402.03300. https://arxiv.org/abs/2402.03300.
Shenfeld, Idan, Jyothish Pari, and Pulkit Agrawal. 2025. RL’s Razor:
Why Online Reinforcement Learning Forgets Less. https://arxiv.org/abs/2509.04259.
Sheng, Guangming, Chi Zhang, Zilingfeng Ye, et al. 2024.
“HybridFlow: A Flexible and Efficient RLHF Framework.”
arXiv Preprint arXiv:2409.19256. https://arxiv.org/abs/2409.19256.
Shojaee, Parshin, Aneesh Jain, Sindhu Tipirneni, and Chandan K. Reddy.
2023. “Execution-Based Code Generation Using Deep Reinforcement
Learning.” Transactions on Machine Learning Research. https://arxiv.org/abs/2301.13816.
Shumailov, Ilia et al. 2023. The Curse
of Recursion: Training on Generated Data Makes Models Forget. https://arxiv.org/abs/2305.17493.
Skalse, Joar, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and David
Krueger. 2022. “Defining and Characterizing Reward
Hacking.” Advances in Neural Information Processing Systems
(NeurIPS). https://arxiv.org/abs/2209.13085.
Snell, Charlie, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2024.
“Scaling LLM Test-Time Compute Optimally Can Be More Effective
Than Scaling Model Parameters.” arXiv Preprint
arXiv:2408.03314. https://arxiv.org/abs/2408.03314.
Song, Yuda et al. 2024. Mind the Gap:
Examining the Self-Improvement Capabilities of Large Language
Models. https://arxiv.org/abs/2412.02674.
Sullivan, Michael, and Alexander Koller. 2025. “GRPO Is Secretly a
Process Reward Model.” arXiv Preprint arXiv:2509.21154.
https://arxiv.org/abs/2509.21154.
Sun, Lin, Chuang Liu, Xiaofeng Ma, Tao Yang, Weijia Lu, and Ning Wu.
2025. “FreePRM: Training Process Reward Models Without Ground
Truth Process Labels.” arXiv Preprint arXiv:2506.03570,
ahead of print. https://doi.org/10.48550/arXiv.2506.03570.
Sun, Weiwei et al. 2025. Scaling
Long-Horizon LLM Agent via Context-Folding. https://arxiv.org/abs/2510.11967.
Sun, Yiyou et al. 2025. RL Grokking
Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs? https://arxiv.org/abs/2509.21016.
Sutskever, Ilya. 2024. Sequence to Sequence Learning with Neural
Networks: What a Decade. NeurIPS 2024 Test of Time Award talk. https://neurips.cc/virtual/2024/test-of-time/105032.
Tajwar, Fahim, Guanning Zeng, et al. 2026.
Maximum Likelihood Reinforcement Learning. https://arxiv.org/abs/2602.02710.
Tan, Zelin, Hejia Geng, Xiaohang Yu, et al.
2025. “Scaling Behaviors of LLM Reinforcement Learning
Post-Training: An Empirical Study in Mathematical Reasoning.”
arXiv Preprint arXiv:2509.25300. https://arxiv.org/abs/2509.25300.
Team OLMo, Allyson Ettinger, Amanda Bertsch, et
al. 2025. “Olmo 3.” arXiv Preprint
arXiv:2512.13961. https://arxiv.org/abs/2512.13961.
TMTPOST Global. 2025. “The Rise of Chinese Startup DeepSeek May
Trigger u.s. Chip Investigation.” January 26. https://en.tmtpost.com/post/7436962.
Turpin, Miles, Julian Michael, Ethan Perez, and Samuel R. Bowman. 2023.
“Language Models Don’t Always Say What They Think: Unfaithful
Explanations in Chain-of-Thought Prompting.” arXiv Preprint
arXiv:2305.04388. https://arxiv.org/abs/2305.04388.
Uesato, Jonathan, Nate Kushman, Ramana Kumar, et al. 2022.
“Solving Math Word Problems with Process- and Outcome-Based
Feedback.” arXiv Preprint arXiv:2211.14275. https://arxiv.org/abs/2211.14275.
Wan, Fanqi, Weizhou Shen, Shengyi Liao, Yingcheng
Shi, et al. 2025. “QwenLong-L1: Towards Long-Context Large
Reasoning Models with Reinforcement Learning.” arXiv Preprint
arXiv:2505.17667. https://arxiv.org/abs/2505.17667.
Wang, Li, Xiaodong Lu, et al. 2026. When
Self-Belief Misleads: Active Label Acquisition for Reinforcement
Learning with Verifiable Rewards. https://arxiv.org/abs/2605.25864.
Wang, Peiyi, Lei Li, Zhihong Shao, et al. 2024. “Math-Shepherd:
Verify and Reinforce LLMs Step-by-Step Without Human
Annotations.” arXiv Preprint arXiv:2312.08935. https://arxiv.org/abs/2312.08935.
Wang, Xuezhi, Jason Wei, Dale Schuurmans, et al. 2022.
“Self-Consistency Improves Chain of Thought Reasoning in Language
Models.” arXiv Preprint arXiv:2203.11171. https://arxiv.org/abs/2203.11171.
Wang, Yiping et al. 2025. Reinforcement
Learning for Reasoning in Large Language Models with One Training
Example. https://arxiv.org/abs/2504.20571.
Wang, Zhun, Nico Schiller, Hongwei Li, Srijiith
Sesha Narayana, et al. 2026. “ExploitGym: Can AI Agents
Turn Security Vulnerabilities into Real Attacks?” arXiv
Preprint arXiv:2605.11086. https://arxiv.org/abs/2605.11086.
Wei, Jason. 2025. Asymmetry of Verification and Verifier’s
Rule. Blog post. https://www.jasonwei.net/blog/asymmetry-of-verification-and-verifiers-law.
Wei, Jason, Xuezhi Wang, Dale Schuurmans, et al. 2022.
“Chain-of-Thought Prompting Elicits Reasoning in Large Language
Models.” arXiv Preprint arXiv:2201.11903. https://arxiv.org/abs/2201.11903.
Werra, Leandro von, Younes Belkada, Lewis Tunstall, et al. 2020.
TRL: Transformers Reinforcement Learning. Released. https://github.com/huggingface/trl.
Wilf, Alex et al. 2025. Propose, Solve,
Verify: Self-Play Through Formal Verification. https://arxiv.org/abs/2512.18160.
Wu, Fang, Weihao Xuan, Ximing Lu, et al. 2025. “The Invisible
Leash: Why RLVR May or May Not Escape Its Origin.” arXiv
Preprint arXiv:2507.14843. https://arxiv.org/abs/2507.14843.
Xie, Myron, Bryan Shan, Harrison Barclay, Minjae Kang, and Dylan Patel.
2026. Long Live the Short King: Why 4-Hi HBM Wins. SemiAnalysis
newsletter. https://newsletter.semianalysis.com/p/long-live-the-short-king-why-4-hi.
Xie, Tianbao, Danyang Zhang, Jixuan Chen, et al. 2024. “OSWorld:
Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer
Environments.” arXiv Preprint arXiv:2404.07972. https://arxiv.org/abs/2404.07972.
Xin, Esther. 2026. Where the Verifier Fails: A Category-Level Audit
of Reward Signals in RLVR. https://arxiv.org/abs/2609.01354.
Xin, Huajian, Daya Guo, Zhihong Shao, et al. 2024.
“DeepSeek-Prover: Advancing Theorem Proving in LLMs Through
Large-Scale Synthetic Data.” arXiv Preprint
arXiv:2405.14333. https://arxiv.org/abs/2405.14333.
Xin, Huajian, Z. Z. Ren, Junxiao Song, et al. 2024.
“DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for
Reinforcement Learning and Monte-Carlo Tree Search.” arXiv
Preprint arXiv:2408.08152. https://arxiv.org/abs/2408.08152.
Yang, Jinming, Zheng Hu, Chuxian Qiu, Zhenyu Deng, Xinshan Jiao, and Tao
Zhou. 2026. “Quantifying and Mitigating Self-Preference Bias of
LLM Judges.” arXiv Preprint arXiv:2604.22891, ahead of
print. https://doi.org/10.48550/arXiv.2604.22891.
Yang, Minglai, Xinyu Guo, et al. 2026.
Rubric Dropout: A Simple Way to Mitigate Reward Hacking in
Rubric-as-Reward RL. https://arxiv.org/abs/2608.11669.
Yao, Feng, Liyuan Liu, Dinghuai Zhang, Chengyu Dong, Jingbo Shang, and
Jianfeng Gao. 2025. Your Efficient RL Framework Secretly Brings You
Off-Policy RL Training. Blog post. https://fengyao.notion.site/off-policy-rl.
Yu, Qiying, Zheng Zhang, Ruofei Zhu, et al.
2025. “DAPO: An Open-Source LLM Reinforcement Learning System at
Scale.” arXiv Preprint arXiv:2503.14476. https://arxiv.org/abs/2503.14476.
Yu, Tianyu et al. 2025. RLPR:
Extrapolating RLVR to General Domains Without Verifiers. https://arxiv.org/abs/2506.18254.
Yuan, Lifan et al. 2024. “Free Process
Rewards Without Process Labels.” arXiv Preprint
arXiv:2412.01981. https://arxiv.org/abs/2412.01981.
Yuan, Lifan et al. 2025. From f(x) and g(x) to f(g(x)): LLMs
Learn New Skills in RL by Composing Old Ones. https://arxiv.org/abs/2509.25123.
Yuan, Suqin et al. 2026. Understanding
Diversity Collapse in RLVR via the Lens of Overtraining. https://arxiv.org/abs/2606.15455.
Yue, Yang, Zhiqi Chen, Rui Lu, et al. 2025. “Does Reinforcement
Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base
Model?” Advances in Neural Information Processing Systems 38
(NeurIPS 2025), Oral. https://arxiv.org/abs/2504.13837.
Zhai, Zhiyuan et al. 2026. Does RL
Expand the Capability Boundary of LLM Agents? A PASS@(k,t)
Analysis. https://arxiv.org/abs/2604.14877.
Zhang, Charlie et al. 2025. On the
Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language
Models. https://arxiv.org/abs/2512.07783.
Zhang, Chenchen. 2026. “From Reasoning to Agentic: Credit
Assignment in Reinforcement Learning for Large Language Models.”
arXiv Preprint arXiv:2604.09459. https://arxiv.org/abs/2604.09459.
Zhang, Chuyifei. 2026. When the Reward Suite Is Leaky: A
Preregistered Causal Contrast of Natural Verifier False Positives in
RLVR. https://arxiv.org/abs/2607.11022.
Zhang, Haiyue. 2026. Credit Without Ground Truth: Auditing
Step-Level Credit Assignment in LLM Agents Against Executed Replay.
https://arxiv.org/abs/2608.19760.
Zhang, Jiajie, Yushi Bai, Xin Lv, et al. 2024. “LongCite: Enabling
LLMs to Generate Fine-Grained Citations in Long-Context QA.”
arXiv Preprint arXiv:2409.02897. https://arxiv.org/abs/2409.02897.
Zhang, Lunjun, Arian Hosseini, Hritik Bansal, Seyed Mehran Kazemi,
Aviral Kumar, and Rishabh Agarwal. 2025. “Generative Verifiers:
Reward Modeling as Next-Token Prediction.” Proceedings of the
International Conference on Learning Representations (ICLR), ahead
of print. https://doi.org/10.48550/arXiv.2408.15240.
Zhang, Yanzhi et al. 2025. No Free
Lunch: Rethinking Internal Feedback for LLM Reasoning. https://arxiv.org/abs/2506.17219.
Zhao, Andrew et al. 2025. Absolute Zero:
Reinforced Self-Play Reasoning with Zero Data. https://arxiv.org/abs/2505.03335.
Zheng, Lianmin, Wei-Lin Chiang, Ying Sheng, et al. 2023. “Judging
LLM-as-a-Judge with MT-Bench and Chatbot Arena.” arXiv
Preprint arXiv:2306.05685. https://arxiv.org/abs/2306.05685.
Zhong, Ziqian et al. 2025.
ImpossibleBench: Measuring LLMs’ Propensity of Exploiting Test
Cases. https://arxiv.org/abs/2510.20270.
Zhou, Shuyan, Frank F. Xu, Hao Zhu, et al. 2023. “WebArena: A
Realistic Web Environment for Building Autonomous Agents.”
arXiv Preprint arXiv:2307.13854. https://arxiv.org/abs/2307.13854.
Zhou, Yujun, Zhenwen Liang, Haolin Liu, et al. 2025. “Evolving
Language Models Without Labels: Majority Drives Selection, Novelty
Promotes Variation.” arXiv Preprint arXiv:2509.15194. https://arxiv.org/abs/2509.15194.
Zhu, Xinyu et al. 2025. The Surprising
Effectiveness of Negative Reinforcement in LLM Reasoning. https://arxiv.org/abs/2506.01347.
Zuo, Yuxin et al. 2025. TTRL: Test-Time
Reinforcement Learning. https://arxiv.org/abs/2504.16084.