Bibliography (6):
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Proximal Policy Optimization Algorithms
https://github.com/ypwang61/One-Shot-RLVR
Wikipedia Bibliography:
Reinforcement learning
Entropy (information theory)