DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning.
Ontology highlight
ABSTRACT: General reasoning represents a long-standing and formidable challenge in artificial intelligence (AI). Recent breakthroughs, exemplified by large language models (LLMs)1,2 and chain-of-thought (CoT) prompting3, have achieved considerable success on foundational reasoning tasks. However, this success is heavily contingent on extensive human-annotated demonstrations and the capabilities of models are still insufficient for more complex problems. Here we show that the reasoning abilities of LLMs can be incentivized through pure reinforcement learning (RL), obviating the need for human-labelled reasoning trajectories. The proposed RL framework facilitates the emergent development of advanced reasoning patterns, such as self-reflection, verification and dynamic strategy ad
SUBMITTER: Guo D
PROVIDER: S-EPMC12443585 | biostudies-literature | 2025 Sep
REPOSITORIES: biostudies-literature
ACCESS DATA