亲爱的研友该休息了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!身体可是革命的本钱,早点休息,好梦!

LLMs Can Get "Brain Rot"!

认知重构 认知 控制(管理) 认知心理学 计算机科学 钥匙(锁) 产量(工程) 心理学 认知负荷 比例(比率) 推论 等值 计算机安全 质量(理念) 安全性令牌 社会心理学 数据质量 数据科学 人工智能 因果推理 应用心理学
作者
Xing Shuo,Hong, Junyuan,Wang, Yifan,Chen, Runjin,Zhang ZhenYu,Grama, Ananth,Tu, Zhengzhong,Wang, Zhangyang
出处
期刊:Cornell University - arXiv [Cornell University]
标识
DOI:10.48550/arxiv.2510.13928
摘要

We propose and test the LLM Brain Rot Hypothesis: continual exposure to junk web text induces lasting cognitive decline in large language models (LLMs). To causally isolate data quality, we run controlled experiments on real Twitter/X corpora, constructing junk and reversely controlled datasets via two orthogonal operationalizations: M1 (engagement degree) and M2 (semantic quality), with matched token scale and training operations across conditions. Contrary to the control group, continual pre-training of 4 LLMs on the junk dataset causes non-trivial declines (Hedges' $g>0.3$) on reasoning, long-context understanding, safety, and inflating "dark traits" (e.g., psychopathy, narcissism). The gradual mixtures of junk and control datasets also yield dose-response cognition decay: for example, under M1, ARC-Challenge with Chain Of Thoughts drops $74.9 \rightarrow 57.2$ and RULER-CWE $84.4 \rightarrow 52.3$ as junk ratio rises from $0\%$ to $100\%$. Error forensics reveal several key insights. First, we identify thought-skipping as the primary lesion: models increasingly truncate or skip reasoning chains, explaining most of the error growth. Second, partial but incomplete healing is observed: scaling instruction tuning and clean data pre-training improve the declined cognition yet cannot restore baseline capability, suggesting persistent representational drift rather than format mismatch. Finally, we discover that the popularity, a non-semantic metric, of a tweet is a better indicator of the Brain Rot effect than the length in M1. Together, the results provide significant, multi-perspective evidence that data quality is a causal driver of LLM capability decay, reframing curation for continual pretraining as a \textit{training-time safety} problem and motivating routine "cognitive health checks" for deployed LLMs.

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
在水一方应助罗添龙采纳,获得30
5秒前
tyh发布了新的文献求助100
13秒前
13秒前
罗添龙发布了新的文献求助10
16秒前
合适雨完成签到,获得积分10
25秒前
26秒前
shadow焓完成签到,获得积分20
28秒前
cxk完成签到 ,获得积分10
29秒前
和谐的青筠完成签到,获得积分10
41秒前
49秒前
1分钟前
1分钟前
欣慰怀梦完成签到,获得积分10
1分钟前
1分钟前
正直的晋鹏完成签到,获得积分10
1分钟前
1分钟前
1分钟前
年轻新晴完成签到,获得积分10
1分钟前
Demi_Ming完成签到,获得积分0
1分钟前
1分钟前
殷勤的岱周完成签到 ,获得积分10
2分钟前
卿亦佳人发布了新的文献求助10
2分钟前
感性的远航完成签到,获得积分10
2分钟前
柒年啵啵完成签到 ,获得积分10
2分钟前
2分钟前
仁爱的鹤轩完成签到,获得积分10
2分钟前
阔达的碧彤完成签到,获得积分10
2分钟前
NattyPoe完成签到,获得积分10
2分钟前
结实智宸完成签到,获得积分0
3分钟前
3分钟前
卿亦佳人发布了新的文献求助10
3分钟前
大胆夏菡发布了新的文献求助10
3分钟前
天真的音完成签到,获得积分10
3分钟前
超帅的半莲完成签到,获得积分10
3分钟前
潘佳琪完成签到 ,获得积分10
3分钟前
研友_nxw2xL完成签到,获得积分0
3分钟前
闪闪雍完成签到,获得积分10
4分钟前
大方的仙人掌完成签到,获得积分10
4分钟前
缓慢怜菡完成签到,获得积分0
4分钟前
飞快的元柏完成签到,获得积分10
4分钟前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Essentials of Carbohydrate Chemistry and Biochemistry, 4th Edition 800
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
CLSI VET01S-2024 Performance Standards for Antimicrobial Disk and Dilution Susceptibility Tests for Bacteria Isolated From Animals (7th Ed) 500
A Case Study on Hotels as Noncongregate Emergency Living Accommodations for Returning Citizens 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 计算机科学 化学工程 工程类 有机化学 物理 复合材料 生物化学 内科学 细胞生物学 基因 遗传学 免疫学 冶金 光电子学 癌症研究
热门帖子
关注 科研通微信公众号,转发送积分 7765667
求助须知:如何正确求助?哪些是违规求助? 9309865
关于积分的说明 20312847
捐赠科研通 7350479
什么是DOI,文献DOI怎么找? 3314969
关于科研通互助平台的介绍 2464376
邀请新用户注册赠送积分活动 2329466