Lv63
1980 积分 2025-09-12 加入
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
15天前
已关闭
Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations
15天前
已关闭
Resolving Discrepancies in Compute-Optimal Scaling of Language Models
15天前
已关闭
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining
15天前
已完结
Scaling Laws for Optimal Data Mixtures
15天前
已完结
DataComp-LM: In search of the next generation of training sets for language models
15天前
已关闭
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
15天前
已关闭
Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
15天前
已完结
Are Emergent Abilities of Large Language Models a Mirage?
15天前
已完结
Synthesizing scientific literature with retrieval-augmented language models
15天前
已完结