蛋白质稳态
转录组
生物
计算生物学
计算机科学
稳健性(进化)
代表(政治)
遗传适应性
基因
语言模型
人工智能
合成生物学
系统生物学
表达式(计算机科学)
细胞老化
基因表达
灵活性(工程)
意义(存在)
细胞模型
生物钟
长寿
自然语言
编码(集合论)
标识
DOI:10.1093/geroni/igaf122.4305
摘要
Abstract Large language models such as GPT have shown impressive performance in capturing structure and meaning from natural language. We adapt this capability to biology by representing single-cell transcriptomes as ordered sequences of gene names, analogous to sentences in text. Each cell’s expression profile is converted into a “cell sentence,” enabling pretrained language models to be fine-tuned directly on single-cell data for age prediction. Trained on millions of cells spanning diverse tissues and life stages, the model learns a universal latent representation of aging. It captures both cell-type–specific trajectories and conserved molecular signatures, generalizing across tissues and species. Unlike other foundational models that must be trained from scratch on omics data, our framework leverages already well-validated LLMs, taking advantage of their robust priors while avoiding the cost and instability of building new architectures. The resulting clock achieves state-of-the-art accuracy in predicting biological age at single-cell resolution. Embeddings highlight aging-associated processes such as immune activation, proteostasis decline, and metabolic shifts, aligning with known hallmarks of aging. Importantly, the model can also identify genes whose modulation decreases predicted transcriptional age, nominating candidate targets for therapeutic intervention. These features make the approach valuable both for measuring aging and for prioritizing interventions that may slow or reverse it. This work establishes the first universal, multi-tissue aging clock at single-cell resolution, demonstrating that pretrained language models can be directly adapted to biological data to map, measure, and potentially modify cellular aging.
科研通智能强力驱动
Strongly Powered by AbleSci AI