亲爱的研友该休息了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!身体可是革命的本钱,早点休息,好梦!

SCREADER: Prompting Large Language Models to Interpret scRNA-seq Data

计算机科学 自然语言处理 人工智能 数据建模 语言学 数据库 哲学
作者
Cong Li,Qingqing Long,Yuanchun Zhou,Meng Xiao
标识
DOI:10.1109/icdmw65004.2024.00092
摘要

Large language models (LLMs) have demonstrated remarkable advancements, primarily due to their capabilities in modeling the hidden relationships within text sequences. This innovation presents a unique opportunity in the field of life sciences, where vast collections of single-cell omics data from multiple species provide a foundation for training foundational models. However, the challenge lies in the disparity of data scales across different species, hindering the development of a comprehensive model for interpreting genetic data across diverse organisms. In this study, we propose an innovative hybrid approach that integrates the general knowledge capabilities of LLMs with domain-specific representation models for single-cell omics data interpretation. We begin by focusing on genes as the fundamental unit of representation. Gene representations are initialized using functional descriptions, leveraging the strengths of mature language models such as LLaMA-2. By inputting single-cell gene-level expression data with prompts, we effectively model cellular representations based on the differential expression levels of genes across various species and cell types. In the experiments, we constructed developmental cells from humans and mice, specifically targeting cells that are challenging to annotate. We evaluated our methodology through basic tasks such as cell annotation and visualization analysis. The results demonstrate the efficacy of our approach compared to other methods using LLMs, highlighting significant improvements in accuracy and interoperability. Our hybrid approach enhances the representation of single-cell data and offers a robust framework for future research in cross-species genetic analysis. 1
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
小贾完成签到 ,获得积分10
刚刚
2秒前
dihou111完成签到,获得积分10
3秒前
dihou111发布了新的文献求助20
7秒前
墨曦发布了新的文献求助10
9秒前
12秒前
单身的曲奇完成签到,获得积分10
16秒前
16秒前
sisi发布了新的文献求助10
20秒前
明子完成签到 ,获得积分10
23秒前
30秒前
勤奋元龙完成签到 ,获得积分10
32秒前
sisi完成签到,获得积分10
37秒前
JM完成签到 ,获得积分10
38秒前
酷波er应助科研通管家采纳,获得10
40秒前
41秒前
奔跑应助科研通管家采纳,获得10
41秒前
Orange应助科研通管家采纳,获得10
41秒前
香蕉觅云应助科研通管家采纳,获得10
41秒前
44秒前
45秒前
jsmmm发布了新的文献求助10
47秒前
48秒前
卡卡东完成签到 ,获得积分10
49秒前
完美世界应助dihou111采纳,获得10
52秒前
53秒前
迷路的缘郡完成签到,获得积分10
54秒前
Owen应助sisi采纳,获得10
54秒前
石宇奇应助obaica采纳,获得10
56秒前
Abhinesh完成签到,获得积分10
56秒前
去去去发布了新的文献求助10
59秒前
润润润完成签到 ,获得积分10
59秒前
所所应助曾生采纳,获得10
1分钟前
去去去完成签到,获得积分20
1分钟前
鲁班大神发布了新的文献求助10
1分钟前
1分钟前
石宇奇应助obaica采纳,获得10
1分钟前
woaichifan完成签到 ,获得积分10
1分钟前
墨曦完成签到,获得积分10
1分钟前
TwentyNine完成签到 ,获得积分10
1分钟前
高分求助中
Markov Chain Monte Carlo 10000
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Common Foundations of American and East Asian Modernisation: From Alexander Hamilton to Junichero Koizumi 5000
Advanced Weaponeering Fourth Edition, Volume 2 1000
Weaponeering: An Introduction Fourth Edition, Volume 1 1000
悉尼大学博士学位论文,题目:Modelling and testing of one-sided stitched laminated composites. 作者:Kristopher P. Plain 700
Matrix Methods in Data Mining and Pattern Recognition Second Edition 610
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7556293
求助须知:如何正确求助?哪些是违规求助? 9138675
关于积分的说明 19533466
捐赠科研通 7147073
什么是DOI,文献DOI怎么找? 3261177
关于科研通互助平台的介绍 2427641
邀请新用户注册赠送积分活动 2250321