计算机科学
一般化
编码(集合论)
任务(项目管理)
过程(计算)
桥(图论)
趋同(经济学)
模态(人机交互)
人工智能
计算机安全
程序设计语言
医学
工程类
外科
数学
数学分析
集合(抽象数据类型)
系统工程
经济
经济增长
作者
Zhanyu Wang,Lingqiao Liu,Lei Wang,Luping Zhou
出处
期刊:
[Elsevier BV]
日期:2023-11-01
卷期号:1 (3): 100033-100033
被引量:98
标识
DOI:10.1016/j.metrad.2023.100033
摘要
Large Language Models (LLMs) have consistently showcased remarkable generalization capa-bilities when applied to various language tasks. Nonetheless, harnessing the full potential of LLMs for Radiology Report Generation (R2Gen) still presents a challenge, stemming from the inherent disparity in modality between LLMs and the R2Gen task. To bridge this gap effectively, we propose R2GenGPT, which is a novel solution that aligns visual features with the word embedding space of LLMs using an efficient visual alignment module. This innovative approach empowers the previously static LLM to seamlessly integrate and process image information, marking a step forward in optimizing R2Gen performance. R2GenGPT offers the following benefits. First, it attains state-of-the-art (SOTA) performance by training only the lightweight visual alignment module while freezing all the parameters of LLM. Second, it exhibits high training efficiency, as it requires the training of an exceptionally minimal number of parameters while achieving rapid convergence. By employing delta tuning, our model only trains 5 M parameters (which constitute just 0.07 % of the total parameter count) to achieve performance close to the SOTA levels. Our code is available at https://github.com/wang-zhanyu/R2GenGPT.
科研通智能强力驱动
Strongly Powered by AbleSci AI