计算机科学
人工智能
计算生物学
基因表达
转录组
代表(政治)
表达式(计算机科学)
基因
桥接(联网)
深度学习
基因调控网络
机器学习
领域(数学分析)
模式识别(心理学)
基因表达谱
基因表达调控
训练集
数据挖掘
基因预测
条件随机场
特征学习
生物
领域知识
作者
Bochong Zhang,Tianyi Zhang,Qiaochu Xue,Zeyu Liu,Dankai Liao,T. Antoni,Yeo Hui Ting Grace,Sicheng Chen,Hwee Kuan Lee,Shangqing Lyu,Yueming Jin
标识
DOI:10.1109/tmi.2026.3688322
摘要
Spatial Transcriptomics (ST) technology detects gene expression from tissue biopsies, playing an emerging role in cancer diagnosis and precision medicine. However, the high cost of ST technology limits its broader application. Recently, deep learning approaches have provided insight into predicting gene expression based on H&E-stained histopathology images. Nevertheless, the relationship between morphological features and gene expression is highly complex. To address these challenges, we propose DiffBulk, a novel two-stage framework that leverages conditional diffusion models to learn expressive image representations enriched with gene expression information. In the first stage, we introduce a gene-to-image conditional diffusion model equipped with a permutationinvariant open-embedding gene encoder, which enables unified training across diverse gene panels. In the second stage, diffusion-derived features are fused with representations from a pathology foundation model, effectively bridging the domain gap and improving downstream gene expression prediction. We evaluate DiffBulk on high-quality Xenium ST data curated from the HEST dataset and the CrunchDAO challenge, constructing tile-level pseudo-bulk datasets for training and evaluation. Extensive experiments demonstrate that DiffBulk consistently outperforms state-of-the-art baselines across all metrics for gene expression prediction. These findings highlight the potential of diffusion-based gene-image representation learning and suggest promising directions for future research.
科研通智能强力驱动
Strongly Powered by AbleSci AI