欧几里德距离
概率逻辑
欧几里得空间
距离测量
计算机科学
相似性(几何)
等价(形式语言)
多项式分布
代表(政治)
人工智能
数学
数字几何
模式识别(心理学)
算法
欧几里德几何
空格(标点符号)
表达式(计算机科学)
测地线
集合(抽象数据类型)
数据挖掘
选择(遗传算法)
合成数据
相似
转录组
相似性度量
计算
选型
交互信息
系统发育中的距离矩阵
协方差
空间分析
理论计算机科学
统计模型
约束(计算机辅助设计)
作者
Jinpu Cai,Yuxuan Wang,Yunhao Qiao,Changhu Wang,Ziqi Rong,Luting Zhou,Haoyang Liu,Meng Jiang,Hong-Bin Shen,Jingyi Jessica Li,Hongyi Xin
出处
期刊:
[Cold Spring Harbor Laboratory]
日期:2026-02-26
标识
DOI:10.64898/2026.02.25.707866
摘要
Abstract Single-cell and spatial transcriptomics provide high-resolution cellular characterization, yet standard analytical approaches remain theoretically misaligned with the probabilistic nature of the data. After UMI normalization, current pipelines rely on Euclidean or log-transformed Euclidean distance for similarity measurement. Both are fundamentally ill-suited to model the multinomial count data. Euclidean distance in normalized space overemphasizes high-variance genes, while log-transformation inverts this bias but at the cost of distorting subtle, continuous expression modulations. Neither approach naturally captures the dual nature of gene expression: both discrete presence/absence transitions and continuous quantitative variation. To overcome these limitations, we introduce GAIA (Geometric Analysis from an Information Aspect), an information-geometric framework for cell representation learning and inter-cell similarity measurement. By anchoring analysis in the true probabilistic model, treating cells as multinomial distributions over genes and projecting cells to a statistical manifold, GAIA organically reconciles both the presence/absence effect and the more continuous expression modulations. Mathematically, GAIA exploits the equivalence between Fisher-Rao distance in multinomial space and geodesic distance on the unit hypersphere, a property that enables both theoretical guarantees and computational efficiency. Experiments in synthetic and real scRNA-seq and spatial transcriptomic datasets demonstrate that GAIA preserves robust and consistent cell-to-cell relationships, delineates biologically nuanced sub-types, mitigates batch effects arising from sequencing depth variation, and eliminates the dependence on knowledge-restricted gene selection for learning meaningful cell representations. Overall, GAIA offers a knowledge-lean, variance-stabilizing framework for analyzing single-cell and spatial transcriptomic data, enhancing discrimination between nuanced cell sub-type and -states.
科研通智能强力驱动
Strongly Powered by AbleSci AI