SaintGSE: Transformer-based efficient and explainable gene set enrichment analysis

预处理器 计算机科学 源代码 自编码 编码(集合论) 基因 钥匙(锁) 计算生物学 数据挖掘 集合(抽象数据类型) 人工智能 克隆(Java方法) 微阵列分析技术 数据预处理 编码 麻省理工许可证 功能(生物学) 数据集 微阵列 情报检索 生物信息学 基因表达谱 机器学习 许可证 基因预测 生物
作者
Min-Seung Jeon,Jiho Nam,Chanmi Cho,Siyoung Yang,Seong-il Eyun
出处
期刊:CERN European Organization for Nuclear Research - Zenodo [European Organization for Nuclear Research]
标识
DOI:10.5281/zenodo.18287022
摘要

SaintGSE: Transformer-based efficient and explainable gene set enrichment analysis SaintGSE is an artificial intelligence model designed to predict human gene–pathway relationships using large-scale differentially expressed gene (DEG) datasets. By leveraging an autoencoder and the SAINT transformer model, SaintGSE addresses practical challenges in gene expression analysis, including data scarcity, model compatibility, and interpretability. This project uses and modifies code from the SAINT project (https://github.com/somepago/saint), licensed under the Apache License 2.0. Key Features * AI-driven pathway prediction: Uses an autoencoder and the SAINT model to analyze DEG signatures and predict associated signaling pathways. * Osteoarthritis study: Applied to osteoarthritis (OA) to identify key pathways and potential therapeutic targets. * Explainability (Integrated Gradients): Produces gene-level attributions via Integrated Gradients (IG) to identify influential genes supporting each pathway prediction. Installation Before installation, we recommend building a conda environment from the attached YAML file and activating it. Our code has been tested with python=3.8 on Linux. ``` $ cd /path/to/SaintGSE $ conda env create -f saintgse_env.yml $ conda activate saintgse_env ``` Clone the repository and install the dependencies: ``` $ git clone https://github.com/MSjeon27/saintgse.git $ cd saintgse ``` The code in this dataset is also accessible via GitHub. You can find the GitHub repository at the following link: https://github.com/MSjeon27/SaintGSE Usage Step 0. Preprocessing the input DEG (from pyDESeq2 result) Currently, SaintGSE has the function of converting mouse genes into human genes. The preprocessing code serves to change the human or mouse DEG data into the format used for SaintGSE. * human DEGs ``` $ preprocessing.py --query_fc /path/to/your/DEGs.tsv --out Preprocessed_fc.tsv ``` * mouse DEGs ``` $ preprocessing.py --query_fc /path/to/your/DEGs.tsv --org mouse --out Preprocessed_fc.tsv ``` Step 1. Training SaintGSE for a target pathway Available pathway names can be found in: /datasets/pathway_list_in_DEG.txt Training: ```bash $ python SaintGSE.py --pathway "Proteins Involved in Osteoarthritis" --pretrain ``` Step 2. Prediction through SaintGSE (Interpretation) Run prediction for a target pathway using a pre-trained or previously trained checkpoint: ```bash $ python SaintGSE.py --pathway "Proteins Involved in Osteoarthritis" --predict Preprocessed_fc.tsv ``` If IG is enabled, SaintGSE also produces gene-level attributions per sample. Typical outputs include: - pathway_IG_sample_scores.tsv (sample-level pathway calibrated score; closer to 1 indicates stronger association, closer to 0 indicates weaker association) - *_IG_gene_contributions.tsv (gene-level IG contributions per sample) Recommended driver gene definition (cumulative |IG|): In our OA analyses, we observed that a compact subset of DEGs (approximately 7–20% of the DEG signature) often explained 50% of the total cumulative |IG|. We recommend defining informative genes as the smallest set of genes that cumulatively accounts for the top 50% of |IG| (top 50% of cumulative |IG|). Note You can control IG computation via optional arguments: - ig_steps (e.g., 10/20/50/100/256) (default: 64) - ig_baseline (zero or mean) (default: zero) How to Cite If you use this model or repository in your research, please cite it as follows: ``` Jeon, MS & Nam, JH et al., "SaintGSE: Transformer-based efficient and explainable gene set enrichment analysis," 2024. GitHub repository. Available at: https://github.com/MSjeon27/SaintGSE ``` For more information or any questions regarding citation, feel free to contact us (msjeon27@cau.ac.kr).
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
Akim应助绝影采纳,获得10
1秒前
tao完成签到,获得积分10
1秒前
11235应助内向面包采纳,获得10
2秒前
3秒前
moshi发布了新的文献求助10
3秒前
keepmoving_12完成签到 ,获得积分10
3秒前
任迷迷发布了新的文献求助10
4秒前
王贝贝完成签到,获得积分10
5秒前
5秒前
5秒前
香蕉觅云应助无情的友容采纳,获得10
5秒前
5秒前
pyrene发布了新的文献求助30
6秒前
7秒前
Sandy完成签到,获得积分10
8秒前
学医没出路完成签到 ,获得积分10
8秒前
chen完成签到 ,获得积分20
9秒前
完美世界应助映城采纳,获得30
10秒前
sjshsbj完成签到,获得积分10
10秒前
li发布了新的文献求助10
10秒前
moshi完成签到,获得积分10
12秒前
12秒前
xubee完成签到,获得积分10
14秒前
超级的续完成签到,获得积分10
14秒前
粗暴的元柏完成签到,获得积分10
15秒前
15秒前
15秒前
所所应助Strange采纳,获得10
16秒前
aaaa应助于洛铱采纳,获得30
16秒前
十九完成签到,获得积分10
17秒前
贾硕完成签到,获得积分10
17秒前
绝影发布了新的文献求助10
17秒前
Cun完成签到 ,获得积分10
18秒前
cxf给cxf的求助进行了留言
18秒前
大个应助任迷迷采纳,获得10
21秒前
隐形曼青应助去看海采纳,获得10
23秒前
久久完成签到 ,获得积分10
23秒前
23秒前
23秒前
24秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Principles of town planning: translating concepts to applications 1000
内視鏡的に摘除しえた十二指腸乳頭部腫瘍の2例 660
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
Positive Obsession: The Life and Times of Octavia E. Butler 500
Interpolation and Regression Models for the Chemical Engineer: Solving Numerical Problems 400
The Neuroscience of Language 400
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7689816
求助须知:如何正确求助?哪些是违规求助? 9251878
关于积分的说明 19973506
捐赠科研通 7262810
什么是DOI,文献DOI怎么找? 3290440
关于科研通互助平台的介绍 2447127
邀请新用户注册赠送积分活动 2295279