自由序列分析
序列(生物学)
可视化
序列比对
生物导体
多序列比对
计算机科学
可扩展性
计算生物学
软件
功能(生物学)
基因组浏览器
数据挖掘
基因组
生物
遗传学
基因组学
肽序列
基因
程序设计语言
数据库
作者
Lang Zhou,Tingze Feng,Shuaimei Xu,Fangluan Gao,Tommy Tsan-Yuk Lam,Qianwen Wang,Tianzhi Wu,Hui-Na Huang,Li Zhan,Lin Li,Yi Guan,Zhigang Dai,Yu Guo
摘要
The identification of the conserved and variable regions in the multiple sequence alignment (MSA) is critical to accelerating the process of understanding the function of genes. MSA visualizations allow us to transform sequence features into understandable visual representations. As the sequence-structure-function relationship gains increasing attention in molecular biology studies, the simple display of nucleotide or protein sequence alignment is not satisfied. A more scalable visualization is required to broaden the scope of sequence investigation. Here we present ggmsa, an R package for mining comprehensive sequence features and integrating the associated data of MSA by a variety of display methods. To uncover sequence conservation patterns, variations and recombination at the site level, sequence bundles, sequence logos, stacked sequence alignment and comparative plots are implemented. ggmsa supports integrating the correlation of MSA sequences and their phenotypes, as well as other traits such as ancestral sequences, molecular structures, molecular functions and expression levels. We also design a new visualization method for genome alignments in multiple alignment format to explore the pattern of within and between species variation. Combining these visual representations with prime knowledge, ggmsa assists researchers in discovering MSA and making decisions. The ggmsa package is open-source software released under the Artistic-2.0 license, and it is freely available on Bioconductor (https://bioconductor.org/packages/ggmsa) and Github (https://github.com/YuLab-SMU/ggmsa).
科研通智能强力驱动
Strongly Powered by AbleSci AI