Computational methods for the analysis of protein structure and function

作者
Lisa Bartoli
出处
期刊:University of Bologna - AMS Dottorato Institutional Doctoral Theses Repository
标识
DOI:10.6092/unibo/amsdottorato/1225
摘要

The vast majority of known proteins have not yet been experimentally characterized and little is known about their function. The design and implementation of computational tools can provide insight into the function of proteins based on their sequence, their structure, their evolutionary history and their association with other proteins. Knowledge of the three-dimensional (3D) structure of a protein can lead to a deep understanding of its mode of action and interaction, but currently the structures of <1% of sequences have been experimentally solved. For this reason, it became urgent to develop new methods that are able to computationally extract relevant information from protein sequence and structure. The starting point of my work has been the study of the properties of contacts between protein residues, since they constrain protein folding and characterize different protein structures. Prediction of residue contacts in proteins is an interesting problem whose solution may be useful in protein folding recognition and de novo design. The prediction of these contacts requires the study of the protein inter-residue distances related to the specific type of amino acid pair that are encoded in the so-called contact map. An interesting new way of analyzing those structures came out when network studies were introduced, with pivotal papers demonstrating that protein contact networks also exhibit small-world behavior. In order to highlight constraints for the prediction of protein contact maps and for applications in the field of protein structure prediction and/or reconstruction from experimentally determined contact maps, I studied to which extent the characteristic path length and clustering coefficient of the protein contacts network are values that reveal characteristic features of protein contact maps. Provided that residue contacts are known for a protein sequence, the major features of its 3D structure could be deduced by combining this knowledge with correctly predicted motifs of secondary structure. In the second part of my work I focused on a particular protein structural motif, the coiled-coil, known to mediate a variety of fundamental biological interactions. Coiled-coils are found in a variety of structural forms and in a wide range of proteins including, for example, small units such as leucine zippers that drive the dimerization of many transcription factors or more complex structures such as the family of viral proteins responsible for virus-host membrane fusion. The coiled-coil structural motif is estimated to account for 5-10% of the protein sequences in the various genomes. Given their biological importance, in my work I introduced a Hidden Markov Model (HMM) that exploits the evolutionary information derived from multiple sequence alignments, to predict coiled-coil regions and to discriminate coiled-coil sequences. The results indicate that the new HMM outperforms all the existing programs and can be adopted for the coiled-coil prediction and for large-scale genome annotation. Genome annotation is a key issue in modern computational biology, being the starting point towards the understanding of the complex processes involved in biological networks. The rapid growth in the number of protein sequences and structures available poses new fundamental problems that still deserve an interpretation. Nevertheless, these data are at the basis of the design of new strategies for tackling problems such as the prediction of protein structure and function. Experimental determination of the functions of all these proteins would be a hugely time-consuming and costly task and, in most instances, has not been carried out. As an example, currently, approximately only 20% of annotated proteins in the Homo sapiens genome have been experimentally characterized. A commonly adopted procedure for annotating protein sequences relies on the "inheritance through homology" based on the notion that similar sequences share similar functions and structures. This procedure consists in the assignment of sequences to a specific group of functionally related sequences which had been grouped through clustering techniques. The clustering procedure is based on suitable similarity rules, since predicting protein structure and function from sequence largely depends on the value of sequence identity. However, additional levels of complexity are due to multi-domain proteins, to proteins that share common domains but that do not necessarily share the same function, to the finding that different combinations of shared domains can lead to different biological roles. In the last part of this study I developed and validate a system that contributes to sequence annotation by taking advantage of a validated transfer through inheritance procedure of the molecular functions and of the structural templates. After a cross-genome comparison with the BLAST program, clusters were built on the basis of two stringent constraints on sequence identity and coverage of the alignment. The adopted measure explicity answers to the problem of multi-domain proteins annotation and allows a fine grain division of the whole set of proteomes used, that ensures cluster homogeneity in terms of sequence length. A high level of coverage of structure templates on the length of protein sequences within clusters ensures that multi-domain proteins when present can be templates for sequences of similar length. This annotation procedure includes the possibility of reliably transferring statistically validated functions and structures to sequences considering information available in the present data bases of molecular functions and structures.

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
刚刚
Owen应助LDX采纳,获得10
刚刚
充电宝应助FAN采纳,获得10
1秒前
Tycoon完成签到,获得积分10
1秒前
YY发布了新的文献求助10
3秒前
沉静的不正完成签到,获得积分20
4秒前
自由滑大王完成签到 ,获得积分10
4秒前
雨打春柳完成签到,获得积分10
4秒前
执着书南发布了新的文献求助10
5秒前
听话的无极完成签到,获得积分10
5秒前
大尾尾发布了新的文献求助10
6秒前
7秒前
7秒前
YT完成签到,获得积分10
7秒前
7秒前
温暖砖头完成签到,获得积分10
8秒前
科研通AI6.2应助哭泣秋蝶采纳,获得10
8秒前
可爱的函函应助哭泣秋蝶采纳,获得10
8秒前
cyy完成签到,获得积分10
9秒前
9秒前
9秒前
温暖砖头发布了新的文献求助10
10秒前
阿布完成签到,获得积分10
11秒前
12秒前
爆米花应助He采纳,获得10
12秒前
12秒前
Verne完成签到,获得积分10
12秒前
汉堡包应助知焉采纳,获得10
13秒前
yang完成签到 ,获得积分10
13秒前
hanhan发布了新的文献求助10
14秒前
dappy完成签到 ,获得积分10
14秒前
英俊的冰棍完成签到 ,获得积分10
15秒前
鹏6发布了新的文献求助30
16秒前
123发布了新的文献求助10
16秒前
18秒前
19秒前
科研通AI6.2应助9977采纳,获得10
19秒前
科研通AI6.2应助还好吧采纳,获得10
19秒前
19秒前
hanhan完成签到,获得积分20
23秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Autoparametric Resonance in Mechanical Systems 1000
Effects of Two Weeks of Red Light Therapy on Choroidal Thickness and Axial Length in Young Adults 700
Cosmos as Art Object: Studies in Plato's Timaeus and Other Dialogues 600
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
Auslegungsgeschichte 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7663815
求助须知:如何正确求助?哪些是违规求助? 9233503
关于积分的说明 19864158
捐赠科研通 7232476
什么是DOI,文献DOI怎么找? 3282577
关于科研通互助平台的介绍 2441937
邀请新用户注册赠送积分活动 2283700