分子置换
蛋白质数据库
蛋白质数据库
计算机科学
算法
序列(生物学)
集合(抽象数据类型)
蛋白质结构
格子(音乐)
结晶学
数据挖掘
晶体结构
化学
物理
程序设计语言
立体化学
生物化学
声学
摘要
Molecular replacement (MR) is the most popular technique to solve the phase problem in macromolecular crystallography. The conventional approach to finding search models for MR is to use the sequence of the target structure to identify a suitable homologue. This approach is based on the assumption that sequence similarity is a useful guide to structural similarity. Whilst largely true, this strategy is not always effective. For example, when a contaminant protein has been crystallised or when the most similar matches sequentially are not the most similar structurally. This thesis describes the development of SIMBAD, a three-step pipeline to perform sequence-independent MR. The first step performs a lattice-parameter search against the entire Protein Data Bank (PDB), rapidly determining whether the protein or a close homologue has been solved in the same crystal form. The second step is designed to screen the data against a database of known contaminants; thus determining if a contaminant protein has been crystallised. The final step is a brute-force search of a non-redundant derivative of the PDB provided by the MoRDa software package. In Chapter 3 the initial implementation of SIMBAD using AMoRe’s fast rotation function is presented, with encouraging results. Testing on a set of structures that covered a wide range of resolution limits, copies in the asymmetric unit, space groups, monomer sizes and secondary-structure types, gave a 40% success-rate with the full MoRDa database search and increased to 52% when combined with the lattice-parameter search. Further validation has come in the form of nine structures deposited to the PDB which used SIMBAD for structure solution. Leading on from the work in Chapter 3, research was carried out on whether the maximum-likelihood enhanced rotation function in Phaser would improve the sensitivity of the full MoRDa database search. Results presented in Chapter 4 show that the use of Phaser yielded a 60% success-rate on the test cases, a marked improvement on the previous iteration of SIMBAD. Combining this method with ensemble search models improved this further to 68%. Lastly, Chapter 5 explores the use of anomalous Fourier maps (AFMs) to validate partial MR solutions obtained from SIMBAD. This was necessary as the absence of sequence information meant that automated model building could not be included in the pipeline as a means to test the correctness of a potential solution. The findings in Chapter 5 demonstrate that when anomalous signal was available, the maximum peak height obtained in AFMs could be combined with R-free to train a classifier with 99% precision and recall.
科研通智能强力驱动
Strongly Powered by AbleSci AI