Dataset Information


Efficient pairwise RNA structure prediction using probabilistic alignment constraints in Dynalign.

ABSTRACT: Joint alignment and secondary structure prediction of two RNA sequences can significantly improve the accuracy of the structural predictions. Methods addressing this problem, however, are forced to employ constraints that reduce computation by restricting the alignments and/or structures (i.e. folds) that are permissible. In this paper, a new methodology is presented for the purpose of establishing alignment constraints based on nucleotide alignment and insertion posterior probabilities. Using a hidden Markov model, posterior probabilities of alignment and insertion are computed for all possible pairings of nucleotide positions from the two sequences. These alignment and insertion posterior probabilities are additively combined to obtain probabilities of co-incidence for nucleotide position pairs. A suitable alignment constraint is obtained by thresholding the co-incidence probabilities. The constraint is integrated with Dynalign, a free energy minimization algorithm for joint alignment and secondary structure prediction. The resulting method is benchmarked against the previous version of Dynalign and against other programs for pairwise RNA structure prediction.The proposed technique eliminates manual parameter selection in Dynalign and provides significant computational time savings in comparison to prior constraints in Dynalign while simultaneously providing a small improvement in the structural prediction accuracy. Savings are also realized in memory. In experiments over a 5S RNA dataset with average sequence length of approximately 120 nucleotides, the method reduces computation by a factor of 2. The method performs favorably in comparison to other programs for pairwise RNA structure prediction: yielding better accuracy, on average, and requiring significantly lesser computational resources.Probabilistic analysis can be utilized in order to automate the determination of alignment constraints for pairwise RNA structure prediction methods in a principled fashion. These constraints can reduce the computational and memory requirements of these methods while maintaining or improving their accuracy of structural prediction. This extends the practical reach of these methods to longer length sequences. The revised Dynalign code is freely available for download.


PROVIDER: S-EPMC1868766 | BioStudies | 2007-01-01

REPOSITORIES: biostudies

Similar Datasets

2006-01-01 | S-EPMC1579236 | BioStudies
2007-01-01 | S-EPMC1904245 | BioStudies
1000-01-01 | S-EPMC3042186 | BioStudies
1000-01-01 | S-EPMC2324098 | BioStudies
1000-01-01 | S-EPMC2248559 | BioStudies
2014-01-01 | S-EPMC4267632 | BioStudies
2020-01-01 | S-EPMC6980424 | BioStudies
2011-01-01 | S-EPMC3120699 | BioStudies
2011-01-01 | S-EPMC3999979 | BioStudies
1000-01-01 | S-EPMC5714223 | BioStudies