Project description:Understanding accurate methods for predicting yields in complex agricultural systems is critical for effective nutrient management and crop growth. Machine learning has proven to be an important tool in this context. Numerous studies have investigated its potential for predicting yields under different conditions. Among these algorithms, Random Forest (RF) has gained prominence due to its ability to manage large data sets with high dimensions, as well as its ability to uncover complicated non-linear relationships and interactions between variables. RF is particularly suitable for scenarios with categorical variables and missing data. Given the complex web of management practices and their nonlinear effects on yield prediction, it is important to investigate new machine learning algorithms. In this context, our study focused on the evaluation of gradient boosting methods, particularly Extreme Gradient Boosting (XGB) and Gradient Boosting Regressor (GBR), as potential candidates for yield estimation of the maize hybrid Zhengdan 958. Our aim was not only to evaluate and compare these algorithms with existing approaches, but also to comprehensively analyze the resulting model uncertainties. Our approach includes comparing multiple machine learning algorithms, developing and selecting suitable features, fine-tuning the models by training and adjusting the hyperparameters, and visualizing the results. Using a recent dataset of over 1700 maize yield data pairs, our evaluation included a spectrum of algorithms. Our results show robust prediction accuracy for all algorithms. In particular, the predictions of XGB (RMSE = 0.37, R2 = 0.87 and MAE = 0.26) and GBR(RMSE = 0.39, R2 = 0.86 and MAE = 0.27), emphasized the central role of weather characteristics and confirmed the high dependence of crop yield prediction on environmental attributes. Utilizing the capabilities of gradient boosting for yield prediction holds immense potential and is consistent with the promise of this method to serve as a catalyst for further investigation in this evolving field.

Project description:BackgroundRNA molecules play critical roles in the cells of organisms, including roles in gene regulation, catalysis, and synthesis of proteins. Since RNA function depends in large part on its folded structures, much effort has been invested in developing accurate methods for prediction of RNA secondary structure from the base sequence. Minimum free energy (MFE) predictions are widely used, based on nearest neighbor thermodynamic parameters of Mathews, Turner et al. or those of Andronescu et al. Some recently proposed alternatives that leverage partition function calculations find the structure with maximum expected accuracy (MEA) or pseudo-expected accuracy (pseudo-MEA) methods. Advances in prediction methods are typically benchmarked using sensitivity, positive predictive value and their harmonic mean, namely F-measure, on datasets of known reference structures. Since such benchmarks document progress in improving accuracy of computational prediction methods, it is important to understand how measures of accuracy vary as a function of the reference datasets and whether advances in algorithms or thermodynamic parameters yield statistically significant improvements. Our work advances such understanding for the MFE and (pseudo-)MEA-based methods, with respect to the latest datasets and energy parameters.ResultsWe present three main findings. First, using the bootstrap percentile method, we show that the average F-measure accuracy of the MFE and (pseudo-)MEA-based algorithms, as measured on our largest datasets with over 2000 RNAs from diverse families, is a reliable estimate (within a 2% range with high confidence) of the accuracy of a population of RNA molecules represented by this set. However, average accuracy on smaller classes of RNAs such as a class of 89 Group I introns used previously in benchmarking algorithm accuracy is not reliable enough to draw meaningful conclusions about the relative merits of the MFE and MEA-based algorithms. Second, on our large datasets, the algorithm with best overall accuracy is a pseudo MEA-based algorithm of Hamada et al. that uses a generalized centroid estimator of base pairs. However, between MFE and other MEA-based methods, there is no clear winner in the sense that the relative accuracy of the MFE versus MEA-based algorithms changes depending on the underlying energy parameters. Third, of the four parameter sets we considered, the best accuracy for the MFE-, MEA-based, and pseudo-MEA-based methods is 0.686, 0.680, and 0.711, respectively (on a scale from 0 to 1 with 1 meaning perfect structure predictions) and is obtained with a thermodynamic parameter set obtained by Andronescu et al. called BL* (named after the Boltzmann likelihood method by which the parameters were derived).ConclusionsLarge datasets should be used to obtain reliable measures of the accuracy of RNA structure prediction algorithms, and average accuracies on specific classes (such as Group I introns and Transfer RNAs) should be interpreted with caution, considering the relatively small size of currently available datasets for such classes. The accuracy of the MEA-based methods is significantly higher when using the BL* parameter set of Andronescu et al. than when using the parameters of Mathews and Turner, and there is no significant difference between the accuracy of MEA-based methods and MFE when using the BL* parameters. The pseudo-MEA-based method of Hamada et al. with the BL* parameter set significantly outperforms all other MFE and MEA-based algorithms on our large data sets.

Dataset Information

Yield prediction for crops by gradient-based algorithms

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets