Cross Validation of Machine Learning Models for Paddy Yield Prediction

Anusha B Rao

Department of Civil Engineering, Alva’s Institute of Engineering & Technology, Moodbidri, Karnataka, India.

P. Subrahmanya V. Bhat

Department of Civil Engineering, Alva’s Institute of Engineering & Technology, Moodbidri, Karnataka, India.

S. G. Yamagar *

Department of Agricultural Engineering, Alva’s Institute of Engineering & Technology, Moodbidri, Karnataka, India.

*Author to whom correspondence should be addressed.


Abstract

Crop-yield regression models support precision agricultural decision-making and regional food-security assessment. This study evaluates the generalisation performance of seven machine-learning algorithms—Linear Regression, Ridge Regression, Lasso Regression, K-Nearest Neighbours (KNN), Decision Tree, Random Forest and Gradient Boosting—against a mean-prediction baseline using an audited dataset of 99 complete observations with seven soil and environmental predictors: nitrogen, phosphorus, potassium, temperature, humidity, pH and rainfall. A leakage-safe pipeline incorporating fold-specific standard scaling was rigorously evaluated using repeated 5-fold cross-validation across 25 validation folds, complemented by Nadeau-Bengio corrected resampled t-tests and permutation feature importance. The mean-prediction baseline attained the lowest overall prediction errors (MAE 0.8121 ± 0.1439, RMSE 1.0317 ± 0.2082, R² -0.0800 ± 0.0941). Among the machine-learning models, Random Forest achieved the lowest errors, with an MAE of 0.8172 ± 0.1241 and an RMSE of 1.0556 ± 0.1785, followed by KNN and Gradient Boosting. Linear, Ridge and Lasso regression showed slightly higher errors, whereas Decision Tree yielded the highest error. Corrected t-tests found no statistically significant differences between Random Forest and the baseline, KNN or Gradient Boosting. An 80:20 holdout evaluation of Random Forest produced an MAE of 0.9057, an RMSE of 1.3106 and an R² of -0.3011. Cross-validated permutation importance identified temperature as the most influential predictor, followed by rainfall. Overall, the findings highlight the essential role of mean baselines and leakage-safe validation in evaluating agricultural predictive models.

Keywords: Paddy yield, machine learning, cross-validation, random forest, regression modelling, yield prediction, permutation importance, data quality, predictive modelling, precision agriculture


How to Cite

Rao, Anusha B, P. Subrahmanya V. Bhat, and S. G. Yamagar. 2026. “Cross Validation of Machine Learning Models for Paddy Yield Prediction”. Archives of Current Research International 26 (10):198-207. https://doi.org/10.9734/acri/2026/v26i102200.

Downloads

Download data is not yet available.