Cross Validation of Machine Learning Models for Paddy Yield Prediction
Anusha B Rao
Department of Civil Engineering, Alva’s Institute of Engineering & Technology, Moodbidri, Karnataka, India.
P. Subrahmanya V. Bhat
Department of Civil Engineering, Alva’s Institute of Engineering & Technology, Moodbidri, Karnataka, India.
S. G. Yamagar *
Department of Agricultural Engineering, Alva’s Institute of Engineering & Technology, Moodbidri, Karnataka, India.
*Author to whom correspondence should be addressed.
Abstract
Crop-yield regression models support precision agricultural decision-making and regional food-security assessment. This study evaluates the generalisation performance of seven machine-learning algorithms—Linear Regression, Ridge Regression, Lasso Regression, K-Nearest Neighbours (KNN), Decision Tree, Random Forest and Gradient Boosting—against a mean-prediction baseline using an audited dataset of 99 complete observations with seven soil and environmental predictors: nitrogen, phosphorus, potassium, temperature, humidity, pH and rainfall. A leakage-safe pipeline incorporating fold-specific standard scaling was rigorously evaluated using repeated 5-fold cross-validation across 25 validation folds, complemented by Nadeau-Bengio corrected resampled t-tests and permutation feature importance. The mean-prediction baseline attained the lowest overall prediction errors (MAE 0.8121 ± 0.1439, RMSE 1.0317 ± 0.2082, R² -0.0800 ± 0.0941). Among the machine-learning models, Random Forest achieved the lowest errors, with an MAE of 0.8172 ± 0.1241 and an RMSE of 1.0556 ± 0.1785, followed by KNN and Gradient Boosting. Linear, Ridge and Lasso regression showed slightly higher errors, whereas Decision Tree yielded the highest error. Corrected t-tests found no statistically significant differences between Random Forest and the baseline, KNN or Gradient Boosting. An 80:20 holdout evaluation of Random Forest produced an MAE of 0.9057, an RMSE of 1.3106 and an R² of -0.3011. Cross-validated permutation importance identified temperature as the most influential predictor, followed by rainfall. Overall, the findings highlight the essential role of mean baselines and leakage-safe validation in evaluating agricultural predictive models.
Keywords: Paddy yield, machine learning, cross-validation, random forest, regression modelling, yield prediction, permutation importance, data quality, predictive modelling, precision agriculture