[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"item-2180":3},{"id":4,"title":5,"url":6,"summary":7,"summary_zh":8,"content":9,"source_name":10,"source_url":6,"published_at":11,"category":12,"cover_url":9,"hotness":13,"is_selected":14,"score":15,"score_detail":16,"sources":23,"tags":25,"view_count":31,"doi":32,"paper":33,"created_at":50},2180,"Crop Yield Estimation with MODIS Derived Normalized Difference Vegetation Index and Comparative Study on Crop Yield Prediction Among Linear Regression, Random Forest and Gradient Boosting as Well as CatBoost","https:\u002F\u002Fdoi.org\u002F10.3390\u002Frs18183107","This paper presents the design, development, and evaluation of a machine-learning system built to forecast agricultural crop yields across Indian states between 2000 and 2026, together with a complementary, national-scale verification of predicted crop yield using a MODIS-derived NDVI time series (MOD13A3.061 Vegetation Indices Monthly L3 Global 1 km SIN Grid). Although many prior studies address crop-yield prediction with linear regression, random forest, gradient boosting, and related methods, a complementary, aggregate-level verification method for predicted crop yield has rarely been proposed. This article contributes such a method, together with a complementary NDVI-based estimation approach for total foodgrain output. Crop yield and MODIS-derived NDVI are strongly correlated (r = 0.84 for annual maximum NDVI; r = 0.78 for annual mean NDVI), and a simple regression of total foodgrains on annual maximum NDVI alone reaches R2 = 0.70. Four modeling approaches—linear regression, random forest, gradient boosting, and CatBoost—were built and compared using a chronology-preserving, expanding-window walk-forward validation procedure with a final, untouched 2024–2026 holdout, rather than a random split; a companion leakage check confirmed that reported production is almost algebraically identical to reported yield and therefore had to be excluded from the feature set. Random forest produced the most reliable and consistent forecasts, reaching a mean absolute percentage error (MAPE) of 11.4% and R2 = 0.982 on the final holdout, ahead of CatBoost (MAPE = 11.6%, R2 = 0.969) and gradient boosting (MAPE = 13.0%, R2 = 0.908), and substantially ahead of linear regression, which failed to generalize to the holdout period (R2 = −10.67); across the walk-forward folds preceding this holdout, however, the three tree ensembles were statistically indistinguishable. A four-configuration ablation study confirms that most of this performance gain is attributable to the inclusion of MODIS-derived NDVI rather than to model choice alone. Prediction error varies considerably by crop, from under 10% MAPE for major staples (rice, wheat, maize, sugarcane, moong) to well over 80% MAPE for several lower-volume crops (soyabean, garlic, Sunn hemp, tobacco, potato). The paper also documents two consequential data-quality findings—a near-perfect algebraic relationship between production and yield, and a structural administrative reporting gap in 2020—and closes with directions for future work, including higher-resolution satellite inputs, temporal deep-learning architectures, additional environmental covariates, and explainable-AI analysis of feature contributions.","本文介绍了一个机器学习系统的设计、开发与评估，该系统用于预测2000年至2026年间印度各邦的农作物产量，并辅以一项基于MODIS衍生的NDVI时间序列（MOD13A3.061植被指数月度L3全球1 km SIN网格）在全国尺度上对预测作物产量的验证。尽管此前已有许多研究采用线性回归、随机森林、梯度提升及相关方法进行作物产量预测，但针对预测作物产量的补充性、聚合层面的验证方法却鲜有提出。本文提出了这样一种方法，并辅以一种基于NDVI的粮食总产量估算方法。作物产量与MODIS衍生的NDVI高度相关（年最大NDVI的r = 0.84；年均NDVI的r = 0.78），仅以年最大NDVI对粮食总产量进行简单回归即可达到R² = 0.70。本文构建了四种建模方法——线性回归、随机森林、梯度提升和CatBoost——并采用保持时间顺序的扩展窗口前向验证程序进行比较，最终以2024—2026年作为未触碰的留出集，而非随机划分；一项配套的泄漏检查证实，报告产量与报告单产在代数上几乎完全相同，因此必须将其从特征集中排除。随机森林产生了最可靠且一致的预测，在最终留出集上达到平均绝对百分比误差（MAPE）为11.4%、R² = 0.982，优于CatBoost（MAPE = 11.6%，R² = 0.969）和梯度提升（MAPE = 13.0%，R² = 0.908），并大幅优于线性回归，后者未能泛化至留出期（R² = −10.67）；然而在此留出集之前的各前向验证折中，三种树集成方法在统计上无法区分。一项四配置消融研究证实，大部分性能提升归因于纳入MODIS衍生的NDVI，而非仅归因于模型选择。预测误差因作物而异，主要粮食作物（水稻、小麦、玉米、甘蔗、绿豆）的MAPE低于10%，而若干低产量作物（大豆、大蒜、菽麻、烟草、马铃薯）的MAPE则远超80%。本文还记录了两项重要的数据质量发现——产量与单产之间近乎完美的代数关系，以及2020年结构性的行政报告缺口——并在结尾提出了未来研究方向。",null,"Remote Sensing","2026-09-10T00:00:00Z","论文",10,false,81,{"impact":17,"substance":18,"depth":17,"authority":19,"freshness":20,"relevant":21,"comment":22},18,23,14,8,1,"基于MODIS NDVI与多种机器学习模型的作物产量预测研究，方法严谨、结论可靠，对农业遥感估产有实质参考价值。",[24],{"name":10,"url":6},[26,27,28,29,30],"农业人工智能","粮食安全","机器学习","遥感","作物估产",0,"10.3390\u002Frs18183107",{"doi":32,"openalex_id":34,"authors":35,"venue":10,"cited_by_count":31,"oa_url":6,"card":42,"direction":48,"ingested_from":49},"W7212191753",[36,39],{"name":37,"orcid":38},"Kohei Arai","https:\u002F\u002Forcid.org\u002F0009-0001-6433-1592",{"name":40,"orcid":41},"Sara Sanwal","https:\u002F\u002Forcid.org\u002F0009-0003-7944-0948",{"tldr":43,"method":44,"finding":45,"direction":46,"opportunity":47},"用MODIS NDVI与四种回归模型预测印度各邦作物产量，并做全国尺度验证。","MODIS NDVI时序、线性回归、随机森林、梯度提升、CatBoost及前向验","随机森林最优（MAPE 11.4%），NDVI贡献主要增益，主粮误差低而小作物误差高。","农业遥感与作物表型","可探索多源遥感与深度模型融合，并针对小作物和行政数据缺口改进预测。","农业人工智能与决策模型","openalex","2026-09-11T23:30:52.943993Z"]