[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"item-2541":3},{"id":4,"title":5,"url":6,"summary":7,"summary_zh":8,"content":9,"source_name":10,"source_url":6,"published_at":11,"category":12,"cover_url":9,"hotness":13,"is_selected":14,"score":15,"score_detail":16,"sources":24,"tags":26,"view_count":32,"doi":33,"paper":34,"created_at":52},2541,"Evidence integrity and review utility in crop-image decision support: an audit-to-review framework for precision agriculture","https:\u002F\u002Fdoi.org\u002F10.1186\u002Fs42269-026-01492-x","Abstract Background Agricultural artificial intelligence benchmarks can overstate decision reliability when multiple files represent the same biological evidence or when split provenance is unclear. We evaluated a classifier-agnostic audit-to-review framework that fixes evidence identity and role provenance before model comparison, declares aggregation estimands, and links predictive uncertainty to review workload. A public wheat-leaf corpus of 7,595 files was hashed to 5,118 unique contents; eligibility criteria defined a 915-content five-class task. Archived file-level benchmarks were separated from three frozen near-duplicate-grouped P2 partition configurations, and semantic, texture and fusion representations were evaluated under partition-weighted and one-content-one-weight estimands. Split conformal prediction was assessed by empirical coverage, review rate, error capture, review yield and dimensionless review utility. Results Macro-F1 was 0.962 under the legacy P0 benchmark, 0.805 ± 0.030 for semantic-only P2 and 0.819 ± 0.029 for direct fusion. The descriptive P0-P2 workflow-sensitivity gap was approximately 15.7% points, versus a 1.4-point partition-weighted fusion increment (95% bootstrap interval − 0.006 to 0.036). The leading fusion-related variants were numerically close, so no inferential ranking is claimed. At α = 0.10, global split conformal prediction achieved 0.896 empirical coverage while reviewing 20.8% of images and capturing 51.4% of top-1 errors; class-conditional Mondrian calibration achieved 0.923 coverage while reviewing 29.7% and capturing 64.8% of errors. Conclusions Defining the evidence unit is part of the statistical specification of a crop-image benchmark, not merely a data-cleaning step. In this corpus, benchmark interpretation was much more sensitive to evidence-workflow specification than to the evaluated representation refinement. The framework supports auditable internal evaluation and review triage but does not establish field, farm-level or intervention performance.","摘要 背景 当多个文件代表同一生物学证据或划分来源不明时，农业人工智能基准可能高估决策可靠性。我们评估了一个与分类器无关的“审计到复核”框架，该框架在模型比较前固定证据身份和角色来源，声明聚合估计目标，并将预测不确定性关联到复核工作量。一个包含7,595个文件的公开小麦叶片语料库经哈希处理得到5,118个唯一内容；资格标准定义了一个915个内容的五分类任务。将归档的文件级基准与三种冻结的近重复分组P2划分配置分离，并在划分加权和“一内容一权重”估计目标下评估了语义、纹理和融合表示。通过经验覆盖率、复核率、错误捕获率、复核产出率和无量纲复核效用评估分裂保形预测。结果 在遗留P0基准下，宏F1为0.962；仅语义P2为0.805 ± 0.030；直接融合为0.819 ± 0.029。描述性的P0-P2工作流敏感性差距约为15.7个百分点，而划分加权融合增量为1.4个百分点（95%自助法区间−0.006至0.036）。领先的融合相关变体在数值上接近，因此不声称推断性排序。在α = 0.10时，全局分裂保形预测实现了0.896的经验覆盖率，同时复核了20.8%的图像并捕获了51.4%的top-1错误；类条件Mondrian校准实现了0.923的覆盖率，同时复核了29.7%并捕获了64.8%的错误。结论 定义证据单元是作物图像基准统计规范的一部分，而不仅仅是数据清理步骤。在该语料库中，基准解释对证据-工作流规范的敏感性远高于对所评估表示改进的敏感性。该框架支持可审计的内部评估和复核分诊，但不能确立田间、农场级或干预性能。",null,"Bulletin of the National Research Centre\u002FBulletin of the National Research Center","2026-09-14T00:00:00Z","论文",10,false,77,{"impact":17,"substance":18,"depth":19,"authority":20,"freshness":21,"relevant":22,"comment":23},15,22,18,13,9,1,"该论文提出面向作物图像决策支持的证据完整性审计框架，揭示基准测试中证据单元定义对性能评估的显著影响，方法新颖、数据规模较大，对农业AI模型评估具有参考价值，但属细分领域方法学研究，产业影响有限。",[25],{"name":10,"url":6},[27,28,29,30,31],"智慧农业","农业人工智能","遥感监测","作物病害识别","模型评估",0,"10.1186\u002Fs42269-026-01492-x",{"doi":33,"openalex_id":35,"authors":36,"venue":10,"cited_by_count":32,"oa_url":6,"card":45,"direction":49,"ingested_from":51},"W7212800882",[37,40,42],{"name":38,"orcid":39},"Xin Li","https:\u002F\u002Forcid.org\u002F0009-0005-0670-5006",{"name":41,"orcid":9},"Bojian Guo",{"name":43,"orcid":44},"A.Dzh. Kartanova","https:\u002F\u002Forcid.org\u002F0000-0003-1479-0747",{"tldr":46,"method":47,"finding":48,"direction":49,"opportunity":50},"提出审计到评审框架，先固定证据身份再比较作物图像分类器，并关联预测不确定性与评审工作量。","对7595个小麦叶片文件哈希去重，构建915内容五类任务，用分裂保形预测评估。","基准解释对证据工作流规范远比对表示优化敏感，P0与P2宏F1差约15.7个百分点。","农业人工智能与决策模型","可将证据单元定义与评审效用纳入农业AI基准标准，并探索田间级验证与主动学习结合。","openalex","2026-09-15T23:30:34.690905Z"]