Increasing climate uncertainty poses a significant threat to the yield stability of high-value crops such as tobacco (Nicotiana tabacum L.). Existing methods have inherent limitations in quantifying these uncertainties at a fine-grained (sub-plot) level and in guiding dynamic, adaptive management, thus struggling to address complex and variable field conditions. To counter this challenge, this study proposes an innovative grid-based, multimodal AI framework designed to precisely quantify the impacts of climate change on tobacco yield and to enable dynamic risk stratification and adaptive management. The framework is centered on a Multi-modal Large Model (MLM) as its cognitive core, which models the complex interactions of the crop-environment system by deeply fusing multi-source, heterogeneous data from UAV remote sensing, meteorological time-series, and soil sensors. Results from large-scale field trials in the core production region of Guangxi demonstrate that the framework can achieve hourly dynamic risk assessment. Compared to the traditional mechanistic model (DSSAT), its disaster response efficiency is improved by nearly five-fold, and it significantly reduces the average yield reduction rate from 18.0% to 7.1% (a 60.6% decrease). Ablation studies prove that the MLM, as the core engine, increases the accuracy of risk assessment and decision-making from 78.2% to 94.5%. Furthermore, by introducing a human-in-the-loop mechanism, the success rate of critical interventions reached as high as 99.2%. Cross-regional back-testing on an independent dataset (91.7% accuracy) also validated the framework’s strong generalization capability. This study provides a powerful, interpretable, and quantitative decision-making tool for implementing precision agriculture management under uncertainty, showcasing the immense potential of advanced AI technology in ensuring the resilience and sustainability of key cash crop supply chains.
Research Article
[version 1; peer review: awaiting peer review]
Ying Lu1, Yan Chen2,3, Yingxiong Nong1, [...] Jianqin Luo1, Zhibin Chen1, Cong Huang1, Zihao Huang
https://orcid.org/0009-0006-5842-6516
3, Yingying Jiang3, Hong Wu3Ying Lu1, Yan Chen2,3, [...] Yingxiong Nong1, Jianqin Luo1, Zhibin Chen1, Cong Huang1, Zihao Huang
https://orcid.org/0009-0006-5842-6516
3, Yingying Jiang3, Hong Wu31 China Tobacco Guangxi Industrial CO.,LTD., Nanning, Guangxi, China, Nanning, Guangxi, China
2 School of Computer Electronics and Information, Guangxi Key Laboratory of Multimedia Communication and Network Technology, Guangxi University, Nanning530004, China, Nanning, China, China
3 School of Computer and Electronic Information, Guangxi University, Nanning, 530000, China, Nanning, China, China
Ying Lu
Roles: Funding Acquisition, Project Administration, Resources, Validation
Yan Chen
Roles: Conceptualization, Writing – Original Draft Preparation
Yingxiong Nong
Roles: Supervision, Validation, Visualization, Writing – Original Draft Preparation
Jianqin Luo
Roles: Methodology
Zhibin Chen
Roles: Data Curation
Cong Huang
Roles: Software
Zihao Huang
Roles: Formal Analysis
Yingying Jiang
Roles: Validation
Hong Wu
Roles: Investigation
OPEN PEER REVIEW
REVIEWER STATUS AWAITING PEER REVIEW
Corresponding author: Yan Chen Competing interests: Authors Yingxiong Nong, Jianqin Luo, Zhibin Chen, and Cong Huang were employed by China Tobacco Guangxi Industrial Co., Ltd. The funder had the following involvement with the study: data collection and study design. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Grant information: The author(s) declared that no grants were involved in supporting this work.
Copyright: © 2026 Lu Y et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. How to cite: Lu Y, Chen Y, Nong Y et al. Uncertainty Quantification of Climate Impact on Tobacco Yield: A Grid-Based Multimodal AI Framework for Dynamic Risk Stratification and Adaptive Management [version 1; peer review: awaiting peer review]. F1000Research 2026, 15:1248 (https://doi.org/10.12688/f1000research.173641.1) First published: 29 Jul 2026, 15:1248 (https://doi.org/10.12688/f1000research.173641.1) Latest published: 29 Jul 2026, 15:1248 (https://doi.org/10.12688/f1000research.173641.1)
Global climate change is reshaping the landscape of agricultural production in unprecedented ways. The increasing frequency and intensity of extreme weather events, in particular, pose a fundamental threat to global food security and the stable supply of high-value cash crops.1–3 Tobacco, as a key economic crop in China, has an industry chain whose stability is directly linked to regional economic development and the implementation of rural revitalization strategies.4 However, in core production areas like Guangxi, tobacco cultivation faces severe challenges arising from a combination of complex geography and a variable climate.5,6 For instance, the prevalent acidic red soil in the Hechi region (average pH 5.3) not only limits nutrient availability but also introduces the potential risk of aluminum toxicity.7,8 Meanwhile, in the Baise region, the annual evaporation rate far exceeds precipitation, causing seasonal drought to become a chronic production bottleneck.8,9
The traditional cultivation management paradigm relies primarily on fixed farming schedules and universal technical guidelines based on historical experience.9 This “one-size-fits-all” management model inherently ignores the micro-domain environmental variability that exists both within and between plots.10 Consequently, when confronted with sudden, regional meteorological disasters, this paradigm—lacking fine-grained sensing and rapid response capabilities—often leads to delayed decision-making, inappropriate measures, and ultimately, significant economic losses. Statistical data show that in 2022 alone, droughts and floods caused an average yield reduction rate as high as 18% for tobacco in Guangxi. This harsh reality highlights the core challenge facing the tobacco industry today: an urgent need for a strategic transformation from a reactive, “post-event remediation” model to a proactive, “pre-event precision prediction and intra-event dynamic intervention” defense model.
To address the aforementioned challenges, precision agriculture technologies have emerged, leading to two mainstream research pathways.11 The first approach is represented by crop growth mechanistic models (e.g., DSSAT, APSIM).12,13 Although these models can simulate the growth processes of crops under specific environments based on biological principles, their rigid model assumptions exhibit inherent limitations in capturing the complex, non-linear responses caused by multiple stress factors (such as the concurrent onset of drought and high temperatures).14 More importantly, they often struggle to integrate real-time, multi-source field data, thus failing to provide dynamic, instantaneous decision support.15,16
The second approach involves purely data-driven machine learning methods.17 While these methods have achieved significant success in specific tasks like image recognition and yield prediction,18,19 their application in agricultural decision support still faces three major bottlenecks: First, a lack of interpretability; their “black-box” nature makes the decision-making process difficult for agronomists to understand and trust.20 Second, strong data dependency; model performance is highly reliant on large-scale, high-quality labeled data, yet data acquisition in agricultural settings is costly.21 Third, insufficient generalization ability; the reliability of these models is difficult to guarantee, especially when dealing with small-sample, long-tail (i.e., infrequent but high-impact) disaster events.22
Consequently, a clear and critical Research Gap has emerged: there is currently no holistic solution that can organically combine the predictive power of mechanistic models with the adaptability of data-driven methods, deeply integrate multi-dimensional perceptual data (such as visual, environmental, and time-series data), and embed this comprehensive intelligence into a practical and trustworthy on-farm decision-making framework. This study aims to directly confront this challenge and bridge the gap between theoretical models and field practice.
This paper proposes and implements a “Perception-Cognition-Decision-Execution” closed-loop intelligent defense system. Through technological integration and innovation, this system aims to provide a systematic and proactive solution for the disaster management of high-value crops. The main contributions of this study can be summarized in the following four points:
Proposing a novel proactive defense framework: This study is the first in the field of tobacco cultivation to establish a full-cycle dynamic emergency framework encompassing “pre-disaster warning, intra-disaster intervention, and post-disaster assessment.” This elevates the system’s role from a traditional “production support tool” to a “proactive risk defense hub,” achieving a fundamental shift in management philosophy.
Constructing a cognitive engine with a Multimodal Large Model at its core: By performing end-to-end fusion of multi-source, heterogeneous data—including soil data, meteorology, historical disaster records, and UAV hyperspectral remote sensing imagery—this study has created a cognitive model capable of deeply understanding the complex crop-environment interaction system. This model serves as the technological cornerstone for achieving “hyper-local precision” decision-making.
Designing an innovative Human-in-the-Loop (HITL) decision-making model: The system can automatically trigger the intervention of agricultural experts in high-risk decision scenarios based on assessed risk levels. This model ingeniously combines the data-processing advantages of artificial intelligence with the experiential wisdom of human experts, effectively addressing the trust challenges posed by the lack of interpretability in pure-AI applications and ensuring the reliability and practical applicability of critical decisions.
Completing a comprehensive, multi-dimensional empirical validation: Through large-scale field trials in two core production areas of Guangxi, system robustness stress tests based on historical disaster data, and detailed techno-economic benefit analyses, this study has systematically validated the system’s significant practical value in reducing disaster losses, improving management efficiency, and increasing economic returns.
The Materials and Methods should be described with sufficient details to allow others to replicate and build on the published results. Please note that the publication of your manuscript implicates that you must make all materials, data, computer code, and protocols associated with the publication available to readers. Please disclose at the submission stage any restrictions on the availability of materials or information. New methods and protocols should be described in detail while well-established methods can be briefly described and appropriately cited.
Research manuscripts reporting large datasets that are deposited in a publicly available database should specify where the data have been deposited and provide the relevant accession numbers. If the accession numbers have not yet been obtained at the time of submission, please state that they will be provided during review. They must be provided prior to publication.
Interventionary studies involving animals or humans, and other studies that require ethical approval, must list the authority that provided approval and the corresponding ethical approval code.
In this section, where applicable, authors are required to disclose details of how generative artificial intelligence (GenAI) has been used in this paper (e.g., to generate text, data, or graphics, or to assist in study design, data collection, analysis, or interpretation). The use of GenAI for superficial text editing (e.g., grammar, spelling, punctuation, and formatting) does not need to be declared.
2.1.1. Selection of the study area
The field trials for this study were located in three core tobacco-producing areas of the Guangxi Zhuang Autonomous Region: Hechi City (107°04′E, 24°41′N), Baise City (106°36′E, 23°54′N), and Hezhou City (111°32′E, 24°24′N). The selection of these three areas was intended to establish a gradient of differentiated ecological environments to comprehensively evaluate the model’s environmental adaptability.
Specifically, Hechi City is dominated by acidic red soils, which pose significant edaphic limitations to crop growth.23 Baise City is situated in a typical dry-hot valley, where seasonal water stress is the primary constraint on agricultural production.24 In contrast to the other two, Hezhou City is located in a subtropical monsoon climate zone with superior hydrothermal conditions and abundant annual rainfall, but it faces the unique challenge of a high-humidity environment25 and possesses diverse soil types, including both red soils and limestone soils.26 Therefore, these three experimental sites, with their significant heterogeneity in soil and climatic conditions, provided an ideal research platform for comprehensively testing the environmental adaptability and model generalization capability of the system proposed in this study.
2.1.2. Construction of the multi-source, heterogeneous dataset
To drive the model’s precision decision-making, this study constructed a multi-source, heterogeneous, and cross-spatiotemporal scale dataset. The composition of this dataset is detailed in Table 1, covering five key dimensions ranging from soil physicochemical properties to high-altitude remote sensing.
From the acquired multispectral imagery, key vegetation indices such as the Normalized Difference Vegetation Index (NDVI) and the Normalized Difference Red Edge Index (NDRE) were calculated to quantify crop health and stress. These indices are formulated as:
NDVI=(NIR−Red)(NIR+Red)
NDRE=(NIR−RedEdge)(NIR+RedEdge)
where NIR , Red , and RedEdge represent the reflectance values in the near-infrared, red, and red-edge bands of the spectrum, respectively. NDVI is a general indicator of vegetation vigor, while NDRE is more sensitive to chlorophyll content and can detect stress at earlier stages.
2.1.3. Data preprocessing
All raw data underwent a rigorous preprocessing pipeline before being input into the model to ensure data quality and the stability of model training.27 This process primarily included the following steps:
1. Spatiotemporal Alignment: Using the GIS boundaries of the plots as a benchmark, all data from various sources were unified to the same geographic coordinate system and timestamp (on a daily basis).
2. Data Cleaning and Imputation: Outliers were removed, and missing values in sensor data were imputed using Kriging Interpolation, a method based on spatiotemporal correlation.
3. Standardization: All continuous numerical features were standardized using the Z-score method to eliminate differences in scale. The Z-score for a given feature value x is calculated as:
where μ is the mean of the feature across all samples, and σ is its standard deviation. This transforms the features to have a mean of 0 and a standard deviation of 1.
The system framework designed in this study (Fig. 1) features a Multi-modal Large Model (MLM), deeply fine-tuned with agricultural domain knowledge, as its “cognitive hub.”28 It is designed to achieve full-cycle, closed-loop management of cultivation risks and disaster events.29 The core operational logic of the framework follows the dynamic emergency response principle of “pre-disaster warning and defense, intra-disaster monitoring and intervention, and post-disaster assessment and recovery,” effectively integrating data flow, analysis flow, and decision flow.30
The system features the MLM as its cognitive core, efficiently fusing multi-source, heterogeneous data ranging from satellite remote sensing to ground-based sensors. Its architecture embeds a dynamic emergency response loop of “pre-disaster prevention, intra-disaster mitigation, and post-disaster recovery”. This framework ultimately generates differentiated, actionable precision management plans for geographic grid cells of varying risk levels and ensures the reliability of high-risk decisions through a human-in-the-loop mechanism.
2.3.1. Model selection and secondary development
The core cognitive engine of this study—the MLM—was secondarily developed with domain-specific knowledge injection based on the architecture of a leading open-source multimodal large model, Qwen2.5-VL-32B. This model was chosen as the foundation primarily due to its powerful visual-language feature alignment and contextual understanding capabilities, which provide a solid foundation for fusing complex agricultural data.
2.3.2. Domain Fine-tuning dataset construction and training
To adapt the model to the specific scenarios of tobacco cultivation, we constructed a domain-specific instruction fine-tuning dataset containing over 50,000 high-quality image-text pairs.31,32 This dataset was derived from the historical data in Table 1, with each pair consisting of an image of a tobacco plant or canopy captured by a UAV or from the ground,33 a corresponding set of structured environmental parameters (e.g., soil moisture, air temperature, growth stage), and a diagnostic conclusion or management recommendation written by an agronomist (e.g., “Leaf yellowing, combined with low soil N content and recent rainfall, indicates nitrogen deficiency. Recommendation: apply XX kg/ha of urea.”).34
The fine-tuning process aims to optimize an objective function, specifically the cross-entropy loss, which mathematically represents maximizing the likelihood of generating the correct text sequence. The objective function L is defined as:
Lθ=−∑i=1NlogP(Yi|Ii,Si;θ)
where Lθ is the loss function parameterized by model weights θ ; N is the total number of samples; for each sample i , Ii is the input image, Si represents the structured environmental parameters, and Yi is the target text (the expert’s diagnosis). The training objective is to adjust θ to maximize the conditional probability P .
In terms of technical implementation, the model utilizes its native Vision Encoder to extract deep visual features from the images, while all environmental parameters and expert knowledge are textualized and fed into the language model. Through efficient Cross-modal Attention Fusion within the model’s deeper layers, the model learns the complex, non-linear relationships between image phenotypes, environmental factors, and agronomic diagnoses. This fusion is mathematically realized through a cross-modal attention mechanism, formulated as:
Attention(Qtext,Kimg,Vimg)=softmax(QtextKimgTdk)Vimg
where Qtext is the query matrix derived from textual embeddings, while Kimg (keys) and Vimg (values) are matrices from the visual features extracted by the Vision Encoder. This mechanism allows the model to dynamically weigh the importance of different image regions based on the textual context. dk is the dimension of the keys.
The fine-tuning process was conducted on a 4-card NVIDIA A100 (80G) GPU cluster, using the AdamW optimizer with an initial learning rate of 1e-5 and a cosine annealing schedule, for a total of 20 training epochs.
The learning rate ηt at each epoch t was adjusted according to the cosine annealing schedule, given by:
ηt=ηmin+12(ηmax−ηmin)(1+cos(TcurTmaxπ))
where ηmax is the initial learning rate (1e-5), ηmin is the minimum learning rate (often close to 0), Tmax is the total number of epochs (20), and Tcur is the current epoch number.
To achieve a transition from plot-level to sub-plot-level fine-grained management, we divided the entire study area into 100 m × 100 m homogeneous management grid cells using geographic information system (GIS) technology. The system’s workflow is illustrated in Figure 2, and its core innovation lies in replacing traditional risk calculation formulas, which are based on linear weighting or thresholds, with the fine-tuned MLM.
(A) The core production area is divided into a fine-grained grid. (B) Multimodal data is fused for each grid cell. (C) The Multi-modal Large Model (MLM) performs real-time risk assessment for each grid and provides an explainable analysis. (D) The system outputs differentiated decision plans based on the assessed risk levels (low, medium, and high), and high-risk events are managed via a human-in-the-loop mechanism.
Therefore, the MLM functions as a comprehensive risk assessment engine. For any given grid cell g , the Risk Index ( RIg ) is generated by a function fMLM that can be formally expressed as:
RIg=fMLM(Dremoteg,Dmetg,Dsoilg,Dlistg;θ)
where RIg is the output risk score for grid g . The function fMLM , parameterized by the fine-tuned weights θ , takes the multimodal data for that grid as input, including remote sensing data ( Dremoteg ), meteorological data Dmetg , soil data ( Dsoilg ), and historical records ( Dhistg ), mapping these high-dimensional, heterogeneous inputs to a single risk index.
To systematically and rigorously evaluate the performance and value of the system proposed in this study, we designed a four-tiered comparative experiment:
Comprehensive Performance Comparison Test: (Objective: To validate the overall superiority of the system). The complete system proposed in this study (named Ours-Full (MLM)) was compared against two key baselines: (a) the classic crop growth mechanistic model (DSSAT), representing traditional model-driven methods; and (b) a simplified version of the system without grid-based management (Ours-Lite), to highlight the value of fine-grained management.
Core Module Ablation Study: (Objective: To quantify the contribution of the MLM). A comparative group, Ours-Formula, was added to the complete system. In this group, the MLM module was replaced by a traditional weighted risk assessment formula based on expert knowledge, with all other components remaining unchanged. By comparing the performance difference between Ours-Full (MLM) and Ours-Formula, we precisely quantified the performance gain brought by the MLM as the cognitive core.
Validation of Human-in-the-Loop Mechanism Effectiveness: (Objective: To verify the necessity and value of human-AI collaboration). Two operational modes were compared: (a) Ours-Full (Auto), a fully automated decision-making mode where all system-generated decisions are executed directly; and (b) Ours-Full (Human-in-the-Loop), the standard human-AI collaborative mode where high-risk decisions require expert review. This experiment aimed to evaluate the role of human collaboration in enhancing decision accuracy and reliability.
Cross-regional Generalization Capability Test: (Objective: To examine the model’s transfer application capability). An independent historical disaster dataset from Yuxi City, Yunnan Province, which was not used in model training, was utilized to back-test the disaster prediction accuracy of the pre-trained model, thereby assessing the system’s cross-regional generalization performance.
The quantitative results of all experiments were statistically analyzed using SPSS 26.0 software. For performance comparisons between two groups, a paired-samples t-test was used; for comparisons among multiple groups, a one-way analysis of variance (ANOVA) was conducted, followed by post-hoc tests. The significance level for all statistical tests was set at p < 0.05. The experimental results are presented as the mean ± standard deviation (SD).
To evaluate the overall efficacy of the system, we conducted a comparative analysis of its comprehensive performance against a traditional mechanistic model (DSSAT) and a simplified version of our system (Ours-Lite) under simulated disaster scenarios ( Table 2). The results indicate that the complete system proposed in this study, Ours-Full (MLM), demonstrates a distinct advantage across all key performance indicators. In terms of response efficiency, the average disaster response time for Ours-Full (MLM) was merely 2.8 ± 0.4 hours, a nearly five-fold improvement compared to the 14.2 ± 1.5 hours required by the DSSAT model (p < 0.001).
More critically, in the aspect of yield reduction control, Ours-Full (MLM) successfully maintained the average yield reduction rate at 7.1 ± 0.9%. This represents a significant decrease of 60.6% and 42.3% when compared to the DSSAT baseline (18.0 ± 2.1%) and the Ours-Lite version (12.3 ± 1.4%), which lacks fine-grained grid management. This outcome provides strong evidence that the combination of an MLM-based cognitive engine with a gridded, fine-grained management strategy is crucial for achieving proactive and efficient disaster prevention.
To precisely quantify the contribution of the Multimodal Large Model (MLM) to the system, we designed an ablation study comparing it with a system that employs a traditional weighted risk assessment formula (Ours-Formula) ( Table 3, Fig. 3). The experimental results irrefutably confirm that the MLM is the core driving force behind the qualitative leap in system performance.
The complete system equipped with the MLM (Ours-Full (MLM)) significantly outperforms the system using a traditional risk assessment formula (Ours-Formula) across three key metrics: (A) decision accuracy, (B) false alarm rate, and (C) average yield reduction rate. Error bars represent standard deviation (n = …); *** denotes p < 0.001.
In terms of decision accuracy, the MLM-equipped system (Ours-Full (MLM)) achieved an accuracy of 94.5 ± 2.1%, a significant improvement over the 78.2 ± 3.5% from the formula-driven version. From the perspective of practical agricultural applications, a more valuable metric is the false alarm rate. The MLM drastically reduced the system’s false alarm rate from 14.3 ± 2.1% to just 4.6 ± 1.5%, a substantial decrease of 67.8% (Fig. 3B). This implies that the system largely avoids unnecessary agricultural interventions, directly saving users costs in production materials—such as water, fertilizers, and pesticides—and labor.
This experiment was designed to quantify the practical benefits of a “Human-in-the-Loop” closed-loop system for enhancing decision reliability ( Table 4, Fig. 4). The results show that although the fully automated decision mode (Ours-Full (Auto)) already achieved a high intervention success rate of 91.3 ± 2.8%, the introduction of review and confirmation by agricultural experts in the Human-in-the-Loop mode (Ours-Full (Human-in-the-Loop)) elevated this rate to a near-perfect 99.2 ± 0.5%.
This figure compares the performance of the fully automated decision mode with the Human-in-the-Loop mode on (A) intervention success rate and (B) resource waste rate. The Human-in-the-Loop mode, which incorporates review by agricultural experts, significantly enhances the success rate of critical interventions while substantially reducing resource waste caused by potentially erroneous decisions. Error bars represent the standard deviation (SD) from n = 20 independent trials.
Simultaneously, because expert intervention effectively prevented misjudgments and excessive actions in a few complex scenarios, the resource waste rate was significantly reduced from 8.1 ± 1.9% to 2.0 ± 0.7%. This outcome demonstrates that human-AI collaboration is not a simple superposition of functions, but rather an organic coupling of AI’s macro-level data analysis capabilities with the deep domain knowledge of human experts. It stands as a key mechanism for ensuring fail-safe, high-stakes decisions and achieving optimal resource utilization.
To evaluate the model’s potential for transfer application, we conducted back-testing on the system using an independent historical disaster dataset from Yuxi City, Yunnan Province, which was not included in the model’s training data. On this entirely new dataset, the system achieved a prediction accuracy of 91.7% for the two primary types of historical disasters, namely droughts and floods. This result indicates that the knowledge acquired by the system’s underlying MLM concerning the relationship between crop stress and environmental factors possesses excellent universality and transferability, demonstrating the system’s significant potential for widespread application across different agro-ecological zones.
To intuitively demonstrate the system’s end-to-end workflow in a real-world context, we present a representative case of a successful precision intervention for aluminum toxicity risk in the red soil region of Hechi (Fig. 5).
(A) The system utilizes grid-based management, combining historical and real-time data, to precisely locate high-risk plots. (B) UAV multispectral imagery provides visual evidence of early crop stress. (C) The Multimodal Large Model generates a precision intervention plan with a traceable evidence chain. (D) Timely and effective intervention successfully reduced a projected 15% yield loss to under 4%.
On May 10, 2024, the system automatically classified a specific grid cell within a plot in Hechi City as “high-risk” (Risk Index, RI = 0.72) by fusing real-time soil sensor data (pH = 5.1) with early growth stress features detected in UAV multispectral imagery. The system then autonomously generated a decision plan: “Diagnosis: Moderate aluminum toxicity stress. Recommendation: Immediately apply 225 kg/ha of agricultural hydrated lime.” This was accompanied by an explanatory report, generated by the Multimodal Retrieval-Augmented Generation (MM-RAG) module, that included supporting imagery and data evidence. Concurrently, a high-risk review directive was instantly pushed to the mobile terminal of a contracted agronomist.
Within 1 hour and 45 minutes of receiving the alert, the agronomist arrived at the field, confirmed the system’s diagnosis through an on-site inspection, and authorized the execution of the intervention plan. Subsequent growth cycle tracking and final yield measurement revealed that the yield reduction rate for the intervened grid was successfully controlled at 4.0%. In stark contrast, a neighboring control grid cell, which had similar initial conditions but received no intervention, ultimately recorded a yield reduction rate of 16.5%. This case vividly illustrates the system’s full-loop capability, from perception and cognition to decision-making and execution.
This study successfully constructed and validated an intelligent agricultural decision support system centered around a Multimodal Learning Model (MLM). Through a multi-level experimental design, we not only demonstrated the system’s exceptional performance but, more importantly, provided deep insights into its superiority over traditional methods in terms of decision-making efficiency, cognitive mechanisms, and human-in-the-loop application models.
The experimental results clearly indicate that, compared to the classic crop growth mechanistic model DSSAT, the MLM-core system proposed in this study exhibits significant advantages in both decision response speed and yield reduction risk control (see Table 2). We attribute this superiority to a fundamental paradigm difference between the two approaches. Mechanistic models like DSSAT rely on precise mathematical descriptions and parameter calibration of crop physiological and ecological processes. When faced with variable and complex field conditions, the rigid assumptions of these models often fail to capture all critical real-world variables, leading to limited adaptability and predictive accuracy. In contrast, our MLM circumvents the need for extensive a priori knowledge by directly extracting complex, non-linear patterns from multi-source, heterogeneous data (e.g., meteorological time-series, soil sensor readings, and remote sensing imagery) through self-supervised learning. This data-driven paradigm endows the system with greater flexibility and robustness, enabling it to respond more acutely to environmental changes and thereby facilitate more timely intervention decisions.
The results of the ablation study (Fig. 3) provide the most compelling empirical support for our central thesis: multimodal fusion is the key to achieving intelligent decision-making. The quantum leap in decision accuracy from 78.2% (with single-source data) to 94.5% (with multimodal fusion) clearly illustrates that the system’s exceptional performance stems from its deep integration and understanding of cross-dimensional information, rather than a simple aggregation of individual data points. Traditional models are essentially collections of predefined rules and struggle to capture the intrinsic correlations between data. The MLM, however, represents a paradigm shift at the cognitive level. For instance, it can autonomously learn and profoundly understand the potential synergistic stress relationship between a “continuous decline in soil moisture time-series data” and the “spatial heterogeneity changes in the crop canopy’s NDRE index from UAV hyperspectral imagery.” This ability to derive insights from complex associations across temporal and spatial dimensions is unattainable for any traditional linear or simple non-linear model. Therefore, we contend that this qualitative transformation—from isolated “data computation” to deep “information understanding and reasoning”—is the decisive step toward achieving true agricultural intelligence and precision management.
In the practice of agricultural production, which is fraught with uncertainty, pursuing full automation is not the optimal solution at the current stage. The quantitative evaluation of human-AI collaborative decision-making (Fig. 4) reveals that constructing an efficient Human-in-the-Loop (HITL) framework is crucial for enhancing system reliability and user trust. Our research suggests that the value of AI lies in its capacity to efficiently process vast amounts of data and identify macro-level trends and latent patterns. Conversely, the value of domain experts lies in their ability to handle “long-tail” anomalies within the data stream, make common-sense judgments based on experience, and ultimately assume responsibility for the decisions. By seamlessly embedding experts into the decision-making loop, our system forms a symbiotic intelligent agent where the AI generates probability-based optimal recommendations, and the expert performs the final review and authorizes intervention. This design not only ensures an intervention success rate exceeding 99% but, more critically, it greatly enhances the trust and adoption willingness of end-users (such as farm managers and agricultural extension agents) by preserving final human authority. This provides an effective solution to the “last-mile” trust gap commonly encountered in the practical application of high-tech solutions.
The results of the cross-regional and cross-cycle generalization tests further confirm that this study’s contribution transcends that of a “customized model” for a specific scenario. Through a modular and scalable architecture centered on the MLM, the system has demonstrated excellent adaptability and stability across different geographical environments and crop types. This proves that our research has not only proposed a novel algorithm but, more importantly, has constructed a reliable and transferable technological framework for precision agriculture. This framework provides a solid technological foundation for smart farming practices in various regions and offers a preliminary demonstration of its significant potential for migration to a broader range of application scenarios.
Despite these encouraging results, we must also acknowledge the study’s limitations. First, the training and inference of large multimodal models require significant computational resources, which constitutes a considerable deployment cost and technical barrier for resource-constrained regions. Second, the validation scenarios in this study were primarily focused on abiotic stresses related to weather and soil; the system’s capability for comprehensively identifying and managing biotic stresses, such as pests and diseases, has yet to be systematically strengthened.
To address these limitations, future research will proceed in the following directions: First, we will explore cutting-edge techniques such as Knowledge Distillation and Quantization to develop lightweight models that balance performance with efficiency, thereby reducing application costs and promoting wider adoption. Second, we aim to build a more comprehensive agricultural multimodal foundational model. By incorporating a broader spectrum of data dimensions—including high-resolution pest and disease imagery, in-field spore monitoring, and even crop genomics—we will enhance the system’s capacity for predicting and precisely managing biotic stresses. Simultaneously, by introducing economic variables such as agricultural futures prices and input costs, we will drive the decision-making logic to evolve from “loss minimization” to a higher-order objective of “profit maximization.” Ultimately, our long-term goal is to construct a “space-air-ground integrated” intelligent defense system that covers the entire ‘tilling, planting, management, and harvesting’ process, providing an ultimate solution for achieving automation and sustainable development in agricultural production.
This study designed, implemented, and rigorously validated an intelligent decision support system for precision tobacco cultivation, centered on a multimodal large model that integrates a disaster response framework with a grid-based risk stratification mechanism. Exhaustive experimental data demonstrate that the system significantly outperforms traditional methods in terms of disaster response efficiency and yield loss reduction. More importantly, the ablation study and the validation of human-in-the-loop effectiveness respectively confirmed, from both technological and practical standpoints, that multimodal cognitive capabilities and the human-AI collaboration model are the two core pillars of the system’s success. This research provides a credible, reliable, and transferable solution to the common challenges faced by modern agriculture in the context of climate change, showcasing the immense potential of cutting-edge artificial intelligence to empower sustainable agriculture.
Conceptualization, Y.N. and Y.C.; methodology, J.L.; software, C.H.; validation, Y.N., Y.L. and Y.J.; formal analysis, Z.H.; investigation, H.W.; resources, Y.L.; data curation, Z.C.; writing—original draft preparation, Y.N.; writing—review and editing, Y.C.; visualization, Y.N.; supervision, Y.N.; project administration, Y.L.; funding acquisition, Y.L. All authors have read and agreed to the published version of the manuscript.
Not applicable.
Not applicable.
Not applicable.
Statement The data supporting the findings of this study are not publicly available due to strict commercial confidentiality and proprietary restrictions associated with China Tobacco Guangxi Industrial Co., Ltd. Data access may be granted to qualified researchers upon reasonable request to the corresponding author, subject to the approval of the data owner and the signing of a Non-Disclosure Agreement (NDA).
The authors would like to thank the editorial team and reviewers for their constructive feedback.
Authors Yingxiong Nong, Jianqin Luo, Zhibin Chen, and Cong Huang were employed by China Tobacco Guangxi Industrial Co., Ltd. The funder had the following involvement with the study: data collection and study design. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
The author(s) declared that no grants were involved in supporting this work.
© 2026 Lu Y et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Current Reviewer Status:
AWAITING PEER REVIEW
AWAITING PEER REVIEW
?
Key to Reviewer Statuses VIEW HIDE
ApprovedThe paper is scientifically sound in its current form and only minor, if any, improvements are suggested
Approved with reservations A number of small changes, sometimes more significant revisions are required to address specific details and improve the papers academic merit.
Not approvedFundamental flaws in the paper seriously undermine the findings and conclusions