Phenotypic and physiological characteristic revealed that genotypic response to different field trials was highly variable (Table 3). The statistically significant difference in the height of the mature plants was from 35.2 to 61.7 cm and the leaf area index from 1.2 to 3.6.
Such great discrepancies in the development of the canopies and the variation of the photosynthetic capacity among the genotypes are the reasons behind the differences in the leaf area index (LAI). In addition, the chlorophyll content as reflected by SPAD measurements among the genotypes varied substantially with an interval of 34.1 to 52.8. These indicate that the genotypes differ in their photosynthetic efficiency and nutrient uptake. Moreover, there was a wide variation in the parameters directly related to the yield level among the different genotypes. The counts of pods per plant were between 18 and 47, while the weight of 100 seeds was from 18.5 g to 28.3 g. These are the main factors, along with the environmental influences present that are reflected in the plants’ reproductive success. The grain yield per plot varied between 1. 2 t ha
-1 to 2. 8 t ha
-1. This variation could be attributed to the contribution of vegetative growth as well as the setting of pods to the level of productivity. Seed quality characteristics such as germination percentage, vigor index and moisture content also depicted differences among genotypes. Germination percentage (Fig 1) was in the range of 78% to 94%, vigor index ranged from 1020 to 1450 and moisture content was between 7. 8% to 12. 5%.
These differences highlight the need to capture those biologically meaningful parameters that are very relevant for AI models. Some plant traits like plant height, chlorophyll index, and pod number, correlated strongly with yield, whereas seed size and germination parameters were good indicators of seed quality (Fig 2). The variety of data thus allowed for creating very efficient prediction models that could work across different genotypes and environments.
As shown in Fig 3, the yield prediction and seed quality classification were very accurate using the HARFO, AI model. The yield predictions had an R of 0.94 with RMSE and MAE values of 0.19 and 0.15 t ha respectively, these values suggest that there was a very good agreement between the predicted and the observed yields. Yield quality classification achieved an accuracy of 95.2%, and the precision and recall values were high for all quality classes, thus the ability of the models to distinguish between seeds of high, medium, and low quality is confirmed. By analysing mechanisms embedded in HARFO, AI, the importance of input features was made known. Yield prediction was most influenced by plant height, number of pods per plant, chlorophyll index, and 100 seed weight, in no particular order. Furthermore, for seed quality determination, germination percentage, vigor index, and moisture content obtained the highest attention weights, thus providing extra evidence of the biological relevance of these traits. Feature visualization brought to light practical guidance to plant breeders in choosing genotypes that have balanced vegetative growth and reproductive traits, which not only yield the highest but also produce the best seed quality.
Fig 3 offers a very informative summary visually of how well the HARFO, AI system performed in both yield prediction and seed quality classification. The scatter diagram exhibits the predicted yield versus the actual yield (t ha
-1). Most of the observations are close to the regression line, which means that the predictive model is good. Besides, an R of 0.94, RMSE = 0.19 t_ha
-1 and MAE = 0.065 t ha
-1 further demonstrate that the model is very trustworthy and the prediction error is very small.
From Fig 3, the very strong diagonal penetration corresponds to the correct predictions and the overall classification accuracy is 95.2%, thus, the good multi, class discrimination has been achieved by the HARFO, AI model. Feature Importance for Yield: A horizontal bar chart visually represents the main features influencing the yield, one of which is plant height, along with the number of pods, chlorophyll index and 100 seed weight. Their relative contributions to yield prediction are disclosed.
Feature importance for seed quality
There is yet another bar graph that depicts the main features of seed quality highlighting the top predictors such as seed moisture, germination percentage, vigor index. The graph points out the relationships that make sense biologically. Combining predictive capability, classification accuracy and explainable feature importance, the figure showcases that HARFO, AI is not only accurate but also understandable and usable by farmers and breeders in agriculture.
HARFO, AI outshined the conventional regression and classification methods in a comparative analysis as shown in Table 4 by a wide margin each time. Linear regression, standard Random Forest, and support vector machine models gave the yield prediction results with lower accuracy (R values ranged from 0.74 to 0.81) and higher RMSE (0.380.45 t ha
-1). Conventional models’ seed quality classification accuracies were in the range of 81% to 87%, therefore, pointing to their weaknesses in disentangling the complex nonlinear interactions between various traits.
Fig 4 shows a comparison between the prediction errors of yield for three different models; HARFO, AI, RF Regressor and LR. The chart basically depicts the size of the prediction error (t ha
-1) and thus, the difference in accuracy and reliability of the models. HARFO, AI has the smallest prediction error and hence, it is the one capable of modelling highly complex, nonlinear relationships between the three types of variables.
The RF method has a medium error level and hence, it reasonably predicts but still not as good as the hybrid architecture in terms of optimization and interpretability. On the other hand, the LR model has the largest error which could be explained by its lack of capability to model nonlinear interactions and complex features of the dependency of agricultural systems. Basically, the graph nicely shows that by using HARFO, AI methodology, one can significantly reduce the level of unpredictability and increase the accuracy of yield forecasting. The superior performance as compared to the other models confirms the computational attention mechanisms and the optimization strategies coordinated within the hybrid model, thus making it the best option for AI, driven decision, making in crop breeding and precision agriculture systems.
Fig 5 displays the performance and feature contribution analysis of the HARFO, AI framework. It compares with non, deep approaches and presents the relative performance gains of the proposed model over LR, SVM, and RF being 20.5%, 11.9%, and 5.6%, respectively, thus depicting the superiority of nonlinear learning and attention, based optimization. The analysis of feature contributions in yield indicates that the yield attributes (44%) and growth traits (38%) are significantly responsible for high productivity, whereas phenological traits (18%) have a moderate effect on stability. Regarding seed quality, the mixture of germination and vigour traits (46%) is the most dominant, then moisture traits (29%), and biochemical protein traits (25%). Prediction of yield (49%) and seed classification (51%) are two examples where consistency is maintained, and both of these cases show very strong reliability.
The integrated experimental outcome shows that the HARFO-AI model is capable of combining the variability of biological traits with the power of advanced artificial intelligence models for improved yield prediction and seed quality evaluation in chickpea. The presence of high phenotypic variability in growth, phenology, yield, and seed quality traits has ensured the availability of a sound dataset for training the model, which has been able to generalize well. The model has shown high predictive accuracy (R
2 = 0.94) with low RMSE and MAE values, ensuring the accuracy of yield prediction. The seed quality classification accuracy of 95.2% has also demonstrated the high multi-class discrimination ability of the model. The comparison study has shown high performance gains over regression, machine learning, and statistical models, emphasizing the advantage of using nonlinear learning and attention-based optimization.