Sensor validation results as illustrated in Table 3 confirmed strong correlations against laboratory reference methods for all three nutrients (N: r = 0.91; P: r = 0.87; K: r = 0.89; all p<0.001), with RMSE ranging from 6.1% (N) to 9.3% (K) of the laboratory mean-within the 10% acceptability threshold. The first published validation of this instrument class against laboratory methods for Indian semi-arid soils as stated by
Van Klompenburg et al. (2020). The lower P correlation (r = 0.87) reflects the known sensitivity of available phosphorus to instantaneous soil pH and moisture (
Anonymous, 2017), which can cause divergence from the time-averaged Olsen extraction.
Crop-class separability (Table 4) was confirmed by one-way ANOVA with Tukey HSD: N (F
3,3596 = 5,412; p<0.001), P (F
3,3596 = 2,871; p<0.001) and K (F
3,3596 = 3,104; p < 0.001)
(Pedregosa et al., 2011). The mango-coconut K pair showed the smallest effect (p = 0.003), explaining most classifier errors between these classes
(Kalimuthu et al., 2020 and
Elbasi et al., 2023). Feature correlation analysis (Fig 3) confirmed N-K (r = 0.85) and N-P (r = 0.71) as the strongest pairs, consistent with balanced NPK fertilisation practices in the district
Velayutham et al., 1999 and
Anonymous, 2017).
Classifier performance (Table 5): 10-fold cross-validation yielded Decision Tree accuracy of 96.7±0.4% and KNN accuracy of 96.5±0.5%
(Pedregosa et al., 2011). These results are consistent with prior ML crop prediction studies reporting accuracies in the 85-97% range
(Elbasi et al., 2023; Gupta et al., 2023; Ramaiah et al., 2023; Musanase et al., 2023; Prity et al., 2024). On the held-out test set, the CWMV ensemble achieved 97.3%, reducing total misclassifications from 30 (Decision Tree) and 27 (KNN) to 24. McNemar’s test confirmed statistically significant improvement over both individual classifiers (p = 0.02 and p = 0.03 respectively). The 27 vs 24 error improvement over naïve voting was not significant (p = 0.09), but all three corrected errors were mango-coconut confusions-the agronomically highest-consequence pair (Section 4.1). Mean inference latency was 187 ms (n = 100 cycles)
(Islam et al., 2023 and
Senapaty et al., 2023). Flask web dashboard showing real-time sensor readings is shown in Fig 4.
Seasonal robustness (Table 6):
kharif CWMV accuracy was 97.1% (95% CI: 95.6-98.3%) and rabi 96.7% (95% CI: 94.9-98.0%); the 0.4 percentage-point difference was not statistically significant (McNemar’s test, p = 0.31), confirming year-round reliability. This seasonal consistency has not been previously reported for comparable sensor-based and ML based crop prediction systems
(Islam et al., 2023; Manju et al., 2024;
Saha et al., 2025; Metagar et al., 2024; Hassan et al., 2026). Hardware and energy results (Table 7): total hardware cost was ₹4,200 (≈US$50), yielding ≈ ₹0.19 per prediction-99.95% less than laboratory testing
(Kapoor et al., 2025). Energy consumption was 0.101 Wh per 1,000 inference cycles, enabling approximately 81 hours on a 10,000 mAh battery or 24-hour autonomy with a 5 W solar panel under Karnataka’s average solar irradiance
(Saha et al., 2025).
Sensor validity and agronomic implications
The validated NPK threshold profiles (Table 4) can be directly mapped to ICAR fertilisation recommendations (
Anonymous, 2017). A ‘banana’ prediction (N > 100, K > 55 mg kg
-1) signals a nutrient-rich soil requiring balanced P supplementation rather than blanket NPK application (ICAR recommendation: 200–220 kg N ha
-1 (
Anonymous, 2017)). A ‘maize’ prediction (N < 50, K < 30 mg kg
-1) indicates nitrogen deficiency requiring urea at 120-150 kg N ha
-1 (
Anonymous, 2017). ‘Mango’ or ‘coconut’ predictions indicate moderate NPK status for perennial orchard establishment (
Anonymous, 2017); management diverges sharply between these species -coconut requires 500 g N palm
-1 yr
-1 and 7-10 years to full production, while mango reaches production in 3-5 years at 50-100 kg N ha
-1 (
Anonymous, 2017).
The mango-coconut confusion carries the greatest agronomic risk of all classifier errors (
Anonymous, 2017), as a misclassification could direct a farmer to a substantially more capital-intensive, longer-gestation planting decision. The CWMV ensemble’s targeted reduction of these confusions from 11 to 7 on the test set is therefore disproportionately valuable relative to its overall accuracy increment. Future work should implement a cost-sensitive classifier that formally penalises this confusion pair
(Dey et al., 2024).
Accuracy compared with prior work
The CWMV accuracy of 97.3% and individual classifier accuracies of 96.-96.7% are directly comparable with
Manju et al., (2024), who achieved 97.3% KNN on a four-crop sensor-collected Indian dataset with the same 75:25 split -providing cross-state empirical validation. Higher accuracies reported by
Dey et al., (2024) (XGBoost, 99.09% on Kaggle benchmark) and
Rani et al. (2023) (Gradient Boosting, 99.27%) were obtained on curated benchmark datasets, which consistently overestimate field performance as noted by
Saha et al., (2025). The accuracy of 97.3% is also consistent with findings of
Elbasi et al., (2023) (Bayes Net, 99.59% on benchmark) and surpasses the 90.4% and 92.1% reported for KNN and SVM by
Ramaiah et al., (2023) on similar Indian multi-crop datasets. The narrow cross-validation SDs (0.4-0.5%) confirm the robustness of these estimates
(Pedregosa et al., 2011). The sensor validation results address a foundational gap identified by
Van Klompenburg et al. (2020) in their 50-study systematic review: in-situ sensor measurements are rarely validated against laboratory reference methods before use in classifier training.
Practical deployment and limitations
At ₹0.19 per prediction-99.95% less than laboratory testing
(Kapoor et al., 2025) -the system is economically accessible to smallholder farmers. The edge-computing architecture, requiring no cloud connectivity, outperforms cloud-dependent systems reviewed by
Islam et al. (2023) and
Senapaty et al. (2023, 2024) in connectivity-limited settings. Energy consumption of 0.101 Wh per 1,000 cycles confirms solar-powered autonomous deployment is feasible
(Saha et al., 2025). The study has several limitations: it covers four crops in one district; sensor calibration was conducted on 20 samples from one soil order
(Velayutham et al., 1999); and the CWMV improvement over naïve voting was not statistically significant (p = 0.09). Multi-region validation-including temperate crops relevant to broader agricultural audiences (
Anonymous, 2023 and
Musanase et al., 2023) is needed before broader generalisation. Future work will expand the sensor suite to include soil pH, Ca and Mg
(Dey et al., 2024); broaden the crop portfolio
(Elbasi et al., 2023; Gupta et al., 2023); implement SHAP explainability
(Rani et al., 2023); and develop a cost-sensitive classifier penalising agronomically high-consequence misclassification pairs
(Babu et al., 2024).