A Deep Residual Learning Framework using ResNet-152 for Accurate Multi-class Classification of Shrimp Diseases

1LIG SYSTEM, 198, Hannam-daero, Yongsan-gu, Seoul, Korea.

Background: Shrimp aquaculture is a major contributor to global seafood production. Disease outbreaks remain a serious challenge, causing severe economic losses and threatening food security. Timely and accurate disease diagnosis is essential for effective farm management. However, conventional diagnostic methods are labor-intensive, time-consuming and require expert knowledge.

Methods: This study proposes a deep learning–based approach for automated multi-class shrimp disease classification using a ResNet-152 convolutional neural network. The publicly available Tiger Shrimp BD dataset was used. It contains 3,574 RGB images of Penaeus monodon grouped into four classes. A substantial portion of the dataset (approximately 72%) consists of synthetically generated images, while the remaining images were collected under real farming conditions across multiple regions of Bangladesh. This combination introduces variability in illumination, background and shrimp orientation but may also affect real-world generalizability. The dataset was divided into training, validation and test sets using a stratified 60:20:20 split to preserve class distribution. Image preprocessing included resizing, normalization and data augmentation. A controlled augmentation pipeline transformed each image using horizontal flip (p=0.5), rotation (-10° to 10°), zoom (0.9-1.1) and contrast adjustment (0.8-1.2), improving robustness and reducing overfitting. Transfer learning was applied and the model was optimized using the Adam optimizer with categorical cross-entropy loss.

Result: The proposed ResNet-152 model achieved strong classification performance. An overall test accuracy of 95.82% (95% CI: 93.17%-96.37%) was obtained, along with a Matthews Correlation Coefficient of 0.9442. High classification metrics were observed across all classes. The Healthy and Yellow Head categories showed particularly strong performance. ROC and precision-recall analyses confirmed robust threshold-independent discrimination, with per-class AUC values exceeding 0.98. Qualitative predictions further demonstrated reliable identification of disease patterns under diverse conditions. The proposed model demonstrates strong classification performance. However, reliance on augmented data and lack of external validation may affect generalizability. Future work will focus on independent dataset validation and efficient deployment.

Shrimp aquaculture is one of the fastest-growing sectors of global food production and plays a crucial role in ensuring food security, rural livelihoods and export income, particularly in Asia, Africa and Latin America. According to the Food and Agriculture Organization of the United Nations, global aquaculture production continues to expand steadily, with crustaceans representing a high-value commodity in international seafood markets (FAO, 2022). Shrimp farming contributes significantly to employment across hatcheries, grow-out farms, feed industries and processing chains. Its economic importance is especially pronounced in developing countries, where aquaculture serves as both a livelihood strategy and a source of foreign exchange (Ray et al., 2021). Among cultivated shrimp species, Penaeus monodon (giant tiger shrimp) remains one of the most commercially valuable. The species is favored for its rapid growth rate, tolerance to a wide range of salinity conditions and adaptability to both extensive and intensive farming systems (Suwoyo et al., 2024; Waiho et al., 2024). P. monodon continues to be widely farmed in South and Southeast Asia, including Bangladesh, India, Thailand and Vietnam, where it supports millions of small-scale and semi-intensive producers. Despite advances in breeding, nutrition and farm management, shrimp aquaculture remains highly vulnerable to disease outbreaks, which continue to be the primary limitation to sustainable production worldwide (Boyd et al., 2021; Asmild et al., 2023).
       
Infectious diseases pose a persistent and severe threat to shrimp farming. Viral, bacterial and parasitic pathogens can spread rapidly under intensive culture conditions, leading to mass mortality events and substantial economic losses. Among these, White Spot Syndrome Virus (WSSV), Yellow Head Disease (YHD) and Black Gill disease are particularly destructive. WSSV is considered the most devastating viral pathogen affecting shrimp aquaculture, with reported mortality rates approaching 100% within a few days of infection (Bhassu et al., 2024; El-Saadony et al., 2022). Yellow Head Disease, caused by yellow head virus (YHV), results in rapid onset of lethargy, yellow discoloration of the cephalothorax and  high mortality, especially in P. monodon populations (AlZubi, 2023; Cho, 2024; Kim and AlZubi, 2024). Black Gill disease, often associated with fungal or bacterial infections and environmental stress, leads to melanization of the gills, impaired respiration and reduced market value, even when mortality is moderate (Senapin et al., 2010; Maezono et al., 2025).
       
The economic consequences of shrimp disease outbreaks are profound. Global losses attributed to shrimp diseases are estimated to reach several billion US dollars annually, affecting farm profitability, supply stability and food security (Asche et al., 2020; FAO, 2022; Patil et al., 2025). Disease outbreaks are exacerbated by high stocking densities, poor water quality, climate variability and biosecurity breaches. Climate change further amplifies disease risks by altering temperature regimes and increasing environmental stress on cultured shrimp (Tamut et al., 2025).  Early and accurate disease diagnosis is therefore essential for effective disease management and mitigation. Conventional diagnostic approaches include visual inspection of clinical signs, histopathological examination and molecular techniques such as polymerase chain reaction (Mohammad et al., 2026; Ozçelik et al., 2026). While PCR-based diagnostics provide high sensitivity and specificity, they require specialized laboratory facilities, trained personnel and significant processing time (Singh et al., 2025; Paek et al., 2026; Souza et al., 2026; Sriram and Kumari, 2026).
       
These limitations highlight the urgent need for rapid, automated and cost-effective diagnostic tools that can operate under real-world farm conditions. Recent advances in artificial intelligence (AI) and deep learning have opened new opportunities for image-based disease diagnosis in agriculture and aquaculture. Convolutional neural networks (CNNs) have demonstrated remarkable performance in extracting discriminative visual features from complex images and have been successfully applied to plant disease detection, livestock health monitoring and aquatic species classification (Lee et al., 2025; Han et al., 2025). In shrimp aquaculture, image-based deep learning systems offer the potential to detect disease symptoms directly from RGB images captured using mobile phones or low-cost cameras (Roy et al., 2025). Recent review studies have highlighted the growing role of deep learning, particularly CNN-based approaches, in aquaculture disease detection and health monitoring, while also noting challenges related to data variability, model generalization and real-world deployment (Tamut et al., 2025).
       
In this context, the present study explores the use of a deep residual network (ResNet-152) for multi-class shrimp disease classification using the TigerShrimpBD dataset. The study focuses on evaluating model performance across multiple disease classes under diverse image conditions representative of practical farming environments. In addition to overall accuracy, emphasis is placed on class-wise evaluation metrics and analysis of misclassification patterns, particularly among visually similar disease categories. Such an approach provides a more comprehensive understand ing of model behavior and its potential applicability in real-world shrimp disease diagnosis. Overall, this study contributes to the ongoing development of AI-based tools for aquaculture by assessing the effectiveness of a widely used deep learning architecture within a structured and application-oriented evaluation framework.
Dataset description
 
This study utilized the publicly available TigerShrimpBD dataset, which contains images of Penaeus monodon collected under real shrimp-farming conditions (Ahmed and Farid, 2025). The dataset consists of 3574 RGB images belonging to four categories: Black Gill (854 images), Yellow Head Disease (896 images), White Spot Syndrome Virus (978 images) and Healthy shrimp (846 images). Of the total images, 1001 were captured manually by the dataset creators, while 2573 were synthetically generated through data-augmentation techniques to increase category balance and variability. All images were taken using mobile phone cameras, introducing natural variations in illumination, orientation, background patterns and shrimp positioning. Data collection occurred across four shrimp-farming regions in Bangladesh: Godanra village in Assasuni Upazila, Bharashimla village in Kaliganj Upazila, Bali Krishnapur village in Debhata Upazila and Fultola village in Bagerhat District. These locations differ in environmental and water-quality characteristics, providing a diverse set of visual conditions that enhance model robustness. Fig 1(a) shows class-wise image distribution for Black Gill, Yellow Head Disease, WSSV and Healthy shrimp.

Fig 1: (a) Class distribution of the dataset. (b) Distribution of training, validation and test splits (60:20:20) with maintained class balance.


       
For model training and evaluation, the dataset was reorganized into three subsets using a stand ard 60:20:20 division. This ensures balanced representation while maintaining the natural class proportions. Using the total image count Ntotal=3574, the images were allocated as follows:
 
Ntrain=0.6 x Ntotal,Nval= 0.2Ntotal,Ntest= 0.2Ntotal
 
After rounding, the final distribution consisted of 2142 training images, 715 validation images and 717 test images. All subsets retained a four-class structure identical to the original dataset. Fig 1(b) presents the dataset split distribution, illustrating the stratified 60:20:20 partitioning into training, validation and testing sets.
 
Image Pre-processing
 
All images were resized to 224 × 224 × 3 pixels so they match the input dimensions required by the ResNet152 architecture. Each image X was normalized using the stand ard ResNet preprocessing function:
 
Xproc = preprocessinput (x)
 
This transformation performs channel-wise scaling based on ImageNet statistics. To visualize images after model transformations, a de-normalization function was applied:

 
This restores pixel values to the range 0-255, enabling clear visualization of feature maps and interpretability outputs. A compact augmentation pipeline was used to improve model generalization, where each input image X was transformed to X′using four operations: rand om horizontal flip (p=0.5), rand om rotation θ∈[-10°, 10°]), rand om zoom α∈[0.9,1.1]) and rand om contrast adjustment β∈[0.8,1.2]). These transformations introduce variations in orientation, scale and illumination, thereby reducing overfitting and improving robustness. Fig 2 illustrates examples of the augmented images.
 
Random Horizontal Flip (X′)= Flip (X, p = 0.5)
Random Rotation (±10%) X′= Rθ (X), θ ∈[-0.1,0.1]
Random Zoom (±10%) X′= Zα (X), α∈ [-0.1,0.1]
Random Contrast Adjustmemnt: X′= Cβ(X), β ∈ [0.8,1.2]

Fig 2: Sample augmented images generated through rand om horizontal flipping, rand om rotation (±10%) and rand om zooming (±10%).


 
ResNet-152 model explanation
 
The architecture shown in Fig 3 represents a ResNet-152 deep convolutional neural network. The model follows a five-stage residual learning framework, where each stage extracts features with increasing semantic complexity. The input to the network is a shrimp image resized to 224 × 224 × 3. The three channels correspond to RGB values and pixel intensities are normalized before convolution to ensure stable learning.

Fig 3: Architecture of the proposed ResNet152-based shrimp disease classification model showing the pretrained backbone and classification layer.


       
As shown in Fig 3, the input image is first processed by an initial convolution layer with a large receptive field. This operation is defined as, X1=ReLU(W7×7 * X0+b) where a 7×7 kernel with 64 filters and a stride of 2 is used. This layer captures low-level features such as edges and textures. The output is then passed through a max-pooling layer, expressed as, X2(i,j) = max(m,n)∈ΩX1(i+m,j+n), with a 3×3 window and stride of 2. This step reduces spatial resolution, suppresses noise and lowers computational cost.
       
The core of the network consists of stacked residual blocks. Each block follows the residual learning principle Y= F(x, W)+X, where X is the input feature map, F(x,W) denotes the residual mapping and y is the output. The element-wise addition enables direct gradient flow across layers, thereby preventing vanishing gradients and stabilizing the training of the very deep ResNet-152 model. Feature extraction is performed across multiple stages. The Conv2 stage contains three residual blocks with a 1×1→3×3→1×1 bottleneck structure and channel dimensions of 64, 64 and 256, respectively, enabling the learning of basic shape features. The Conv3 stage includes eight residual blocks with channel sizes of 128, 128 and 512, allowing the network to capture mid-level patterns. The Conv4 stage is the deepest, comprising thirty-six residual blocks with channel dimensions of 256, 256 and 1024 and it extracts fine, disease-specific features. Finally, the Conv5 stage contains three residual blocks with channel sizes of 512, 512 and 2048, producing high-level semantic representations.
       
Following the convolutional backbone, a global average pooling layer aggregates spatial information across feature maps. This operation is defined as:

 
Where,
Ak(i,j)= Activation of channel k.
H and W= Spatial dimensions. This step converts feature maps into a compact feature vector and reduces overfitting.
       
The pooled feature vector is then passed to a fully connected dense layer, defined as z=Wf+b, with c=4 output neurons corresponding to healthy and unhealthy classes. A softmax activation function is applied to compute class probabilities using: 

 
Where,
The probabilities sum to one. The model was compiled using the Adam optimizer, which adaptively updates model parameters during training. The parameter update rule is given by

 
Where,
θrepresents the model parameters at iteration t, mt and νt are the first and second-order moment estimates of the gradients and ∈ is a small constant for numerical stability. The learning rate (5 × 10-5 ) was selected to ensure stable fine-tuning of pretrained weights and prevent divergence. Early stopping (patience = 5) was used to avoid overfitting while maintaining computational efficiency. Categorical cross-entropy was used as the loss function to measure the difference between predicted and true class distributions and it is defined as:

 
Where,
Yk = Ground truth label.
P= Predicted probability for class. Model accuracy was used as the primary evaluation metric during training.
       
Training was performed for a maximum of 100 epochs. Model checkpointing was applied to save the weights corresponding to the highest validation accuracy. Early stopping was employed to prevent overfitting and training was terminated if validation accuracy did not improve for five consecutive epochs. After training, model performance was evaluated using several quantitative metrics. The confusion matrix was computed as:

 
Cij= I {x:y(x) = i, y(x) = j} I
 
Where,
y(x)= True class.
ŷ(x) = Predicted class. From the confusion matrix, precision, recall and F1-score were derived. Precision and recall are defined as:







 
In addition, receiver operating characteristic (ROC) curves were generated to analyze the trade-off between true positive and false positive rates. The true positive rate and false positive rate are defined as:


The area under the ROC curve (AUC) was used as a threshold-independent performance indicator. Precision-recall curves were also evaluated and the average precision (AP) was computed as:

AP = n ∑ (Rn - Rn - 1) Pn

Where,
Pn and Rn represent precision and recall at the n-th threshold. Together, these metrics provide a comprehensive assessment of the model’s classification performance.
       
An ablation study was conducted to assess the impact of different backbone architectures on classification performance. Three pretrained models-MobileNetV2, ResNet50 and ResNet152-were evaluated under identical training conditions. Each model was trained using the same dataset split, preprocessing steps and hyperparameters to ensure a fair comparison. Performance was assessed using accuracy, Matthews Correlation Coefficient (MCC), parameter count, training time and inference time. This analysis was used to examine the trade-off between model complexity and predictive performance and to support the selection of the final architecture.
       
A 5-fold cross-validation strategy was employed to ensure robust and reliable evaluation of model performance across different data splits.
The ResNet-152 model was trained for a maximum of 100 epochs, with early stopping triggered at epoch 49 (Fig 4). The best model weights were restored from epoch 44, which achieved the highest validation accuracy. During early training, performance improved rapidly. At epoch 1, training accuracy was 26.05%, while validation accuracy reached 42.10%, indicating effective transfer learning. By epoch 5, training accuracy increased to 77.83%, with validation accuracy of 78.60%. This shows fast convergence and good feature reuse from pretrained weights. Loss values consistently decreased, confirming stable optimization.

Fig 4: Training and validation performance of the ResNet-152 model, showing epoch-wise accuracy and loss curves.


       
From epochs 10 to 20, the model showed strong generalization. Training accuracy improved from 86.48% to 91.60%, while validation accuracy increased from 85.87% to 92.87%. Validation loss steadily declined, indicating reduced overfitting. The residual connections helped maintain gradient flow across deep layers. The best validation performance was achieved at epoch 44, with a validation accuracy of 95.52% and a validation loss of 0.1504. Training accuracy at this stage was 96.04%. The small gap between training and validation accuracy suggests good generalization. After this point, validation accuracy plateaued, while training accuracy continued to increase slightly, indicating the onset of overfitting. Early stopping effectively prevented performance degradation. Overall, the results demonstrate that ResNet-152 is highly effective for multi-class classification of shrimp diseases.
       
Fig 5 presents the confusion matrix obtained from the ResNet-152 model on the test dataset with four shrimp classes. Each row represents the true class and each column represents the predicted class. Diagonal values indicate correct predictions, while off-diagonal values represent misclassifications. For the Black Gill class, 159 samples were correctly classified. A small number were misclassified as WSSV (11 samples) and Healthy (1 sample). No Black Gill sample was misclassified as Yellow head. This indicates strong discriminative ability, with limited confusion mainly with WSSV, likely due to visual symptom similarity. For the Healthy class, 169 samples were correctly identified. Only one sample was misclassified as Yellow head. No confusion occurred with Black Gill or WSSV. This shows that healthy shrimp features were clearly separated from diseased samples. For the WSSV class, 182 samples were correctly classified. Minor confusion occurred with Black Gill (12 samples), Healthy (1 sample) and Yellow head (1 sample).

Fig 5: Confusion matrix of the ResNet152 model on the test dataset, where each row represents the true class and each column represents the predicted class.


       
The higher confusion with Black Gill suggests overlapping texture or lesion patterns in some images. For the Yellow head class, 177 samples were correctly classified. Only three samples were misclassified, one each as Black Gill, Healthy and WSSV. This indicates robust recognition of Yellow Head disease features. Overall performance is strong, as most predictions lie along the diagonal. The confusion matrix supports high classification accuracy and balanced performance across classes. The low off-diagonal values confirm that the ResNet-152 backbone effectively learned disease-specific visual patterns. The remaining errors are mainly between visually similar disease classes, which is expected in real-world shrimp disease diagnosis.
       
Table 1 summarizes the classification performance of the ResNet-152 model on the test dataset. For Black Gill, the model achieved a precision of 0.9244 and a recall of 0.9298. This indicates reliable detection with limited false positives and false negatives. Minor confusion with WSSV affected the recall slightly. For the Healthy class, the model performed exceptionally well. Precision reached 0.9826 and recall reached 0.9941. This shows that healthy shrimp were almost perfectly separated from diseased samples. The F1-score of 0.9883 confirms robust discrimination. For WSSV, precision and recall were 0.9381 and 0.9286, respectively. The slightly lower recall reflects confusion mainly with Black Gill. The F1-score of 0.9333 still indicates strong performance on this clinically important disease. For Yellow head, the model achieved very high precision (0.9888) and recall (0.9833). The F1-score of 0.9861 shows that Yellowhead disease features were learned effectively with minimal misclassification. An overall test accuracy of 95.82% (95% CI: 93.17%–96.37%) was obtained, along with a Matthews Correlation Coefficient (MCC) of 0.9442. MCC accounts for true and false predictions across all classes. A value close to 1 indicates strong agreement between predictions and ground truth. This high MCC confirms that the ResNet-152 model is robust and reliable for multi-class shrimp disease classification.

Table 1: Classification report of the ResNet152 model including precision, recall, F1-score and support for each class.


       
Fig 6 illustrates the ROC and precision-recall (PR) curves of the ResNet-152 model for the four shrimp classes. In Fig 6(a), the ROC curves for all classes are positioned close to the top-left corner, indicating a high true positive rate with a low false positive rate. The Healthy and Yellowhead classes exhibit near-perfect discrimination, as their curves approach the upper boundary of the plot. Black Gill and WSSV also demonstrate strong separability, with only slight deviation from the ideal curve. The dashed diagonal line represents rand om classification and all curves lie well above this line, confirming that the model performs significantly better than chance across all classes. Per-class AUC was computed, yielding values of 0.992 for Black Gill, 1.000 for Healthy, 0.991 for WSSV and 1.000 for Yellowhead.

In Fig 6(b), the P-R curves remain high across most recall values. This shows that the model maintains high precision even when recall increases. In practical terms, the model detects diseased shrimp accurately while producing very few false alarms. Healthy and Yellowhead classes show almost flat curves near the top, indicating extremely reliable predictions. Black Gill and WSSV show a small drop at very high recall levels, which suggests limited confusion between visually similar disease patterns.

Fig 6: ROC and Precision-Recall curves for all classes, demonstrating the model’s discriminative performance across thresholds.


       
Fig 7 presents representative qualitative prediction results of the ResNet-152 model on test images from all four shrimp classes. The displayed examples were selected to include high-confidence correct predictions as well as representative misclassifications, providing a balanced assessment of model performance across classes. For each image, the true label (T), predicted label (P) and prediction confidence are shown. For the Black Gill class, most samples are correctly classified with high confidence values above 85%. The visual symptoms, such as darkened gills and body discoloration, are consistently recognized by the model. A few samples show moderate confidence, indicating natural variation in disease appearance, but predictions remain correct. For the Healthy class, the model shows very strong performance. All displayed healthy shrimp are correctly classified with confidence values close to 99%. The clean body surface, normal coloration and absence of lesions are clearly learned features. This confirms that healthy shrimp are well separated from diseased classes.

Fig 7: Qualitative prediction examples of the ResNet-152 model on test shrimp images, showing true labels (T), predicted labels (P) and corresponding confidence scores.


       
For the WSSV class, most samples are correctly predicted, though confidence scores are comparatively lower than for Healthy and Yellowhead. One example is misclassified as Black Gill, highlighted in red. This error suggests visual overlap between WSSV symptoms and Black Gill characteristics, such as body darkening and texture changes. Despite this, the majority of WSSV samples are correctly identified. For the Yellow head class, predictions are highly accurate with strong confidence values, mostly above 95%. Even in cases with lower confidence, the predicted class remains correct. This indicates that Yellowhead disease features are distinct and consistently captured by the model.
       
The ablation study results, comparing the performance of different backbone architectures, are summarized in Table 2.

Table 2: Ablation study comparing MobileNetV2, ResNet50 and ResNet152 in terms of performance and computational efficiency.


       
The ablation study indicates that ResNet152 outperformed both ResNet50 and MobileNetV2 in terms of classification performance, demonstrating its ability to capture more complex and discriminative features. While MobileNetV2 provided faster inference and lower computational cost and ResNet50 offered a balanced trade-off between efficiency and accuracy, ResNet152 achieved the most reliable predictions. Therefore, ResNet152 was selected as the final model, as its superior performance outweighs the increased computational complexity for the given application. Using 5-fold cross-validation, the model achieved an overall accuracy of 94.84% and an MCC of 0.9312, indicating strong and consistent classification performance.
     
Recent studies have explored diverse deep learning strategies for shrimp disease detection, including ensemble CNNs, lightweight object detectors, capsule networks and specialized convolutional architectures (Table 3).  However, direct comparison across studies is inherently limited because datasets differ significantly in size, class distribution, imaging conditions and disease definitions. Therefore, performance values should be interpreted as relative indicators rather than strict benchmarks. Büyükarıkan (2025) demonstrated that ensemble learning with heterogeneous CNNs can significantly improve performance, achieving up to 97.3% accuracy using a weighted combination of MobileNet and DenseNet models. While effective, ensemble methods increase computational complexity and deployment cost.  In addition, such models typically have higher inference latency due to multiple forward passes, which may limit real-time applicability in field conditions. Yuhuan et al., (2025) focused on real-time detection using an improved YOLOv8n model, achieving 92.7% mAP@0.5 with reduced parameters, making it suitable for edge deployment, though primarily designed for object detection rather than multi-class image classification. Raj et al., (2025) proposed an Enhanced Recurrent Capsule Network with hybrid metaheuristic optimization, achieving 95.2% accuracy, but at the expense of increased architectural and training complexity. Ramachandran et al. (2023) targeted single-disease diagnosis and reported 97.22% accuracy for WSSV detection using a Dense Inception CNN, demonstrating strong performance but limited generalizability across multiple diseases.

Table 3: Comparison of recent shrimp disease detection studies, including dataset size, model architecture and reported performance.


       
In contrast, the present study employs a single deep residual network, ResNet-152, to perform multi-class classification of shrimp diseases using the TigerShrimpBD dataset. Despite relying on a single model, the proposed approach achieved 95.82% accuracy and a high MCC of 0.9442, indicating strong agreement between predictions and ground truth across all classes. The strong ROC and P-R performance further confirms reliable threshold-independent discrimination. In terms of efficiency, the single-model design also reduces computational overhead compared to ensemble-based approaches. ResNet-152 requires only one forward pass during inference, making it more practical for deployment scenarios where latency and hardware resources are constrained, although it still remains heavier than lightweight architectures such as YOLOv8n variants. These results highlight that deep residual architectures remain highly competitive for shrimp disease diagnosis and offer practical advantages for scalable aquaculture health monitoring systems.
 
Limitations
 
This study has several limitations. A substantial portion of the dataset (approximately 72%) consists of synthetically generated images, which may introduce bias and limit generalizability to real-world farm conditions. The model was evaluated using a single stratified 60:20:20 split without cross-validation, which may affect robustness assessment. Although basic augmentation techniques were applied, more diverse transformations (e.g., illumination and color variations) were not fully explored. In addition, while an ablation study was conducted, the computational trade-offs between model complexity and efficiency require further investigation. Finally, misclassification analysis was limited and deeper interpretability methods could provide more insight into model behavior.
 
Future work
 
Future research will focus on validating the proposed model using independent real-world farm datasets to assess generalizability. More advanced data augmentation techniques, including illumination and color variations, will be explored to improve robustness. Lightweight architectures and model optimization will be investigated to enable deployment on resource-constrained devices. Additionally, enhanced interpretability methods, such as Grad-CAM and saliency analysis, will be further utilized to better understand model decisions and reduce misclassification.
This study proposed a deep learning–based framework for automated multi-class shrimp disease classification using a ResNet-152 architecture. The model was evaluated on the TigerShrimpBD dataset. Through effective preprocessing, data augmentation and transfer learning, the model achieved stable convergence and strong generalization. Experimental results show high classification performance across all four classes. The model achieved an overall accuracy of 95.82% and an MCC of 0.9442, indicating excellent agreement between predictions and ground truth. Class-wise metrics confirm balanced performance, with especially strong results for Healthy and Yellow Head disease. Most misclassifications occurred between visually similar diseases such as Black Gill and WSSV. ROC and precision-recall analyses further validated robust threshold-independent performance. Despite these results, the model relies on image-based data only and uses a computationally heavy architecture. Future work will focus on lightweight models for edge deployment, larger and more diverse datasets and real-time systems integrating environmental or sensor data for improved disease monitoring.
 
Disclaimers
 
The views and conclusions expressed in this article are solely those of the authors and do not necessarily represent the views of their affiliated institutions. The authors are responsible for the accuracy and completeness of the information provided, but do not accept any liability for any direct or indirect losses resulting from the use of this content.
Funding details
 
This research received no external funding.
 
Authors’ contributions
 
All authors contributed toward data analysis, drafting and revising the paper and agreed to be responsible for all the aspects of this work.
 
Data availability
 
The data analysed/generated in the present study will be made available from corresponding authors upon reasonable request.
 
Availability of data and materials
 
Not applicable
 
Use of artificial intelligence
 
Artificial intelligence techniques (deep learning models) were used only for research purposes in this study. No AI tools were used in the writing, editing, or preparation of this manuscript.
 
Declarations
 
Authors declare that all works are original and this manuscript has not been published in any other journal.
The authors declare that they have no conflict of interest.

  1. Ahmed, S.I. and Farid, D.M. (2025). TigerShrimpBD: A tiger shrimp image dataset (Version 1) [Data set]. Mendeley Data. https://doi.org/10.17632/9dj4sk5d55.1.

  2. AlZubi, A.A. (2023). Artificial intelligence and its application in the prediction and diagnosis of animal diseases: A Review. Indian Journal of Animal Research. 57(10): 1265-1271. doi: 10.18805/IJAR.BF-1684.

  3. Asche, F. and erson, J.L., Botta, R., Kumar, G., Abrahamsen, E.B., Nguyen, L.T. and Valderrama, D. (2020). The economics of shrimp disease. Journal of Invertebrate Pathology. 186: 107397. https://doi.org/10.1016/j.jip.2020.107397.

  4. Asmild, M., Hukom, V., Nielsen, R. and Nielsen, M. (2023). Is economies of scale driving the development in shrimp farming from Penaeus monodon to Litopenaeus vannamei? The case of Indonesia. Aquaculture. 579: 740178. https:// doi.org/10.1016/j.aquaculture.2023.740178.

  5. Bhassu, S., Shama, M., Tiruvayipati, S., Soo, T.C.C., Ahmed, N. and Yusoff, K. (2024). Microbes and pathogens associated with shrimps-implications and review of possible control strategies. Frontiers in Marine Science. 11. https://doi.org/ 10.3389/fmars.2024.1397708.

  6. Boyd, C.E., Davis, R.P. and McNevin, A.A. (2021). Perspectives on the mangrove conundrum, land use and benefits of yield intensification in farmed shrimp production: A review. Journal of the World Aquaculture Society. 53(1): 8-46. https://doi.org/10.1111/jwas.12841.

  7. Büyükarıkan, B. (2025). Robust shrimp disease detection using multi-model convolutional neural networks-based ensemble strategies. Aquacultural Engineering. 111: 102616. https:// doi.org/10.1016/j.aquaeng.2025.102616.

  8. Cho, O.H. (2024). Machine learning algorithms for early detection of legume crop disease. Legume Research. 47(3): 463-469. doi: 10.18805/LRF-788.

  9. El-Saadony, M.T., Swelum, A.A., Ghanima, M.M.A., Shukry, M., Omar, A.A., Taha, A.E., Salem, H.M., El-Tahan, A.M., El- Tarabily, K.A. and El-Hack, M.E.A. (2022). Shrimp production, the most important diseases that threaten it and the role of probiotics in confronting these diseases: A review. Research in Veterinary Science. 144: 126-140. https://doi.org/10.1016/j.rvsc.2022.01.009.

  10. FAO (2022). The State of World Fisheries and Aquaculture 2022: Towards Blue Transformation. Food and Agriculture Organization of the United Nations, Rome, Italy. ISBN: 978-92-5-136364-5.

  11. Han, G., Goncharov, A., Eryilmaz, M., Ye, S., Palanisamy, B., Ghosh, R., Lisi, F., Rogers, E., Guzman, D., Yigci, D., Tasoglu, S., Di Carlo, D., Goda, K., McKendry, R.A. and Ozcan, A. (2025). Machine learning in point-of-care testing: innovations, challenges and opportunities. Nature Communications. 16(1): 3165. https://doi.org/10.1038/s41467-025-58527-6.

  12. Kim, S.Y. and AlZubi, A.A. (2024). Blockchain and artificial intelligence for ensuring the authenticity of organic legume products in supply chains. Legume Research. 47(7): 1144-1150. doi: 10.18805/LRF-786.

  13. Lee, S., Park, J.S., Hong, J.H., Woo, H., Lee, C., Yoon, J.H., Lee, K., Chung, S., Yoon, D.S. and Lee, J.H. (2025). Artificial intelligence in bacterial diagnostics and antimicrobial susceptibility testing: Current advances and future prospects. Biosensors and Bioelectronics. 280: 117399. https:// doi.org/10.1016/j.bios.2025.117399.

  14. Maezono, M., Nielsen, R., Buchmann, K. and Nielsen, M. (2025). The current state of knowledge of the economic impact of diseases in global aquaculture. Reviews in Aquaculture. 17(3). https://doi.org/10.1111/raq.70039.

  15. Mohammad, A.A.S., Mohammad, S.I., Vasudevan, A. and Heindric, A. (2026). Intelligent decision support system for crop management using data analytics. Journal of Experimental Biology and Agricultural Sciences. 14(3): 9-14. https:// jebas.org/ojs/index.php/jebas/article/view/3796.

  16. Ozçelik, H., Sağlam, B. and Öcal, A.S. (2026). New natural methods for combating mold fungi in agricultural products: The Isparta example. Bulletin of Pure and Applied Sciences- Botany. 45B(2): 1-17. https://bpasjournals.com/botany/ index.php/journal/article/view/174/150.

  17. Paek, J.J., Kim, J., Bae, S., Paek, K. and Lee, Y. (2026). Chive- fortified fermentation enhances gamma-aminobutyric acid production by Levilactobacillus brevis PL9014. International Journal of Probiotics and Prebiotics. 21(1): 1-6. https://doi.org/10.37290/ijpp.v21i1.12.

  18. Patil, P.K., Geetha, R., Mishra, S.S., Abraham, T.J., Solanki, H.G., Sharma, S.R.K., Pradhan, P.K. et al (2025). Unveiling the economic burden of diseases in aquatic animal food production in India. Frontiers in Sustainable Food Systems. 9. https://doi.org/10.3389/fsufs.2025.1480094.

  19. Raj, A.S., Senthilkumar, S., Radha, R. and Muthaiyan, R. (2025). Enhanced recurrent capsule network with hyrbid optimization model for shrimp disease detection. Scientific Reports. 15(1): 10400. https://doi.org/10.1038/s41598-025-94413-3.

  20. Ramachandran, L., Mohan, V., Senthilkumar, S. and Ganesh, J. (2023). Early detection and identification of white spot syndrome in shrimp using an improved deep convolutional neural network. Journal of Intelligent and Fuzzy Systems. 45(4): 6429-6440. https://doi.org/10.3233/jifs-232687.

  21. Ray, S., Mondal, P., Paul, A.K., Iqbal, S., Atique, U., Islam, M.S., Mahboob, S., Al-Ghanim, K.A., Al-Misned, F. and Begum, S. (2021). Role of shrimp farming in socio-economic elevation and professional satisfaction in coastal communities. Aquaculture Reports. 20: 100708. https://doi.org/10. 1016/j.aqrep.2021.100708.

  22. Roy, S.M., Beg, M.M., Bhagat, S.K., Charan, D., Pareek, C., Moulick, S. and Kim, T. (2025). Application of artificial intelligence in aquaculture-recent developments and prospects. Aquacultural Engineering. 111: 102570. https://doi.org/ 10.1016/j.aquaeng.2025.102570.

  23. Senapin, S., Thaowbut, Y., Gangnonngiw, W., Chuchird, N., Sriurairatana, S. and Flegel, T.W. (2010). Impact of yellow head virus outbreaks in the whiteleg shrimp, Penaeus vannamei (Boone), in Thailand. Journal of Fish Diseases. 33(5): 421-430. https://doi.org/10.1111/j.1365-2761.2009. 01135.x

  24. Singh, J., Kashyap, R., Bansal, K., Das, R., Sand hu, K., Singh, G. and Saini, D.K. (2025). Transition from conventional to AI-based methods for detection of foliar disease symptoms in vegetable crops: A comprehensive review. Journal of Plant Pathology. 107(4): 1791-1814. https://doi.org/10. 1007/s42161-025-01983-2.

  25. Souza, K.F., Gonçalves, L.C.O. and Espindola, F.S. (2026). A broad overview of the multifaceted royal jelly and its applicability in health and performance. Current Topics in Nutraceutical Research. 24(1): 4-14. https://doi.org/10.37290/ctnr.v24i 1.8.

  26. Sriram, P. and Kumari, Y.S. (2026). Game theory in urban birds: A Nash equilibrium analysis using crows and pigeons. Bio- Science Research Bulletin. 42(1): 1-6. https://bpasjournals. com/life-sciences/index.php/journal/article/view/150/140.

  27. Suwoyo, H.S., Sahabuddin, S., Nawang, A., Makmur, M., Sahrijanna, A., Mulyaningrum, S. R.H. and Ilham, I. (2024). Growth performance of selected and non-selected black tiger shrimp (Penaeus monodon) on pond cultivation. BIO Web of Conferences. 136: 01002. https://doi.org/10.1051/ bioconf/202413601002.

  28. Tamut, H., Ghosh, R., Gosh, K. and Siddique, M.A.S. (2025). Enhancin disease detection in the aquaculture sector using convolutional neural networks analysis. Aquaculture Journal. 5(1): 6. https://doi.org/10.3390/aquacj5010006.

  29. Waiho, K., Ling, Y., Ikhwanuddin, M., Shu Chien, A. C., Afiqah Aleng, N., Wang, Y., Hu, M., Liew, H., Kasan, N. A., Peh, J.H. and Fazhan, H. (2024). Current Advances in the Black Tiger Shrimp Penaeus monodon Culture: A Review. Reviews in Aquaculture. 17(1). https://doi.org/10.1111/ raq.12958.

  30. Yuhuan, F., Gengchen, W., Fenghao, L., Ran, Z., Xufei, S. and  Hao, C. (2025). Lightweight shrimp disease detection research based on YOLOv8n.arXiv preprint. https:// doi.org/10.48550/arxiv.2507.02354.

A Deep Residual Learning Framework using ResNet-152 for Accurate Multi-class Classification of Shrimp Diseases

1LIG SYSTEM, 198, Hannam-daero, Yongsan-gu, Seoul, Korea.

Background: Shrimp aquaculture is a major contributor to global seafood production. Disease outbreaks remain a serious challenge, causing severe economic losses and threatening food security. Timely and accurate disease diagnosis is essential for effective farm management. However, conventional diagnostic methods are labor-intensive, time-consuming and require expert knowledge.

Methods: This study proposes a deep learning–based approach for automated multi-class shrimp disease classification using a ResNet-152 convolutional neural network. The publicly available Tiger Shrimp BD dataset was used. It contains 3,574 RGB images of Penaeus monodon grouped into four classes. A substantial portion of the dataset (approximately 72%) consists of synthetically generated images, while the remaining images were collected under real farming conditions across multiple regions of Bangladesh. This combination introduces variability in illumination, background and shrimp orientation but may also affect real-world generalizability. The dataset was divided into training, validation and test sets using a stratified 60:20:20 split to preserve class distribution. Image preprocessing included resizing, normalization and data augmentation. A controlled augmentation pipeline transformed each image using horizontal flip (p=0.5), rotation (-10° to 10°), zoom (0.9-1.1) and contrast adjustment (0.8-1.2), improving robustness and reducing overfitting. Transfer learning was applied and the model was optimized using the Adam optimizer with categorical cross-entropy loss.

Result: The proposed ResNet-152 model achieved strong classification performance. An overall test accuracy of 95.82% (95% CI: 93.17%-96.37%) was obtained, along with a Matthews Correlation Coefficient of 0.9442. High classification metrics were observed across all classes. The Healthy and Yellow Head categories showed particularly strong performance. ROC and precision-recall analyses confirmed robust threshold-independent discrimination, with per-class AUC values exceeding 0.98. Qualitative predictions further demonstrated reliable identification of disease patterns under diverse conditions. The proposed model demonstrates strong classification performance. However, reliance on augmented data and lack of external validation may affect generalizability. Future work will focus on independent dataset validation and efficient deployment.

Shrimp aquaculture is one of the fastest-growing sectors of global food production and plays a crucial role in ensuring food security, rural livelihoods and export income, particularly in Asia, Africa and Latin America. According to the Food and Agriculture Organization of the United Nations, global aquaculture production continues to expand steadily, with crustaceans representing a high-value commodity in international seafood markets (FAO, 2022). Shrimp farming contributes significantly to employment across hatcheries, grow-out farms, feed industries and processing chains. Its economic importance is especially pronounced in developing countries, where aquaculture serves as both a livelihood strategy and a source of foreign exchange (Ray et al., 2021). Among cultivated shrimp species, Penaeus monodon (giant tiger shrimp) remains one of the most commercially valuable. The species is favored for its rapid growth rate, tolerance to a wide range of salinity conditions and adaptability to both extensive and intensive farming systems (Suwoyo et al., 2024; Waiho et al., 2024). P. monodon continues to be widely farmed in South and Southeast Asia, including Bangladesh, India, Thailand and Vietnam, where it supports millions of small-scale and semi-intensive producers. Despite advances in breeding, nutrition and farm management, shrimp aquaculture remains highly vulnerable to disease outbreaks, which continue to be the primary limitation to sustainable production worldwide (Boyd et al., 2021; Asmild et al., 2023).
       
Infectious diseases pose a persistent and severe threat to shrimp farming. Viral, bacterial and parasitic pathogens can spread rapidly under intensive culture conditions, leading to mass mortality events and substantial economic losses. Among these, White Spot Syndrome Virus (WSSV), Yellow Head Disease (YHD) and Black Gill disease are particularly destructive. WSSV is considered the most devastating viral pathogen affecting shrimp aquaculture, with reported mortality rates approaching 100% within a few days of infection (Bhassu et al., 2024; El-Saadony et al., 2022). Yellow Head Disease, caused by yellow head virus (YHV), results in rapid onset of lethargy, yellow discoloration of the cephalothorax and  high mortality, especially in P. monodon populations (AlZubi, 2023; Cho, 2024; Kim and AlZubi, 2024). Black Gill disease, often associated with fungal or bacterial infections and environmental stress, leads to melanization of the gills, impaired respiration and reduced market value, even when mortality is moderate (Senapin et al., 2010; Maezono et al., 2025).
       
The economic consequences of shrimp disease outbreaks are profound. Global losses attributed to shrimp diseases are estimated to reach several billion US dollars annually, affecting farm profitability, supply stability and food security (Asche et al., 2020; FAO, 2022; Patil et al., 2025). Disease outbreaks are exacerbated by high stocking densities, poor water quality, climate variability and biosecurity breaches. Climate change further amplifies disease risks by altering temperature regimes and increasing environmental stress on cultured shrimp (Tamut et al., 2025).  Early and accurate disease diagnosis is therefore essential for effective disease management and mitigation. Conventional diagnostic approaches include visual inspection of clinical signs, histopathological examination and molecular techniques such as polymerase chain reaction (Mohammad et al., 2026; Ozçelik et al., 2026). While PCR-based diagnostics provide high sensitivity and specificity, they require specialized laboratory facilities, trained personnel and significant processing time (Singh et al., 2025; Paek et al., 2026; Souza et al., 2026; Sriram and Kumari, 2026).
       
These limitations highlight the urgent need for rapid, automated and cost-effective diagnostic tools that can operate under real-world farm conditions. Recent advances in artificial intelligence (AI) and deep learning have opened new opportunities for image-based disease diagnosis in agriculture and aquaculture. Convolutional neural networks (CNNs) have demonstrated remarkable performance in extracting discriminative visual features from complex images and have been successfully applied to plant disease detection, livestock health monitoring and aquatic species classification (Lee et al., 2025; Han et al., 2025). In shrimp aquaculture, image-based deep learning systems offer the potential to detect disease symptoms directly from RGB images captured using mobile phones or low-cost cameras (Roy et al., 2025). Recent review studies have highlighted the growing role of deep learning, particularly CNN-based approaches, in aquaculture disease detection and health monitoring, while also noting challenges related to data variability, model generalization and real-world deployment (Tamut et al., 2025).
       
In this context, the present study explores the use of a deep residual network (ResNet-152) for multi-class shrimp disease classification using the TigerShrimpBD dataset. The study focuses on evaluating model performance across multiple disease classes under diverse image conditions representative of practical farming environments. In addition to overall accuracy, emphasis is placed on class-wise evaluation metrics and analysis of misclassification patterns, particularly among visually similar disease categories. Such an approach provides a more comprehensive understand ing of model behavior and its potential applicability in real-world shrimp disease diagnosis. Overall, this study contributes to the ongoing development of AI-based tools for aquaculture by assessing the effectiveness of a widely used deep learning architecture within a structured and application-oriented evaluation framework.
Dataset description
 
This study utilized the publicly available TigerShrimpBD dataset, which contains images of Penaeus monodon collected under real shrimp-farming conditions (Ahmed and Farid, 2025). The dataset consists of 3574 RGB images belonging to four categories: Black Gill (854 images), Yellow Head Disease (896 images), White Spot Syndrome Virus (978 images) and Healthy shrimp (846 images). Of the total images, 1001 were captured manually by the dataset creators, while 2573 were synthetically generated through data-augmentation techniques to increase category balance and variability. All images were taken using mobile phone cameras, introducing natural variations in illumination, orientation, background patterns and shrimp positioning. Data collection occurred across four shrimp-farming regions in Bangladesh: Godanra village in Assasuni Upazila, Bharashimla village in Kaliganj Upazila, Bali Krishnapur village in Debhata Upazila and Fultola village in Bagerhat District. These locations differ in environmental and water-quality characteristics, providing a diverse set of visual conditions that enhance model robustness. Fig 1(a) shows class-wise image distribution for Black Gill, Yellow Head Disease, WSSV and Healthy shrimp.

Fig 1: (a) Class distribution of the dataset. (b) Distribution of training, validation and test splits (60:20:20) with maintained class balance.


       
For model training and evaluation, the dataset was reorganized into three subsets using a stand ard 60:20:20 division. This ensures balanced representation while maintaining the natural class proportions. Using the total image count Ntotal=3574, the images were allocated as follows:
 
Ntrain=0.6 x Ntotal,Nval= 0.2Ntotal,Ntest= 0.2Ntotal
 
After rounding, the final distribution consisted of 2142 training images, 715 validation images and 717 test images. All subsets retained a four-class structure identical to the original dataset. Fig 1(b) presents the dataset split distribution, illustrating the stratified 60:20:20 partitioning into training, validation and testing sets.
 
Image Pre-processing
 
All images were resized to 224 × 224 × 3 pixels so they match the input dimensions required by the ResNet152 architecture. Each image X was normalized using the stand ard ResNet preprocessing function:
 
Xproc = preprocessinput (x)
 
This transformation performs channel-wise scaling based on ImageNet statistics. To visualize images after model transformations, a de-normalization function was applied:

 
This restores pixel values to the range 0-255, enabling clear visualization of feature maps and interpretability outputs. A compact augmentation pipeline was used to improve model generalization, where each input image X was transformed to X′using four operations: rand om horizontal flip (p=0.5), rand om rotation θ∈[-10°, 10°]), rand om zoom α∈[0.9,1.1]) and rand om contrast adjustment β∈[0.8,1.2]). These transformations introduce variations in orientation, scale and illumination, thereby reducing overfitting and improving robustness. Fig 2 illustrates examples of the augmented images.
 
Random Horizontal Flip (X′)= Flip (X, p = 0.5)
Random Rotation (±10%) X′= Rθ (X), θ ∈[-0.1,0.1]
Random Zoom (±10%) X′= Zα (X), α∈ [-0.1,0.1]
Random Contrast Adjustmemnt: X′= Cβ(X), β ∈ [0.8,1.2]

Fig 2: Sample augmented images generated through rand om horizontal flipping, rand om rotation (±10%) and rand om zooming (±10%).


 
ResNet-152 model explanation
 
The architecture shown in Fig 3 represents a ResNet-152 deep convolutional neural network. The model follows a five-stage residual learning framework, where each stage extracts features with increasing semantic complexity. The input to the network is a shrimp image resized to 224 × 224 × 3. The three channels correspond to RGB values and pixel intensities are normalized before convolution to ensure stable learning.

Fig 3: Architecture of the proposed ResNet152-based shrimp disease classification model showing the pretrained backbone and classification layer.


       
As shown in Fig 3, the input image is first processed by an initial convolution layer with a large receptive field. This operation is defined as, X1=ReLU(W7×7 * X0+b) where a 7×7 kernel with 64 filters and a stride of 2 is used. This layer captures low-level features such as edges and textures. The output is then passed through a max-pooling layer, expressed as, X2(i,j) = max(m,n)∈ΩX1(i+m,j+n), with a 3×3 window and stride of 2. This step reduces spatial resolution, suppresses noise and lowers computational cost.
       
The core of the network consists of stacked residual blocks. Each block follows the residual learning principle Y= F(x, W)+X, where X is the input feature map, F(x,W) denotes the residual mapping and y is the output. The element-wise addition enables direct gradient flow across layers, thereby preventing vanishing gradients and stabilizing the training of the very deep ResNet-152 model. Feature extraction is performed across multiple stages. The Conv2 stage contains three residual blocks with a 1×1→3×3→1×1 bottleneck structure and channel dimensions of 64, 64 and 256, respectively, enabling the learning of basic shape features. The Conv3 stage includes eight residual blocks with channel sizes of 128, 128 and 512, allowing the network to capture mid-level patterns. The Conv4 stage is the deepest, comprising thirty-six residual blocks with channel dimensions of 256, 256 and 1024 and it extracts fine, disease-specific features. Finally, the Conv5 stage contains three residual blocks with channel sizes of 512, 512 and 2048, producing high-level semantic representations.
       
Following the convolutional backbone, a global average pooling layer aggregates spatial information across feature maps. This operation is defined as:

 
Where,
Ak(i,j)= Activation of channel k.
H and W= Spatial dimensions. This step converts feature maps into a compact feature vector and reduces overfitting.
       
The pooled feature vector is then passed to a fully connected dense layer, defined as z=Wf+b, with c=4 output neurons corresponding to healthy and unhealthy classes. A softmax activation function is applied to compute class probabilities using: 

 
Where,
The probabilities sum to one. The model was compiled using the Adam optimizer, which adaptively updates model parameters during training. The parameter update rule is given by

 
Where,
θrepresents the model parameters at iteration t, mt and νt are the first and second-order moment estimates of the gradients and ∈ is a small constant for numerical stability. The learning rate (5 × 10-5 ) was selected to ensure stable fine-tuning of pretrained weights and prevent divergence. Early stopping (patience = 5) was used to avoid overfitting while maintaining computational efficiency. Categorical cross-entropy was used as the loss function to measure the difference between predicted and true class distributions and it is defined as:

 
Where,
Yk = Ground truth label.
P= Predicted probability for class. Model accuracy was used as the primary evaluation metric during training.
       
Training was performed for a maximum of 100 epochs. Model checkpointing was applied to save the weights corresponding to the highest validation accuracy. Early stopping was employed to prevent overfitting and training was terminated if validation accuracy did not improve for five consecutive epochs. After training, model performance was evaluated using several quantitative metrics. The confusion matrix was computed as:

 
Cij= I {x:y(x) = i, y(x) = j} I
 
Where,
y(x)= True class.
ŷ(x) = Predicted class. From the confusion matrix, precision, recall and F1-score were derived. Precision and recall are defined as:







 
In addition, receiver operating characteristic (ROC) curves were generated to analyze the trade-off between true positive and false positive rates. The true positive rate and false positive rate are defined as:


The area under the ROC curve (AUC) was used as a threshold-independent performance indicator. Precision-recall curves were also evaluated and the average precision (AP) was computed as:

AP = n ∑ (Rn - Rn - 1) Pn

Where,
Pn and Rn represent precision and recall at the n-th threshold. Together, these metrics provide a comprehensive assessment of the model’s classification performance.
       
An ablation study was conducted to assess the impact of different backbone architectures on classification performance. Three pretrained models-MobileNetV2, ResNet50 and ResNet152-were evaluated under identical training conditions. Each model was trained using the same dataset split, preprocessing steps and hyperparameters to ensure a fair comparison. Performance was assessed using accuracy, Matthews Correlation Coefficient (MCC), parameter count, training time and inference time. This analysis was used to examine the trade-off between model complexity and predictive performance and to support the selection of the final architecture.
       
A 5-fold cross-validation strategy was employed to ensure robust and reliable evaluation of model performance across different data splits.
The ResNet-152 model was trained for a maximum of 100 epochs, with early stopping triggered at epoch 49 (Fig 4). The best model weights were restored from epoch 44, which achieved the highest validation accuracy. During early training, performance improved rapidly. At epoch 1, training accuracy was 26.05%, while validation accuracy reached 42.10%, indicating effective transfer learning. By epoch 5, training accuracy increased to 77.83%, with validation accuracy of 78.60%. This shows fast convergence and good feature reuse from pretrained weights. Loss values consistently decreased, confirming stable optimization.

Fig 4: Training and validation performance of the ResNet-152 model, showing epoch-wise accuracy and loss curves.


       
From epochs 10 to 20, the model showed strong generalization. Training accuracy improved from 86.48% to 91.60%, while validation accuracy increased from 85.87% to 92.87%. Validation loss steadily declined, indicating reduced overfitting. The residual connections helped maintain gradient flow across deep layers. The best validation performance was achieved at epoch 44, with a validation accuracy of 95.52% and a validation loss of 0.1504. Training accuracy at this stage was 96.04%. The small gap between training and validation accuracy suggests good generalization. After this point, validation accuracy plateaued, while training accuracy continued to increase slightly, indicating the onset of overfitting. Early stopping effectively prevented performance degradation. Overall, the results demonstrate that ResNet-152 is highly effective for multi-class classification of shrimp diseases.
       
Fig 5 presents the confusion matrix obtained from the ResNet-152 model on the test dataset with four shrimp classes. Each row represents the true class and each column represents the predicted class. Diagonal values indicate correct predictions, while off-diagonal values represent misclassifications. For the Black Gill class, 159 samples were correctly classified. A small number were misclassified as WSSV (11 samples) and Healthy (1 sample). No Black Gill sample was misclassified as Yellow head. This indicates strong discriminative ability, with limited confusion mainly with WSSV, likely due to visual symptom similarity. For the Healthy class, 169 samples were correctly identified. Only one sample was misclassified as Yellow head. No confusion occurred with Black Gill or WSSV. This shows that healthy shrimp features were clearly separated from diseased samples. For the WSSV class, 182 samples were correctly classified. Minor confusion occurred with Black Gill (12 samples), Healthy (1 sample) and Yellow head (1 sample).

Fig 5: Confusion matrix of the ResNet152 model on the test dataset, where each row represents the true class and each column represents the predicted class.


       
The higher confusion with Black Gill suggests overlapping texture or lesion patterns in some images. For the Yellow head class, 177 samples were correctly classified. Only three samples were misclassified, one each as Black Gill, Healthy and WSSV. This indicates robust recognition of Yellow Head disease features. Overall performance is strong, as most predictions lie along the diagonal. The confusion matrix supports high classification accuracy and balanced performance across classes. The low off-diagonal values confirm that the ResNet-152 backbone effectively learned disease-specific visual patterns. The remaining errors are mainly between visually similar disease classes, which is expected in real-world shrimp disease diagnosis.
       
Table 1 summarizes the classification performance of the ResNet-152 model on the test dataset. For Black Gill, the model achieved a precision of 0.9244 and a recall of 0.9298. This indicates reliable detection with limited false positives and false negatives. Minor confusion with WSSV affected the recall slightly. For the Healthy class, the model performed exceptionally well. Precision reached 0.9826 and recall reached 0.9941. This shows that healthy shrimp were almost perfectly separated from diseased samples. The F1-score of 0.9883 confirms robust discrimination. For WSSV, precision and recall were 0.9381 and 0.9286, respectively. The slightly lower recall reflects confusion mainly with Black Gill. The F1-score of 0.9333 still indicates strong performance on this clinically important disease. For Yellow head, the model achieved very high precision (0.9888) and recall (0.9833). The F1-score of 0.9861 shows that Yellowhead disease features were learned effectively with minimal misclassification. An overall test accuracy of 95.82% (95% CI: 93.17%–96.37%) was obtained, along with a Matthews Correlation Coefficient (MCC) of 0.9442. MCC accounts for true and false predictions across all classes. A value close to 1 indicates strong agreement between predictions and ground truth. This high MCC confirms that the ResNet-152 model is robust and reliable for multi-class shrimp disease classification.

Table 1: Classification report of the ResNet152 model including precision, recall, F1-score and support for each class.


       
Fig 6 illustrates the ROC and precision-recall (PR) curves of the ResNet-152 model for the four shrimp classes. In Fig 6(a), the ROC curves for all classes are positioned close to the top-left corner, indicating a high true positive rate with a low false positive rate. The Healthy and Yellowhead classes exhibit near-perfect discrimination, as their curves approach the upper boundary of the plot. Black Gill and WSSV also demonstrate strong separability, with only slight deviation from the ideal curve. The dashed diagonal line represents rand om classification and all curves lie well above this line, confirming that the model performs significantly better than chance across all classes. Per-class AUC was computed, yielding values of 0.992 for Black Gill, 1.000 for Healthy, 0.991 for WSSV and 1.000 for Yellowhead.

In Fig 6(b), the P-R curves remain high across most recall values. This shows that the model maintains high precision even when recall increases. In practical terms, the model detects diseased shrimp accurately while producing very few false alarms. Healthy and Yellowhead classes show almost flat curves near the top, indicating extremely reliable predictions. Black Gill and WSSV show a small drop at very high recall levels, which suggests limited confusion between visually similar disease patterns.

Fig 6: ROC and Precision-Recall curves for all classes, demonstrating the model’s discriminative performance across thresholds.


       
Fig 7 presents representative qualitative prediction results of the ResNet-152 model on test images from all four shrimp classes. The displayed examples were selected to include high-confidence correct predictions as well as representative misclassifications, providing a balanced assessment of model performance across classes. For each image, the true label (T), predicted label (P) and prediction confidence are shown. For the Black Gill class, most samples are correctly classified with high confidence values above 85%. The visual symptoms, such as darkened gills and body discoloration, are consistently recognized by the model. A few samples show moderate confidence, indicating natural variation in disease appearance, but predictions remain correct. For the Healthy class, the model shows very strong performance. All displayed healthy shrimp are correctly classified with confidence values close to 99%. The clean body surface, normal coloration and absence of lesions are clearly learned features. This confirms that healthy shrimp are well separated from diseased classes.

Fig 7: Qualitative prediction examples of the ResNet-152 model on test shrimp images, showing true labels (T), predicted labels (P) and corresponding confidence scores.


       
For the WSSV class, most samples are correctly predicted, though confidence scores are comparatively lower than for Healthy and Yellowhead. One example is misclassified as Black Gill, highlighted in red. This error suggests visual overlap between WSSV symptoms and Black Gill characteristics, such as body darkening and texture changes. Despite this, the majority of WSSV samples are correctly identified. For the Yellow head class, predictions are highly accurate with strong confidence values, mostly above 95%. Even in cases with lower confidence, the predicted class remains correct. This indicates that Yellowhead disease features are distinct and consistently captured by the model.
       
The ablation study results, comparing the performance of different backbone architectures, are summarized in Table 2.

Table 2: Ablation study comparing MobileNetV2, ResNet50 and ResNet152 in terms of performance and computational efficiency.


       
The ablation study indicates that ResNet152 outperformed both ResNet50 and MobileNetV2 in terms of classification performance, demonstrating its ability to capture more complex and discriminative features. While MobileNetV2 provided faster inference and lower computational cost and ResNet50 offered a balanced trade-off between efficiency and accuracy, ResNet152 achieved the most reliable predictions. Therefore, ResNet152 was selected as the final model, as its superior performance outweighs the increased computational complexity for the given application. Using 5-fold cross-validation, the model achieved an overall accuracy of 94.84% and an MCC of 0.9312, indicating strong and consistent classification performance.
     
Recent studies have explored diverse deep learning strategies for shrimp disease detection, including ensemble CNNs, lightweight object detectors, capsule networks and specialized convolutional architectures (Table 3).  However, direct comparison across studies is inherently limited because datasets differ significantly in size, class distribution, imaging conditions and disease definitions. Therefore, performance values should be interpreted as relative indicators rather than strict benchmarks. Büyükarıkan (2025) demonstrated that ensemble learning with heterogeneous CNNs can significantly improve performance, achieving up to 97.3% accuracy using a weighted combination of MobileNet and DenseNet models. While effective, ensemble methods increase computational complexity and deployment cost.  In addition, such models typically have higher inference latency due to multiple forward passes, which may limit real-time applicability in field conditions. Yuhuan et al., (2025) focused on real-time detection using an improved YOLOv8n model, achieving 92.7% mAP@0.5 with reduced parameters, making it suitable for edge deployment, though primarily designed for object detection rather than multi-class image classification. Raj et al., (2025) proposed an Enhanced Recurrent Capsule Network with hybrid metaheuristic optimization, achieving 95.2% accuracy, but at the expense of increased architectural and training complexity. Ramachandran et al. (2023) targeted single-disease diagnosis and reported 97.22% accuracy for WSSV detection using a Dense Inception CNN, demonstrating strong performance but limited generalizability across multiple diseases.

Table 3: Comparison of recent shrimp disease detection studies, including dataset size, model architecture and reported performance.


       
In contrast, the present study employs a single deep residual network, ResNet-152, to perform multi-class classification of shrimp diseases using the TigerShrimpBD dataset. Despite relying on a single model, the proposed approach achieved 95.82% accuracy and a high MCC of 0.9442, indicating strong agreement between predictions and ground truth across all classes. The strong ROC and P-R performance further confirms reliable threshold-independent discrimination. In terms of efficiency, the single-model design also reduces computational overhead compared to ensemble-based approaches. ResNet-152 requires only one forward pass during inference, making it more practical for deployment scenarios where latency and hardware resources are constrained, although it still remains heavier than lightweight architectures such as YOLOv8n variants. These results highlight that deep residual architectures remain highly competitive for shrimp disease diagnosis and offer practical advantages for scalable aquaculture health monitoring systems.
 
Limitations
 
This study has several limitations. A substantial portion of the dataset (approximately 72%) consists of synthetically generated images, which may introduce bias and limit generalizability to real-world farm conditions. The model was evaluated using a single stratified 60:20:20 split without cross-validation, which may affect robustness assessment. Although basic augmentation techniques were applied, more diverse transformations (e.g., illumination and color variations) were not fully explored. In addition, while an ablation study was conducted, the computational trade-offs between model complexity and efficiency require further investigation. Finally, misclassification analysis was limited and deeper interpretability methods could provide more insight into model behavior.
 
Future work
 
Future research will focus on validating the proposed model using independent real-world farm datasets to assess generalizability. More advanced data augmentation techniques, including illumination and color variations, will be explored to improve robustness. Lightweight architectures and model optimization will be investigated to enable deployment on resource-constrained devices. Additionally, enhanced interpretability methods, such as Grad-CAM and saliency analysis, will be further utilized to better understand model decisions and reduce misclassification.
This study proposed a deep learning–based framework for automated multi-class shrimp disease classification using a ResNet-152 architecture. The model was evaluated on the TigerShrimpBD dataset. Through effective preprocessing, data augmentation and transfer learning, the model achieved stable convergence and strong generalization. Experimental results show high classification performance across all four classes. The model achieved an overall accuracy of 95.82% and an MCC of 0.9442, indicating excellent agreement between predictions and ground truth. Class-wise metrics confirm balanced performance, with especially strong results for Healthy and Yellow Head disease. Most misclassifications occurred between visually similar diseases such as Black Gill and WSSV. ROC and precision-recall analyses further validated robust threshold-independent performance. Despite these results, the model relies on image-based data only and uses a computationally heavy architecture. Future work will focus on lightweight models for edge deployment, larger and more diverse datasets and real-time systems integrating environmental or sensor data for improved disease monitoring.
 
Disclaimers
 
The views and conclusions expressed in this article are solely those of the authors and do not necessarily represent the views of their affiliated institutions. The authors are responsible for the accuracy and completeness of the information provided, but do not accept any liability for any direct or indirect losses resulting from the use of this content.
Funding details
 
This research received no external funding.
 
Authors’ contributions
 
All authors contributed toward data analysis, drafting and revising the paper and agreed to be responsible for all the aspects of this work.
 
Data availability
 
The data analysed/generated in the present study will be made available from corresponding authors upon reasonable request.
 
Availability of data and materials
 
Not applicable
 
Use of artificial intelligence
 
Artificial intelligence techniques (deep learning models) were used only for research purposes in this study. No AI tools were used in the writing, editing, or preparation of this manuscript.
 
Declarations
 
Authors declare that all works are original and this manuscript has not been published in any other journal.
The authors declare that they have no conflict of interest.

  1. Ahmed, S.I. and Farid, D.M. (2025). TigerShrimpBD: A tiger shrimp image dataset (Version 1) [Data set]. Mendeley Data. https://doi.org/10.17632/9dj4sk5d55.1.

  2. AlZubi, A.A. (2023). Artificial intelligence and its application in the prediction and diagnosis of animal diseases: A Review. Indian Journal of Animal Research. 57(10): 1265-1271. doi: 10.18805/IJAR.BF-1684.

  3. Asche, F. and erson, J.L., Botta, R., Kumar, G., Abrahamsen, E.B., Nguyen, L.T. and Valderrama, D. (2020). The economics of shrimp disease. Journal of Invertebrate Pathology. 186: 107397. https://doi.org/10.1016/j.jip.2020.107397.

  4. Asmild, M., Hukom, V., Nielsen, R. and Nielsen, M. (2023). Is economies of scale driving the development in shrimp farming from Penaeus monodon to Litopenaeus vannamei? The case of Indonesia. Aquaculture. 579: 740178. https:// doi.org/10.1016/j.aquaculture.2023.740178.

  5. Bhassu, S., Shama, M., Tiruvayipati, S., Soo, T.C.C., Ahmed, N. and Yusoff, K. (2024). Microbes and pathogens associated with shrimps-implications and review of possible control strategies. Frontiers in Marine Science. 11. https://doi.org/ 10.3389/fmars.2024.1397708.

  6. Boyd, C.E., Davis, R.P. and McNevin, A.A. (2021). Perspectives on the mangrove conundrum, land use and benefits of yield intensification in farmed shrimp production: A review. Journal of the World Aquaculture Society. 53(1): 8-46. https://doi.org/10.1111/jwas.12841.

  7. Büyükarıkan, B. (2025). Robust shrimp disease detection using multi-model convolutional neural networks-based ensemble strategies. Aquacultural Engineering. 111: 102616. https:// doi.org/10.1016/j.aquaeng.2025.102616.

  8. Cho, O.H. (2024). Machine learning algorithms for early detection of legume crop disease. Legume Research. 47(3): 463-469. doi: 10.18805/LRF-788.

  9. El-Saadony, M.T., Swelum, A.A., Ghanima, M.M.A., Shukry, M., Omar, A.A., Taha, A.E., Salem, H.M., El-Tahan, A.M., El- Tarabily, K.A. and El-Hack, M.E.A. (2022). Shrimp production, the most important diseases that threaten it and the role of probiotics in confronting these diseases: A review. Research in Veterinary Science. 144: 126-140. https://doi.org/10.1016/j.rvsc.2022.01.009.

  10. FAO (2022). The State of World Fisheries and Aquaculture 2022: Towards Blue Transformation. Food and Agriculture Organization of the United Nations, Rome, Italy. ISBN: 978-92-5-136364-5.

  11. Han, G., Goncharov, A., Eryilmaz, M., Ye, S., Palanisamy, B., Ghosh, R., Lisi, F., Rogers, E., Guzman, D., Yigci, D., Tasoglu, S., Di Carlo, D., Goda, K., McKendry, R.A. and Ozcan, A. (2025). Machine learning in point-of-care testing: innovations, challenges and opportunities. Nature Communications. 16(1): 3165. https://doi.org/10.1038/s41467-025-58527-6.

  12. Kim, S.Y. and AlZubi, A.A. (2024). Blockchain and artificial intelligence for ensuring the authenticity of organic legume products in supply chains. Legume Research. 47(7): 1144-1150. doi: 10.18805/LRF-786.

  13. Lee, S., Park, J.S., Hong, J.H., Woo, H., Lee, C., Yoon, J.H., Lee, K., Chung, S., Yoon, D.S. and Lee, J.H. (2025). Artificial intelligence in bacterial diagnostics and antimicrobial susceptibility testing: Current advances and future prospects. Biosensors and Bioelectronics. 280: 117399. https:// doi.org/10.1016/j.bios.2025.117399.

  14. Maezono, M., Nielsen, R., Buchmann, K. and Nielsen, M. (2025). The current state of knowledge of the economic impact of diseases in global aquaculture. Reviews in Aquaculture. 17(3). https://doi.org/10.1111/raq.70039.

  15. Mohammad, A.A.S., Mohammad, S.I., Vasudevan, A. and Heindric, A. (2026). Intelligent decision support system for crop management using data analytics. Journal of Experimental Biology and Agricultural Sciences. 14(3): 9-14. https:// jebas.org/ojs/index.php/jebas/article/view/3796.

  16. Ozçelik, H., Sağlam, B. and Öcal, A.S. (2026). New natural methods for combating mold fungi in agricultural products: The Isparta example. Bulletin of Pure and Applied Sciences- Botany. 45B(2): 1-17. https://bpasjournals.com/botany/ index.php/journal/article/view/174/150.

  17. Paek, J.J., Kim, J., Bae, S., Paek, K. and Lee, Y. (2026). Chive- fortified fermentation enhances gamma-aminobutyric acid production by Levilactobacillus brevis PL9014. International Journal of Probiotics and Prebiotics. 21(1): 1-6. https://doi.org/10.37290/ijpp.v21i1.12.

  18. Patil, P.K., Geetha, R., Mishra, S.S., Abraham, T.J., Solanki, H.G., Sharma, S.R.K., Pradhan, P.K. et al (2025). Unveiling the economic burden of diseases in aquatic animal food production in India. Frontiers in Sustainable Food Systems. 9. https://doi.org/10.3389/fsufs.2025.1480094.

  19. Raj, A.S., Senthilkumar, S., Radha, R. and Muthaiyan, R. (2025). Enhanced recurrent capsule network with hyrbid optimization model for shrimp disease detection. Scientific Reports. 15(1): 10400. https://doi.org/10.1038/s41598-025-94413-3.

  20. Ramachandran, L., Mohan, V., Senthilkumar, S. and Ganesh, J. (2023). Early detection and identification of white spot syndrome in shrimp using an improved deep convolutional neural network. Journal of Intelligent and Fuzzy Systems. 45(4): 6429-6440. https://doi.org/10.3233/jifs-232687.

  21. Ray, S., Mondal, P., Paul, A.K., Iqbal, S., Atique, U., Islam, M.S., Mahboob, S., Al-Ghanim, K.A., Al-Misned, F. and Begum, S. (2021). Role of shrimp farming in socio-economic elevation and professional satisfaction in coastal communities. Aquaculture Reports. 20: 100708. https://doi.org/10. 1016/j.aqrep.2021.100708.

  22. Roy, S.M., Beg, M.M., Bhagat, S.K., Charan, D., Pareek, C., Moulick, S. and Kim, T. (2025). Application of artificial intelligence in aquaculture-recent developments and prospects. Aquacultural Engineering. 111: 102570. https://doi.org/ 10.1016/j.aquaeng.2025.102570.

  23. Senapin, S., Thaowbut, Y., Gangnonngiw, W., Chuchird, N., Sriurairatana, S. and Flegel, T.W. (2010). Impact of yellow head virus outbreaks in the whiteleg shrimp, Penaeus vannamei (Boone), in Thailand. Journal of Fish Diseases. 33(5): 421-430. https://doi.org/10.1111/j.1365-2761.2009. 01135.x

  24. Singh, J., Kashyap, R., Bansal, K., Das, R., Sand hu, K., Singh, G. and Saini, D.K. (2025). Transition from conventional to AI-based methods for detection of foliar disease symptoms in vegetable crops: A comprehensive review. Journal of Plant Pathology. 107(4): 1791-1814. https://doi.org/10. 1007/s42161-025-01983-2.

  25. Souza, K.F., Gonçalves, L.C.O. and Espindola, F.S. (2026). A broad overview of the multifaceted royal jelly and its applicability in health and performance. Current Topics in Nutraceutical Research. 24(1): 4-14. https://doi.org/10.37290/ctnr.v24i 1.8.

  26. Sriram, P. and Kumari, Y.S. (2026). Game theory in urban birds: A Nash equilibrium analysis using crows and pigeons. Bio- Science Research Bulletin. 42(1): 1-6. https://bpasjournals. com/life-sciences/index.php/journal/article/view/150/140.

  27. Suwoyo, H.S., Sahabuddin, S., Nawang, A., Makmur, M., Sahrijanna, A., Mulyaningrum, S. R.H. and Ilham, I. (2024). Growth performance of selected and non-selected black tiger shrimp (Penaeus monodon) on pond cultivation. BIO Web of Conferences. 136: 01002. https://doi.org/10.1051/ bioconf/202413601002.

  28. Tamut, H., Ghosh, R., Gosh, K. and Siddique, M.A.S. (2025). Enhancin disease detection in the aquaculture sector using convolutional neural networks analysis. Aquaculture Journal. 5(1): 6. https://doi.org/10.3390/aquacj5010006.

  29. Waiho, K., Ling, Y., Ikhwanuddin, M., Shu Chien, A. C., Afiqah Aleng, N., Wang, Y., Hu, M., Liew, H., Kasan, N. A., Peh, J.H. and Fazhan, H. (2024). Current Advances in the Black Tiger Shrimp Penaeus monodon Culture: A Review. Reviews in Aquaculture. 17(1). https://doi.org/10.1111/ raq.12958.

  30. Yuhuan, F., Gengchen, W., Fenghao, L., Ran, Z., Xufei, S. and  Hao, C. (2025). Lightweight shrimp disease detection research based on YOLOv8n.arXiv preprint. https:// doi.org/10.48550/arxiv.2507.02354.
In this Article
Published In
Indian Journal of Animal Research

Editorial Board

View all (0)