Fig 1 shows the whole framework of the weed detection system design. The approach used in designing the system comprises five main steps, namely; data collection, image preprocessing, data augmentation, feature extraction and classification through CCNN and model performance evaluation. First, images of cassava fields are collected under different conditions and are divided into four categories, which include: cassava plants, broadleaf weeds, grassy weeds and Sedges. Image preprocessing is done through resizing and normalizing the images obtained, after which data augmentation processes are done to increase model generalization through techniques such as rotation, flipping and scaling. The final stage involves training a CCNN model using the images that have undergone data augmentation, after which model evaluation is done using several performance metrics.
Data collection and preprocessing
A total of 2,500 images in RGB format were obtained from cassava plant cultivation fields from Sundarapuram, Attur, Salem, Tamil Nadu, India and SRM College of Agricultural Sciences, Baburayanpettai-Chengalpattu. These images were acquired through a digital high-definition camera in different environments based on illumination intensity, weed density, stage of growth of crops and various angles of view. In this way, we obtained a dataset consisting of four semantic classes, namely, cassava plants, broadleaf weeds, grassy weeds and soil. Pixel-wise annotations of these images were created using the LabelMe annotation tool. In order to have effective model training, we have divided the data into training, and testing sets, each consisting of 2,000 and 500 images, respectively. Fig 2 shows sample images with diverse lighting conditions.
This dataset contains images of cassava crops and various weed species. After collection, the data undergoes a preparation stage aimed at removing poor-quality samples and ensuring that only clean, high-quality data is retained. This process is essential for selecting and training the most effective model. Fig 3 shows data pre-processing stages.
Data augmentation
Images are resized, normalized and augmented like rotation, scaling and flipping to improve the model’s robustness. To overcome over-fitting problems and improve model generalization, extensive online data augmentation was applied during training. This included random rotations (±30°), horizontal and vertical flips, brightness and contrast adjustments (±20%) and gaussian blurring. Table 1 shows data augmentation techniques.
Model architecture
The convolutional neural networks are deep learning architectures that are specifically designed to operate directly on grid-like or image data
(Hasan et al., 2021). Convolutional layers apply learnable filters to incoming data, constructing spatial hierarchies based on patterns such as edges, shapes and textures. Each convolution layer has an associated activation function, typically ReLU, which adds non-linearity to the model. Pooling layers, like max pooling, reduce the spatial dimensions of data to control the over-fitting and complexity. CNNs typically end by using fully connected layers that combine the learned features for regression or classification
(Chepuri et al., 2025). A major strong point of CNNs is that their shared weights and local connection enable them to be computationally efficient as well as being applicable to high-dimensional data, such as 2D image plane. Fig 4 shows the general layout of Convolution Neural Network and Fig 5 presents the architecture of novel customized CNN employed for cassava weed classification.
Convolutional operation
During the convolution operation, the filter slides over the input image. A dot product is then taken at each position between the filter and a region of the input it is covering. The resultant value is saved into the output feature map at that (x, y) position. This carries on throughout the whole image. For instance, if the size of filter is 33 and that of input image is 256*256, then it slides over as a result producing an output feature map of 256*256. The output of the convolution operation is a feature map, which is essentially a transformed version of the original output, highlighting the features that the filter was designed to detect such as edges or textures. The stride of ta convolution is the number of row and column steps that are taken when you shift the kernel. Increasing stride speeds up processing but may cause the network to miss small features. Padding adds extra pixels around each input edge, typically zeros, to maintain output size after convolution. Without padding, feature maps shrink after each layer, potentially losing important edge information. Starting with small filters, stride one, padding helps preserve details in early layers. Fig 6 shows a) Convolutional layer, b) Max pooling layer.
The CNN’s convolution operation is described as follows:
Yi,j = (X * W)i,j = ∑m ∑n Xi + m, j + n Wm,n ...(1)
Where,
X= Input image.
Y (i,j)= Output activation at position (i, j).
W= Filter weights (kernel).
i,j= Indices over the output dimensions.
m, n= Indices over the filter dimensions.
*= Convolution operation.
The activation map calculations and output dimension calculations are given below:
Where,
X= Size of the input image.
F= Size of the filter (height or width).
P= Padding applied to the input.
S= Stride of the convolution.
Pooling operation
After initializing the CNN, a pooling operation is applied to reduce the spatial dimensions of the feature maps. Pooling down-samples the input, thereby decreasing the computational load and reducing the number of learnable parameters in the network. A popular activation function in the DL algorithm adds non-linearity to the model is the rectifier linear unit (ReLU).
ReLU (x) = max (0, x) ...(3)
Mathematically, ReLU returns the input itself if then is positive; otherwise, it returns zero. ReLU is efficient because it only needs a threshold operation and it makes the network sparser by only activating some of the neurons.
Dropout layer
The dropout layer randomly shuts off a fraction of neurons from the previous layer, cutting down the number of active units. This forces the network to learn more robust features instead of relying too much on any single neuron. It is used to deep CNNs, especially when the architecture gets complicated.
Flatten layer
This layer flattens the multi-dimensional data from convolutional blocks to one-dimensional vector, preparing it for fully connected layers. This one-dimensional vector is suitable as input for fully connected layers, which typically work with flat data for classification or regression tasks.
Fully connected layers
The actual classification or prediction happens in these layers. The normal fully connected layers take input from the flatten layer and carry out the classification or prediction task. The output layer has the appropriate function like SoftMax, based on the task, which is usually multi class classification (
Pakruddin et al. 2025). We combine all the parts sequentially to build a complete CNN network. Multiple convolution blocks, which will extract features, then a set of fully connected (Dense) Layers, which will use those features to make classification.
Algorithm 1: Cassava weed classification using proposed Customized CNN.
Require:
Epochs = 50.
Batch _ size = 16.
Input _ shape = (256, 256, 3).
Classes = 4.
Activation = ReLu, Softmax.
Optimizer = Adam.
Function: Cassava_Weed_Detection (epochs, batch_size, input_shape, classes, activation, optimizer).
#1: Data preparation
1. Load cassava weed image dataset.
2. Resize images to 256 × 256 × 3.
3. Normalize pixel values.
4. Split datasets into training and validation sets.
train_data, val_data←load_and_preprocess_data (input_ shape).
#2: Model definition
1. nitialize customized CNN architecture.
2. Add convolutional, pooling and dense layers.
3. Apply ReLU activation in hidden layers.
4. Apply Softmax activation in output layer for multi-class classification.
model¬define_CNN_model (input_shape, classes, activation).
#3: Model training
1. Compile model using Adam optimizer.
2. Train model with training dataset for 50 epochs using batch size 16.
3. Validate model using validation dataset during training.
trained_model¬ train_model (model, train_data, val_data, epochs, batch_size, optimizer).
#4: Model evaluation
1. Evaluate trained model on validation dataset.
2. Compute classification accuracy, loss, precision, recall and F1-score.
evaluate_model (trained_model, val_data).
#5: Result interpretation
1. Predict weed categories from test images.
2. Analyze performance metrics and classification results.
3. Identify correctly and incorrectly classified weed samples.
End function.
Experimental setup
The experimental Parameter setup is shown in Table 2. The weight decay is 0.005. The batch size is set to 16. The initial learning rate is 0.001, which maximum training is fixed at the epoch 50.
Hyperparameter for customized CNN model
The hyperparameter values of the customized CNN Model are shown in Table 3.
The process of hyperparameter optimization was carried out by conducting experiments. The optimal combination of hyperparameters was found when the learning rate was set at 0.001, batch size was equal to 16, dropout rate was 0.5 and the value of the weight decay coefficient was 0.005. For the model training, the Adam optimizer was chosen owing to fast convergence and adaptive learning.