The LeafScan classification model is a lightweight Convolutional Neural Network (CNN) designed to classify plant leaf images into 38 plant disease and healthy-leaf categories. The model takes an RGB leaf image resized to 128 × 128 pixels and progressively transforms it from low-level visual information such as edges and textures into higher-level disease-specific features.
The model begins with two standard convolutional layers containing 32 filters, followed by batch normalization and ReLU activation. A max-pooling layer reduces the spatial resolution from 128 × 128 to 64 × 64 while retaining prominent features. Spatial dropout is applied to improve generalization.
The second block replaces standard convolutions with depthwise-separable convolutions and increases the number of feature channels to 64. This allows the network to learn richer representations while keeping the model computationally lightweight. Another pooling operation reduces the representation to 32 × 32.
The third block further increases the feature depth to 128 channels using separable convolutions. At this stage, the network has moved from detecting simple visual patterns toward learning more complex structures associated with different plant diseases
Instead of flattening the final feature maps into a very large vector, the model uses Global Average Pooling. The resulting 32 × 32 × 128 feature representation is reduced to a compact vector of 128 values. This significantly reduces the number of parameters in the classification stage.
The feature vector is passed through a 128-unit fully connected layer with ReLU activation, followed by 50% dropout for regularization. The final 38-unit softmax layer produces a probability distribution across all supported plant disease and healthy-leaf classes.
The architecture is intentionally lightweight, containing approximately 67,046 trainable parameters. It combines several techniques to improve both efficiency and generalization:
Data augmentation: random horizontal flips, rotations, and zooms expose the model to variations of training images.
Batch normalization: stabilizes activations during training.
Separable convolutions: reduce computational and parameter costs compared with conventional convolutions.
Max pooling: progressively reduces spatial dimensions while retaining salient features.
Spatial dropout and dropout: reduce over-reliance on particular features and help mitigate overfitting.
Global average pooling: provides a compact representation before classification.
Early stopping: monitors validation accuracy and restores the best-performing weights.
The model was trained using the Adam optimizer with a learning rate of 0.001 and sparse categorical cross-entropy as the loss function.