(a) Original image, (b) corresponding ground truth (GT) mask, and (c) predicted segmentation mask generated by the proposed MFPN-ASPP framework.
The proposed MFPN–ASPP further augments MobileNetV2-FPN by incorporating an Atrous Spatial Pyramid Pooling head, significantly enriching contextual understanding and multiscale representation.
Table 4 Per-image segmentation performance of the proposed MFPN–ASPP framework on the test set.
These findings demonstrate that ASPP effectively captures multi-scale contextual information and compensates for the limitations of the lightweight backbone, leading to superior photovoltaic defect segmentation performance.
The proposed MFPN-ASPP model achieved a BF Score of 0.8436 for 2-pixel tolerances, indicating accurate boundary delineation and strong agreement with the ground-truth annotations.
Qualitative results
Figure 6 presents representative samples from the photovoltaic infrared test set, showing the input image, the ground truth (GT) defect mask, and the predictions of U-Net, FPN, and the proposed MFPN-ASPP. As illustrated, the proposed model efficiently delineates hot spot defects and closely matches the GT annotations. The MFPN-ASPP method yields cleaner, better-localized masks than both U-Net and FPN and is consistently the closest to the GT. Figure 7 contrasts GT defect masks with predictions from the MFPN-ASPP model qualitatively; these observations are consistent with the strong performance of Dice and IoU observed in our evaluation, indicating a close agreement between the predicted masks and GT.
Fig. 6 Full size image Qualitative comparison of segmentation results for representative photovoltaic thermal test samples (IDs 15, 34, 42, 57, 63, 195, and 275). (a) Original thermal image, (b) ground truth (GT) defect mask, (c) U-Net prediction, (d) FPN prediction, and (e) prediction generated by the proposed MFPN-ASPP framework.
Fig. 7 Full size image Representative segmentation results of the proposed MFPN-ASPP framework on photovoltaic infrared (IR) images for test samples with IDs 20, 33, 45, and 50. (a) Original image, (b) corresponding ground truth (GT) mask, and (c) predicted segmentation mask generated by the proposed MFPN-ASPP framework. The corresponding Dice and IoU scores are reported for each sample.
Quantitative results
In this study, several deep learning segmentation architectures were explored and systematically evaluated. This stepwise approach allowed for a clear understanding of how architectural complexity influences segmentation accuracy and generalization performance. Thirteen architectures were implemented and trained under identical experimental conditions. All benchmark models were trained using the same dataset split, optimizer, learning rate, batch size, number of epochs, loss function, and evaluation protocol to ensure a fair comparison. These models were categorized into three groups according to their architectural design: (1) classical convolutional and U-Net–based networks, (2) YOLO–based segmentation frameworks, and (3) multiscale feature pyramid and hybrid models. Table 3 summarizes their quantitative results in terms of Dice, IoU, Precision, and Recall. The following discussion analyzes the quantitative behavior of each group to highlight the progression of performance improvements leading to the proposed MFPN–ASPP model.
The first group of models includes the baseline CNN and several encoder–decoder variants derived from the U-Net family. The simple CNN employs convolutional and upsampling layers to reconstruct spatial features, but exhibits weak generalization on complex photovoltaic textures, reflected by its low Dice (0.602) and IoU (0.431) values. U-Net improves structural preservation through skip connections, resulting in higher segmentation accuracy (Dice = 0.6624, IoU = 0.4952), but still lacks semantic depth to detect small or low-contrast defects. Incorporating pre-trained encoders, such as ResNet18 and ResNet34, enhances feature representation and robustness; however, the latter’s increased complexity slightly reduces performance (Dice = 0.6085, IoU = 0.4373) due to potential overfitting. The EfficientNet-B3 encoder balances depth and width scaling, achieving better Dice (0.7112) and IoU (0.5518) scores, but its computational cost limits real-time UAV deployment. Replacement of the encoder with MobileNetV2 further reduces the parameters while maintaining competitive accuracy (Dice = 0.7319, IoU = 0.5787). Finally, ResUNet introduces residual connections to stabilize training and reuse features, but despite moderate success (Dice = 0.6449, IoU = 0.4759), it still struggles to capture fine-grained contextual variations in PV thermal imagery. The second group comprises detection-oriented architectures adapted for segmentation. YOLOv8 integrates object detection and mask prediction into a single framework, achieving high precision (0.7898) but lower recall (0.5793), which limits boundary fidelity in defect segmentation. To refine mask quality, YOLOv8–SAM combines YOLOv8 with the Segment Anything Model, achieving improved Dice (0.6798) and IoU (0.5691) scores; however, the two-stage design increases computational overhead. The YOLO11–SAM hybrid enhances localization accuracy for small and irregular faults, producing the highest recall among this group (0.7617). However, the added complexity and memory requirements hinder its practicality for real-time UAV-based inspection. The final group involves multiscale feature fusion networks that leverage contextual learning. The Feature Pyramid Network (FPN) effectively integrates multi-level representations, achieving strong Dice (0.7493) and IoU (0.6958) scores with balanced precision (0.7616) and recall (0.7393). When the lightweight MobileNetV2 encoder was integrated into the FPN framework (MobileNetV2-FPN), the overall accuracy decreased slightly (Dice = 0.6872, IoU = 0.5234). Although this configuration reduces computational cost and model size, the smaller encoder limits the network’s receptive field and its ability to capture deep contextual features, resulting in marginally lower segmentation performance compared to the original FPN. The proposed MFPN–ASPP further augments MobileNetV2-FPN by incorporating an Atrous Spatial Pyramid Pooling head, significantly enriching contextual understanding and multiscale representation. This hybrid architecture achieves the best overall performance (Dice = 0.8440, IoU = 0.7564, Precision = 0.8771, Recall = 0.8428), maintaining computational efficiency suitable for UAV-based photovoltaic inspection.
Table 3 Comparative analysis for segmentation results (Dice, IoU, Precision, and Recall) using the Photovoltaic Thermal Images Dataset. Full size table
To further evaluate the consistency of the proposed model in individual test samples, the segmentation metrics per image were analyzed, as summarized in Table 4. The proposed MFPN–ASPP framework achieved a mean Dice score of 0.844 ± 0.165, a mean IoU of 0.757 ± 0.187, a mean Precision of 0.877 ± 0.187, and a mean Recall of 0.843 ± 0.155 across 101 test images. The corresponding missed detection rate (MDR), calculated as \(1-\textrm{Recall}\), was 0.157. The observed variation reflects differences in defect size, shape, and thermal contrast among samples, while the high mean values indicate robust segmentation performance across diverse operating conditions. As demonstrated in Table 5, the proposed MFPN–ASPP framework achieved the highest overlap performance among the methods evaluated on the same Photovoltaic Thermal Images Dataset reported by Pierdicca et al. 3. The proposed method achieved a Dice score of 0.8440 and an IoU of 0.7564, outperforming U-Net (0.841/0.741), LinkNet (0.825/0.748), FPN (0.825/0.734), and Mask-RCNN (0.605/0.499). These results indicate that the integration of MobileNetV2, FPN, and ASPP provides more accurate hotspot localization and segmentation than existing approaches evaluated on the same benchmark dataset.
Table 4 Per-image segmentation performance of the proposed MFPN–ASPP framework on the test set. Full size table
Table 5 Comparison between the proposed method and the recent segmentation methods that use the same Photovoltaic Thermal Images Dataset3. Full size table
Table 6 Statistical stability analysis of the proposed MFPN-ASPP model, five times. Full size table
In addition, the False Alarm Rate (FAR), computed from the global (micro) confusion matrix aggregated over the entire test set, was 0.0422%, indicating a very low rate of background pixels incorrectly classified as defects.
To compute the stability of the proposed model, the training process was repeated five times under the same experimental settings. The results of the Dice and IoU are reported in Table 6. The low variation among the runs indicates consistent and reliable model performance.
Table 7 presents the ablation analysis of the proposed MFPN-ASPP architecture. The baseline FPN achieved a Dice score of 0.7493 and an IoU of 0.6958. Replacing the original backbone with MobileNetV2 reduced the segmentation performance due to the lightweight nature of the encoder and its limited ability to preserve fine spatial details. However, incorporating the ASPP module significantly improved the segmentation results, increasing the Dice score from 0.6872 to 0.8440 and the IoU from 0.5234 to 0.7564. Replacing the conventional FPN with BiFPN results in a slight improvement in the Dice coefficient, indicating that the bidirectional feature fusion of BiFPN enhances feature representation. These findings demonstrate that ASPP effectively captures multi-scale contextual information and compensates for the limitations of the lightweight backbone, leading to superior photovoltaic defect segmentation performance.
Table 7 Ablation study of the proposed MFPN-ASPP architecture using the Photovoltaic Thermal Images Dataset. Full size table
Table 8 Effect of ASPP dilation rates on segmentation performance. Full size table
Table 9 Computational complexity and boundary localization analysis of the proposed MFPN-ASPP model. Full size table
To further investigate the effect of ASPP design choices, an additional ablation study was conducted using different dilation-rate configurations. As reported in Table 8, To evaluate the computational cost of the proposed MFPN–ASPP model, the number of trainable parameters, FLOPs, inference time, and GPU memory consumption were measured. As reported in Table 9, the model contains 4.54 million parameters and requires 18.32 GFLOPs for a 512 × 640 input image. The average inference time was 15.26 ms per image (65.52 FPS) on a GPU platform, while the peak GPU memory consumption during inference was 147.63 MB. To further evaluate boundary localization performance, the Boundary F-score (BF Score) was computed using tolerance-based matching. The proposed MFPN-ASPP model achieved a BF Score of 0.8436 for 2-pixel tolerances, indicating accurate boundary delineation and strong agreement with the ground-truth annotations. These results demonstrate that the proposed framework achieves a favorable balance between segmentation performance and computational efficiency.