
Figure 1:
Framework of the proposed LWA-MoDUNet.

Figure 2:
Architecture of the encoder block of the proposed LWA-ModUNet.

Figure 3:
Architecture of the inception block of the proposed LWA-ModUNet [54].

Figure 4:
Decoder architecture of the proposed LWA-MoDUnet.

Figure 5:
Diagram of additive attention gate [55].

Figure 6:
Flowchart of the proposed LWA-MoDUNet depicting the encoding and decoding phases with inception and attention modules.

Figure 7:
Sample images of considered datasets (Row 1) along with their corresponding GT (Row 2). GT, ground truth; IDD-Lite, Indian driving dataset-lite; MADS, martial arts, dancing and sports.
Table 1:
Ablation study of LWA-MoDUNet on autorickshaw dataset
| Measure | LWA-MoDUNet | W/o inception module | W/o attention |
|---|---|---|---|
| IoU score | 0.8768 | 0.6729 | 0.8173 |
| Error | 0.0663 | 0.1452 | 0.0927 |
| Accuracy | 93.3691 | 78.7966 | 86.4217 |
| F-score | 0.9344 | 0.7913 | 0.8786 |

Figure 8:
Generated output masks by UNet, E-Net, SegNet, UNet with ResNet-18 encoder, UNet with ResNet-34 encoder, and LWA-ModUNet models on exemplary images of autorickshaw dataset. GT, ground truth.

Figure 9:
Generated output masks by UNet, E-Net, SegNet, UNet with ResNet-18 encoder, UNet with ResNet-34 encoder, and LWA-ModUNet on exemplary images of IDD Lite dataset. GT, ground truth; IDD-Lite, Indian driving dataset-lite.

Figure 10:
Generated output masks by UNet, E-Net, SegNet, UNet with ResNet-18 encoder, UNet with ResNet-34 encoder, and LWA-ModUNet models on exemplary images of PETS dataset. GT, ground truth.

Figure 11:
Generated output masks by UNet, E-Net, SegNet, UNet with ResNet-18 encoder, UNet with ResNet-34 encoder, and LWA-ModUNet models on exemplary images of MADS dataset. GT, ground truth; MADS, martial arts, dancing and sports.

Figure 12:
Generated output masks by UNet, E-Net, SegNet, UNet with ResNet-18 encoder, UNet with ResNet-34 encoder, and LWA-ModUNet models on exemplary images of CT-Liver dataset. GT, ground truth.
Table 2:
Comparison of key performance indicators of LWA-MoDUNet and SOTA models on autorickshaw dataset
| Measures | UNet | UNet-ResNet34 | UNet-ResNet18 | E-Net | SegNet | LWA-MoDUNet |
|---|---|---|---|---|---|---|
| IoU score | 0.7665 | 0.8459 | 0.8356 | 0.8480 | 0.7592 | 0.8768 |
| Err | 0.1326 | 0.0849 | 0.0912 | 0.0857 | 0.1403 | 0.0663 |
| Acc | 86.7412 | 91.5061 | 90.8837 | 91.4349 | 85.9677 | 93.3691 |
| Sp | 0.8697 | 0.9303 | 0.9241 | 0.9514 | 0.8789 | 0.9429 |
| Ss | 0.8651 | 0.9009 | 0.8946 | 0.8829 | 0.8423 | 0.9248 |
| F-score | 0.8678 | 0.9165 | 0.9104 | 0.9177 | 0.8631 | 0.9344 |
| Cc | 0.7348 | 0.8306 | 0.8182 | 0.8315 | 0.7203 | 0.8676 |
Table 3:
Comparison of key performance indicators of LWA-MoDUNet and compared models on IDD-Lite dataset
| Measures | UNet | UNet-ResNet34 | UNet-ResNet18 | E-Net | SegNet | LWA-MoDUNet |
|---|---|---|---|---|---|---|
| IoU score | 0.6031 | 0.6174 | 0.5981 | 0.566 | 0.3076 | 0.9283 |
| Acc | 0.9203 | 0.9398 | 0.9356 | 0.9321 | 0.8971 | 0.9616 |
| Err | 0.0716 | 0.0601 | 0.0643 | 0.0679 | 0.1028 | 0.00435 |
| Sp | 0.9500 | 0.9484 | 0.9469 | 0.9395 | 0.8975 | 0.9328 |
| Ss | 0.8534 | 0.8617 | 0.8472 | 0.8669 | 0.8896 | 0.9436 |
| F-score | 0.7056 | 0.7635 | 0.7485 | 0.7229 | 0.4705 | 0.8909 |
| Cc | 0.6794 | 0.7371 | 0.7186 | 0.6979 | 0.4965 | 0.8153 |
Table 4:
Comparison of key performance indicators of LWA-MoDUNet and compared models on PETS dataset
| Measures | UNet | UNet-ResNet34 | UNet-ResNet18 | E-Net | SegNet | LWA-MoDUNet |
|---|---|---|---|---|---|---|
| IoU score | 0.7665 | 0.8459 | 0.8356 | 0.8480 | 0.7592 | 0.8768 |
| Err | 0.1326 | 0.0849 | 0.0912 | 0.0857 | 0.1403 | 0.0663 |
| Acc | 86.7412 | 91.5061 | 90.8837 | 91.4349 | 85.9677 | 93.3691 |
| Sp | 0.8697 | 0.9303 | 0.9241 | 0.9514 | 0.8789 | 0.9429 |
| Ss | 0.8651 | 0.9009 | 0.8946 | 0.8829 | 0.8423 | 0.9248 |
| F-score | 0.8678 | 0.9165 | 0.9104 | 0.9177 | 0.8631 | 0.9344 |
| Cc | 0.7348 | 0.8306 | 0.8182 | 0.8315 | 0.7203 | 0.8676 |
Table 5:
Comparison of key performance indicators of LWA-MoDUNet and compared models on MADS dataset
| Measures | UNet | UNet-ResNet34 | UNet-ResNet18 | E-Net | SegNet | LWA-MoDUNet |
|---|---|---|---|---|---|---|
| IoU score | 0.8968 | 0.9448 | 0.9084 | 0.9609 | 0.9721 | 0.9768 |
| Err | 0.0545 | 0.0290 | 0.0496 | 0.0201 | 0.0142 | 0.0118 |
| Acc | 94.5520 | 97.1042 | 95.0428 | 97.9919 | 98.5781 | 98.8225 |
| Sp | 0.9464 | 0.9905 | 0.9818 | 0.9872 | 0.9916 | 0.9934 |
| Ss | 0.9447 | 0.9531 | 0.9229 | 0.9728 | 0.9801 | 0.9832 |
| F-score | 0.9456 | 0.9716 | 0.9520 | 0.9801 | 0.9859 | 0.9883 |
| Cc | 0.8910 | 0.9428 | 0.9028 | 0.9599 | 0.9716 | 0.9765 |
Table 6:
Comparison of key performance indicators of LWA-MoDUNet and compared models on CT-Liver dataset
| Measures | UNet | UNet-ResNet34 | UNet-ResNet18 | E-Net | SegNet | LWA-MoDUNet |
|---|---|---|---|---|---|---|
| IoU score | 0.8839 | 0.9445 | 0.9565 | 0.9349 | 0.9495 | 0.9650 |
| Err | 0.0616 | 0.0287 | 0.0223 | 0.0338 | 0.0262 | 0.0178 |
| Acc | 93.8400 | 97.1308 | 97.7686 | 96.6167 | 97.3808 | 98.2162 |
| Sp | 0.9384 | 0.9758 | 0.9811 | 0.9719 | 0.9845 | 0.9855 |
| Ss | 0.9384 | 0.9669 | 0.9743 | 0.9605 | 0.9635 | 0.9788 |
| F-score | 0.9384 | 0.9714 | 0.9778 | 0.9664 | 0.9741 | 0.9822 |
| Cc | 0.8768 | 0.9427 | 0.9554 | 0.9324 | 0.9478 | 0.9643 |

Figure 13:
Comparative analysis of IoU score on training and validation images taken from autorickshaw dataset. IoU, intersection over union.

Figure 14:
Comparative analysis of IoU score on training and validation images taken from IDD-Lite dataset. IDD-Lite, Indian driving dataset-lite; IoU, intersection over union.

Figure 15:
Comparative analysis of IoU score on training and validation images taken from PETS dataset. IoU, intersection over union.

Figure 16:
Comparative analysis of IoU score on training and validation images taken from MADS dataset. IoU, intersection over union; MADS, martial arts, dancing and sports.

Figure 17:
Comparative analysis of IoU score on training and validation images taken from CT-Liver dataset. IoU, intersection over union.
Table 7:
Statistical performance analysis of LWA-MoDUNeT against other compared SOTA models
| Models | Samples size = 30 p-value | Samples size = 50 p-value |
|---|---|---|
| UNet and LWA-MoDUNet | 0.0029 × E−05 | 3.54 × E−09 |
| ResNet34 and LWA-ModUNet | 8.49 × E−04 | 7.28 × E−06 |
| ENet and LWA-MoDUNet | 3.04 × E−02 | 9.51 × E−03 |
| SegNet and LWA-ModUNet | 4.85 × E−02 | 6.39 × E−04 |

Figure 18:
Model type 1 and type 2 errors comparison across datasets. IDD-Lite, Indian driving dataset-lite; MADS, martial arts, dancing and sports.