Table 1.
Comparison of Recent EEG-Based Schizophrenia Detection Methods.
| Reference | Method | Representation | Strengths | Limitations |
|---|---|---|---|---|
| [9] | RNN | Raw EEG | Learns temporal dependencies directly from EEG sequences | Limited spatial feature extraction and weak discrimination capability |
| [10] | Multi-domain Connectome CNN | EEG connectivity matrices | Captures functional brain connectivity across multiple domains | High computational cost due to connectivity estimation |
| [11] | Effective Connectivity Analysis | Effective connectivity features | Provides interpretable neurophysiological insights | Complex preprocessing and limited real-time applicability |
| [12] | LSTM | Raw EEG | Captures long-term temporal dependencies effectively | Ignores spatial relationships among EEG channels |
| [13] | Hybrid Deep Neural Network | Raw EEG | Learns complementary deep feature representations | High computational complexity and training cost |
| [14] | CNN–LSTM | Raw EEG | Extracts both spatial and temporal features | Large model size and longer training time |
| [15] | CNN (Scalogram-based) | Scalogram | Effectively captures time-frequency information | Uses only a single representation |
| [16] | Wavelet Feature Engineering | Wavelet coefficients | Good for non-stationary EEG signal representation | Depends on handcrafted features |
| [17] | Deep Residual Network (ResNet) | EEG feature maps | Strong hierarchical feature learning capability | High computational and training cost |
| [20] | Machine Learning with Sensor + Source Features | Sensor & source EEG | Combines complementary EEG information | Requires source localization and preprocessing complexity |
| [21] | Nonlinear Signal Processing | Nonlinear EEG features | Captures nonlinear brain dynamics | Feature-engineering dependent approach |
| [22] | Deep Learning with Connectivity Images | Effective connectivity images | Integrates connectivity with deep learning | Computationally expensive preprocessing pipeline |
| [23] | Wavelet-Based Single-Channel EEG | Single-channel EEG | Low hardware complexity | Limited spatial information |
| [24] | 1D Transformer | Raw EEG | Captures long-range temporal dependencies | Requires large datasets and high computation |
| [25] | LeViT-based Spatial–Temporal Mapping | EEG feature maps | Efficient local + global feature extraction | Does not exploit multimodal EEG representations |

Figure 1.
Schematic diagram of the proposed attention-based fusion framework for schizophrenia detection using EEG spectrograms and scalograms with a Vision Transformer architecture.

Figure 2.
Illustration of the proposed attention-based cross-representation fusion framework.

Figure 3.
Schematic diagram of the ViT.
Table 2.
List of Simulation Parameters.
| Category | Parameter | Value / Description |
|---|---|---|
| EEG Acquisition | Sampling Frequency | 256 Hz |
| Number of Channels | 20 | |
| Epoch Length | 2–5 seconds | |
| Frequency Range | 0.5 – 50 Hz | |
| Dataset Split | 80% Train / 20% Validation | |
| Preprocessing | Band-pass Filter | 0.5 – 50 Hz |
| Normalization | Z-score | |
| Artifact Removal | Optional (ICA / Filtering) | |
| Spectrogram (STFT) | Window Type | Hamming |
| Window Length | 256 samples | |
| Overlap | 50% | |
| FFT Points | 256 | |
| Output Size | 224 × 224 | |
| Scalogram (CWT) | Mother Wavelet | Morlet |
| Scales | 1 – 128 | |
| Frequency Resolution | Adaptive | |
| Output Size | 224 × 224 | |
| Image Processing | Image Type | RGB |
| Image Size | 224 × 224 × 3 | |
| Data Augmentation | Rotation, Flipping, Scaling | |
| Patch Embedding | Patch Size | 16 × 16 |
| Number of Patches (N) | 196 | |
| Embedding Dimension (d) | 768 | |
| Attention Fusion | Fusion Type | Cross-Attention |
| Attention Heads | 8 | |
| Fusion Strategy | Bidirectional (S↔C) | |
| Projection Matrix | Learnable | |
| Vision Transformer | Model Type | ViT-Base |
| Number of Layers | 12 | |
| Hidden Dimension | 768 | |
| MLP Size | 3072 | |
| Dropout | 0.1 | |
| Training | Optimizer | Adam |
| Learning Rate | 1e-4 | |
| Batch Size | 16 / 32 | |
| Epochs | 50 – 100 | |
| Weight Decay | 1e-5 | |
| Classification | Activation | Softmax |
| Loss Function | Categorical Cross-Entropy |

Figure 4.
Time–Frequency Representations of EEG Signals for Normal and Schizophrenia Classes: (a) Spectrogram of the Normal class (b) Scalogram of the Normal class (c) Spectrogram of the Schizophrenia class (d) Scalogram of the Schizophrenia class.

Figure 5.
Training and Validation Loss Curves of the Proposed Model During Learning.

Figure 6.
Training and Validation Accuracy Curves of the Proposed Model During Learning.

Figure 7.
Attention Map Visualization of the Proposed Model for EEG Signal Classification (a) Normal (b) SZ.

Figure 8.
Confusion Matrices of the Proposed Model for Schizophrenia Classification on Dataset 1(a) Training Confusion Matrix (b) Validation Confusion Matrix.
Table 3.
Performance Metrics Comparison Using Random Data Splitting (Dataset: 1).
| Dataset | Precision | Recall | F1-score | Accuracy |
|---|---|---|---|---|
| Training | 0.9928 ± 0.0012 | 0.9911 ± 0.0015 | 0.9920 ± 0.0013 | 0.9920 ± 0.0012 |
| Validation | 0.9928 ± 0.0024 | 0.9786 ± 0.0041 | 0.9857 ± 0.0033 | 0.9857 ± 0.0032 |

Figure 9.
Confusion Matrices of the Proposed Model for Schizophrenia Classification on Dataset 2 (a) Training Confusion Matrix (b) Validation Confusion Matrix.
Table 4.
Performance Metrics Comparison Using Random Data Splitting (Dataset: 2).
| Dataset | Precision | Recall | F1-score | Accuracy |
|---|---|---|---|---|
| Training | 0.9937 ± 0.0011 | 0.9922 ± 0.0014 | 0.9930 ± 0.0012 | 0.9930 ± 0.0011 |
| Validation | 0.9937 ± 0.0023 | 0.9875 ± 0.0038 | 0.9906 ± 0.0030 | 0.9906 ± 0.0028 |

Figure 10.
ROC and AUC curves (a) Dataset 1 (b) Dataset 2.
Table 5.
Subject-Level 5-Fold Cross-Validation Results (Proposed Method: Dataset:1).
| Fold | Accuracy (%) | Precision (%) | Recall (%) | F1-Score (%) |
|---|---|---|---|---|
| Fold 1 | 98.21 | 98.35 | 98.02 | 98.18 |
| Fold 2 | 98.47 | 98.62 | 98.28 | 98.44 |
| Fold 3 | 98.05 | 98.19 | 97.91 | 98.03 |
| Fold 4 | 98.63 | 98.74 | 98.45 | 98.59 |
| Fold 5 | 98.32 | 98.48 | 98.11 | 98.29 |
| Mean ± SD | 98.34 ± 0.21 | 98.48 ± 0.20 | 98.15 ± 0.19 | 98.31 ± 0.20 |
Table 6.
Subject-Level 5-Fold Cross-Validation Results (Proposed Method: Dataset:2).
| Fold | Accuracy (%) | Precision (%) | Recall (%) | F1-Score (%) |
|---|---|---|---|---|
| Fold 1 | 98.58 | 98.69 | 98.41 | 98.55 |
| Fold 2 | 98.76 | 98.88 | 98.60 | 98.74 |
| Fold 3 | 98.49 | 98.61 | 98.33 | 98.47 |
| Fold 4 | 98.91 | 99.02 | 98.76 | 98.89 |
| Fold 5 | 98.67 | 98.80 | 98.52 | 98.65 |
| Mean ± SD | 98.68 ± 0.15 | 98.80 ± 0.16 | 98.52 ± 0.17 | 98.66 ± 0.16 |
Table 7.
Statistical Significance Analysis of the Proposed Method Using Paired t-Test and Wilcoxon Signed-Rank Test.
| Metric | Paired t-test (t) | Paired t-test (p-value) | Wilcoxon W | Wilcoxon p-value | Significant (p < 0.05) |
|---|---|---|---|---|---|
| Accuracy | 7.66 | 0.0016 | 15 | 0.0313 | Yes |
| Precision | 8.00 | 0.0013 | 15 | 0.0313 | Yes |
| Recall | 8.72 | 0.0010 | 15 | 0.0313 | Yes |
| Fl-score | 7.82 | 0.0015 | 15 | 0.0313 | Yes |
Table 8.
Ablation Study of Different Model Variants on Dataset 1 Using Subject-Level 5-Fold Cross-Validation.
| Model Variant | Precision | Recall | F1-score | Accuracy |
|---|---|---|---|---|
| ViT (Spectrogram) | 97.21 ± 0.29 | 96.93 ± 0.31 | 97.07 ± 0.30 | 97.10 ± 0.30 |
| ViT (Scalogram) | 97.74 ± 0.25 | 97.46 ± 0.27 | 97.60 ± 0.26 | 97.63 ± 0.26 |
| Feature Fusion (Without Attention) | 98.13 ± 0.22 | 97.86 ± 0.23 | 97.99 ± 0.22 | 98.02 ± 0.22 |
| Feature Fusion + Attention (Proposed) | 98.48 ± 0.20 | 98.15 ± 0.19 | 98.31 ± 0.20 | 98.34 ± 0.21 |
Table 9.
Ablation Study of Different Model Variants on Dataset 2 Using Subject-Level 5-Fold Cross-Validation.
| Model Variant | Precision | Recall | F1-score | Accuracy |
|---|---|---|---|---|
| ViT (Spectrogram) | 97.56 ± 0.24 | 97.28 ± 0.25 | 97.42 ± 0.24 | 97.45 ± 0.24 |
| ViT (Scalogram) | 98.04 ± 0.20 | 97.81 ± 0.21 | 97.92 ± 0.20 | 97.96 ± 0.21 |
| Feature Fusion (Without Attention) | 98.43 ± 0.18 | 98.16 ± 0.19 | 98.29 ± 0.18 | 98.31 ± 0.18 |
| Feature Fusion + Attention (Proposed) | 98.80 ± 0.16 | 98.52 ± 0.17 | 98.66 ± 0.16 | 98.68 ± 0.15 |
Table 10.
Comparison with the state-of-the-art methods (K-fold Cross Validation).
| Category | Study | Model Description | Dataset Composition | Accuracy (%) |
|---|---|---|---|---|
| Deep CNN | Bagherzadeh et al. [22] | DenseNet121 | 14 SZ – 14 Healthy | 96.26 |
| Classical ML | Sharma et al. [23] | KNN | 14 SZ – 14 Healthy | 97.20 |
| Transformer | Shoeibi et.al [24] | Transformer | 26 SZ – 30 Healthy | 97.62 |
| LeViT | Beilin et.al [25] | Transformer | - | 85.04 |
| Proposed Variants | This Work | ViT (Spectrogram) | 14 SZ – 14 Healthy | 96.80 |
| Proposed Variants | This Work | ViT (Scalogram) | 14 SZ – 14 Healthy | 97.27 |
| Proposed Variants | This Work | Fusion (Without Attention) | 14 SZ – 14 Healthy | 97.80 |
| Proposed Method | This Work | Fusion with Attention Mechanism | 14 SZ – 14 Healthy | 98.13 |