Figure 1.

Figure 2.

Figure 3.

Figure 4.

Figure 5.

LSTM Performance Stability Across 5 Independent Runs
| Run | Recall (%) | F1-Score | AUC-ROC |
|---|---|---|---|
| Run 1 | 97.0 | 0.965 | 0.987 |
| Run 2 | 96.5 | 0.961 | 0.984 |
| Run 3 | 97.0 | 0.963 | 0.986 |
| Run 4 | 96.0 | 0.960 | 0.983 |
| Run 5 | 97.5 | 0.966 | 0.985 |
| Mean ± SD | 96.8 ± 0.4 | 0.963 ± 0.003 | 0.985 ± 0.002 |
Key Confusion Matrix Terms for the Three Evaluated Models
| Metric | LSTM | Isolation Forest | One-Class SVM |
|---|---|---|---|
| TP (True Positive) | 97 | 93 | 90 |
| FN (False Negative) | 3 | 7 | 10 |
| TN (True Negative) | 396 | 388 | 384 |
| FP (False Positive) | 4 | 12 | 16 |
Simulation parameters
| Parameter | Value |
|---|---|
| System | IEEE 118-bus |
| Library | Pandapower v2.13 |
| Total sample | 2000 |
| Normal / FDIA | 1600 / 400 (80% / 20%) |
| Test set | 500 (400 normal + 100 FDIA) |
| Feature size | 12 features per bus |
| Training / Test split | 75% / 25% |
| Target bus | Bus 69 |
Recall Comparison by Attack Type
| Attack Scenario | LSTM | Isolation Forest | One-Class SVM |
|---|---|---|---|
| Scenario A (Gradual Drift) | 96.8% | 89.4% | 83.2% |
| Scenario B (Sudden Injection) | 98.9% | 97.1% | 94.6% |
| Scenario C (Coordinated Corruption) | 97.4% | 91.3% | 87.0% |
| Scenario D (Low Magnitude) | 95.6% | 85.7% | 81.9% |
5-Fold Walk-Forward Validation Results
| Fold | LSTM Recall (%) | IF Recall (%) | OC-SVM Recall (%) |
|---|---|---|---|
| Fold 1 | 96.0 | 92.0 | 89.0 |
| Fold 2 | 97.0 | 93.0 | 90.0 |
| Fold 3 | 96.0 | 92.5 | 89.5 |
| Fold 4 | 97.0 | 93.5 | 91.0 |
| Fold 5 | 96.5 | 93.0 | 90.5 |
| Mean ± SD | 96.5 ± 0.6 | 92.8 ± 0.6 | 90.0 ± 0.7 |
Comparison of FDIA detection performances for a test set of 500 samples
| Metric | LSTM | Isolation Forest | One-Class SVM |
|---|---|---|---|
| Accuracy | 98.60% | 96.20% | 94.80% |
| [97.14%–99.32%] | [94.14%–97.55%] | [92.49%–96.43%] | |
| Recall | 97.00% | 93.00% | 90.00% |
| [91.55%–98.97%] | [86.25%–96.57%] | [82.56%–94.48%] | |
| Precision | 96.04% | 88.57% | 84.91% |
| [90.26%–98.45%] | [81.08%–93.34%] | [76.88%–90.49%] | |
| F1-Score | 0.965 | 0.907 | 0.874 |
| [0.949–0.981] | [0.882–0.933] | [0.845–0.903] | |
| AUC-ROC | 0.987 | 0.965 | 0.942 |
| [0.971–1.000] | [0.940–0.990] | [0.910–0.974] | |
| False Positive Rate | 1.0% | 3.0% | 4.0% |
| Training Time (s) | 42.3 | 0.8 | 4.2 |
| 95% Recall Benchmark | ✓ PASSES | ✗ FAILS | ✗ FAILS |
McNemar Test Results
| Comparison | b | c | χ2 | p | Significant? |
|---|---|---|---|---|---|
| LSTM vs Isolation Forest | 12 | 0 | 10.083 | 0.0015 | Yes** |
| LSTM vs One-Class SVM | 19 | 0 | 17.053 | <0.001 | Yes*** |
| Isolation Forest vs One-Class SVM | 7 | 0 | 5.143 | 0.0233 | Yes* |