
Fig. 1
Illustration of LOF.

Fig. 2
Pipeline for generating adversarial examples.
Table 2
Performance of the classifiers against adversarial attacks before the implementation of the LOF technique.
| Dataset | Model | Accuracy (%) |
|---|---|---|
| AG NEWS | BERT | 21.09 |
| AG NEWS | WordCNN | 13.68 |
| AG NEWS | LSTM | 11.56 |
| MR | BERT | 12.97 |
| MR | WordCNN | 20.59 |
| MR | LSTM | 19.29 |
| Yelp | BERT | 9.98 |
| Yelp | WordCNN | 9.64 |
| Yelp | LSTM | 7.88 |
Table 3
Performance of the classifiers against adversarial attacks after the implementation of the LOF technique.
| Dataset | Model | Accuracy (%) |
|---|---|---|
| AG NEWS | BERT | 85.12 |
| AG NEWS | WordCNN | 72.47 |
| AG NEWS | LSTM | 65.78 |
| MR | BERT | 88.39 |
| MR | WordCNN | 74.83 |
| MR | LSTM | 68.55 |
| Yelp | BERT | 92.59 |
| Yelp | WordCNN | 81.34 |
| Yelp | LSTM | 78.45 |
Table 4
Illustrates the performance of our three classifiers against the Deepwordbug attack technique prior to the implementation of LOF technique.
| Dataset | Model | Accuracy |
|---|---|---|
| AG NEWS | BERT | 21.09 |
| AG NEWS | WordCNN | 13.68 |
| AG NEWS | LSTM | 11.56 |
| MR | BERT | 12.97 |
| MR | WordCNN | 20.59 |
| MR | LSTM | 19.29 |
| Yelp | BERT | 9.98 |
| Yelp | WordCNN | 9.64 |
| Yelp | LSTM | 7.88 |

Fig. 3
ROC curves for BERT under Deepwordbug (DWB) and Textbugger (TB) attacks.