
Fig. 1.
Diagram illustrating the role of the phagobiome in the host's gut. ARG – antibiotic resistance gene

Fig. 2.
The trend in the number of publications over the years 1947–2025
Table 1.
Number of publications in selected databases
| Database | Number of publications |
|---|---|
| Web of Science | 240 |
| Scopus | 88 |
| PubMed | 69 |
| Total before deduplication | 397 |
| Total after deduplication | 244 |

Fig. 3.
Network map showing co-occurrence of terms in swine gut phageome publications. FFT – faecal filtrate transplantation – the extracellular bacteriophage particles constituting the active component of the phageome; ARGs– antibiotic resistance genes; PWD– post-weaning diarrhoea in piglets; ETEC– enterotoxin-producing strains of Escherichia coli; NEC– necrotising enterocolitis

Fig. 4.
Key effects of phages on the pig host

Fig. 5.
Selected tools used in phagobiome analyses with assigned analysis stages
Table 2.
Summary of limitations and development direction for virome analysis
| Analysis stage | Limitation | Development directions |
|---|---|---|
| Viral DNA identification | Limited representation of poorly characterised viruses in reference databases can lead to low sensitivity to novel and highly divergent viruses in similarity-based approaches | Expansion of virus reference databases; deep learning models trained on wider datasets |
| Genome assembly | Uneven coverage and population variability in viral and metagenomic datasets (de novo assembly); assembly of genomes with terminal repeats often results in fragmented assemblies; strain-level reconstruction remains challenging | Hybrid assemblers (short and long reads); using approximate rather than exact k-mer alignment |
| Genome annotation | Prone to false positives due to bacterial contamination; machine-learning-based annotations depend on the quality and representativeness of training datasets | Integration of records from multiple databases for cross-validation; improvement of contamination detection |
| Taxonomic classification | Lack of universal viral markers; resolution depends on genome completeness; genomic diversity and rearrangements limit classification accuracy; classification based on entire genomes is not applicable in bulk metagenomes | Standardised virus taxonomy; graph, trees and network-based classification |
| Functional analysis | Large proportion of phage proteins remain hypothetical or uncharacterised; large numbers of ORFans and ‘viral dark matter; rapid evolution and incomplete reference databases give uncertain results | Expansion of virus-specific functional databases; improvement of protein family clustering; integration of machine learning-based structural prediction |
| Host prediction | High uncertainty, especially below the species level; affected by genome variability and database bias; CRISPR-based methods limited by the absence of CRISPRs in host and reference databases; high false-positive rate | Development of hybrid approaches combining multiple signals (CRISPR, sequence similarity, k-mer composition and network-based inference) |
| Statistical analysis | Data sparsity and multidimensionality with excess zeros due to rare taxa; undersampling; compositional effects | Development of statistical methods specific for virome |
| Multianalysis | Individual modules may perform worse than specialised tools; less flexibility and transparency; heavy memory usage | Design of modular pipelines enabling tool substitution, benchmarking and further optimisation |