Skip to main content
Have a personal or library account? Click to login
Diversity analysis and cross-species amplification of custard apple (Annona squamosa) using simple sequence repeats developed via next-generation sequencing technology Cover

Diversity analysis and cross-species amplification of custard apple (Annona squamosa) using simple sequence repeats developed via next-generation sequencing technology

Open Access
|Jul 2025

Full Article

Introduction

India is considered as a secondary centre of origin for custard apple (Annona squamosa L.), which belongs to the family Annonaceae. The chromosome number of A. squamosa is 2n = 2x = 14 (Anuragi et al., 2016). Annonaceae, a family in the class Magnolideae, comprises 129 genera and more than 2000 species. Of these, Annona cherimola Mill. (cherimoya), Annona glabra L. (pond apple), Annona muricata L. (soursop), Annona reticulata L. (custard apple), Annona atemoya (a hybrid of A. squamosa and A. cherimola) and A. squamosa L. (sweetsop or sugar apple) are of major commercial importance (Vinay et al., 2017).

A. squamosa L. is commonly known as ‘custard apple’ or ‘sugar apple’. While most species of Annona are thought to have originated in South America and the Antilles, wild soursop (A. muricata) is believed to have originated in Africa (Pinto et al., 2005). Now, these important species are found in nearly all continents, with soursop and sugar apple being widely distributed, particularly in the tropical regions.

The pulp of Annona fruits is rich in minerals and vitamins (Gyamfi et al., 2011) and is a potential source of dietary fibre (up to 50% w/w dry basis). The seeds, especially those of A. squamosa, contain a significant amount of oil, which can be used for industrial purposes (Mariod et al., 2010). Custard apple is a versatile plant with multiple uses, and it is hardy and deciduous by nature. However, its cultivation is still in the early stages of domestication (Van Zonneveld et al., 2012).

The genetic resources and plant diversity of Annona species are being eroded due to agricultural modernisation, urbanisation and land-use changes. Therefore, the genetic resources of edible custard apples, mostly found in situ and in natural populations, need to be conserved. Characterising genetic diversity is essential for the efficient conservation and utilisation of genetic resources. Despite this, few efforts have been made to identify the diverse germplasm of custard apples and assess the diversity using molecular markers.

There are limited reports on the use of molecular markers to assess genetic diversity in Annona species. Some of these markers include random amplified polymorphic DNA (RAPD) markers (Ronning et al., 1995; Bharad et al., 2009), amplified fragment length polymorphism (AFLP) markers (Rahaman et al., 1998; Zhichang et al., 2011) and simple sequence repeat (SSR) markers derived from related species like A. cherimola (Escribano et al., 2004; Kwapata et al., 2007; Pereira et al., 2008; Escribano et al., 2008; Van Zonneveld et al., 2012; Anuragi et al., 2016; Thanachseyan et al., 2017). However, no species-specific SSR markers have been developed for A. squamosa, which would provide more precise information on the diversity of this economically important species. In light of this, in this study, we have attempted to develop SSR markers and assess the genetic diversity of A. squamosa in the Indian subcontinent and to examine transferability across related species.

MATERIALS AND METHODS

Plant material and DNA isolation

For molecular diversity analysis, a total of 40 A. squamosa cultivars, along with five related species – A. cherimola, A. reticulata, A. glabra, A. muricata and A. atemoya – were selected from the field gene bank of Indian Council of Agricultural Research-Indian Institute of Horticultural Research (ICAR-IIHR), Bengaluru, India (Supplementary Table 1). The 40 cultivars were collected from different regions of India and are maintained in the field germplasm bank. Total genomic DNA was isolated from young, tender leaves using a modified CTAB method (Ravishankar et al., 2000). The quantity and quality of DNA were assessed using a UV spectrophotometer (NABI, Microdigital) at 260 nm and visualised via agarose gel electrophoresis (0.8%) under a UV transilluminator. DNA samples were then diluted to 20 ng ⋅ μL−1 with sterile water and stored at 4°C for further analysis.

Genome sequencing and microsatellite identification

Total genomic DNA from the A. squamosa cultivar Balanagar was used for partial genome sequencing and microsatellite marker identification. Genome sequencing was performed using the Illumina HiSeq 2500 platform at M/Sandor Specialty Diagnostics Pvt. Ltd., Hyderabad, following the manufacturer’s instructions. The reads were assembled into scaffolds using the de novo assembly tool SOAPdenovo2 (version 2.04; Altschul et al., 1990). Microsatellites (SSRs) were identified from the assembled genome fragments using MISA software (Beier et al., 2017), and primers for the identified microsatellites were designed using PRIMER3 software (Rozen and Skaletsky, 2000). Sequence data have been submitted to NCBI (Bioproject PRJNA682654).

PCR amplification and fragment size analysis

A total of 100 SSR primers were randomly selected for PCR standardisation. Pooled DNA was initially used to standardise PCR conditions at various annealing temperatures. Primers were tailed with M13 sequences, and PCR was conducted using M13-tailed primers labelled with fluorophores (Schuelke, 2000). M13 tails were added to both the forward (GTAAAACGACGGCCAGT) and reverse (GTTTCTT) primers at the 5′ end. We tested the 40 A. squamosa genotypes and the 5 related species (A. cherimola, A. reticulata, A. glabra, A. muricata and A. atemoya).

DNA amplification was performed in a 20 μL reaction volume, containing 2.5 μL of template DNA (20 ng · μL−1), 1 μL of each forward and reverse primers (5 pM), 0.4 μL of Taq DNA polymerase (3 U · μL−1), 2.0 μL of Taq buffer A (with 15 mM MgCl2) (10X), 2 μL of dNTPs (1 mM), 1 μL of M13 primers with fluorescent labels (5 pM) (fluorophores such as FAM, ATTO 550, ATTO 565 and YAKAMA YELLOW were synthesised using M/S Eurofins, Bengaluru) and 9.6 μL of sterile water. PCR amplification was carried out using the T100 Thermal Cycler (BIO-RAD, California, USA). The temperature profile was 94°C for 2 min, followed by 35 cycles of 94°C for 30 s, annealing at 50°C/55°C/60°C (Table 3) for 30 s, 72°C for 1 min and a final extension at 72°C for 10 min. Amplified products were confirmed via 1.5% agarose gel electrophoresis, and the primers with strong, distinct amplification products were selected based on Tm values.

The PCR products from the four different fluorescent dyes were multiplexed and resolved using an automated ABI 3730 DNA analyser (Applied Biosystems, USA) at M/S Eurofins, Bengaluru. The resulting data were analysed using Peak Scanner software (Applied Biosystems) to determine the exact fragment size of the PCR products in base pairs (bp).

Statistical analysis

Using the fragment size data, polymorphic information content (PIC), probability of identity (PI), observed heterozygosity (Ho), expected heterozygosity (He) and the number of alleles per locus were calculated using CERVUS software (version 3.0; Kalinowski et al., 2007). SSR genotypic data were used to construct a dendrogram via unweighted pair group method with arithmetic mean (UPGMA), employing a neighbour-joining (NJ) tree and simple matching (SM) dissimilarity matrix using DARwin software. Bootstrap analysis was done with 10000 iterations (version 6.0.10; Perrier and Jacquemoud-Collet, 2003).

A Bayesian model-based clustering was performed using STRUCTURE software (version 2.3.1; Pritchard et al., 2000). The number of subgroups (K) was set between 2 and 10, with 10 iterations for each K value. The project parameters included a burn-in period of 10000, followed by 100000 Monte Carlo Markov Chain (MCMC) replications. Structure Harvester software was used to generate ΩK (Evanno et al., 2005) to estimate the number of populations. Additionally, analysis of molecular variance (AMOVA) was conducted to assess genetic variability among and within populations using GenAlEx software (version 6.5; Peakall and Smouse 2006).

RESULTS

Genome sequencing and microsatellite identification

Next-generation sequencing (NGS) technology, using the Illumina HiSeq 2500 platform, was employed to sequence the total genomic DNA isolated from the A. squamosa cultivar Balanagar. The sequencing run produced 3.9 million bases from 47476322 reads, after low-quality reads were filtered out. A total of 1388525 contigs were generated, and three assemblies were created based on different K-mer lengths of 40, 62 and 89. These assemblies were compared using QUAST software, and the K-mer length of 89 was selected as the best genome assembly based on the number of contigs, overall size, N50 value (954) and contig length distribution. A total of 58527 scaffolds were assembled and screened for microsatellite identification using MISA software (Supplementary Table 2).

The results revealed that the scaffolds contained 179080 microsatellites in total. Among the microsatellite repeats, mono-nucleotide repeats were the most abundant, accounting for 56.2% of all repeat types, followed by di-nucleotide repeats (25.96%), tri-nucleotide repeats (8.7%), tetra-nucleotide repeats (1.26%), penta-nucleotide repeats (0.27%), hexanucleotide repeats (0.11%) and compound nucleotide repeats (7.4%) (Table 1).

Table 1.

Distribution of microsatellite types in A. squamosa genome.

Motif typesNumber of SSRsFrequency (%)
Mono-nucleotide10065856.20
Di-nucleotide4649325.96
Tri-nucleotide155908.70
Tetra-nucleotide22661.26
Penta-nucleotide4950.27
Hexa-nucleotide2020.11
Compound nucleotide133767.4
Total179080

1 SSR, simple sequence repeat.

Using Primer 3.0 software, 58527 primers were designed (excluding mono-nucleotide repeats). Of these, 100 microsatellite markers were randomly selected, and primers were synthesised. After initial screening for amplification of reproducible PCR products, 70 SSR primers were selected for future analysis. Seventy of these primers were used to amplify DNA from 40 A. squamosa genotypes and five related Annona species. The product sizes for the 70 microsatellites ranged from 200 bp to 450 bp. The analysis of data from these microsatellite markers showed a total of 1878 alleles, with an average of 26.82 alleles per locus. The number of alleles ranged from 15 (ASIIHR38) to 46 (ASIIHR19, ASIIHR31 and ASIIHR35) per locus. The three markers with the highest number of alleles (ASIIHR19, ASIIHR31 and ASIIHR35) had simple di-nucleotide repeat motifs. The PIC ranged from 0.78 (ASIIHR61) to 0.96 (ASIIHR40 and ASIIHR19), with a mean of 0.90.

Ho ranged from 0.250 to 1.000, with a mean of 0.647, while He ranged from 0.781 to 0.978, with a mean of 0.926. The PI showed a maximum value of 0.0758 for the ASIIHR21 locus (16 alleles) and a minimum value of 0.0022 for the ASIIHR19 locus (46 alleles), with an average value of 0.01283 across all SSR loci. The characteristics of the 70 polymorphic SSR markers are summarised in Table 2.

Table 2.

Genetic analysis of SSR primers data using. A. squamosa genotypes.

LocusForward sequence and reverse sequence 5′–3′Repeat motif*Tm (°C)No. of alleles per locus (k)HoHePICPI
ASIIHR01 F
ASIIHR01 R
ACATGCCCTCAATCATCTCC
AGTTAAAAATTAATGAATGAGCCAGG
(TA)655230.500.930.910.0103
ASIIHR02 F
ASIIHR02 R
CTGAATTCTAACAGATGTGCTGGG
GCTTTGGACGTCTACGCAT
(TA)860200.600.910.900.0150
ASIIHR03 F
ASIIHR03 R
TCAGACTTCAACTCAAGTCGATCC
TAATACCCAAGTCATGAAGACCAA
(AT)855250.620.930.910.0114
ASIIHR04 F
ASIIHR04 R
ACTCGTACTGCTATAAAAGTGGGT
ATGAGAGCTGCCCACGAC
(AT)755370.650.970.950.0032
ASIIHR05 F
ASIIHR05 R
AAAACTGTCGGCTTCCATGT
AGA AC C A AC AC AGC TTC C ATT
(TA)655200.800.840.820.0364
ASIIHR06 F
ASIIHR06 R
CTTCTGTTGTCATCACTCGCA
CACGATAGAAATCGAAGAACAGA
(TA)655220.520.900.880.0186
ASIIHR07 F
ASIIHR07 R
AAATATCAAGTTTAAAGCAGTATTTGC
TACAAGCGGGACATAGAAGC
(TA)655300.970.950.930.0064
ASIIHR08 F
ASIIHR08 R
GATAGGAAGACAAACAGTTAGTTTAGG
CGGCAGCTTTCTTCTTGTTC
(AG)655280.920.910.900.0143
ASIIHR09 F
ASIIHR09 R
CCAAGAAATCTCACGTTCGC
TCTTTCTTTAGCAAGTAGTGGTTTTT
(AG)1360300.820.950.930.0065
ASIIHR10 F
ASIIHR10 R
TGTGGGTTTATATTGACCATCATT
GTGGAGCTAGGGTTCTTCGTC
(GA)860310.970.930.940.0057
ASIIHR11 F
ASIIHR11 R
ACGCTTTTCTTCTCCGGC
CTCGCCTGCTCCTCTCAC
(GA)955260.520.950.910.0110
ASIIHR12 F
ASIIHR12 R
GCTTTGAGAGAAAATGAGAGACAA
CTACCTTTCCGGCGAATCC
(AG)1155300.920.910.930.0065
ASIIHR13 F
ASIIHR13 R
GATATTCAAAGAGCACGAGAGGA
AAAAGTTAGTCGGCAAATCCC
(GA)655330.700.880.890.0147
ASIIHR14 F
ASIIHR14 R
GTGAGAGAGAGAGAAGGAAGGC
GCGAATCTCTCTTCCGACAG
(GA)1055190.700.930.850.0284
ASIIHR15 F
ASIIHR15 R
TTTTCTCTTTTCTTCGTTCTTGC
GAGGGGGCTGGTGACCTAT
(GA)1260250.820.920.910.0107
ASIIHR16 F
ASIIHR16 R
CCTAATCGGAAAGGTGCAAA
TGGCTTTATTGGATGTGTTTGA
(AC)755230.700.850.910.0126
ASIIHR17 F
ASIIHR17 R
GCTAAGACGGGGCCAACC
TGCCTTATTTCTTTAAACAGGGTC
(AC)660160.750.870.830.0388
ASIIHR18 F
ASIIHR18 R
CTCTCTCTTGTGCTTCTCCCA
CTTTCTCTCTCTCCCCCTCC
(AC)655180.820.970.840.0338
ASIIHR19 F
ASIIHR19 R
TGACGAGATCGAATTAAGTACCC
TGATGCCTATAAAAGGGGCA
(CA)860460.920.940.960.0022
ASIIHR21 F
ASIIHR21 R
ACCAGCAAATCCTGGGAAG
TGAATCTGCAACTCAAAACTGA
(AC)655291.000.780.930.0076
ASIIHR21 F
ASIIHR21 R
TTGATGCAATTCTTCAGTTTGA
TACAAGGGTTGGGAATTTGG
(AC)855160.920.920.740.0758
ASIIHR22 F
ASIIHR22 R
CATACATTTTGCCCACGACC
CACAGATACACACACACGAACAA
(GT)860250.770.950.900.0132
ASIIHR23 F
ASIIHR23 R
AAAAAGTCCATTCTTTTTCTCCA
TCATTTCTTCATCCCATTGC
(GT)655280.920.940.930.0066
ASIIHR24 F
ASIIHR24 R
CACATCACCCATATAAAAAGCG
GTTGGGACAACTCTTCACCC
(TG)660350.820.930.930.0065
ASIIHR25 F
ASIIHR25 R
CAGCGATGGTTGCTTAATTTG
AGTAGGTGGAGAGACCCACG
(GT)655250.770.930.920.0099
ASIIHR26 F
ASIIHR26 R
AGCAAAAGTGGTCATCCGAA
TTTTAGATCGCTAAGAAGATCACA
(TG)1360300.800.960.910.0102
ASIIHR27 F
ASIIHR27 R
TCGCTATTTCAAAATTAAGTAAAAGAA
TGTTTTAGTTGAGCAAGGCTAGG
(TG)855370.820.960.950.0044
ASIIHR28 F
ASIIHR28 R
TCTTGTTTTTGCCAGTTCCC
GCTTGAACTCAAAGCATGTTG
(GT)860340.870.870.950.0038
ASIIHR29 F
ASIIHR29 R
TGTTACTGTTGGGCATGGAA
GGTTGGGTTGAAAATTTAAGCA
(GT)655230.670.950.850.0263
ASIIHR30 F
ASIIHR30 R
CCTTCCACCCTTGGATCTTA
GATCAATGTGGATAAAACTTCGC
(CT)655330.820.970.940.0058
ASIIHR31 F
ASIIHR31 R
CTTTTCTTCTCCATTTTCCCG
CGTTGGAGATTCCAAGAGAAA
(CT)855440.950.950.950.0031
ASIIHR32 F
ASIIHR32 R
AGGTGGATCGCTTAAGATGAA
TCTGAGCAGGTGTGTAGTCCT
(TC)655290.800.920.930.0069
ASIIHR33 F
ASIIHR33 R
ACTGGCCGAGGAAAGGGT
GAGAGGGAGGAAAAAGTTAATCG
(TC)1255280.970.970.910.0117
ASIIHR34 F
ASIIHR34 R
CTCCCCGTTACCCAACTG
GGTTCCTTCCGTCCTCCTG
(CT)955350.720.960.950.0034
ASIIHR35 F
ASIIHR35 R
TTTC ATAGC TTTTATTGC TTTCTTAG
AGTCGGGTCGAATCCACA
(GA)855400.850.960.950.0036
ASIIHR36 F
ASIIHR36 R
CTTCTGTCTTCCTCATTTTCTCG
TGGGAGAGAAGTGAGAGAGC TT
(AG)860350.770.940.940.0047
ASIIHR37 F
ASIIHR37 R
GGCCACACTTGCTCAAAAAT
TTCTTATTGATGCGTAGCGG
(GC)755190.920.900.920.0090
ASIIHR38 F
ASIIHR38 R
GGGAGGAAACTTGATCCCTT
AAAATTATGGTGCAGTGGCG
(GC)660150.270.940.880.0196
ASIIHR39 F
ASIIHR39 R
GGCCACACTTGCTCAAAAAT
TTCTTATTGATGCGTAGCGG
(GC)755280.270.970.920.0087
ASIIHR40 F
ASIIHR40 R
CCAATCCCTTTATCCAAGCA
GTGCAATTTGAAAAGCAGCA
(GC)755350.270.930.960.0027
ASIIHR41 F
ASIIHR41 R
CATCTCCGCAACACCAGATA
GCCAGAAGAGGCAGTGATTC
(GC)855270.670.940.920.0092
ASIIHR42 F
ASIIHR42 R
AGAGGAAAACTTACAAAAACATAGACG
GCCCTAAGAGAGCAGTTTACCC
(ATG)655260.300.930.920.0089
ASIIHR43 F
ASIIHR43 R
GTATGTCATGGAGGATACAGGGA
CATCCTCATCATCTCCCACA
(GAA)555230.250.900.920.0091
ASIIHR44 F
ASIIHR44 R
ACTGCTGCTGAGATGTGCG
CTGCTGCTGCTGTTGACG
(GAT)755230.300.920.910.0112
ASIIHR45 F
ASIIHR45 R
TTATTGTATAAAACACCCCAAAGAA
TTCTTATCATTTTGCCCGTCT
(TTC)1060180.450.960.890.0183
ASIIHR46 F
ASIIHR46 R
TTGGCAACCATCAGAATAAGA
ACAACCCTGCTTCCATTCAA
(AAT)560170.350.940.900.0152
ASIIHR47 F
ASIIHR47 R
AAAAACCTTGGGCTTGTGC
CCTTGCCCCTATTATTTTCC
(TCT)655320.370.950.950.0037
ASIIHR48 F
ASIIHR48 R
TTGGTGAAGCATTCAAAAATTC
CCTTGCCCCTATTATTTTCC
(ATG)1055320.520.930.930.0072
ASIIHR49 F
ASIIHR49 R
TCAAACGCCCGCATATTTA
GCTGGAGAAAGACGGCAAG
(AGC)555260.400.940.930.0070
ASIIHR50 F
ASIIHR50 R
ACCTCAAAGCTAGGGGGTAAA
CCGAAGTAGAGATACGCCTCTT
(GAA)555260.350.930.920.0103
ASIIHR51 F
ASIIHR51 R
AAAATGAGCATGAAGAAAAGAAAAA
GATTGTAAGACAAATTGAGATGTAATG
(ATT)1055220.300.940.930.0081
ASIIHR52 F
ASIIHR52 R
TCCCATTTTCTGATCGAGTTG
TAACCCTCGCCGTGAATAG
(TCA)555230.420.930.920.0101
ASIIHR53 F
ASIIHR53 R
AGACTAATCTAAGTTTAAAGCAAGCAA
TGATC TTTGTTGAAGCGTCTCT
(AAG)555210.450.940.920.0090
ASIIHR54 F
ASIIHR54 R
TTCAAGAATCATCTTTTAAGTCAACC
TTTTCATGGAAACAACCAAATG
(GTG)555300.450.940.920.0079
ASIIHR55 F
ASIIHR55 R
TCACTTGGAATAATGTGGAACG
CAGTCCCCAGCTCCAAAC
(GTC)1155190.400.910.890.0181
ASIIHR56 F
ASIIHR56 R
CCCCAATCCCAATCCTTAGT
GGAGTTCGTGTGCTTTACCG
(AGC)555280.750.960.940.0047
ASIIHR57 F
ASIIHR57 R
GACGTGCTGCTG
GGATTCTTCACCAGGCAGTT
(CGA)555240.350.920.900.0134
ASIIHR58 F
ASIIHR58 R
AAAATGCATGCCTTGTTTGT
TAGTCCTTGAGCAACACATGC
(ATTC)760230.550.910.900.0146
ASIIHR59 F
ASIIHR59 R
AAGGGCATGTTGTCTTCTCAA
TTTGCAAGTTTATTTCTGCTCAA
(TTGT)660290.450.950.930.0070
ASIIHR60 F
ASIIHR60 R
TCCCGACCTTTCTTACGGAT
TTTATCTCTATCTCTCCACCGGA
(ATCTC)660220.670.870.850.0276
ASIIHR61 F
ASIIHR61 R
AATATGTTAACCCGAAACTCAACC
AATAATTACATAATTGATGAGGGTCAA
(TAGGGT)555160.600.950.780.0571
ASIIHR62 F
ASIIHR62 R
CTCTGTTTCTATCTCTCTCAAACTCA
GGAATGGGAGAGTATTTGAAGG
(TC)9tgtctctcatt(TC)655340.500.870.940.0058
ASIIHR63 F
ASIIHR63 R
TACCGGATCTCTCATTTTCG
GCTGGGAGAGTGAGCTGAAA
(TC)5cgccatttctatctct ctctccgtcattcct ctctctctttctctgttttccttt ggaaaaatcggca aacccaaat(TC)655240.620.810.900.0139
ASIIHR64 F
ASIIHR64 R
CATGTTGGACATGTGAGCCA
ATGCCTATTTTAGGCTGGGT
(A)10g(A)1055280.850.950.910.0101
ASIIHR65 F
ASIIHR65 R
GGCGTCCAAAAATTGAGATT
TTTGGGGAGTATCTACTCAGGC
(AT)7gtatttgccccatggg ccccaaaaaaaataaa(AT)7 ttttagactttcaa(AT)755310.670.920.920.0082
ASIIHR66 F
ASIIHR66 R
TCTACGCTACCCAGCAAATG
AATGACAATATGATTTGCCTTGA
(T)11(TA)6*55280.670.930.900.0128
ASIIHR67 F
ASIIHR67 R
CCGACTTCAACCTTCTGAGC
TCAATGGAATATTCACTTTTCTAGG
(A)11ttaaaacagattta tcaaaaatgttctta cgagaaagggaaaa taggagaaaaaggt agaatgagggttttc tttttgtgcgtgttttgga(AG)855270.600.940.910.0075
ASIIHR68 F
ASIIHR68 R
GTGGATACTCCCCGACTGG
TTCACATACTTTTGCCTGAGTAGA
(AAT)655260.550.950.920.0066
ASIIHR69 F
ASIIHR69R
TTCACATACTTTTGCCTGAGTAGA
GCCATCTTGGGCTTTTTAGA
(AT)6gacctcttcgctcaatccaagcctcaatgtc(A)1055210.670.920.900.0143
ASIIHR70F
ASIIHR70 R
GAAGGGAGATGCAAACGTTAAG
AACTGCTTGCTTTATGCACTTT
(A)11(T)1055270.600.920.910.0118

* Indicates the number of times a particular motif was repeated in the microsatellite locus. For example: (TA6 – means TA has been repeated six times in the sequence results analysed during SSR identification. He, expected heterozygosity; Ho, observed heterozygosity; PI, probability of identity; PIC, polymorphic information content; SSR, simple sequence repeat.

Genetic diversity and population structure analysis of A. squamosa genotypes

A dendrogram (Figure 1) was generated using the UPGMA method, based on a shared allele matrix, revealing two major clusters. Of the 40 cultivars, 23 (57.5%) were grouped in Cluster I, while the remaining 17 (42.5%) were grouped in Cluster II, based on their genetic relatedness. Cultivars native to Andhra Pradesh were distributed across both clusters. A bar plot of ΔK, following the method described by Evanno et al. (2005), indicated that the optimal ΔK was at K = 2 (Figure 2).

Figure 1.

Dendrogram analysis of A. squamosa genotypes using SSR markers data. SSR, simple sequence repeat.

Figure 2.

Structure analysis of Annona genotypes using SSR data.

An AMOVA revealed limited variation among populations, with 30% of the variation occurring among individuals and 70% within individuals of the populations. The Fst value from AMOVA was 0.065, indicating low genetic differentiation (Table 4).

DISCUSSION

In this study, we utilised the Illumina HiSeq 2500, an NGS platform, to sequence the genome and isolate microsatellites for A. squamosa (L.). NGS technologies have become essential tools in plant biology for wholegenome sequencing, marker development and gene identification. Our sequencing efforts yielded 3.9 million bp sequences, and the assembly of 47476322 raw reads resulted in 58527 scaffolds. Microsatellite analysis revealed that mono-nucleotide repeats were the most abundant (56.20%), followed by di-nucleotides (25.96%), tri-nucleotides (8.7%) and tetra-nucleotides (1.26%). Mono-, di-, tri- and tetra-nucleotide repeats accounted for the majority (92.12%) of microsatellites in custard apple. This pattern of mono-repeat predominance and followed by di-repeats is consistent with findings in other plant species (Ravishankar et al., 2015, 2017a, 2017b).

Among the 179080 identified SSRs, 46493 were dinucleotide repeats, 15590 were tri-nucleotide repeats, 2266 were tetra-nucleotide repeats, and the remaining were penta- and hexa-nucleotide repeats (Table 2). From the 58527 designed primers, 70 SSRs were randomly selected and standardised for PCR conditions. These SSRs were employed for DNA amplification of 40 Annona cultivars and 5 related species (A. cherimola, A. reticulata, A. glabra, A. muricata and A. atemoya) (Table 3).

Table 3.

Cross species amplification of 70 SSR loci derived from A. squamosa.

No.LocusA. cherimolaA. reticulataA. glabraA. muricataA. atemoya
1ASIIHR01AAAAA
2ASIIHR02AAAAA
3ASIIHR03AAAAA
4ASIIHR04AAAAA
5ASIIHR05AAAAA
6ASIIHR06AAAAA
7ASIIHR07NAANAANA
8ASIIHR08AAAAA
9ASIIHR09AAAAA
10ASIIHR10AAAAA
11ASIIHR11AAAAA
12ASIIHR12AAAAA
13ASIIHR13AAAAA
14ASIIHR14AAAAA
15ASIIHR15AAAAA
16ASIIHR16AAAAA
17ASIIHR17AAAAA
18ASIIHR18AAAAA
19ASIIHR19AAAAA
20ASIIHR20AAAAA
21ASIIHR21AAAAA
22ASIIHR22AAANANA
23ASIIHR23AAAAA
24ASIIHR24AAAAA
25ASIIHR25AAAAA
26ASIIHR26AAAAA
27ASIIHR27NAAAANA
28ASIIHR28AAAAA
29ASIIHR29AAAAA
30ASIIHR30AAAAA
31ASIIHR31AAAAA
32ASIIHR32AAAAA
33ASIIHR33AAAAA
34ASIIHR34AAAAA
35ASIIHR35AAAANA
36ASIIHR36AAAAA
37ASIIHR37AAAAA
38ASIIHR38AAAAA
39ASIIHR39AAAAA
40ASIIHR40AAAAA
41ASIIHR41AANAANA
42ASIIHR42AAAAA
43ASIIHR43AAAAA
44ASIIHR44AAAAA
45ASIIHR45AAAAA
46ASIIHR46AAAAA
47ASIIHR47AAAAA
48ASIIHR48AAAANA
49ASIIHR49AAAAA
50ASIIHR50AAAAA
51ASIIHR51NAAAAA
52ASIIHR52AAANANA
53ASIIHR53NAAAANA
54ASIIHR54AAAAA
55ASIIHR55AAAAA
56ASIIHR56AAAAA
57ASIIHR57AAAAA
58ASIIHR58ANAANANA
59ASIIHR59AAAAA
60ASIIHR60AAANANA
61ASIIHR61AAAAA
62ASIIHR62ANAANANA
63ASIIHR63AAAAA
64ASIIHR64AAAAA
65ASIIHR65AAANANA
66ASIIHR66AAAAA
67ASIIHR67AAAAA
68ASIIHR68ANAANANA
69ASIIHR69AAAAA
70ASIIHR70AAAAA
Transferability %94.295.797.190.0080.00

1 A, amplified, NA, not amplified; SSR, simple sequence repeat.

The SSR markers showed high PIC values, ranging from 0.78 to 0.96, which were notably higher than those reported in earlier studies using SSRs from A. cherimola in A. squamosa (Anuragi et al., 2016). For example, Anuragi et al. (2016) examined molecular diversity among 20 A. squamosa genotypes using 12 A. cherimola SSR markers, which had PIC values ranging from 0.169 to 0.694. Thus, species-specific SSR are more efficient in assessing the diversity. Similarly, Thanachseyan et al. (2017) studied 34 accessions of A. muricata using 8 A. cherimola SSR markers, reporting an average PIC value of 0.0131, which is much lower than our findings. Thus, it appears that analysis employing species-specific SSR markers gives a better indication of diversity and heterozygosity in the population.

Table 4.

AMOVA analysis of genetic variances within and among populations of 40 custard apple accessions.

Source of variationdfSSMSEstimated variance% Variation
Among populations5213.76942.740.0210
Among individuals341447.24442.569.87030
Within individuals40913.0022.82522.82570
Total100

1 AMOVA, analysis of molecular variance; df, degrees of freedom; MS, mean sum of squares; SS, sum of squares.

In this study, all 70 SSR markers were highly polymorphic, with PIC values exceeding 0.5 (Table 3), indicating a high level of genetic diversity among the 40 Annona genotypes. This was further supported by the high Ho values, which ranged from 0.250 to 1.000, and He values, which ranged from 0.781 to 0.978 (Table 3). These results demonstrate the effectiveness of speciesspecific SSR markers for assessing genetic diversity in A. squamosa. Additionally, the Ho being higher than expected suggests the presence of at least two isolated populations among the genotypes studied.

Among the markers, three SSRs (ASIIHR19, ASIIHR31 and ASIIHR35) showed the highest number of alleles (46), all of which were derived from dinucleotide perfect repeat motifs (CA)8, (CT)8 and (GA)8, respectively. Previous studies (Liu et al., 2016 in kiwifruit; Biswas et al., 2014 in sweet orange; Chapman, 2019 in lablab) have also reported that microsatellite markers with di-nucleotide repeat motifs exhibit a significantly higher number of alleles compared with those with tri-, tetra- and penta-nucleotide repeats, likely due to the ease of mutation via DNA slippage during replication.

Furthermore, seven markers (ASIIHR04, ASIIHR19, ASIIHR31, ASIIHR34, ASIIHR35, ASIIHR40 and ASIIHR47) with low PI values ranging from 0.0022 to 0.0037 were found to be suitable for DNA fingerprinting of A. squamosa genotypes (Table 3). SSR markers with low PI values are highly useful for DNA fingerprinting (Ravishankar et al., 2017a, 2017b), making these markers ideal for the identification of custard apple varieties.

The dendrogram analysis classified the A. squamosa genotypes into two main groups, further subdivided into sub-clusters based on genetic relatedness (Figure 1). Low bootstrap values observed here may be due to only too few SSR markers employed to estimate meaningful genetic differentiation or from a weak population structure due to recent divergence or gene flow. The later explanation is plausible given the recent introduction of Annona species to the Indian subcontinent. This clustering aligns with earlier studies, such as Anuragi et al. (2016), where 20 A. squamosa genotypes were grouped into 7 clusters using 12 A. cherimola SSR markers. The high genetic diversity observed in this study is likely due to the collection of germplasm from diverse locations and species-specific markers used, reflecting the presence of a broad genetic base in A. squamosa. A similar pattern of high diversity was reported for A. cherimola (Escribano et al., 2007) and A. senegalensis (Kwapata et al., 2007).

A Bayesian model-based analysis further supported the clustering results, grouping the 40 A. squamosa genotypes into two major clusters (Figure 2A). The ΔK value derived from Evanno’s algorithm predicted K = 2, consistent with the dendrogram results. Cluster I comprised genotypes from Andhra Pradesh, Tamil Nadu, Maharashtra, Karnataka, the USA and Taiwan, while Cluster II comprised genotypes from Andhra Pradesh, Telangana, Odisha and the USA.

Cross-species amplification of SSRs was highly successful, with a transferability rate of 80.0%–97.1% across five Annona species (Table 3). This high transferability is likely due to the conserved nature of flanking sequences around microsatellites in Annona species. Previous studies, such as those by Anuragi et al. (2016) and Kwapata et al. (2007), reported similar results, with SSR markers showing high cross-species transferability among different Annona taxa.

Although SSR development is expensive, once established, these markers are cost-effective and timeefficient for genetic analysis. In A. squamosa, SSRs have not been extensively used, but their high polymorphism suggests that they are valuable for assessing genetic diversity. Most of the genetic variation in this study was due to differences within populations, likely because A. squamosa is a recently introduced cultivated species. The AMOVA results, with an Fst value of 0.065, indicate significant genetic differentiation among the 40 custard apple genotypes, reflecting the coexistence of different genotypes within the same region (Ravishankar et al., 2015).

CONCLUSIONS

NGS proved to be an efficient method for identification and development of microsatellite markers. Compared with previous studies, this study revealed relatively higher heterozygosity and PIC values, emphasising the usefulness of species-specific SSR markers. The high PIC, polymorphism and Ho values indicate that the SSR markers developed in this study are effective for genetic diversity analysis in A. squamosa. These markers also demonstrated high transferability to closely related species, including A. cherimola, A. reticulata, A. glabra, A. muricata and A. atemoya. The microsatellite markers generated in this study will be valuable for genetic diversity studies, mapping, development of linkage map and other crop improvement programmes, like gene discovery, in A. squamosa.

ACKNOWLEDGEMENTS

We would like to thank the Indian Council of Agricultural Research-Indian Institute of Horticultural Research (ICAR-IIHR), Hesaraghatta Lake Post, Bangalore for providing facilities.

Notes

[5] Conflicts of interest CONFLICT OF INTEREST

The authors report that there are no competing interests to declare. AI tool Kimi has been used to correct English language only. Sequence data has been submitted to NCBI (Bioproject PRJNA682654).

SUPPLEMENTARY MATERIALS

Supplementary Table 1.

List of 40 custard apple (A. squamosa) genotypes/cultivars belonging to A. squamosa and different species of Annona used in the study.

Sl. No.Cultivars/accessions No.Place of collectionState
1.BalanagarSangareddyTelangana
2.RaidurgAnanthapurAndhra Pradesh
3.TaiwanTaiwanTaiwan
4.Arka_Sahan (IC No. 0632061)IIHRKarnataka
5.NMK_1GarmaleMaharashtra
6.APK_1AruppukottaiTamil Nadu
7.WASHINGTON_05FloridaUSA
8.BarbadosFloridaUSA
9.Washington_97FloridaUSA
10.19/26 (IC no. 0632061)IIHR, BangaloreKarnataka
11.Arka Neelanchal VikramBhubaneswarOdisha
12.MammothFloridaUSA
13.Red_sitaphalSangareddyTelangana
14.6_8NalakadoddiAndhra Pradesh
15.2_1VengalampalliAndhra Pradesh
16.3_1VengalampalliAndhra Pradesh
17.1_1VengalampalliAndhra Pradesh
18.8_8PythotaAndhra Pradesh
19.8_9PythotaAndhra Pradesh
20.2_2VengalampalliAndhra Pradesh
21.27_1JambugumpalaAndhra Pradesh
22.5_8MolakalmuruKarnataka
23.3_2VengalampalliAndhra Pradesh
24.5_1MolakalmuruKarnataka
25.8_18PythotaAndhra Pradesh
26.1_10VengalampalliAndhra Pradesh
27.8_16PythotaAndhra Pradesh
28.11_8YercaudTamil Nadu
29.3_3VengalampalliAndhra Pradesh
30.4_11MolakalmuruKarnataka
31.8_17PythotaAndhra Pradesh
32.1_5VengalampalliAndhra Pradesh
33.2_13VengalampalliAndhra Pradesh
34.4_1MolakalmuruKarnataka
35.2_10VengalampalliAndhra Pradesh
36.4_10MolakalmuruKarnataka
37.6_10NalakadoddiAndhra Pradesh
38.4_12MolakalmuruKarnataka
39.8_7PythotaAndhra Pradesh
40.RoleniaMoodubidireKarnataka
41.A. cherimolaSangareddyTelangana
42.A. reticulataSangareddyTelangana
43.A. glabraSangareddyTelangana
44.A. muricataSangareddyTelangana
45.A. atemoyaSangareddyTelangana
Supplementary Table 2.

Summary of sample Balanagar assembled genome.

Plant WGSBalanagar
Contigs generated1388525
Maximum contig length21462
Minimum contig length100
Average contig length466.126
Total contigs length647227553 (647.2 MB)
Total number of non-ATGC characters54429
Percentage of non-ATGC characters8.40956E-05
Contigs >100 b:1379816
Contigs >500 b:350135
Contigs >1 Kb:170254
Contigs >10 Kb:19
Contigs >100 Kb:0
n50 value:954
n90_value:151
kmer length89
Total number of scaffolds58528
Total number of identified SSRs1790080

1 SSR, simple sequence repeat.

DOI: https://doi.org/10.2478/fhort-2025-0009 | Journal eISSN: 2083-5965 | Journal ISSN: 0867-1761
Language: English
Page range: 139 - 155
Submitted on: Feb 6, 2025
Accepted on: May 20, 2025
Published on: Jul 10, 2025
Published by: Polish Society for Horticultural Sciences (PSHS)
In partnership with: Paradigm Publishing Services
Publication frequency: 2 issues per year
Related subjects:

© 2025 Abhilasha Krishnamurthy, Thandavarayan Sakthivel, Kundapura V. Ravishankar, published by Polish Society for Horticultural Sciences (PSHS)
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 3.0 License.