Work overview

Section 05 of 09

Results and discussion

A comprehensive analysis of Y-chromosomal diversity in Colombian populations

Julyana Ribeiro, Zehra Köksal, Adriana Castillo, Adriana Ibarra, Beatriz Martinez, Humberto Ossa, Masinda Nguidi, Juan David Granda, Verónica Gomes, Pedro Rodrigues, Maria Inês Machado, Vania Pereira, and Leonor Gusmão · 2026

Contents

Section 05 of 09

  1. 01Introduction
  2. 02Materials and methods
  3. 03Y-STR genotyping
  4. 04Y-SNP genotyping
  5. 05Results and discussion
  6. 06Comparison of Y-STR haplotypes in Colombian populations
  7. 07Comparisons of admixed South American populations
  8. 08Conclusions
  9. 09Supplementary Information
Text size
Work overview

Section 5 of 9

Results and discussion

Julyana Ribeiro, Zehra Köksal, Adriana Castillo, Adriana Ibarra, Beatriz Martinez, Humberto Ossa, Masinda Nguidi, Juan David Granda, Verónica Gomes, Pedro Rodrigues, Maria Inês Machado, Vania Pereira, and Leonor Gusmão · about 6 minutes

Haplotype diversity and allelic variation

Haplotype profiles and haplogroups obtained in this study are shown in Supplementary Table S2.

A high value of HD (0.9998 ± 0.0001) was observed in the general sample from Colombia, consistent with those reported in other admixed South American populations [39–41, 43, 54, 55]. In the Colombian samples analyzed, 913 distinct haplotypes were identified: 861 were singletons (observed only once), 44 were shared by two individuals, six by three individuals, and two haplotypes were shared by four individuals.

An allele 6 at DYS391 locus was detected in four Colombian samples. Although rare, this allele is represented in YHRD in samples from East Asia, and with a lower frequency in admixed populations from South America. It was previously reported in the Andean region of Colombia and associated with a novel Native American lineage within Q1a2-M346*(xM3) [16]. In this study, the allele DYS391*6 was also detected in the Andes, as well as in both the Caribbean and Amazon regions in samples predicted to belong to haplogroup Q1b1a2-Z780, according to results from the NEVGEN online tool (three samples with probabilities ≥ 98.1% and one sample with a probability of 42.2%).

Intermediate, duplicated and null alleles were detected in 64 out of the 975 unrelated male samples genotyped in this study (Table 1). To investigate the origin of these alleles, haplogroups of four samples were attributed based on Y-SNP genotyping using the Ion AmpliSeq™ HID Y-SNP Research Panel v1. For the other 60 samples, haplogroups were predicted from each corresponding haplotype using the NEVGEN software. In 20 samples, haplogroups were predicted with less than 92% probability. Of these, 16 samples were genotyped for relevant Y-SNPs by SNaPshot, and the results were compatible with the haplogroup predicted by NEVGEN in all cases. No reliable results were obtained for four samples from the Andes, namely (i) one carrying a DYS45817.2 allele, assigned to J1a2a1a2-P58; (ii) one carrying a DYS45815,16 duplication, with no haplogroup predicted (100% Unsupported subclade); (iii) one carrying a DYS44819,20 duplication, assigned to E1a-M132; and (iv) one carrying a DYS 45816.2 allele, assigned to J1a2a1a2-P58.

Seventeen different intermediate alleles were observed in the complete dataset at the following loci: DYS437, DYS458, DYS385, DYS570, DYS635, and DYS438 (Table 1). Intermediate alleles at Y-STRs have been widely reported at various loci in different populations. In the present dataset, DYS458 exhibited the highest number of intermediate alleles, followed by DYS385. For both markers, intermediate alleles were detected in samples from different regions, with no geographic specificity being found (Table 1).

Y-STR | Genotype | Region (n) | NEVGEN - Haplogroup Predicted#
Intermediate alleles | 
DYS385 | 17,17.2 | Caribbean (2) | A1a-M31
DYS385 | 10.2,12 | Andean (1) | R1b-DF27
DYS385 | 11,13.2 | Andean (1) / Caribbean (1) | *R1b1a1a2a1a2a1b1a / R1b-U152
DYS385 | 11.2,12 | Andean (1) | R1b-DF27
DYS385 | 13,14.2 | Caribbean (1) | J2a1-L26
DYS385 | 13.2,16 | Caribbean (1) | R1b-V88
DYS437 | 13.2 | Andean (1) | R1b-DF27
DYS437 | 14.2 | Andean (1) | R1b-U106
DYS438/DYS458 | 9.2/19.2 | Andean (3) | J1a3-Z1828
DYS458 | 16.2 | Andean (3) / Caribbean (3) | 1 R1b-Z2103 / 5 J1a2a1a2-P58
DYS458 | 17.2 | Andean (6) / Caribbean (6) / Orinoquía (1) | 12 J1a2a1a2-P58 / 1 *J1a2a1a2d2b2b2c
DYS458 | 18.2 | Andean (5) / Caribbean (1) / Pacific (1) | 6 J1a2a1a2-P58 / 1 *J1a2a1a2d2b2b2c4c
DYS458 | 19.2 | Andean (1) / Caribbean (2) / Pacific (1) | 3 J1a3-Z1828 / 1 J1a-PH77
DYS458 | 20.1 | Caribbean (2) | 1 E1b1b-M81 /1 J1a-PH77
DYS458 | 20.2 | Caribbean (1) | J1a3-Z1828
DYS570 | 17.2 | Andean (1) | R1b-DF27
DYS635 | 23.1 | Andean (1) | R1b-DF27
Duplications | 
DYS389I/DYS437 | 12,13/14,15 | Caribbean (1) | R1b-L51
DYS448 | 18,20 | Andean (1) | R1b-U106
DYS448 | 20,21 | Caribbean (2) | E1a M132
DYS448 | 19,20 | Andean (1) / Caribbean (1) / Orinoquía (1) | 1 R1b-P312 / 2 E1a-M132
DYS448 | 19,21 | Caribbean (1) | E1a-M132
DYS456 | 15,16 | Andean (1) | 100% Unsupported subclade
DYS576 | 15,17 | Pacific (2) | E1b1a-V38
Null alleles | 
DYS19/DYS392 | 0 | Andean (1) | R1b-L21
DYS390 | 0 | Andean (1) | R1b-L21
DYS448 | 0 | Andean (3) / Caribbean (1) | 3 R1a-M198 / 1 *R1b1a1a2a1a1c2b2a1b1b1a

Samples with DYS385 intermediate alleles were assigned to haplogroups A1a, J2a1, and different sub-lineages within R1b, indicating that these intermediate alleles likely arose from multiple independent events. It is noteworthy that the two haplotypes classified within haplogroup A1a, in addition to allele 17.2 in DYS385 also had allele 6 at DYS533, a rare variant found just in one sample in the YHRD Admixed Metapopulation. Most samples with non-consensus alleles at DYS458, including three individuals simultaneously carrying 9.2 at DYS438 and 19.2 at DYS458, were classified within J1a, except one that was assigned to the R1b haplogroup. This finding aligns with previous associations of DYS458 intermediate alleles with J1a-M267 and R1b3-M405 haplogroups [56]. The remaining intermediate alleles at DYS437, DYS570 and DYS635 were assigned to sub-lineages within haplogroup R1b, which was the most frequent in the studied Colombian samples. The low number of observations does not allow inferences regarding the origin of these alleles.

Duplicated alleles were observed at DYS448, DYS456, DYS576, DYS389I and DYS437 (Table 1). Deletions, duplications, and triplications are frequently observed in ampliconic regions of the Y chromosome, since this chromosome is under low selective pressure [57]. On the long arm of the Y chromosome, the Azoospermia Factor regions (AZFa, AZFb, AZFc) are susceptible to frequent structural rearrangements. Y-STRs used in forensic genetics that are located inside these regions (e.g. DYF387S1, DYS385, DYS389I/II, DYS391, DYS392, DYS437, DYS438, DYS439, DYS448, DYS460, DYS549 and DYS635) tend to present duplications and/or deletions, as a result of non-allelic homologous recombination between palindromes [58–61]. This mechanism can explain the duplications involving both DYS437 and DYS389I (Table 1), located in the AZFa [62], as well as the seven duplications at DYS448, located in the AZFc. The lack of association of the DYS448 duplications with a specific haplogroup point to at least three independent duplication events. Balaresque et al. [58] identified DYS448 duplications in samples belonging to haplogroups O3e, C*, C3c, G, O2, D*, and E3b, supporting independent deletion and duplication events within AZF regions, making it challenging to assign a specific geographic origin to the duplications and null alleles observed in this study.

DYS456 and DYS576 are both located on the short arm of the Y chromosome. For the sample with the duplication at DYS456, the NEVGEN software could not predict any haplogroup. The two samples with duplication at DYS576 are from the Pacific region, and they were predicted as belonging to the African haplogroup E1b1a. This finding is consistent with the previous detection of this duplication in a sample from Benin [39] and with the high levels of African ancestry in the Colombian Pacific region [5].

Null alleles were also observed at DYS448, DYS390, DYS19 and DYS392. In the case of DYS448, the null alleles can also be due to non-allelic homologous recombination within the AZFc, which explains the 4 instances observed in Andes and Caribbean within two different haplogroups (R1a and R1b).

Null alleles at DYS390, DYS19 and DYS392 are outside palindromic regions and, therefore, they most likely resulted from mutations at the primer binding regions. Null alleles at these loci were described in Ghana [44] and South Africa [63, 64], and reported in YHRD in admixed South American, Native American and African populations. In our samples, the haplotype carrying null alleles simultaneously at DYS19 and DYS392, as well as the haplotype with a null allele at DYS390, were predicted to belong to haplogroup R1b-L21.