Work overview

Section 07 of 09

Comparisons of admixed South American populations

A comprehensive analysis of Y-chromosomal diversity in Colombian populations

Julyana Ribeiro, Zehra Köksal, Adriana Castillo, Adriana Ibarra, Beatriz Martinez, Humberto Ossa, Masinda Nguidi, Juan David Granda, Verónica Gomes, Pedro Rodrigues, Maria Inês Machado, Vania Pereira, and Leonor Gusmão · 2026

Contents

Section 07 of 09

  1. 01Introduction
  2. 02Materials and methods
  3. 03Y-STR genotyping
  4. 04Y-SNP genotyping
  5. 05Results and discussion
  6. 06Comparison of Y-STR haplotypes in Colombian populations
  7. 07Comparisons of admixed South American populations
  8. 08Conclusions
  9. 09Supplementary Information
Text size
Work overview

Section 7 of 9

Comparisons of admixed South American populations

Julyana Ribeiro, Zehra Köksal, Adriana Castillo, Adriana Ibarra, Beatriz Martinez, Humberto Ossa, Masinda Nguidi, Juan David Granda, Verónica Gomes, Pedro Rodrigues, Maria Inês Machado, Vania Pereira, and Leonor Gusmão · about 10 minutes

For comparative purposes, PPY23 haplotype datasets from the Colombian regions were compared with those from admixed South American, Native American, European, East Asian and African populations, for which data were available for the same set of markers. The results are presented in Supplementary Table S4. As shown in Fig. 3A, populations from the Andean, Caribbean, and Orinoquía regions cluster with European and other admixed populations, indicating stronger genetic affinities with Argentina, Paraguay and Brazil, all showing high genetic proximity to Iberian populations [41, 43, 60]. These results are historically consistent, as the colonization of South America was predominantly carried out by settlers from the Iberian Peninsula beginning in the early 15th century. It is also important to highlight that Germany appears more distant from this cluster, particularly according to the _R_ST results.

Although pairwise genetic distances (_F_ST and _R_ST) indicate statistically significant differentiation between Colombian regions and African groups, the Pacific region still shows a greater proximity to the western African populations. As previously reported for Y-SNPs in the Caribbean in this study, this region exhibits a significantly higher frequency of sub-Saharan lineages, highlighting stronger genetic affinities with African populations.

The Amazon population clusters with Peru and Ecuador, showing no significant differences, likely due to the strong Native American component shared among these groups. However, although the Amazonian population is genetically closer to Bolivia, the increased distance and significant differences relative to Bolivian natives suggest that the Andean Native component of Bolivia is distinct from the Amazonian component found in Colombia.

Fig. 3: MDS plot of populations comparisons. A) MDS plot based on FST genetic distances between Colombian regions and other populations, based on 23 Y-STRs (B) MDS plot based on RST genetic distances between Colombian regions and other populations, based on 21 Y-STRs, by excluding DYS385 loci

Fig. 3: MDS plot of populations comparisons. A) MDS plot based on FST genetic distances between Colombian regions and other populations, based on 23 Y-STRs (B) MDS plot based on RST genetic distances between Colombian regions and other populations, based on 21 Y-STRs, by excluding DYS385 loci

The Amazon population clusters with Peru and Ecuador, showing no significant differences, likely due to the strong Native American component shared among these groups. However, although the Amazonian population is genetically closer to Bolivia, the increased distance and significant differences relative to Bolivian natives suggest that the Andean Native component of Bolivia is distinct from the Amazonian component found in Colombia.

Y-SNP analysis and haplogroup diversity in the Andes and Caribbean regions

Haplogroup frequencies and diversity

In agreement with what was previously reported [66], seven of the 602 amplicons failed to meet the 20x threshold in ≥ 90% of the samples and were excluded from downstream analyses. Despite this, haplogroup assignment was still possible since the panel contained alternative Y-SNPs for the same haplogroup, or downstream Y-SNPs relative to those that failed.

From the 859 Y-SNPs claimed to be included in the panel, two of the Y-SNPs are in regions outside the amplified amplicons, 16 Y-SNPs were excluded for lower coverage, and 569 Y-SNPs were not variable in the samples analyzed, presenting the ancestral allele in all samples. The final dataset consisted of 272 variable Y-SNPs (i.e., with ancestral and derived states in the 175 samples), defining 64 haplogroups.

Yleaf prediction results for the 175 samples are provided in Supplementary Table S2. Five Andean samples presented an additional derived allele that has been observed within the context of other haplogroup phylogenies. One sample classified as I-M36 (AN475 in Supplementary Figure S1) presented a derived allele at M7 usually observed within haplogroup O. Two samples from haplogroup Q-M3 (AN437 and AN444), one from haplogroup R-V88 (AN471), and one from haplogroup J-L283 (AN424) presented a derived allele in variant L729, originally reported within haplogroup N [31].

These positions were all above the coverage threshold, but the majority base threshold was below 95% for three of them. Further inspection on the IGV software confirmed that these positions were not sequencing artefacts and may represent recurrent mutations (Supplementary Figure S1). Indeed, in the ISOGG Tree (https://isogg.org/tree/index.html) L729 is associated to sub-lineages inside haplogroups R (L729.2) and J (L729.2), in the UYSD (https://ysnp.erasmusmc.nl/ysnp_database/) and YFull (https://www.yfull.com/tree/) it is described within haplogroup N, and Yleaf included this mutation as diagnostic for haplogroup R1b1a1b1a1a1b1a1a. To the best of our knowledge, position M7 has not been previously reported as recurrent, and L729 has not been observed in haplogroup Q or R-V88 backgrounds. More studies will be needed to understand if these positions define novel branches within haplogroups I2a1a1, Q1b1a1a and R1b1a2.

The frequencies of the haplogroups detected in the Andean, Caribbean and Amazon regions are shown in Fig. 4A. Since only nine Y chromosomes were analyzed from the Amazon region, these samples were not used to make population inferences.

The macrohaplogroup R-M173 was the most prominent haplogroup in both Colombian regions, with frequencies of 64% in the Andean region and 39% in the Caribbean. This lineage is subdivided in R1a-M420, found in Eastern Europe, and R1b-M269, more frequent in Western and Central European populations [67–69]. Lineages inside the R1b-P312 branch predominate in both the Caribbean and the Andes, with a frequency of 85% within haplogroup R1b-M269 (Fig. 4B). This is one of the most frequent haplogroups in Western Europe and particularly in the Iberian Peninsula [67, 70]. A high diversity of R1b-P312 sub-lineages were found in our data.

Fig. 4: Phylogenetic tree of Y-haplogroups analyzed and their absolute frequencies in Colombian regions

Fig. 4: Phylogenetic tree of Y-haplogroups analyzed and their absolute frequencies in Colombian regions

A sample from haplogroup R-V88 was observed in our Colombian dataset. This lineage is more frequent in Africa [46, 71], although it has been also described in current populations from Sardinia and Corsica islands [52, 53, 71]. To determine the most likely origin of this R-V88 lineage, a network was constructed using haplotypes with 8 Y-STRs from Sub-Saharan African and European populations (Supplementary Figure S2). The Colombian haplotype shows greater affinity with African haplotypes and, therefore, its continental origin was inferred to be African.

Other European lineages were also identified in Colombia. The G sub-lineages represented 1.81% of our dataset, while haplogroup I represented 13.86%. Sub-haplogroup I showed the highest incidence in Caribbean. Haplogroup J was identified in 7.83% of the Colombia samples, and haplogroup T was the less frequent (1.2%). Although haplogroups G, I, J and T can be found in other continents [72–77], in the historical context of Caribbean and Andean populations, these lineages were most likely introduced by Europeans.

Haplogroup E was more frequent in the Caribbean (27%) than in Andean region (10.6%). In the Caribbean, the most frequent sub-haplogroup was E-M2 (Fig. 4C), widely distributed in sub-Saharan Africa and associated with the Bantu expansion [46, 78]. The presence of E-M2 lineages in Colombia most likely resulted from transatlantic slave trade, mainly from West and Central-West Africa, as well as from Mozambique on the East African coast. On the other hand, Andes exhibited a higher frequency of lineages inside sub-haplogroup E-M35 (Fig. 4C). E-M35 lineages (E-M78; E-M310 and E-M34) are frequent both in Eurasia and Africa [74, 79, 80].

Concerning clade Q, only the Q-M3* lineage was detected in the present study with a low frequency in the total dataset (6.02%). Haplogroup Q is found approximately in 85% of Native Americans [75, 81, 82]. The low frequencies observed in both the Caribbean and Andes are consistent with historical records documenting low Y-chromosomal Native ancestry, due to sex-biased admixture between Native women and non-Native men, mostly from Europe [9].

The high diversity of lineages found in our samples reflects the intense migratory flow and admixture resulting from the arrival of colonizers from the Iberian Peninsula, accompanied by enslaved Africans. Future studies with a broader analysis of the distribution of Y-STR haplogroups, including the Pacific, Amazonian, and Orinoco regions, would be relevant for a more comprehensive view of Colombian history, particularly for African lineages linked to the transatlantic slave trade and for Native American lineages associated with early migrations.

Colombian ancestry

Lineages belonging to haplogroups E-M35, G, I, J, R-M173*(xV88) and T are consistent with a predominant European paternal contribution in Colombia and were inferred to have been introduced during colonization and later reinforced by more recent European migrations into the country. Although a Middle Eastern contribution of lineages from haplogroups E, G, J and T cannot be fully excluded, considering the historical and migratory movements to Colombia, the introduction of these lineages was most likely mediated by males coming from Europe. An African ancestry was attributed to the remaining E-lineages, all of them carrying the M2 derived allele, as well as to haplogroup R-V88 detected in one sample (Fig. 4). Native American ancestry was inferred by the proportion of samples belonging to haplogroup Q-M3, and no sub-lineages were detected inside this haplogroup.

The largest proportion of paternal lineages in Andean and Caribbean regions was of European origin, 87.9% and 79.0%, respectively. In the Andes, Native American ancestry (7.6%) exceeded African ancestry (4.5%), whereas in the Caribbean, African lineages (16.0%) were more frequent than the Native American ones (5.0%).

These findings are in accordance with previous studies showing that European lineages represent the main contributor of the paternal ancestry in Colombia due to the Spanish colonization of this country [16]. The higher frequency of sub-Saharan African lineages in the Caribbean region, also previously reported by others [15], can be explained by the role of Cartagena as one of the most important ports in South America during the transatlantic slave trade.

Forensic relevant parameters

Molecular diversity parameters were recalculated using the complete haplotype dataset for Colombia, considering the regions and the department’s populations. High levels of haplotype diversity (HD ≥ 99.92%; Table 2) were observed for all regions of Colombia, except for Orinoquía. The Andean region was the only region that shared haplotypes with the others (six haplotypes with the Caribbean, one with Orinoquía, and three with Amazon). In the studied populations, the HD values were similar to those previously reported from Colombia admixed individuals analyzed with YFiler [14, 16, 17, 29, 30]. Urban populations such as those analyzed in this study, are expected to show high levels of admixture and uniparental-lineage diversity due to colonization, the slave trade, and more recent immigration waves [5, 9, 16].

Population | Y-STR haplotypes | Y-SNP haplogroups
N | Different haplotypes | Haplotype diversity | N | Different haplogroups | Haplogroup diversity
Andes | 480 | 454 | 0.9997 ± 0.0002 | 66 | 38 | 0.9068 ± 0.0323
Caribbean | 364 | 343 | 0.9997 ± 0.0002 | 10 | 41 | 0.9539 ± 0.0093
Orinoquía | 35 | 29 | 0.9832 ± 0.0131 | - | - | -
Pacific | 52 | 51 | 0.9992 ± 0.0040 | - | - | -
Amazon | 44 | 44 | 1.0000 ± 0.0048 | 9 | 6 | 0.8333 ± 0.1265
Total sample | 975 | 913 | 0.9998 ± 0.0001 | - | - | -

Given that different marker sets commonly used in forensics include Y-STRs with differential mutation rates, the probability of finding at least one haplotype difference between related individuals due to mutation depends on the average mutation rate and the number of analyzed loci. In Table 3 are the proportion of related males (separated by 1 to 5 meiosis) that are expected to present different haplotypes, considering previously reported mutation rates [83–85]. The results show that the average mutation rate among loci is higher for PPY23 compared to previous kits including fewer markers, but lower than for larger kits. The Yfiler Plus marker set shows the highest average mutation rate, reflecting the inclusion of rapid mutating markers. As expected, the probability of differentiation between related individuals increases with the inclusion of a greater number of markers.

Marker set | N | Av. mutation rate | Expected frequency of different haplotypes between individuals separated by one to five meiosis*
1 | 2 | 3 | 4 | 5
Minimal | 8 | 0.22% | 1.76% | 3.48% | 5.18% | 6.84% | 8.48%
PPY12 | 11 | 0.22% | 2.41% | 4.77% | 7.07% | 9.31% | 11.50%
YFiler | 16 | 0.27% | 4.18% | 8.18% | 12.01% | 15.69% | 19.21%
PPY23 | 22 | 0.34% | 7.29% | 14.05% | 20.32% | 26.13% | 31.52%
Yfiler Plus andSTRtyper-27 | 25 | 0.49% | 11.64% | 21.93% | 31.02% | 39.06% | 46.15%
Argus-28 | 27 | 0.47% | 12.02% | 22.60% | 31.91% | 40.10% | 47.30%
AGCU Y37 | 34 | 0.45% | 14.30% | 26.56% | 37.06% | 46.07% | 53.78%
Yfiler Platinum | 35 | 0.44% | 14.44% | 26.79% | 37.36% | 46.41% | 54.15%
PathFinder Plusand GoldenEye | 37 | 0.43% | 14.66% | 27.17% | 37.85% | 46.96% | 54.74%