Work overview

Section 04 of 09

Y-SNP genotyping

A comprehensive analysis of Y-chromosomal diversity in Colombian populations

Julyana Ribeiro, Zehra Köksal, Adriana Castillo, Adriana Ibarra, Beatriz Martinez, Humberto Ossa, Masinda Nguidi, Juan David Granda, Verónica Gomes, Pedro Rodrigues, Maria Inês Machado, Vania Pereira, and Leonor Gusmão · 2026

Contents

Section 04 of 09

  1. 01Introduction
  2. 02Materials and methods
  3. 03Y-STR genotyping
  4. 04Y-SNP genotyping
  5. 05Results and discussion
  6. 06Comparison of Y-STR haplotypes in Colombian populations
  7. 07Comparisons of admixed South American populations
  8. 08Conclusions
  9. 09Supplementary Information
Text size
Work overview

Section 4 of 9

Y-SNP genotyping

Julyana Ribeiro, Zehra Köksal, Adriana Castillo, Adriana Ibarra, Beatriz Martinez, Humberto Ossa, Masinda Nguidi, Juan David Granda, Verónica Gomes, Pedro Rodrigues, Maria Inês Machado, Vania Pereira, and Leonor Gusmão · about 3 minutes

A total of 175 samples were randomly selected from Andes (n = 66), Caribbean (n = 100) and Amazon (n = 9) regions. This subset is representative of the Andean and Caribbean regions, where more than 90% of the population is concentrated [4], with differences expected compared to other regions, based on Y-STR results. Samples were genotyped for Y-SNPs using the Ion AmpliSeq™ HID Y-SNP Research Panel v1 [31] developed for massively parallel sequencing in the IonS5 System (Thermo Fisher Scientific). The panel contains 884 Y-SNPs located in 602 amplicons with an average amplicon size of 130 bp [31].

DNA amplification was performed by combining 4 µL of Ion AmpliSeq™ HiFi Mix, 10 µL of Ion AmpliSeq™ primer pool (2X), 1–5 µL of gDNA template (1-5ng) and the remaining volume of water to complete the 20 µL PCR reaction. A negative control was used in each PCR batch to confirm the absence of contamination. Library preparation was performed according to the manufacturer’s protocol.

Eluted libraries were quantified using the TaqMan™ Library Quantitation Kit (Thermo Fisher Scientific) and pooled to a final concentration of 50–80 pM. Libraries with concentrations below 80 pM were re-amplified with 8 cycles according to the manufacturer’s manual (Precision ID Library Kit on the Ion Torrent system by Thermo Fisher Scientific Pub. No. MAN0015802) to guarantee enough genomic material for normalization.

Template preparation and chip loading were performed on the Ion Chef™ Instrument (Thermo Fisher Scientific). 32 libraries were loaded into each Ion 530™ Chip (Thermo Fisher Scientific). Sequencing was performed on the Ion PGM S5™ instrument (Thermo Fisher Scientific) with 650 run flows, using Ion S5™ Sequencing Kit (Thermo Fisher Scientific).

Statistical analysis

Data analyses were performed after grouping data from this study with published data from Colombia. From the available studies, we have selected those from admixed populations (excluding Native or Afro-descent communities) with complete haplotype profiles for the Y-Filer STRs (Supplementary Table S1) [14, 16, 17, 29, 30].

The NEVGEN software (https://www.nevgen.org/) was used to predict haplogroups from PPY23 haplotypes in samples with intermediate, duplicated, or null alleles. In cases where NEVGEN attributed a probability of less than 90%, the haplogroup assignment was confirmed using previously described SNaPshot multiplexes [34].

For Y-SNP data, coverage statistics for the 602 target regions were generated using the Coverage Analysis plugin v5.10.0.3. A minimum coverage threshold of 20 reads per amplicon was applied for quality trimming. The Y-leaf software version 3.0 [35] was run for Y-SNP variant calling and to infer the haplogroups using BAM files. The threshold criteria for the haplogroup inference were set as follows: minimum of 20 reads for each base, quality threshold of 20 for each base read, and majority base threshold of 95%. The Integrative genomics viewer (IGV) software was used to inspect ambiguous variation [36].

Haplogroup frequencies were calculated by direct counting. The haplotype/haplogroup diversities (HD) and the total number of observed and unique haplotypes were calculated using the Arlequin software v3.5.1.2 [37]. The same software was used in population differentiation analyses of Colombian samples, based on pairwise genetic distances (_F_ST and _R_ST), and in the analysis of molecular variance (AMOVA). In multiple tests, the significance level of 0.05 was adjusted by sequential Bonferroni procedure [38].

In pairwise genetic distance analyses, the number of repeats in DYS389I was subtracted in DYS389II. Moreover, for _R_ST analyses DYS385 was excluded, and microvariant alleles, duplications, and null alleles were coded as “?”.

Pairwise genetic distances (_F_ST and _R_ST) based on 23 Y-STR haplotypes were also calculated between our population samples and populations from Europe and Africa, as well as Native American and other admixed populations from South America (Supplementary Table S1) [39–44].

Multi-Dimensional Scaling (MDS) plots of pairwise genetic distances were generated with the software STATISTICA v.14.0.0.15 (TIBCO Software Inc.).

A Y-STR-based phylogenetic network was built for R-V88 chromosomes from African and European populations [45–53] for haplotypes defined by 8 loci shared with the available datasets and using the software Network 4.6.1.1 (Fluxus-Engineering), with the median-joining algorithm.