Work overview

Section 01 of 03

Methods and Materials

A Genome-Wide Association Study of Premenstrual Symptoms in Two Nordic Populations

Elgeta Hysaj, Piotr Jaholkowski, Alexey A. Shadrin, Jacob Bergstedt, Yi Lu, Elizabeth Bertone-Johnson, Cynthia M. Bulik, Mikael Landén, Sven Sandin, Kaarina Kowalec, Sara Hägg, Arianna Di Florio, David Goldman, Peter J. Schmidt, Unnur A. Valdimarsdóttir, Ole A. Andreassen, and Donghao Lu · 2026

Contents

Section 01 of 03

  1. 01Methods and Materials
  2. 02Results
  3. 03Discussion
Text size
Work overview

Section 1 of 3

Methods and Materials

Elgeta Hysaj, Piotr Jaholkowski, Alexey A. Shadrin, Jacob Bergstedt, Yi Lu, Elizabeth Bertone-Johnson, Cynthia M. Bulik, Mikael Landén, Sven Sandin, Kaarina Kowalec, Sara Hägg, Arianna Di Florio, David Goldman, Peter J. Schmidt, Unnur A. Valdimarsdóttir, Ole A. Andreassen, and Donghao Lu · about 8 minutes

Study Population

We conducted a GWAS of 72, 297 European-ancestry women nested from the LifeGene cohort in Sweden and the MoBa (The Norwegian Mother, Father and Child Cohort Study) cohort in Norway.

LifeGene is a large-scale Swedish prospective cohort launched in 2009 with longitudinal follow-ups (22). It enrolled 39,862 people (24,265 women) ages 18 to 50 years who were randomly selected from the Swedish population and their household members. A thorough web-based questionnaire for collecting information on lifestyle, physical, mental, and social well-being was administered at baseline and in 5 annual follow-up cycles. Blood samples were collected during the in-person testing at baseline. Participants were linked to the national population and health registers using their unique Swedish personal identification number, a lifelong identifier assigned at birth or upon immigration to Sweden. Informed consent was obtained either electronically from all participants upon registration online or in writing at the in-person testing center. The current study was approved by the Swedish Ethical Review Authority (2021-02775).

MoBa is a population-based pregnancy cohort study conducted by the Norwegian Institute of Public Health (23, 24, 25, 26). Participants were recruited from across Norway from 1999 to 2008. Women consented to participation in 41% of pregnancies. The cohort includes approximately 114,500 children, 95,200 mothers, and 75,200 fathers. The current study is based on version 12 of the quality-assured data files released for research in January 2019. The establishment of MoBa and initial data collection were approved via a license from the Norwegian Data Protection Agency and after review of the Regional Committees for Medical and Health Research Ethics. The MoBa cohort is currently regulated by the Norwegian Health Registry Act. The current study was approved by the Regional Committees for Medical and Health Research Ethics (2016/1226/REK).

Assessment of PSs

Both questionnaire assessments and nationwide health care registers were sourced for case assessment. Cases with PSs were defined as either having a clinical diagnosis of PMD recorded in the registers or having met the criteria for a probable PMD based on self-report questionnaires (as described in detail in LifeGene). The controls were women with no clinical diagnosis of PMDs in registers and not meeting the PMD criteria during any available questionnaire cycles.

LifeGene

A modified version of the Premenstrual Symptom Screening Tool (PSST) (27) was used to assess premenstrual symptoms at baseline and annual follow-ups for 5 years (28). The original PSST has been validated with a sensitivity of 79% (29). The PSST was modified to start with 3 screening questions: 1) “During most menstruation cycles during the last year, have you experienced mood changes and/or physical symptoms during the week before menstruation?” 2) “Have your premenstrual symptoms been so severe that they have affected your relationships with others or your ability to perform work or other activities?” and 3) “Are you absolutely certain that the symptoms are limited to the premenstrual period, meaning that you are always completely symptom-free approximately a week after menstruation begins?” Upon confirmation of all screening questions, participants were prompted to rate the severity of 15 physical and affective symptoms from 1 (none), 2 (moderate), 3 (considerably severe), to 4 (severe). As described elsewhere (28), participants were classified as cases if they met 1) ≥1 of 4 affective symptoms rated as considerably severe to severe and 2) ≥4 other symptoms rated as moderate to severe.

To complement the questionnaire assessment in LifeGene, for the Swedish participants, we further identified clinical diagnoses of PMDs, as described elsewhere (7). According to the Swedish guidelines, a clinical diagnosis of PMD should be based on prospective daily symptom ratings for at least 2 consecutive menstrual cycles (30). Briefly, we identified PMD diagnoses based on ICD codes (Table S1) from the National Patient Register (NPR) (1987–2023) and the Stockholm Primary Care Register (2001–2021), since 82% of the participants lived in Stockholm County. Primary care data were unavailable for residents in other counties; we also obtained information on filled prescriptions for antidepressants (Anatomical Therapeutic Chemical codes: N06AB, N06AX, N06AA) and hormonal contraceptives (G03A) with a written indication for PMD treatment from the National Prescribed Drug Register (2006–2023).

The Norwegian Mother, Father and Child Cohort Study

Questionnaire assessment was based on the women’s responses to the following questions at the 15th week of gestation: “Are you usually depressed or irritable before your period?” and “If yes, does this feeling disappear after you get your period?” As described elsewhere (31), the cases were defined based on the response “yes, noticeably” or “yes, very much” to the first question and “yes” to the second question. We excluded individuals whose symptoms of PMD did not resolve after the onset of menses.

Clinical diagnoses of PMDs were derived from the Primary Care Registry of Norway (2006–2023), which contains diagnoses given at the primary care level. We identified cases based on the Premenstrual Tension Syndrome (X89) diagnosis according to the International Classification of Primary Care, Second Edition (32).

Genotyping, Quality Control, and Imputation

LifeGene

In LifeGene, DNA was extracted from blood samples collected at baseline. Participants from the LifeGene cohort were genotyped by 4 substudies including the current study (Table S2). Quality control (QC) was performed using the Ricopili bioinformatics pipeline for each substudy (33). Briefly, these steps included a call rate threshold of ≥0.98 for cases and controls, heterozygosity (FHET) within ±0.20, and exclusion of sex mismatches. Single nucleotide polymorphism (SNP) QC required a call rate of ≥0.98, missingness difference ≤0.02, and minor allele frequency (MAF) ≥0.01. Ungenotyped SNPs were then imputed based on the Haplotype Reference Consortium (HRC) reference panel (r1.1) (34) via the Sanger imputation server and pooled after removing duplicated individuals. Information on a total of 7,135,674 SNPs was available for analysis. Admixture analysis was performed using the ADMIXTURE software to estimate genetic ancestry proportions, leveraging reference populations from the 1000 Genomes Project (phase 3 v5) (35). Analysis was restricted to autosomal variants. Individuals estimated with <90% probability of European ancestry were excluded. Briefly, genotype data were available for 8826 women (36%); after excluding 862 participants who were first-degree relatives (having identity-by-descent sharing ≥0.2), had significant non-European ancestry (n = 896), or had no information on phenotype (n = 1839), 5229 participants were included in this analysis.

The Norwegian Mother, Father and Child Cohort Study

In MoBa, venous blood was collected from women at approximately the 15th week of gestation and immediately after giving birth. Genomic DNA was extracted and stored at the Norwegian Institute of Public Health (36). The MoBa cohort genotyping was conducted through multiple research projects over several years (37). A novel family-based pipeline (MoBaPsychGen genotype QC pipeline) was implemented to handle the relatedness structure of the MoBa dataset, while appropriately accounting for the differences resulting from array and batch effects (37). The pipeline (38) includes preimputation QC, phasing, imputation, and postimputation QC and prioritizes retaining individuals over SNPs. After QC procedure, 6,981,748 SNPs were available for further analysis. Our analysis was restricted to individuals of European ancestry, selected based on visual comparison of the first 7 genetic principal components (PCs) with PCs from 1000 Genomes phase 1 unrelated samples.

Statistical Analysis

Genome-Wide Association Study

Characteristics of individuals with and without PSs (age, educational level, civil status, history of depression or anxiety disorder, age at menarche, and parity) in LifeGene and MoBa were summarized using descriptive statistics. p Values were calculated using χ2 tests.

In LifeGene, GWAS was performed using logistic regression using PLINK2 (39). Analysis was restricted to 6,508,434 genetic variants with an MAF ≥0.01 and imputation quality score (INFO) ≥0.90. In MoBa, logistic regression was used for GWAS analysis using Regenie (version 3.2.5) (40). We excluded SNPs with an imputation info score <0.80 and SNPs with minor allele count <20. For both cohorts, the estimates were adjusted for substudy membership and the top 10 PCs in model 1. To understand the potential pleiotropic effect of the top loci on other common psychiatric conditions among women with PMD, model 2 was also adjusted for history of clinically diagnosed depression and anxiety disorder (ICD codes in Table S1). Briefly, we identified depression and anxiety diagnoses recorded before PS assessment or diagnosis from the Swedish NPR in LifeGene and the Norwegian NPR in MoBa. In an additional analysis, we restricted analysis to cases confirmed by both questionnaire assessment and clinical diagnosis.

We generated the quantile-quantile (QQ) plot using the qqman (41) package in R version 4.2.2 (2022-10-31). Summary statistics from both LifeGene and MoBa were then meta-analyzed using inverse variance-weighted analysis as implemented in METAL (version 2020-05-05) (42). SNPs with p value < 5 × 10−8 were considered genome-wide significant and with p value < 1 × 10−6 were considered borderline significant. The odds ratio (OR), 95% CI, and p value were reported for any independent SNPs (LD _r_2 < 0.6) above marginal significance (i.e., lead SNPs). Functional annotation was conducted using FUMA (43) integrating expression quantitative trait loci data from brain and blood tissues.

SNP-Based Heritability and Genetic Correlation

Based on the summary statistics, we estimated the SNP-based heritability on the observed scale and on the liability scale (for a prevalence of 0.2) of PSs using the linkage disequilibrium score regression (LDSC) software package (version 2.0.1) (44). Several studies have suggested that PMDs are linked to a broad range of other disorders and traits (31,45,46). We used LDSC to undertake genetic correlation analyses with major psychiatric traits (e.g., major depression), gynecological conditions (e.g., endometriosis), sex hormones, and known risk factors (e.g., age at menarche) for PMDs as described in Table S3. We also examined genetic correlations with circulating hormone levels (sex hormone binding globulin, total and bioavailable testosterone), hypothesizing a null finding because PMDs are more likely related to differential sensitivity to hormonal fluctuations than to variants in hormone levels. Briefly, we used GWAS summary statistics from the UK Biobank, where the hormone levels were measured in serum samples collected at baseline from female participants (no information on menstrual phase) using standardized immunoassays, and bioavailable testosterone was calculated using the Vermeulen equation.

Additional Analysis

We performed complementary analyses to evaluate the consistency of genetic signals across the LifeGene and MoBa cohorts. First, we conducted cross-cohort polygenic prediction analyses by testing whether polygenic scores derived from MoBa summary statistics predicted probable cases in LifeGene across multiple association thresholds (p < .05, .005, and .0005). Second, we estimated the genetic correlation between the 2 cohorts to assess shared genetic architecture. Finally, consistency of effect estimates across cohorts was formally evaluated using heterogeneity statistics from the meta-analysis.