Section 3 of 7
Methods
Nuria Domínguez-Pérez, Irene Sevilla-Arrabal, Beatriz Navarro-Brazález, María Torres-Lacomba, and Javier Courel-Ibáñez · about 9 minutes
For the umbrella review, we identified and screened systematic reviews and meta-analyses according to predefined eligibility criteria focused on UI prevalence in physically active or athletic women populations. From eligible reviews, we extracted the cited primary observational studies reporting prevalence data. To ensure a comprehensive and updated synthesis of the available evidence, we also applied a snowball search strategy (forward citation tracking) to locate additional updated primary studies not captured in the original systematic reviews. For the meta-analysis, all primary studies meeting the inclusion criteria were subjected to independent quality appraisal and then included in a new meta-analysis to generate updated prevalence estimates. This dual-layered approach allowed us to summarise the synthesised evidence while expanding and refining it using original primary data. The study forms part of a larger overview project exploring the prevalence of PFD in nulliparous women athletes. The protocol was prospectively registered in PROSPERO (#CRD420251055357) and follows the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) reporting guidelines [31, 32].
Literature Search
Records were retrieved from five electronic databases: MEDLINE (via PubMed), Scopus, Web of Science, CINAHL (via EBSCOhost) and the Cochrane Library. The search was limited to reviews published from January 2010 to 29 September, 2025. Search strategies for each database combined Medical Subject Headings (or equivalent) and relevant free-text terms related to UI, athletic populations and prevalence. Forward citation tracking (via Google Scholar) was manually screened on all included systematic reviews to identify additional primary studies. The full search strategy was developed as part of a broader overview-of-reviews project on the prevalence of PFD in women athletes and is available in Appendix 1 of the Electronic Supplementary Material (ESM). Analyses were restricted to adult women athletes (aged 18–45 years) to reduce developmental and hormonal heterogeneity, as adolescence and perimenopause represent distinct life stages with respect to pelvic floor health and training exposure [25].
Eligibility Criteria for Systematic Reviews and Meta-analyses
Study Population
Adult women athletes aged 18–45 years. Para athletes were excluded.
Exposure
Engagement in regular sport participation (training and/or competition), irrespective of sport modality or performance level.
Outcomes and Instruments
Prevalence of any type of UI, (overall UI, stress UI [SUI], urgency UI [UUI], mixed UI [MUI] and exercise-related/exertional UI [EXUI]), assessed using any measurement approach (e.g. validated questionnaires, clinical assessment, self-report).
Study Design
Systematic reviews or meta-analyses.
Publication Status, Language Restrictions and Timeframe
Studies published in peer-reviewed academic journals, in English, French, Portuguese or Spanish (reflecting the language proficiency of the review team), published after 1 January, 2010 to reduce redundancy and excessive overlap. Conference abstracts, theses and non-peer-reviewed reports were excluded.
Eligibility Criteria for Primary Studies
Study Population
Adult women athletes aged 18–45 years. Para athletes were excluded.
Exposure
Participation in a specific sport discipline (e.g. soccer, CrossFit, rugby, cycling). Studies were included when the sport discipline was clearly reported, even if the specific specialty was not specified (e.g. athletics or track and field, gymnastics, dance). Studies were excluded when sport participation was reported only at a broad or non-specific level (e.g. collective sports, endurance sports, or physically active) without discipline-level classification, as such classifications were not sufficiently granular for quantitative synthesis.
Outcomes and Instruments
Prevalence of any type of UI, eligible for quantitative synthesis if measured with (a) validated questionnaires, (b) clinician-administered structured assessments (e.g. standardised clinical screening, pad tests) or (c) clearly defined self-report items aligned with the International Continence Society definitions (with an explicit recall period). Studies using lifetime/unspecified recall or not validated questionnaires were included but flagged and examined in sensitivity analyses. Where available, we extracted subtype-specific prevalence (SUI, UUI, MUI and EXUI).
Study Design
Prospective or retrospective cohorts, case–control, cross-sectional or intervention studies.
Publication Status, Language Restrictions and Timeframe
Studies published in peer-reviewed academic journals, in English, French, Portuguese or Spanish (reflecting the language proficiency of the review team), with no restriction on year of publication. Conference abstracts, editorials, commentaries, dissertations and other forms of grey literature were excluded.
Study Selection
Data management was conducted using the systematic review management software Rayyan. Duplicate entries were identified and removed using Rayyan’s semi-automated duplicate detection tool, followed by manual verification to ensure accuracy. Two reviewers (ISA and NDP) independently screened titles and abstracts against predefined inclusion and exclusion criteria. Full texts were then retrieved and assessed for eligibility. Disagreements at any stage were resolved through discussion or, if needed, by a third reviewer (JCI). Additional studies identified through forward citation tracking were screened using the same process.
Data Extraction
A standardised data extraction form was developed a priori and piloted on a subset of studies to ensure clarity and consistency. Data extraction was independently performed by two reviewers (ISA and NDP), with disagreements resolved through discussion or, if needed, adjudicated by a third reviewer (JCI). Authors were contacted to clarify missing or unclear data when necessary.
For systematic reviews and meta-analyses included in the umbrella review, extracted data comprised study characteristics (authors, year and design), population characteristics (number of included studies, sample size, age range, parity status, level of competition and sport discipline/exposure, when reported), and UI-related outcomes. Reviews that did not restrict inclusion criteria by parity status, competitive level or sport discipline/exposure were classified as mixed.
For primary studies included in the meta-analysis, extracted data included study characteristics (authors, year, country, design), participant characteristics (sample size, age, parity status, race/ethnicity and socioeconomic status), exposure details (sport discipline, sport modality, competitive level), weekly training volume, UI instruments and quantitative prevalence (events and total sample). The primary outcome variable was UI and UI subtype prevalence, extracted as the number of UI events and total sample size. Instruments were coded as Validated (named, referenced tool) versus International Continence Society Criteria aligned item (single/multi-item, not validated, maps to International Continence Society Criteria ICS definition) for sensitivity analyses. Exposure and classification variables extracted or derived for analysis included were classified as follows: (1) Sport discipline, as reported in the original study; (2) Sport modalities, defined as groups of sport disciplines sharing common biomechanical and physiological characteristics and classified into Technical, Endurance, Aesthetic, Weight-dependent, Ball Games, Power and Gravity sports [33]; (3) Sport impact, based on the mechanical load transferred to the pelvic floor through increased intra-abdominal pressure and ground reaction forces, according to pelvic floor-specific frameworks and consensus [34, 35]: High-impact sports (e.g. gymnastics, basketball, volleyball, high jump, trampoline, powerlifting), Medium-impact sports (e.g. tennis, running, karate, soccer), Low-impact sports (e.g. swimming, cycling and walking, where pelvic floor strain is minimised because of limited ground impact or absence of explosive movements); (4) Competitive level, conceptualised as a proxy for long-term cumulative exposure and training context rather than performance status alone, based on training volume, performance level and competitive context [36]: Amateurs (individuals engaged in structured training without elite or national-level competition), High-level group (athletes competing at the national level with structured training and competition schedules; this category also included athletes with at least 3 years of sub-elite competitive involvement), and Professionals (athletes competing at the highest national or international levels, training in elite or Olympic centres or explicitly identified as elite by the original study authors); and (5) Training volume, extracted as mean or median hours per week as the central tendency [37].
Quality Assessment
The methodological quality of included systematic reviews and meta-analyses was evaluated using the AMSTAR 2 (A Measurement Tool to Assess Systematic Reviews) instrument [38]. AMSTAR 2 evaluates 16 methodological domains, including protocol registration, comprehensiveness of the literature search, duplicate study selection and data extraction, risk-of-bias assessment of primary studies, appropriateness of meta-analytic methods and consideration of risk of bias when interpreting results. Overall confidence in review findings was determined based on the presence and severity of methodological limitations across domains, classified into three categories: Non-critical weaknesses: Items 1, 3, 5, 6, 8, 10, 14 and 16, Critical weaknesses: Items 2, 4, 7, 11 and 15, and Critical flaws: Items 9, 12 and 13. Reviews presenting one or more critical flaws were rated as having critically low confidence. Reviews without critical flaws but presenting at least one critical weakness were rated as having low confidence. Reviews with no critical weaknesses or flaws were initially rated as having moderate confidence; however, the presence of multiple non-critical weaknesses was considered sufficient to further downgrade confidence from moderate to low when appropriate. Reviews with no or only one non-critical weakness were rated as having high confidence.
For primary studies included in the meta-analysis, methodological quality was evaluated using the Appraisal Tool for Cross-sectional Studies (AXIS) [39], a 20-item checklist developed for observational research. AXIS assesses three broad domains: reporting quality (clarity of objectives, methods and results), study design quality (sampling procedures, sample size justification and handling of non-responders) and risk of bias (validity of outcome measurement, consideration of confounding and appropriateness of statistical analyses). Item fulfilment was coded as Yes (1 point), Partial (0.5 points) or No (0 points), and total AXIS scores were used to classify studies as high quality (≥ 16.5 points), moderate quality (13.0–16.0 points) or low quality (< 13.0 points).
Small-study effects were assessed when k ≥ 10 using funnel plots and Egger’s test. To avoid double-counting of primary studies, we identified systematic reviews with non-overlapping evidence for each outcome. When multiple reviews were available, we calculated the corrected covered area to quantify overlap [40]. If overlap was high (corrected covered area > 10%), older or lower-quality reviews were excluded to ensure that each primary study contributed only once to the synthesis.
Statistical Analysis
Each study was rated independently by two reviewers (ISA, NDP), with disagreements resolved through discussion (JCI). Data from primary studies were aggregated to one cohort per study before pooling (k denotes cohorts). Random-effects models pooled logit-transformed prevalence (PLOGIT) with 95% confidence intervals (CIs) and 95% prediction intervals; heterogeneity was quantified by _τ_2, _I_2 and Cochran’s Q. Pre-specified subgroup analyses compared sport discipline, modality, impact and competitive level; only subgroups with three or more cohorts were meta-analysed. Subgroup differences were tested with the _χ_2 test for subgroup differences (unadjusted comparisons of pooled subgroup means). We also fitted random-effects meta-regressions with moderators entered categorically (QM). Because these approaches use different models and weighting, they may yield different inferences; meta-regression was considered primary. Pairwise comparisons used re-levelling, reported as log-odds with 95% CIs and odds ratios (ORs). Moderator pseudo-_R_2 was the proportional reduction in _τ_2 versus the null model. Training volume was entered as continuous predictors, reported as OR per 1 standard deviation (SD) and per unit (h·week⁻1). Egger’s test was applied only where k ≥ 10. Analyses were run in R (Version 4.1.2).
Sensitivity Analysis
To verify the robustness of pooled and subgroup estimates, we re-ran the entire analysis pipeline under two prespecified restrictions: (1) validated instruments only, retaining studies that explicitly used standardised UI tools (e.g. ICIQ-UI SF, UDI-6), and (2) nulliparous samples only, retaining studies that reported data exclusively for nulliparous women. Across sensitivity runs, we report whether the direction and magnitude of the global prevalence, subgroup contrasts and continuous associations remained consistent with the primary analyses.