Section 3 of 7
Methods
Hunter Bennett, Henry Blake, Noah d’Unienville, James Murray, and Jordan Fox · about 6 minutes
This study employed a cross-sectional design. Original research articles published between 1 January and 31 December 2024 in journals in the SCImago category “Sport Sciences” were eligible for inclusion. Articles were selected from journals across all four quartiles, aiming for a similar number of articles from each quartile. SCImago calculates journal quartiles on the basis of the SCImago Journal Rank (SJR) indicator, which reflects both the number of citations received by a journal and the prestige of the journals where those citations come from. Within each subject category, journals are ranked by SJR score and then divided into four quartiles. Once the journals were identified, Altmetric Explorer was used to identify all articles published from those journals within the specified date range, and to extract their titles, author lists, and DOIs. Journals were chosen at random. However, to account for variations in publication numbers (expecting higher-quartile journals to publish more articles), in the first round of data extraction, articles were extracted from three Q1 journals, four Q2 journals, five Q3 journals, and five Q4 journals. To balance the numbers between quartiles, additional articles were then extracted from one Q1 journal and one Q3 journal (total = 19). Any non-original research articles (i.e., reviews, editorials, consensus statements, letters to the editors, etc.) were excluded, as were articles that implemented a Delphi or qualitative design, as these were considered less likely to have a formal hypothesis.
Preregistration
This study was prospectively preregistered on the Open Science Framework (https://osf.io/x7eju/overview).
Extracted Outcomes
The author guidelines of all included journals were examined to describe what recommendations they provide around the inclusion of a hypothesis, and requirements for preregistration (Supplementary Digital Content 1), and to support the forthcoming discussion. Data pertaining to preregistration status, hypothesis testing, and hypothesis testing results were also extracted (Table 1). Finally, the preregistrations of included studies were examined to assess whether the stated hypotheses were consistent with those in the published article.
Was the study preregistered (identified via a statement within the manuscript)? | Yes/no
Did the study have a formal hypothesis (defined as a statement in which the article specifically predicted a testable outcome, rather than posing a general research question or aim) provided in the Introduction or Methods? | Yes/no
If no formal hypothesis was provided
Did the study state its analysis was exploratory, or did it clearly state no formal hypothesis was provided? | Yes/no
If formal hypothesis was provided
Did the results of the study align with the hypothesis (yes/no/partially [partially being considered where multiple hypotheses are presented without a clear “primary” hypothesis, and some are supported, and some are not]/unclear [results are not provided for the hypothesized outcome])? | Yes/no/partially/unclear
Study Procedures
Altmetric Explorer was used to identify all relevant data for articles published within the eligible timeframe from the selected journals, which were then downloaded to a custom Microsoft Excel spreadsheet for eligibility assessment and data extraction. All eligible articles were independently examined by the lead author to extract all outcome measures (n = 2006). Once initial data extraction was complete, all included articles were cross-checked by one other member of the research team to evaluate extraction accuracy (JF = 670; JM = 683; ND = 347; HB = 306). Any discrepancies were resolved by consensus among the authorship team. Of the 6018 data items extracted (3 items per study), 176 discrepancies were identified (97.1% agreement), all of which were resolved through consensus. Finally, for each article that reported at least one hypothesis, the total number of hypotheses within the paper and the number that were supported were recorded and reported descriptively.
Statistical Analysis
To test the first hypothesis, the following data were reported descriptively: the number of total studies examined (across the whole cohort and each quartile), and the proportion that were preregistered and not preregistered (across the whole cohort and each quartile). Then, a one-sample binomial test comparing the observed proportion of preregistered studies with the lower bound of the hypothesized range (15%) was conducted to evaluate whether the observed rate differed significantly from the expected proportion. To test the second hypothesis, the proportion of supported hypotheses between preregistered and non-preregistered studies were compared using a chi-squared test of independence. Cramér’s V effect size was also reported to provide insight into the extent to which any differences are practically meaningful, and interpreted using the following thresholds: < 0.1 = negligible, < 0.2 = weak, < 0.4 moderate, < 0.6 = relatively strong, < 0.8 = strong, and > 0.80 very strong [29]. To test the third hypothesis, a logistic regression was conducted to assess whether preregistration frequency and the proportion of supported hypotheses varied across journal quartiles. For preregistration, the dependent variable was preregistration status (yes/no), with journal quartile (Q1, Q2, Q3/Q4) as the independent variable. For hypothesis support, the dependent variable was supported hypothesis (yes/no), with journal quartile as the independent variable. For this analysis, effect sizes were quantified using odds ratios (OR) and considered trivial (0.77–1.00 or 1.00–1.29), small (0.51–0.78 or 1.30–1.99), moderate (0.25–0.50 or 2.00–3.99), and large (≤ 0.24 or ≥ 4.00) [30]. All analysis was conducted in R studio (version 4.3.1). All data and code can be found on the Open Science Framework project linked to this study ( https://osf.io/c8jfm/files/osfstorage).
Assuming that a ~ 10% difference in successful hypothesis rates is practically meaningful, and using the ~ 80% positive hypothesis rate presented in prior research [14], we aimed to detect a difference between proportions of 80% and 70% with an alpha of 0.05 and 80% power. The sample size calculation was based on a two-proportion chi-squared test (independent groups) in G*Power (version 3.1.9.7). Expecting a sampling ratio of 1:6 between preregistered and non-preregistered studies (i.e., ~ 20% of studies being preregistered), we determined that a minimum of 142 preregistered and 708 non-preregistered studies would be required (total n = 850). We note that the a priori power calculation was based on the chi-squared comparison. As such, the logistic regression analyses may be at greater risk of false negatives.
Deviations from Planned Protocol
It was initially proposed that for all formal analysis examining differences between quartiles, quartiles 1–4 would be analyzed independently. However, owing to there being much fewer articles published in Q4 journals than initially anticipated, Q3 and Q4 were combined for analysis. For complete transparency, the preregistered analysis was still conducted, and the results are provided in Supplementary Digital Content 2. Furthermore, the decision to conduct the exploratory descriptive analysis was made after preregistration occurred. Additionally, early data extraction identified a lower preregistration rate than hypothesized, and therefore, to ensure the minimum number of preregistered studies was met, a larger total number of studies was included. Lastly, rather than crosschecking all data, it was initially proposed that 10% of articles would be crosschecked to obtain a measure of extraction accuracy. However, this initial check resulted in 96% agreement between authors, which the authors believed could result in categorization errors for a notable number of the 2006 articles identified. These all related to whether a hypothesis was “fully” or “partially” supported, and occurred when a study had multiple outcome measures contribute to a single hypothesis and was on a topic outside the research team’s direct area of expertise. As such, the authors opted to cross-check all data for discrepancies, which were then resolved through consensus agreement to enhance data extraction and categorization accuracy.