Section 4 of 5
Methods
Alistair G. Auffret, Emma Ladouceur, Natalie S. Haussmann, Petr Keil, Eirini Daouti, Tatiana G. Elumeeva, Ineta Kačergytė, Jonas Knape, Dorota Kotowska, Matthew Low, Vladimir G. Onipchenko, Matthieu Paquet, Diana Rubene, and Jan Plue · about 11 minutes
The global seed bank database
To study the effects of spatial scale, habitat type and quality on soil seed banks, we used a global database of seed bank richness, abundance and density31. Briefly, the database is based on a Web of Science search using the search term “seed bank*” OR seedbank*, including papers included in the Web of Science until the end of 2019 (a few papers returned by the search were subsequently published in journals in 2020). All resulting papers were checked for being both community seed bank studies (i.e., taking soil samples and counting and identifying all seeds, rather than focussing on one or more specific species), and containing information about species richness (number of species), seed abundance (number of seeds), and/or seed density (number of seeds per unit area or volume) in the seed bank. Most of the papers were in English, but data were also extracted from papers published in, for example, Spanish and Portuguese (often South American studies), as well as French and German. Due to a large data gap, the database was supplemented by an additional database and literature search featuring studies published in Russian and carried out in a number of former Soviet states. In total, this resulted in 7224 unique studies to be checked, and 1442 studies from which records could be extracted or obtained. A record refers to a row in the database, and represents a known number of soil samples with specified (by the study authors) geographical coordinates. Depending on the detail in which study location and design, and seed bank results were reported in a specific study, a record can represent one or more locations or environments. Each record can also provide one or more commonly recorded seed bank metrics (mainly total species, total number of seeds, seeds m−2 density), which we term observations. A study can provide multiple observations by reporting more than one metric, and/or by (for example) reporting metrics across several distinct records (i.e. different locations or environments). In total, the 1442 studies provided 8087 observations across 3096 records (Fig. 1, Tables 1 and S1). The database includes soil seed bank data from 94 countries and all seven continents. Full details regarding which data were extracted and how are available in Auffret et al.31. Below, we describe only the variables in the database that were used for this analysis.
Summary information and data variation in the global seed bank
From the 1185 studies and 2585 observations contributing values of seed density, the total number of seeds or seedlings counted across all studies was 35,932,626. Five studies (or 0.2% of all observations) found no seeds within samples. This included an investigation of aquatic systems in Canada79, deserts and xeric shrublands in Egypt80, temperate forests in Germany and the Netherlands81,82, and tundra in Antarctica83. Ten studies observed only one species in their seed bank samples, while the maximum was 246 species in 88 m2 of soil across 156 arable sites in the Midwest of the United States84. The minimum (non-zero) observed density of seeds in the seed bank was 3 seeds m−2 in a tropical forest in South Africa85, with the maximum being 501,119 seeds m−2 observed in a study from temperate forests in Germany86. The minimum number of seeds greater than zero recorded in a soil seed bank observation was two seeds in two studies in temperate forests in the United States, and tropical and subtropical forests in Indonesia with 0.67 seeds m−2 and 0.16 m−2 of soil sampled, respectively87,88, with the maximum being 5,907,132 seeds in samples taken in arable systems in 462 m2 soil from 924 samples across 22 sites in the United Kingdom89.
Naturally, such raw maximum and minimum values are at least partly an outcome of each investigation’s scientific focus and study design. Indeed, there was a large variation in sampling effort across records. The minimum total number of samples was one sample, used in 19 studies, with the maximum being 114,000 samples taken in arable systems representing 224 m2 of soil sampled across 228 sites in China90. The minimum total sampled area was 0.0015 m2 (5 sites, 15 total samples) in temperate broadleaf and mixed forests from the United States91, with the maximum being 816 m2 (20 sites, 20,400 samples) in another large arable study in the UK92. The minimum number of sites was one, representing the majority of our samples (1832 observations from 692 studies), with the maximum being 595 sites (total sampled area of 59.5 m2) in arable fields across Finland93.
Modelled variables
From the global seed bank database, we extracted the two most common metrics from the soil seed bank literature as response variables: [1] the total number of species recorded in a reported area soil seed bank sampled; [2] the density of seeds m−2 in the soil. Soil seed bank studies often involve the collection of multiple samples of soil with a soil corer. These samples are then usually pooled, in order to account for the spatial patchiness of species and seeds in the soil72,94,95. Samples can also be collected (and pooled) at several sites in order to capture variation across different locations representing the same habitat or treatment category, often described as ‘replicates’ within a study. We assume that authors of component papers chose a combination of soil samples and sites that they expected would characterise the soil seed bank in their study system. In our analysis, we used the total sampled area in m2 or the sample grain as our measure of sampling scale, calculated as the size of the individual soil samples multiplied by the total number of samples taken for a particular observation. We also included the number of sites contributing to each observation as a measure of geographical extent. While this is an imperfect measure, we found it to be a parsimonious covariate, as well as the only such metric that was available in all records. More than half of our observations represent samples collected within a single site (n = 4763).
In order to categorise the global seed bank observations to an ecosystem in a systematic way, we used a combination of authors self-declared habitat descriptions coming from the global seed bank database (aquatic, open grassy systems, forest, arable, or wetland), broad categories of the WWF terrestrial ecoregions based on an observation’s geographic coordinates32, as well as realms in the fashion of the function-based typology for the Earth’s ecosystems33. Observations with habitat defined as arable or aquatic were placed in an arable and aquatic ‘biome’, respectively. Habitat and biome were used to categorise all observations into one of 15 categories (Table 1), with the aim to maintain ecologically-relevant distinctions of ecosystems (the combination of biome and ecoregion) while also including a reasonable number of observations within each group. Each ecosystem was then assigned into either the aquatic, terrestrial or transitional realm. Within the terrestrial realm, we assigned biomes as tundra, forests, grasslands, Mediterranean and desert, and arable systems, while we included all aquatic observations within an aquatic biome within the aquatic realm. The transitional realm included wetlands in the Mediterranean and desert, temperate and boreal, and tropical ecoregions (Table 1). Finally, the database labelled non-arable observations as coming from undisturbed or degraded habitats. Habitat degradation was included as a binary classification within each ecosystem, based on descriptions from the component studies. During compilation of the database, the contemporary habitat being assigned to one of the habitat categories listed above, with an additional ‘target habitat’ was also added where necessary. Target habitat refers to the ideal natural or semi-natural habitat of the system, according to the study’s authors, interpreted from site descriptions and/or due to the common practice in the seed bank literature of comparing assemblages across habitat gradients. If there was no entry for target habitat for a record in the database, we used the (contemporary) habitat entry and the record was labelled as being undisturbed. When there was an entry for target habitat, this was used for the habitat definition and the record was labelled as degraded. Note that the target habitat can have the same category as the contemporary habitat if it is still—for example—a forest, but in a degraded state. Examples of degraded habitats include drained wetlands, early-mid successional secondary forests, abandoned semi-natural grassland or any habitat under restoration, see Auffret et al.31 for further details.
Statistical models
For each response (i.e. species richness and seed density) and biome (e.g. terrestrial grasslands and transitional wetlands), we fit a univariate mixed effects linear model. This amounted to 14 models in total (two responses × seven biomes; Table S2), allowing estimates to vary independently for each biome without well-sampled biomes and realms (e.g. temperate forests) influencing the estimates of undersampled biomes and realms (e.g. tundra, aquatic). In general, if there were different ecoregions represented within a biome, then these were treated as categorical fixed effects that interacted with habitat degradation. If not (e.g. tundra), then degradation was the only categorical fixed effect. In the arable biome, there was no fixed effect for degradation, because we consider arable to already represent a degraded system. For each model, intercepts were allowed to vary for every record, nested within study identity as random effects.
In the species richness models, we quantified richness as a function of sampling area across the different biomes by fitting total sampled area (m2; sample grain) as a continuous predictor, ecoregion as an interacting categorical predictor (where relevant), degradation as an interacting categorical predictor (a three-way interaction where relevant), and the number of sites as a continuous covariate as a proxy for study extent. This approach, allowing drivers of diversity to vary with sampling grain within a single model, has been found to effectively unify disparate results from previous studies, enabling efficient integration of heterogeneous biodiversity data29. We assumed a Poisson distribution with a log-link function and we log-transformed and centred the total sampling area and the number of sites. We set the lower bound of slopes to be zero as informative prior information. This prevented uncertainty estimates for shallow slopes (aquatic realm) from overlapping with negative values, which does not logically hold. Seed density models included only degradation interacting with ecoregion as fixed effects, and assumed a lognormal distribution. For all model specifications and details, see Table S2.
For Bayesian inference and estimates of uncertainty, we fit models using the Hamiltonian Monte Carlo (HMC) sampler Stan96 using the R package brms (citations for all software are given at the end of the ‘Methods’). We fit all models with four chains, differing iterations and differing assumed distributions (Model Details Table S2). We used weakly regularising priors, visually inspected HMC chains for convergence, and conducted posterior predictive checks to ensure the model represented the distribution of the data for each (Table S2, Figs. S2–S3). We selected model structure and distributions based on this and on model performance (Rhats, sample sizes) (Table S2, Figs. S1–S2).
Quantifying scale-dependent richness
Relative differences in local species richness (alpha diversity) between habitats or biomes do not always match patterns of richness when larger areas are considered (gamma diversity). This indicates spatial variation in beta diversity (Whittaker’s beta = gamma/alpha)97. To investigate this divergence between scales, we first predicted the overall average species richness at two spatial scales for every biome using the species richness models to keep richness estimates for every biome consistent. We predicted richness at 0.01 m2 (alpha diversity), and at a larger scale, gamma diversity at 15 m2, keeping modelling parameters constant. These are arbitrary values, but we consider 0.01 m2 to be a commonly sampled size in soil seed bank studies, that falls within the interquartile range of seed bank sampling in the database31 and has been previously used for pattern prediction of soil seed bank richness and density26. Gamma diversity at 15 m2 was chosen as the largest scale that we felt confident predicting across biomes, given the input data and model outputs. In the main results we report diversity at both alpha and gamma scales assuming each observation was carried out at one site, as this represents the majority of our values, and acts as our baseline estimates for diversity.
We estimated the average number of seeds m−2 globally for observations within non-disturbed natural terrestrial, arable, wetland and aquatic realms. To do so, we combined estimates of the fixed effect (i.e. fitted values) values across all models (i.e. biomes) from each realm, and predicted the mean seed density m−2 across all fitted values from all models within each broad group. Finally, to look at the joint relationship between the density of seed in the soil seed bank and species richness, we predicted species richness at the 1 m2 scale used for seed density for direct comparison between the two. We predicted richness at this third scale 1 m2, following the same methods as for our alpha and gamma diversity. For visualisation, Fig. 2 density curves were built from full posterior predictive draws, while Fig. 3 density curves were built from population-level fitted values only. This reflects the scale-dependent nature of the richness analysis (alpha and gamma scales), which required predictive draws, whereas density was estimated at a single scale, for which fitted values were sufficient. For all data cleaning, management, wrangling and analyses, we used the R Project for Statistical Computing version 4.4.098. To handle data, quantify metrics, summarise results, and visualise data and estimates we used the brms, tidyverse, bayesplot, and ggplot2 packages99–102.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.