Work overview

Section 02 of 07

Package Specification

oceandatr: An R Package to Acquire and Process Geospatial Ocean Data

Jason Flower, Echelle S. Burns, Daniel C. Dunn, Andy Estep, Jason D. Everett, Jeffrey O. Hanson, Sarah E. Lester, and Anthony J. Richardson · 2026

Contents

Section 02 of 07

  1. 01Introduction
  2. 02Package Specification
  3. 03Package Usage
  4. 04Conclusion
  5. 05Author Contributions
  6. 06Funding
  7. 07Conflicts of Interest
Text size
Work overview

Section 2 of 7

Package Specification

Jason Flower, Echelle S. Burns, Daniel C. Dunn, Andy Estep, Jason D. Everett, Jeffrey O. Hanson, Sarah E. Lester, and Anthony J. Richardson · about 14 minutes

The oceandatr R package (hereafter, the package) extends the functionality of the R statistical computing environment (R Core Team 2025). The package is designed to be compatible with the existing ecosystem of R packages for processing spatial data and supports the two primary types of geospatial data: raster data in the SpatRaster format of the terra R package (Hijmans 2025) and vector data in sf (Pebesma 2018) format. Details of how to install the package can be found on GitHub where the package is hosted (https://github.com/emlab‐ucsb/oceandatr). The website also contains vignettes demonstrating uses of the package and detailed help pages for each function.

Workflow

The package has three general functions that: (1) obtain a boundary for an area of interest, (2) create a grid and (3) transfer data into the grid (Figure 1). In addition to these three functions, it has several functions designed to obtain specific datasets (see Section 2.2). The general workflow involves using the function get_boundary() to obtain a boundary for the area of interest, such as a country, ocean, or a country's exclusive economic zone (EEZ). To achieve this, the get_boundary() function uses the rnaturalearth R package (Massicotte and South 2023) to retrieve land boundaries, and the mrp_get() function from the mregions2 R package (Fernandez‐Bejarano and Pohl 2023) to obtain marine boundaries. The advantage of using the oceandatr function for retrieving boundary data is that it provides a single function for retrieving both terrestrial and marine boundaries (a capability not currently available in other packages), and the get_boundary() function is simpler to use than the mregions2 functions, which require the user to write a Contextual Query Language filter to select the country or area they want. The boundary can then be used with the get_grid() function to create a grid for the area with a specified resolution, coordinate reference system and format. Grids can be created in vector or raster format and can be composed of squares (for raster or vector formats) or hexagons (for vector format). Finally, the get_data_in_grid() function is used to transfer vector or raster data into the specified grid. The grid and data supplied to get_data_in_grid() can be in different coordinate reference systems, can cross the antimeridian, and can be any combination of raster or vector format. This flexibility is a key strength of the package since data from different sources are often in different formats and poor decisions concerning the spatial transformation and gridding process can result in distorted and inaccurate data. By using the package, users can ensure they have a reproducible and precise approach to processing and standardizing geospatial data.

FIGURE 1: Example usage of the oceandatr R package for the Pacific island of Samoa: The EEZ boundary is acquired using the get_boundary() function, then a raster grid at a resolution of 10,000 m (10 km × 10 km) covering the EEZ is created using the get_grid() function, and data on the distribution of oceanic ridges is placed into that grid using the get_data_in_grid() function.

FIGURE 1: Example usage of the oceandatr R package for the Pacific island of Samoa: The EEZ boundary is acquired using the get_boundary() function, then a raster grid at a resolution of 10,000 m (10 km × 10 km) covering the EEZ is created using the get_grid() function, and data on the distribution of oceanic ridges is placed into that grid using the get_data_in_grid() function.

Data Access

The package provides access to a wide range of geospatial data for the marine environment, including bathymetry and geomorphology; ecoregions and ecology; and human use data (Table 1). To minimize file size downloads, data are downloaded—where possible—in spatial subsets from public servers based on the area of interest. As some data are only available for download as a single, global dataset (i.e., the global seamounts, knolls, coral habitat and geomorphology data), these data are housed in a separate data package, oceandatrsets, purpose built to store such data (Flower 2026). This separate data package is automatically installed when installing the oceandatr package, and keeping the large datasets separate from the package functions avoids having to re‐download them each time oceandatr is updated.

Data | Description | Function | Source | Relevance to spatial modeling and planning
Bathymetry | Global ocean bathymetry data | get_bathymetry() | GEBCO 2026 global terrain model (GEBCO Bathymetric Compilation Group 2026) or user provided | Species distributions are strongly depth dependent (Brown and Thatje 2014; Ready et al. 2010)
Depth zones | Ocean depths classified into 5 depth zones: Continental shelf: 0–200 m depthUpper bathyal: 200–800 m depthLower bathyal: 800–3500 m depthAbyssal: 3500–6500 m depthHadal: > 6500 m depth | get_bathymetry(classify_bathymetry = TRUE) | GEBCO 2026 global terrain model (GEBCO Bathymetric Compilation Group 2026) or user provided. Classification done by package | Depth zones contain different pelagic and benthic habitats (Costello 2009; Vinogradov 1997)
Geomorphology | Global geomorphological features: abyssal hills, abyssal plains, basins, bridges, canyons, escarpments, fans, glacial troughs, guyots, plateaus, ridges, rift valleys, rises, shelf valleys, sills, spreading ridges, terraces, trenches, and troughs | get_geomorphology() | www.bluehabitats.org (Harris et al. 2014) | Feature types represent distinct seafloor habitats that are proxies for biodiversity (Ceccarelli et al. 2021)
Seamounts | Global distribution of seamount peaks | get_seamounts(raw = TRUE) | Yesson et al. (2021) | Seamounts are known to be biodiversity hotspots (Morato et al. 2010; Rowden et al. 2010)
Seamount peaks buffered | Seamount point locations buffered to a defined distance, i.e., a circle is drawn around each seamount peak | get_seamounts(buffer = …) | Yesson et al. (2021) | Morato et al. (2010) found higher biodiversity within 30–40 km of seamounts
Knolls | Global distribution of knolls (base areas). Knolls are small seamounts, < 1000 m but > 200 m above the seafloor (full definition in Morato et al. (2008)) | get_knolls() | Yesson et al. (2011) | Knolls, as small seamount‐like features, are also likely to host elevated biodiversity relative to surrounding areas
Coral habitat | Predicted ranges of: Antipatharia (black coral), octocorals, and Scleractinia (cold‐water corals) from 3 global distribution datasets | get_coral_habitat() | Davies and Guinotte (2011); Yesson et al. (2012, 2017) | Species are vulnerable to damage and are habitat for other species including commercially important fish (Davies and Guinotte 2011; Yesson et al. 2012)
Environmental zones | Environmental zones generated using sea surface data from Bio‐Oracle: chlorophyll concentration (mean), dissolved oxygen concentration (mean), nitrate concentration (mean), pH (mean), phosphate concentration (mean), total phytoplankton (mean), salinity (mean), sea surface temperature (mean, min, and max), and silicate concentration (mean) | get_enviro_zones() | Bio‐Oracle data: www.bio‐oracle.org (Assis et al. 2018, 2024). Accessed via the biooracler package (Fernandez 2024) | Zones are proxies for different pelagic species assemblages (Magris et al. 2021)
Ecoregions | Global marine ecoregion classifications: Marine ecoregions of the world (Spalding et al. 2007)Longhurst biogeographical provinces (Longhurst 2007)Large marine ecosystems (LMEs) (Sherman and Alexander 1986)Mesopelagic ecoregions (Sutton et al. 2017) | get_ecoregion() | Data accessed via the mregions2 R package (Fernandez‐Bejarano and Pohl 2023) | Ecoregions are commonly used in large‐scale spatial planning to ensure representation of biodiversity (Brito‐Morales et al. 2022; Ceccarelli et al. 2021)
Distance from shore, port or anchorage | Distance from shore, port or anchorage for each grid cell | get_dist() | Calculated using shore from Natural Earth (Natural Earth 2024), ports from World Port Index (National Geospatial‐Intelligence Agency 2024) and anchorages from Global Fishing Watch (Global Fishing Watch 2022) | Distance from shore can be used as a proxy for fishing effort (Caddy and Carocci 1999) for small‐scale, nearshore fishing
Fishing effort | Mean total annual fishing effort (hours) | get_gfw() | Global Fishing Watch (Kroodsma et al. 2018) | Used as an opportunity cost in spatial planning to avoid excessive fisheries impact (Chollett et al. 2022)

Most of the data oceandatr provides access to have global coverage and can be used at a range of spatial scales, from global to country level, for modeling and spatial planning. However, much of the data that the package provides direct access to may not be ideal for analyses at finer spatial scales, such as spatial planning in coastal environments. In these cases, higher‐resolution local data, where available, might be more appropriate. Any spatial data, regardless of whether it was accessed via the package, provided by local experts, or found online, can be prepared using the get_data_in_grid() function.

All functions for obtaining data require the user to provide either a spatial grid, in raster or vector format, or a polygon for the area of interest, and can return either gridded data or data within the polygon. The advantage of using the package to obtain data within a user‐provided polygon is that it handles the transformation between coordinate reference systems and can also return data for polygons that cross the antimeridian, a process that can be confusing and potentially lead to problems with the data if not done correctly. If the requested gridded data has multiple data classes (e.g., depth zones in the ocean), a multi‐layer raster will be returned with each class as a separate layer. Conversely, if the grid is in vector format, each class will be a column in the associated data frame. Some functions have further arguments to specify the classification of data, their source, and whether the area of interest crosses the antimeridian. In the following sections, we provide greater detail on the data that can be accessed through the package.

Bathymetric and Geomorphological Data

Bathymetry and geomorphology are widely used in modeling and spatial planning in the marine environment (Brito‐Morales et al. 2022; Harris et al. 2014; McQuaid et al. 2023). Bathymetry data can be retrieved using the get_bathymetry() function, and automated routines are available for classifying these data into five commonly used ocean depth zones: continental shelf (0–200 m depth), upper bathyal (200–800 m), lower bathyal (800–3500 m), abyssal (3500–6500 m), and hadal (> 6500 m). By default, these data are retrieved from the GEBCO 2026 global terrain model (GEBCO Bathymetric Compilation Group 2026) hosted at the Natural Environment Research Council's Centre for Environmental Data Analysis (CEDA), allowing access to the latest bathymetry data, a functionality not currently offered by other R packages. To allow downloading of larger‐than‐memory files, the data are downloaded in chunks that are saved sequentially to a single file on the user's hard drive, which is an important capability given the potentially large size of the downloads. Alternatively, get_bathymetry() provides the option for the user to specify a local raster file as the source of the bathymetry data, which is useful in cases where local datasets with higher accuracy and resolution are available or the CEDA server is unavailable.

The package has multiple functions to retrieve geomorphological data (Table 1). For instance, the get_geomorphology() function returns data from the global seafloor geomorphic features map (Harris et al. 2014). We also include specific functions to access data for seamounts and knolls (smaller seamounts) since these features are known to attract and aggregate biodiversity (Morato et al. 2010; Rowden et al. 2010) and are commonly a key feature for spatial plans (Brito‐Morales et al. 2022; Government of Bermuda and Bermuda Ocean Prosperity Programme 2024). The get_seamounts() function provides access to a global seamount dataset (Yesson et al. 2021), while get_knolls() can retrieve data on the area of knolls at their base (Yesson et al. 2011). The seamount dataset contains only the peaks of the seamounts (a single point). However, for spatial planning or modeling at local scales, it is useful to define an area around the seamount peak where biodiversity is expected to be higher than the surrounding ocean (Morato et al. 2010). This functionality is available via a “buffer” parameter, whereby the area within a specified radius of the seamount peak is included in the returned polygon or raster. Although the buffer radius can be defined by the user, a 30–40 km buffer radius is recommended as a default value (Morato et al. 2010).

Ecoregions and Ecological Data

Ecoregions are commonly used in spatial planning and modeling to delineate areas of differing environmental conditions (Brito‐Morales et al. 2022; Sala et al. 2021), under the assumption that different conditions support different biological communities. We provide access to four commonly used global ecoregional classification schemes via the get_ecoregion() function: marine ecoregions of the world (Spalding et al. 2007); the Longhurst biogeographical provinces (Longhurst 2007); Large Marine Ecosystems (Sherman and Alexander 1986); and mesopelagic ecoregions of the world (Sutton et al. 2017). These data are from the Marine Regions database (https://www.marineregions.org).

Although global ecoregion data are useful at larger scales, regional classifications are likely to be more valuable for planning and modeling at a national scale. The package allows the user to create their own ecoregionalisation, which can be used to delineate environmental zones (i.e., zones based on physical and chemical data). Building on an approach used in Magris et al. (2021), the zones are created using global ocean environmental data from Bio‐Oracle for 11 variables, including sea surface temperature (minimum, maximum and mean), pH and chlorophyll concentration (Assis et al. 2024; Tyberghein et al. 2012). These data are clustered using a kmeans clustering algorithm and the optimal number of clusters is estimated using the NbClust R package (Charrad et al. 2014), but can also be specified by the user. Each cluster represents a unique environmental ‘zone’. Both the raw environmental data and the clustered environmental zones can be retrieved using the get_enviro_zones() function.

The package can also access output from a limited number of species distribution models (SDMs) or habitat suitability models that predict the spatial distribution of species of interest. These models are commonly used in spatial planning as biodiversity layers to inform which areas to protect. While global SDMs for marine species exist (O'Hara et al. 2017), these require manual download for each species and cannot be included in the package due to usage restrictions and file size constraints. However, global cold water coral habitat suitability maps can be accessed for the species groups Antipatharia (Yesson et al. 2017), octocorals (Yesson et al. 2012) and Sclerarctinia (Davies and Guinotte 2011) using the get_coral_habitat() function. For inclusion in spatial planning, it is common to convert habitat suitability values to presence‐absence format. The original Scleractinia data are made available in presence‐absence format, while the habitat suitability values for Antipatharia are converted to presence‐absence by the package using the threshold value proposed by the dataset authors (Yesson et al. 2017), and octocorals are considered present in cells where at least 2 octocoral suborders are present (range of values: 0–7), though these thresholds can be changed by the user.

Human Use

Data on human use of the ocean is used extensively in spatial planning and fisheries and ecosystem models (Frawley et al. 2022; Lotze et al. 2019; Sala et al. 2021). In spatial planning, maps of fishing intensity often serve as a proxy for the opportunity cost of marine protection (Chollett et al. 2022). A widely used source of these data is Global Fishing Watch (GFW), which provides global data on apparent fishing effort (hours of fishing) at hourly, 0.01° resolution based on automatic identification system (AIS) data, a satellite tracking system for large vessels (Kroodsma et al. 2018). The gfwr R package allows access to many GFW data products via an API key (Clavelle et al. 2024), but fishing effort data can only be obtained for a single year per query. The get_gfw() function in oceandatr obtains GFW fishing effort data for multiple years and allows users to sum or average the data across selected years, but still requires an API key, which can be created for free on the GFW website (https://globalfishingwatch.org).

Only larger, industrial‐scale vessels are required to use AIS, and thus these data may not reflect fishing effort in areas where fishing is mainly by small, nearshore fishing vessels. In these cases, users could import other sources of fishing effort data if available. If no such data are available, distance to port or shore may be a reasonable proxy for fishing effort (Caddy and Carocci 1999). The get_dist() function can be used to get gridded distance from ports, anchorages, or shore. Port data are downloaded directly from the World Port Index, which is updated monthly (National Geospatial‐Intelligence Agency 2024), and the distance to shore is calculated from the Natural Earth country boundaries layer (Natural Earth 2024). Anchorages are from GFW, which includes > 167,000 geospatial points (Global Fishing Watch 2022). As calculating distance from ports and anchorages to each grid cell is computationally expensive, oceandatr provides two derived datasets. One groups anchorages with the same name into a single point, which is then assigned the mean longitude and latitude coordinates from the aggregated points. The second takes this aggregated dataset and removes anchorages that fall within land boundaries of countries, as defined by Natural Earth data (Natural Earth 2024), such as ports on rivers, assuming that users of the package will be principally interested in distance to coastal ports.