neural_lam.datastore.mdp#

Datastore implementation wrapping mllam-data-prep outputs.

Module Contents#

class neural_lam.datastore.mdp.MDPDatastore(config_path, n_boundary_points=30, reuse_existing=True)#

Bases: neural_lam.datastore.base.BaseRegularGridDatastore

Datastore class for datasets made with the mllam_data_prep library (mllam/mllam-data-prep). This class wraps the mllam_data_prep library to do the necessary transforms to create the different categories (state/forcing/static) of data, with the actual transform to do being specified in the configuration file.

Construct a new MDPDatastore from the configuration file at config_path. A boundary mask is created with n_boundary_points boundary points. If reuse_existing is True, the dataset is loaded from a zarr file if it exists (unless the config has been modified since the zarr was created), otherwise it is created from the configuration file.

Parameters:
  • config_path (str) – The path to the configuration file, this will be fed to the mllam_data_prep.Config.from_yaml_file method to then call mllam_data_prep.create_dataset to create the dataset.

  • n_boundary_points (int) – The number of boundary points to use in the boundary mask.

  • reuse_existing (bool) – Whether to reuse an existing dataset zarr file if it exists and its creation date is newer than the configuration file.

get_dataarray(category: str, split: str | None, standardize: bool = False) xarray.DataArray | None#

Return the processed data (as a single xr.DataArray) for the given category of data and test/train/val-split that covers all the data (in space and time) of a given category (state/forcing/static). “state” is the only required category, for other categories, the method will return None if the category is not found in the datastore.

The returned dataarray will at minimum have dimensions of (grid_index, {category}_feature) so that any spatial dimensions have been stacked into a single dimension and all variables and levels have been stacked into a single feature dimension named by the category of data being loaded.

For categories of data that have a time dimension (i.e. not static data), the dataarray will additionally have (analysis_time, elapsed_forecast_duration) dimensions if is_forecast is True, or (time) if is_forecast is False.

If the data is ensemble data, the dataarray will have an additional ensemble_member dimension.

Parameters:
  • category (str) – The category of the dataset (state/forcing/static).

  • split (str) – The time split to filter the dataset (train/val/test).

  • standardize (bool) – If the dataarray should be returned standardized

Returns:

The xarray DataArray object with processed dataset.

Return type:

xr.DataArray or None

get_num_data_vars(category: str) int#

Return the number of variables in the given category.

Parameters:

category (str) – The category of the dataset (state/forcing/static).

Returns:

The number of variables in the given category.

Return type:

int

get_standardization_dataarray(category: str) xarray.Dataset#

Return the standardization dataarray for the given category. This should contain a {category}_mean and {category}_std variable for each variable in the category. For category==”state”, the dataarray should also contain a state_diff_mean_standardized and state_diff_std_standardized variable for the one-step differences of the state variables.

Parameters:

category (str) – The category of the dataset (state/forcing/static).

Returns:

The standardization dataarray for the given category, with variables for the mean and standard deviation of the variables (and differences for state variables).

Return type:

xr.Dataset

get_vars_long_names(category: str) List[str]#

Return the long names of the variables in the given category.

Parameters:

category (str) – The category of the dataset (state/forcing/static).

Returns:

The long names of the variables in the given category.

Return type:

List[str]

get_vars_names(category: str) List[str]#

Return the names of the variables in the given category.

Parameters:

category (str) – The category of the dataset (state/forcing/static).

Returns:

The names of the variables in the given category.

Return type:

List[str]

get_vars_units(category: str) List[str]#

Return the units of the variables in the given category.

Parameters:

category (str) – The category of the dataset (state/forcing/static).

Returns:

The units of the variables in the given category.

Return type:

List[str]

get_xy(category: str, stacked: bool) numpy.ndarray#

Return the x, y coordinates of the dataset.

Parameters:
  • category (str) – The category of the dataset (state/forcing/static).

  • stacked (bool) – Whether to stack the x, y coordinates.

Returns:

The x, y coordinates of the dataset, returned differently based on the value of stacked: - stacked==True: shape (n_grid_points, 2) where

n_grid_points=N_x*N_y.

  • stacked==False: shape (N_x, N_y, 2)

Return type:

np.ndarray

CARTESIAN_COORDS = None#
SHORT_NAME = 'mdp'#
property boundary_mask: xarray.DataArray#

Produce a 0/1 mask for the boundary points of the dataset, these will sit at the edges of the domain (in x/y extent) and will be used to mask out the boundary points from the loss function and to overwrite the boundary points from the prediction. For now this is created when the mask is requested, but in the future this could be saved to the zarr file.

Returns:

A 0/1 mask for the boundary points of the dataset, where 1 is a boundary point and 0 is not.

Return type:

xr.DataArray

property config: mllam_data_prep.Config#

The configuration of the dataset.

Returns:

The configuration of the dataset.

Return type:

mdp.Config

property coords_projection: cartopy.crs.Projection#

Return the projection of the coordinates.

NOTE: currently this expects the projection information to be in the extra section of the configuration file, with a projection key containing a class_name and kwargs for constructing the cartopy.crs.Projection object. This is a temporary solution until the projection information can be parsed in the produced dataset itself. mllam-data-prep ignores the contents of the extra section of the config file which is why we need to check that the necessary parts are there.

Returns:

The projection of the coordinates.

Return type:

ccrs.Projection

property grid_shape_state#

The shape of the cartesian grid for the state variables.

Returns:

The shape of the cartesian grid for the state variables.

Return type:

CartesianGridShape

property root_path: pathlib.Path#

The root path of the dataset.

Returns:

The root path of the dataset.

Return type:

Path

property step_length: datetime.timedelta#

The length of the time steps as a time interval.

Returns:

The length of the time steps as a datetime.timedelta object.

Return type:

timedelta