Skip to content

utils

annual_mean_from_monthly(ds: xr.Dataset) -> xr.Dataset

Day-weighted annual mean of a 12-month climatology.

Weights are bound by month label, not position, so a permuted month coordinate still weights correctly; anything other than a full 1..12 coordinate is rejected rather than misweighted.

Source code in src/climate_data/generate/utils.py
def annual_mean_from_monthly(ds: xr.Dataset) -> xr.Dataset:
    """Day-weighted annual mean of a 12-month climatology.

    Weights are bound by month label, not position, so a permuted month
    coordinate still weights correctly; anything other than a full 1..12
    coordinate is rejected rather than misweighted.
    """
    months = np.arange(1, 13)
    if not np.array_equal(np.sort(ds["month"].to_numpy()), months):
        msg = f"expected a full 1..12 month coordinate, got {ds['month'].to_numpy()}"
        raise ValueError(msg)
    weights = xr.DataArray(cdc.DAYS_IN_MONTH, dims=["month"], coords={"month": months})
    return ds.weighted(weights).mean("month")

buck_vapor_pressure(temperature_c: xr.Dataset) -> xr.Dataset

Approximate vapor pressure of water.

https://en.wikipedia.org/wiki/Arden_Buck_equation https://journals.ametsoc.org/view/journals/apme/20/12/1520-0450_1981_020_1527_nefcvp_2_0_co_2.xml

Parameters

temperature_c Temperature in Celsius

Returns

xr.Dataset Vapor pressure in hPa

Source code in src/climate_data/generate/utils.py
def buck_vapor_pressure(temperature_c: xr.Dataset) -> xr.Dataset:
    """Approximate vapor pressure of water.

    https://en.wikipedia.org/wiki/Arden_Buck_equation
    https://journals.ametsoc.org/view/journals/apme/20/12/1520-0450_1981_020_1527_nefcvp_2_0_co_2.xml

    Parameters
    ----------
    temperature_c
        Temperature in Celsius

    Returns
    -------
    xr.Dataset
        Vapor pressure in hPa
    """
    over_water = 6.1121 * np.exp(
        (18.678 - temperature_c / 234.5) * (temperature_c / (257.14 + temperature_c))
    )
    over_ice = 6.1115 * np.exp(
        (23.036 - temperature_c / 333.7) * (temperature_c / (279.82 + temperature_c))
    )
    vp = xr.where(temperature_c > 0, over_water, over_ice)  # type: ignore[no-untyped-call]
    return vp  # type: ignore[no-any-return]

daily_accumulation_last(ds: xr.Dataset) -> xr.Dataset

Collapse a within-day cumulative variable to its daily total.

For ERA5-Land total_precipitation, which accumulates since 00Z, the day's total is the window's closing sample, so last reads it directly. Prefer this to a maximum: the two agree only while the window rises monotonically, and int16 packing can make it tick down in its final step, in which case the maximum is an earlier, quantisation-inflated sample. Measured over 741,270 land day-pixels, last matched the true close for 100.0000% against 99.9864% for max.

Note resample emits a regular axis with NaN for empty bins, unlike groupby, which emits observed groups only. Callers must trim the partial bins at each end.

skipna=False because the default steps back past an absent closing sample to hour 23, reintroducing the incomplete window this function exists to eliminate. NaN is the honest answer for a window that never closed. It is not, however, a detectable one: generate_historical_daily_main fills ERA5-Land NaNs from the interpolated single-level field -- that is how ocean pixels are supplied -- so such a pixel carries a complete 0.25 degree value rather than a truncated 0.1 degree one, and never reaches validate_output. Preferring the coarse-but-whole value is the point; detecting the substitution would need a separate check against the sea mask.

Source code in src/climate_data/generate/utils.py
def daily_accumulation_last(ds: xr.Dataset) -> xr.Dataset:
    """Collapse a within-day cumulative variable to its daily total.

    For ERA5-Land `total_precipitation`, which accumulates since 00Z, the day's total is
    the window's closing sample, so `last` reads it directly. Prefer this to a maximum:
    the two agree only while the window rises monotonically, and int16 packing can make
    it tick down in its final step, in which case the maximum is an earlier,
    quantisation-inflated sample. Measured over 741,270 land day-pixels, `last` matched
    the true close for 100.0000% against 99.9864% for `max`.

    Note `resample` emits a regular axis with NaN for empty bins, unlike `groupby`, which
    emits observed groups only. Callers must trim the partial bins at each end.

    `skipna=False` because the default steps back past an absent closing sample to hour
    23, reintroducing the incomplete window this function exists to eliminate. NaN is the
    honest answer for a window that never closed. It is not, however, a detectable one:
    `generate_historical_daily_main` fills ERA5-Land NaNs from the interpolated
    single-level field -- that is how ocean pixels are supplied -- so such a pixel carries
    a complete 0.25 degree value rather than a truncated 0.1 degree one, and never reaches
    `validate_output`. Preferring the coarse-but-whole value is the point; detecting the
    substitution would need a separate check against the sea mask.
    """
    resampled = ds.resample(
        time=ACCUMULATION_FREQ,
        closed=ACCUMULATION_CLOSED,
        label=ACCUMULATION_LABEL,
    ).last(skipna=False)
    return resampled.rename({"time": "date"})

daily_accumulation_sum(ds: xr.Dataset) -> xr.Dataset

Collapse a variable of per-hour increments to its daily total.

For ERA5 single-levels, "accumulations are over the hour ending at the validity date/time", so the increments sum. The same interval convention applies as for daily_accumulation_last, so the two stay on a common date axis -- they are merged downstream in generate_historical_daily_main.

Source code in src/climate_data/generate/utils.py
def daily_accumulation_sum(ds: xr.Dataset) -> xr.Dataset:
    """Collapse a variable of per-hour increments to its daily total.

    For ERA5 single-levels, "accumulations are over the hour ending at the validity
    date/time", so the increments sum. The same interval convention applies as for
    `daily_accumulation_last`, so the two stay on a common `date` axis -- they are merged
    downstream in `generate_historical_daily_main`.
    """
    resampled = ds.resample(
        time=ACCUMULATION_FREQ,
        closed=ACCUMULATION_CLOSED,
        label=ACCUMULATION_LABEL,
    ).sum()
    return resampled.rename({"time": "date"})

identity(ds: xr.Dataset) -> xr.Dataset

Identity transformation

Source code in src/climate_data/generate/utils.py
def identity(ds: xr.Dataset) -> xr.Dataset:
    """Identity transformation"""
    return ds

interpolate_to_target_latlon(ds: xr.Dataset, method: str = 'nearest', target_lon: xr.DataArray = cdc.TARGET_LONGITUDE, target_lat: xr.DataArray = cdc.TARGET_LATITUDE) -> xr.Dataset

Interpolate a dataset to a target latitude and longitude grid.

Parameters

ds Dataset to interpolate method Interpolation method target_lon Target longitude grid target_lat Target latitude grid

Returns

xr.Dataset Interpolated dataset

Source code in src/climate_data/generate/utils.py
def interpolate_to_target_latlon(
    ds: xr.Dataset,
    method: str = "nearest",
    target_lon: xr.DataArray = cdc.TARGET_LONGITUDE,
    target_lat: xr.DataArray = cdc.TARGET_LATITUDE,
) -> xr.Dataset:
    """Interpolate a dataset to a target latitude and longitude grid.

    Parameters
    ----------
    ds
        Dataset to interpolate
    method
        Interpolation method
    target_lon
        Target longitude grid
    target_lat
        Target latitude grid

    Returns
    -------
    xr.Dataset
        Interpolated dataset
    """
    return (
        ds.interp(longitude=target_lon, latitude=target_lat, method=method)  # type: ignore[arg-type]
        .interpolate_na(dim="longitude", method="nearest", fill_value="extrapolate")
        .sortby("latitude")
        .interpolate_na(dim="latitude", method="nearest", fill_value="extrapolate")
        .sortby("latitude", ascending=False)
    )

kelvin_to_celsius(temperature_k: xr.Dataset) -> xr.Dataset

Convert temperature from Kelvin to Celsius

Parameters

temperature_k Temperature in Kelvin

Returns

xr.Dataset Temperature in Celsius

Source code in src/climate_data/generate/utils.py
def kelvin_to_celsius(temperature_k: xr.Dataset) -> xr.Dataset:
    """Convert temperature from Kelvin to Celsius

    Parameters
    ----------
    temperature_k
        Temperature in Kelvin

    Returns
    -------
    xr.Dataset
        Temperature in Celsius
    """
    return temperature_k - 273.15

meter_to_millimeter(rainfall_m: xr.Dataset) -> xr.Dataset

Convert rainfall from meters to millimeters

Parameters

rainfall_m Rainfall in meters

Returns

xr.Dataset Rainfall in millimeters

Source code in src/climate_data/generate/utils.py
def meter_to_millimeter(rainfall_m: xr.Dataset) -> xr.Dataset:
    """Convert rainfall from meters to millimeters

    Parameters
    ----------
    rainfall_m
        Rainfall in meters

    Returns
    -------
    xr.Dataset
        Rainfall in millimeters
    """
    return 1000 * rainfall_m

parse_reference_years(reference_years: str) -> slice

Turn a START-END year string (inclusive, four-digit years) into a date slice.

Source code in src/climate_data/generate/utils.py
def parse_reference_years(reference_years: str) -> slice:
    """Turn a START-END year string (inclusive, four-digit years) into a date slice."""
    start, sep, end = reference_years.partition("-")
    valid = (
        bool(sep)
        and len(start) == 4  # noqa: PLR2004
        and len(end) == 4  # noqa: PLR2004
        and start.isdigit()
        and end.isdigit()
        and int(start) <= int(end)
    )
    if not valid:
        msg = (
            "reference-years must be 'START-END' with four-digit years and "
            f"START <= END, got {reference_years!r}"
        )
        raise ValueError(msg)
    return slice(f"{start}-01-01", f"{end}-12-31")

precipitation_flux_to_rainfall(precipitation_flux: xr.Dataset) -> xr.Dataset

Convert precipitation flux to rainfall

Parameters

precipitation_flux Precipitation flux in kg m-2 s-1

Returns

xr.Dataset Rainfall in mm/day

Source code in src/climate_data/generate/utils.py
def precipitation_flux_to_rainfall(precipitation_flux: xr.Dataset) -> xr.Dataset:
    """Convert precipitation flux to rainfall

    Parameters
    ----------
    precipitation_flux
        Precipitation flux in kg m-2 s-1

    Returns
    -------
    xr.Dataset
        Rainfall in mm/day
    """
    seconds_per_day = 86400
    mm_per_kg_m2 = 1
    return seconds_per_day * mm_per_kg_m2 * precipitation_flux

rh_percent(temperature_c: xr.Dataset, dewpoint_temperature_c: xr.Dataset) -> xr.Dataset

Calculate relative humidity from temperature and dewpoint temperature.

Parameters

temperature_c Temperature in Celsius dewpoint_temperature_c Dewpoint temperature in Celsius

Returns

xr.Dataset Relative humidity as a percentage

Source code in src/climate_data/generate/utils.py
def rh_percent(
    temperature_c: xr.Dataset, dewpoint_temperature_c: xr.Dataset
) -> xr.Dataset:
    """Calculate relative humidity from temperature and dewpoint temperature.

    Parameters
    ----------
    temperature_c
        Temperature in Celsius
    dewpoint_temperature_c
        Dewpoint temperature in Celsius

    Returns
    -------
    xr.Dataset
        Relative humidity as a percentage
    """
    # saturation vapour pressure
    svp = buck_vapor_pressure(temperature_c)
    # actual vapour pressure
    vp = buck_vapor_pressure(dewpoint_temperature_c)
    return 100 * vp / svp

scale_wind_speed_height(wind_speed_10m: xr.Dataset) -> xr.Dataset

Scaling wind speed from a height of 10 meters to a height of 2 meters

Reference: Bröde et al. (2012) https://doi.org/10.1007/s00484-011-0454-1

Parameters

wind_speed_10m The 10m wind speed [m/s]. May be signed (ie a velocity component)

Returns

xr.DataSet The 2m wind speed [m/s]. May be signed (ie a velocity component)

Source code in src/climate_data/generate/utils.py
def scale_wind_speed_height(wind_speed_10m: xr.Dataset) -> xr.Dataset:
    """Scaling wind speed from a height of 10 meters to a height of 2 meters

    Reference: Bröde et al. (2012)
    https://doi.org/10.1007/s00484-011-0454-1

    Parameters
    ----------
    wind_speed_10m
        The 10m wind speed [m/s]. May be signed (ie a velocity component)

    Returns
    -------
    xr.DataSet
        The 2m wind speed [m/s]. May be signed (ie a velocity component)
    """
    scale_factor = np.log10(2 / 0.01) / np.log10(10 / 0.01)
    return scale_factor * wind_speed_10m  # type: ignore[no-any-return]

vector_magnitude(x: xr.Dataset, y: xr.Dataset) -> xr.Dataset

Calculate the magnitude of a vector.

Source code in src/climate_data/generate/utils.py
def vector_magnitude(x: xr.Dataset, y: xr.Dataset) -> xr.Dataset:
    """Calculate the magnitude of a vector."""
    return np.sqrt(x**2 + y**2)  # type: ignore[return-value]