Skip to content

historical_daily

drop_noncore_coords(ds: xr.Dataset) -> xr.Dataset

Drop coordinates the older and newer CDS extract formats disagree about.

Extracts pulled through the newer CDS API carry number (ensemble member) and expver ('0001' final ERA5, '0005' preliminary ERA5T) alongside the grid coords; older extracts carry neither, and are packed int16 rather than float32. xr.concat refuses to join datasets whose coordinates differ, so a look-ahead that crosses from an old extract into a new one has to be normalised first.

Which era a file belongs to tracks its download date, not its data year: surveying the archive, valid_time covers 1950-1989 and 2024 while time covers 1990-2023. So there are two seams, not the single 2023/2024 one this used to describe -- 1989->1990 crosses new format into old, 2023->2024 crosses old into new -- and the 1950-2023 regeneration passed through both.

Source code in src/climate_data/generate/historical_daily.py
def drop_noncore_coords(ds: xr.Dataset) -> xr.Dataset:
    """Drop coordinates the older and newer CDS extract formats disagree about.

    Extracts pulled through the newer CDS API carry `number` (ensemble member) and
    `expver` ('0001' final ERA5, '0005' preliminary ERA5T) alongside the grid coords;
    older extracts carry neither, and are packed int16 rather than float32. `xr.concat`
    refuses to join datasets whose coordinates differ, so a look-ahead that crosses from
    an old extract into a new one has to be normalised first.

    Which era a file belongs to tracks its *download* date, not its data year: surveying
    the archive, `valid_time` covers 1950-1989 and 2024 while `time` covers 1990-2023. So
    there are two seams, not the single 2023/2024 one this used to describe -- 1989->1990
    crosses new format into old, 2023->2024 crosses old into new -- and the 1950-2023
    regeneration passed through both.
    """
    extra = []
    for coord in ds.coords:
        if coord not in CORE_COORDS:
            extra.append(coord)
    return ds.drop_vars(extra)

load_variable_with_lookahead(cdata: ClimateData, variable: str, year: str, month: str, dataset: str) -> xr.Dataset

Load a month's hourly data plus the one sample that closes its final day.

Accumulation windows are stamped by their end, so the last day of a month is closed by the next month's first sample -- for December, by the next year's January. That single extra sample completes the final bin without opening a new one.

Raises if the look-ahead file is absent rather than degrading to the incomplete 23-hour window, which should stop the run rather than quietly shorten one day. Checked before loading the target month so the run fails fast. Note this reaches one year past the history range -- regenerating the last history year needs the following January -- which is why the extract tasks span cdc.EXTRACT_YEARS rather than cdc.HISTORY_YEARS.

Source code in src/climate_data/generate/historical_daily.py
def load_variable_with_lookahead(
    cdata: ClimateData,
    variable: str,
    year: str,
    month: str,
    dataset: str,
) -> xr.Dataset:
    """Load a month's hourly data plus the one sample that closes its final day.

    Accumulation windows are stamped by their end, so the last day of a month is closed
    by the next month's first sample -- for December, by the next *year's* January. That
    single extra sample completes the final bin without opening a new one.

    Raises if the look-ahead file is absent rather than degrading to the incomplete
    23-hour window, which should stop the run rather than quietly shorten one day.
    Checked before loading the target month so the run fails fast. Note this reaches one
    year past the history range -- regenerating the last history year needs the following
    January -- which is why the extract tasks span `cdc.EXTRACT_YEARS` rather than
    `cdc.HISTORY_YEARS`.
    """
    next_year, next_month = _next_month(year, int(month))
    lookahead_path = cdata.extracted_era5_path(dataset, variable, next_year, next_month)
    if not lookahead_path.exists():
        msg = (
            f"Cannot close {year}-{month} for {variable}: the accumulation window of its"
            f" final day ends in the following month, whose extract is missing at"
            f" {lookahead_path}. Extract it before generating {year}:\n"
            f"  cdtask extract era5_download -d {dataset} -x {variable}"
            f" -y {next_year} -m {next_month}\n"
            f"  cdtask extract era5_compress -d {dataset} -x {variable}"
            f" -y {next_year} -m {next_month}\n"
            f"Use the single-job tasks, not `cdrun extract era5`: the runner's --year"
            f" stops at the last history year, and it decides what to fetch by file"
            f" existence."
        )
        raise FileNotFoundError(msg)

    ds = drop_noncore_coords(load_variable(cdata, variable, year, month, dataset))
    lookahead = drop_noncore_coords(
        load_variable(cdata, variable, next_year, next_month, dataset).isel(time=[0])
    )

    # Taking sample zero assumes the next month opens on midnight, which is what closes
    # the target month's final day. ERA5-Land 1950_01 opens at 01:00 instead, so the
    # assumption does fail somewhere in the archive -- silently, at one day per month.
    stamp = pd.Timestamp(lookahead.time.to_index()[0])
    if (stamp.hour, stamp.minute) != (0, 0):
        msg = (
            f"Cannot close {year}-{month} for {variable}: the look-ahead at"
            f" {lookahead_path} opens at {stamp}, not midnight, so its first sample does"
            f" not close the final day of {year}-{month}."
        )
        raise ValueError(msg)

    # `join="exact"` rather than the default `"outer"`: a look-ahead on a different grid
    # should fail here, not silently NaN-pad both months onto a union grid. Note this
    # guards the single-level datasets only -- `load_variable` overwrites ERA5-Land's
    # coordinates from `cdc.ERA5_LAND_*`, so a genuine land-grid change would be relabelled
    # before reaching this point. The real land files differ across the CDS format change
    # by ~1e-5 degrees, which is why that overwrite exists.
    return xr.concat([ds, lookahead], dim="time", join="exact")

trim_to_month(ds: xr.Dataset, year: str, month: int) -> xr.Dataset

Drop collapsed bins that fall outside the target month.

The interval-aware collapse labels its first bin with the previous month's last day, since that bin holds that day's closing sample. Left in place, every month contributes a duplicate date at its seam.

Source code in src/climate_data/generate/historical_daily.py
def trim_to_month(ds: xr.Dataset, year: str, month: int) -> xr.Dataset:
    """Drop collapsed bins that fall outside the target month.

    The interval-aware collapse labels its first bin with the *previous* month's last
    day, since that bin holds that day's closing sample. Left in place, every month
    contributes a duplicate date at its seam.
    """
    dates = pd.to_datetime(ds.date.to_index())
    keep = np.flatnonzero((dates.year == int(year)) & (dates.month == month))
    return ds.isel(date=keep)