extract
Climate Data Extraction
This module contains pipelines for extracting climate data from various sources.
cmip6
CMIP6 Data Extraction
check_encoding_covers(data_min: float, data_max: float, offset: float, scale: float, variable: str, dtype: str = 'int16') -> None
Refuse to write values the declared encoding cannot represent.
Packing to a 16-bit integer rounds (value - offset) / scale, and anything outside
the type's range wraps modulo 65536. to_netcdf does this silently, so the corruption
is invisible until someone plots the result and finds negative rainfall. pr shipped
with scale_factor=1e-9 for two years -- a 2.83 mm/day ceiling -- and produced 295
files in which 26.4% of sampled cells were wrong and 12.4% were negative.
Raising here makes the next such mistake a failed extract rather than a corrupt archive. An unsigned dtype also makes a negative value an error rather than a wrap, which for a flux like precipitation is the honest outcome.
Source code in src/climate_data/extract/cmip6.py
extract_cmip6(cmip6_source: list[str], cmip6_experiment: list[str], cmip6_variable: list[str], output_dir: str, queue: str, overwrite: bool, dry_run: bool) -> None
Extract CMIP6 data.
Extracts CMIP6 data for the given source, experiment, and variable. We use the
the table at https://www.nature.com/articles/s41597-023-02549-6/tables/3 to determine
which CMIP6 source_ids to include. See ClimateData.load_koppen_geiger_model_inclusion
to load and examine this table. The extraction criteria does not completely
capture model inclusion criteria as it does not account for the year range avaialable
in the data. This determiniation is made when we proccess the data in later steps.
Fans out one job per ensemble member rather than one per (source, experiment). The
member counts are wildly uneven -- MIROC6 has 50 pr members where most sources have
one -- so grouping them made three jobs carry fifty times the work of a typical one.
More importantly, extract_cmip6_main re-raises on failure, so a member the encoding
guard rejects used to abandon every member behind it in the same job. One job per
member contains that to the member that failed, and makes a resumed run skip the
members already written instead of redoing whole groups.
The member space is not a cartesian product -- not every source publishes every
variant for every experiment -- so it is enumerated from the metadata and passed as
flat_node_args.
Source code in src/climate_data/extract/cmip6.py
254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 | |
load_cmip_data(zarr_path: str) -> xr.Dataset
Loads a CMIP6 dataset from a zarr path.
Source code in src/climate_data/extract/cmip6.py
select_members(meta: pd.DataFrame, cmip6_source: str, cmip6_experiment: str, cmip6_variable: str, gcm_member: str | None = None) -> dict[str, str]
The {member_id: zstore} this job should extract.
With gcm_member given, narrows to that one ensemble member so the runner can put
each member in its own job. Shared with the runner, which enumerates the same space
to build its task list -- if these two disagreed, the runner would submit jobs whose
member does not exist and they would silently extract nothing.
Source code in src/climate_data/extract/cmip6.py
elevation
extract_elevation(model_name: str, output_dir: str, queue: str, dry_run: bool) -> None
Download elevation data from Open Topography.
Source code in src/climate_data/extract/elevation.py
extract_elevation_task(model_name: str, lat_start: int, lon_start: int, output_dir: str) -> None
Download elevation data from Open Topography.
Source code in src/climate_data/extract/elevation.py
era5
ERA5 Data Extraction
check_extract_year_floor(years: Sequence[str], *, allow_pre_floor: bool) -> None
Refuse a run that reaches below cdc.EXTRACT_YEAR_FLOOR unless asked to.
build_task_lists treats a missing output file as work to do, so after ERF's Sep2026
deletion every pre-1980 extract looks like a gap. The runner's --year defaults to
ALL, which means the bare invocation would refill 3,238 files and silently undo the
reclamation -- days of Copernicus queue to recover from. Guarding the outcome rather
than the default also catches an explicitly typed --year ALL.
Source code in src/climate_data/extract/era5.py
variables_for_full_expansion(era5_variables: Sequence[str]) -> list[str]
Drop never-read variables when the whole variable set was requested.
--era5-variable ALL resolves to every declared variable, so the default invocation
downloads and stores surface_pressure for every month of every year although no
stage opens it. Naming a variable explicitly still extracts it; this only narrows the
meaning of "all" to "all the ones we use".