Data selection options¶
What this shows: the options that control which part of a dataset a forcing object uses. Most of them are inherited from rompy's data classes, so they work the same way for every wave, wind and water level class.
Prerequisites: Tutorial 4: Adding forcing and Data sources.
You will learn:
- how to inspect the data a forcing object will use, before writing any files
- how
coordsandvariablestell rompy what the dataset looks like - how
filtercrops the data when it is opened - how the run period crops the data with
crop_dataandtime_buffer
Data used: ssh_gridded.nc (sea-surface height).
Setup¶
from pathlib import Path
from rompy.core.time import TimeRange
from rompy.logging import config as logging_config
from rompy_xbeach.grid import RegularGrid
from rompy_xbeach.source import SourceCRSFile
logging_config.update(level="WARNING") # rompy logs every step, show warnings only
DATA_DIR = Path("../data")
OUT_DIR = Path("_output") / "data_selection"
OUT_DIR.mkdir(parents=True, exist_ok=True)
grid = RegularGrid(
ori={"x": 115.594239, "y": -32.641104, "crs": 4326},
alfa=347.0,
dx=10.0,
dy=15.0,
nx=230,
ny=220,
crs=28350,
)
period = TimeRange(start="2023-01-01T06", end="2023-01-01T12", interval="1h")
ssh_source = SourceCRSFile(
uri=DATA_DIR / "ssh_gridded.nc", crs=4326, x_dim="lon", y_dim="lat"
)
1. Where the options come from¶
Every wave, wind and water level class in rompy-xbeach is built on rompy's data classes and inherits their fields. rompy-xbeach adds the fields that choose a location on the XBeach grid:
| Field | From | Purpose |
|---|---|---|
source |
rompy | Where the data comes from |
coords |
rompy | Names of the time, x, y and site coordinates |
variables |
rompy | Variables to read |
filter |
rompy | Sort, subset, crop or rename when reading |
crop_data, time_buffer |
rompy | Crop the data to the run period, plus extra time steps |
buffer |
rompy | Spatial margin for cropping. No effect in rompy-xbeach, which samples points |
location |
rompy-xbeach | Sample at the grid centre or the offshore boundary midpoint |
sel_method, sel_method_kwargs |
rompy-xbeach | How station data is interpolated |
The inherited fields are documented with rompy's data objects.
2. Inspecting the data before writing files¶
The ds property returns the dataset exactly as the object sees it: source, variables
and filters applied. It is the quickest way to check a configuration.
from rompy_xbeach.data.waterlevel import WaterLevelGrid
waterlevel = WaterLevelGrid(
source=ssh_source,
coords={"x": "lon", "y": "lat"},
variables=["ssh"],
)
waterlevel.ds
<xarray.Dataset> Size: 208kB
Dimensions: (time: 25, lat: 45, lon: 23)
Coordinates:
* time (time) datetime64[ns] 200B 2023-01-01 ... 2023-01-02
* lat (lat) float64 360B -33.97 -33.9 -33.83 ... -31.17 -31.1 -31.03
* lon (lon) float64 184B 114.5 114.6 114.7 ... 115.9 115.9 116.0
spatial_ref int64 8B 0
Data variables:
ssh (time, lat, lon) float64 207kB ...
Attributes:
description: Tide elevation and currents time seriescoords names the dimensions. The default is longitude/latitude/time, so a
dataset with lon/lat must say so. variables limits what is read.
from rompy.core.filters import Filter
cropped = WaterLevelGrid(
source=ssh_source,
coords={"x": "lon", "y": "lat"},
variables=["ssh"],
filter=Filter(crop={"lon": slice(115.4, 115.8), "lat": slice(-32.8, -32.5)}),
)
print("Full: ", dict(waterlevel.ds.sizes))
print("Cropped:", dict(cropped.ds.sizes))
Full: {'time': 25, 'lat': 45, 'lon': 23}
Cropped: {'time': 25, 'lat': 5, 'lon': 6}
Other filters¶
Filter also provides sort (sort by coordinates), subset (keep some variables)
and rename (rename variables). Filters are applied after variables are selected,
so they transform the data a class reads but cannot create the variables it asks for.
4. Cropping to the run period¶
With crop_data=True (the default), get() limits the data to the run period plus
time_buffer source time steps on each side. The margin lets start and end times
that fall between source time steps be interpolated. After get() the time crop
appears in the filter:
print("Before:", waterlevel.filter.crop)
waterlevel.get(destdir=OUT_DIR, grid=grid, time=period)
print("After: ", waterlevel.filter.crop)
print("Times used:", waterlevel.ds.time.values[[0, -1]])
Before: {}
After: {'time': Slice(start=datetime.datetime(2023, 1, 1, 5, 0), stop=datetime.datetime(2023, 1, 1, 13, 0))}
Times used: ['2023-01-01T05:00:00.000000000' '2023-01-01T13:00:00.000000000']
Summary¶
dsshows the data a forcing object will use.coords,variablesandfilterare rompy options shared by all forcing classes.filter.croplimits what is read from large datasets. Filters apply aftervariablesare selected.- The run period crops the data automatically, with a margin set by
time_buffer.