7. Configuration as YAML and the rompy CLI¶
What this shows: the hindcast from Tutorial 6 written as a YAML file, loaded in Python, and generated and run from the command line.
Prerequisites: 6. A tide and wind hindcast.
You will learn:
- how Python objects map to YAML, and what the type fields are for
- how to load, check and generate a YAML configuration in Python
- how to validate, generate and run it with the
rompycommand-line tool
Data used: the same files as Tutorial 6, referenced from
07_yaml_and_cli.yml.
Why YAML?¶
A YAML file holds the whole model in one place. It can be version-controlled, shared, edited without Python, and run on another machine or in a pipeline. It is checked against the same classes as the Python objects, so mistakes are caught before any file is written.
Setup¶
import shutil
from pathlib import Path
import yaml
from pydantic import ValidationError
from rompy.logging import config as logging_config
from rompy.model import ModelRun
logging_config.update(level="ERROR")
CONFIG_FILE = Path("07_yaml_and_cli.yml")
shutil.rmtree("_output/07_yaml_and_cli", ignore_errors=True)
1. The YAML file¶
The file has the fields of a ModelRun: run_id, output_dir and period, plus
config holding the SCHISM model, with its grid, data and nml. Each nested
object has the same fields as its Python class. The friction is a constant here, so
the file is self-contained.
print(CONFIG_FILE.read_text())
# The tide and wind hindcast of Tutorial 6 as a rompy ModelRun, with a constant
# drag coefficient. Paths are relative to this file's folder.
run_id: hindcast_yaml
output_dir: _output/07_yaml_and_cli
period:
start: 2023-01-01T00:00
end: 2023-01-05T00:00
interval: 1h
config:
model_type: schism
grid:
grid_type: schism
hgrid:
model_type: data_blob
source: ../data/hgrid.gr3
drag: 0.0025
data:
data_type: schism
boundary_conditions:
data_type: boundary_conditions
tidal_data:
tidal_database: ../data/tides
tidal_model: TPXO9-perth
constituents: [M2, S2, N2, K2, K1, O1, P1, Q1]
nodal_corrections: true
default_boundary:
elev_type: 3 # tidal elevation
vel_type: 3 # tidal currents
atmos:
data_type: sflux
air_1:
data_type: sflux_air
source:
model_type: file
uri: ../data/era5-perth-20230101-05.nc
uwind_name: u10
vwind_name: v10
prmsl_name: msl
filter:
sort:
coords: [latitude]
nml:
param:
core:
ibc: 1 # barotropic
ibtp: 0
dt: 120.0
nspool: 10 # output every 20 minutes
ihfskip: 720 # one output file per day
opt:
wtiminc: 3600.0 # hourly ERA5
schout:
iof_hydro__1: 1 # water level
iof_hydro__14: 1 # wind
iof_hydro__16: 1 # depth-averaged velocity
iof_hydro__26: 0 # no 3D velocity in a 2D model
model_type selects the model (schism), and grid_type and data_type name the
classes of the parts, so rompy can tell which class a block of YAML describes. Paths
are relative to the folder you run from, here the folder of the file.
2. Load and check in Python¶
Loading the YAML into a ModelRun checks everything and builds the same objects as
in Tutorial 6.
conf = yaml.safe_load(CONFIG_FILE.read_text())
modelrun = ModelRun(**conf)
setup = modelrun.config.data.boundary_conditions.default_boundary
print(type(modelrun.config.data.atmos.air_1).__name__)
print(f"open boundaries: elev_type={setup.elev_type}, vel_type={setup.vel_type}")
SfluxAir open boundaries: elev_type=3, vel_type=3
Invalid values are reported with the path to the field. Here a misspelt parameter in
param.nml's OPT group:
conf["config"]["nml"]["param"]["opt"]["wtimninc"] = 3600.0
try:
ModelRun(**conf)
except ValidationError as error:
for err in error.errors():
print(".".join(str(loc) for loc in err["loc"]), "->", err["msg"])
config.schism.nml.param.opt.wtimninc -> Extra inputs are not permitted
3. Generate from Python¶
Calling the ModelRun writes the workspace to output_dir/run_id.
workspace = Path(ModelRun(**yaml.safe_load(CONFIG_FILE.read_text()))())
sorted(p.name for p in workspace.iterdir())
['README', 'albedo.gr3', 'bctides.in', 'datasets', 'diffmax.gr3', 'diffmin.gr3', 'drag.gr3', 'hgrid.gr3', 'hgrid.ll', 'hgrid_WWM.gr3', 'outputs', 'param.nml', 'sflux', 'tvd.prop', 'vgrid.in', 'watertype.gr3', 'windrot_geo2proj.gr3', 'wwmbnd.gr3']
4. The rompy command-line tool¶
rompy installs a rompy command that works directly on the YAML file, so no Python
code is needed to produce a workspace or run the model. In a notebook, the ! prefix
runs a shell command. In a terminal, type the same command without the !.
rompy validate checks the file and exits with an error if anything is wrong:
!rompy validate 07_yaml_and_cli.yml > /dev/null 2>&1 && echo "Configuration is valid"
Configuration is valid
rompy generate writes the workspace:
!rompy generate 07_yaml_and_cli.yml > /dev/null 2>&1 && ls _output/07_yaml_and_cli/hindcast_yaml
albedo.gr3 diffmin.gr3 hgrid_WWM.gr3 sflux windrot_geo2proj.gr3 bctides.in drag.gr3 outputs tvd.prop wwmbnd.gr3 datasets hgrid.gr3 param.nml vgrid.in diffmax.gr3 hgrid.ll README watertype.gr3
rompy run generates the workspace and runs the model with a backend, described in
its own small YAML file. 07_docker_backend.yml runs SCHISM
in the public Docker image, with 6 processes of which 2 are scribes:
print(Path("07_docker_backend.yml").read_text())
# Run SCHISM in the public Docker image with rompy's Docker backend: 6 processes, # 2 of them scribes that write the output. type: docker image: ghcr.io/rom-py/schism:5.13.0 executable: schism 2 mpiexec: mpirun cpu: 6
The run takes about a minute, and needs Docker:
!rompy run 07_yaml_and_cli.yml --backend-config 07_docker_backend.yml > /dev/null 2>&1
As always with SCHISM, the log says whether the run completed, whatever the command returned:
!tail -n 1 _output/07_yaml_and_cli/hindcast_yaml/outputs/mirror.out
Run completed successfully at 20260929, 013638.886
Other useful commands are rompy schema (print the JSON schema of the
configuration) and rompy --help. Running SCHISM
covers other backends.
Summary¶
- A YAML file mirrors the Python objects;
model_type,grid_typeanddata_typeselect the classes. ModelRun(**yaml.safe_load(...))loads and checks it in Python.rompy validate,rompy generateandrompy rundo the same from the command line;runtakes the backend from a second YAML file.- Check
outputs/mirror.outafter a run, from the command line too.
This completes the tutorial. The examples cover each feature in more depth.