Seibert et al. (2026) Setting the bar: benchmarks for model performances in large-sample hydrology
Identification
- Journal: Hydrology and earth system sciences
- Year: 2026
- Date: 2026-09-16
- Authors: Jan Seibert, Marc J. P. Vis, Sandra Pool
- DOI: 10.5194/hess-30-5857-2026
Research Groups
- Department of Geography, University of Zurich, Switzerland
- Eawag, Swiss Federal Institute of Aquatic Science and Technology, Department Water Resources and Drinking Water, Switzerland
Short Summary
This study develops and provides lower and upper benchmarks for hydrological model performance across 13 large-sample datasets, offering guidance on their computation and advocating for their broader use in assessing model performance relative to local hydroclimatic conditions.
Objective
- To evaluate different approaches for deriving lower benchmark values, including determining appropriate ensemble sizes, assessing the effects of parameter ranges, deciding between random or regional parameter sets, and evaluating ensemble aggregation methods.
- To compute and provide upper and lower benchmark values for streamflow simulations across existing large-sample datasets.
- To examine the relationships between lower and upper benchmarks and catchment characteristics.
Study Configuration
- Spatial Scale: Over 6000 catchments from 13 large-sample datasets across Australia, Brazil, Central Europe, Chile, Denmark, France, Germany, Great Britain, Luxembourg, Spain, Sweden, Switzerland, and the US. Catchments were disaggregated into 200 meter (m) elevation bands.
- Temporal Scale: Daily hydro-meteorological time series. Simulations included a 1-2 year warming-up period and were run for the longest possible concurrent data period, which varied across datasets.
Methodology and Data
- Models used: HBV (Hydrologiska Byråns Vattenavdelning) bucket-type model, specifically HBV light, version 4, with 13 free parameters. Random forest regression trees were used to predict benchmark values based on catchment characteristics.
- Data sources:
- Thirteen large-sample datasets (e.g., CAMELS-Australia, CAMELS-Brazil, LamaH-CE, CAMELS-Chile, CAMELS-Denmark, CAMELS-France, CAMELS-Germany, CAMELS-GB, CAMELS-Luxembourg, CAMELS-Spain, CAMELS-SE, CAMELS-CH, CAMELS-US).
- Catchment-averaged daily precipitation, temperature, and potential evapotranspiration time series.
- Observed daily catchment streamflow time series.
- Digital Elevation Models (DEMs) (SRTM, ASTER, EarthEnv) for elevation band disaggregation.
- 20 catchment characteristics (climatic, hydrological, and topographic) for random forest analysis.
- Performance measures: Nash-Sutcliffe efficiency (NSE), Kling-Gupta efficiency (KGE), and non-parametric Kling-Gupta efficiency (NPE).
Main Results
- Lower benchmark ensemble size: Using 1000 or more randomly generated parameter sets ensured stable and robust lower benchmark performance values.
- Lower benchmark aggregation: The performance of the ensemble mean discharge simulation consistently outperformed the median performance of individual simulations in 94-99% of catchments, particularly for reasonably well-performing catchments (NPE > 0).
- Effect of parameter ranges: While the median model performance across all catchments did not vary significantly, the mean model performance increased monotonically with increasing parameter range width. Performance values generally agreed when parameter ranges were not drastically altered.
- Regional vs. random parameter sets: Regional parameter set ensembles (calibrated sets from other catchments) generally resulted in slightly higher NPE values (approximately 0.05 units) for the lower benchmark compared to randomly generated sets in catchments where random sets already yielded NPE > 0.5.
- Benchmark values: Median upper benchmark values were 0.86 (NPE), 0.85 (KGE), and 0.75 (NSE). Median lower benchmark values were 0.68 (NPE), 0.49 (KGE), and 0.47 (NSE).
- Spatial patterns: Lower benchmark values exhibited greater variability across and within datasets than upper benchmark values, with higher values observed in humid coastal regions (e.g., US east and west coasts) and lower values in drier central regions.
- Predictability of benchmarks: Runoff ratio, aridity, high flows (Q95), and mean flow (Qmean) were the most important catchment attributes for predicting both upper and lower benchmark performance, with performance generally increasing with wetness and decreasing with aridity.
Contributions
- Provides a comprehensive framework and practical guidance for computing robust lower and upper benchmarks for hydrological model performance.
- Delivers a valuable resource of pre-computed lower and upper benchmark values for over 6000 catchments across 13 widely used large-sample datasets.
- Demonstrates the inadequacy of using fixed performance thresholds for model evaluation across diverse hydroclimatic conditions.
- Offers trained random forest models, enabling the estimation of benchmarks for catchments not included in the study.
- Recommends specific methodological choices for benchmark computation, such as using 1000 ensemble members and aggregating by ensemble mean.
Funding
This study was supported by the availability of various large-sample datasets and cloud computing infrastructure provided by Science IT (S3IT) at the University of Zurich. No specific project or program reference codes were listed.
Citation
@article{Seibert2026Setting,
author = {Seibert, Jan and Vis, Marc J. P. and Pool, Sandra},
title = {Setting the bar: benchmarks for model performances in large-sample hydrology},
journal = {Hydrology and earth system sciences},
year = {2026},
doi = {10.5194/hess-30-5857-2026},
url = {https://doi.org/10.5194/hess-30-5857-2026}
}
Original Source: https://doi.org/10.5194/hess-30-5857-2026