
Fig. 1
The three validation metrics – RPS score (top row), spread-skill comparison (centre row) and rank histogram (bottom rows) – illustrated with an under-dispersive EPS (left column), a rather well-tuned EPS (centre column) and an over-dispersed EPS (right column).

Fig. 2
Examples of ensembles launched at different times for the first observed model component (third element in the state vector). The grey envelope is the 95% confidence envelope estimated from the ensemble. Daily observations (red dots) are used for constructing the likelihood.

Fig. 3
The negative log-likelihood values for the first two parameters, the initial-value perturbation magnitude and the variance parameter of the stochastic forcing σ e .

Fig. 4
The conditional likelihood values for the stochastic physics parameters with (left) and (right).

Fig. 5
Negative log-likelihood values vs. RPS scores. The colours indicate the value of the initial-value perturbation parameter , red means a large value.

Fig. 6
All 1000 RPS curves with different parameter values (grey lines) and the curves corresponding to the best (green) and 10 best (red) parameter values according to the likelihood calculation.

Fig. 7
Difference between the ensemble spread and the forecast error of the mean for all 1000 tested parameter values (grey lines) and the curves corresponding to the best (green) and 10 best (red) parameter values according to the likelihood calculation.

Fig. 8
Rank histograms for all 1000 tested parameter values (grey lines) and the curves corresponding to the best (green) and 10 best (red) parameter values according to the likelihood calculation. The four different plots represent different forecast lead times.
