Back

Does the model actually work?

Everything else here works out what you should provide. This page checks whether the model that produced it has ever been right. Give it a period that has already finished — the PDs you used at the start, and what actually happened by the end.

Failed1,200 borrowers · worked example

Failed on calibration. The allowance this model produces is understated.

This example is meant to fail. The book is 1,200 synthetic borrowers with one grade deliberately wrong: BB was priced at 2.00% and defaulted at 6.00%.

Notice what still passes. The ranking is sound, and the portfolio total — 3.75% observed against 2.75% predicted — reads as acceptable. That is the whole point: a portfolio-level test nets an understated grade against a conservative one and reports nothing wrong, so calibration is tested grade by grade. Load your own file above to replace this.

Your data

One row per borrower. Two columns are essential: the PD you predicted, and whether it defaulted.

This is an example book of 1,200 synthetic borrowers, not your data. It is here so the page shows what it does before you have fetched anything. Replace it when you are ready.

Read 1,200 rows and 5 columns.

Which column is which

Corrected automatically where the heading is recognisable. Change anything that is wrong.

Whatever identifies the account. Used only to report which rows failed.
The probability of default the model gave BEFORE the outcome was known. A PD refreshed during the period tests hindsight, not prediction.
What actually happened by the end of the period.
Lets the test run per grade. Without it only a portfolio-level result is possible, and a portfolio total can net an understated grade against an overstated one.
Used to weight the average predicted PD. Rows count equally without it.
The PD column is being read as decimals

Every value sits in the decimal range. Read the wrong way round, a model would look a hundred times more conservative than it is and pass every test on this page.

Can it tell good from bad?

Whether the model RANKS risk. A model that cannot rank is not repairable by recalibration.

Passed

Gini 0.568, confidence interval from 0.408 to 0.729, on 45 defaults. The model separates defaulting from performing obligors.

Gini
0.568
AUC
0.7842
95% interval
0.704 – 0.865
KS
0.522
Defaults observed
45
Default rate
3.75%

Are the levels right?

Whether the PDs match what happened — tested per grade, because a total can net an understated grade against an overstated one.

Failed

The portfolio total is acceptable, but 1 grade (BB) understate default risk at the 1% level. A total that nets out an understated grade against an overstated one is not a calibrated model.

GradeBorrowersDefaultsPredictedObservedPredicted vs observedJeffreys p
A30010.30%0.33%
0.3847Passed
BBB35030.90%0.86%
0.4931Passed
BB300182.00%6.00%
<0.0001Failed
B180115.51%6.11%
0.3453Passed
CCC701218.08%17.14%
0.5680Passed
Hosmer–Lemeshow p
0.0002
Spiegelhalter p
0.0084
Brier score
0.03473
— calibration error
0.00041
— resolution
0.00179

Correlation assumption

Used for the bound that allows for defaults arriving together rather than independently.

0.15Set this to the asset correlation your PDs were built with. Higher correlation means a bad year is less surprising, so the test is more forgiving — which is why it must match the measurement rather than be chosen here.