REFERENCE CASE

Issued quarterly by Reservoir GmbH, Hamburg

PAGE

Method

A written-down model, so the number can be rebuilt.

The method is published with the first case.

Scoring rule

Version 2. Dated 2026-09-17. Supersedes version 1, archive/SCORING_RULE.md, dated 2026-09-09, which is not edited and stays in force for every case that names it.

This document is fixed. It is never edited. If it is found to be wrong, a new version is written as a new document and applies only to cases issued after that version's date. Cases are scored under the version in force when they were issued.

This is a specification. It defines how a stored case's forecast is compared with what happened, precisely enough that two parties working independently reach the same number.

What changed from version 1, and why

On 2026-09-17 FRED served DCOILBRENTEU with blank values for 63 historical days it had previously published: 33 in 2005, 26 in 2006, 3 in 2007 and 1 in 2014. No published value changed; the days are absent. March 2006 now carries 9 observed values, March 2005 12, and April 2005 and April 2006 13 each, all below the minimum in section 2.3, so a score for any of them would defer.

Version 1, section 2.2, described the observed count as ranging from 17 to 23 values a month. That described one fetch of the series, not the rule, and the revision made it false. Version 1 is fixed and was not edited; this version replaces it. Three changes, and no other:

  1. Section 2.2 no longer states an observed range. The rule is the minimum in section 2.3, or defer. The minimum, the earliest computation date, the aggregation, the scoring month and the band comparison are unchanged.
  2. Section 2.4 records the revision above.
  3. Section 7 adds three fields every score records -- the daily values it averaged with their dates, the series identifier, and when the series was fetched -- so a score can be re-derived from its own file after the source changes.

A case names the rule it was issued under in its scoring_rule field, and is scored under that rule (section 9).


1. Source series

Series DCOILBRENTEU
Title Crude Oil Prices: Brent — Europe
Publisher US Energy Information Administration, relayed by FRED
Host Federal Reserve Bank of St. Louis (FRED)
Units US dollars per barrel
Frequency Daily
Seasonal adjustment None

No other series is used. No substitute is adopted without a new version of this document.


2. Realised price for a month

The realised price of a calendar month is:

the arithmetic mean of all observed daily values of DCOILBRENTEU whose observation date falls within that calendar month.

2.1 Which days count

Every day the publisher reports a value for. In practice these are weekdays; weekends never appear in the series.

A day with no published value does not count and is not filled. It is not interpolated, not carried forward from the previous day, not carried back, and not replaced by any estimate. It is absent from both the sum and the count. Public holidays are the usual reason a weekday is absent.

2.2 Observation counts vary and must be read, not assumed

The number of observed values differs from month to month, and for the same month from one fetch of the series to the next. No count is assumed. The divisor is the count of values actually observed in that month, whatever it is. Whether that count is enough is decided by section 2.3 and by nothing else.

2.3 Minimum observation count

A month is not scorable with fewer than 15 observed values.

When a month has fewer than 15, no score is computed and the scoring is deferred to a later date. Deferral does not change the scoring date, the month scored, or any figure in the case; it changes only when the computation runs.

A month that still has fewer than 15 observed values one year after its end is recorded as unscorable, with its observation count and the date that determination was made.

2.4 The revision of 2026-09-17

FRED blanked 63 historical days of DCOILBRENTEU on 2026-09-17 (33 in 2005, 26 in 2006, 3 in 2007, 1 in 2014) with no value changed. March 2006 now carries 9 observed values and would defer; March 2005 (12), April 2005 (13) and April 2006 (13) would defer too. The rule is at least 15 observed values, or defer.


3. Which month is scored

A stored case carries a scoring_dates map from horizon to date, and each scoring date is a month end.

The month scored at horizon h is the calendar month that scoring_dates[h] is the last day of.

3.1 Worked example

Case brent-2026-08-31-live-0ed0022-dirty-d0133d91c37e83d3, horizon 3.

  1. Read scoring_dates["3"] from the case file → 2026-11-30.
  2. That date is the last day of November 2026. November 2026 is the month scored.
  3. Collect every observed DCOILBRENTEU value with observation date from 2026-11-01 to 2026-11-30 inclusive.
  4. If fewer than 15 values, stop and defer (section 2.3).
  5. The realised price is their arithmetic mean (section 4).
  6. Read structural.paths from the same case file, take the row where horizon_months == 3, giving p10 = 50.0589, p50 = 87.0993, p90 = 131.8011.
  7. Compare (section 5).

The earliest this may be computed is 2026-12-10 (section 6).


4. Arithmetic and precision

The realised price is computed as:

realised = sum(values in observation-date order) / count(values)

in IEEE-754 binary64.

Source values carry one or two decimal places. The realised price is published to four decimal places. The comparison in section 5 uses the unrounded value, not the published one.

Two independent computations of the same month must agree to within 1e-9. A larger difference means the two parties used a different set of observations, not a different summation order.


5. Comparison with the band

Let realised be the value from section 4, and p10, p90 the values read from the case's structural.paths row for that horizon.

The outcome is inside the band when p10 <= realised <= p90.

The band is inclusive at both ends. A realised price exactly equal to p10 or exactly equal to p90 is inside.

The band values are read from the stored case. They are never recomputed.

The error at that horizon is:

error = realised - p50


6. Earliest computation date

A score for month M may be computed no earlier than the tenth day of the month following M.

The publisher's lag on the single vintage available at the time of writing was two days: the series held observations to 2026-09-01 when fetched on 2026-09-03. The tenth is set well beyond that. The minimum observation count in section 2.3 is the binding guard; this date is a floor beneath it.

Both conditions must hold. A score is computed only when the date has passed and the month has at least 15 observed values.


7. A published score is never recomputed

Once a score has been computed and written, it is final. It is not recomputed, not corrected, not superseded in place, and not recomputed when the source series later revises.

Every written score records:

field meaning
computed_at the date the computation was run
source_as_of the date of the most recent observation present in the series at that time
observations the count of daily values used
realised the value from section 4
daily_values every daily value averaged, with its observation date
source_series the series identifier, DCOILBRENTEU
fetched_at when the series was fetched

With daily_values a score can be re-derived from its own file, whatever the source says later.

If DCOILBRENTEU later restates a value inside a month already scored, the restatement is recorded as input drift against the case (<case_id>@<date>.drift.json). The score itself does not change. config/input_revisions.py declares this series as one that may restate.

A score written under this rule is a statement about what the series said on source_as_of, not about what it says today.


8. If the source is discontinued

No substitute series is adopted automatically, and no series is silently swapped in.

If DCOILBRENTEU ceases publication:

  1. Months whose full span was published before the last available observation are scored normally under this document.
  2. Months whose span extends beyond the last available observation are recorded as unscorable, with the reason, the last available observation date, and the observation count reached.
  3. Cases with outstanding horizons in that state remain in the archive unchanged. They are not withdrawn and their forecasts are not restated.
  4. Adopting a replacement series requires a new version of this document. The replacement applies only to cases issued on or after that version's date. Cases issued earlier are never rescored against a different series.

9. Applicability

A case names its scoring rule in its scoring_rule field, and the scorer reads the rule from the case. This version applies to every case whose scoring_rule is archive/SCORING_RULE_v2.md, which is every case issued on or after 2026-09-17. A case that names version 1 is scored under version 1. A case that names no rule was issued before the field existed, on or after 2026-09-09, and is scored under version 1.

The two versions compute the same realised price, the same minimum and the same comparison; they differ in what section 2.2 says and in the fields a score records.

Cases issued before 2026-09-09 were issued without frozen inputs and carry "inputs_frozen": false. They may be scored under version 1, and a score so computed states which case it scored; the absence of frozen inputs affects whether the case's forecast can be re-derived, not how the realised price is computed.

SCORING_RULE_v2.md
The file above is the authoritative text.

Calibration rule

Version 1. Dated 2026-09-09.

This document is fixed. It is never edited. If it is found to be wrong, a new version is written as a new document and applies only to cases issued after that version's date. Cases are counted under the version in force when they were issued.

This is a specification. It defines which stored cases enter the calibration record, precisely enough that two parties working independently reach the same count.

It governs archive/calibration.json only. It does not govern scoring, which is archive/SCORING_RULE.md, and it changes no score.


1. What is counted

The calibration record counts scoreable claims, not stored cases.

A scoreable claim is one unique fit: one pair of data_hash and as_of_effective.

Two cases carrying the same pair are the same fit. However many cases were issued from it, they are one claim.

data_hash and as_of_effective are read from the case file. No other field enters the identity: not code_version, not issued_at, not kind, not raw_input_hash.


2. Which case counts

When more than one case shares a fit:

The case with the earliest issued_at is the one that counts.

The others are re-issues. They are excluded from the calibration count.

2.1 Tie-break

If two cases sharing a fit carry an identical issued_at, the one whose case_id sorts first as a byte string counts. This is deterministic and needs no information outside the case files.

2.2 What exclusion does not mean

An excluded case is not deleted, not edited, not hidden and not unscorable.

  • It remains in archive/cases/ unchanged.
  • It may be scored individually under SCORING_RULE.md, and its score is written and kept like any other.
  • Its score is not counted in calibration.json.
  • It is listed in calibration.json by id, with the case that supersedes it and the reason.

3. Rationale

Recorded because a count is only as trustworthy as the reason behind it.

A re-issue under a new commit is not a new forecast. The model said one thing about one month on one set of inputs. Issuing that same fit again — after a refactor, a rename, a code change that did not alter the numbers — produces a second case file but not a second claim about the world.

The first issue is the one that was committed to at that time. That is what makes it the claim: it was published before the outcome was known, and nothing later can change when it was said.

The direction of the error decides the rule. Counting one fit twice inflates scores_counted and computes coverage over a claim counted twice. If that fit scored inside its band, coverage rises on a duplicate; if it scored outside, the sample is padded. Either way the record reads better-founded than it is, in the author's favour. That is the error that destroys a calibration record, so the rule resolves against it.


4. What the record states

archive/calibration.json carries, alongside its counts:

field meaning
cases_in_archive how many case files exist
claims_counted how many unique fits those cases represent
excluded_as_reissue how many cases were excluded under section 2
claims per counted claim: the fit, the case that counts, its issue date
reissues per excluded case: which case supersedes it, and why

A reader must be able to see the exclusion, not infer it from a number that is smaller than the file count.


5. This rule does not change a score

Scores are written under SCORING_RULE.md and are final. This rule decides which of them are counted, and nothing else. Excluding a case from the count does not withdraw, amend or annotate its score.


6. Applicability

This version applies to the calibration record built on or after 2026-09-09, over all cases in the archive whatever their issue date.

The record is rebuilt in full from the cases and scores on disk every time. It is never incremented, so applying this rule cannot require any stored value to be revised.

CALIBRATION_RULE.md
The file above is the authoritative text.

Archive policy

Version 1. Dated 2026-09-09.

This document is fixed. It is never edited. If it is found to be wrong, a new version is written as a new document.

This is a specification. It states what may be done to the case archive, and when.


1. The archive was cleared once

On 2026-09-09 the case archive was cleared. Five files were deleted:

bytes file
14,762 brent-2026-08-31-live-4b00be9-dirty-d0133d91c37e83d3.json
14,698 brent-2026-08-31-live-0ed0022-dirty-d0133d91c37e83d3.json
183,560 brent-2026-08-31-live-0ed0022-dirty-d0133d91c37e83d3.inputs.json
3,492 brent-2026-08-31-live-0ed0022-dirty-...@2026-09-09.drift.json
1,931 calibration.json

Both cases were produced while the machinery that writes them was being built. Neither was published. Neither was signed. Neither was relied on by anything. No public repository existed at the time of the clearing.

The clearing was a normal deletion commit. Git history was not rewritten, nothing was force-pushed and nothing was amended. The commits that introduced those files stand, and every deleted file can still be recovered from history. That these cases existed, and that they were removed deliberately and when, is itself part of the record.


2. From the first case issued after this date, the archive is append-only

Nothing in archive/cases/ is ever deleted, edited, or rewritten.

This applies to case files, frozen inputs, drift observations, scores, deferral logs and score revisions alike.

archive/calibration.json is the single exception and is not a record: it is a view, rebuilt in full from the cases and scores on every run, holding nothing they do not. Overwriting it restates what they already say.


3. A wrong case is corrected by issuing the next one

A case that turns out to be wrong — wrong inputs, wrong configuration, a defect in the model — is not removed and not amended.

It is corrected by issuing the next case, which carries its own inputs, its own timestamp and its own note stating what was wrong with the earlier one. The earlier case stays exactly as issued.

A record that can be edited proves nothing, whether or not it ever is. The value of the archive is that what was said before an outcome was known cannot be changed after it.


4. Clearing is not available again

Section 1 records a one-time act, performed before any case had been published, under conditions that no longer hold and cannot recur: there was no published record to protect, because there was no record.

From the first case issued after 2026-09-09 there is one, and section 2 governs it without exception. A future clearing would not be a housekeeping decision; it would be the destruction of the thing the archive exists to be.


5. What enforces this

Nothing in this repository deletes from archive/. There is no delete function, no cleanup routine, no retention policy and no expiry. tests/test_archive_policy.py asserts it by scanning the source.

That is a guard, not a guarantee. A person with a shell can delete any file. What makes the archive append-only in practice is that it is committed to git: a deletion is itself a commit, visible in the history, and the deleted content remains recoverable from it. The policy is enforced by the record of breaking it being permanent.

ARCHIVE_POLICY.md
The file above is the authoritative text.