Method
A written-down model, so the number can be rebuilt.
The method is published with the first case.
Scoring rule
Version 2. Dated 2026-09-17. Supersedes version 1, archive/SCORING_RULE.md,
dated 2026-09-09, which is not edited and stays in force for every case that names
it.
This document is fixed. It is never edited. If it is found to be wrong, a new version is written as a new document and applies only to cases issued after that version's date. Cases are scored under the version in force when they were issued.
This is a specification. It defines how a stored case's forecast is compared with what happened, precisely enough that two parties working independently reach the same number.
What changed from version 1, and why
On 2026-09-17 FRED served DCOILBRENTEU with blank values for 63 historical days
it had previously published: 33 in 2005, 26 in 2006, 3 in 2007 and 1 in 2014. No
published value changed; the days are absent. March 2006 now carries 9 observed
values, March 2005 12, and April 2005 and April 2006 13 each, all below the minimum
in section 2.3, so a score for any of them would defer.
Version 1, section 2.2, described the observed count as ranging from 17 to 23 values a month. That described one fetch of the series, not the rule, and the revision made it false. Version 1 is fixed and was not edited; this version replaces it. Three changes, and no other:
- Section 2.2 no longer states an observed range. The rule is the minimum in section 2.3, or defer. The minimum, the earliest computation date, the aggregation, the scoring month and the band comparison are unchanged.
- Section 2.4 records the revision above.
- Section 7 adds three fields every score records -- the daily values it averaged with their dates, the series identifier, and when the series was fetched -- so a score can be re-derived from its own file after the source changes.
A case names the rule it was issued under in its scoring_rule field, and is
scored under that rule (section 9).
1. Source series
| Series | DCOILBRENTEU |
| Title | Crude Oil Prices: Brent — Europe |
| Publisher | US Energy Information Administration, relayed by FRED |
| Host | Federal Reserve Bank of St. Louis (FRED) |
| Units | US dollars per barrel |
| Frequency | Daily |
| Seasonal adjustment | None |
No other series is used. No substitute is adopted without a new version of this document.
2. Realised price for a month
The realised price of a calendar month is:
the arithmetic mean of all observed daily values of
DCOILBRENTEUwhose observation date falls within that calendar month.
2.1 Which days count
Every day the publisher reports a value for. In practice these are weekdays; weekends never appear in the series.
A day with no published value does not count and is not filled. It is not interpolated, not carried forward from the previous day, not carried back, and not replaced by any estimate. It is absent from both the sum and the count. Public holidays are the usual reason a weekday is absent.
2.2 Observation counts vary and must be read, not assumed
The number of observed values differs from month to month, and for the same month from one fetch of the series to the next. No count is assumed. The divisor is the count of values actually observed in that month, whatever it is. Whether that count is enough is decided by section 2.3 and by nothing else.
2.3 Minimum observation count
A month is not scorable with fewer than 15 observed values.
When a month has fewer than 15, no score is computed and the scoring is deferred to a later date. Deferral does not change the scoring date, the month scored, or any figure in the case; it changes only when the computation runs.
A month that still has fewer than 15 observed values one year after its end is recorded as unscorable, with its observation count and the date that determination was made.
2.4 The revision of 2026-09-17
FRED blanked 63 historical days of DCOILBRENTEU on 2026-09-17 (33 in 2005, 26 in
2006, 3 in 2007, 1 in 2014) with no value changed. March 2006 now carries 9
observed values and would defer; March 2005 (12), April 2005 (13) and April 2006
(13) would defer too. The rule is at least 15 observed values, or defer.
3. Which month is scored
A stored case carries a scoring_dates map from horizon to date, and each
scoring date is a month end.
The month scored at horizon h is the calendar month that
scoring_dates[h]is the last day of.
3.1 Worked example
Case brent-2026-08-31-live-0ed0022-dirty-d0133d91c37e83d3, horizon 3.
- Read
scoring_dates["3"]from the case file →2026-11-30. - That date is the last day of November 2026. November 2026 is the month scored.
- Collect every observed
DCOILBRENTEUvalue with observation date from2026-11-01to2026-11-30inclusive. - If fewer than 15 values, stop and defer (section 2.3).
- The realised price is their arithmetic mean (section 4).
- Read
structural.pathsfrom the same case file, take the row wherehorizon_months == 3, givingp10 = 50.0589,p50 = 87.0993,p90 = 131.8011. - Compare (section 5).
The earliest this may be computed is 2026-12-10 (section 6).
4. Arithmetic and precision
The realised price is computed as:
realised = sum(values in observation-date order) / count(values)
in IEEE-754 binary64.
Source values carry one or two decimal places. The realised price is published to four decimal places. The comparison in section 5 uses the unrounded value, not the published one.
Two independent computations of the same month must agree to within 1e-9. A
larger difference means the two parties used a different set of observations,
not a different summation order.
5. Comparison with the band
Let realised be the value from section 4, and p10, p90 the values read
from the case's structural.paths row for that horizon.
The outcome is inside the band when
p10 <= realised <= p90.
The band is inclusive at both ends. A realised price exactly equal to p10
or exactly equal to p90 is inside.
The band values are read from the stored case. They are never recomputed.
The error at that horizon is:
error = realised - p50
6. Earliest computation date
A score for month M may be computed no earlier than the tenth day of the month following M.
The publisher's lag on the single vintage available at the time of writing was two days: the series held observations to 2026-09-01 when fetched on 2026-09-03. The tenth is set well beyond that. The minimum observation count in section 2.3 is the binding guard; this date is a floor beneath it.
Both conditions must hold. A score is computed only when the date has passed and the month has at least 15 observed values.
7. A published score is never recomputed
Once a score has been computed and written, it is final. It is not recomputed, not corrected, not superseded in place, and not recomputed when the source series later revises.
Every written score records:
| field | meaning |
|---|---|
computed_at |
the date the computation was run |
source_as_of |
the date of the most recent observation present in the series at that time |
observations |
the count of daily values used |
realised |
the value from section 4 |
daily_values |
every daily value averaged, with its observation date |
source_series |
the series identifier, DCOILBRENTEU |
fetched_at |
when the series was fetched |
With daily_values a score can be re-derived from its own file, whatever the
source says later.
If DCOILBRENTEU later restates a value inside a month already scored, the
restatement is recorded as input drift against the case
(<case_id>@<date>.drift.json). The score itself does not change.
config/input_revisions.py declares this series as one that may restate.
A score written under this rule is a statement about what the series said on
source_as_of, not about what it says today.
8. If the source is discontinued
No substitute series is adopted automatically, and no series is silently swapped in.
If DCOILBRENTEU ceases publication:
- Months whose full span was published before the last available observation are scored normally under this document.
- Months whose span extends beyond the last available observation are recorded as unscorable, with the reason, the last available observation date, and the observation count reached.
- Cases with outstanding horizons in that state remain in the archive unchanged. They are not withdrawn and their forecasts are not restated.
- Adopting a replacement series requires a new version of this document. The replacement applies only to cases issued on or after that version's date. Cases issued earlier are never rescored against a different series.
9. Applicability
A case names its scoring rule in its scoring_rule field, and the scorer reads
the rule from the case. This version applies to every case whose scoring_rule
is archive/SCORING_RULE_v2.md, which is every case issued on or after
2026-09-17. A case that names version 1 is scored under version 1. A case that
names no rule was issued before the field existed, on or after 2026-09-09, and is
scored under version 1.
The two versions compute the same realised price, the same minimum and the same comparison; they differ in what section 2.2 says and in the fields a score records.
Cases issued before 2026-09-09 were issued without frozen inputs and carry
"inputs_frozen": false. They may be scored under version 1, and a score so
computed states which case it scored; the absence of frozen inputs affects
whether the case's forecast can be re-derived, not how the realised price is
computed.
Calibration rule
Version 1. Dated 2026-09-09.
This document is fixed. It is never edited. If it is found to be wrong, a new version is written as a new document and applies only to cases issued after that version's date. Cases are counted under the version in force when they were issued.
This is a specification. It defines which stored cases enter the calibration record, precisely enough that two parties working independently reach the same count.
It governs archive/calibration.json only. It does not govern scoring, which
is archive/SCORING_RULE.md, and it changes no score.
1. What is counted
The calibration record counts scoreable claims, not stored cases.
A scoreable claim is one unique fit: one pair of
data_hashandas_of_effective.
Two cases carrying the same pair are the same fit. However many cases were issued from it, they are one claim.
data_hash and as_of_effective are read from the case file. No other field
enters the identity: not code_version, not issued_at, not kind, not
raw_input_hash.
2. Which case counts
When more than one case shares a fit:
The case with the earliest
issued_atis the one that counts.
The others are re-issues. They are excluded from the calibration count.
2.1 Tie-break
If two cases sharing a fit carry an identical issued_at, the one whose
case_id sorts first as a byte string counts. This is deterministic and needs
no information outside the case files.
2.2 What exclusion does not mean
An excluded case is not deleted, not edited, not hidden and not unscorable.
- It remains in
archive/cases/unchanged. - It may be scored individually under
SCORING_RULE.md, and its score is written and kept like any other. - Its score is not counted in
calibration.json. - It is listed in
calibration.jsonby id, with the case that supersedes it and the reason.
3. Rationale
Recorded because a count is only as trustworthy as the reason behind it.
A re-issue under a new commit is not a new forecast. The model said one thing about one month on one set of inputs. Issuing that same fit again — after a refactor, a rename, a code change that did not alter the numbers — produces a second case file but not a second claim about the world.
The first issue is the one that was committed to at that time. That is what makes it the claim: it was published before the outcome was known, and nothing later can change when it was said.
The direction of the error decides the rule. Counting one fit twice
inflates scores_counted and computes coverage over a claim counted twice. If
that fit scored inside its band, coverage rises on a duplicate; if it scored
outside, the sample is padded. Either way the record reads better-founded than
it is, in the author's favour. That is the error that destroys a calibration
record, so the rule resolves against it.
4. What the record states
archive/calibration.json carries, alongside its counts:
| field | meaning |
|---|---|
cases_in_archive |
how many case files exist |
claims_counted |
how many unique fits those cases represent |
excluded_as_reissue |
how many cases were excluded under section 2 |
claims |
per counted claim: the fit, the case that counts, its issue date |
reissues |
per excluded case: which case supersedes it, and why |
A reader must be able to see the exclusion, not infer it from a number that is smaller than the file count.
5. This rule does not change a score
Scores are written under SCORING_RULE.md and are final. This rule decides
which of them are counted, and nothing else. Excluding a case from the count
does not withdraw, amend or annotate its score.
6. Applicability
This version applies to the calibration record built on or after 2026-09-09, over all cases in the archive whatever their issue date.
The record is rebuilt in full from the cases and scores on disk every time. It is never incremented, so applying this rule cannot require any stored value to be revised.
Archive policy
Version 1. Dated 2026-09-09.
This document is fixed. It is never edited. If it is found to be wrong, a new version is written as a new document.
This is a specification. It states what may be done to the case archive, and when.
1. The archive was cleared once
On 2026-09-09 the case archive was cleared. Five files were deleted:
| bytes | file |
|---|---|
| 14,762 | brent-2026-08-31-live-4b00be9-dirty-d0133d91c37e83d3.json |
| 14,698 | brent-2026-08-31-live-0ed0022-dirty-d0133d91c37e83d3.json |
| 183,560 | brent-2026-08-31-live-0ed0022-dirty-d0133d91c37e83d3.inputs.json |
| 3,492 | brent-2026-08-31-live-0ed0022-dirty-...@2026-09-09.drift.json |
| 1,931 | calibration.json |
Both cases were produced while the machinery that writes them was being built. Neither was published. Neither was signed. Neither was relied on by anything. No public repository existed at the time of the clearing.
The clearing was a normal deletion commit. Git history was not rewritten, nothing was force-pushed and nothing was amended. The commits that introduced those files stand, and every deleted file can still be recovered from history. That these cases existed, and that they were removed deliberately and when, is itself part of the record.
2. From the first case issued after this date, the archive is append-only
Nothing in
archive/cases/is ever deleted, edited, or rewritten.
This applies to case files, frozen inputs, drift observations, scores, deferral logs and score revisions alike.
archive/calibration.json is the single exception and is not a record: it is a
view, rebuilt in full from the cases and scores on every run, holding nothing
they do not. Overwriting it restates what they already say.
3. A wrong case is corrected by issuing the next one
A case that turns out to be wrong — wrong inputs, wrong configuration, a defect in the model — is not removed and not amended.
It is corrected by issuing the next case, which carries its own inputs, its own timestamp and its own note stating what was wrong with the earlier one. The earlier case stays exactly as issued.
A record that can be edited proves nothing, whether or not it ever is. The value of the archive is that what was said before an outcome was known cannot be changed after it.
4. Clearing is not available again
Section 1 records a one-time act, performed before any case had been published, under conditions that no longer hold and cannot recur: there was no published record to protect, because there was no record.
From the first case issued after 2026-09-09 there is one, and section 2 governs it without exception. A future clearing would not be a housekeeping decision; it would be the destruction of the thing the archive exists to be.
5. What enforces this
Nothing in this repository deletes from archive/. There is no delete
function, no cleanup routine, no retention policy and no expiry.
tests/test_archive_policy.py asserts it by scanning the source.
That is a guard, not a guarantee. A person with a shell can delete any file. What makes the archive append-only in practice is that it is committed to git: a deletion is itself a commit, visible in the history, and the deleted content remains recoverable from it. The policy is enforced by the record of breaking it being permanent.