Methodology
Avow Methodology
Version 1.0.0 · Every published score states the methodology version that produced it.
Avow measures how accurate a public figure's on-the-record predictions have been since January 1, 2011. A score is never an opinion about a person — it is arithmetic over a published list of verifiable records, each traceable to a verbatim quote, a dated source, and an archived snapshot. Anyone can re-derive any score by hand from the published records and the formula in §5.
1. What counts as a prediction
Our full inclusion criteria are documented in full; in summary — a statement is scoreable only if all of the following hold:
- Future-oriented at the time it was made.
- Falsifiable with public evidence.
- Committed — the speaker asserts a direction or an explicit probability. Pure "could/might happen" musings don't count.
- Time-bound — a deadline stated or reasonably inferable under fixed rules. A bare price/level target with no stated date qualifies via a default 3-year horizon (and is resolved on whether the target was reached at any point within that window).
- Own voice, serious register — not quoting someone else, not a hypothetical, not comedy.
Statements that are prediction-flavored but fail the bar (vague doom-saying, "eventually...") are classified implied: they are verified and displayed on the pundit's timeline, but never scored. This makes hedging itself visible without pretending we can grade it.
2. How predictions are gathered
Two routes per pundit, both disclosed in a per-pundit coverage ledger:
- Open search — agents search the public web, retrospectives, interviews, and social archives. Every candidate must be traced back to a primary source.
- Systematic sweep, by medium — every medium the pundit materially uses is enumerated as its own corpus under a documented sampling rule: the text site/column archive, the video channel (transcripts/captions are the primary source), podcast appearances (published transcripts or directly-quoting show notes), and books (dated, falsifiable theses). Video, podcast, and book sweeps are first-class, never folded into "open search" — a TV/radio figure's words live in audio and video, so we go there rather than settling for paraphrase.
Two honesty mechanisms run on top. A recall audit reads a random sample of the items a forecast-keyword filter *didn't* flag, to estimate roughly what fraction of forecasts the filter missed — so the sampling rule is honest about false negatives, not just about what it caught. And a completeness loop keeps expanding (deeper sampling, more media, fresh angles) across bounded rounds until the resolved-scoreable base reaches the provisional threshold (20 resolved — see §5) or a documented ceiling fires (corpus exhausted / round cap / budget); the ledger records the round count and the stop reason. The loop is mechanical and fully autonomous — a shortfall marks the pundit provisional and is logged, never paused on.
The ledger states exactly what was swept per medium, how much of it, the estimated recall, the completeness trail, and what was not. An Avow score is therefore a claim about a documented sample of the public record, not about everything a person has ever said. We treat selection bias as the primary threat to validity and publish coverage statistics rather than hiding them.
3. Verification
Before anything is scored or displayed, independent verification agents attempt to reject each record:
- Does the verbatim quote exist at the cited source (or its archive)?
- Is it genuinely a prediction by this person, in their own voice?
- Are the date, deadline, and attribution correct?
- Does the cited outcome evidence actually support the proposed resolution?
Records emerge verified, corrected (with the correction logged), rejected, or flagged for human review. A record whose primary source cannot be fetched is unverifiable and is neither scored nor displayed until a human confirms it.
Quotes are tiered by provenance. A quote fetched from the pundit's own channel, or a first-person direct quotation carried by a reputable outlet, can be scored. A third-person paraphrase ("the host says Cramer expects…") is never treated as a verbatim and cannot score until the pundit's actual words are found — this matters most for TV/radio figures, whose words live in video rather than fetchable text.
4. Quantification
Three independent judges score each verified record against fixed, versioned rubrics; we take the median and store all three judgments, including dissent:
- Importance (1–5) — stakes of the claim at the time it was made, per our published importance rubric (1–5). No hindsight.
- Stated certainty (p) — the pundit's verbal confidence mapped to a probability via a fixed, published certainty-to-probability table. Explicit numbers ("30% chance") are used as stated. The verbatim certainty phrase is always stored so readers can dispute the mapping.
- Outcome (o, 0–1) — the degree to which the asserted outcome occurred, judged only against the resolution criterion written before evidence review, per our published resolution rubric. Partial credit follows fixed bands; deadlines are strict (late = miss, annotated as a timing miss).
Judges disagreeing beyond fixed thresholds send the record to human review; it is excluded from scoring until resolved.
5. The score
Computed only by our scorer (formula v1.0.0), a deterministic rule specified in full here:
per prediction: s = 1 − 2·(p − o)² Brier-style proper scoring rule
weight: w = importance (1–5; certainty is NOT in the weight —
the scoring rule already prices it)
raw skill: S = Σ w·s / Σ w
grade (0–100): G = 100 · max(0, (N·S + k·0.5) / (N + k)), k = 10
headline (Avow): A = (G − 50) · 2 signed score in [−100, +100]Properties of this scale:
- 0 = a chronic hedger who never commits beyond 50/50 — also the score of an empty record. Provisional pundits sit near 0, because shrinkage pulls them there.
- +100 = bold and consistently right. A negative score is an *anti-signal*: the person's confident calls subtracted information, and a reader would have done better fading them; −100 is maximally, confidently wrong.
- The signed headline is a pure recentering of the underlying Brier skill on its neutral point (the hedger) — presentation only, never the ranking or the data. The intermediate skill values (before and after shrinkage) are recorded in the published score record, alongside a description of the scale, so anyone can re-derive the number.
- The shrinkage term (k = 10 pseudo-predictions at the prior) keeps a handful of lucky calls from outscoring a long solid record. Scores from fewer than 20 resolved predictions are labeled provisional (and should not be shown with a definitive verbal band), and N is always displayed.
Worked example: a pundit says a recession is "almost certain" within a year (p = 0.90, importance 5) and none occurs (o = 0): s = 1 − 2(0.81) = −0.62 — a heavily weighted hit to the score. The same miss said as "I think" (p = 0.65) gives s = 0.155 — hedging is punished less when wrong, but also earns less when right.
Alongside the headline score we publish the components: a calibration table ("when they said 'definitely', they were right X% of the time"), the high-confidence hit rate, N, coverage statistics, and each record's individual contribution.
Per-tag sub-scores
Every prediction is tagged with 1–3 domains from a small controlled vocabulary (a small controlled vocabulary of domains). The scorer computes a separate Avow score for each tag, using the identical formula restricted to that tag's predictions, and surfaces any tag with at least a few resolved predictions (provisional under 20, like the overall). A prediction with multiple tags counts toward each of its tags, so the per-tag scores are not a partition of the overall — and the overall is not their average; it counts every prediction exactly once. This lets a multi-domain forecaster — sharp on geopolitics, weak on market timing — be read honestly per domain instead of blurred into a single number that describes neither.
6. Exclusions
Excluded from scoring (but listed with reasons in the published score record):
- Pending predictions (deadline not passed) — shown on the timeline as pending.
- Conditionals whose condition was not met — void.
- Records awaiting human review, unverifiable records, implied-class statements.
- Statements made before 2011-01-01.
- Near-duplicate repetitions of one claim within 90 days are consolidated into a single record (annual re-assertions score separately — a broken clock is many separate wrong clocks).
7. Known limitations
- Selection bias. Even with systematic sweeps, the gathered record is a sample. The coverage ledger is the honest disclosure, not a cure.
- Verbal probability mapping is lossy. "Likely" is not exactly 0.80. The mapping is fixed, public, and applied uniformly, and the verbatim phrase is always shown.
- LLM judgment noise. Research, verification, and quantification use AI agents. Mitigations: adversarial verification, three-judge medians with stored dissent, human review queues, and full audit trails. Errors will still occur; see corrections.
- Resolution ambiguity. Some outcomes resist clean resolution. Such records carry partial credit with rationale or are flagged rather than force-scored.
8. Corrections and right of reply
Every prediction record has a stable ID and a public audit trail. Anyone — including the pundit — can dispute a quote, a resolution, or a mapping. Disputes are logged, re-verified, and corrections are published with the record's history preserved. Scores recompute automatically from corrected data.
9. Versioning
The methodology, rubrics, certainty mapping, and scoring formula carry a single current version. Once scores are published, any change to them bumps the version and triggers recomputation of affected pundits, and every published score states the version that produced it. Every published score records the methodology version that produced it.