Methodology

Avow Methodology

Version 1.0.0 · Every published score states the methodology version that produced it.

Avow measures how accurate a public figure's on-the-record predictions have been since January 1, 2011. A score is never an opinion about a person — it is arithmetic over a published list of verifiable records, each traceable to a verbatim quote, a dated source, and an archived snapshot. Anyone can re-derive any score by hand from the published records and the formula in §5.

1. What counts as a prediction

Our full inclusion criteria are documented in full; in summary — a statement is scoreable only if all of the following hold:

  1. Future-oriented at the time it was made.
  2. Falsifiable with public evidence.
  3. Committed — the speaker asserts a direction or an explicit probability. Pure "could/might happen" musings don't count.
  4. Time-bound — a deadline stated or reasonably inferable under fixed rules. A bare price/level target with no stated date qualifies via a default 3-year horizon (and is resolved on whether the target was reached at any point within that window).
  5. Own voice, serious register — not quoting someone else, not a hypothetical, not comedy.

Statements that are prediction-flavored but fail the bar (vague doom-saying, "eventually...") are classified implied: they are verified and displayed on the pundit's timeline, but never scored. This makes hedging itself visible without pretending we can grade it.

2. How predictions are gathered

Two routes per pundit, both disclosed in a per-pundit coverage ledger:

  • Open search — agents search the public web, retrospectives, interviews, and social archives. Every candidate must be traced back to a primary source.
  • Systematic sweep, by medium — every medium the pundit materially uses is enumerated as its own corpus under a documented sampling rule: the text site/column archive, the video channel (transcripts/captions are the primary source), podcast appearances (published transcripts or directly-quoting show notes), and books (dated, falsifiable theses). Video, podcast, and book sweeps are first-class, never folded into "open search" — a TV/radio figure's words live in audio and video, so we go there rather than settling for paraphrase.

Two honesty mechanisms run on top. A recall audit reads a random sample of the items a forecast-keyword filter *didn't* flag, to estimate roughly what fraction of forecasts the filter missed — so the sampling rule is honest about false negatives, not just about what it caught. And a completeness loop keeps expanding (deeper sampling, more media, fresh angles) across bounded rounds until the resolved-scoreable base reaches the provisional threshold (20 resolved — see §5) or a documented ceiling fires (corpus exhausted / round cap / budget); the ledger records the round count and the stop reason. The loop is mechanical and fully autonomous — a shortfall marks the pundit provisional and is logged, never paused on.

The ledger states exactly what was swept per medium, how much of it, the estimated recall, the completeness trail, and what was not. An Avow score is therefore a claim about a documented sample of the public record, not about everything a person has ever said. We treat selection bias as the primary threat to validity and publish coverage statistics rather than hiding them.

3. Verification

Before anything is scored or displayed, independent verification agents attempt to reject each record:

  • Does the verbatim quote exist at the cited source (or its archive)?
  • Is it genuinely a prediction by this person, in their own voice?
  • Are the date, deadline, and attribution correct?
  • Does the cited outcome evidence actually support the proposed resolution?

Records emerge verified, corrected (with the correction logged), rejected, or flagged for human review. A record whose primary source cannot be fetched is unverifiable and is neither scored nor displayed until a human confirms it.

Quotes are tiered by provenance. A quote fetched from the pundit's own channel, or a first-person direct quotation carried by a reputable outlet, can be scored. A third-person paraphrase ("the host says Cramer expects…") is never treated as a verbatim and cannot score until the pundit's actual words are found — this matters most for TV/radio figures, whose words live in video rather than fetchable text.

Deleted posts. Some pundits delete their posts or wipe their account, and a prediction does not stop having been made because the evidence of it was removed. When the original post is unretrievable and no archive capture of it exists, we score the quote only if its identical wording is carried by at least three mutually independent sources, or by two where one of them is a dated capture showing the handle and timestamp. The test is independence, not volume: a dozen outlets reproducing a single screenshot count as one source, as does syndicated wire copy or anything visibly derived from another item on the list. If the wording diverges between sources, only the span they agree on can be quoted. Every corroborating source is listed on the record so a reader can recount them, and anything short of the bar stays unscored. The order is always archive capture first, reputable direct quote second, corroboration last.

4. Quantification

Three independent judges score each verified record against fixed, versioned rubrics; we take the median and store all three judgments, including dissent:

  • Importance (1–5) — stakes of the claim at the time it was made, per our published importance rubric (1–5). No hindsight.
  • Stated certainty (p) — the pundit's verbal confidence mapped to a probability via a fixed, published certainty-to-probability table. Explicit numbers ("30% chance") are used as stated. The verbatim certainty phrase is always stored so readers can dispute the mapping.
  • Outcome (o, 0–1) — the degree to which the asserted outcome occurred, judged only against the resolution criterion written before evidence review, per our published resolution rubric. Partial credit follows fixed bands; deadlines are strict (late = miss, annotated as a timing miss).

Judges disagreeing beyond fixed thresholds send the record to human review; it is excluded from scoring until resolved.

5. The score

Computed only by our scorer (formula v1.0.0), a deterministic rule specified in full here:

per prediction:   s = 1 − 2·(p − o)²        Brier-style proper scoring rule
weight:           w = importance             (1–5; certainty is NOT in the weight —
                                              the scoring rule already prices it)
raw skill:        S = Σ w·s / Σ w
shrunk skill:     S' = (N·S + k·0.5) / (N + k),   k = 10
headline (Avow):  A = 100 · (S' − 0.5) / span,    signed score in [−100, +100]
                  where span = 0.5  if S' ≥ 0.5   (the hedger→perfect arm)
                             = 1.5  if S' < 0.5   (the hedger→worst arm)

The two arms carry different slopes because the scoring rule is asymmetric about the hedger, and that asymmetry is a fact about the rule rather than a presentation choice: s runs over [−1, 1] and a hedger scores 0.5, so perfect play sits 0.5 above the hedger while maximal wrongness sits 1.5 below it. Being confidently wrong is three times as far from a hedger as being perfectly right. No single linear map reaches both ±100 anchors, so each arm is scaled over its own span.

Properties of this scale:

  • 0 = a chronic hedger who never commits beyond 50/50 — also the score of an empty record. Provisional pundits sit near 0, because shrinkage pulls them there.
  • +100 = bold and consistently right. A negative score is an *anti-signal*: the person's confident calls subtracted information, and a reader would have done better fading them; −100 is maximally, confidently wrong.
  • The signed headline is a pure recentering of the underlying Brier skill on its neutral point (the hedger) — presentation only, never the ranking or the data. The intermediate skill values (before and after shrinkage) are recorded in the published score record, alongside a description of the scale, so anyone can re-derive the number.
  • Every score is published with an interval. A percentile bootstrap (4,000 resamples, fixed seed, so the scorer stays a pure function of the data) resamples the scored records *together with* the k shrinkage pseudo-predictions and re-runs the identical arithmetic, giving the 95% range the score could plausibly sit in. Two deliberate choices: it is not a t-interval, because per-record s is bounded in [−1, 1] and skewed, and the headline map bends at the hedger, so a normal approximation drawn on one arm's slope runs off the end of the scale; and the prior is resampled as pseudo-data rather than applied afterwards, because a bootstrap over the records alone collapses to zero width when they all agree, claiming certainty from four data points. Weighting each pseudo-prediction at the mean importance makes the pool's weighted mean identical to the published score, so the interval widens nothing but the honesty. Records are resampled independently; restatements of one thesis that survive dedupe are not independent, so every interval is a floor on the true uncertainty, never a ceiling.
  • Verbal bands are a reading aid over the same number, never a separate calculation, and they must clear two bars: a big enough effect (|A| ≥ 10 net, |A| ≥ 50 strong) and enough evidence (the 90% one-sided bound on the score's own side of zero). Either test alone mislabels. At N = 700 the standard error is about 3.6 points, so a +8 would be "statistically significant" while describing someone 4% better than a coin; at N = 37 a +19 sits inside its own noise. Failing the evidence bar is not a lesser band, it is a coin flip: we cannot tell this person from a hedger, and saying so is the honest answer. A provisional score additionally shows no verbal band at all.
  • Provisional means one of two things, and they are reported separately because they are different uncertainties and a reader needs to tell them apart. Small sample (thin_record) fires under 20 resolved predictions: sampling noise in the records we *have*, which the interval already measures. Partial record (incomplete_coverage) fires when the coverage ledger itself reports a gap: a medium we could not enumerate, a completeness loop that stopped on a documented ceiling rather than the threshold, or a recall audit estimating we captured less than half of the pundit's datable calls. The interval cannot see that second one at all, because it is a fact about what never entered the sample. A score can carry both, and the published score names which apply. This is deliberately keyed on the ledger's own evidence rather than on a single free-text field: a pundit whose corpus came from third-party promise trackers should be flagged whether or not a research agent remembered to write the reason down.
  • The shrinkage term (k = 10 pseudo-predictions at the prior) keeps a handful of lucky calls from outscoring a long solid record. Scores from fewer than 20 resolved predictions are labeled provisional / small sample (and are never shown with a definitive verbal band), and N is always displayed.

Worked example: a pundit says a recession is "almost certain" within a year (p = 0.90, importance 5) and none occurs (o = 0): s = 1 − 2(0.81) = −0.62 — a heavily weighted hit to the score. The same miss said as "I think" (p = 0.65) gives s = 0.155 — hedging is punished less when wrong, but also earns less when right.

Alongside the headline score we publish the components: a calibration table ("when they said 'definitely', they were right X% of the time"), the high-confidence hit rate, N, coverage statistics, and each record's individual contribution.

Per-tag sub-scores

Every prediction is tagged with 1–3 domains from a small controlled vocabulary (a small controlled vocabulary of domains). The scorer computes a separate Avow score for each tag, using the identical formula restricted to that tag's predictions, and surfaces any tag with at least a few resolved predictions (provisional under 20, like the overall). A prediction with multiple tags counts toward each of its tags, so the per-tag scores are not a partition of the overall — and the overall is not their average; it counts every prediction exactly once. This lets a multi-domain forecaster — sharp on geopolitics, weak on market timing — be read honestly per domain instead of blurred into a single number that describes neither.

6. Exclusions

Excluded from scoring (but listed with reasons in the published score record):

  • Pending predictions (deadline not passed) — shown on the timeline as pending.
  • Conditionals whose condition was not met — void.
  • Records awaiting human review, unverifiable records, implied-class statements.
  • Statements made before 2011-01-01.
  • Near-duplicate repetitions of one claim within 90 days are consolidated into a single record (annual re-assertions score separately — a broken clock is many separate wrong clocks).

7. Known limitations

  • Selection bias. Even with systematic sweeps, the gathered record is a sample. The coverage ledger is the honest disclosure, not a cure.
  • Verbal probability mapping is lossy. "Likely" is not exactly 0.80. The mapping is fixed, public, and applied uniformly, and the verbatim phrase is always shown.
  • LLM judgment noise. Research, verification, and quantification use AI agents. Mitigations: adversarial verification, three-judge medians with stored dissent, human review queues, and full audit trails. Errors will still occur; see corrections.
  • Resolution ambiguity. Some outcomes resist clean resolution. Such records carry partial credit with rationale or are flagged rather than force-scored.

8. Corrections and right of reply

Every prediction record has a stable ID and a public audit trail. Anyone — including the pundit — can dispute a quote, a resolution, or a mapping. Disputes are logged, re-verified, and corrections are published with the record's history preserved. Scores recompute automatically from corrected data.

9. Versioning

The methodology, rubrics, certainty mapping, and scoring formula carry a single current version. Once scores are published, any change to them bumps the version and triggers recomputation of affected pundits, and every published score states the version that produced it. Every published score records the methodology version that produced it.

Every public score traces back to sourced prediction records.

Avow is an independent project. The pundits scored here are not affiliated with Avow and do not endorse it. Names and likenesses are used for identification and commentary.

Evidence firstMethodologyPrivacy