Version 1.4 · Effective 2026-09-07
This document is published before it is used, and it is never applied retroactively. Every claim in the Ledger carries the version of this rulebook that was current when it was logged, and is scored by that version forever. When the rules change, the change is forward-only, dated, and listed in the changelog at the end.
That single property is what separates a scoreboard from an opinion with a number attached. A rulebook that can be adjusted after the outcomes are known can prove anything.
What this is
We log the checkable predictions made by a fixed roster of crypto creators, and we score them against what actually happened. The record is public, the source of every claim is linked to the exact moment in the exact video, and the arithmetic behind every verdict is shown.
What this is not: investment advice, a recommendation to follow or avoid any creator, or a prediction of our own. Lambda Chi Capital makes no market calls. The moment we did, we would forfeit the standing to score anyone else’s.
What gets logged, and what gets scored
Everything a tracked creator says about an asset can be logged. Only a narrow subset is scored.
A claim is scoreable only if it has all three of:
- A named asset — a specific coin or token, resolved to a canonical id.
- A direction — up or down, or a price target that implies one.
- A resolvable horizon — a stated timeframe, or the published default.
Everything else is logged as COMMENTARY and never enters any score.
When the classifier is uncertain, it defaults to COMMENTARY. We would
rather under-count a real prediction than manufacture one.
Claim types
| Type | Meaning |
|---|---|
DIRECTIONAL |
“X is going up / down” |
TARGET |
“X hits $N by date D” |
CONDITIONAL |
“If A happens, X goes up” — scored only if A actually happened |
COMMENTARY |
Everything else. Never scored. |
Rule 10 below extends this to nine types — LEVEL, EVENT, PROBABILITY, ACTION, VOLATILITY and REGIME — each with its own published resolution test.
The known weakness, stated plainly
Automated extraction misclassifies. The VideoConviction paper (KDD 2025) found that current models “frequently misclassify general commentary as definitive recommendations.” That is the central technical risk of this product.
Our mitigations: an explicit directional verb plus a named asset is required
before anything is scoreable; uncertainty defaults to COMMENTARY; a random
weekly sample is checked by a human; and the measured error rate is published
on this page. A scoreboard that hides its own error rate is asking for trust it
has not earned.
Measured classification error rate: not yet established. The first sample is taken after 100 claims are logged. This line will carry a number and a date.
The rules
1. Asset, direction, horizon — or it is commentary
Stated above. No exceptions, and uncertainty resolves downward.
2. The horizon default is 30 days
A creator who says “this is going up” without a timeframe is scored over 30 calendar days from the video’s publication. This number was fixed and published before the first claim was scored, and a horizon is never assigned or adjusted after an outcome is known.
Where a creator states a timeframe, we use theirs. The row records which of the two applied.
3. No grade below 25 scored claims
A creator’s hit rate is not published until they have 25 scored claims. Below that, the Ledger shows the raw claims and the count, and no rate. Small samples produce dramatic numbers and mean nothing.
4. Score against the benchmark, never against zero
“Solana is going up” during a market-wide rally is not a skilled call.
Every verdict compares the asset’s return to Bitcoin’s return over the same window. A call that went up 10% while Bitcoin went up 20% is a MISS — being long anything would have done better.
Bitcoin calls are the exception: BTC cannot be its own benchmark, so a BTC claim is scored on its absolute move.
5. Boldness adjustment
A correct prediction that everyone else also made is worth less than a correct prediction nobody else made.
Each claim is tagged with its stance against the roster’s consensus that day —
contrarian, neutral, or consensus — and weighted 1.5 / 1.0 / 0.75 in
the boldness score. The raw hit rate is published alongside it, unweighted, so
both numbers are visible.
6. The universe is snapshotted at prediction time
Assets that die between the claim and its resolution are the single biggest source of false flattery in this kind of scoring. If the coin list is built from today and read backwards, everything that went to zero and got delisted simply vanishes — and the bias runs one way: it makes predictors look better than they were.
A dead or delisted asset resolves as −100%, never as “no data.”
7. Liquidity floor — below it, we do not pretend to know
Below roughly the top 300 by market capitalisation, or below a minimum order-book depth, a verdict reads “unscoreable — insufficient liquidity” rather than producing a number.
This is not caution for its own sake. Research in Management Science (Cong, Li, Tang & Yang, 2023) found wash trading averaged over 70% of reported volume on unregulated exchanges. A scoreboard is only as honest as its price data, and on the long tail that data is frequently fiction.
The practical effect is that some of the boldest microcap calls are never scored at all — in either direction. We think that is more honest than scoring them on numbers we do not believe.
8. The methodology is versioned and never retroactive
Changes are forward-only, dated, and logged below. A claim scored under v1.0 is scored under v1.0 forever.
9. A channel is scored on what it BROADCASTS, not only on what its host says
(new in v1.1)
If a channel airs a checkable forecast, that forecast enters the channel’s record — whoever originally made it. Nobody broadcasts a call they believe is worthless, so the decision to put a forecast in front of an audience is itself a forecast. This is the standard already applied to weather: a station owns the forecast it airs, whoever’s model produced it.
Every scored row therefore carries two additional fields:
source— who made the call, as named in the video (Cathie Wood,Grayscale, or the channel itself for a first-person call). A call with no stated attribution is recorded as the channel’s own. We never invent a name.host_stance— endorsed, neutral, or disputed: what the host did with the call on air.
A disputed call still counts as broadcast, and is always displayed as
disputed. A host who airs a forecast while rejecting it has still put it in
front of the audience, but the record must show that he pushed back, and any
published figure that includes disputed calls must say so.
Two views are published from the same rows, and they are never summed:
- the channel’s own calls (
source= the channel), and - the calls it amplified (
source= anyone else).
The record of the sources — the people being quoted — is maintained on the same basis and under the same rules.
Thread selection is mechanical. Where a summary covers only the week’s leading topics, those topics are the ones covered by the largest number of tracked channels, counted from the published day plans. They are never chosen editorially. A scoreboard that picks which predictions to grade is not a scoreboard.
Rule 9 widens whose calls are scored. Rule 10 widens what counts.
10. A claim is anything checkable, not only a price call
(new in v1.2)
A claim is any statement about the future that the speaker could later be shown to have got right or wrong. It does not have to name a price, and it does not have to name an asset.
This rule exists because of a measurement. One nine-minute episode was read by hand and by machine on the same day: the reader found roughly eighteen checkable statements, the extractor found two. The gap was not tuning. Every missed statement was a shape the rulebook had no name for.
Nine types are now recognised. Each carries its own resolution test, and each test is published here before it is used on anybody.
| Type | Example | How it resolves |
|---|---|---|
| TARGET | “Bitcoin hits 150k by Q4” | Window extreme against the target; ≥50% of the distance is a PARTIAL |
| DIRECTIONAL | “Bitcoin grinds higher” | Excess return over Bitcoin, per rule 4 |
| LEVEL | “The 58k bottom is in” | The floor holds or it breaks. No direction, no benchmark |
| EVENT | “The Clarity Act passes” | Public record, recorded by a human. No asset needed |
| PROBABILITY | “Sixty percent chance of a hike” | Brier score, plus HIT/MISS on which side of 50% |
| ACTION | “Strategy will not sell any Bitcoin” | Filings and public record, recorded by a human |
| VOLATILITY | “Prepare for volatility” | Realised 30-day volatility against its trailing 90-day median |
| REGIME | “Crypto winter is over” | Only against a definition published below |
| CONDITIONAL | “If the Act passes, a rally” | Split into two claims and scored separately |
The numbers, fixed in advance. A LEVEL claim with no stated timeframe runs 90 days. A VOLATILITY claim resolves HIT when realised volatility reaches 1.25× its trailing 90-day median. Both were published before the first claim of either type was logged.
Brier scoring for stated probabilities. A forecaster who says “sixty percent” and is wrong scores 0.36; one who says “ninety-five percent” and is wrong scores 0.90. Lower is better, zero is perfect. This rewards honest uncertainty and punishes confident wrongness, and it is the reason a stated probability is worth more to this Ledger than a bare call.
Conditionals are split, always. “If the Clarity Act passes, we get a huge rally” becomes an EVENT claim and a DIRECTIONAL claim, scored independently. A forecaster can read Washington well and read the market badly, and a single verdict would hide that.
Two things this rule does NOT do.
It does not let us score causation. “Bitcoin moved because of Iran” is not a claim we grade, in either direction. We score “expect volatility” against realised volatility. Who caused what stays an argument, and the disagreement is the story.
It does not let us invent a definition after the fact. A REGIME claim is scoreable only against a test written down in advance. The published tests are:
- “bull market” / “crypto winter is over” / “the bottom is in” (as a regime claim rather than a price level) — Bitcoin closes at least 20% above its trailing 12-month low and above its 200-week moving average, measured on the resolution date.
- Any other regime claim, including “the four-year cycle has ended”, has no published test yet and resolves UNRESOLVABLE. We log it, we show it, and we say plainly that we cannot grade it. Inventing a test once we know the answer is the exact failure rule 8 exists to prevent.
11. Extreme-magnitude claims may be scored below the liquidity floor
Rule 7 keeps us off the long tail because wash trading makes the price untrustworthy. That reasoning is about precision. It does not apply equally to every question we might ask of a price.
Deciding whether an asset beat a benchmark by four points needs a trustworthy price. Deciding whether an asset multiplied by ten does not. Fake volume can move a printed price by tens of percent. It cannot manufacture a 900% move that did not happen, and it cannot conceal one that did.
So: a claim whose stated magnitude is at least a 5× gain, or a fall of 80% or more, may be scored even if the asset sits below the liquidity floor — subject to all four of the following, which are the whole point of the rule.
- The robustness test. Compute the verdict three times: at the observed resolution price, at half that price, and at one and a half times it. If all three give the same verdict, it stands. If any disagree, the claim resolves UNRESOLVABLE. A call that survives being wrong about the price by fifty percent in either direction was never really a question about the price.
- A price must exist at both ends, from our published source. No price is still no verdict — rule 7’s other half is untouched.
- The exception is stated on the row. Every claim scored this way is
flagged
scored below the liquidity floorin the Ledger and on the site, with the asset’s rank at claim time shown next to it. A reader can find every use of this rule and judge it. - The threshold does not move. 5× is fixed here, in advance. A 3× call below the floor is unscoreable, and will stay unscoreable, however much we might later wish otherwise.
Why this is being added now, and what it is worth. It was written on 7 September 2026 in response to a video we had not yet processed — a creator naming several coins he said would 10× within the month. We knew the shape of the claim and not which assets it named or where they ranked. Publishing the rule first was the only way to add it honestly; the commit history of this document shows the order, and anyone can check it.
Had we run the extraction first, seen five microcaps, and then written this rule, it would be indistinguishable from moving the goalposts to keep a story we liked. That is the precise failure this publication exists to document in other people. The sequence is the safeguard, not our good intentions.
What it does not do. It does not lower the bar for ordinary calls, it does not make thinly traded assets suitable for anything, and it does not change a single claim logged before today. Rule 8 still holds: v1.2 claims are scored under v1.2 forever.
Our voice
Lambda Chi Capital does not say “we think”.
Not as a matter of caution, and not because anyone made us. Reacting to a public video is ordinary commentary, and we could offer an opinion on whether a call will land without any difficulty at all. We don’t, because of what it would cost.
The moment this publication says “that one won’t happen”, it holds a position. Every verdict it later issues on that claim is then open to a fair question: was it graded on the evidence, or on the view we had already taken? A scoreboard answers that question by never having a view.
So the rule is simple and it is absolute:
- We report what was said, who said it, the reason they gave, and when it comes due.
- We never say whether we agree.
- We never tell anyone to buy or sell anything.
- Where a prediction-shaped statement is wanted, we use the record — “calls of this shape have landed four times in nineteen” — which is made entirely of history and requires no opinion from us.
This is about market calls, not about our own reasoning. Rule 7 says we think a liquidity floor is more honest than scoring prices we do not believe — that is us explaining a choice we made, and we will keep doing it. What we never do is tell you where a price is going, or whether someone else is right about it.
Reporting is not endorsing. Quoting a creator saying “we will not sell any Bitcoin” is not a recommendation, and describing the reasoning someone offered is not adopting it. The line is between what they said and what we think of it, and we only ever publish the first.
This is enforced in the software that writes our scripts, not left to discipline: a script containing a first-person opinion, or an instruction to a viewer, is rejected before it can be recorded.
Prices
| Role | Source |
|---|---|
| Scoring price (canonical) | CoinGecko — volume-weighted across venues, which is what a viewer would have seen on a public price site |
| Rank / liquidity | CoinGecko market-cap rank at claim time |
Every verdict on the site shows its price source, both timestamps, and the arithmetic. A verdict whose provenance is not auditable is just an opinion with a number attached.
We store our own daily snapshots from day one. Free tiers give roughly a year of history and can vanish — two major free crypto data sources closed within the last fourteen months. Our own append-only archive is the only history we control.
Verdicts
| Verdict | Meaning |
|---|---|
HIT |
The claim’s direction beat the benchmark |
MISS |
It did not |
PARTIAL |
A price target that travelled ≥50% of the way. Counts as half a hit. |
UNRESOLVABLE |
Below the liquidity floor, no price data, or a conditional whose trigger never fired |
VOID |
Malformed — a scoreable type with no asset or no direction |
COMMENTARY |
Not a prediction. Never scored. |
OPEN |
The horizon has not closed yet |
A price target counts as hit if the asset reached the target inside the window, even if it closed lower. “Bitcoin hits 120k this month” is true if it touched 120k.
What we do not claim
- That this predicts anything. The Ledger is a record of what was said and what happened. It is not a model, and past accuracy does not forecast future accuracy.
- That a low score means a creator is dishonest. Being wrong about markets is the normal condition. The Ledger measures accuracy, not integrity.
- That our extraction is perfect. See the error rate above.
- That the roster is representative. It is twelve channels chosen for prediction density and audience size. It is not a survey of the space.
Corrections
Errors get fixed visibly. A corrected row keeps its claim id, records what changed and when, and the correction is listed here. We do not quietly edit history — a project about other people’s receipts has to keep its own.
2026-09-04 — eleven claims were stalled by a software fault
A logged claim carries two fields: its type and its resolution status.
When a borderline row is promoted out of commentary by human review, the code
changed the type and left the status alone. The resolver only looks at rows
marked OPEN, so every hand-promoted claim sat permanently at COMMENTARY —
visible in the record, counted as a live claim, and unreachable by scoring.
Eleven rows were affected. Ten of them were the oldest claims in the Ledger, and therefore the first ones due to be scored.
What was changed. Ten complete claims were returned to OPEN with their
original horizons intact. One — LCC-00043 — was found to have been promoted
without a direction, which made it unscoreable in principle; it was returned to
commentary rather than left as a claim we could only ever void. Every affected
row records the change in its notes field.
What was not changed. No horizon was moved, no outcome was assigned, and no resolved claim was touched. The fault was caught 25 days before the first horizon closes, so nothing had yet been scored either way.
The fix. Promotion now re-opens the row it promotes, and refuses outright to promote a claim that lacks the fields its type requires. Both behaviours are covered by tests that fail if the old behaviour returns.
We are recording this in as much detail as an error that had reached a published verdict, because the standard has to be set while the stakes are low.
Changelog
v1.4 — 2026-09-07. Adds no scoring rule. Publishes the editorial commitment that governs everything this brand puts out: Lambda Chi Capital never states a view on whether a call will land, and never instructs anyone to buy or sell. It is recorded here rather than kept as a house style because a promise made in public and dated is worth more than one that is not, and because the value of the Ledger depends on it. Enforced in code; see “Our voice” above.
v1.3 — 2026-09-07. Added rule 11: a claim of 5× or more may be scored below the liquidity floor, gated by a robustness test — the verdict must survive the resolution price being halved or raised by half — and flagged on every row that uses it. Written and published before the video that prompted it had been processed, so the assets and their ranks were unknown at the time. Forward-only; claims logged under v1.0 to v1.2 are untouched.
v1.2 — 2026-09-03. Added rule 10: a claim is anything checkable, not only a price call. Nine claim types with published resolution tests, Brier scoring for stated probabilities, mandatory splitting of conditionals, and published definitions for the one regime claim we are willing to grade. Prompted by a hand-versus-machine read of a single episode: 18 checkable statements found by a person, 2 by the extractor. Forward-only; v1.0 and v1.1 claims are untouched.
v1.1 — 2026-08-31. Added rule 9, the broadcast rule: a channel is scored on
what it airs, not only on what its host asserts in the first person, with
source and host_stance recorded on every row. Forward-only — claims logged
under v1.0 keep a blank source and are treated as the channel’s own calls.
Rules 1–8 are unchanged.
v1.0 — 2026-08-27. First published version. Rules 1–8 as above; horizon default 30 days; grading floor 25 claims; liquidity floor top-300; boldness weights 1.5 / 1.0 / 0.75; PARTIAL threshold 50%.