ΦBitcoin Field Theory

Admission Records AR-001 / AR-002

The Cost-Basis Verdict

MVRV Z-Score · RejectedNUPL · Rejected

Abstract. We tested Bitcoin’s two most famous on-chain valuation indicators - MVRV Z-Score and NUPL - for admission into Bitcoin Field Theory as measurements of the cost-basis field. Formulas, parameters, and failure conditions were pre-registered before testing. Both instruments reproduce cleanly across independent data sources and both are directionally valid in isolation: historically, extreme cheap readings preceded strongly above-median returns. But neither adds measurable out-of-sample information beyond a properly built, lookahead-free Power Law trend model. Each won only 2 of 5 walk-forward test windows (NUPL: 1 of 5 at the two-year horizon) against a pre-registered bar of 75 percent. Verdict: both rejected. No composite valuation index ships until an instrument genuinely earns its place beside the time-trend field.

Why we ran this test

MVRV and NUPL appear on virtually every Bitcoin dashboard as independent gauges of whether Bitcoin is cheap or expensive. The theory behind them is genuinely attractive: every coin was last moved at some price, so the blockchain itself records what the market actually paid - the realized capitalization. When price floats far above the aggregate cost basis, holders sit on unrealized profit and the temptation to sell grows. When price sinks below it, sellers are exhausted. That is a real economic mechanism, observed from real on-chain data.

Our framework’s first admitted instrument is the Power Law - Bitcoin’s long-run growth trajectory as a function of time. The question this admission run asked is precise: does the cost-basis field carry information the time-trend field doesn’t already have? If yes, a composite valuation index becomes possible. If no, the famous indicators are echoes - the same field measured twice.

One identity most dashboards never mention

Before any data touched the tests, one mathematical fact was declared up front: NUPL = 1 - 1/MVRV. NUPL is a strictly monotonic transform of the MVRV ratio. Ranked against each other, they are the same number. Every dashboard displaying both as separate gauges is displaying one measurement twice. The only genuine difference between the candidates is normalization: MVRV Z standardizes against its own history; NUPL is the raw bounded ratio. So this was always a two-horse race with one horse.

How the test was designed

Everything below was fixed and dated before any result was computed - the formulas, the thresholds, and what failure would look like.

What we found

The candidates are real - in isolation

Using decile thresholds fixed on pre-2016 data only, extreme readings behaved exactly as the cost-basis mechanism predicts, entirely out of sample:

Behavioral test - forward log returns after extreme readings (out of sample, 2016-2026)
Signal state1y forward (mean)2y forward (mean)Test-period median
MVRV Z bottom decile (“cheap”)+0.84+2.02+0.53 / +1.09
MVRV Z top decile (“expensive”)-0.03+0.03+0.53 / +1.09
NUPL bottom decile+0.82+1.93+0.53 / +1.09
NUPL top decile-0.00-0.06+0.53 / +1.09

Cheap readings preceded returns far above the period median; expensive readings preceded roughly zero. The mechanism is real. If these instruments existed alone, they would be respectable.

But they don’t exist alone

The redundancy screen already hinted at the problem: within any single cycle, cost-basis signals and the Power Law deviation rank-correlate around 0.80. The full-sample correlation drops to 0.58 only because the relationship shifts between cycles - and that instability is precisely what kills the candidates in the decisive test:

Walk-forward test - does adding the candidate improve out-of-sample accuracy? (1y horizon)
Test window+ MVRV Z+ NUPL
2016-2017worse (-84%)worse (-45%)
2018-2019better (+69%)worse (-31%)
2020-2021worse (-18%)better (+2%)
2022-2023better (+30%)better (+16%)
2024-2026worse (-98%)worse (-262%)
Score2 / 5 (bar: 75%)2 / 5 (bar: 75%)

The two-year horizon was no kinder: 2 of 5 for MVRV Z, 1 of 5 for NUPL. The relationship each candidate learned in past cycles not only failed to help in the next one - it frequently made forecasts substantially worse. An instrument whose calibration must be relearned every cycle, with only four cycles of history to learn from, is not yet delivering independent information. It is delivering hindsight.

The verdict

On 14 years of Bitcoin history, cost-basis valuation indicators - MVRV Z-Score and NUPL - are directionally valid in isolation but add no measurable out-of-sample information beyond a point-in-time Power Law trend model at one-to-two-year horizons under a pre-registered walk-forward test.

Both candidates: REJECTED. And the consequence is bigger than the verdict: our composite Bitcoin Valuation Index - the product this test was meant to enable - does not ship. Not until an instrument genuinely earns a second seat. We wanted the composite. The data said no. Publishing the no is the point.

What this does and does not claim

Reproduce it

Data source: CoinMetrics Community API (metrics PriceUSD, CapMrktCurUSD, CapMVRVCur), snapshot frozen 2026-07-21, SHA-256 57A0D97C...803A1DE1. Cross-check source: bitcoin-data.com MVRV. The full admission records - pre-registrations, parameters, failure conditions, complete result tables, and analysis code - are the permanent record behind this article. Every number above can be rebuilt from the named sources; if you find an error, we want it public.

Get the next verdict first