§19. Layer 4: Truth & Work
Copy/paste (plain text):
Jason St George. "§19. Layer 4: Truth & Work" in Next Generation Stores of Value: Privacy, Proofs, Compute. Version v3.1. /v/3.1/read/part-iv/19-layer-4/ Layer 4: Truth & Work
Layer 4 is the Prove/Verify spine of the stack.
Its job is to take arbitrary claims—“this media object came from camera C at time T,” “this inference was computed by model M on input X,” “this corridor swap executed atomically”—and turn them into portable proofs that:
-
anyone can check with cheap, public verification, and
-
can be standardized as commodities (canonical workloads, SLAs, prices).
What Layer 4 Is (and Isn’t)
Layer 4 is:
-
The layer of circuits, proofs, and verification economics.
-
The place where verification asymmetry is engineered and measured.
-
The home of Proofs-as-a-Library (PaL), proof factories, and canonical workload registries.
Layer 4 is not:
-
A specific proof system (SNARK vs. STARK vs. something else); it assumes multi-ZK.
-
A single PoUW design; it supports several patterns as long as VerifyPrice and anti-capture constraints are met.
-
A truth oracle. Proofs do not prove truth; they prove specific claims under stated assumptions—origin, custody, computation, policy compliance, and settlement finality. The thesis is not that cryptography tells us what is true, but that it reduces the surface area over which institutions must be trusted.
Verification Asymmetry Revisited
Recall the definition from Part I. For a workload :
-
= cost (time, energy, hardware) to produce a result + proof.
-
= cost to verify that result + proof.
-
= verification asymmetry.
Both and are aggregated into the same numeraire — USD via the per-run price model of §19: Layer 4: Truth & Work () — so the ratio is dimensionless and both halves carry the same measurement contract: under the adversarial harness, under the reference production model (Appendix A: Formal Model of Verification Asymmetry & VerifyPrice). Read the denominator before celebrating the ratio: improves in two ways, verification getting cheaper or production getting dearer, and a passing achieved by expensive production is a cost problem wearing an asymmetry credential. The ratio is therefore published with both components; a falling driven by a rising is flagged as such.
Layer 4’s goal is simple:
Make for the workloads that matter, and keep it that way in production.
This is the touchstone of the gold analogy, now carrying a price: expensive to smelt, cheap to assay. is the assay, priced.
Why this matters:
-
If is small and stable, anyone (including small nodes) can check proofs.
-
That makes proofs and verified FLOPs commodities: units of work that any counterparty can accept without trusting a platform.
Layer 4 introduces:
\text{VerifyPrice}(W) = \{\underbrace{\text{p50 time},\ \text{p95 time},\ \text{p50 cost},\ \text{p95 cost},\ \text{failure rate}}_{\text{canonical five-field tuple }([Appendix A: Formal Model of Verification Asymmetry & VerifyPrice](/v/3.1/read/appendix/a-verifyprice-model/))},\ \text{hardware profile mix}\}where the first five fields are the canonical tuple of Appendix A: Formal Model of Verification Asymmetry & VerifyPrice and the hardware profile mix is Layer 4’s reporting context: every published series names the verifier classes it aggregates over, and the mix travels with the tuple rather than replacing any field.
Verification Modalities, and Which One Carries the Monetary Claim
The ratio has been carried through this document as a single scalar against a single ceiling. That was adequate while cheap verification was an operational requirement. It is not adequate now that §10: Work Credits: Energy-Anchored Claims makes cheap verification the mechanism the monetary claim runs on, because the ceiling as written does not protect the property the monetary claim needs.
The ceiling is not binding where it matters.
A workload can satisfy and be nowhere near the regime the touchstone image describes. At checking a claim costs thirty per cent of producing it. Nobody checks everything at that price. Participants check samples, or delegate, or accept a counterparty’s assurance — which is precisely the symmetric ignorance described in §10: Work Credits: Energy-Anchored Claims, arrived at by an SLO the network is passing. An asset whose verification sits at the ceiling is accepted because checking is too expensive to bother with, which is the old mechanism with better tooling, and it fails the way the old mechanism fails.
The ceiling is therefore an engineering floor and should be labelled as one. What the monetary argument requires is a statement about modality, because the three ways of establishing that work was done correctly differ by orders of magnitude and differ in kind.
Verification Modality Bands
M1 — Succinct. A cryptographic proof checked by a verifier whose cost is sublinear in, and typically independent of, the work proved. Target , stretch . Everyone can afford to check everything.
M2 — Algebraic or probabilistic. A structured check cheaper than re-execution by a factor of the problem size, accepting a stated soundness error — Freivalds-style matrix verification [Freivalds 1977] at against production is the canonical case. Each Freivalds round costs three matrix–vector products, so rounds at soundness give in arithmetic alone (a single round over a large field, with soundness , gives ); the input/output and hashing cost of reading the committed operands adds a constant. Target with the soundness parameter published beside the ratio. Everyone can afford to check everything, in expectation.
M3 — Replicated or attested. Correctness established by independent re-execution, committee agreement, or hardware attestation. Verification cost does not fall asymptotically below production cost; what is bought is fault tolerance and accountability, not asymmetry. No target applies, because the ratio is not the operative quantity. Someone checked. You are trusting who.
The aggregate ceiling is retained as a portfolio-level engineering floor and as the fallback for workloads not yet assigned a band. It is not a monetary threshold and no monetary claim reads on it.
Only M1 and M2 carry the monetary mechanism, and this is a consequence rather than a preference.
§10: Work Credits: Energy-Anchored Claims locates moneyness in no-questions-asked acceptance reached by symmetric knowledge — checking is free enough, and public enough, that no counterparty can acquire an informational edge by performing it. M1 delivers that. M2 delivers it up to a published soundness error, which is a quantified and therefore acceptable dilution. M3 does not deliver it at all: under replication the participant is not checking, a committee is, and the participant is trusting the committee’s composition, independence, and continued availability.
That is not a criticism of M3. Replication is the right construction for a great many workloads and much of the useful capacity this thesis describes will be M3. It is a statement about what M3 capacity can be used for. M3 supports service delivery, Work Credits, and SLA tiers. It does not support the pledgeability claim, because a position whose correctness rests on a committee requires the lender to form a view about that committee — which is due diligence, which is the thing information-insensitivity is defined as the absence of.
Replication buys fault tolerance. Succinctness buys moneyness. The thesis needs both and has been pricing them as though they were the same good.
The modality is what the haircut was always pricing.
This resolves a loose end rather than creating one. The collateral haircut schedule of §14: Layer 0: Verifiable Machines & Energy already discounts positions by how legible their substrate is, and §10: Work Credits: Energy-Anchored Claims identifies those haircuts as prices on residual unit-quality uncertainty. Modality is the second axis of the same schedule and the more fundamental one: an M3-verified position carries irreducible delegated-trust exposure regardless of how good its hardware grade is, and an M1-verified position on modest hardware may carry less. The haircut table should therefore be indexed on modality grade, and the two should not be collapsed. A later pass owes the joint table; this document should not pretend it is there yet.
What this changes in the instruments.
Three reporting consequences, none of which relaxes anything:
-
VerifyPrice series are published banded by modality. An aggregate across a mixed M1/M3 portfolio is a number with no interpretation, and its improvement can be produced by changing the mix rather than by improving anything.
-
The M1+M2 share of verified capacity becomes a published series in its own right. It is the fraction of the network’s output that carries the monetary mechanism at all, and it belongs beside fee coverage on the Value Capture Board (§23: Extended Telemetry) rather than among the engineering metrics.
-
Red Line 3’s monoculture condition acquires a second reading. Concentration of verification in a single vendor is one failure; migration of the capacity mix from M1 toward M3 is a quieter version of the same failure, because both end at “trust the verifier” and only the first is currently instrumented.
And the rhetoric can now be correct.
Claims that checking costs a small fraction of one per cent of producing are true of M1 and of M2 at realistic problem sizes — one Freivalds round at costs roughly seven parts in ten thousand of production (), and twenty rounds at soundness cost about , which is outside the M2 band. The band therefore admits MATMUL_4096 under Freivalds only at rounds (, soundness ) or under the large-field single-round variant, and the registry entry must say which; larger relaxes the constraint linearly. An earlier draft quoted “two parts in ten thousand” here, which is and counts one matrix–vector product where the check needs three. The claims are false of M3 and of the aggregate ceiling. Any statement of the form “checking costs one ten-thousandth of producing” is a claim about succinct verification specifically and must name the modality wherever it appears, in this document and in every derivative artifact. The looser figure was never wrong as a floor. It was wrong as a description, and it was quoted as a description.
Canonical Workloads and Proof Types
To avoid an unbounded zoo of bespoke proofs, Layer 4 maintains canonical workloads:
-
MatMul() – matrix multiplication at specified dimensions and componentwise relative error bound.
-
Inference(; policy) – run model on input under constraints.
-
Provenance(, chain) – provenance chain for content type .
-
Settlement(, policy) – settlement and refund logic for a corridor.
For each canonical workload, the stack defines:
-
Circuits / arithmetization: how the workload is represented for proving systems.
-
Proof schemas: which proof systems are supported.
-
Verification modality: M1, M2, or M3 (§19: Layer 4: Truth & Work); a computation that can be checked under two modalities is registered as two SKUs, never as one SKU carrying both.
-
SLO tiers: latency and assurance classes (§19: Layer 4: Truth & Work).
Canonical workloads matter economically:
-
They become the SKUs for proof and compute markets.
-
Work Credits are minted against units of these workloads.
-
VerifyPrice and are tracked per workload and tier.
Every SKU this document names lives in one table, the workload registry of §19: Layer 4: Truth & Work; the definition template, the starter set (§19: Layer 4: Truth & Work), the constitutional SLO table (§19: Layer 4: Truth & Work), and the primitive catalog (§21: The Modular Stack) all point at it rather than carrying their own lists.
The precedent for standardized work.
The requirement that demand standardize into canonical units before it can become an economic object has a long pedigree, and it is worth naming because it is easy to mistake for an engineering preference. The standardization thesis in the print literature [Eisenstein 1979] documents the sequence: typographic uniformity—every letter identical to every other letter of its kind, modular, movable, recombinable—was daily practice at scale for roughly three centuries before the machine shop embodied the same template in screws, gauges, and interchangeable parts. The claim is causal and well-documented as mechanism: standardization of the cognitive module preceded and enabled standardization of the physical one. The stronger claim—that industrialization was possible only because of movable type—is not adopted here. The analogy to this layer is exact at the level of function: canonical workloads and PIDL receipts are what make two providers’ units of work identical, comparable, and combinable; diffuse demand for “compute” becomes an economic unit only when the units are interchangeable. The registry is the type case; the receipt is the printed page.
Canonical Workload Definition Template
For a workload to become a tradable SKU (and eligible for Work Credit issuance), it must be defined precisely enough that two providers’ units are interchangeable—and that is the whole reason this template exists. Free-form compute cannot be a unit of account. If one vendor’s “matmul” means fixed dimensions under a relative error bound and another’s means whatever their kernel happens to produce, no counterparty can compare the two, no market can price them, and no registry can mint against them. The template fixes the fields that turn work into a unit: what is claimed, what is public, what is private, what gets checked, what it costs to check, and where it may run.
One field deserves to be stated in the open rather than left inside a parenthesis, because two natural drafts of this template get it wrong in opposite directions. The error bound is relative, componentwise, and normalized to the operands that produced each entry — never absolute, and never a matrix-norm inequality. An absolute bound fails honest provers: standard BLAS kernels reorder floating-point arithmetic, and the discrepancy that reordering introduces grows with the magnitude of the operands, so a correct product computed by a different honest kernel would still fail the test. A matrix-norm bound fails in the other direction, and this document carried one through v3.0. Written as with the matrix -norm (maximum row sum), the right-hand side is a per-row budget of order . At on i.i.d. Gaussian FP32 inputs with that budget is about against an honest FP32 error of : a product computed from INT8-quantized inputs errs by about and passes, and a single output entry corrupted by about (typical ) also passes. A bound that admits the next-cheaper precision class and a two-hundred-fold corruption of one entry is not an acceptance test; it is a formality. The correct form is the componentwise backward-error bound of floating-point analysis [Higham 2002] (Higham, Accuracy and Stability of Numerical Algorithms):
with set from the pinned reference kernel’s precision — for strict FP32 accumulation — and calibrated so that the next-cheaper precision class the template forbids (TF32, FP16, INT8) fails. The relative componentwise bound forgives honest implementations and still fails results that are actually wrong, entry by entry. Strict with dishonesty, tolerant of implementation—that is the register the entire template is written in, and a bound has to be tested against dishonesty before it earns the register: every workload registration ships an adversarial acceptance suite (INT8 input substitution, single-entry corruption, row zeroing, precision downgrade of the accumulator) that the bound must reject, and a registration whose bound passes any of them is returned.
Canonical Workload Definition Template
WorkloadID: Unique identifier from the registry of §19: Layer 4: Truth & Work (e.g., MATMUL_4096_FP32_FREIVALDS, MATMUL_4096_FP32_SNARK, INFER_LM_70B_256TOK). One ID names one modality; the same computation under a second modality is a second ID.
Statement: What is being proved (natural language + formal predicate)
- Example: “The matrix product was computed correctly, where , and for every ” — the componentwise, scale-normalized bound of §19: Layer 4: Truth & Work. Neither an absolute bound nor a matrix-norm bound is the template’s form: the first fails honest BLAS kernels, the second admits INT8 execution.
Public Inputs:
-
Commitment to , (hash or Merkle root)
-
Claimed result commitment (hash of )
-
Error bound and the reference kernel ID that sets it
-
Timestamp range
Private Inputs (Witness):
-
Full matrices , ,
-
Intermediate computation trace (if required by proof system)
What Is Verified:
-
Correctness (computation matches claim)
-
Bounds (result within specified limits)
-
Freshness (timestamp within allowed window)
-
Liveness (proof generated within epoch, not precomputed)
Adversarial Acceptance Suite: the named set of dishonest transcripts the acceptance test must reject (for MatMul: INT8 input substitution, TF32/FP16 accumulation, single-entry corruption, row zeroing). Registration is refused while any suite member passes.
Verifier Complexity Class:
-
Time: two regimes, and the registry entry must say which it prices. Succinct-proof (M1) workloads verify in after the committed output is in hand; Freivalds-style (M2) audit workloads verify in — reading and hashing the committed operands, then matrix–vector rounds — but avoid the recompute. The template’s complexity field names the regime. A computation checkable in both regimes is two SKUs with two entries (
MATMUL_4096_FP32_FREIVALDSandMATMUL_4096_FP32_SNARKin §19: Layer 4: Truth & Work); an entry never publishes both regimes under one ID, because the two have different modalities, tiers, and prices. -
Memory: specified peak (e.g., 512MB for Laptop-Class)
-
Proof size: max allowed (e.g., 1MB)
Policy Hooks: (machine-verifiable predicates)
-
Hardware profile requirements: the minimum grade in the Allowed Hardware Profiles field below (L0-A by default), with a higher floor named per SLO tier where the entry requires one (e.g., L0-B or higher for the Gold latency tier)
-
Prover stake requirements
-
Membership/non-membership predicates (allowlist proofs, not graph inspection)
Allowed Hardware Profiles:
-
L0-A: Yes (with issuance weight 0.9)
-
L0-B: Yes (1.0)
-
L0-C/D: Yes (1.0; open-hardware premium via collateral grade and routing priority, never via issuance above delivered work)
SLA Tiers: Latency and Assurance
The v3.0 draft published one three-row table — Bronze/Silver/Gold at 60 s/10 s/2 s prove-side latency with // verification — and it priced the wrong thing. “ verification + audit” is replication: modality M3 in the taxonomy of §19: Layer 4: Truth & Work, the one modality that carries no monetary mechanism. A 2 s prove-side target is reachable today only under M2 or M3, because succinct (M1) proving of any workload in the registry runs minutes to hours on current provers. The highest-fee tier as written therefore rested on the least monetary modality and rewarded provers for choosing it. Latency and assurance are separate goods with separate prices, and the tiers are now two orthogonal axes: a latency tier that prices how fast the attested result arrives, and an assurance tier that prices how it was verified. A quote names one of each.
| Latency tier | Prove-side p95 to attested result | Fee multiplier |
|---|---|---|
| Bronze | 60 s | 1.0 |
| Silver | 10 s | 1.3 |
| Gold | 2 s | 1.6 |
Latency tiers for canonical workloads. The target is the prove-side time to produce the attested result (generation plus attestation or proof), not verify-side time: verification SLOs are constitutional and set separately in §19: Layer 4: Truth & Work, and no prove-side tier can relax them. Gold is attainable today only with M2 or M3 assurance; an M1 quote at Gold latency is refused by the registry until a prover benchmark on record supports it.
| Assurance tier | Modality | What the buyer receives | Fee multiplier |
|---|---|---|---|
| Replicated | M3 | One independent re-execution or attestation; committee is the trust root | 1.0 |
| Replicated + audit | M3 | Three independent re-executions plus randomized third-party audit; Tier C/B service claims | 1.4 |
| Probabilistic | M2 | Algebraic check with published soundness parameter (e.g., Freivalds , error ); Tier B Work Credits | 1.6 |
| Succinct | M1 | Cryptographic proof checked in milliseconds independent of workload size; Tier A Work Credits when the maturity gate lifts. Prove-side latency SLO set per workload from the current prover benchmark on record, not from the latency table | 2.5 |
Assurance tiers for canonical workloads, indexed by verification modality (§19: Layer 4: Truth & Work). The highest multiplier attaches to the modality that carries the monetary mechanism, and its latency promise is whatever current provers can honestly meet. A quote’s total multiplier is the product of its latency and assurance multipliers.
Example: MATMUL_4096_FP32, two SKUs
The same computation is registered twice, because it is checkable under two modalities and the two are different goods.
| Field | MATMUL_4096_FP32_FREIVALDS |
MATMUL_4096_FP32_SNARK |
|---|---|---|
| Statement | Matrix multiply C = A × B, dimensions 4096×4096, FP32; componentwise bound |Cij − (A ⋅ B)ij| ≤ εr∑k|Aik| |Bkj| for every (i, j), εr = 10−5 (strict FP32 accumulation, pinned kernel) | |
| Modality / tier ceiling | M2 / Tier B | M1 / Tier A (maturity-gated; [§19: Layer 4: Truth & Work](/v/3.1/read/part-iv/19-layer-4/)) |
| Public Inputs | Hash(A), Hash(B), Hash(C), εr, reference kernel ID, timestamp | |
| Private Inputs | A, B, C | A, B, C, arithmetization trace |
| Verified | Correctness (probabilistic, soundness 2−k published), Bounds, Freshness | Correctness, Bounds, Freshness |
| Verifier work | O(n2): stream and hash A, B, C against the public commitments (192 MB of FP32 operands), then k rounds of three matrix–vector products with the componentwise tolerance applied per row; k = 13 default (r ≈ 9.5 × 10−3) | O(log n) proof check after the committed output is in hand; the verifier never touches A or B. The buyer still receives C (∼64 MB) and hashes it once against the commitment |
| Memory | ≤256MB with streaming (operands read once per round, never held whole); a verifier that materializes all three matrices needs ≥192MB and sits at the ceiling | ≤256MB |
| Proof / receipt size | Receipt only (≤500KB); no proof artifact | ≤500KB proof plus receipt |
| Prove-side feasibility | Production: the prover computes C once with the pinned kernel | Research: no production prover exists for a 40963 FP32 product at a price a buyer would pay; registered so the target is public |
| L0 Requirement | L0-A minimum; L0-B+ for Gold latency | L0-A minimum |
| VerifyPrice Target | t95 ≤ 5s (Laptop-Class; I/O-bound on the 192 MB read) | t95 ≤ 2s (Laptop-Class) |
Why this matters: Without this template, “canonical workload” risks becoming “whatever the prover says it is.” With it, workloads are precisely specified (anyone can implement a conforming prover/verifier), auditable (claims can be checked against the template), and tradable (markets can price and exchange standardized units).
Workload Registry
Three lists in the v3.0 draft named workloads — the template’s examples, the starter set, and the primitive catalog — and no two agreed on the identifiers. A registry that is quoted three ways is not a registry. §19: Layer 4: Truth & Work is now the single table; every other mention of a SKU in this document is a pointer into it, and an identifier that does not appear here is not a canonical workload.
| WorkloadID | Class | Modality | Tier ceiling | Acceptance criterion | Reference verifier class |
|---|---|---|---|---|---|
PROOF_2^20 |
Proof | M1 | A | Succinct proof of a 220-constraint statement verifies against the published verifying key | Laptop |
|
MatMul | M2 | B | k = 13 Freivalds rounds pass the componentwise bound εr = 10−5 per row; operands hash to the public commitments; soundness 2−13 published | Laptop |
|
MatMul | M1 | A (gated) | Succinct proof that C satisfies the componentwise bound against committed A, B; buyer hashes delivered C | Laptop |
|
Inference | M3 (sampled re-execution) | B | Pinned-kernel transcript; ≥ 10% of inferences re-executed by random auditors; divergence beyond εr fails; stake slashed on fraud | Server (Laptop re-execution is minutes; published, not constitutional) |
|
Inference | M1 | A (gated) | Succinct proof of the quantized forward pass; quantization gap to the advertised model published separately | Laptop |
|
Inference | M3 | B | As INFER_LM_7B_512TOK; re-execution bound to Server-Class (model does not fit Laptop-Class memory) |
Server |
|
Provenance | M1 | A | Proof that the asset’s hash chain is rooted in a Layer-0 attested capture and each hop is signed by the stated party | Laptop, Mobile |
|
Settlement | M1 receipt over chain finality | A | Both legs finalized, or the no-loss exit of [§20: Layer 5: Value & Settlement](/v/3.1/read/part-iv/20-layer-5/) taken within the refund window; receipt carries both chain proofs | Laptop |
|
Settlement | M1 receipt over chain finality | A (gated) | As BTC↔︎XMR, over the shielded ZEC pool; registered, not admissible: no production shielded-ZEC adaptor-signature swap exists ([§20: Layer 5: Value & Settlement](/v/3.1/read/part-iv/20-layer-5/)) | Laptop |
PoUW Design Patterns
Layer 4 doesn’t pick a single consensus recipe. It supports several proof-of-useful-work patterns, as long as they satisfy:
-
Open admission: commodity participants can join.
-
Unpredictable leader election: no one can cheaply bias the lottery.
-
Useful work binding: you can’t precompute offline or reuse stale work.
-
Proof quality & anti-spam: junk proofs can’t flood the system without penalty.
Pattern 1: Hash-gated useful work
-
Miners perform a cheap hash race; crossing a threshold gives short-lived eligibility.
-
To produce a valid block, the miner attaches a PoUW artifact seeded from header randomness.
Pattern 2: Proof-first selection
-
Provers race to produce useful-work proofs and post them to a mempool.
-
A lightweight mechanism selects which proofs get included and rewarded.
Both patterns are compatible with MatMul-PoUW, Inference-PoUW, and hybrid schemes.
The Usefulness–Unpredictability Dilemma
Neither pattern resolves the problem that sits under every proof-of-useful-work design, and this document has until now stated the problem’s two halves in separate places as though each were solved. It is one problem, it is open, and it should be named as such.
A work function must be unpredictable to secure consensus: if the input is known in advance, a miner can precompute and the lottery is not a lottery (property P1 below). A work function must be useful to have a buyer: someone must want the result of this input, not a random one. The two requirements pull on the same field of the block. Header-seeded inputs — Pattern 1 as written — are unpredictable and useless: nobody buys the product of two matrices drawn from a block hash, so the “useful” work is heat with better branding, and the demand curve the thesis relies on (§19: Layer 4: Truth & Work) does not exist for it. Client-supplied inputs — the only kind a buyer pays for — are predictable to whoever supplied them: a client and a miner who are the same party, or who collude, can precompute the work, submit it as fresh, and bias the lottery by exactly the amount P1 forbids. The hybrid designs — random-seeded verification challenges layered on client work, in the manner of Ambient’s proof-of-logits [Ambient 2024] — move the randomness from the input to the audit: the client’s input is taken as given and the header seeds which slices of the work get checked. That is the right shape, and it is not a solution until it comes with a stated collusion bound: for a client–miner coalition controlling fraction of demand and fraction of hashrate, what reward advantage does it obtain, and under what audit rate does that advantage fall below the grinding tolerance of P2? No design this document cites publishes that bound.
This is the flagship research problem for a build lab, and it is stated here as one rather than folded into an engineering table. Verification asymmetry is a theorem [Freivalds 1977] and is not the hard part. The hard part is making the work simultaneously worth buying and worth mining, and the honest current status is: unsolved, with a promising shape and no proof. Until a hybrid design carries a collusion bound, MatMul-PoUW and Inference-PoUW secure the service market (Work Credits minted against delivered, verified work) and do not by themselves secure consensus; the consensus role remains a conjecture in §19: Layer 4: Truth & Work, and the starter-set rows inherit that label.
PoUW Security Properties (Must Hold)
P1. Precomputation Resistance: Proofs seeded by unpredictable randomness; valid only for blocks after seed.
P2. Grinding Resistance: Attacker with hashrate gains at most rewards, where .
P3. Network Advantage Bound: p95 propagation delay s for proof announcements.
P4. Spam Resistance: Deposit expected verification cost; 100% slash for invalid proofs.
P5. Cartel Detection: Top-1 share ; Top-5 share ; entry latency days.
P6. Useful Work Binding: of proofs independently re-verified by random auditors.
Research Maturity Labels:
| Component | Maturity | Status |
|---|---|---|
| MatMul verification asymmetry (Freivalds, check) | Production | Theorem [Freivalds 1977] (Freivalds 1977); the asymmetry itself is not in question |
| MatMul-PoUW consensus security (grinding / precomputation bounds under useful inputs) | Experimental | Conjecture; no published collusion bound (§19: Layer 4: Truth & Work); not yet deployed |
| Inference PoUW (Proof-of-Logits) | Pilot | Tested in limited networks |
Succinct MatMul proving at FP32 (_SNARK SKU) | Research | No production prover at buyer-acceptable cost |
| Full ZKML for large models | Research | Not yet practical at production scale |
| Hash-gated useful work (Pattern 1) | Experimental | Devnet demonstrations |
| Proof-first selection (Pattern 2) | Experimental | Early implementations |
PoUW component research maturity.
Implication for service claims: The “AI Money” label is an analytical lens on verified-compute demand, not an instrument classification. As components mature from Experimental to Production, inference Work Credits can receive stronger assurance grades and smaller service haircuts. Their issuance remains separate from the base-asset monetary constitution (see Tier taxonomy in §21: The Modular Stack).
| Verification Type | Service-Claim Eligibility |
|---|---|
| Full cryptographic proof (Tier A) | Full eligibility for high-assurance Work Credit issuance |
| Probabilistic audit with explicit error bounds (Tier B) | Discounted or limited eligibility |
| TEE/vendor attestation only (Tier C) | Service market only; not eligible for high-assurance capacity claims |
| Unverified output | No eligibility |
Verification type and typed service-claim eligibility.
Proof Factories and PaL
Most developers don’t want to think about circuits; they want to say:
“Prove that this computation happened, then get me a receipt and pay whoever did the proving.”
Layer 4 provides this via:
-
Proof factories: infrastructure clusters specialized in generating proofs.
-
Proofs-as-a-Library (PaL): an SDK that compiles high-level claims into proofs.
PaL exposes interfaces like:
prove_compute(f, inputs, policy)
prove_provenance(asset_id, lineage)
prove_settlement(tx, corridor_policy)
Under the hood, PaL:
-
Maps the request to a canonical workload .
-
Selects suitable proof systems and hardware profiles.
-
Submits the job via neutral routers to proof factories.
-
Returns a PIDL receipt plus a proof artifact.
VerifyPrice Observatory
Verification asymmetry is a design goal; VerifyPrice turns it into a dashboard.
The VerifyPrice observatory continuously measures:
-
p50/p95 verify times,
-
estimated energy per verification,
-
failure rates and mismatch rates,
-
diversity metrics (how many independent verifiers are actually checking).
Reference Verifier Classes:
| Class | Hardware Spec | Use Case |
|---|---|---|
| Laptop-Class | 4-core CPU, 16GB RAM, SSD, no GPU | Baseline for “anyone can verify” |
| Mobile-Class | ARM SoC, 8GB RAM, flash | Edge verification, IoT |
| Server-Class | 32-core CPU, 128GB RAM, optional GPU | High-throughput |
Reference verifier classes for VerifyPrice measurement.
VerifyPrice Measurement Specification
VerifyPrice is a service-assurance KPI and one necessary input to the base asset’s conditional monetary test. It determines whether Work Credits retain a verifiable service proposition; it does not make them monetary. That requires a rigorous measurement harness, not just a dashboard slogan.
The harness exists because the observatory is not a disinterested instrument. It is a publisher with interests, run by operators who can be pressured, and the numbers it publishes are exactly the kind a protocol under strain would prefer not to publish. Assume the observatory is fallible, and further assume it is adversarially fallible—inclined toward flattering measurements when the truth is inconvenient. The specification below is not written to make the observatory honest. It is written so that when the observatory lies, anyone can catch it: the samplers are named, the harness is open, the raw data is archived, and independent operators run the same tests on the same proofs. A lie has to survive all of that. The spec does not produce trust in the measurement; it produces detection of the absence of it.
Input View (What the Harness Records):
VerifyPrice has exactly one definition: the canonical five-field tuple of Appendix A: Formal Model of Verification Asymmetry & VerifyPrice, introduced at §19: Layer 4: Truth & Work,
The v3.0 draft restated it here as a different five-tuple and again in §23: Extended Telemetry as a six-tuple, each introduced as the definition; those were the harness’s inputs wearing the output’s name. What the harness records, per run, for workload on verifier class , is the input view:
Where:
-
: verification time (seconds); and are its percentiles over the window
-
: energy per verification (Joules, measured via hardware counters or watt-meter); enters below
-
: peak memory (MB) — published, not priced: a run exceeding the registry entry’s memory ceiling is recorded as a failure and lands in ; observed peak is published beside the tuple
-
: bandwidth consumed (bytes) — published, not priced, on the same rule as memory
-
: pass, or a failure code (timeout, invalid proof rejected, crash, resource ceiling exceeded); is the failure fraction
USD cost is derived per run, then percentiled: , where is the run’s measured energy in Joules, is the reference energy price (USD 0.10/kWh, so the first term divides by J/kWh), and is opportunity cost (USD 0.001/s), both published quarterly. The cost percentiles of the canonical tuple — and — are the percentiles of across the measurement window, never recomputed from percentile inputs; in particular, deriving a cost figure from alone is not the construction.
Verifier Implementations:
| Requirement | Rationale |
|---|---|
| 2 independent implementations per workload tier | Prevents single-implementation bugs from corrupting measurements |
| Open-source, reproducible builds | Anyone can audit and rebuild |
| Deterministic output | Same proof same result, always |
| Versioned and tagged | Measurements tied to specific verifier version |
Verifier implementation requirements.
Sampling Methodology:
-
Random selection: Proofs selected uniformly at random from recent submissions (not cherry-picked).
-
Stratified by region/profile: Measurements cover 10 regions and 3 hardware profiles per workload.
-
Adversarial corpus: 10% of test proofs are malformed or worst-case (max witness size, pathological inputs) to measure failure handling.
-
Continuous measurement: Not periodic snapshots; rolling 24-hour windows published hourly.
Adversarial Conditions:
| Condition | How Tested |
|---|---|
| Network latency | 200ms RTT injected for proof fetch |
| Packet loss | 5% random packet loss during verification |
| Worst-case proof size | Max allowed witness for workload class |
| Malformed proofs | 10% of test corpus intentionally invalid |
| Resource exhaustion | Verification under 80% memory pressure |
Adversarial conditions for VerifyPrice testing.
Reproducibility Standard:
-
Benchmark harness: Open-source, versioned, deterministic. Anyone can run the same tests.
-
Signed results: Each measurement batch is signed by 2 independent measurement operators.
-
Divergence alerts: If independent operators diverge by 5%, investigation triggered.
-
Archived raw data: All proof samples and timing logs archived for 1 year.
Why this matters: Without this spec, “VerifyPrice observatory” is an oracle claim. With it, VerifyPrice becomes reproducible consensus—any skeptic can run the harness, check the measurements, and falsify the dashboard if it’s wrong.
Physical VerifyPrice SLOs (Constitutional):
| Workload | Modality | Target (Laptop) | Sev-1 | Prove-side feasibility |
|---|---|---|---|---|
PROOF_2^20 | M1 | s | s for 7 days | Production; seconds to minutes on Server-Class |
MATMUL_4096_FP32_ FREIVALDS | M2 | s (I/O-bound: 192 MB read and hashed) | s for 7 days | Production; one pinned-kernel product |
MATMUL_4096_FP32_ SNARK | M1 | s | registry target; not constitutional until the maturity gate lifts | Research; no production prover at buyer-acceptable cost |
INFER_LM_7B_512TOK | M3 (sampled re-execution) | s on Server-Class (registry-bound); Laptop-Class re-execution of a 7B forward pass over 512 tokens is minutes and is published, not targeted | s Server-Class for 7 days | Production; the prover runs the model once |
INFER_LM_7B_512TOK_ ZKML | M1 | s (proof check is milliseconds, independent of model size) | registry target; not constitutional until the maturity gate lifts | Research; current ZKML proves quantized transformers at small scale with hours of overhead per inference |
INFER_LM_70B_256TOK | M3 | Server-Class only (model exceeds Laptop-Class memory); aspirational as an M1 SKU | n/a until ZKML matures | M3: production on Server-Class; M1: not currently achievable |
| All | — | failure rate | for 7 days | — |
Constitutional VerifyPrice targets, banded by modality. The band is not decoration: M1 verification is milliseconds independent of model size, and M3 verification of a language model is running the language model, so a single “INFER_LM_7B s” row was either an M1 promise no prover can yet keep or an M3 promise no laptop can keep. Rows marked “registry target” are published so the destination is public; placing them in the constitutional set as though they were deliverable would trip Red Line 1 on publication day, and they become constitutional only when the maturity gate of §19: Layer 4: Truth & Work lifts for their component. Where the reference class is Server-Class, the Laptop-Class figure is still published.
VerifyPrice in Practice
At the North-Star level, VerifyPrice answers a simple allocator question:
“If I hold this asset for a cycle, does one unit still buy at least as much verification as it used to?”
We care about VerifyPrice in three dimensions, designed to avoid circular reasoning:
The Constitutional SLO Band
These are non-negotiable targets, measured in real resources on reference hardware. They are exogenous to token price. Whether the token rallies or dumps, these targets must hold. If verification takes 30s on a laptop, the “anyone can verify” promise is broken. The table below is the one referenced as §19: Layer 4: Truth & Work’s SLO band; the measurement specification that produces its inputs is §19: Layer 4: Truth & Work.
What “Exogenous” Does and Does Not Cover
Exogeneity to token price is the property we need here, and it is the one being claimed: the target must not be redefined because the market moved.
It is not exogeneity to everything. Measuring in real resources on reference hardware makes these SLOs dependent on two conditions a state can adjust—whether unprivileged reference hardware remains obtainable, and what power costs at the edge and at the prover. Both are policy surfaces (§4: Threat Model).
The consequence is that Red Line 1, the thesis’s hinge, is externally triggerable. This is monitored upstream via sovereign optionality (§14: Layer 0: Verifiable Machines & Energy) and Red Line 13 rather than left to be discovered at the hinge. §14: Layer 0: Verifiable Machines & Energy separates the verification-side exposure (hardware obtainability) from the proving-side exposure (bulk energy), which behave differently and should not be conflated.
There is also a reason the exogeneity requirement runs deeper than engineering preference. If the framing of §1: The Failure of Soft Guarantees is right that a medium restructures only those who habituate to it, then verification must be cheap enough to become a habit before it can become anything else. If its cost tracks the token’s price, habit never forms, and neither does whatever the habit would have built. The SLO is not a performance target; it is the precondition of the medium having any effect at all.
Protocol Affordability SLOs (Operational)
These measure whether verification remains affordable as a fraction of typical transaction costs:
-
Target: AffordabilityRatio for all workload classes.
-
Settlement: Verification cost of median transaction value.
Read the denominator before celebrating the ratio.
The ratio improves in two ways: verification gets cheaper, or the fee gets bigger. A protocol that raises passes this SLO while becoming less affordable to use, so the ratio must never be published without its components, and a falling ratio driven by a rising denominator is a deterioration wearing an improvement’s clothes. The settlement variant has the mirrored exposure: if median transaction value is measured in token terms, appreciation mechanically lowers the ratio and re-imports token price through the back door — the exact circularity the physical SLOs exist to exclude. Three guards hold the family honest, each of a different kind. The first is component disclosure: the fee-based ratio is always published alongside the fee level itself, so a denominator-driven improvement is visible on inspection. The second is a unit convention: the settlement ratio is computed on USD-referenced transaction value, never token-referenced. The third is a monotonicity rule: the affordability SLO is credited only when the ratio falls while the fee level is flat or falling, and improvements achieved on a rising fee are reported as neutral, not as passes. The affordability family is operational, not constitutional, and these guards are why it stays that way honestly.
Token-Quoted VerifyPrice (Market Signal, Not Target)
Token-quoted VerifyPrice is a useful market signal, but not a constitutional target. For a scalar reading, fix a single reference workload and take the cost-percentile field only, with the fee schedule expressed as a fixed token quantity:
where is the fee schedule for denominated in units of the native asset , and is the asset’s USD reference price at the observation — so the signal carries the token price explicitly on the right-hand side, which is what makes it a market signal rather than a disguised operational one. A vector has no reciprocal; the scalar must be constructed by naming the workload and the field. Which workload is the reference is a publication choice, not a free parameter — it is declared, pinned, and changed only with notice, so the signal cannot be dressed by re-picking . Note the family relationship: this quantity is the token-side image of the AffordabilityRatio, and the two must be published together so neither is read alone; the affordability ratio is the operational SLO, this is its purchasing-power translation.
Token price is endogenous. Making this a target creates circular reasoning. The protocol commits to physical SLOs and affordability ratios. Token-quoted metrics are dashboards for market participants, not governance constraints.
SLO Hierarchy:
| Level | What It Measures | Who Enforces | Consequence of Breach |
|---|---|---|---|
| Physical (Constitutional) | Real-world verification cost | Protocol governance | Sev-1; remediation required |
| Affordability (Operational) | Verification as % of fees | Fee policy | Fee schedule review |
| Token-quoted (Market) | Purchasing power signal | Market participants | Informational only |
Physical VerifyPrice: Auditing the Infrastructure Behind a Workload
Note on terminology: §19: Layer 4: Truth & Work uses “Physical VerifyPrice SLOs” to mean the real-resource cost of verifying a proof or compute claim on reference hardware, as distinct from token-quoted metrics. This section introduces a related but distinct metric: the cost of verifying the physical infrastructure claims (energy, hardware, jurisdiction) behind that proof or compute claim, i.e., the Facility Capacity Receipts introduced in §14: Layer 0: Verifiable Machines & Energy.
VerifyPrice tells us whether a proof or compute claim is cheap to check. Physical VerifyPrice tells us whether the physical claims behind that proof are cheap to audit. A Work Credit cannot be high-assurance if verifying the proof takes seconds but verifying the facility, hardware profile, energy receipt, and jurisdictional risk takes months of trust-the-vendor diligence.
Concretely, Physical VerifyPrice asks:
-
Can an allocator verify the facility’s claimed energy profile?
-
Can an auditor verify outage and curtailment history?
-
Can the network verify that the hardware profile was not silently substituted?
-
Can a user verify that compute was not routed through a sanctioned or coerced jurisdiction?
Physical VerifyPrice
The time, cost, and confidence required to verify the physical infrastructure claims (energy, hardware, jurisdiction) behind a unit of verified work.
For each facility profile, the stack should publish p50/p95 time and cost to verify FCR claims, sampling frequency, third-party challenge success rates, and unresolved discrepancy counts. If Physical VerifyPrice rises above threshold, credits from that facility receive reduced issuance weights or issuance caps (§22: Layer 6: Governance & Telemetry). This metric is formalized alongside the VerifyPrice model in Appendix A: Formal Model of Verification Asymmetry & VerifyPrice.
Layer-4 Stress Tests
Layer 4 passes its SoV audition if it survives several adversarial scenarios:
Circuit bloat.
Can we detect when proofs become too expensive to verify? Is there a migration path to leaner circuits?
Prover cartel.
Can neutral routers and SLA/slashing mechanisms prevent a small set of proof factories from monopolizing high-value workloads?
Proof system break / new attack.
Can we deprecate a proof system, rotate to alternatives, and quarantine affected Work Credits?
If Layer 4 remains cheap to verify, open to participate, and instrumented enough to handle change, then Proofs and Compute deservedly move closer to “monetary primitive” rather than “platform feature.”
Minimum Viable Layer-4 Economy
Canonical Workload Starter Set:
The starter set is the subset of the registry (§19: Layer 4: Truth & Work) that a minimum viable Layer-4 economy launches with. Identifiers are the registry’s; the columns here are the ones a market maker reads.
| WorkloadID | Description | Modality | Commercial target | WC tier ceiling |
|---|---|---|---|---|
PROOF_2^20 | Generic ZK proof | M1 | s | A (full) |
MATMUL_4096_FP32_ FREIVALDS | Matrix multiply 40964096, Freivalds-audited | M2 | s | B (discounted) |
MATMUL_4096_FP32_ SNARK | Matrix multiply 40964096, succinctly proved | M1 | s | A (gated) |
INFER_LM_7B_512TOK | 7B-parameter inference, sampled re-execution | M3 | s (Server-Class) | B (discounted) |
PROVENANCE_MEDIA_V1 | Media provenance chain | M1 | s | A (full) |
SETTLE_ATOMIC_BTC_XMR_V1 | Atomic swap settlement, BTCXMR | M1 receipt over chain finality | s | A (full) |
Canonical workload starter set, drawn from §19: Layer 4: Truth & Work. Two target levels are distinct by design and must not be read as one: the constitutional floor (§19: Layer 4: Truth & Work SLO table) is the minimum a network must hold to remain constitutional, while this starter-set target is the commercial bar the registry holds itself to for the named SKU. Where the two coincide, as they now do for the inference row, it is because the constitutional table was rebanded by modality and the earlier “s” commercial figure for a 7B re-execution on Laptop-Class was not a target any verifier could meet. WC tiers shown are the ceiling each workload can reach once its verification component is Production-mature; current effective tier is the lower of the ceiling and the maturity-gated tier (§19: Layer 4: Truth & Work tier rules). The _SNARK SKU is listed so the M1 destination for matrix work is public; it issues nothing until a production prover exists.
Maturity gate.
No workload may issue Tier A (full-weight, pristine-collateral) Work Credits while the maturity label of its verification component (§19: Layer 4: Truth & Work) is Experimental or Pilot. Until a component reaches Production, its workloads issue at Tier B or below regardless of the mathematical tier their proof type would support. As of this writing that gate is binding on every PoUW-based row in the starter set: MatMul-PoUW consensus security is a conjecture (the verification asymmetry itself is Freivalds’ theorem [Freivalds 1977] and is not what the gate reads on; §19: Layer 4: Truth & Work), succinct MatMul proving is research, inference PoUW is pilot-stage, and pricing pristine collateral over any of them would define a AAA rating on an instrument whose proof system has not survived production. For starter-set rows whose verification component does not appear in §19: Layer 4: Truth & Work at all (PROOF_2^20, PROVENANCE_MEDIA_V1, SETTLE_ATOMIC_BTC_XMR_V1), the gate reads on the maturity of the underlying proof system family — generic succinct proofs, provenance attestation, atomic-swap verification — each of which the maturity table tracks at the family level; a family not yet tracked is treated as Pilot until it is. The gate lifts only when the maturity table lifts, and the maturity table is moved by deployment evidence, not by argument.
Tier Rules for Work Credit Issuance:
| Tier | Verification Type | WC Weight | Collateral Grade |
|---|---|---|---|
| Tier A | Full cryptographic proof (ZK-SNARK/STARK) and component at Production maturity | 1.0 | Pristine |
| Tier B | Probabilistic verification (audited transcripts) | 0.6 | Standard |
| Tier C | Attestation-backed (TEE + sampling) | 0.3 (service-grade claims only) | Ineligible |
Work Credit issuance tiers.
Fee + Burn + Slashing Logic:
-
70% of fee to prover (reward)
-
20% burned (supply reduction)
-
10% to protocol treasury (security budget)
The 20% burn share here is a deliberately conservative variant of the 30–50% reference band in §6: The Triad and the Monetary Candidate: this worked example also funds a protocol treasury, so the three shares are set to sum to 100%. The reference band remains the design target; a 20% burn is the harder case for the value-capture claim, not the easier one.
Fee formula (reference):
New here? Start with the one-minute version.
Tip: hover a heading to reveal its permalink symbol for copying.