Proof of Intelligence
Evidence preview, not an economic proof system. Production core can issue assignment-bound text probes and aggregate scorecards/quorum state, but the public validator distribution, richer capability batteries, media lanes, routing effects, rewards, stake, and slashing are not live. Current evidence has no economic authority.
Centralized AI APIs ask you to trust the model behind the endpoint. The AI Power Grid is designed to measure it. As validator stages roll out, models on the network will be continuously benchmarked for real intelligence — the goal is measured capability, not a precision label.
Why this exists
On a decentralized network, you can’t take a worker’s word for what it’s running. A node could claim it serves a full-precision model at long context while actually running a badly quantized version at a fraction of the context. And quantization is a spectrum: a clean quant can be indistinguishable from the original, while a broken one produces fluent-sounding nonsense that fails the moment real reasoning is required.
So the design has the grid stop trusting the label and measure the model itself. A great quantized model that passes is welcome; a “full-precision” model that fails is not. The grid will rank by delivered intelligence.
How it is being built
The current text lane can target a Grid-issued assignment and record evidence. The broader design expands that into unpredictable probe batteries graded automatically:
- Structured generation — render a chess position to SVG, emit JSON to a schema, write code that has to compile. Quantization damage shows up first in structured and spatial tasks, and a parser catches it instantly.
- Reasoning — math and logic with verifiable answers.
- Long-context recall — hide a fact deep in a long prompt and ask for it. This also verifies the model’s real context length, not just the claimed one.
- Instruction following & perplexity — format adherence and a low-level signal of degradation.
Production-grade probes should be procedurally generated, randomly timed, and indistinguishable from ordinary work where practical. The current preview does not yet prove that every challenge is unfingerprintable or that every model is continuously covered.
What it should produce
- A live quality score and capability tier for every model on the grid — effective context length, structured-output reliability, reasoning accuracy.
- Smart routing: requests will go to nodes that measurably perform; persistent under-performers will be downranked and removed.
- Collateralized quality: operators will eventually stake AIPG to serve. Repeated objective failures should lower routing trust first; slashing should be reserved for objective fraud after assignment-bound evidence, quorum, and dispute tooling exist.
What it will mean for you
- Pick capability, not quants. Choose a measured tier; the network is designed to ensure the model behind it actually performs.
- Trust that’s verifiable. Few centralized providers publish continuous, independent measurements of the model you’re hitting. The grid is designed to.
- Honest hardware. Declared GPUs will be cross-checked against measured throughput, so capacity is real, not advertised.
The goal is measurable capability with explicit confidence and evidence, not a claim that every token or model weight can be cryptographically proven.
See also: Validator Node · Architecture Overview