0%
RHAURUM://FORGE
[ INIT ] CALIBER_SCORING_ENGINE ........ OK
[ LOAD ] PRISM_RESERVOIR ................ 2,418,002 ITEMS
[ SYNC ] PILOTS 214 ▸ AUDITORS 37 ....... OK
[ SEAL ] NEXT_SORTIE #-- .............. LOCKED
[ DONE ] ALL_SYSTEMS_NOMINAL
0 %TAP_TO_SKIP
ENTER_MISSION SYSTEMS_READY ▸ AWAITING_PILOT_INPUT
PRISM://EVAL_NETWORK ▸ LIVE ▸ SORTIE #--

Where AI quality is proven, not claimed

RHAURUM is a decentralized evaluation network for open-weight AI. Every 60 minutes the Forge ships a fresh, sealed test bundle that no one can study for. Pilots fly their models against it. Auditors replay every run. The network pays for measurable quality — not marketing.

NET: ROBINHOOD_CHAIN 4663 ▸ ONLINE ▸ PROTOCOL: ◈ SIM PRE-LAUNCH ▸ CADENCE: 60 MIN
SORTIE // -- ◈ SIM IN_FLIGHT
NEXT_SORTIE--:--:--
ACTIVE_PILOTS214/312
BUNDLE_SIZE160 ITEMS
SEALED_ANCHORS40
LAST_WINNER5FgH…xQ2k
TOP_SCORE0.9372
PROTOCOL_TELEMETRY ▸ ◈ SIM SIMULATED — PRE-LAUNCH · REAL CHAIN DATA BELOW
0
SORTIES / DAY
0
SEALED BUNDLES / YR
0
PILOTS LIVE
0
TOP SCORE
CHAIN_TELEMETRY ● LIVE
--BLOCK --BLOCK TIME --GAS AVG --ADDRESSES
CHAIN_FEED_OFFLINE — RETRYING…
MISSION 01 // THE PROBLEM

The evaluation crisis

Benchmarks decide which models get bought, deployed and trusted. Today, almost all of them are broken in the same three ways.

[ 01 ] CONTAMINATION

Public tests leak

Test sets end up inside training corpora. Models score 90% on a test they have already memorized — and nobody can tell the difference between ability and recall.

[ 02 ] SELF-GRADING

Every lab grades its own homework

Model creators publish their own numbers. There is no independent replay, no third-party audit, no penalty when the number fails to hold up in the field.

[ 03 ] FROZEN RANKINGS

Leaderboards reward one lucky run

A static ranking measures one good afternoon, not sustained quality. Nothing measures how a model performs tomorrow, on questions it has never seen.

MISSION 02 // THE MECHANISM

Four steps. Every hour. Forever.

A sortie is one complete evaluation cycle. The loop repeats 24 times a day, and no item ever flies twice.

01

FORGE

PRISM draws a fresh bundle of 160 items from the reservoir — 40 of them sealed anchors whose answers are committed before the bundle ships.

02

FLY

Pilots run their open-weight models against the bundle and submit outputs with per-item confidence. Scored and unscored items look identical.

03

AUDIT

Auditors replay every pilot run in isolated containers and score only the hidden anchors with the CALIBER engine. Verdicts are reproducible, not reported.

04

REWARD

Emissions flow to the top five qualifying pilots, weighted by score squared. The network pays for skill that survives contact with the unknown.

MISSION 03 // THE SHIFT

What changes with RHAURUM

BEFORE — CLAIMED QUALITY
AFTER — PROVEN QUALITY
Benchmarks the lab wrote and ran itself
Blind sorties, audited by independent replays
Public test sets, saturated by training data
Sealed, burn-once bundles no model has ever seen
One-shot leaderboard runs, gamed in an afternoon
24 fresh sorties a day, quality tracked continuously
A slide in a deck that says "state of the art"
An on-chain score nobody can edit after the fact
LIVE_SORTIE // #-- ◈ SIM STREAMING
WINNING_MODELQwen-3-32B / VLLM-0.8
ANCHORS_RESOLVED38 / 40
REPLAY_HASH9f3a…c21d ✓
DESIGN_PROPERTIES ENFORCED
BLIND_ANCHORS40/160 scored, hidden
REPLAYABLE_AUDITevery run re-executed
BURN-ONCE_ITEMSno repeat, ever
STAKE-GATEDforgery forfeits stake
SEALED_COMMITSanswers hashed pre-sortie
SCORE_SMOOTHING12-sortie window
FLIGHT_CLEARANCE_OPEN

Run your model where nobody can game the test

Stake, fly, and let the network pay for what your model can actually do.

ENTER MISSION CONTROL →