Analysis

Numinous documents a two-week test of forecasting reasoning

Bittensor's Numinous describes how a hindsight ledger could grade miner reasoning. Its scoring documentation lists the Reasoning pool as inactive.

Written by Nora Blake Platforms and products correspondent
Format
News report
Read time
3 min
Source trail
3 links
Review
Tao Outsider Engine
Conceptual notebooks representing reasoning on day zero and review on day fourteen, with inactive pool status.
Tao Outsider conceptual AI illustration, generated with Imagine Bridge and edited with OpenAI image generation.

Numinous has documented a way to revisit a forecasting agent’s explanation two weeks after it was written. The Bittensor project compares the original reasoning with a shared record of subsequent events, looking for developments the agent anticipated.

The September 18 documentation describes a scoring design whose economic role remains on hold. In the same repository revision, Numinous lists its Reasoning pool as inactive and receiving no emissions, with reintroduction expected.

An explanation with a deadline

The design starts with the miner’s reasoning on day zero. Fourteen days later, a ledger is assembled for a market that was forecast, moved in price and left enough evidence to examine.

A model first turns the market’s title into search queries. Retrieval combines its associated news stream with a date-bounded corpus search. A second model reads those sources alongside the price path and produces the ledger. Every miner addressing that market and window is assessed against the same record.

At least three developments beyond the price move itself must support the ledger. A thinner record is discarded, leaving that market without a reasoning score for the window.

That threshold gives evidence collection a consequential role. Markets with sparse reporting can disappear from this part of the evaluation, while well-documented events offer more material for both the ledger and its grader.

How the grade works

The grader assigns an argument score of zero, one or two. It assesses whether the text connects evidence to the forecast probability; a list of headlines receives zero. It also identifies later events whose substance the original explanation anticipated.

The documented formula adds up to two points for anticipating selected key events, producing a score between zero and four. An argument grade of zero makes the whole score zero.

Key events are developments after day zero that fall within a day of the two largest daily price moves. If none meets that timing rule, the fallback uses all later ledger events. This makes the market’s path part of the evaluation criteria. Temporal proximity alone cannot establish what caused a price move.

The design also treats missing and very short explanations differently. Text under 50 characters scores zero. An agent returning no reasoning receives the 25th-percentile score for that ledger. That distinction deserves scrutiny before any return to paid scoring, because the treatment of omission can shape what agents submit.

The question for Bittensor

A shared ledger gives validators a common reference for evaluating explanations. Its quality depends on which sources retrieval finds, which events a model includes, and how another model decides that an earlier explanation anticipated them.

Those choices can reward concrete, testable reasoning. They can also favor widely reported narratives or penalize a sound argument about a driver the ledger missed. Public examples of original forecasts, evidence records and resulting grades would make that tradeoff easier to evaluate.

For subnet builders, the useful contribution is a documented evaluation design with explicit inputs and exclusions. Its eventual incentive effect will depend on the rules used when, and if, it enters a rewarded pool.

Scope and verification

This analysis covers documentation merged September 18 at commit 4f6e55b. The scoring document explicitly separates the inactive Reasoning pool from its active reforecasting pool. Current production settings, validator agreement and forecasting performance were not independently tested. The documentation supports a description of the mechanism, with activation and reward outcomes still requiring separate evidence.

Sources

Follow the Bittensor desk

Read the latest Bittensor stories with the same source discipline.