New ORO SN15 repository code makes an evaluation fail when even one expected reasoning judgment is missing. A separate code change gives weight setting a controlled public fallback after private standings have been unavailable for a full epoch.
The commits address two different failure surfaces. One stops an incomplete reasoning sample from becoming a score. The other keeps validators from moving straight to a burn-only vector when the preferred standings feed remains unavailable. Neither commit describes an exploit. Both are mechanism maintenance around evidence, timing and fallback behavior.
That distinction matters for ORO because its product claim depends on evaluation. The subnet says miners submit Python agents that search products, compare options and make purchase decisions. Validators run those agents in isolated Docker environments against ShoppingBench, then score the results. If the path from an agent run to an on-chain weight can accept missing judgments or handle missing standings unpredictably, the shopping task is only part of the system under test.
What changed in reasoning scoring
The August 28 commit marks each scorable trajectory as expecting a reasoning judgment. The validator then compares the number of completed judgments with the number expected. Full coverage allows the reasoning average to be calculated. Any gap produces a failed evaluation run.
That is stricter than the previous threshold. The earlier logic failed a run when there were at least three judge calls and more than half had failed. Under that rule, a run could retain a reasoning score from the judgments that did arrive, even when some expected calls were absent. The new logic does not average the surviving subset.
The implementation also separates a real zero from a judge call that never produced a usable result. When every model attempt fails, the reasoning score is recorded as missing rather than as zero. The run still fails because coverage is incomplete, but the reason remains visible. A confirmed payment-required response is labeled as insufficient miner credits. Other missing judgments are labeled as infrastructure failures.
ORO also narrowed how HTTP 402 responses are interpreted. A response carrying structured payment_required metadata stops immediately and is attributed to credits. An in-flight reservation response is retried with model rotation and backoff. An unclassified 402 is also retried instead of automatically blaming the miner.
This is a useful accounting change. A zero should mean the judge evaluated the trajectory and assigned zero. A missing value should mean the judgment did not complete. Collapsing those states would make later diagnosis harder, even if both paths prevent the run from scoring.
Why complete coverage is the right boundary
A reasoning average built from partial coverage can be misleading. The missing calls may be random, or they may cluster around longer, harder or more expensive trajectories. The validator cannot know which case it has from the remaining scores alone.
Failing closed avoids that ambiguity. It does not establish that the judge is correct, unbiased or resistant to every strategy a miner might use. It establishes a narrower rule. All trajectories designated for reasoning review need a judgment before the evaluation can produce a valid aggregate.
The tests added with the commit cover complete judgments, a confirmed credit failure and other missing judgments. They also verify that a failed judge call opens the credit circuit without turning the failure into a scored zero. That is evidence for the intended software behavior. It is not evidence that every SN15 validator has installed or is running the change.
The weight fallback uses public information after a delay
The second commit concerns what happens when a validator cannot obtain the authenticated, epoch-pinned standings used to build weights.
During a transient miss, the validator keeps its last accepted on-chain vector. It does not submit a replacement immediately. Once a full epoch has elapsed without a successful private standings update, the validator can request ORO’s post-embargo public top miner and build a fallback vector from that result.
The fallback still has guardrails. The public top must be revealed, its hotkey must remain registered in the metagraph and the emission policy must be usable. If the public endpoint is unavailable, the top remains embargoed, the hotkey is missing or the policy is malformed, the setter preserves the last good weights. When private standings recover, they take priority again.
This replaces the prior full burn response after a sustained miss. The public vector assigns weight to the revealed top miner alongside the configured burn share, rather than sending everything to the burn UID. The commit tests the grace period, an ineligible validator, unusable public data and recovery of private standings.
The design is conservative in a different way from the reasoning change. Evaluation integrity fails closed because an unknown judgment cannot support a score. Weight continuity degrades to a public, delayed input because validators still need a deterministic basis for setting weights. The fallback is bounded by time, registration and data checks.
What the commits do not prove
Repository code and unit tests do not prove validator-wide deployment. They also do not show that ORO now produces fairer outcomes, higher benchmark scores or better shopping recommendations. Measuring those claims would require version adoption data, comparable evaluation runs and observed validator behavior.
The code says nothing by itself about product adoption, revenue or demand. At 10:45 BRT on August 29, the supplied TaoSwap snapshot showed 69 active miners and an emission_value of 0.023974307 for SN15. Those numbers provide status context for the subnet at one point in time. They do not identify users, purchases, commercial revenue or sustained demand.
There is also no exploit claim in the commits. The useful story is that ORO found two places where validator behavior needed a sharper rule and added tests around those rules. Treating routine hardening as an attack narrative would obscure the actual mechanism work.
Tao Outsider read
ORO is trying to reward agents that can operate inside a messy commercial environment. That makes scoring discipline part of the product. Search quality, product matching and recommendation accuracy matter, but validators also need a defensible answer when a judge disappears or a private standings feed goes dark.
The latest code gives clearer answers. Missing reasoning coverage invalidates the run. Missing private standings trigger patience first, then a limited public fallback. Neither change proves better outcomes today, but both make the conditions for accepting a score or submitting a weight vector easier to audit.
The next evidence should come from operations. Validator version uptake, the frequency and causes of incomplete judge coverage, how often the public fallback activates and whether independent runs preserve the intended behavior are still unknown. Until then, this is a credible mechanism update with an open deployment question.
Sources
ORO commit: Fail closed on incomplete reasoning judge coverage
ORO commit: Use revealed top as validator weight fallback
ORO repository: ORO AI shopping agents on Bittensor
Was this article useful?
One tap feedback helps us improve each post.