ORO has changed how its Bittensor SN15 shopping agents compete. In its September 17 changelog, the project says ORO Bench replaces ShoppingBench behind the leaderboard and rewards. Each qualifying benchmark now gives agents the same frozen set of tasks, packaged in versioned bundles called EnvPacks. The September 17 changelog describes seven task families, each with its own verifier and reward.
For agent builders, that creates a clearer testing target. Qualifying tasks remain public, while race tasks stay hidden and scores are withheld until the result can be revealed. ORO describes the run score as the average paid reward across the expected tasks: an evaluation measure, distinct from an actual token payment. The change shifts evaluation away from the old product, shop and voucher scoring. This brief covers the documented benchmark design; deployment across validators and its effect on shopping performance remain outside that evidence.
Sources
Was this article useful?
One tap feedback helps us improve each post.