Subnet Deep Dive

Harnyx SN67 changes how research agents reach a novel score

Harnyx SN67 code adds fast-mode task routing, a second novelty challenge and clearer failure attribution to its Bittensor scoring path.

Written by Nora Blake Platforms and products correspondent
Format
News report
Read time
6 min
Source trail
7 links
Review
Tao Outsider Engine
Harnyx SN67 scoring path with a primary novelty result passing through an independent challenge before the lower classification becomes final.
Tao Outsider original editorial diagram based on Harnyx validator and miner-task commits. AI-assisted base image produced with Imagine Bridge.

Harnyx SN67 has changed several parts of the pipeline that turns a research-agent submission into a miner score.

The most important change concerns the route to the word novel, rather than a new model or leaderboard result.

When a submission receives a primary novel classification, Harnyx can now send it through a separate challenge against a different structural reference. The lower of the two classifications becomes final. A miner therefore cannot secure the strongest novelty label only by looking different from the first comparison selected by the system.

That rule sits inside a wider four-day sequence of repository changes. Harnyx also enabled fast-mode generation in repository code for some source tasks, routed canonical-answer benchmarks through the same fast path, added content-free progress signals to similarity judging and separated platform-owned tool failures from miner-owned failures.

These are intended code behaviors. Tao Outsider did not reproduce the Harnyx validator pipeline end to end, and the commits are not proof that its research agents outperform other systems.

The second comparison changes the burden of novelty

Harnyx miners submit research-agent code. The validator needs to score output quality, but it also needs to decide whether a new artifact is meaningfully different from previous work.

The repository describes four similarity outcomes. They are duplicate, near duplicate, notable change and novel. A cosmetic change should not receive the same treatment as a different controller, evidence flow and answer-production path.

Before the August 29 change, a pairwise comparison could still miss a shared branch. Imagine an agent with a new primary controller whose minority route preserves an older answer pipeline. If the selected reference did not cover that route, the first comparison might be too generous.

The new challenge is designed for that edge case. The validator first checks the closest symmetric structural reference. Only a primary novel result can open the second round. That round uses a different reference selected to cover more of the candidate’s own structure. The more conservative result wins.

Harnyx also expanded its benchmark cases to include hash-routed and minority branches. The repository says every reachable successful branch counts, even when a hash or another routing decision selects it only some of the time.

This does not establish global uniqueness. Harnyx says the label remains pairwise. It means the candidate survived the comparisons the validator selected, not that no similar agent exists anywhere.

Fast mode now reaches generated tasks and canonical answers

Harnyx has two broad scoring needs.

Some tasks have a canonical answer. They can be scored by checking expected and excessive answer components, then computing a deterministic precision and recall result. Other tasks depend on weighted rubrics where evidence and broader response quality remain part of the evaluation.

The August 31 commit sends canonical-answer benchmarks through fast mode while keeping weighted-rubric benchmarks in ordinary mode. That split follows the scoring semantics instead of applying one query mode to every benchmark.

Three days earlier, the repository code enabled fast mode in source-task generation. Each output slot independently selects fast mode with a probability of 0.5.

Independent selection does not guarantee a 50/50 batch. Ten slots can produce five fast tasks, three or eight. The code deliberately avoids a fixed per-batch quota.

For miners, the distinction matters because fast scoring drops citation benefit and focuses on answer components. Ordinary mode remains appropriate where the rubric needs evidence-sensitive judgment. A system tuned only for one path may not behave the same way on the other.

The commits show how Harnyx intends to diversify its task surface. They do not independently establish that fast mode is cheaper, more accurate or better for users in production.

A slow judge should look alive before it times out

Similarity judging can take time. A streamed model response may be making progress even while the outer request has not completed.

Harnyx added an observability layer that records when response headers arrive, when a stream produces output and how an attempt moves across candidate judge models. The new events are content-free. The logging identifies an invocation and its operational state without inserting the candidate or reference payload into the progress record.

The scoring formula stays intact here. The diagnostic change helps an operator distinguish a slow active judge from a dead request.

That distinction is practical. A validator that mistakes progress for a hang may retry work unnecessarily or terminate a comparison before the judge finishes. Better progress evidence can make failure handling more precise.

The repository tests the event path across Chutes and OpenAI-compatible providers. Those tests demonstrate the expected software contract. They do not show production uptime or prove that every provider-side failure is now observable.

The validator should own its own failures

Another August 28 commit addresses failure attribution.

Harnyx uses platform-controlled tooling around miner tasks. If the platform fails to grant tool-proxy access or denies a request at that control layer, the miner did not cause the delivery failure.

The project added explicit error codes for those cases and classifies them as validator-owned. The goal is to prevent an infrastructure failure from becoming a negative event attached to the miner.

This sounds administrative, but attribution is part of incentive integrity. A miner score is only meaningful if the validator separates output failure from its own transport, sandbox and permission failures.

The commit also preserves the opposite boundary. A genuinely invalid miner response remains a miner-owned error. Clearer labels should reduce ambiguous terminal states without turning every failure into platform fault.

Again, this is repository logic. Tao Outsider has not observed these codes moving live weights on Finney.

What the current subnet snapshot tells us

A TaoSwap snapshot captured at 10:04 BRT on August 31 showed 127 active miners for Harnyx SN67 and an emission_value of 0.000727912.

That confirms an active subnet context. It does not tell us whether the new scoring sequence is fully deployed, how often a novelty challenge changes a result, whether fast tasks improve research quality or whether users prefer the output.

The useful evidence today is narrower. Harnyx is making its evaluation path harder to satisfy by accident and easier to diagnose when infrastructure stalls.

The next useful evidence would be operational. Challenge frequency, classification reversals, failure ownership rates and benchmark results under both query modes would show what the changes do in practice. Without that data, the code is a serious mechanism update rather than a performance victory.

Sources

Harnyx commit: Route canonical-answer benchmarks through fast mode

Harnyx commit: Instrument similarity-judge stream progress

Harnyx commit: Require an independent challenge for miner-task novelty

Harnyx commit: Activate independent fast task generation

Harnyx commit: Fix miner-task terminal failure attribution

Harnyx repository: Harnyx SN67

TaoSwap subnet API: Current subnet status

Follow the Bittensor desk

Read the latest Bittensor stories with the same source discipline.