Subnet Deep Dive

Reliquary SN81 hardens proof throughput and restart recovery

Seven merged Reliquary changes target proof batching, checkpoint races, restart-safe weights and stalled duplicate prompts without proving a public V1 launch.

Written by Nora Blake Platforms and products correspondent
Format
News report
Read time
7 min
Source trail
10 links
Review
Tao Outsider Engine
A Reliquary training proof pipeline moving one full batch while restart checkpoints remain ordered and recoverable.
Tao Outsider original editorial composition based on an authentic Reliquary product visual and pull requests 242 through 248. Source image edited with Imagine Bridge.

Reliquary merged seven changes in less than nine hours across the most fragile parts of a decentralized training system. The work spans proof transport, checkpoint installation, weight submission and restart recovery.

The sequence began with bounded proof and code-execution capacity. It then measured a throughput bottleneck, changed one training group from four proof round trips to one, repaired a checkpoint race, stopped duplicate weight attempts after restarts and fixed a streaming queue that could stall behind a duplicate prompt.

No public V1 launch is established by these merges. The engineering targets failure modes that appear when a validator, proof worker, signer, trainer and chain clock have to keep agreeing under load.

Reliquary’s repository supplies detailed measurements and tests. It also records where the evidence stops. The projected throughput improvement was not validated by a clean local end-to-end remote-proof environment.

The first problem was transport capacity

Pull request 242 bounded the shared code-execution pool to service capacity and introduced authenticated proof batching. It also added an optional pick target independent of the immutable journal stride, allowing a window to train fewer selected groups while preserving the existing miner submission contract.

The next pull request added phase timing around expensive proof verification. Reliquary’s initial observation was counterintuitive. The worker reported roughly 96 to 119 milliseconds for a proof job while the scheduler saw around five seconds.

The instrumentation separated preparation, remote proof, degeneracy checks, sketch commitments, termination checks and log-probability validation. According to the project, deployed measurements then placed 92% to 97% of total wall time inside the proof RPC path.

The forward computation was not consuming those seconds. Round trips were.

A training group contains 16 rollouts while the maximum proof batch was four. The controller therefore split one group across four network calls. Reliquary measured roughly 1.2 seconds of wall time for about 0.44 seconds of forward work per trip and attributed around 60% of proof wall time to round-trip overhead.

One group now travels in one request

Pull request 244 raises the request item count to 16 so one training group can cross the proof boundary in a single call. It keeps byte ceilings fixed rather than scaling memory allocation with the larger item count.

The project says production batches measured 29 to 124 KB against a 64 MB ceiling. It estimates that removing three round trips can move the pipeline from roughly 12.5 to roughly 24 groups per minute, above the rate needed by the configured window.

Those numbers need three labels.

The 92% to 97% share is a project measurement. The 12.5-group baseline is a project measurement of the observed pipeline. The 24-group result is a projection from the changed transport shape, not an independently reproduced production benchmark.

The pull request is unusually explicit about its test limitation. Its local remote-proof integration suite failed on both the branch and unmodified main because the required environment was absent. CI, paired controller/worker deployment and observed production behavior remain the meaningful gates.

The transport hash also changes, which means the controller and proof worker need compatible images. A mixed deployment is expected to refuse the connection rather than silently combine protocols.

Checkpoints had a race after successful installation

The throughput work was followed by a smaller but dangerous checkpoint bug.

When checkpoint N was being installed, the intake path could already be staging N+1. Completion compared the newly installed N against the newer pending candidate and raised a rollback error after the model and checkpoint store had already advanced.

That ordering skipped intake confirmation and accumulator reset. Pull request 245 changes the final check to compare against the installed checkpoint, while preserving monotonic checks when a new candidate first enters the queue.

Accepting N while N+1 downloads does not constitute a rollback. Installing an older candidate over a newer installed checkpoint still does.

Reliquary reports 21 focused tests for the checkpoint and detached-flow cases. No training-quality or model-improvement claim follows from those tests. They cover state transition, not model capability.

Restart recovery touched both weights and miner records

Controller restarts exposed another class of ordering problem.

The validator could immediately retry an epoch already reserved by its signer. If Bittensor’s strict weight-rate limit refused the attempt, the software could return an empty reason. Pull request 246 makes the controller read the durable signer epoch and check chain eligibility before preparing another submission.

It also aligns fill-closed miner verdicts with the groups committed by the final batch assembler. Previously, telemetry could read an inactive auction selection and describe paid submissions as unselected even though the archived payment set said otherwise.

Pull request 247 adds one catch-up submission at the first unattempted epoch allowed by the chain, then returns to the normal lead-window cadence. A consumed signer epoch is skipped rather than delaying the first valid refresh by another full epoch.

These changes preserve monotonic reservations and uncertain-outcome protections. They do not prove that weights were accepted on Finney after every restart or that the resulting allocation improved rewards.

One late duplicate could stop later proofs

The final merge in the sequence addresses a streaming scheduler edge case.

If a duplicate of an already proven prompt arrived after the original verdict, it could remain pending even though it was no longer dispatchable. Because proof results are applied in rank order, the impossible item prevented later valid results from advancing.

Pull request 248 marks those appended duplicates as skipped using the existing prompt-claim decision. Reliquary says all 38 scheduler and streaming tests pass with the regression fixed.

Prompt deduplication policy and miner contracts remain unchanged. The fix lets the queue skip an item that cannot legitimately run instead of leaving every later proof blocked behind it.

What the package says about Reliquary’s product stage

The seven merges do not introduce one headline feature. They reduce the number of ways a live training market can disagree with itself.

Proof capacity has to match window capacity. Controller and worker protocols have to match. A checkpoint must remain ordered while the next one downloads. A restarted controller must respect the signer’s durable epoch. Paid miner groups must match archived verdicts. Duplicate prompts must not halt unrelated proofs.

That is a strong engineering signal. It is not evidence that a public V1 product is released, that projected throughput is sustained or that the training process produces better models.

At TaoSwap block 9,044,495, Reliquary SN81 showed 80 active miners, 0.1007278% of subnet emission and 24.71% miner burn. This confirms an active subnet context. It does not link miners to the merged revision, reveal how many proofs completed or establish revenue and adoption.

The next proof should be operational. It would show compatible images across the full proof plane, windows filling inside their declared budget, restart drills preserving checkpoint and weight state, and public product access tied to the same revision.

Until that appears, the accurate headline is about hardening. Reliquary has made its training pipeline more explicit about capacity, ordering and recovery. The launch claim remains unproven.

September 16 update: checkpoint files and submission feedback

Update by Iris Vale. Original reporting by Nora Blake.

Reliquary merged a change that replaces matching model weight files in a single Hugging Face commit, then compares the published set of weight filenames with the expected set. Its publication-recovery path applies the same check.

This addresses a specific problem for evolving checkpoints: old shard files can remain beside a new weight layout. The added check raises a publication conflict when the sets differ. It verifies which matching files are present, with a scope limited to the patterns defined in the code.

The same change makes miner admission responses more explicit. A busy precommit signature path supplies a one-second retry hint. A terminal queue-full response identifies the admission queue and omits that hint, because repeating the same reveal returns its cached verdict.

The distinction gives callers better information about which stage rejected their submission. It also keeps temporary capacity pressure separate from a completed decision on that request.

Limits and verification

Tao Outsider read the implementation and added tests without executing them. A filename-set check alone cannot certify the numerical integrity of model weights. The patch establishes neither prior production corruption nor a verified production rollout. Live training performance and miner earnings remain outside this evidence.

Sources

Reliquary checkpoint publication and admission changes

Commit-pinned publication implementation

Bounded V1 proof transport pull request

Proof timing instrumentation pull request

One-batch-per-training-group pull request

Checkpoint installation race fix

Restart-safe weight cadence pull request

Archive replay and catch-up weight pull request

Streaming duplicate prompt fix

TaoSwap subnet status API

Follow the Bittensor desk

Read the latest Bittensor stories with the same source discipline.