Fugal has released version 0.2.0 of its Bittensor LLM-routing software after finding that a genuine Intel TDX quote was not enough to prove the result a validator thought it was checking.
The project reports closing five verifier exploit classes and three defects that could prevent the software from setting a single weight. Its local-chain dress rehearsal now passes 28 of 28 project-defined scenarios.
This is pre-Finney developer and security analysis. Tao Outsider found no matching Fugal identity in the TaoSwap subnet snapshot at block 9,009,564, and the repository’s provisioning examples still use placeholder netuids. We did not verify an assigned netuid, a Finney deployment or an independent security audit.
Fugal’s release is still worth examining because it exposes a recurring problem in decentralized AI. Authentic hardware can produce an authentic quote about the wrong claim.
What Fugal is trying to measure
Fugal describes a market for small model-routing policies.
Miners train a compact linear head on top of a frozen Qwen3-0.6B backbone. Given a question, the head selects a model from a defined pool. The target is not maximum accuracy at any cost. Fugal’s project-defined score combines quality relative to the best single model with thrift relative to a reference cost.
Validators build benchmark slices, while miners execute those slices inside an Intel TDX trusted execution environment. The miner returns a hardware-attested proof containing the selected head, answers, costs and execution evidence.
In theory, the validator can check a result without rerunning every model call. In practice, that only works when every relevant object is bound to the signed evidence.
Version 0.2.0 is mainly about those bindings.
A valid quote can still describe attacker-controlled fields
Fugal says it modeled five exploit classes against the previous verifier while assuming the attacker owned genuine Intel TDX hardware. The DCAP signature could therefore pass.
The first weakness was image identity. Approved-image matching read a source_hash field written by the workload instead of deriving identity from the quote’s own MRTD and RTMR measurement registers. A genuine enclave could state the hash a verifier expected without proving that the approved runtime image produced the work.
The second weakness concerned the assigned slice. Results were not checked tightly against the nonce-selected benchmark questions, while the proof could carry the larger answer pool. Without that binding, hardware attestation does not prove the miner answered the challenge the epoch assigned.
The third weakness separated the submitted router head from its on-chain commitment. A weights_hash field existed in the proof but was not compared with the committed artifact. That opens room to evaluate one head while claiming another.
The fourth weakness affected the downloaded evidence bundle. The verifier did not compare the bytes it received with the advertised proof hash.
The fifth turned cost inconsistency into a warning. In Fugal’s score, understating cost can improve the result. A field that changes rank cannot be advisory if miners control it.
Fugal reports that v0.2.0 now reads hardware measurement registers, checks answers against the assigned slice, binds the head to its on-chain commitment, verifies the bundle hash and rejects inconsistent costs.
These are repository and release claims. Tao Outsider did not recreate the attacks on TDX hardware.
Three ordinary bugs stopped the economic loop earlier
Not every consequential failure was cryptographic.
The release says miners and validators formatted epoch IDs differently. That caused benchmark slices to overlap incorrectly and proofs to fail nonce checks.
A miner also imported a function that did not exist. The epoch failed inside an exception path that logged the error without restoring progress.
Finally, the harness passed a raw loader dictionary into the grader rather than the object the scoring function expected. Fugal says every TEE answer received a score of zero without a loud failure.
Together, those defects meant the project could describe a complete incentive mechanism while the shipped path could not set weights correctly.
The release’s local-chain dress rehearsal therefore carries more weight than a large unit-test count. The project says scripts/dress_rehearsal.py starts the shipped components against a real local subtensor node and surfaced eight defects that in-process tests had missed.
The reported 28 of 28 result is encouraging within that test boundary. Its scope excludes Finney execution, security certification and agreement among production validators under hostile load.
The score is an explicit product choice
Fugal defines quality using a Wilson lower confidence bound on task accuracy relative to the best single model. It defines thrift from reference cost divided by miner cost. The combined score weights quality more heavily than thrift.
That design attempts to prevent a cheap but unusable router from winning solely on cost. It also avoids treating one lucky epoch as durable evidence by accumulating observations for the same committed head.
The benchmark mix includes MMLU, MATH, GSM8K, AIME, IFEval, GPQA-Diamond, HumanEval and optional LiveCodeBench data. Dataset revisions are pinned where the project provides them.
None of that makes the score objective truth. Benchmark choice, model pool, price table, confidence treatment and quality-versus-cost exponent all encode preferences. A router that wins Fugal’s frame may not minimize cost or maximize quality for a different workload.
The release also says the price table is hash-pinned rather than a stub. Reproducibility improves when validators agree on the exact table, but a pinned price can still diverge from what an external provider actually bills.
Why the pre-Finney label cannot be hidden
Fugal’s README includes scripts for local testnet, mock inference, test-network provisioning and a mainnet launch sequence. Those operational materials are not evidence that the subnet is live on Finney.
The missing TaoSwap identity match cannot rule out registration under another name. This article therefore makes no live-subnet claim.
No reader should interpret v0.2.0 as an invitation to deploy, fund a validator or connect credentials. The repository notes that paid operation requires explicit live and budget settings. Tao Outsider did not run those paths and will not reproduce setup instructions here.
Fugal should earn stronger language through public state. That evidence would include a declared netuid, a verifiable runtime, independent review of the quote bindings and reproducible epochs on the intended network.
Until then, its most credible contribution is the failure record itself. It documents how a TEE design can pass hardware verification while missing the software relationships that make the proof meaningful.
Bittensor builders can use the failure record before emissions depend on the result.
Sources
Fugal repository and mechanism documentation
Was this article useful?
One tap feedback helps us improve each post.