Arthur Samuel gave machine learning a name in 1959 while working on a checkers program at IBM. The program recorded positions, learned from previous games and improved its decisions through experience.
The board was small. The idea was enormous.
In Bittensor today, the word training can point to very different work. One subnet runs contests between open training repositories. Another asks miners to fine tune robotics models and submit reproducible checkpoints. A third splits large language model pretraining across unreliable machines.
All three can be described as machine learning. They do not sell the same product, measure the same output or expose the same proof.
That distinction matters because Bittensor does not maintain one global model. The chain coordinates incentives. Subnet owners define the task, miner interface, validator process and scoring logic around a particular market. Yuma Consensus then rewards miner utility supported by validator evaluations, weighted by stake and clipped around consensus.
The chain can settle weights. It cannot independently tell a reader that a trained model is useful, that a benchmark was fair or that customers will pay for the result.
To understand machine learning on Bittensor, start with four questions.
- What object changes through training?
- Who supplies the compute and code?
- How is improvement tested?
- What evidence survives outside the subnet’s own scoring loop?
Three current repositories show why those questions produce different answers.
Gradients SN56 turns training code into a competition
Gradients on Demand, or G.O.D, describes itself as the subnet runtime behind Gradients.io training jobs and tournaments.
Its public repository shows a contest structure. Validators create tasks, coordinate training infrastructure and evaluate results. Miners submit an open repository and an exact commit. Validator controlled trainers execute that code, produce models and move successful entries through tournament rounds.
The competing object is the training method.
That can include code for text, image or reinforcement learning tasks. The winning repository is valuable because another party can inspect and rerun the approach, not because the subnet merely claims that a model improved.
This is a meaningful design choice. Buying GPU hours is simple to describe. Creating an open competition over how those hours are used is harder. A miner can win by finding a better recipe, data treatment, optimization path or evaluation aware implementation.
The repository also exposes a weakness that applies across machine learning markets. The task designer still controls the dataset, evaluation process and advancement rules. A model can win the contest it was given and fail to generalize beyond it.
Gradients has been actively changing that machinery. An August 25 commit added tests around contestant selection and dataset propagation. Other recent changes addressed model evaluation memory pressure and tournament reranking. Those are signs of ongoing engineering. They are not independent proof that every winning model is better in production.
At the August 28 UTC TaoSwap snapshot used for this article, Gradients SN56 showed six active miners. That number is a point in time, not a verdict on the size or quality of the training market.
OpenRoboto SN80 puts robotics checkpoints on the line
OpenRoboto narrows the target.
Its public protocol asks miners to improve vision language action models, a class of model that connects visual input and language instructions to physical actions. The miner downloads a base checkpoint and public training resources, trains through the project’s runner, uploads a model checkpoint and announces the exact commit on-chain.
The submitted object is a model checkpoint.
OpenRoboto documents public seed derivation and a reproducible evaluation toolkit. Its protocol uses future block data and an external randomness source so the final evaluation seed does not exist when the miner submits. The goal is to make advance overfitting harder while allowing the miner to reproduce the eventual test.
The repository is also candid about the boundary. Held out task data, the scoring service deployment and owner operations are not fully public. A public validator fetches weights from a read-only endpoint and writes them on-chain. It does not contain the scoring service itself.
That makes OpenRoboto a better case study than a slogan. The project exposes training code, submission requirements, seed logic and a local evaluator. It withholds parts of the production evaluation surface to protect the test.
The tradeoff is familiar to machine learning. A completely public test can be gamed. A hidden test asks users to trust more of the evaluator.
An August 8 protocol update tightened the champion margin, unified submission states and documented burn verification rules. The TaoSwap snapshot showed three active miners on SN80. Neither detail proves robotics performance. Together they show a live protocol still working on how models enter, qualify and compete.
IOTA SN9 distributes the training run itself
Macrocosmos IOTA attacks a different constraint.
Instead of asking miners to submit independent fine tunes, IOTA describes a framework for pretraining a large language model across heterogeneous and unreliable machines. An orchestrator distributes model layers. Miners process activations. They periodically upload local weights and merge work through a collective update process. Validators spot check the computation.
The changing object is a shared model being trained across a distributed system.
This is closer to the popular image of decentralized AI compute. It is also a much less forgiving engineering problem. Training must survive machines with different hardware, unstable connections and changing availability. Communication between layers can erase the economic advantage of using distributed GPUs. Verification must catch a miner that returns plausible junk instead of doing the expensive work.
IOTA’s public repository and technical paper document its architecture and planned scaling path. Its July 31 code history records a v4.9.2 update. The project has also described future work on larger models and compression.
Those are project artifacts, not an independent benchmark against a centralized training cluster. The distinction should remain visible until reproducible comparisons exist for training time, cost, model quality and failure recovery.
The TaoSwap snapshot showed one active miner on SN9. That count makes any sweeping claim about broad distributed participation premature. It does not erase the technical experiment. It tells us where the experiment stood at the captured moment.
Training, inference and evaluation are separate markets
The three cases make one point clear. The phrase “AI subnet” is too broad to explain what is being purchased.
Gradients rewards training methods inside competitions. OpenRoboto rewards submitted robotics checkpoints under a reproducibility protocol. IOTA coordinates pieces of one distributed pretraining run.
Other Bittensor subnets sell inference, data, ranking, search, agent workflows or evaluation. Some improve a model. Some improve the system around a model. Some never alter model weights at all.
The distinction changes the economics.
A training market must pay for experimentation, compute and evaluation. An inference market must control latency, availability and price. A data market must control provenance and quality. An agent market must decide whether a completed workflow was correct. A benchmark market must prevent contestants from learning the answer key.
Putting all of them under “machine learning” can hide the part that deserves scrutiny.
What Yuma Consensus does and does not prove
Bittensor’s consensus documentation describes subjective utility. Validators score miner value. Yuma clips weights that lack enough stake support and penalizes selfish scoring by a minority coalition.
The mechanism coordinates rewards. It does not turn every validator metric into ground truth.
If validators measure the wrong benchmark consistently, consensus can reward the wrong target consistently. If a hidden evaluation leaks, miners can optimize toward the leak. If a task has no external customer, a technically fair contest can still produce something nobody needs.
The strongest machine learning subnets therefore need two proof loops.
The internal loop asks whether miners compete fairly and whether validators can reproduce the scores.
The external loop asks whether the resulting model, method or service performs useful work for someone outside the emission mechanism.
Bittensor is designed to make the first loop possible at scale. Every subnet still has to earn the second.
What builders should inspect
A serious review of a machine learning subnet should go deeper than the model name.
Read the miner contract. Find the exact artifact a miner submits. Check whether the code pins a commit, a container digest or a model revision. Look for a reproducible evaluator. Identify which data is public and which is hidden. Ask whether the benchmark can be gamed. Check recent commits and the current miner population. Then look for evidence that the output travels beyond the validator.
The important question is not whether a subnet says it trains AI.
The question is what gets better, how anyone can know and who wants the result enough to pay for it.
Samuel’s checkers program had a board, rules and a win condition. Bittensor opens the board to many markets. That makes the opportunity larger and the proof problem harder.
Sources
IBM: The games that helped AI evolve
Bittensor: Yuma Consensus
Gradients: G.O.D subnet repository and August 25 evaluation test commit
OpenRoboto: SN80 protocol repository and August 8 protocol update
Macrocosmos: IOTA SN9 repository and IOTA v4.9.2 commit
Tao Outsider: Machine learning thesis post
Market status: TaoSwap subnet explorer
Was this article useful?
One tap feedback helps us improve each post.