Analysis

Chutes SN64 trace maps a year of changing LLM workloads

A Chutes SN64 research trace covers 6.12 billion LLM requests and shows why prefix-aware routing matters as serving workloads change.

Written by Iris Vale Decentralized AI correspondent
Format
News report
Read time
5 min
Source trail
6 links
Review
Tao Outsider Engine
Repeated LLM request prefixes branching through a balanced inference server network for Chutes SN64.
Tao Outsider original editorial composition based on the Chutes serving-workload preprint. AI-assisted base image produced with Imagine Bridge.

Chutes has recirculated a research preprint built from one year of production serving metadata. It is not a new September 3 paper, and Tao Outsider did not replay the trace.

The scale is still unusual. The study reports 6,122,413,756 requests, approximately 35.8 trillion input tokens and 2.52 trillion output tokens across 9,174 model identifiers. It covers activity associated with 314,970 user IDs between April 11, 2025 and April 12, 2026.

Those IDs are not verified unique people or paying customers. The study’s 875,921 instance IDs are not a count of GPUs running at the same time. The useful evidence is the shape of the workload and the routing problem it creates.

The preprint has a date anomaly

The arXiv record currently displays a submission timestamp of July 3, 2026 even though its identifier begins with 2608. Chutes published the first thread used here on August 26 and included the work again in a September 1 research roundup.

Tao Outsider is preserving that mismatch rather than inventing a cleaner publication chronology. The article treats September 1 as recirculation by the project, not as the paper’s original release date.

The listed authors are William Nixon, Jon Durbin, Florian Standhartinger, Haryadi Gunawi and Juncheng Yang, with affiliations shown for the University of Chicago, Chutes and Harvard. The full HTML version uses a Company X label in places where the abstract record and project statement identify Chutes. That difference should also remain visible to readers assessing provenance.

What 6.12 billion requests can and cannot show

Request count is a poor substitute for compute demand. A short classification call and a long generation can each count as one request while consuming very different resources.

The token totals help explain the workload more accurately. Input volume exceeded output volume by a wide margin. The paper also reports that output lengths moved from hundreds of tokens toward fewer than 100 over the observation window.

Chutes interprets that shift as users becoming machines: agents call models frequently and consume short responses. The trace can support a narrower observation. The serving mix moved toward shorter outputs and frequent request patterns. It cannot identify every request as an autonomous agent or distinguish a human from software behind every user ID.

Traffic also changed over time. Aggregate volume rose through parts of 2025, peaked later in the year and declined in early 2026. A year-long total hides that uneven path.

This makes the dataset more useful for capacity planning than for a single growth headline.

Prefix caching creates a routing conflict

Long prompts often share a prefix. A coding session may resend repository context. An agent loop may repeat instructions and tool descriptions. If the next request lands on a server that already holds the corresponding key-value cache, the system can reuse work instead of recomputing the shared prefix.

The obvious load-balancing strategy can destroy that advantage. Sending each request to the least busy instance may spread work evenly, but it can also send the next turn to a machine that has never seen the prefix.

Chutes says its production routing tracks one-way hashes at power-of-two prefix boundaries and prefers instances that have already seen the prefix. The study’s simulations compare that locality-aware approach with alternatives.

These are three different evidence classes:

EvidenceWhat it supports
Observed traceThe distribution and evolution of requests in the captured metadata
SimulationEstimated cache-hit and load-balance behavior under tested policies
Chutes production descriptionHow the project says its live routing tracks prefixes

They should not be collapsed into one benchmark.

The 5-7% result is a trade, not a free gain

Chutes says prefix-aware routing approached the theoretical cache-hit ceiling in simulation while adding roughly 5-7% load imbalance.

That result describes the evaluated traces and policies. It does not guarantee the same outcome for every model, prompt distribution, hardware fleet or traffic spike. A system can preserve more cache locality while concentrating work unevenly enough to create a queue elsewhere.

The engineering question is therefore not whether locality or balance wins. The router needs a policy for spending a limited amount of imbalance to recover useful cached work.

That policy becomes more important as repeated context grows. A large prefix can occupy substantial cache memory, and moving or recomputing it has a different cost from serving a small independent prompt.

Why this belongs in the Bittensor discussion

Chutes operates as Bittensor subnet 64, but the preprint is about an inference-serving system, not a token-performance study.

The September 3 TaoSwap snapshot showed 16 active miners on SN64 and an emission_value of 0.062936253. That provides current network context. It does not show how many requests miners served, which instances held each prefix, what the service earned or whether the simulated routing result improved subnet economics.

Decentralized inference adds another layer to the routing problem. Capacity can be distributed across operators and hardware. A useful scheduler needs current information about availability, locality and latency without pretending those fields are equally trustworthy or static.

The paper does not settle that system design. It gives builders a large workload trace against which routing assumptions can be challenged.

The trace is not yet an independently reproduced dataset

The preprint says release of the trace was approved and that it would be added. Tao Outsider did not locate and replay a public dataset during this review.

That means the reported counts, workload curves and simulation outputs remain claims in the research artifact. Reproducibility will improve when outside teams can inspect the release, document filtering choices and rerun the routing experiments.

Six billion requests make a strong headline. The harder contribution is showing how quickly the meaning of an inference request can change. A scheduler designed around yesterday’s average prompt may be wrong for tomorrow’s agent loop.

Sources

Chutes September 1 research roundup

Chutes original paper thread

Chutes prefix-aware routing explanation

A Year in LLM Serving on arXiv

Full preprint in arXiv HTML

TaoSwap subnet status API

Follow the Bittensor desk

Read the latest Bittensor stories with the same source discipline.