Lium has patched an ambiguity that sits at the intersection of distributed systems and a customer’s cloud bill.
One lium up command could create two GPU pods.
The failure could occur without a malicious server. A rent request might reach Lium and create a pod while its response times out or returns through a failing gateway. The SDK would see an error and retry the POST. The server could treat the second request as another rental.
Lium’s September 7 commit changes that sequence. The client now checks whether the first request produced a new pod before it sends anything again. If it must retry, it does so once with the same client-generated idempotency key.
The repair first landed on Lium’s main branch after v0.0.33. Lium then published numbered packages through v0.0.37 on September 8. The release pages include binary and Python artifacts, closing the earlier package-version gap without proving operator deployment.
September 8 update: the patch reaches numbered releases
Lium published v0.0.34, v0.0.35, v0.0.36 and v0.0.37 on September 8. The latest observed release, v0.0.37, was published at 12:48:53 UTC.
The release pages are sparse. They identify packaged artifacts rather than documenting production adoption or a separate user-facing feature. Their value in this story is narrower: the duplicate-rent repair is no longer only a post-v0.0.33 change on main.
Packaging does not show which SDK version customers have installed, whether every service component honors Idempotency-Key or whether an ambiguous rent has occurred in production. Tao Outsider still found no public incident count, verified overbilling or customer-impact report.
Why retrying a read is different from retrying a rent
Automatic retry is usually helpful for a read request. If a list of pods fails with a temporary server error, asking for the same list again does not create another resource.
A rent POST is different. The client cannot infer from a timeout that the server did nothing.
There are two possible realities behind the same error:
- The request never reached the service, so no pod exists.
- The service created the pod, but the client never received the confirmation.
Blindly sending the same payload again handles the first case and makes the second worse.
Lium’s old _request path applied a retry wrapper broadly. The new code adds a retry=False option for calls that must not be repeated automatically. Lium.up() uses that path for the rent endpoint.
The distinction is small in code and material in product behavior. A clean HTTP response alone cannot finish a cloud command. Completion requires the client and service to agree on which resource was created.
The new flow takes an inventory before the order
Before renting, the SDK asks for the IDs of pods that already exist.
That snapshot addresses another ambiguity. GPU splitting can place several pods on one node, and a default pod name can match an older resource. If the client searched by name alone after a failed request, it could return a stale pod and pretend the new rent succeeded.
The new helper excludes IDs captured before the request. When the rent response is uncertain, the client looks for a matching pod name on the selected executor that was not in the earlier set.
If such a pod appears, Lium.up() returns it and does not send a second rent request.
If nothing appears, the client waits and sends one more POST. Both attempts carry the same UUID in the Idempotency-Key header.
That key gives a server an identifier it can use to collapse two attempts into one operation. A server that does not support the header can ignore it, which is why the client-side inventory check remains important.
The exact pod ID now wins over a name guess
Lium’s rent route can return a successful response containing a pod_id rather than a complete pod record.
The updated SDK uses that ID to read the created pod back from the listing. A name search is the fallback when the response lacks the resource identifier.
When the API has already returned a pod ID and the listing endpoint is temporarily unavailable, the client returns a minimal pending record with that ID. It no longer turns a confirmed resource identifier into a generic failure simply because the next read did not work.
That design is useful beyond Lium. In any paid infrastructure API, a create response should give the client a stable identifier. A display name is not enough when names can repeat.
What the tests cover
The commit adds a focused test module around the rent path.
It checks that the first request carries an idempotency key, that the same key is reused for the second attempt and that a pod created during a timeout prevents a retry.
It also covers stale same-name pods, pods on another executor, a successful response containing only pod_id, a failed pre-rent snapshot and a listing endpoint that fails after the rent request.
One test states the customer-facing consequence directly. One lium up must create one pod.
Tests are evidence of intended behavior, not a production incident report. The repository does not establish how many duplicate rentals occurred, whether any customer was overbilled or whether the bug was exploited.
Tao Outsider found no public incident count or audited financial impact and makes no such claim.
What the patch still depends on
The fix reduces the duplicate path, but it does not make a universal exactly-once guarantee.
If the first rent succeeds, the response fails and the pod listing is also unavailable, the client can still reach its one permitted retry. Preventing a second server-side creation in that case depends on the service honoring the repeated Idempotency-Key.
The public commit adds the header on the client. It does not independently verify production support across every backend version.
Client adoption is another boundary. The patch is merged after v0.0.33. Users pinned to that numbered tag do not automatically receive code that landed later on main.
The safest release claim is therefore precise. Lium has merged a client-side repair and a reusable idempotency key for an ambiguous rent retry. The numbered package now shows how the client fix reaches users; server-side confirmation is still needed to establish end-to-end idempotency support.
Why this matters to a Bittensor compute subnet
Lium SN51 sells a familiar product category: rentable GPU infrastructure.
That makes operational correctness part of the subnet thesis. A buyer does not care that capacity is coordinated through Bittensor if a single command can ambiguously create two billable resources. Reliability, reconciliation and cleanup are product features, not background engineering.
At TaoSwap block 9,015,736, Lium had 62 active miners and a nonzero emission_value of 0.072413006. That is active-subnet context. It does not prove how many rentals run through the open SDK, how much customer revenue the service produces or whether all operators use the patched path.
The next evidence should be explicit server documentation for idempotency-key behavior, version-adoption evidence and a customer-visible way to reconcile an ambiguous create request.
For now, the P0 label is justified by the class of failure. The repair does not promise perfect distributed transactions. It gives one GPU order a much better chance of remaining one GPU order when the network stops answering clearly.
Sources
Lium P0 duplicate-rent repair commit
Lium changelog fragment for the fix
Lium v0.0.33, the numbered tag that preceded the fix
Was this article useful?
One tap feedback helps us improve each post.