Model pools In testing
A pool is a group of machines serving the same model with the same settings. The gateway sends each request to the ready machine with the lightest load, so the answer does not depend on whose hardware took it. Putting a machine into a pool is how it earns while nobody is renting it.
What you need
- A host that already works — see Become a host. Pools run on the same agent.
- An accelerator that fits the pool: enough memory for the model, and the right compute platform. NVIDIA with CUDA is what the live pools use today.
- Room for the model files and the machine's disk, and a link that can fetch several gigabytes without you minding.
The public pools are defined by the platform. Which ones are open, what they serve and what they
require is shown in the panel and through GET /v1/models. You can also create a private pool
from a platform pool's template In testing: only your keys can
call it, its tokens cost nothing, and you pay rent for the machines you put into it on other people's
hosts.
Joining a pool
-
Create the machine
In the panel, create a machine as usual but pick a pool for it. On your own host that costs you nothing; on somebody else's it is an ordinary rental, with its contract and its per-minute price.
-
It prepares itself
The machine is built from the pool's environment image, fetches the model files and starts the engine that will serve them.
-
It reports ready
Once the model is loaded and answering, the machine joins the rotation and starts receiving requests from the gateway.
What is checked
Before the machine is created, not after — a mismatch you would only discover at load time costs an hour of downloading for nothing:
- The compute platform of your accelerator against the pool's.
- The accelerator's memory against the pool's minimum — per card, not added up across cards.
- The number of accelerators, when the pool splits a model across several.
- Processor, memory and disk against the environment image's minimums.
- Your driver against the version the image needs.
What fails is named, so you know whether it is a setting to change or hardware that does not fit.
The machine is closed
A pool machine is built without your key, without a session user and without a console. You can stop it, restart it and delete it — it is your hardware — but you cannot get inside it while it serves.
That is deliberate: the network answers requests from people who never chose your machine, and the only way that can be offered honestly is if the machine's holder cannot read what passes through it.
Getting the weights
- Model files come from the platform's own mirror, and from your host's cache when a machine of that pool was built on it before — the second machine of the same pool downloads far less.
- Every file is checked against its expected checksum whatever it came from, so a truncated or swapped file is refused rather than served.
- It is gigabytes over your own link, so the first machine of a pool takes a while.
- Once loaded, the model stays loaded: the engine runs as a service and keeps it warm rather than starting it per request.
Rating
Each machine carries a rating out of 100, built from what it actually did:
| What is measured | Weight |
|---|---|
| Requests that finished successfully | 30 |
| Availability | 25 |
| Speed against the pool's own median | 20 |
| How quickly it becomes ready after a start | 15 |
| Agreement with the rest of the pool on check questions | 10 |
- What has no measurements yet simply does not count. A new machine is not judged on a handful of requests; it gets less traffic until it has a record.
- A pool may require a minimum rating. Below it the machine stays in the pool but stops receiving requests, and earns its way back.
- Machines are asked identical check questions now and then and compared with each other. Read it for what it is: a sampling check that catches a machine drifting out of line, not protection against a holder who sets out to game it.
Pause and drain
You need your machine back at any moment, and a request that is mid-answer should not die for it:
- Pause stops new requests. What is already running finishes, and the machine reports draining until it does.
- Then it can be stopped or deleted. If you cannot wait, stopping it outright is allowed too.
- Resume puts it back only after a fresh check that it is actually answering.
- Availability hours do the same on a schedule you set: outside them the machine drains and rests, and it returns on its own In testing.
What you earn
- Requests are priced by the tokens they consume and produce, at the price fixed when the request was sent. Your share is that amount minus the platform's commission.
- The tokens are counted by the platform, not reported by your machine: what a machine says about its own work is not what anybody is billed for.
- A request that failed, or one the platform could not count properly, is not charged to anyone — the loss is ours, not the renter's and not yours.
- The panel's Pool machines section shows, per machine, the requests and tokens served, what you earned, and what the machine cost if it runs on a rented host — over 7, 30 or 90 days.
- Test credits only, with no monetary value and no payout, like everything else during the test.
What is not there yet
- Pools on Apple silicon: the engine exists, but no such pool is open In development.
- Pools with a model of your own, rather than one from a platform pool's template Planned.
- Keeping a pool machine's outbound traffic to an approved list is In development.
Questions about a pool your hardware could serve? Ask us in Telegram.