Gateway API In testing
The ServerShare gateway speaks the OpenAI API. Point the official SDK at our base URL, use a ServerShare API key, and your requests run on open models served by machines in the network.
Quick start
The gateway is open to invited developers during the test. Ask for access in Telegram, sign in to the panel, and create an API key in the Developer console (see API keys). Then change the base URL — nothing else.
openai packagefrom openai import OpenAI
client = OpenAI(
base_url="https://servershare.io/v1",
api_key="YOUR_SERVERSHARE_KEY",
)
stream = client.chat.completions.create(
model="example-chat-model",
messages=[{"role": "user", "content": "Hello!"}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")
openai packageimport OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://servershare.io/v1",
apiKey: process.env.SERVERSHARE_API_KEY,
});
const answer = await client.chat.completions.create({
model: "example-chat-model",
messages: [{ role: "user", content: "Hello!" }],
});
console.log(answer.choices[0].message.content);
curl https://servershare.io/v1/chat/completions \
-H "Authorization: Bearer $SERVERSHARE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "example-chat-model", "messages": [{"role": "user", "content": "Hello!"}]}'
example-chat-model is a placeholder: take a real model name from
GET /v1/models.
API keys
Every request carries Authorization: Bearer <key>. Only an API key opens the gateway: the
session of the panel does not.
- A key looks like
ssk_followed by 43 characters. It is shown once, when it is created or rotated. We keep only its hash and first characters, so a lost key cannot be recovered — rotate it. - Give each application its own key with a name. Rotating issues a new key with the same name and limits and revokes the old one at the same moment; there is no overlap.
- Revoking stops a key on the very next request, on every server.
- A key may carry its own requests-per-minute and tokens-per-minute limits, set when it is created.
- You can hold up to 20 active keys.
Keys are created, rotated and revoked in the Developer console of the panel, which also shows your usage by day and model, your credits and the model catalogue with prices.
Models GET /v1/models
Lists the models of active pools in the OpenAI shape; GET /v1/models/{id} returns one of them,
which is what client.models.retrieve() calls. A pool is a set of machines serving the same model with
the same settings, so the answer does not depend on which machine took the request. Besides the OpenAI fields,
each model carries:
| Field | Meaning |
|---|---|
type | What kind of model it is: chat for every model today. |
context_length | The real maximum context of the pool, not the one the model was trained for. |
max_output_tokens | The longest answer the model can give: the context, which the prompt and the answer share. |
quantization | How the weights are quantized, if at all. |
license | The model's license. |
pricing | input_per_million and output_per_million, in test credits, as strings. |
ready_machines | How many machines can take a request right now. |
status | available, or unavailable when no machine is ready — the model stays listed. |
supports_tools | Whether a request with tools is served: the engine parses tool calls. |
supports_structured_output | Whether response_format with a JSON schema is served. |
pool, engine | The pool's name and its serving engine. |
Chat completions POST /v1/chat/completions
model and a non-empty messages array are required. The rest of the body goes to the
model's engine as you sent it, and its answer comes back untouched — whole, or as a server-sent event stream
with "stream": true.
- As with OpenAI, a stream ends with a chunk carrying
usageonly when you set"stream_options": {"include_usage": true}. Usage is counted by the platform either way. - If a machine fails before sending the first byte, the gateway retries on another machine of the pool, up to three attempts. If a stream breaks after it has started, the gateway continues it on another machine of the pool without repeating tokens; the client sees one uninterrupted response.
- Errors of the engine itself (a 400 for a bad parameter, for example) keep their status and message
and come in the same
{"error": {...}}shape as every other error. - Tool calls and structured output work where the pool's engine supports them; the model catalogue says so
in
supports_toolsandsupports_structured_output, and a request asking for a feature the model lacks is refused withcapability_not_supported. Image generation has its own endpoint,POST /v1/images/generations; embeddings and speech endpoints are not available yet. - When no machine of the pool can take a request, the platform may send it to a third-party model provider, if it has switched this on for the model. Every such request is recorded with the provider's name. See the Test terms.
Context tiers
One model can be served by more than one pool, differing in how much context they hold — say 4096 tokens on
one card and 32768 on another. The catalogue lists the model once and shows every context it has in
contexts; context_length is the largest of them.
- The size of your request is the prompt plus the answer you asked for. With no
max_tokensthe engine generates until the context runs out, so the gateway reserves a small allowance instead of zero. - Your request goes to the tightest pool that holds it. A large context is expensive — it eats the cache the machine serves everybody from — so spending it on a short request would deny a long one that has nowhere else to go.
- A request that fits no pool of the model is refused with
context_length_exceededbefore any machine is tried; the message names what you asked for and the largest context available. - Price follows the pool. The price in the catalogue is the tightest pool's — what you pay in the ordinary case. Under load, when the tight pool has no room, a short request may be served by a wider pool and billed at that pool's price. It is the usual price, not a guaranteed one.
Asynchronous jobs POST /v1/jobs
Some work does not fit a request: it runs for minutes or hours, and holding a connection open for it helps nobody. Such work is a job. You post it, get an id back at once, and ask about it later or let a webhook tell you.
| Request | What it does |
|---|---|
POST /v1/jobs | Places a job: its kind, the list of parts to compute, and optionally a webhook_url. Answers with the job's id and state. |
GET /v1/jobs | Your jobs, newest first. |
GET /v1/jobs/{job_id} | One job: its state and the state of its parts. |
- A job is cut into parts, and each part is handed to exactly one executor. A part that fails goes back to the queue and is retried; when the attempts run out, the part is marked failed and the job finishes incomplete rather than hanging.
- Jobs have no tokens to count, so they are not priced per token — see Credits.
- If you gave a
webhook_url, the platform calls it once when the job reaches its final state, and signs the call so you can tell it apart from anything else knocking at that address.
Usage labels
Your consumption is yours in one lump unless you say otherwise. Two labels let you split it up — the platform stores them with every attempt and adds up spending by them.
| Label | Where you put it |
|---|---|
end_user_id | The user field of the request body — the same field the OpenAI SDKs already have. |
project_id | The X-ServerShare-Project request header. |
- The platform does not read or interpret a label: what it means is known to your application alone. It only records it and lets you filter your consumption by it in the Developer console.
- Send an opaque identifier, not a name or an email address: a label is stored with the request record, and there is no reason for your users' personal data to be here.
Limits
Limits count across all our servers, over a sliding minute. A request over a limit gets a
429 with Retry-After in seconds.
| Limit | Default |
|---|---|
| Requests per minute, per account (all keys together) | 120 |
| Tokens per minute, per account | 200000 |
| Burst: requests per account within | 30 in 10 seconds |
| Requests and tokens per minute, per key | Set when the key is created; none by default |
- Tokens are counted after an answer, so a long answer can take you past a token limit once; the next request is refused until the minute slides.
- Every answer carries
x-ratelimit-limit-requests,x-ratelimit-remaining-requests,x-ratelimit-reset-requestsand the same fortokens, for the tightest limit that applies. - Only chat completions are limited;
/v1/modelsis not. - The same prompt sent many times a minute is flagged for review. It is not refused — evaluations repeat prompts too — and the prompt itself is not stored.
Credits
Requests are paid for in test credits, which have no monetary value. Credits are issued by an administrator; ask in Telegram.
- A model's price is per million input tokens and per million output tokens (
pricingin/v1/models). The price that applies is the one when you sent the request. - Tokens are counted by the gateway with the model's own tokenizer, not taken from the machine that
answered. The
usagein your response is that count — the one your credits are charged by, so your own logs reconcile with your consumption. Where the gateway could not count (a very long answer, or a response two machines finished between them) the engine's own numbers are passed through instead. - You pay only for the attempt that answered. Attempts that failed and were retried elsewhere cost nothing, and neither does an answer a machine broke off. If you close a stream yourself, you pay for what was generated until then.
- Most of the price goes to the owner of the machine that served you; the platform keeps a commission.
- Credits are deducted shortly after the answer. With no credits left, a request to a paid model is refused
with
insufficient_quotabefore it reaches any machine. Requests already running when the balance ran out still complete, so a balance can end slightly below zero. - Asynchronous jobs have no tokens to count, so they are priced per completed part, at a
rate the platform sets; a job that lost parts pays for the ones it kept. The rate is zero unless an
administrator has set one — while it is zero, jobs cost nothing, and your consumption log shows no entry
for them. Placing a job your balance cannot cover is refused with
insufficient_quotaat once, rather than after the hours of work it would take.
Errors
Errors come in the OpenAI shape, so the SDKs raise their usual exceptions:
{"error": {"message", "type", "param", "code"}}.
| Status | code | When | What to do |
|---|---|---|---|
| 400 | — | The body is not JSON, or model or messages is missing; param names the field. | Fix the request. |
| 400 | context_length_exceeded | The prompt plus the answer asked for exceeds the largest context this model has. | Shorten the prompt or lower max_tokens. |
| 400 | capability_not_supported | No pool of this model serves what the request asks for; param is tools or response_format. | Pick a model whose entry states that support. |
| 400 | content_policy_violation | The platform's content review refused the request; param is messages. | Change the request. Review runs on the platform's own node, and prompts are not sent to any third party. |
| 400 | invalid_image_request | An image request the platform cannot run: no prompt, a size off the 64-pixel grid or above the model's limit, n outside 1–4, an unknown style. param names the field; extensions are named as servershare.<field>. | Fix the named field. |
| 401 | invalid_api_key | No key, an unknown or revoked key, or a panel session instead of a key. | Send a valid API key. |
| 403 | account_blocked | The account owning the key is blocked. | Contact us. |
| 403 | key_scope_forbidden | The key is limited to part of the API — a studio key, which opens only /v1/models, /v1/chat/completions and /v1/images/generations. | Use a developer key from the panel. |
| 404 | model_not_found | No active pool serves this model. | Pick a model from /v1/models. |
| 429 | rate_limit_exceeded | A per-minute limit; type is requests or tokens. | Wait Retry-After seconds. |
| 429 | burst_limit_exceeded | Too many requests within a few seconds. | Wait Retry-After seconds and spread requests out. |
| 429 | insufficient_quota | No credits left for a paid model. Carries x-should-retry: false. | Do not retry; ask for credits. |
| 503 | model_overloaded | Every machine of the model is busy. | Retry after Retry-After. |
| 503 | no_available_machine | No machine is ready, or every attempt failed. | Retry after Retry-After. |
Any other status comes from the model's engine and is passed on unchanged.
Test phase
- Credits have no monetary value. Nothing is charged and nothing is paid out. Balances, prices and limits may change or be reset.
- The API may change. We keep to the OpenAI shape, but fields ServerShare adds, and limits, can change during the test.
- Hosts are individual hardware owners. Requests run in isolated virtual machines on Community hosts In testing. The memory of a running request is not protected from a host's owner. Verified hosts with hardware attestation are Planned; Confidential hosts with hardware-protected memory are Planned.
- Please don't send confidential data during the test — no personal data, secrets or anything you would not show a stranger.