Open test. ServerShare runs on test credits only. Nothing is charged, nothing is paid out, and credits have no monetary value. Features, APIs and data may change or be reset. Test terms →

Gateway API In testing

The ServerShare gateway speaks the OpenAI API. Point the official SDK at our base URL, use a ServerShare API key, and your requests run on open models served by machines in the network.

Quick start

The gateway is open to invited developers during the test. Ask for access in Telegram, sign in to the panel, and create an API key in the Developer console (see API keys). Then change the base URL — nothing else.

Python — official openai package
from openai import OpenAI

client = OpenAI(
    base_url="https://servershare.io/v1",
    api_key="YOUR_SERVERSHARE_KEY",
)

stream = client.chat.completions.create(
    model="example-chat-model",
    messages=[{"role": "user", "content": "Hello!"}],
    stream=True,
)

for chunk in stream:
    if chunk.choices:
        print(chunk.choices[0].delta.content or "", end="")
JavaScript — official openai package
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://servershare.io/v1",
  apiKey: process.env.SERVERSHARE_API_KEY,
});

const answer = await client.chat.completions.create({
  model: "example-chat-model",
  messages: [{ role: "user", content: "Hello!" }],
});

console.log(answer.choices[0].message.content);
curl
curl https://servershare.io/v1/chat/completions \
  -H "Authorization: Bearer $SERVERSHARE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "example-chat-model", "messages": [{"role": "user", "content": "Hello!"}]}'

example-chat-model is a placeholder: take a real model name from GET /v1/models.

API keys

Every request carries Authorization: Bearer <key>. Only an API key opens the gateway: the session of the panel does not.

Keys are created, rotated and revoked in the Developer console of the panel, which also shows your usage by day and model, your credits and the model catalogue with prices.

Models GET /v1/models

Lists the models of active pools in the OpenAI shape; GET /v1/models/{id} returns one of them, which is what client.models.retrieve() calls. A pool is a set of machines serving the same model with the same settings, so the answer does not depend on which machine took the request. Besides the OpenAI fields, each model carries:

Fields ServerShare adds to a model
FieldMeaning
typeWhat kind of model it is: chat for every model today.
context_lengthThe real maximum context of the pool, not the one the model was trained for.
max_output_tokensThe longest answer the model can give: the context, which the prompt and the answer share.
quantizationHow the weights are quantized, if at all.
licenseThe model's license.
pricinginput_per_million and output_per_million, in test credits, as strings.
ready_machinesHow many machines can take a request right now.
statusavailable, or unavailable when no machine is ready — the model stays listed.
supports_toolsWhether a request with tools is served: the engine parses tool calls.
supports_structured_outputWhether response_format with a JSON schema is served.
pool, engineThe pool's name and its serving engine.

Chat completions POST /v1/chat/completions

model and a non-empty messages array are required. The rest of the body goes to the model's engine as you sent it, and its answer comes back untouched — whole, or as a server-sent event stream with "stream": true.

Context tiers

One model can be served by more than one pool, differing in how much context they hold — say 4096 tokens on one card and 32768 on another. The catalogue lists the model once and shows every context it has in contexts; context_length is the largest of them.

Asynchronous jobs POST /v1/jobs

Some work does not fit a request: it runs for minutes or hours, and holding a connection open for it helps nobody. Such work is a job. You post it, get an id back at once, and ask about it later or let a webhook tell you.

Job endpoints
RequestWhat it does
POST /v1/jobsPlaces a job: its kind, the list of parts to compute, and optionally a webhook_url. Answers with the job's id and state.
GET /v1/jobsYour jobs, newest first.
GET /v1/jobs/{job_id}One job: its state and the state of its parts.

Usage labels

Your consumption is yours in one lump unless you say otherwise. Two labels let you split it up — the platform stores them with every attempt and adds up spending by them.

Labels you can put on a request
LabelWhere you put it
end_user_idThe user field of the request body — the same field the OpenAI SDKs already have.
project_idThe X-ServerShare-Project request header.

Limits

Limits count across all our servers, over a sliding minute. A request over a limit gets a 429 with Retry-After in seconds.

Default limits
LimitDefault
Requests per minute, per account (all keys together)120
Tokens per minute, per account200000
Burst: requests per account within30 in 10 seconds
Requests and tokens per minute, per keySet when the key is created; none by default

Credits

Requests are paid for in test credits, which have no monetary value. Credits are issued by an administrator; ask in Telegram.

Errors

Errors come in the OpenAI shape, so the SDKs raise their usual exceptions: {"error": {"message", "type", "param", "code"}}.

Gateway errors
StatuscodeWhenWhat to do
400—The body is not JSON, or model or messages is missing; param names the field.Fix the request.
400context_length_exceededThe prompt plus the answer asked for exceeds the largest context this model has.Shorten the prompt or lower max_tokens.
400capability_not_supportedNo pool of this model serves what the request asks for; param is tools or response_format.Pick a model whose entry states that support.
400content_policy_violationThe platform's content review refused the request; param is messages.Change the request. Review runs on the platform's own node, and prompts are not sent to any third party.
400invalid_image_requestAn image request the platform cannot run: no prompt, a size off the 64-pixel grid or above the model's limit, n outside 1–4, an unknown style. param names the field; extensions are named as servershare.<field>.Fix the named field.
401invalid_api_keyNo key, an unknown or revoked key, or a panel session instead of a key.Send a valid API key.
403account_blockedThe account owning the key is blocked.Contact us.
403key_scope_forbiddenThe key is limited to part of the API — a studio key, which opens only /v1/models, /v1/chat/completions and /v1/images/generations.Use a developer key from the panel.
404model_not_foundNo active pool serves this model.Pick a model from /v1/models.
429rate_limit_exceededA per-minute limit; type is requests or tokens.Wait Retry-After seconds.
429burst_limit_exceededToo many requests within a few seconds.Wait Retry-After seconds and spread requests out.
429insufficient_quotaNo credits left for a paid model. Carries x-should-retry: false.Do not retry; ask for credits.
503model_overloadedEvery machine of the model is busy.Retry after Retry-After.
503no_available_machineNo machine is ready, or every attempt failed.Retry after Retry-After.

Any other status comes from the model's engine and is passed on unchanged.

Test phase

Test terms Acceptable use