Get API key

GLM API: Switch your client in three lines

Connect your existing code to an uncensored GLM API in minutes by updating two configuration values. This quickstart guides you through the standard OpenAI-compatible endpoints using your prepaid credit.

Prerequisites: Get Your Key

Before writing code, you need an API key. Visit the Get API key page on this site to sign up with an email and password. No phone number or credit card is required to start. Upon registration, you receive a trial credit of $0.50 valid for 7 days, or you can top up with crypto (USDT or USDC). Your key is displayed immediately after signup. Store it securely, as you can regenerate it at any time to revoke the old one.

This key authenticates all requests to our glm api endpoint. We do not use your prompts for training, and the key is tied to a single account. If you need higher volume, you can purchase prepaid credits that never expire.

Configure Base URL and Key

To use our service as a drop-in replacement for OpenAI, you must point your SDK to our base URL. This is the only infrastructure change required. The base URL is https://api.glmapikey.com/v1. You also need to set your authorization header or SDK client key to the value you received during signup.

Most OpenAI-compatible clients, including the official Python and Node SDKs, allow you to override the base URL. This allows you to keep your existing application logic while switching the underlying model. Ensure your key is set correctly before making requests to avoid authentication errors.

Make Your First Chat Completion Request

Send a POST request to /v1/chat/completions to generate text. Our model accepts standard message structures with roles like user and assistant. The model ID you must specify is uncensored. This triggers the open-weight model hosted on our GPUs, which is tuned to answer without content refusals for lawful adult use.

The response will be JSON containing the generated text. If you encounter a 401 error, verify your key. If you see 402, your prepaid credit has been exhausted. You can top up starting from $10. Note that we do not support embeddings, images, or fine-tuning in this endpoint.

Use the Python SDK

The official OpenAI Python library works directly with our API. Initialize the client with your key and the specific base URL. This approach is ideal for backend services or scripts where you need structured JSON responses. The client handles serialization and retry logic automatically.

Pass the model name uncensored to ensure you are hitting our specific endpoint. The context window supports up to 100,000 tokens for both prompt and completion combined. This allows for substantial input data without truncation, provided you stay within the 8 MB request body limit.

Use the Node SDK

For JavaScript and TypeScript applications, the openai NPM package is compatible. Configure the client with the same base URL and key as the Python example. This enables server-side rendering or API endpoints that require uncensored text generation. The SDK streams data efficiently if configured correctly.

Tool calling is supported in this SDK. You can define function schemas in the request, and the model will return structured JSON outputs that match your schema. This makes it easy to integrate external data sources or actions into your application flow without parsing raw text.

Enable Streaming Responses

Set stream: true in your request to receive server-sent events (SSE). This allows you to display text to users as it is generated, improving perceived latency. The API returns chunks of text, which your client must concatenate or render incrementally.

Streaming is supported for both Python and Node SDKs. It is particularly useful for chat interfaces where users expect immediate feedback. Each chunk contains a portion of the response. The model continues generating until it completes or hits the context window limit. Ensure your client handles partial JSON structures correctly if using tool calling with streaming.

Limits, Errors, and Models

Our API enforces a rate limit of 300 requests per minute per key. If you exceed this, you will receive a 429 error. The maximum request body size is 8 MB. For model discovery, use the GET /v1/models endpoint. It returns a list of available models, including the uncensored model ID.

Common errors include 401 for invalid keys and 402 for insufficient credit. Credits do not expire, but trial credits are valid for 7 days. We do not offer SLAs or certifications. The model blocks sexual content involving minors but is otherwise uncensored. Use these limits to design your retry logic and user notifications appropriately.

cURL

curl https://api.glmapikey.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Python

from openai import OpenAI

client = OpenAI(base_url="https://api.glmapikey.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Node.js

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.glmapikey.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Streaming

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

API specifications

One table with every limit, feature and price that applies to your key.

FeatureSupport
API formatOpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key
AuthenticationBearer token in the Authorization header
Model IDuncensored
MethodsPOST /v1/chat/completions · GET /v1/models
Base URLhttps://api.glmapikey.com/v1
Other parameterstemperature, top_p, stop, seed and the two penalties are passed through
Max outputup to 16,000 tokens per request (default 2,048)
Tools / tool callsSupported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages
Context window100,000 tokens (prompt + completion together)
SSE streamingYes — server-sent events; the last chunk carries token usage
Structured outputresponse_format: {"type": "json_object"}
Request sizeup to 8 MB per request
Requests per minute300/min per key
Parallel requests8 requests at the same time per key
HeadersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Volume bonus+5% on $50+, +10% on $100+
Billingpay as you go from prepaid credit; nothing is charged for failed or refused requests
Trial credit$0.50 for 7 days, no card
Subscriptionpaid credit never expires, no subscription
Token pricesinput $0.25 / 1M tokens, output $1.00 / 1M tokens
Paymentcrypto: USDT on TRON or USDC on Base, $10–$500, any whole sum
Key managementone key per account, regenerate any time (the old one stops working)
Accountsign in with Google or with e-mail + password
Content policyuncensored for adults; the only hard rule: no sexual content involving minors

Error codes

Errors come back as JSON with a stable type; failed and refused requests are not billed.

HTTPTypeWhat to do
400bad_requestmalformed request or too long for the context window
401missing_key · invalid_key · key_revokedno key, wrong key, or a key replaced by a newer one
402no_creditout of credit; add credit and retry
403content_blockedrefused by the content policy
404not_foundonly /v1/chat/completions and /v1/models exist
413request_too_largerequest body larger than 8 MB
429rate_limited · concurrencyslow down: rate or parallel limit reached
503upstream_busytemporary overload, retry shortly

Questions and answers

Is this the official GLM API?

No, we are an independent service. We host an open-weight uncensored model and provide an OpenAI-compatible endpoint. We are not affiliated with the original GLM creators or any other major AI vendor.

Can I use my existing OpenAI key?

No, you need a key from our site. However, you can use the same SDK code you use for OpenAI by simply changing the base URL and key in your client configuration.

What happens if I run out of credits?

Requests will return a 402 error. You can top up your prepaid credit at any time using crypto (USDT or USDC). Your credits never expire, so you can add funds whenever you need them.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key