API Documentation: Uncensored OpenAI-Compatible Endpoint
Get started with our uncensored openai-compatible endpoint in minutes. Use standard OpenAI SDKs to send text and receive unrestricted responses with predictable pricing and no vendor lock-in.
Base URL and Authentication
The API follows the standard OpenAI interface, making integration straightforward. Use the base URL https://api.uncensoredendpoint.com/v1 for all requests. Authentication relies on a single API key generated during signup. Include this key in the Authorization header as a Bearer token. Each account is limited to one active key, which can be regenerated at any time to revoke access instantly. This setup works with any client that supports standard OpenAI-compatible endpoints, ensuring you avoid proprietary lock-in.
First Request
Send a standard chat completion request to test the uncensored llm hosted on our infrastructure. The model ID is simply "uncensored". This endpoint does not filter lawful adult content, creative writing, or controversial topics, providing a clean output for unrestricted use. You can send up to 8 MB per request body. Ensure your client is configured to handle the standard JSON response structure.
curl https://api.uncensoredendpoint.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Verify the response includes the generated text. If you receive an error, check your key validity or account balance. The model operates with a 100,000 token context window, allowing for substantial input and output within a single request.
Python SDK Integration
Use the official OpenAI Python library to interact with the API. Configure the client with your specific base URL and API key. The library handles serialization and error parsing automatically. This approach is ideal for backend services or data processing pipelines where you need reliable, synchronous calls to the uncensored ai api.
from openai import OpenAI
client = OpenAI(base_url="https://api.uncensoredendpoint.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Remember that the model ID must be set to "uncensored". The SDK will manage the HTTP connections efficiently. You can adjust timeouts and retries as needed for your specific deployment environment. This method ensures you are using a verified, maintained client library.
Node SDK Integration
The Node.js OpenAI SDK works identically to the Python version. Initialize the client with your base URL and key. This is useful for server-side rendering or API gateways in JavaScript environments. The package handles stream handling and JSON parsing natively.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.uncensoredendpoint.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Set the model parameter to "uncensored". The SDK will return a response object containing the completion data. You can integrate this directly into Express routes or other Node.js frameworks. Ensure your environment supports the required Node.js version for the SDK you choose.
Streaming Responses
Enable streaming by setting the stream parameter to true. The API returns a Server-Sent Events (SSE) stream, not a single JSON object. This allows you to display text tokens as they are generated, improving perceived latency for end-users. Each event in the stream contains a partial chunk of the response.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Parse each SSE event to extract the delta content. This is crucial for chat interfaces where users expect real-time typing effects. The uncensored openai-compatible endpoint supports this standard streaming behavior, ensuring compatibility with most modern web clients.
Rate Limits and Constraints
Each API key is limited to 300 requests per minute. If you exceed this limit, the API returns a 429 status code indicating rate limiting. The maximum request body size is 8 MB. For authentication errors, such as an invalid key, you will receive a 401 status. While prepaid credit issues may occur, specific HTTP codes for payment states are not guaranteed to be consistent across all client implementations.
The context window is 100,000 tokens for both prompt and completion combined. This is a hard limit; exceeding it will result in an error. There are no SLAs or uptime guarantees provided. The service is designed for straightforward, high-volume usage without complex routing or multi-model aggregation.
Technical reference
Everything the endpoint can and cannot do, in one place — check it before you top up.
| Feature | Support |
|---|---|
| Compatibility | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| Methods | POST /v1/chat/completions · GET /v1/models |
| Model | uncensored |
| Authentication | Authorization: Bearer YOUR_KEY |
| Base URL | https://api.uncensoredendpoint.com/v1 |
| Context window | 100,000 tokens, input and output combined |
| SSE streaming | Supported (stream: true), usage included at the end |
| Sampling parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| Tools / tool calls | Supported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages |
| Completion length | up to 16,000 tokens per request (default 2,048) |
| Structured output | response_format: {"type": "json_object"} |
| Max body | 8 MB request body |
| Rate limit | 300/min per key |
| Response headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Parallel requests | 8 requests at the same time per key |
| Free trial | $0.50 for 7 days, no card |
| Billing | prepaid credit, charged by real token usage; errors and refusals are free |
| Credit expiry | paid credit never expires, no subscription |
| Top-up | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Token prices | input $0.25 / 1M tokens, output $1.00 / 1M tokens |
| Volume bonus | +5% on $50+, +10% on $100+ |
| Sign-in | Google or e-mail and password |
| Key management | one key per account, regenerate any time (the old one stops working) |
| Content policy | uncensored for adults; the only hard rule: no sexual content involving minors |
Error codes
Every error is JSON with a type you can switch on. You are never charged for an error.
| Code | Type | Meaning |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | balance is empty — top up, requests resume at once |
403 | content_blocked | refused by the content policy |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | request body larger than 8 MB |
429 | rate_limited · concurrency | over 300/min or 8 parallel — back off and retry |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
Is the uncensored ai api key tied to a specific model version?
Yes, the key works exclusively with the "uncensored" model ID. This is a single, open-weight model run on our own GPU servers, not a routing service between multiple vendors. You do not need to manage model versioning manually.
How does the context window work?
The 100,000 token limit applies to the sum of your input prompt and the generated output. If your combined token count exceeds this, the request will fail. This allows for large documents or long conversation histories in a single API call.
Can I use this endpoint for function calling?
Yes, the chat-completions endpoint supports tool and function calling. You can define tools in your request, and the model will return structured JSON output when appropriate. This works with standard OpenAI-compatible client libraries.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.