Quickstart: Integrate the Uncensored LLM API

Integrate our uncensored LLM API in minutes using standard OpenAI-compatible clients. This guide covers authentication, completion requests, streaming, and tool usage with zero content refusals.

Base URL & Authentication

To begin integration, point your client to the base URL https://api.unfilteredaichatbot.cc/v1. The API follows the standard OpenAI chat-completions structure, meaning you can use existing SDKs by simply updating the base URL and providing your API key.

Authentication relies on the Authorization header. Include your unique key in the format Bearer YOUR_API_KEY. The system validates this key on every request. If the key is missing or invalid, you will receive a 401 error. Each account is restricted to one active key; generating a new one immediately invalidates the previous key. This ensures straightforward access control without complex permission hierarchies.

The API key is displayed immediately upon signup via Google or email. No phone number is required, and the key remains valid as long as your prepaid credit covers usage. Errors and refusals do not consume credit, so you can test authentication safely.

Basic Completion Request

Send a POST request to /v1/chat/completions with your model ID set to uncensored. The payload requires a messages array containing your conversation history. This endpoint accepts text input and returns text output, adhering strictly to the chat-completions format.

Below is a standard curl example demonstrating a simple two-turn conversation. The model will respond without blocking on adult or controversial topics, provided they are lawful. Ensure you include the model field explicitly to target the correct uncensored instance.

For Python developers, the official OpenAI SDK works out of the box. Just initialize the client with your base URL and key.

from openai import OpenAI

client = OpenAI(base_url="https://api.unfilteredaichatbot.cc/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

And for Node.js users, the same structure applies. Initialize the client and call chat.completions.create.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.unfilteredaichatbot.cc/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Streaming Responses

Enable streaming by setting stream: true in your request payload. The API returns a Server-Sent Events (SSE) stream, delivering tokens incrementally as they are generated. This is ideal for real-time user interfaces where latency matters.

The final chunk of the stream contains the complete usage statistics, including input and output token counts. You can use this data to update your UI or track billing accurately. Note that token usage is only reported in the last chunk, so you must parse the entire stream to get accurate counts.

Here is a curl example showing the streaming request structure. Replace the placeholder with your actual key and message content.

curl https://api.unfilteredaichatbot.cc/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Streaming does not affect pricing or rate limits; each token generated still counts toward your usage and limits.

Function Calling

The API supports function calling via the tools and tool_choice parameters. Define your functions in the tools array, and the model will return structured JSON when appropriate. You can enforce tool usage by setting tool_choice to a specific function or auto for model decision.

This feature works within the same uncensored context, meaning the model will not refuse to call tools related to adult or niche topics unless they violate the hard content limit. Use response_format set to {"type": "json_object"} if you need strict JSON output for parsing.

Function calling is part of the standard chat-completions endpoint, so no additional configuration is needed beyond the request payload. Ensure your function definitions follow the OpenAI schema format for compatibility.

JSON Mode

For applications requiring structured data, set response_format to {"type": "json_object"}. This instructs the model to output valid JSON, reducing parsing errors in your downstream systems. This mode is useful for extracting data, generating configurations, or powering agent workflows.

JSON mode works alongside function calling but can also be used for direct text-to-JSON conversion. The model is uncensored, so it will generate JSON even for complex or unusual data structures without refusing due to content filters.

Remember that JSON mode still counts toward your token limits and billing. Ensure your prompt clearly defines the JSON structure to improve output quality. The 64,000 token context window applies to the total input and output combined.

Rate Limits & Errors

Each API key is limited to 300 requests per minute and 8 concurrent requests. The request body must not exceed 8 MB. If you exceed the rate limit, the API returns a 429 error. If your credit is exhausted, you receive a 402 error. Authentication failures result in a 401 error.

The context window is 64,000 tokens total, with a maximum output of 16,000 tokens per request (2,048 if max_tokens is unset). Errors and refusals are free, so you can test limits without consuming credit. Credit never expires, and prepaid amounts are charged based on real usage.

For detailed error handling, check the status codes. A 401 indicates an invalid or missing key, while 402 requires a top-up. Use the support page to resolve billing issues. No card is needed for the trial, but crypto is required for top-ups.

Streaming

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Questions and answers

How do I get an API key?

Sign up using 'Continue with Google' or email and password on the Get API key page. Your key is displayed immediately after registration. No phone number is required, and the key is active instantly.

What happens if I exceed my rate limit?

The API returns a 429 error. Each key allows 300 requests per minute and 8 concurrent requests. Errors and refusals do not count toward your credit balance, but they do count toward rate limits.

Is the trial credit renewable?

No, the $0.50 trial credit is one-time per person and valid for 7 days. After expiration or usage, you must top up with crypto (USDT TRC20 or USDC Base) to continue using the API.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key