First request in five minutes
Get your uncensored API key and make your first request in under five minutes. This guide covers the OpenAI-compatible endpoints, streaming, and integration steps for the grok nsfw API.
Authentication & Base URL
To start using the uncensored model API, you need an API key and the correct base URL. Sign up on the Get API key page using your Google account or an email and password. The key is displayed immediately. No phone number is required, and the trial credit ($0.50) activates instantly without a card.
Configure your client to point to the official endpoint. The base URL is https://api.groknsfw.cc/v1. Pass your key in the Authorization header as a Bearer token. The model identifier is uncensored, an open-weight model tuned for lawful adult content without standard refusals.
Basic Chat Completion
The core functionality is the POST /v1/chat/completions endpoint. Send a JSON body with the model ID and a messages array containing your prompt. The API returns a text response. If you do not specify max_tokens, the output is capped at 2,048 tokens. Otherwise, you can request up to 16,000 tokens per request.
This is a text-only API. There are no embeddings, image, or audio endpoints. It is designed for developers who need raw text generation for adult or unrestricted themes.
curl https://api.groknsfw.cc/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Streaming Responses
For real-time output, enable streaming by setting stream: true. The API returns Server-Sent Events (SSE). Each chunk contains partial text. The final chunk includes the complete token usage statistics, including input and output token counts. This allows you to track costs accurately as the response generates.
Streaming is essential for chat interfaces or applications where latency matters. The context window supports up to 64,000 tokens for the combined prompt and completion.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Function Calling & Tools
The API supports function calling via the tools parameter. Define your functions in the tools array and set tool_choice to auto or a specific function name. The model will return a response with tool calls if the prompt requires it. You can then execute the function and send the results back in a subsequent message.
This feature works alongside streaming and JSON mode. It is useful for building agents or applications that need to interact with external systems while maintaining the uncensored nature of the model.
from openai import OpenAI
client = OpenAI(base_url="https://api.groknsfw.cc/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
JSON Mode
If your application requires structured data, set response_format to {"type": "json_object"}. The model will output strictly valid JSON. This is useful for parsing data, generating configuration files, or feeding structured inputs into other systems. Ensure your prompt clearly instructs the model to output JSON.
JSON mode is compatible with all other features, including streaming and function calling. It ensures that your integration receives predictable, parseable output without extra text.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.groknsfw.cc/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Token Usage & Limits
The API enforces strict limits: 300 requests per minute per key and 8 concurrent requests. The maximum request body size is 8 MB. If you exceed these limits, you will receive a 429 rate limit error. Authentication errors return 401, and insufficient credit returns 402. Errors and refusals are free; you only pay for successful token usage.
Pricing is $0.25 per 1M input tokens and $1.00 per 1M output tokens. Credit is topped up via crypto (USDT TRC20 or USDC Base) and never expires. The context window is 64,000 tokens total.
Questions and answers
Is this the official Grok API?
No, this is an independent service running an open-weight uncensored model. It is not XAI, Grok, or any other vendor's model. It uses the same OpenAI-compatible API format for easy integration.
What happens if I get a 402 error?
A 402 error means your prepaid credit is exhausted. You can top up immediately using USDT or USDC. Credit is charged by real token usage, and errors are free.
Can I use this for commercial projects?
Yes. The API is uncensored and allows lawful adult content. There are no subscription fees, only pay-as-you-go pricing. Your prompts are not used for training.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.