llmocrapi.comDocs
Quickstart: Integrate Uncensored LLM into Your OCR App
Integrate our uncensored LLM into your OCR pipeline to extract, clean, and summarize raw text from documents without content filters interfering with your data. This guide covers the essential steps to connect, authenticate, and stream responses using the OpenAI-compatible API.
Prerequisites: API Key and Base URL
Before making any requests, you need an API key and the correct base URL. Sign up on the Get API key page with just an email and password. Your key is displayed immediately after signup, and no credit card is required for the trial.
The base URL for this API is https://api.llmocrapi.com/v1. This endpoint is compatible with standard OpenAI SDKs. You must configure your client to use this base URL and include your API key in the Authorization header. The model identifier you will use for all requests is uncensored, which runs on our own GPU servers and is designed to handle raw text extraction without refusing lawful adult or controversial content.
Basic Text Extraction Request
Start by sending a simple chat completion request to test the connection. This endpoint accepts text input and returns extracted or cleaned text. Use the POST /v1/chat/completions endpoint with the uncensored model. Ensure your request body includes the message content and the model ID.
This basic request verifies that your API key works and that the uncensored model responds as expected. It is the first step in building your OCR post-processing pipeline.
curl https://api.llmocrapi.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'If you receive a response with choices containing text, your integration is successful. If you get a 401 error, check your API key. If you get a 402 error, you may need to add prepaid credit.
Python SDK Integration
For Python developers, using the official OpenAI SDK is the recommended approach. Install the openai package and configure the base URL to point to our API. This allows you to use familiar methods like chat.completions.create().
Set the base_url parameter to https://api.llmocrapi.com/v1 and pass your API key. This ensures that all requests are routed correctly to our uncensored model. You can then send your OCR-extracted text as the user message and define a system prompt to instruct the model on how to clean or summarize the data.
from openai import OpenAI
client = OpenAI(base_url="https://api.llmocrapi.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)This method is ideal for batch processing documents where streaming is not required. It provides a straightforward way to integrate the LLM into your existing Python OCR scripts.
Node SDK Integration
Node.js developers can also use the OpenAI SDK to interact with the API. Install the openai package for Node and configure the base URL similarly to the Python setup. This allows you to leverage the same chat.completions endpoint with familiar JavaScript syntax.
Configure the client with your API key and the base URL https://api.llmocrapi.com/v1. This ensures compatibility with your existing Node.js applications. You can then send requests to extract or summarize text from your OCR pipeline.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.llmocrapi.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);This approach is suitable for server-side processing in Node environments, allowing you to handle OCR outputs efficiently without leaving your Node ecosystem.
Streaming Responses for Large Documents
When processing large documents, streaming responses can improve user experience by delivering text in chunks as it is generated. Use the stream: true parameter in your request to enable Server-Sent Events (SSE). This allows you to process text in real-time, which is useful for long OCR outputs or when displaying progress to users.
Streaming helps manage memory usage by not requiring the entire response to be held in memory before processing. It is particularly useful for documents with significant text that might otherwise cause latency issues.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Ensure your client handles SSE events correctly to parse the incremental text chunks. This method is ideal for applications that require real-time feedback during the OCR post-processing stage.
Rate Limits, Errors, and Context Windows
Be aware of the API limits to avoid service interruptions. You are limited to 300 requests per minute per API key. Each request body must not exceed 8 MB. If you exceed the rate limit, you will receive a 429 error. Ensure your application implements exponential backoff to handle these errors gracefully.
The context window is 100,000 tokens for both prompt and completion combined. If your OCR output exceeds this limit, you may need to split the document or truncate earlier parts. Common errors include 401 for invalid keys, 402 for insufficient credit, and 429 for rate limits. Monitor your usage and add prepaid credit as needed to maintain uninterrupted service.
API facts in one table
The numbers below are the real limits of this API, not marketing. Compare them with what your app needs.
| Item | Value |
|---|---|
| Protocol | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| Base URL | https://api.llmocrapi.com/v1 |
| API key | Authorization: Bearer YOUR_KEY |
| Model | uncensored |
| Methods | POST /v1/chat/completions · GET /v1/models |
| Max output | 16,000 tokens max; 2,048 if max_tokens is not set |
| Tools / tool calls | Supported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages |
| Sampling parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| JSON mode | response_format: {"type": "json_object"} |
| SSE streaming | Supported (stream: true), usage included at the end |
| Max context | 100,000 tokens, input and output combined |
| Max body | up to 8 MB per request |
| Headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Rate limit | 300/min per key |
| Parallel requests | 8 requests at the same time per key |
| Top-up | USDT (TRC20) or USDC (Base), any whole amount from $10 to $500 |
| Trial credit | $0.50 for 7 days, no card |
| Credit expiry | no monthly fee; paid credit does not expire |
| Price | input $0.25 / 1M tokens, output $1.00 / 1M tokens |
| How you pay | prepaid credit, charged by real token usage; errors and refusals are free |
| Volume bonus | +5% on $50+, +10% on $100+ |
| Keys | one active key per account; a new key replaces the old one |
| Sign-in | Google or e-mail and password |
| Content | adult content allowed; sexual content involving minors is refused |
Errors and what to do
Errors come back as JSON with a stable type; failed and refused requests are not billed.
| HTTP | Type | What to do |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | check the Authorization header or use your current key |
402 | no_credit | out of credit; add credit and retry |
403 | content_blocked | sexual content involving minors — refused, not billed |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | body over 8 MB |
429 | rate_limited · concurrency | slow down: rate or parallel limit reached |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
Is the uncensored model the same as GPT or Claude?
No. The <code>uncensored</code> model is an open-weight model run on our own GPU servers, tuned to answer without content refusals for lawful adult use. It is not GPT, Claude, Gemini, Grok, or any other vendor's model.
What happens if my prepaid credit runs out?
If your prepaid credit runs out, you will receive a 402 error on subsequent requests. You can add more credit via crypto (USDT or USDC) starting from $10. Your trial credit expires after 7 days, but prepaid credit never expires.
Can I use this API for image or video generation?
No. This API is a text chat-completions API. It does not support image, audio, or video generation. It is designed for text input and output, such as post-processing OCR text, summarization, or extraction.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.