Uncensored LLM for OCR post-processinghttps://api.llmocrapi.com/v1

llmocrapi.comDocs

Quickstart: Integrate Uncensored LLM into Your OCR App

Integrate our uncensored LLM into your OCR pipeline to extract, clean, and summarize raw text from documents without content filters interfering with your data. This guide covers the essential steps to connect, authenticate, and stream responses using the OpenAI-compatible API.

Prerequisites: API Key and Base URL

Before making any requests, you need an API key and the correct base URL. Sign up on the Get API key page with just an email and password. Your key is displayed immediately after signup, and no credit card is required for the trial.

The base URL for this API is https://api.llmocrapi.com/v1. This endpoint is compatible with standard OpenAI SDKs. You must configure your client to use this base URL and include your API key in the Authorization header. The model identifier you will use for all requests is uncensored, which runs on our own GPU servers and is designed to handle raw text extraction without refusing lawful adult or controversial content.

Basic Text Extraction Request

Start by sending a simple chat completion request to test the connection. This endpoint accepts text input and returns extracted or cleaned text. Use the POST /v1/chat/completions endpoint with the uncensored model. Ensure your request body includes the message content and the model ID.

This basic request verifies that your API key works and that the uncensored model responds as expected. It is the first step in building your OCR post-processing pipeline.

curl https://api.llmocrapi.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

If you receive a response with choices containing text, your integration is successful. If you get a 401 error, check your API key. If you get a 402 error, you may need to add prepaid credit.

Python SDK Integration

For Python developers, using the official OpenAI SDK is the recommended approach. Install the openai package and configure the base URL to point to our API. This allows you to use familiar methods like chat.completions.create().

Set the base_url parameter to https://api.llmocrapi.com/v1 and pass your API key. This ensures that all requests are routed correctly to our uncensored model. You can then send your OCR-extracted text as the user message and define a system prompt to instruct the model on how to clean or summarize the data.

from openai import OpenAI

client = OpenAI(base_url="https://api.llmocrapi.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

This method is ideal for batch processing documents where streaming is not required. It provides a straightforward way to integrate the LLM into your existing Python OCR scripts.

Node SDK Integration

Node.js developers can also use the OpenAI SDK to interact with the API. Install the openai package for Node and configure the base URL similarly to the Python setup. This allows you to leverage the same chat.completions endpoint with familiar JavaScript syntax.

Configure the client with your API key and the base URL https://api.llmocrapi.com/v1. This ensures compatibility with your existing Node.js applications. You can then send requests to extract or summarize text from your OCR pipeline.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.llmocrapi.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

This approach is suitable for server-side processing in Node environments, allowing you to handle OCR outputs efficiently without leaving your Node ecosystem.

Streaming Responses for Large Documents

When processing large documents, streaming responses can improve user experience by delivering text in chunks as it is generated. Use the stream: true parameter in your request to enable Server-Sent Events (SSE). This allows you to process text in real-time, which is useful for long OCR outputs or when displaying progress to users.

Streaming helps manage memory usage by not requiring the entire response to be held in memory before processing. It is particularly useful for documents with significant text that might otherwise cause latency issues.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Ensure your client handles SSE events correctly to parse the incremental text chunks. This method is ideal for applications that require real-time feedback during the OCR post-processing stage.

Rate Limits, Errors, and Context Windows

Be aware of the API limits to avoid service interruptions. You are limited to 300 requests per minute per API key. Each request body must not exceed 8 MB. If you exceed the rate limit, you will receive a 429 error. Ensure your application implements exponential backoff to handle these errors gracefully.

The context window is 100,000 tokens for both prompt and completion combined. If your OCR output exceeds this limit, you may need to split the document or truncate earlier parts. Common errors include 401 for invalid keys, 402 for insufficient credit, and 429 for rate limits. Monitor your usage and add prepaid credit as needed to maintain uninterrupted service.

API facts in one table

The numbers below are the real limits of this API, not marketing. Compare them with what your app needs.

ItemValue
ProtocolOpenAI Chat Completions schema; official openai SDKs work unchanged
Base URLhttps://api.llmocrapi.com/v1
API keyAuthorization: Bearer YOUR_KEY
Modeluncensored
MethodsPOST /v1/chat/completions · GET /v1/models
Max output16,000 tokens max; 2,048 if max_tokens is not set
Tools / tool callsSupported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages
Sampling parameterstemperature, top_p, stop, seed, presence_penalty, frequency_penalty
JSON moderesponse_format: {"type": "json_object"}
SSE streamingSupported (stream: true), usage included at the end
Max context100,000 tokens, input and output combined
Max bodyup to 8 MB per request
HeadersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Rate limit300/min per key
Parallel requests8 requests at the same time per key
Top-upUSDT (TRC20) or USDC (Base), any whole amount from $10 to $500
Trial credit$0.50 for 7 days, no card
Credit expiryno monthly fee; paid credit does not expire
Priceinput $0.25 / 1M tokens, output $1.00 / 1M tokens
How you payprepaid credit, charged by real token usage; errors and refusals are free
Volume bonus+5% on $50+, +10% on $100+
Keysone active key per account; a new key replaces the old one
Sign-inGoogle or e-mail and password
Contentadult content allowed; sexual content involving minors is refused

Errors and what to do

Errors come back as JSON with a stable type; failed and refused requests are not billed.

HTTPTypeWhat to do
400bad_requestinvalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend
401missing_key · invalid_key · key_revokedcheck the Authorization header or use your current key
402no_creditout of credit; add credit and retry
403content_blockedsexual content involving minors — refused, not billed
404not_foundonly /v1/chat/completions and /v1/models exist
413request_too_largebody over 8 MB
429rate_limited · concurrencyslow down: rate or parallel limit reached
503upstream_busymodel busy — retry in a few seconds

Questions and answers

Is the uncensored model the same as GPT or Claude?

No. The <code>uncensored</code> model is an open-weight model run on our own GPU servers, tuned to answer without content refusals for lawful adult use. It is not GPT, Claude, Gemini, Grok, or any other vendor's model.

What happens if my prepaid credit runs out?

If your prepaid credit runs out, you will receive a 402 error on subsequent requests. You can add more credit via crypto (USDT or USDC) starting from $10. Your trial credit expires after 7 days, but prepaid credit never expires.

Can I use this API for image or video generation?

No. This API is a text chat-completions API. It does not support image, audio, or video generation. It is designed for text input and output, such as post-processing OCR text, summarization, or extraction.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key