
Native Inference built into SearchAI 12.2. Install SearchAI and supported models run inside the application - no external inference service to deploy, and no per-token costs.
Until now, using an LLM with SearchAI meant running a separate inference service or paying for an external API. With 12.2, the model runs inside the platform, so AI search, RAG, chatbots and agents can all run on hardware you own or rent. This post covers why that matters, how it works, and everything else that ships in the release.
One less service to manage
Inference is built into SearchAI. Nothing external to deploy, integrate or monitor.
Data stays on your servers
Prompts, context and answers are processed inside your environment.
Scale on your own CPUs or GPUs
Start on CPU. Add NVIDIA GPUs or scale out in cluster mode as demand grows.
No per-token costs
Run on hardware you own or rent, with a fixed cost instead of a usage meter.
HOW IT WORKS
Native Inference changes where inference runs - not how your applications connect. Here is what that means for your architecture, your capacity and your costs.
01
ARCHITECTURE
Chat and generation, embeddings, reranking, vision and speech normally mean a separate provider and endpoint for each one. In 12.2 they all run inside SearchAI, on hardware you own - so nothing leaves your network, and there are no keys or per-token costs to manage.

02
CAPACITY
Real work runs on the CPUs you already have - grounded answers, summarization, extraction, function calling, vision and speech. When more people and more applications arrive, cluster mode spreads that work across additional nodes. An NVIDIA GPU is an option, not a requirement.

03
COST
Your inference cost is the hardware you run - and that is the whole bill. It does not move when usage doubles, when an agent loops, or when someone summarizes a long document. There is no token meter running in the background, so your cost stays predictable as adoption grows.


WHAT VISION MODELS CAN READ
Photos
Forms
Charts
Video
Vision model
Gemma 4 · Qwen 3.5 · Qwen 3.8
Extracted fields
Grounded answer
Image search
EXTRACTED FIELDS
Understands photos, forms, charts and video
Ask about a chart in a report, a field on a form, a product photo or a moment in a video, and get an answer based on what the image actually shows. Point it at a pile of forms and the same reading comes back as fields - the invoice number, the date, the name - ready for whatever system needs them.
IMAGE UNDERSTANDING
Reads images in milliseconds
Gemma 4 understands images natively. Its 12B model takes about 19 ms to read an image, so visual questions do not slow the conversation down.
DESCRIPTION AND IMAGE SEARCH
Search text and images together
Every image is described and tagged as it arrives, so words and pictures live in the same search index and one question can return the right photo alongside the right paragraph. A reranker then puts the best matches first - all on a single server.
MODEL CHOICE
Choose the right vision model
Pick the model that fits your hardware and quality needs: Qwen 3.5 (2B up to 35B), Qwen 3.8 27B or Gemma 4 (E2B up to 26B).
Beyond text and images, the same endpoint your applications already call can listen, speak, use your tools and think through harder problems.
Speech
Turn recordings into text, read answers out loud, and create a voice from an audio sample.
Takes action with your tools
Let the AI call your systems to get things done. It works with every model, connects to MCP clients, and can return clean JSON your apps can use straight away.
Thinks harder when needed
For tougher questions, the model can reason longer before it answers. You choose how much effort to spend (Qwen 3.8 and Spark X2.5).
Try it before you build
A built-in console with 380 ready-made prompts across 13 industries, side-by-side model comparisons, and a Prompt Optimizer that suggests better prompts in one click.
Secure and predictable
HTTPS and API keys are built in. The same question gets the same answer, and every response shows how long it took.
Easy to scale out
Add servers and they find each other automatically - no extra load balancer to set up. Each request goes to the server that has the right model, is least busy, and already holds the conversation.
Key Enhancements - Highlights
Download SearchAI 12.2
NEW
Supported models run inside SearchAI on CPU or NVIDIA GPUs. No external inference service, no per-token costs.
UPDATED
Grounded answers with citations - now with a streaming REST API and an MCP tool.
NEW
Direct access to the underlying LLM through /v1/chat/completions, with streaming, for MCP tool-calling integrations.
NEW
Scanned documents, including multi-page PDFs, are read by a vision LLM instead of Tesseract.
UPDATED
Five chunking strategies, with token-aware sizing and per-collection overrides.
NEW
Volume-scoped file storage with automatic indexing, versioning, static-site hosting and sandboxed scripts.
NEW
Jira, Alfresco, Drupal, Adobe Edge Delivery Services, 130+ SaaS apps via REST/OData, CMIS and 15+ more databases.
NEW
Amazon Cognito, Microsoft Entra ID, Google Workspace, generic OIDC and generic SAML 2.0.
UPDATED
Per-deployment session keys, hardened XML parsing, authenticated admin endpoints and updated dependencies.
NEW
New in 12.2
UPDATED
Improved from earlier releases
See the full SearchAI 12.2 changelog

LEGAL
Contract analytics
Pull clauses, dates and obligations into structured fields without documents leaving your environment.

SUPPORT
Customer support chatbots
Answer from your knowledge base with citations, then automate the ticket or follow-up.

COMPLIANCE
Compliance review
Check policies and documents against internal rules, with human review where required.

IT OPS
Incident triage
Summarize incidents and logs, find the matching runbook and route the next step.

WORKPLACE
Employee knowledge assistants
Give teams grounded answers from intranets, wikis and file shares.

DATA
Document intelligence
Run large extraction and enrichment jobs on your own schedule and hardware.













