NEW! SearchBlox is now an Adobe Platinum Technology Partner. I Get Started.

NEW! SearchBlox is now an Adobe Platinum Technology Partner. I Get Started.

NEW! SearchBlox is now an Adobe Platinum Technology Partner. I Get Started.

Abstract purple SearchBlox release background

ARTIFICIAL INTELLIGENCE

AI INFERENCE

RAG

ENTERPRISE SEARCH

SEARCHBLOX SEARCHAI 12.2

Native AI Inference.

Built into SearchAI.

Native AI Inference.

Built into SearchAI.

Native AI Inference.

Built into SearchAI.

September 18, 2026

Native Inference built into SearchAI 12.2. Install SearchAI and supported models run inside the application - no external inference service to deploy, and no per-token costs.

Until now, using an LLM with SearchAI meant running a separate inference service or paying for an external API. With 12.2, the model runs inside the platform, so AI search, RAG, chatbots and agents can all run on hardware you own or rent. This post covers why that matters, how it works, and everything else that ships in the release.

WHY IT MATTERS

WHY IT MATTERS

Native Inference is now part of the SearchAI stack.

Native Inference is now part of the SearchAI stack.

Native Inference is now part of the SearchAI stack.

SearchAI can retrieve enterprise data, generate grounded answers and automate the next step without a separate inference service.

SearchAI can retrieve enterprise data, generate grounded answers and automate the next step without a separate inference service.

Download SearchAI 12.2

One less service to manage

Inference is built into SearchAI. Nothing external to deploy, integrate or monitor.

Data stays on your servers

Prompts, context and answers are processed inside your environment.

Scale on your own CPUs or GPUs

Start on CPU. Add NVIDIA GPUs or scale out in cluster mode as demand grows.

No per-token costs

Run on hardware you own or rent, with a fixed cost instead of a usage meter.

HOW IT WORKS

No external inference APIs. No per-token costs.

No external inference APIs. No per-token costs.

No external inference APIs. No per-token costs.

Native Inference changes where inference runs - not how your applications connect. Here is what that means for your architecture, your capacity and your costs.

01

ARCHITECTURE

No external services in the path

No external services in the path

Chat and generation, embeddings, reranking, vision and speech normally mean a separate provider and endpoint for each one. In 12.2 they all run inside SearchAI, on hardware you own - so nothing leaves your network, and there are no keys or per-token costs to manage.

Animated enterprise architecture overview

02

CAPACITY

Starts on your CPUs. Scales with cluster mode.

Starts on your CPUs. Scales with cluster mode.

Real work runs on the CPUs you already have - grounded answers, summarization, extraction, function calling, vision and speech. When more people and more applications arrive, cluster mode spreads that work across additional nodes. An NVIDIA GPU is an option, not a requirement.

Animated enterprise architecture overview

03

COST

A fixed cost, not a per-token meter

A fixed cost, not a per-token meter

Your inference cost is the hardware you run - and that is the whole bill. It does not move when usage doubles, when an agent loops, or when someone summarizes a long document. There is no token meter running in the background, so your cost stays predictable as adoption grows.

Animated enterprise architecture overview

VISION MODELS

VISION MODELS

Search and answer from images, not just text

Search and answer from images, not just text

A lot of business knowledge lives in pictures: scanned forms, product photos, charts in reports and recorded video. Vision models look at these the way a person would, pick out what matters and use it to answer questions.

A lot of business knowledge lives in pictures: scanned forms, product photos, charts in reports and recorded video. Vision models look at these the way a person would, pick out what matters and use it to answer questions.

WHAT VISION MODELS CAN READ

Photos

Forms

Charts

Video

Vision model

Gemma 4 · Qwen 3.5 · Qwen 3.8

Extracted fields

Grounded answer

Image search

EXTRACTED FIELDS

Understands photos, forms, charts and video

Ask about a chart in a report, a field on a form, a product photo or a moment in a video, and get an answer based on what the image actually shows. Point it at a pile of forms and the same reading comes back as fields - the invoice number, the date, the name - ready for whatever system needs them.

IMAGE UNDERSTANDING

Reads images in milliseconds

Gemma 4 understands images natively. Its 12B model takes about 19 ms to read an image, so visual questions do not slow the conversation down.

DESCRIPTION AND IMAGE SEARCH

Search text and images together

Every image is described and tagged as it arrives, so words and pictures live in the same search index and one question can return the right photo alongside the right paragraph. A reranker then puts the best matches first - all on a single server.

MODEL CHOICE

Choose the right vision model

Pick the model that fits your hardware and quality needs: Qwen 3.5 (2B up to 35B), Qwen 3.8 27B or Gemma 4 (E2B up to 26B).

MORE CAPABILITIES

MORE CAPABILITIES

More ways to put AI to work

More ways to put AI to work

More ways to put AI to work

Beyond text and images, the same endpoint your applications already call can listen, speak, use your tools and think through harder problems.

Speech

Turn recordings into text, read answers out loud, and create a voice from an audio sample.

Takes action with your tools

Let the AI call your systems to get things done. It works with every model, connects to MCP clients, and can return clean JSON your apps can use straight away.

Thinks harder when needed

For tougher questions, the model can reason longer before it answers. You choose how much effort to spend (Qwen 3.8 and Spark X2.5).

Try it before you build

A built-in console with 380 ready-made prompts across 13 industries, side-by-side model comparisons, and a Prompt Optimizer that suggests better prompts in one click.

Secure and predictable

HTTPS and API keys are built in. The same question gets the same answer, and every response shows how long it took.

Easy to scale out

Add servers and they find each other automatically - no extra load balancer to set up. Each request goes to the server that has the right model, is least busy, and already holds the conversation.

Key Enhancements - Highlights

What’s new - at a Glance

What’s new - at a Glance

What’s new - at a Glance

Download SearchAI 12.2

NEW

Native Inference

Native Inference

Supported models run inside SearchAI on CPU or NVIDIA GPUs. No external inference service, no per-token costs.

UPDATED

AI Overview

AI Overview

Grounded answers with citations - now with a streaming REST API and an MCP tool.

NEW

OpenAI-Compatible Endpoint

OpenAI-Compatible Endpoint

Direct access to the underlying LLM through /v1/chat/completions, with streaming, for MCP tool-calling integrations.

NEW

Vision-Based OCR

Vision-Based OCR

Scanned documents, including multi-page PDFs, are read by a vision LLM instead of Tesseract.

UPDATED

RAG Chunking

RAG Chunking

Five chunking strategies, with token-aware sizing and per-collection overrides.

NEW

Virtual Filesystem

Virtual Filesystem

Volume-scoped file storage with automatic indexing, versioning, static-site hosting and sandboxed scripts.

NEW

New Collections

New Collections

Jira, Alfresco, Drupal, Adobe Edge Delivery Services, 130+ SaaS apps via REST/OData, CMIS and 15+ more databases.

NEW

New SSO Providers

New SSO Providers

Amazon Cognito, Microsoft Entra ID, Google Workspace, generic OIDC and generic SAML 2.0.

UPDATED

Security Hardening

Security Hardening

Per-deployment session keys, hardened XML parsing, authenticated admin endpoints and updated dependencies.

NEW

New in 12.2

UPDATED

Improved from earlier releases

See the full SearchAI 12.2 changelog

USE CASES

USE CASES

Where teams put Native Inference to work

Where teams put Native Inference to work

Where teams put Native Inference to work

Start where keeping models close to your data matters most.

Start where keeping models close to your data matters most.

LEGAL

Contract analytics

Pull clauses, dates and obligations into structured fields without documents leaving your environment.

SUPPORT

Customer support chatbots

Answer from your knowledge base with citations, then automate the ticket or follow-up.

COMPLIANCE

Compliance review

Check policies and documents against internal rules, with human review where required.

IT OPS

Incident triage

Summarize incidents and logs, find the matching runbook and route the next step.

WORKPLACE

Employee knowledge assistants

Give teams grounded answers from intranets, wikis and file shares.

DATA

Document intelligence

Run large extraction and enrichment jobs on your own schedule and hardware.

Put it on a real workload

Put it on a real workload

See Native Inference on your own infrastructure.

See Native Inference on your own infrastructure.

Book a demo and bring a real workload, or download SearchAI 12.2 and run Native Inference today.

Book a demo and bring a real workload, or download SearchAI 12.2 and run Native Inference today.

{ "@context": "https://schema.org", "@type": "FAQPage", "mainEntityOfPage": { "@type": "WebPage", "@id": "https://www.searchblox.com/searchblox-12-2-native-ai-inference" }, "inLanguage": "en-US", "publisher": { "@type": "Organization", "name": "SearchBlox", "url": "https://www.searchblox.com" }, "mainEntity": [ { "@type": "Question", "name": "What is Native AI Inference in SearchBlox 12.2?", "acceptedAnswer": { "@type": "Answer", "text": "Native AI Inference means SearchBlox 12.2 runs AI models inside the SearchAI platform itself, rather than calling an external inference provider. Chat and generation, embeddings, reranking, vision and speech all execute on the same servers that run search and retrieval. There is no separate inference service to deploy, no API keys to manage, and no request leaves your network." } }, { "@type": "Question", "name": "Does SearchBlox 12.2 need a GPU to run AI inference?", "acceptedAnswer": { "@type": "Answer", "text": "No. SearchBlox 12.2 runs real inference workloads purely on CPUs, including grounded answers, summarization, extraction, function calling, vision and speech. An NVIDIA GPU is an option, not a requirement. Organizations typically add NVIDIA GPUs only when they want higher throughput for long generation or many concurrent users." } }, { "@type": "Question", "name": "Does SearchBlox 12.2 charge per token for AI inference?", "acceptedAnswer": { "@type": "Answer", "text": "No. SearchBlox 12.2 has no per-token costs. Because inference runs natively on hardware you own or rent, your inference cost is tied to the servers you run rather than to the volume of tokens your users generate. Cost does not rise when usage doubles, when an agent loops, or when someone summarizes a long document." } }, { "@type": "Question", "name": "Which AI capabilities run natively inside SearchAI 12.2?", "acceptedAnswer": { "@type": "Answer", "text": "SearchAI 12.2 runs five inference capabilities natively: chat and generation, embeddings, reranking, vision, and speech. Each of these would normally require a separate external provider and endpoint. In 12.2 they all run inside SearchAI, powering AI Search, RAG, ChatBot, AI Agents, AI Assistants and hyper-personalization as part of the same platform." } }, { "@type": "Question", "name": "How do you scale AI inference in SearchBlox 12.2?", "acceptedAnswer": { "@type": "Answer", "text": "SearchBlox 12.2 scales through cluster mode. Native Inference starts built into SearchAI on a single node, and when more people and more applications arrive, cluster mode spreads the work across additional nodes. Nodes can be CPU or NVIDIA GPU. Scaling adds capacity on your own infrastructure without introducing external inference APIs." } }, { "@type": "Question", "name": "Where can SearchBlox 12.2 Native Inference be deployed?", "acceptedAnswer": { "@type": "Answer", "text": "SearchBlox 12.2 Native Inference runs on-premises, on AWS or on GCP. You own or rent the hardware and SearchBlox supplies the software. Because inference is native to the platform, prompts, documents and responses stay inside your own network on every deployment option." } }, { "@type": "Question", "name": "What can the vision models in SearchBlox 12.2 do?", "acceptedAnswer": { "@type": "Answer", "text": "Vision models in SearchBlox 12.2 read photos, scanned forms, charts and recorded video, then answer questions based on what the image actually shows. They extract fields from documents such as an invoice number, date or name, describe and tag every image as it arrives so words and pictures live in the same search index, and return grounded answers alongside matching images." } }, { "@type": "Question", "name": "Which vision models does SearchBlox 12.2 support?", "acceptedAnswer": { "@type": "Answer", "text": "SearchBlox 12.2 supports Gemma 4, Qwen 3.5 and Qwen 3.8 for vision. You choose the model that fits your hardware and quality needs: Qwen 3.5 ranges from 2B to 35B parameters, Qwen 3.8 is available at 27B, and Gemma 4 ranges from E2B to 26B. Gemma 4 understands images natively, and its 12B model reads an image in roughly 19 milliseconds." } }, { "@type": "Question", "name": "Does any data leave the network when using SearchBlox 12.2 AI features?", "acceptedAnswer": { "@type": "Answer", "text": "No. Because SearchBlox 12.2 runs inference natively inside SearchAI on infrastructure you control, no prompt, document or response is sent to an external inference provider. There are zero external inference APIs in the path, which removes the API keys, third-party contracts and cross-network data exposure that separate model providers normally introduce." } }, { "@type": "Question", "name": "How is SearchBlox 12.2 different from connecting search to external model APIs?", "acceptedAnswer": { "@type": "Answer", "text": "Connecting search to external model APIs usually means a different provider for every capability: one API for chat and generation, another for embeddings, another for reranking, vision and speech. That is five services to connect, secure and pay for by the token. SearchBlox 12.2 replaces all five with inference that runs natively inside SearchAI, on hardware you already control." } }, { "@type": "Question", "name": "How do I get SearchBlox 12.2?", "acceptedAnswer": { "@type": "Answer", "text": "SearchBlox 12.2 is available to download from searchblox.com/downloads. Installing SearchAI installs Native Inference with it, so there is no separate inference component to set up. The full list of changes in the release is published in the SearchBlox developer changelog at developer.searchblox.com/changelog." } } ] }