Skip to content

Build1 publisher3 min readPublished

Cloudflare AI Search goes GA with pixel-level image search and a November 1 billing start

Cloudflare made AI Search generally available with native image embeddings and PDF OCR, and will start billing for it on November 1, 2026. Which embedding model an instance runs now decides whether an image query searches pixels or a caption.

The Engineer · Build desk

Illustration accompanying Cloudflare AI Search goes GA with pixel-level image search and a November 1 billing start

What happened

  • Cloudflare will charge for three things: the content a team ingests, the data it stores and the queries it runs.
  • A free tier stays available on all Workers plans after billing begins.
  • Text files such as Markdown, HTML, CSV and JSON, and PDFs, can now be up to 10 MiB, up from 4 MiB.
  • OCR for scanned PDFs is open to every account and is billed as image processing ingestion tokens.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost A file limit 2.5 times larger lets bigger documents into an index, and from November each of them adds to both the ingestion and the storage lines of the bill.
  • decision Turning on OCR for an archive of scanned PDFs becomes a spending decision, because every page it reads is charged as ingestion.
  • constraint Instances left on a text-only embedding model will match image queries only on what an auto-generated caption happened to describe.

AI Search is built from Workers AI, Vectorize, R2 and Browser Run [1]. A query is optionally rewritten, then embedded. Vector and keyword search run in parallel. The results are fused and optionally reranked, and the top chunks go back to the caller or to a generation model that writes an answer [10].

General availability changes the embedding step. At query time, AI Search checks whether the instance's embedding model supports images [6]. If it does, that model embeds the query image directly, into the same vector space as the indexed images and text [6]. Today the model that does this is Qwen3-VL-Embedding [6]. On a text-only model, AI Search turns the query image into text with ToMarkdown and searches on the caption [7].

That fallback looks a lot like the old design. Before GA, Cloudflare ran object detection on an image, generated a caption and embedded the text, so images were searchable only through details the caption captured [8]. The company's own argument against captions is that catching every useful detail takes long or specialised captions, written with the eventual query in mind [9]. Keeping the caption path as a fallback is still good engineering. According to the post, it gives every model basic multimodal support, and models with native image support get the full visual signal [20].

The post's showcase query is "a bird perched on a leafy branch with plums" [18]. The use cases it lists are more ordinary: product discovery, screenshot matching, charts, diagrams and scanned documents [19].

On the native path both signals are kept. Pixels are embedded for visual retrieval and captions are retained for text [5]. I think Matryoshka Representation Learning is the right choice for a service that bills on stored data [11]. Cloudflare says MRL lets smaller embeddings keep useful information while storage stays manageable and search stays fast [5].

OCR sits at the front of ingestion. It reads the text from each page before chunking and embedding [13]. Cloudflare says the pricing is designed so a bill can be estimated before a single file is indexed [15]. In its description, parsing, chunking, embedding with Workers AI models, keyword indexing and reranking are the work that happens between ingestion and queries [16]. Preview pricing was announced at the August 2026 Agents Week [17]. Cloudflare says it will email a reminder before billing is switched on [3].

Migration is the cost still open. Native retrieval depends on the query image and the indexed content sharing one model's vector space [6]. I'd expect moving an existing text-only instance to Qwen3-VL-Embedding to mean re-embedding the corpus. The post does not say how that re-ingestion is billed.

What to watch

  • Whether the per-unit rates that take effect on November 1 match the preview pricing Cloudflare announced in August.
  • Which other Workers AI embedding models gain native image support alongside Qwen3-VL-Embedding.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories