Cloudflare Web Search API Tutorial: REST, Workers and Citations

• Tapovan

Cloudflare's Web Search API went into open beta on October 2, 2026. Its three providers are priced 28 times apart: Ceramic costs $0.25 per 1,000 requests, Linkup $5.00 and Exa $7.00 (Cloudflare Providers page, retrieved 2026-10-06). The Cloudflare Web Search API runs through AI Gateway. You send a query and get back up to 10 results, each with a title, URL and description. Searches are billed to the same credits as your model calls, which makes it an easy web search API for AI agents that already run on Cloudflare. One more detail stands out on launch day. The changelog says all three providers support Zero Data Retention, but the Providers page, also dated October 2, lists Exa as "No". So the one-word provider parameter decides more than it appears to.

Beta product. Everything below reflects the docs as of 2026-10-06. Cloudflare has not published an API reference page for websearch yet (the expected URL returned 404 on that date), so the how-to-use page is the only spec. Parameters, prices and limits may change.

Cloudflare Web Search API: how to use it with curl

Before the first call you need three things:

  • A Cloudflare account and an AI Gateway. Every account has a gateway called default, and AI Gateway creates it automatically on the first authenticated request that uses it.
  • AI Gateway credits, or a provider API key stored on that gateway.
  • An API token with both Account > Workers AI > Read and Account > AI Gateway > Read. If either is missing, the request will fail with an auth error.

Export CLOUDFLARE_ACCOUNT_ID and CLOUDFLARE_API_TOKEN, then run this:

curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/websearch/ \
  --request POST \
  --header "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \
  --header "Content-Type: application/json" \
  --data '{
    "query": "What are some fun things to do in Salt Lake City as fall approaches?",
    "provider": "ceramic",
    "limit": 5,
    "options": {
      "gateway": { "id": "default" }
    }
  }'

Untested: verified against the How to use Web Search API docs. I made no live calls for this post.

The response body has two keys:

  • items: an array of results. Each one always has url and title. It may also have description (an excerpt of the page), imageUrl, faviconUrl and lastModifiedDate. Optional fields only appear when the provider returns them. The workers-types package says lastModifiedDate is a naive ISO-8601 datetime with no timezone, such as 2025-11-30T04:39:48.
  • metadata: query, requestId and latencyMs. Log requestId, because it is what you match against the gateway logs.
ParameterType and rulesDefault
querystring, required, 1 to 1,024 charactersnone
providerceramic, exa or linkupceramic
limitinteger, 1 to 1010
byokAliasstring matching ^[A-Za-z0-9_-]{1,64}$unset (see BYOK below)
options.gateway.id (REST) / gatewayId (Worker binding)string, requirednone

The same request in Windows PowerShell

$headers = @{ Authorization = "Bearer $env:CLOUDFLARE_API_TOKEN" }
$body = @{
  query    = "What are some fun things to do in Salt Lake City as fall approaches?"
  provider = "ceramic"
  limit    = 5
  options  = @{ gateway = @{ id = "default" } }
} | ConvertTo-Json -Depth 5
$r = Invoke-RestMethod -Method Post `
  -Uri "https://api.cloudflare.com/client/v4/accounts/$env:CLOUDFLARE_ACCOUNT_ID/ai/websearch/" `
  -Headers $headers -ContentType "application/json" -Body $body
$r.items | Select-Object title, url

Tested on Windows PowerShell 5.1.26100.9444 (parsing and the ConvertTo-Json output only). The HTTP call itself is untested. Keep -Depth 5. I tried a lower depth, and with -Depth 1 the nested object serialises as "gateway":"System.Collections.Hashtable", which the API cannot read.

Python, for apps that don't run on Workers

import os, requests
acct, token = os.environ["CLOUDFLARE_ACCOUNT_ID"], os.environ["CLOUDFLARE_API_TOKEN"]
r = requests.post(
    f"https://api.cloudflare.com/client/v4/accounts/{acct}/ai/websearch/",
    headers={"Authorization": f"Bearer {token}"},
    json={"query": "Cloudflare Web Search API providers", "provider": "linkup", "limit": 5,
          "options": {"gateway": {"id": "default"}}},
    timeout=30,
)
r.raise_for_status()
for i, item in enumerate(r.json()["items"], 1):
    print(f"[{i}] {item['title']} - {item['url']}")

Compiled with py_compile on Python 3.12.10. The API call is untested: verified against the how-to-use docs.

Where it sits: the Web Search API inside Cloudflare AI Gateway

Every search goes through a gateway you name, so in practice this is a Cloudflare AI Gateway web search API rather than a separate service. For each request, AI Gateway forwards the query to the chosen provider and normalises the response into the shape above. It then logs the request next to your model calls and bills it. Because of that, a search gets the same logging, analytics, billing and access controls as an inference request. This is a model-and-search gateway. It is a different layer from a tool gateway for MCP servers, and the MCP registry vs MCP gateway post explains that distinction.

Each provider has also agreed to Cloudflare's crawler rules. That means "verified bot compliance" (the crawler identifies itself and respects robots.txt) and source attribution: every result must link to where the content was crawled from.

Watch: AI Gateway's next evolution: an inference layer designed for agents by Cloudflare

This video was uploaded on 2026-04-20, about five and a half months before Web Search API launched, and it never mentions search. It still explains the gateway that searches now run through. Craig Dennis and AI Gateway PM Ming Lu cover the unified inference layer (0:52), failover between providers serving the same model (4:16), and unified billing (8:58), including loading money into one AI Gateway wallet (10:02). They also cover custom metadata for splitting spend per agent or per customer (11:31). Web search calls now land in that same wallet and log view.

Video by the Cloudflare YouTube channel.

Call it from a Worker with env.AI.websearch()

Inside a Worker, you don't need a token. Add the AI binding to your Wrangler config. Use either format:

{
  "$schema": "./node_modules/wrangler/config-schema.json",
  "name": "web-search-worker",
  "main": "src/index.ts",
  "compatibility_date": "2026-10-05",
  "ai": { "binding": "AI" }
}
name = "web-search-worker"
main = "src/index.ts"
compatibility_date = "2026-10-05"

[ai]
binding = "AI"

Untested: verified against the Workers binding docs. Wrangler 4.147.0 is the current version on npm, but I did not install or run it. The docs' sample uses 2026-10-05 with a comment telling you to set compatibility_date to today's date.

The minimal call looks like this. Note that the binding takes a flat gatewayId, whereas REST nests the id under options.gateway.id:

const response = await env.AI.websearch({
  gatewayId: "default",
  query: "What are some fun things to do in Salt Lake City as fall approaches?",
  provider: "exa",
  limit: 5,
});
const results = await response.json();

Untested: verified against the docs and the @cloudflare/workers-types 5.20261005.1 signature websearch(request: AiWebSearchRequest): Promise<Response>.

A common mistake is to treat the return value as data. websearch() returns a standard Response, so check response.ok before you call .json(). If you skip the check, a 400 or 429 becomes a confusing JSON parse error or an empty result further down your code.

A complete Cloudflare Web Search API example: grounded answers with citations

The docs show web search tool calling with Workers AI. Their example passes raw JSON back to the model and stops there, with no citation prompt and no source list. The Worker below goes further. The model decides whether it needs to search. If it does, the Worker runs one search and numbers the results, and the model answers with [n] markers that point to a sources array returned alongside the answer.

// Grounded answers with numbered citations: Workers AI + Web Search API.
// Env.AI is the Workers AI binding ([ai] binding = "AI" in wrangler config).
interface Env {
  AI: any; // `Ai` from @cloudflare/workers-types in a real project
}

const MODEL = "@cf/google/gemma-4-26b-a4b-it";
const GATEWAY = "default";
const MAX_DESC = 600; // trim each result description (Ceramic can return up to 8,000 chars)

type Item = { url: string; title: string; description?: string };

const tools = [
  {
    type: "function",
    function: {
      name: "web_search",
      description: "Search the web for current information.",
      parameters: {
        type: "object",
        properties: { query: { type: "string" } },
        required: ["query"],
      },
    },
  },
];

const SYSTEM = `Answer the user's question.
If you need current information, call web_search once.
When you use search results, cite them inline as [n] using ONLY the numbers given in the results.
Never invent URLs. If the results do not answer the question, say so.`;

// Gemma 4's published output schema is OpenAI-style (choices[0].message.tool_calls[].function,
// arguments = JSON string). Older Workers AI models return top-level tool_calls[{name, arguments}].
function firstToolCall(c: any): { id: string; name: string; args: any } | null {
  const oa = c?.choices?.[0]?.message?.tool_calls?.[0];
  if (oa?.function?.name) {
    const raw = oa.function.arguments;
    return { id: oa.id ?? "call_0", name: oa.function.name, args: typeof raw === "string" ? JSON.parse(raw || "{}") : raw };
  }
  const legacy = c?.tool_calls?.[0];
  if (legacy?.name) {
    const raw = legacy.arguments;
    return { id: legacy.id ?? "call_0", name: legacy.name, args: typeof raw === "string" ? JSON.parse(raw || "{}") : raw };
  }
  return null;
}

function textOf(c: any): string {
  return c?.choices?.[0]?.message?.content ?? c?.response ?? "";
}

function numbered(items: Item[]): string {
  return items
    .map((it, i) => `[${i + 1}] ${it.title}\nURL: ${it.url}\n${(it.description ?? "").slice(0, MAX_DESC)}`)
    .join("\n\n");
}

export default {
  async fetch(request: Request, env: Env): Promise<Response> {
    const question = new URL(request.url).searchParams.get("q") ?? "What did Cloudflare launch for web search in October 2026?";
    const messages: any[] = [
      { role: "system", content: SYSTEM },
      { role: "user", content: question },
    ];

    const first = await env.AI.run(MODEL, { messages, tools, tool_choice: "auto" }, { gateway: { id: GATEWAY } });
    const call = firstToolCall(first);
    if (!call || call.name !== "web_search") {
      return Response.json({ answer: textOf(first), sources: [] });
    }

    const query = String(call.args?.query ?? question).slice(0, 1024); // API max 1,024 chars
    const sr = await env.AI.websearch({ gatewayId: GATEWAY, query, provider: "ceramic", limit: 5 });
    if (!sr.ok) {
      const body = await sr.text();
      return Response.json({ error: "web search failed", status: sr.status, body }, { status: 502 });
    }
    const { items = [], metadata } = (await sr.json()) as { items: Item[]; metadata?: any };
    if (items.length === 0) {
      return Response.json({ answer: "No web results found for that query.", sources: [], query });
    }

    const final = await env.AI.run(
      MODEL,
      {
        messages: [
          ...messages,
          {
            role: "assistant",
            content: "",
            tool_calls: [{ id: call.id, type: "function", function: { name: "web_search", arguments: JSON.stringify({ query }) } }],
          },
          { role: "tool", tool_call_id: call.id, content: numbered(items) },
        ],
        max_completion_tokens: 600,
      },
      { gateway: { id: GATEWAY } },
    );

    return Response.json({
      answer: textOf(final),
      sources: items.map((it, i) => ({ n: i + 1, title: it.title, url: it.url })),
      searchRequestId: metadata?.requestId,
      searchLatencyMs: metadata?.latencyMs,
    });
  },
};

How this was tested: I ran this exact file on Node v24.15.0, using native type stripping and with env.AI mocked to return the documented response shapes. All 5 scenarios behaved as expected:

  1. An OpenAI-shape tool call led to a search, and the second model call received the tool message with the right tool_call_id. A 1,640-character description was trimmed to 600 characters, and the response came back with sources [1] and [2] and the search request id.
  2. The legacy top-level tool_calls shape produced the same result.
  3. When the model made no tool call, the Worker returned a plain answer with an empty sources list.
  4. When the search returned 400, the Worker returned HTTP 502 with the upstream status and body.
  5. When items was empty, the Worker returned "No web results found".

None of this ran against real Workers AI or Web Search. The live behaviour is untested. The code was checked against the tool-use docs and the gemma-4 input and output JSON schemas. I also skipped a full tsc type check.

Why the tool-call parser handles two shapes

As of 2026-10-06, the docs' tool example reads completion.tool_calls?.[0].name and toolCall.arguments.query. That is the older Workers AI shape. The published output schema for @cf/google/gemma-4-26b-a4b-it, the model the example itself uses, is OpenAI-style. Tool calls live at choices[0].message.tool_calls[].function, and arguments is a JSON string. The input schema also requires tool_call_id on role: "tool" messages, and the docs example leaves it out. I found this by comparing the documented schemas, not by running the docs example, so it may work in practice. Still, firstToolCall() accepts both shapes, and the follow-up messages include the assistant tool_calls turn and the tool_call_id. Either way, you are covered.

The citation prompt and why descriptions get trimmed

The system prompt does three jobs. It limits the model to one search. It tells the model to cite only the numbers it was given. And it tells the model to say when the results don't answer the question. The numbers map to the sources array, so the URLs the user sees come from your code, not from model output, and the model has no URL to invent. If you want better ordering before the second call, you can rank results with yes/no probabilities and keep only the top few.

MAX_DESC = 600 exists because Ceramic can return descriptions of up to 8,000 characters per result. When you call Ceramic directly, the description length is set by maxDescriptionLength (default 3,000, range 1,000 to 8,000). Cloudflare doesn't expose that parameter or say which value it sends. Five untrimmed results could add 40,000 characters to every prompt. The model has a 256,000-token context, so it would fit, but you would pay for those tokens on every call.

The model decides when to search, so treat the search tool like any other agent tool. Cap the calls, validate the arguments and log every request. The post on giving AI agents tools safely covers those guardrails in more depth.

Ceramic vs Exa vs Linkup: which provider to pick

If you have been comparing Tavily vs Exa vs Linkup, note that Tavily is not one of the options here. The API only accepts ceramic, exa and linkup. I included Tavily in the table (direct API, figures from tavily.com) because many readers compare against it.

Ceramic (default)ExaLinkupTavily (direct only)
Price per 1,000 searches$0.25$7.00$5.00$8 basic / $16 advanced (PAYG, $0.008 per credit)
ZDR via Cloudflare (Providers page)YesNo (changelog says yes)Yesn/a
Mode Cloudflare usesNot statedauto, with page highlights as the descriptionfast depth, raw results, no generated answern/a
Index or methodOwn index of more than 40 billion pagesKeyword search combined with embeddings-based searchPositioned for quick, cited results in agent tool callsBasic, advanced, fast, ultra-fast
Description styleLong, up to 8,000 charactersHighlightsRaw result snippetsn/a
Best forHigh-volume agent loops where cost mattersSemantic queries, as long as the queries aren't sensitiveQuick cited lookups with ZDRDomain and date filters, or up to 20 results

Prices from the Cloudflare Providers page and tavily.com, retrieved 2026-10-06. The Cloudflare prices match each vendor's own direct list price.

The Exa ZDR conflict, and what ZDR actually covers

Two pages dated October 2, 2026 disagree. The changelog states that all three providers offer Zero Data Retention "for requests made through Cloudflare." The Providers page lists Exa's Zero Data Retention as No and Ceramic's and Linkup's as Yes. Both pages were still live and unchanged on 2026-10-06. The table is the more specific statement, so treat Exa as non-ZDR and keep sensitive queries away from provider: "exa" until Cloudflare reconciles the two pages.

The scope of ZDR matters too. According to the Unified Billing docs, Cloudflare's ZDR applies only to requests billed to credits with Cloudflare-managed credentials. It does not apply to BYOK. If you use your own key, retention depends on your own contract with the provider. Directly, Exa offers ZDR only on Enterprise plans, Linkup turns it on only when you request it, and Ceramic asks you to contact them. ZDR also does not control AI Gateway logging. Logs are on by default for each gateway, and you turn them off separately in the gateway settings.

What it costs per 1,000 agent tasks

The per-search price is only part of the bill. Here is what 1,000 agent tasks cost under these assumptions:

  • Each task makes 3 searches with limit: 5.
  • Each task uses 6,000 input tokens and 800 output tokens on gemma-4-26b-a4b-it, priced at $0.10 per million input tokens and $0.30 per million output tokens.
  • The model cost is therefore 1,000 × (6,000 × $0.10 + 800 × $0.30) / 1,000,000 = $0.84.
  • The 10,000 free Workers AI neurons per day are ignored.

The token counts are my assumptions, so substitute your own.

Search providerSearches (3,000)+ model ($0.84)All on credits, incl. 5% fee
Ceramic via Cloudflare$0.75$1.59$1.67
Linkup via Cloudflare$15.00$15.84$16.63
Exa via Cloudflare$21.00$21.84$22.93
Tavily basic, direct PAYG$24.00n/an/a
Tavily advanced, direct PAYG$48.00n/an/a

Computed with a small Python script on Python 3.12.10, using list prices retrieved 2026-10-06.

With Exa or Linkup, search is more than 90% of the total. With Ceramic it is about 47%. Two billing details are easy to miss:

  • "No additional markup" applies to each search, not to the credits. Cloudflare bills searches at the provider's list price, but Unified Billing adds a 5% fee on credit purchases, so $100 of credits costs $105. The last column includes that fee. It assumes the model calls are also paid with credits. Workers AI tokens are normally billed as neurons on your Workers plan unless the gateway is set to unified billing for Workers AI.
  • Exa is locked to auto through Cloudflare. That costs $7 per 1,000, while Exa's own instant search type costs $4 per 1,000 if you call it directly.

If you are also deciding where to route the model calls, see how Jev pricing compares across Cloudflare AI Gateway and OpenRouter.

Bring your own key: the alias rule and the 400 with no fallback

To be billed by the provider directly, go to AI Gateway, select your gateway, open Provider Keys, and add a Ceramic, Exa or Linkup key with an alias such as default. Cloudflare encrypts stored keys with Secrets Store and never sends the key in the request. Then reference the alias in your call:

const response = await env.AI.websearch({
  gatewayId: "default",
  query: "What is Cloudflare Workers?",
  provider: "exa",
  byokAlias: "default",
});

Untested: verified against the BYOK docs.

Which credential gets used depends on two rules:

  • byokAlias set: if that provider or alias isn't configured on the gateway, the request fails with 400. It does not fall back to credits. This is useful when you never want surprise credit spend.
  • byokAlias omitted: the gateway uses a stored key with the alias default if one exists. Otherwise the search is billed to your credits.

AI Gateway also has a gateway-wide Require provider credentials setting (byok_only: true) that returns 400 instead of using Cloudflare-managed credentials. The docs don't say whether this setting covers /ai/websearch/, so test it before you rely on it.

Gateway or direct: Web Search API vs calling Exa, Linkup or Tavily yourself

Choosing a search API for an LLM app comes down to whether you need the provider's extra features.

What you gain by going through Cloudflare:

  • One bill and one credits balance.
  • Gateway logs and analytics next to your model calls.
  • ZDR terms through Cloudflare for Ceramic and Linkup.
  • No provider keys in your code.
  • Switching providers is a one-parameter change.

What you lose:

  • A hard cap of 10 results. Directly, Exa allows up to 100 and Tavily up to 20.
  • Exa's instant, deep and contents options.
  • Linkup's deep depth and its sourcedAnswer and structured outputs.
  • Ceramic's maxDescriptionLength.
  • Tavily's domain include and exclude lists (up to 300 and 150 domains), date ranges and include_answer.

If you only need a query in and ranked links out, the Cloudflare Web Search API is the simpler route. If you need filters or full page contents, call the provider directly.

Limits, beta caveats and troubleshooting

The docs list only two limits: queries of up to 1,024 characters and up to 10 results per request. A few more apply in practice or may apply:

  • Request rate: AI Gateway limits Unified Billing requests to 200 per 60 seconds per gateway and returns 429 when you exceed that. BYOK requests are exempt. The limits page doesn't mention Web Search, so this likely applies but is not confirmed.
  • Provider rate limits: directly, Ceramic allows 20 QPS on pay-as-you-go, Linkup 10 QPS, and Exa 10 QPS on Starter or 25 on Developer. Cloudflare doesn't say whether these apply behind the gateway.
  • Type definitions: as of 2026-10-06, @cloudflare/workers-types 5.20261005.1 says limit "Defaults to 10 and is capped at 20", while the docs say the maximum is 10. Follow the docs. The types also declare an undocumented WEBSEARCH binding with a search() method. I spotted it in the types but it isn't documented, so I would not build on it yet.
  • Billing for failed or empty searches: Linkup doesn't charge for errors or empty results when you call it directly. Cloudflare's docs don't say how errored or empty searches are billed.
SymptomLikely causeWhat to do
Auth error on the REST callToken is missing Workers AI Read or AI Gateway Read (the exact status code isn't documented for websearch)Edit the token so it has both Account-level read permissions
400 with byokAlias setThat provider or alias isn't configured on the gatewayAdd the key under Provider Keys, or remove byokAlias to fall back to credits
400 with no BYOK in the requestThe gateway may have Require provider credentials turned on (coverage for websearch is unverified)Store a key or turn the setting off
429Probably the Unified Billing limit of 200 requests per 60 seconds per gatewayBack off, spread traffic across gateways, or use BYOK
Request rejectedQuery longer than 1,024 characters, or limit above 10Slice model-generated queries with .slice(0, 1024) and keep limit at 10 or less
items: []No results (this case isn't documented)Return a "no results" answer instead of calling the model with nothing
PowerShell request fails, body shows System.Collections.HashtableConvertTo-Json depth too lowUse -Depth 5
Tool call never detectedYour code reads the legacy shape while the model returns the OpenAI shape, or the other way roundParse both shapes, as firstToolCall() does
Prompt costs spikeLong Ceramic descriptionsTrim each description, for example to 600 characters

FAQ

Is the Cloudflare Web Search API free?

No free tier is documented. Searches are billed to AI Gateway credits at list price: Ceramic $0.25, Linkup $5.00 and Exa $7.00 per 1,000. On top of that, credit purchases carry a 5% fee. With BYOK, the provider bills you under your own plan. The docs don't say whether provider free tiers, such as Ceramic's 1,000 free queries, apply when you use BYOK.

Can I call it outside Cloudflare Workers?

Yes. Any backend can call POST /client/v4/accounts/{account_id}/ai/websearch/ with a token that has Workers AI Read and AI Gateway Read. The curl, PowerShell and Python examples above all use that endpoint.

Is my search data retained?

It depends on the provider and on how you pay. The Providers page lists ZDR for Ceramic and Linkup but not for Exa, even though the changelog says all three have it. ZDR covers only credit-billed requests, not BYOK. AI Gateway logging is a separate setting that is on by default.

Can I pair it with OpenAI or Anthropic models?

Yes. The same AI binding runs third-party models through the gateway, for example env.AI.run("openai/gpt-4.1-mini", {...}, { gateway: { id } }), and those calls are billed to the same credits. Swap the model id in the Worker, then check how that model shapes its tool calls.

Is it a Tavily alternative?

For the basic job of sending a query and getting ranked links and snippets back, yes. It is also cheaper at list price: Ceramic costs $0.25 and Linkup $5 per 1,000 searches, against $8 for Tavily basic on pay-as-you-go. It is narrower, though. You get at most 10 results, no domain or date filters, no generated answer and fixed provider modes.

Last updated: 2026-10-06; tested on Node v24.15.0 (Worker logic with a mocked env.AI, 5/5 scenarios), Windows PowerShell 5.1.26100.9444 (parse and serialisation only) and Python 3.12.10 (py_compile and the cost script). Docs checked against @cloudflare/workers-types 5.20261005.1 and wrangler 4.147.0 (npm). No live Cloudflare API calls were made.

Last updated: October 06, 2026
an "open and free" initiative. Powered by Blogger.