Meta's AI search index advances beyond Google and Bing dependence

Pieter Levels says a Meta employee described a plan to keep AI search queries from Google, adding an unverified motive to a project reported in 2024.

By · Published

Why it matters

A proprietary web index would give Meta control over retrieval costs, ranking and AI query data while reducing dependence on Google and Microsoft. It would also make Meta a new gatekeeper between publishers and billions of users across its apps.

Fragment of highly magnified source code text, indicating web data crawling (Duotone press photo, two-ink newspaper reproduction, heavy grain, dramatic halftone shadows, fine microdots visible)

Pieter Levels (@levelsio) said in an August 6th post on X that a person he identified as a Meta employee privately described Meta's effort to build its own web index, claiming the project would keep Meta AI's search traffic and resulting data away from Google.

Levels said he published the message with permission. The source was unnamed, and the post supplied no documents or other evidence establishing the employee's identity or the specific claim about Google's access to Meta AI search activity. That rationale remains uncorroborated.

Meta's underlying search project is established and dates back at least to 2024. The Information reported in October 2024 that Meta was developing a search engine capable of crawling the web and supplying current information to Meta AI. Reuters reported on October 28th, 2024 that the work was intended to reduce Meta's dependence on Google and Microsoft's Bing, which were supplying information about news, stocks and sports.

The distinction matters. The verified reporting supports a web index built for Meta AI's retrieval infrastructure. It does not establish that Meta plans to release a standalone, general-purpose search destination that directly mirrors Google Search.

Levels, a solo software founder whose portfolio includes Photo AI, Interior AI, Remote OK and Nomads, framed the private message as new information about the incentive behind Meta's investment. According to his account, sending AI-generated searches through Google would expose valuable query activity that Google could use to improve its own systems. The post does not specify what data Google would receive, what contractual safeguards apply, or whether the alleged concern involves model training, search ranking, product analytics or a combination of those uses.

Meta has continued assembling the pieces of an independent retrieval stack since the initial 2024 report. When Meta launched its standalone Meta AI app in April 2025, Meta said the assistant could search across the web, without identifying the underlying providers for each type of query.

Meta also added licensed sources rather than relying on crawling alone. In a December 2025 announcement, Meta named publishers including CNN, Fox News, Le Monde Group, People Inc. and USA Today as sources for timely information and links in Meta AI. Meta later updated that list in March 2026 with additional partners.

In June 2026, Meta introduced AI Mode inside Facebook search, grounding answers in public material from Facebook products including Groups and Reels. That gives Meta another proprietary corpus alongside licensed news feeds and the open web.

The crawling infrastructure has a visible public footprint. The user agent meta-webindexer/1.1 has appeared in website logs during 2026, and crawler references citing Meta's webmaster documentation describe it as an indexer used to make public pages eligible for Meta AI search answers and citations. Meta's search build is therefore well beyond a newly floated internal idea, even though Meta has not publicly detailed the index's coverage, ranking system or role across individual Meta AI products.

Owning the index gives Meta strategic advantages without requiring the specific Google-training explanation to be true. Meta can control freshness, ranking, retrieval costs and query data while reducing reliance on two competitors that operate their own AI assistants. It can also decide which publishers appear in answers, when users receive outbound links and how licensed material competes with crawled pages and content posted directly to Meta's platforms.

Levels' post adds an alleged internal justification to that established strategy. The evidence supports Meta's push toward search independence. It does not yet support the stronger claim that preventing Google from training on Meta AI searches is the reason driving it.

Reader comments

Conversation for this story loads after sign-in.