AERIOXFLUX
← Agents & Jarvis
Agents & Jarvis · memory rag

Firecrawl Wants to Pay the Sources Your Agent Reads

A $75M Series B led by Smash Capital funds Alexandria, a catalog of official data providers and specialized indexes that Firecrawl says lifted agent answer quality 21% across 845 tasks.

Flux Desk·2026-09-23·5 min read

Firecrawl made its name turning messy web pages into clean text for language models. On September 22 it said scraping was never the goal. The company announced a $75 million Series B led by Smash Capital, with Altos Ventures, Nexus Venture Partners, Y Combinator, Freestyle and Offline Ventures participating. It also launched Alexandria, which it calls a knowledge library for AI agents.

The pitch, from CEO Caleb Peffer: "People have licensed data for training, but almost no one gets paid when an AI agent actually uses what they know."

That sentence explains why the product exists.

What Alexandria actually is

Alexandria puts four kinds of source behind one interface: official data providers, custom connectors, Firecrawl's own indexes, and the live web. The goal is that an agent has one way to find a source, see what it holds, and pull from it.

At launch, Firecrawl lists 88 official data providers, registries and publishers, exposing 504 capabilities across 28 categories, plus more than 113 million indexed sources. Three of the indexes are its own:

  • a Research Index of tens of millions of scientific paper abstracts
  • a Developer Index of tens of millions of documentation pages, READMEs, issues and merged pull requests
  • a Government Index of laws, regulations and ordinances

The docs name Particle for podcasts and CoinGecko for crypto data as example providers, and the blog says Firecrawl already pays Wikimedia Enterprise.

The workflow has three steps: discover, inspect, execute. An agent searches with Alexandria as a source, reads a tool's contract (inputs, response shape and a listed credit price), then runs it. Discovery and inspection are free. Only execution costs money, at the price the tool publishes. It is available through Firecrawl's MCP server, its API and SDKs, and its CLI.

The 21% claim, read carefully

Firecrawl says agents using Alexandria scored 21% higher on answer quality than agents using built-in web tools, across 845 tasks, with the same models and prompts and blind AI judging.

The setup is sound: same model, same prompts, only the retrieval layer changes. But it is Firecrawl's own evaluation, the judges are models, and the company reports the result "across the verticals we tested" without publishing the per-vertical breakdown here. Treat it as a strong vendor claim, not a settled benchmark.

The direction is still easy to believe. Built-in web search sends an agent to whatever ranks, which is often an SEO page summarizing a primary source. An agent that can reach the regulation, the repo issue or the provider's structured feed directly skips a layer of noise. For most retrieval-augmented work, answer quality depends less on the model than on whether the right document was retrieved at all.

Why this is a memory story, not a scraping story

Agent builders have had two ways to ground an agent. Build your own vector store over documents you control, or give the agent a search tool and trust the open web. The first is accurate but narrow. The second is broad but noisy.

Alexandria is a third option: a curated, priced catalog of external knowledge that the agent browses the same way it browses its own tools. The design idea is that a data source is described like a tool, with a schema and a price, so the agent can decide whether it is worth calling before spending anything. That is the MCP pattern applied to knowledge instead of actions.

If it works, retrieval starts to look less like search and more like procurement. The agent is not asking which page ranks highest. It is asking which authoritative source answers the question, and what it costs.

The economics are the real bet

The funding round will get the attention, but the payment model is the more consequential part. Firecrawl says it will use the money to expand Alexandria's coverage and to pay the researchers, publishers and data providers who contribute. It plans to open a self-service system soon so individuals, creators and organizations can earn when agents use their knowledge.

This targets an open problem in the agent economy. Publishers have spent two years watching AI companies take their content for training and answer questions that used to send traffic their way. Licensing deals have mostly been large lump sums between large companies. Very little has reached the long tail, and almost nothing is paid per use at inference time.

A per-call price, attached to a provider, that an agent pays when it pulls the data is a different arrangement. It turns a source from something to scrape into something to sell. If agents end up doing most of the reading on the internet, a working marketplace for that reading matters more than any single model release.

The risks are equally clear. A two-sided marketplace needs supply and demand at once. Providers join if agents pay, and agents route there only if the catalog beats free search on enough queries. Firecrawl has an unusual advantage here: 1.5 million users already building with its tools, by the company's and Smash Capital's count, and a very popular open-source repository. That is a ready distribution channel on the demand side.

What to watch

Three things will show whether Alexandria becomes infrastructure or stays a feature.

First, supply. Eighty-eight providers is a start. The number that matters is how many publishers join once the self-service program opens.

Second, independent evaluation. A third-party benchmark showing a lift over built-in web tools would do more for adoption than any launch post.

Third, pricing. If a task that uses Alexandria costs more than a model plus free search, many developers will stay with the noisier option. Free discovery helps, but execution prices decide the outcome.

For now, agent builders get a structured way to reach primary sources, and publishers get what could become a revenue line tied to agent use. Firecrawl is betting that the web crawler of the agent era will be a marketplace that pays its sources, not a better scraper.

#firecrawl#alexandria#series-b#agent-retrieval#data-licensing

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.