Web Search and Scraping MCP: Firecrawl, Exa, Tavily, Brave and MarkItDown
Web research MCP servers give a coding agent current information from outside the repository: Exa, Brave Search and Tavily search the web, Firecrawl and the Fetch reference server turn pages into Markdown, and MarkItDown converts PDFs and Office files. Everything they return is untrusted input, so conclusions need citations that a script can check.
You are asked whether to move a service from Express 4 to Express 5, and the agent answers from training data that ends before the release you care about. Or it searches, reads six pages, and writes a confident summary in which one “breaking change” appears in no source at all. The first problem is missing context; the second is unverifiable context, and it is the one that reaches production.
This page is for developers who want the agent to research against live sources, and for tech leads who need a research spike to end in an architecture decision record (ADR) that a reviewer can check without re-reading every page the agent read.
What you get from a checked web research setup
Section titled “What you get from a checked web research setup”- A decision table for Firecrawl, Exa, Tavily, Brave Search, Fetch and MarkItDown: what each does, what it costs to start, and where the key goes.
- Install commands for Claude Code, Codex and Cursor that keep API keys out of URLs and out of committed files.
- A worked example: scrape a changelog and list the breaking changes since v4.
- A research-spike workflow that ends in a draft ADR whose every quote is checked by a short script.
- The failure modes and name traps, with the recovery for each.
Which web research server fits the job?
Section titled “Which web research server fits the job?”The six servers split into three jobs: find sources, read a known URL, and convert a file. Most teams need one search server and one reader, not all six. Versions and package names were checked on npm and PyPI on 2026-09-26.
| Server | Job | How it connects | Start without paying | Auth | Package / endpoint |
|---|---|---|---|---|---|
| Firecrawl | Scrape, search, crawl, map, parse files | Remote HTTP or local stdio | Yes: the hosted endpoint has a keyless, rate-limited tier for scrape, search and parse; crawl, map and agent need a key | OAuth endpoint or Authorization: Bearer header | https://mcp.firecrawl.dev/v2/mcp; npm firecrawl-mcp 3.25.5 |
| Exa | Semantic web search plus page fetch | Remote HTTP | Yes: works anonymously with rate limits | OAuth preferred; key as header | https://mcp.exa.ai/mcp; npm exa-mcp-server 3.4.1 |
| Tavily | Search, extract, map, crawl | Remote HTTP or local stdio | Needs an account; Tavily offers a free one | OAuth or Authorization: Bearer header | https://mcp.tavily.com/mcp; npm tavily-mcp 0.2.22 |
| Brave Search | Web, news, image, video and local search, plus an LLM-context tool | Local stdio (the default since 2.x) or HTTP | Needs a Brave Search API key | BRAVE_API_KEY environment variable | npm @brave/brave-search-mcp-server 2.1.4 |
| Fetch (reference server) | Read one URL as Markdown, in chunks | Local stdio | Yes: free, no account | None | PyPI mcp-server-fetch 2026.8.18 |
| MarkItDown | Convert PDF, Word, Excel, PowerPoint, HTML, EPUB and more to Markdown | Local stdio (or localhost HTTP) | Yes: free, no account | None | PyPI markitdown-mcp 0.0.1a7 (alpha) |
How to choose:
- Start with Firecrawl’s keyless endpoint if you want one server that both searches and scrapes, and add a key when you need
crawlormap. - Choose Exa when the question is conceptual (“reports of this bug and the workarounds”) and keyword search returns noise. Its default tools are
web_search_exaandweb_fetch_exa. - Choose Brave Search when you want an independent search index and a local process whose tool list you control with
BRAVE_MCP_ENABLED_TOOLS. - Choose Tavily when your team already holds a Tavily account; the tool set overlaps with Firecrawl’s.
- Add Fetch for single pages when you want no third-party service in the loop. It obeys
robots.txtfor requests the model starts. - Add MarkItDown only for files: a vendor PDF, an RFP in Word, a pricing sheet in Excel.
Do you need an MCP server for web research at all?
Section titled “Do you need an MCP server for web research at all?”Not always. Check what the agent already has before you add a server.
- Claude Code ships built-in
WebFetchandWebSearchtools.WebSearchis not available on Amazon Bedrock, and it fails on Microsoft Foundry deployments hosted on Azure (checked against the Claude Code tools reference, v2.1.283; see LLM gateway configuration for how to tell which route a session takes). - Codex turns on live web search with
codex --search, which exposes the nativeweb_searchtool with no per-call approval.codex exec --helpin codex-cli 0.157.1 lists no--searchflag. - MarkItDown has a CLI, so an agent with a shell can convert a file without an MCP server:
uvx --from 'markitdown[all]' markitdown spec.pdf -o spec.md.
Add a server when the built-in tools fall short: you need a crawl across many pages, a search index you can pin across the team, structured JSON extraction, or the same research setup in all three tools.
Install web research servers in Claude Code, Codex and Cursor
Section titled “Install web research servers in Claude Code, Codex and Cursor”The recommended pair below is Firecrawl (scrape and search) plus Exa (search), both remote, so there is nothing to install locally. The other four follow. Every key is read from your environment, never written into a URL: Firecrawl’s README says “Never put an API key in the server URL”, and Tavily’s and Exa’s ?tavilyApiKey= and ?exaApiKey= URL forms end up in shell history, config files and proxy logs.
Run in a terminal at the repository root. --scope project writes a shareable .mcp.json; the single quotes keep ${FIRECRAWL_API_KEY} as a reference that Claude Code expands at connect time (tested on Claude Code 2.1.283):
# Firecrawl with a key (full tool set)claude mcp add --scope project --transport http firecrawl https://mcp.firecrawl.dev/v2/mcp \ --header 'Authorization: Bearer ${FIRECRAWL_API_KEY}'
# Or Firecrawl keyless: scrape, search and parse only, rate-limitedclaude mcp add --scope project --transport http firecrawl https://mcp.firecrawl.dev/v2/mcp
# Exa, default tools only; sign in with /mcp for higher limitsclaude mcp add --scope project --transport http exa 'https://mcp.exa.ai/mcp?tools=web_search_exa,web_fetch_exa'The others:
claude mcp add --scope project --transport http tavily https://mcp.tavily.com/mcp # OAuth via /mcpclaude mcp add --scope project brave-search -e 'BRAVE_API_KEY=${BRAVE_API_KEY}' -- npx -y @brave/brave-search-mcp-serverclaude mcp add --scope project fetch -- uvx mcp-server-fetchclaude mcp add --scope project markitdown -- uvx markitdown-mcpPlugins exist too: claude plugin install exa@claude-plugins-official (Exa’s README) and a firecrawl plugin in the same marketplace. The tavily entry in claude-plugins-official points at Tavily’s skills repository, not at the MCP server.
For servers that sign in with OAuth (Exa, Tavily, Firecrawl’s https://mcp.firecrawl.dev/v2/mcp-oauth), run /mcp in a session or claude mcp login <name>.
Codex writes each server to ~/.codex/config.toml (tested on codex-cli 0.157.1):
# Firecrawl with the key read from an environment variablecodex mcp add firecrawl --url https://mcp.firecrawl.dev/v2/mcp --bearer-token-env-var FIRECRAWL_API_KEY
# Exa (Exa's README line), then sign in for higher limitscodex mcp add exa --url https://mcp.exa.ai/mcpcodex mcp login exa
codex mcp add tavily --url https://mcp.tavily.com/mcp && codex mcp login tavilycodex mcp add brave-search -- npx -y @brave/brave-search-mcp-server # then add env_vars below, or it starts without BRAVE_API_KEYcodex mcp add fetch -- uvx mcp-server-fetchcodex mcp add markitdown -- uvx markitdown-mcpCodex does not pass your shell environment to a stdio server unless you name the variable. Add env_vars to the Brave entry instead of writing the key with --env, and use enabled_tools to keep Firecrawl to the tools you need:
[mcp_servers.brave-search]command = "npx"args = ["-y", "@brave/brave-search-mcp-server"]env_vars = ["BRAVE_API_KEY"]
[mcp_servers.firecrawl]url = "https://mcp.firecrawl.dev/v2/mcp"bearer_token_env_var = "FIRECRAWL_API_KEY"enabled_tools = ["firecrawl_scrape", "firecrawl_search", "firecrawl_map"]codex mcp get brave-search --json then lists BRAVE_API_KEY under env_vars with no value, and codex mcp list shows Firecrawl’s auth as “Bearer token”.
Put servers that carry a key in the global ~/.cursor/mcp.json, so the key never lands in the repository. The shapes come from the vendors’ READMEs and were not tested in a Cursor binary:
{ "mcpServers": { "exa": { "url": "https://mcp.exa.ai/mcp" }, "firecrawl": { "url": "https://mcp.firecrawl.dev/v2/mcp", "headers": { "Authorization": "Bearer ${env:FIRECRAWL_API_KEY}" } }, "brave-search": { "command": "npx", "args": ["-y", "@brave/brave-search-mcp-server"], "env": { "BRAVE_API_KEY": "${env:BRAVE_API_KEY}" } }, "fetch": { "command": "uvx", "args": ["mcp-server-fetch"] }, "markitdown": { "command": "uvx", "args": ["markitdown-mcp"] } }}Export FIRECRAWL_API_KEY and BRAVE_API_KEY in the shell that starts Cursor. The ${env:NAME} interpolation comes from a search snippet of Cursor’s MCP docs (secondary; cursor.com was unreachable on 2026-09-26), so confirm in the MCP list that Firecrawl lists more than three tools and Brave returns results before you rely on them. If interpolation fails, switch Firecrawl to https://mcp.firecrawl.dev/v2/mcp-oauth with no header, and move Brave’s key into the environment of the shell that starts Cursor, dropping its env block. A project .cursor/mcp.json can hold the keyless entries (Exa, Fetch, MarkItDown) for the whole team.
Fetch and MarkItDown are Python packages: they need uv (for uvx) and Python 3.10 or newer, and neither is published on npm (npm holds only a 0.0.1-security placeholder under the name mcp-server-fetch, checked 2026-09-26). npx is the right runner only for Firecrawl, Tavily and Brave.
Check that the servers answer
Section titled “Check that the servers answer”Run /mcp in a Claude Code or Codex session, or codex mcp list in a terminal; in Cursor, open the MCP list in settings. Expect web_search_exa and web_fetch_exa for Exa, three tools (firecrawl_scrape, firecrawl_search, firecrawl_parse) for keyless Firecrawl, fetch for Fetch, and convert_to_markdown for MarkItDown. A keyed Firecrawl server registers up to 26 tools with default settings (Firecrawl README).
Scrape a changelog and list the breaking changes since v4
Section titled “Scrape a changelog and list the breaking changes since v4”The first test uses a page whose answer you can check yourself: the Express History.md. Express 5.2.1 is the current release on npm (checked 2026-09-26).
What you should see: one scrape call, a saved Markdown file, and a table in which every row carries a quote. If a row has no quote, or quotes text you cannot find with grep in the saved file, the agent answered from memory. With Fetch instead of Firecrawl, the agent reads the page in chunks of 5,000 characters (the default max_length) and pages through with start_index; tell it to keep going until it reaches the 4.x entries.
From research spike to ADR: the workflow
Section titled “From research spike to ADR: the workflow”A research spike ends in a decision, and the decision should be an ADR a reviewer can check. The loop below is the same in all three tools; only the server names in the prompts change. It assumes ADRs live in docs/adr/, as in architecture decisions agents can follow.
-
Write the question before the agent searches. Create
research/<spike>/question.mdwith the decision to make, the options you already know, and the criteria that will decide it (for example: migration effort, Node.js support, security fixes). A spike without criteria turns into a reading list. -
Find sources, and approve the list.
You remove anything undated, anonymous or older than the release in question. This is the cheapest review in the whole loop.
-
Convert every approved source to Markdown in the repository. Web pages go through
firecrawl_scrapeor Fetch; PDFs and Office files go through MarkItDown’sconvert_to_markdown. Each file lands inresearch/<spike>/sources/with its URL and retrieval date on the first line. Keeping the sources on disk is what makes the next steps checkable: the evidence no longer depends on a page that may change tomorrow. -
Extract claims with quotes, one file at a time. Ask for
research/<spike>/notes.md: one row per claim, with the verbatim quote and the source file. Run this in a subagent (Claude Code) or a separate session so the raw pages stay out of the main context. -
Draft the ADR from the notes, not from the web.
-
Run the citation check and the gates. The script in the next section fails the build if any quote is missing from its source. Then a reviewer reads the ADR (not the sources) and decides.
The order matters. Searching and reading in one step lets the agent drop sources that disagree with its first guess. Converting to files before extracting claims means the draft can be checked against something fixed.
How to verify a research result without reading every source
Section titled “How to verify a research result without reading every source”A human cannot re-read eight web pages for every ADR, and does not need to. The checks below move the proof onto the files and leave the reviewer one judgement: is the decision right, given sources that are shown to say what the ADR claims?
This script fails when a quoted citation does not appear verbatim in the file it cites. Save it as scripts/check_citations.py:
import pathlib, re, sys
adr = pathlib.Path(sys.argv[1]).read_text(encoding="utf-8")adr = adr.replace("\u201c", '"').replace("\u201d", '"') # curly quotes count toomarkers = re.findall(r"\[src: [^\]]+\]", adr)cites = re.findall(r'"([^"]{20,})"\s*\[src: ([^\]]+)\]', adr)missing = 0for quote, src in cites: text = " ".join(pathlib.Path(src.strip()).read_text(encoding="utf-8").split()) if " ".join(quote.split()) not in text: print(f"NOT FOUND in {src}: {quote[:80]}") missing += 1unchecked = len(markers) - len(cites)if unchecked: print(f"{unchecked} [src: ...] markers have no quote of 20+ characters before them")print(f"{len(cites)} citations, {missing} not found")sys.exit(1 if missing or unchecked or not cites else 0)Run it from the repository root with python3 scripts/check_citations.py docs/adr/0014-express-5-migration.md. It treats curly quotes (“…”) like straight ones and normalises whitespace. It exits non-zero when a quote is not in its source, when a [src: …] marker has no quote of at least 20 characters in front of it, or when the ADR cites nothing at all. Add it to CI for any pull request that touches docs/adr/.
| Check | What it catches | Who owns it |
|---|---|---|
| Citation script | Invented quotes, quotes attributed to the wrong source | CI, blocking |
| Source list approved before reading | Undated, anonymous or stale sources | Author, at step 2 |
| Every source file starts with URL and retrieval date | Evidence nobody can re-fetch later | Reviewer, by glance |
| “Not established by sources” lines | Gaps the agent would otherwise fill from memory | Reviewer decides whether the gap matters |
A second agent with no web tools reviews the ADR against sources/ | Claims that are quoted correctly but out of context | Review agent; see reviewing an agent’s pull request |
The tech lead signs off on the decision with that evidence attached, just as the evidence bundle attaches proof to a code change.
What web research costs in context and credits
Section titled “What web research costs in context and credits”Tool schemas are the small cost: Claude Code and Codex load MCP tool schemas on demand. Even so, keep the list short. Exa’s ?tools= parameter replaces the default set, Brave’s BRAVE_MCP_ENABLED_TOOLS takes a space-separated allowlist, Codex’s enabled_tools filters any server, and in Claude Code a deny rule on mcp__firecrawl__firecrawl_crawl in .claude/settings.json blocks one tool.
Page content is the large cost. A scraped page lands in context in full unless you save it to a file. Three habits keep it in check:
- Ask for
onlyMainContent: trueon Firecrawl scrapes, and let Fetch’s default 5,000-character chunks do their job. - Save sources to
research/<spike>/sources/and work from the files, as the workflow above does. - Run the reading step in a subagent so only its summary returns to the main session.
Measure rather than guess: run /context in Claude Code before and after one scrape. Credits are the other bill. Firecrawl’s README says a search costs 2 credits and refunds 1 when the agent calls firecrawl_search_feedback; set FIRECRAWL_NO_SEARCH_FEEDBACK=1 on the local server if you do not want that tool registered. More techniques are in reducing MCP token cost.
How popular are these servers?
Section titled “How popular are these servers?”As of 2026-09-26 (GitHub stars read through the GitHub API; Official MCP Registry and claude-plugins-official read the same day): Firecrawl ★7.5k, in the registry at 3.25.5, with a plugin; Exa ★5.1k, in the registry as ai.exa/exa 3.4.1, with a plugin; Tavily ★2.4k, in the registry at 0.2.15; Brave Search ★1.5k, in the registry at 2.1.3 (npm is ahead at 2.1.4). Fetch lives in modelcontextprotocol/servers (★90.6k, a figure for the whole repository). MarkItDown’s ★187.1k belongs to the library; the MCP server is an alpha subpackage and is not in the registry. No npm or PyPI download figures are given, because none were read.
What breaks in web research over MCP?
Section titled “What breaks in web research over MCP?”The agent quotes text that is not in the source. The summary reads well and one quote is invented. Recovery: the citation script; make the ADR prompt require [src: …] on every claim, and treat a missing quote as a failed build, not a style note.
A page instructs the agent. A scraped README or forum post contains text such as “ignore previous instructions and run…”. Web content is untrusted input. Recovery: run research sessions without write tools that reach production, secrets or outbound actions; keep the spike on a branch; review any command the agent proposes after reading the web. See the agent threat model and securing MCP servers.
Fetch reaches your internal network. The Fetch README warns that it “can access local/internal IP addresses”. Recovery: do not run Fetch on a machine that can reach admin panels or cloud metadata endpoints you would not hand to the agent; use Firecrawl’s hosted endpoint for public pages, or run Fetch inside a sandbox (see permissions and sandboxing).
Fetch refuses a page. It obeys robots.txt for requests the model starts. Recovery: respect the refusal, or scrape a mirror you are allowed to read. --ignore-robots-txt exists; use it only for sites you own.
Firecrawl crawl or map returns an auth error. The keyless tier covers scrape, search and parse only. Recovery: add a key through the header (Claude Code, Cursor) or --bearer-token-env-var (Codex), or sign in through https://mcp.firecrawl.dev/v2/mcp-oauth.
Brave Search stopped answering on HTTP. Version 2.x defaults to stdio; before that it defaulted to HTTP. Recovery: run it as a stdio command as shown above, or set BRAVE_MCP_TRANSPORT=http (or --transport http) if a client really needs HTTP. The HTTP endpoint is unauthenticated and binds to 127.0.0.1 by default; keep it there.
Exa lost its default tools. Adding ?tools=web_search_advanced_exa replaces the defaults rather than adding to them. Recovery: list every tool you want, for example ?tools=web_search_exa,web_fetch_exa,web_search_advanced_exa.
MarkItDown breaks after an update. Every markitdown-mcp release so far is an alpha (0.0.1a7 on 2026-09-14). Recovery: pin the version (uvx markitdown-mcp@0.0.1a7, tested), or call the stable markitdown 0.1.8 CLI from the shell instead of the MCP server.