Web Data & Scraping
Scraping, crawling, web search, and structured extraction servers.
Benchmark run 20260717T113200Z-7113 promoted 2026-07-17 — servers with a Tested badge carry measured results.
Methodology.
Benchmarked reviews
- Firecrawl MCP Server (100%)
- exa (100%)
- tavily-mcp (86%)
- apify-mcp-server (86%)
- brightdata-mcp (57%)
| Server | Status | Stars | Namespace | Transports | Description |
|---|---|---|---|---|---|
| apify-mcp-server | Tested | 2,068 | Domain-verified | remote | Extract data from any website with thousands of scrapers, crawlers, and automations on Apify Store ⚡ |
| brightdata-mcp | Tested | 2,516 | GitHub | stdio | Bright Data's Web MCP server enabling AI agents to search, extract & navigate the web |
| exa | Tested | 4,746 | Domain-verified | remote | Fast, intelligent web search and web crawling. New mcp tool: Exa-code is a context tool for coding |
| Firecrawl MCP Server | Tested | 7,000 | GitHub | stdio | MCP server for Firecrawl — search, scrape, and interact with the web. |
| tavily-mcp | Tested | 2,232 | GitHub | stdio | MCP server for advanced web search using Tavily |
| apify | Listed | 28 | Domain-verified | stdio | Build, deploy, and run web scraping and automation actors in the cloud |
| blockrun-mcp | Listed | 475 | GitHub | stdio | Web search, deep research, prediction markets & crypto data for AI agents. Pay per call via x402. |
| brave | Listed | 1,311 | Domain-verified | remote | Visit https://brave.com/search/api/ for a free API key. Search the web, local businesses, images,… |
| CRW Web Scraper | Listed | 437 | GitHub | stdio | Open-source web scraper for AI agents with scrape, crawl, and map tools |
| fastCRW | Listed | 437 | GitHub | remote | Scrape, crawl, map & search the web. Open-source, self-hostable Rust crawler & search for AI agents. |
| fortress | Listed | 381 | GitHub | stdio | Stealth browser for AI agents: fetch pages behind Cloudflare/DataDome/CAPTCHA, extract clean data. |
| luminati-io-brightdata-mcp | Listed | — | Domain-verified | remote | One MCP for the Web. Easily search, crawl, navigate, and extract websites without getting blocked.… |
| SearXNG Search | Listed | 1,055 | GitHub | stdio | MCP server for SearXNG — privacy-respecting web search with pagination, URL reading |
| mcp-server-search | Listed | 38,134 | GitHub | stdio | MCP server for web search operations |
| firecrawl | Listed | 28 | Domain-verified | stdio | Scrape, crawl, and extract structured data from websites at scale |
| Octocode MCP - AI Context Platform | Listed | 896 | GitHub | stdio | AI code research platform. Search, analyze, and extract insights from any GitHub repository. |
| OpenBrand | Listed | 769 | GitHub | stdio | Extract brand assets (logos, colors, backdrop images, brand name) from any website URL |
| Exa | Listed | 0 | GitHub | remote | Exa MCP — neural/semantic web search + content retrieval (exa.ai) |
| Firecrawl | Listed | 0 | GitHub | remote | Firecrawl MCP — wraps the Firecrawl API (firecrawl.dev) for web |
| Scrapling MCP Server | Listed | 70,243 | GitHub | stdio | Web scraping with stealth HTTP, real browsers, and Cloudflare bypass. CSS selectors supported. |
| Tavily MCP Server | Listed | 0 | GitHub | stdio | MCP server for advanced web search using Tavily API. |
| Skyvern | Listed | 22,516 | GitHub | remote | AI-powered browser automation — navigate, click, fill forms, and extract data from any website. |
| acrawl | Listed | 10 | GitHub | stdio | Autonomous web crawler. 17 browser tools + goal-driven run_goal agent. Single binary, stealth. |
| AEO Audit | Listed | 0 | Domain-verified | remote | AEO audit: score any website 0-100 for AI visibility. Checks schema, meta, content, AI crawlers. |
| aeo-mcp | Listed | 0 | GitHub | stdio | AEO tools for AI agents: crawler permissions, llms.txt, structured data, AI-readiness audit. |
| Aetheris-MCP — x402 Web Scraping MCP Server | Listed | 0 | GitHub | remote | x402-gated web scraping MCP server. Pay 0.01 USDC per page on Base. Returns clean Markdown via JSDOM |
| Afterpaths | Listed | 2 | GitHub | stdio | Session memory for AI coding agents. Search past sessions, extract rules, and track what works. |
| agent-data-api | Listed | — | Domain-verified | remote | Pay-per-call (x402/USDC-Base) web + crypto data tools for AI agents: audit, extract, crypto, DeFi. |
| AgentScrape | Listed | 2 | GitHub | remote | Pay-per-call web scraping for AI agents via x402 on Base USDC. Six tools, no signup. |
| Agent Utility API | Listed | — | GitHub | remote | x402 pay-per-call tools: company enrichment, PDF extraction, Amazon/KDP data, YouTube transcripts. |
| agent-vending-factory | Listed | — | GitHub | remote | Pay-per-call MCP tools via x402 USDC: ZAR prices, data extraction, Python sandbox, SA flights. |
| agentcast | Listed | 1 | GitHub | stdio | Structured-output enforcer: extract and validate JSON from messy LLM text. |
| agentfetch | Listed | 2 | GitHub | stdio | Token-budgeted web fetch for AI agents — auto-routes Jina, FireCrawl, Trafilatura, PDF. |
| agentql | Listed | 28 | Domain-verified | stdio | Query webpages and extract structured data using natural language |
| agentsbase — Email for AI Agents | Listed | — | Domain-verified | stdio | Email for AI agents. Create mailboxes, send/receive emails, and auto-extract verification codes. |
| ai-crawler-policy | Listed | 0 | GitHub | remote | Cloudflare Workers MCP server: ai-crawler-policy |
| ActableSite AI Crawler Monitor | Listed | 0 | Domain-verified | remote | Check AI crawler robots.txt policy and monitor public-site policy, sitemap, and llms.txt changes. |
| AI-First Scraper | Listed | 2 | GitHub | stdio | Three MCP tools: fetch_page, fetch_pages_batch, search_web. Ad-free Markdown for AI agents. |
| ai-hr-management-toolkit | Listed | 1 | GitHub | stdio | AI HR toolkit: 24 MCP tools for resume parsing, skill extraction & ATS management. |
| search | Listed | 0 | GitHub | remote | Collaborative, cache-first web search for agents — cited answers from a shared live-web pool. |
| AiPayGen — 65+ AI Tools as an MCP Server | Listed | 2 | GitHub | stdio, remote | 65+ AI tools as MCP: research, write, code, scrape, translate, RAG, agent memory, workflows |
| AIR SDK | Listed | 1 | GitHub | stdio | Collective intelligence for browser automation agents. Site capabilities, selectors, and extraction. |
| Alembica MCP | Listed | 1 | GitHub | stdio | MCP server for Alembica validation, extraction, cost estimation, and schema queries. |
| mcp-server | Listed | 0 | Domain-verified | stdio | Web scraping MCP server — scrape, extract structured data, screenshot any site with anti-bot bypass. |
| Amazon Scraper API | Listed | 1 | GitHub | stdio | Scrape Amazon products, search, and async batch ASIN lookups across 20 marketplaces |
| mcp-webgate | Listed | 3 | GitHub | stdio | Web search that doesn't wreck your AI's memory. |
| Kai AGI - Autonomous AI Agent | Listed | — | Domain-verified | remote | AI predictions, model comparison, research briefs and web search from an autonomous AI |
| anyapi | Listed | 0 | GitHub | remote | Hundreds of scraping & data APIs through one key. USD pay-per-request, normalized schemas, failover. |
| apexapi-mcp | Listed | 0 | GitHub | stdio | Call 120+ AI models and live web context (scrape/crawl/extract) from any MCP client with one key |
| Apollo Intelligence | Listed | 2 | Domain-verified | stdio | 36 tools: intel feeds, DeFi, crypto, OSINT, NLP, scraping, proxy. x402 micropayments. |
| apple-voice-memo-mcp | Listed | 6 | GitHub | stdio | Access Apple Voice Memos on macOS. List, get audio, extract and generate transcripts. |
| Arachne MCP | Listed | — | GitHub | remote | Web scraping, browser RAG, vision, transcription, and 20 AI tools. Universal data intelligence. |
| Archiet | Listed | — | GitHub | stdio | Archiet MCP server: scaffold full-stack apps from PRD text and extract capabilities. |
| arjunkmrm-brave-search-mcp-server | Listed | — | Domain-verified | remote | Search the web, images, videos, news, and local businesses with robust filters, freshness controls… |
| arjunkmrm-fetch | Listed | — | Domain-verified | remote | Fetch web pages and extract exactly the content you need. Select elements with CSS and retrieve co… |
| arjunkmrm-perplexity-search | Listed | 10 | Domain-verified | remote | Enable AI assistants to perform web searches using Perplexity's Sonar Pro. |
| arjunkmrm-scrapermcp_el | Listed | — | Domain-verified | remote | Extract and parse web pages into clean HTML, links, or Markdown. Handle dynamic, complex, or block… |
| ATEX AI Gateway | Listed | 2 | GitHub | remote | One API key for 6 AI models. Pay-per-use. MCP protocol support with web search. |
| Audioscrape Audio Intelligence | Listed | — | Domain-verified | remote | The audio intelligence layer. Search podcast transcripts, speakers, and entities across 250K+ shows. |
| BadRooBot-test_m | Listed | 0 | Domain-verified | remote | Send quick greetings, scrape website content, and generate text or images on demand. Perform web s… |
| BAILII UK Case Law | Listed | 3 | GitHub | stdio | Search UK case law on BAILII — court judgments with section extraction. Runs locally. |
| PDF Kit | Listed | — | GitHub | stdio | AI-powered PDF tools: fill forms, merge, extract data, and split PDFs |
| biolit | Listed | 2 | GitHub | stdio | LLM-assisted biomedical literature screening and extraction for PubMed, GEO, and preprints. |
| blacklotusdev8-test_m | Listed | 0 | Domain-verified | remote | Greet anyone by name with a friendly hello. Scrape webpages to extract content for quick reference… |
| Listed | 1 | GitHub | stdio | MCP server for Blip disposable email — create inboxes, receive emails, extract OTP codes | |
| bluesky-mentions-scraper | Listed | — | GitHub | remote | Track brand mentions & keywords on Bluesky. Sentiment, engagement, author reach. Pay per result. |
| bluesky-profile-scraper | Listed | — | GitHub | remote | Bulk Bluesky profiles plus full follower/following exports via the open AT Protocol. Pay per record. |
| bluesky-scraper | Listed | — | GitHub | remote | Scrape Bluesky posts, profiles, followers, threads and keyword search. Clean JSON, pay per result. |
| bol-ai | Listed | — | GitHub | remote | Extract structured data from Bills of Lading: parties, ports, containers, incoterms. EU-hosted. |
| Web Scraper to Markdown API | Listed | 0 | GitHub | remote | Extract clean markdown from any URL. Removes boilerplate. For RAG pipelines. x402. |
| Web Search API | Listed | 0 | GitHub | remote | Web search returning structured results — titles, URLs, snippets. x402 micropayment. |
| Brandcode MCP | Listed | 9 | GitHub | stdio | Make your brand machine-readable. Extract identity from any website into tokens and policies. |
| browser-use | Listed | 0 | GitHub | stdio | AI browser automation - navigate, click, type, extract content, and run autonomous web tasks |
| IRONCLAW BTC Node | Listed | 0 | GitHub | remote | Real BTC full node: fees, mempool, txs, portfolio, trace, whales, SEC, scraping, Reddit via x402. |
| caesar | Listed | — | Domain-verified | remote | Web search and page-reading for AI agents. One-click OAuth connect, or a Caesar API key. |
| Caliper | Listed | — | Domain-verified | remote | Geometry and CAD file metadata extraction for STL, OBJ, PLY, PCD, LAS/LAZ, glTF/GLB. |
| CalmSEO | Listed | — | Domain-verified | remote | SEO MCP server for keyword research, SERP analysis, audits, and Search Console workflows. |
| camoufox-mcp | Listed | 7 | GitHub | stdio | Anti-detection browser automation with Camoufox - stealth Firefox for web scraping |
| CatchAll | Listed | 1 | Domain-verified | remote | Web search API: find every relevant event across the open web, not just the top results. |
| ccbot | Listed | — | Domain-verified | remote | Validate CommonCrawl CCBot IP addresses. Remote MCP validate_ip tool. |
| ClawPage | Listed | 0 | GitHub | stdio | Extract and structure any web page into clean JSON. Free tier: 10/day. |
| ClicheFactory Document Intelligence | Listed | 0 | GitHub | stdio | Extract structured JSON from PDFs, images, DOCX, XLSX, CSV, EML attachments, and DSPy pipelines. |
| google-maps-mcp-server | Listed | — | Domain-verified | remote | Local business lead extraction with email + phone enrichment from Google Maps. |
| seo-web-analysis-mcp-server | Listed | — | Domain-verified | remote | Site crawl + tech stack + DNS + SSL + WHOIS — five web-intel layers in one MCP. |
| Content Intelligence API | Listed | 0 | GitHub | remote | 9 MCP tools: extract, analyze, research, compare, monitor, brief. Pay-per-call x402 or subscribe. |
| content-optimizer | Listed | 0 | GitHub | stdio | SERP-based content scoring and optimization with 7 SEO categories. |
| contractoracle | Listed | — | Domain-verified | remote | ContractOracle - 10 contract analysis tools: clause extraction, redlines, DORA mappings. |
| CoreWise | Listed | — | GitHub | remote | Extract structured insights from videos, podcasts, articles, and PDFs with multi-model AI |
| Crawlbase | Listed | 0 | GitHub | remote | Crawlbase MCP — wraps the Crawlbase Crawling API (crawlbase.com, formerly |
| Crawlberg | Listed | 141 | GitHub | stdio | Scrape, crawl, and map websites to Markdown or JSON via local CLI. |
| CrawlConsole | Listed | — | Domain-verified | stdio, remote | Backlink Analysis MCP for AI Agents: authority, referring domains, and competitor link gaps. |
| Crawleo-MCP | Listed | 11 | GitHub | remote | Hosted Crawleo MCP (remote streamable HTTP endpoint). |
| crawlforge-mcp-server | Listed | 1 | GitHub | stdio | 26-tool MCP server for web scraping, crawling, deep research & autonomous extraction |
| crawlgraph-mcp | Listed | 1 | GitHub | stdio | Backlink lookups and competitor gap analysis on Common Crawl's open web graph (4.4B links). |
| Crawlie | Listed | 91 | Domain-verified | remote | Technical SEO + GEO (AI-search) site audits: hosted crawls, prioritized fixes, report diffs. |
| crawlinx-mcp | Listed | 0 | GitHub | remote | Free technical-SEO audit MCP: crawl a site, run checks, return an LLM-ready shareable report. |
| Crawlora MCP | Listed | 1 | Domain-verified | remote | Hosted MCP: 319 structured web-data tools for search, maps, commerce, social & finance. |
| customjs-mcp | Listed | 0 | Domain-verified | remote | Host static HTML pages, generate PDFs, screenshots, scrape JS sites, run sandboxed JavaScript. |
| DACIX — The Store for AI Agents | Listed | 0 | Domain-verified | remote | AI-agent marketplace: agent templates, web search & crawl, SEO audit, RO company data, city info. |
| DarkWeb-Breach | Listed | — | GitHub | remote | OSINT dark web database search tool detecting credential leaks. |
Showing 100 of 419 servers in this category.