Search & Data Extraction
186 MCP servers and agent skills in the Search & Data Extraction category, ranked by quality score — 186 results.
joelio/stocky
An MCP server for searching and downloading royalty-free stock photography from Pexels and Unsplash. Features multi-provider search, rich metadata, pagination support, and async performance for AI assistants to find and access high-quality images.
cameronrye/gopher-mcp
Modern, cross-platform MCP server enabling AI assistants to browse and interact with both Gopher protocol and Gemini protocol resources safely and efficiently. Features dual protocol support, TLS security, and structured content extraction.
OctagonAI/octagon-deep-research-mcp
Lightning-Fast, High-Accuracy Deep Research Agent
jhomen368/overseerr-mcp
Integrate AI assistants with Overseerr and the Seerr (the unified successor) for automated media discovery, requests, and management in Plex, Jellyfin, and Emby ecosystems.
scavio-ai/scavio-mcp
Unified real-time search API for AI agents. Google, YouTube, Amazon, Walmart, Reddit, and TikTok through one endpoint. 21 tools for web search, e-commerce, product data, social media, and video platforms. Free tier included.
qune-tech/ocds-mcp
German public procurement data (OCDS) — semantic search, tender matching with company profiles, and structured filtering.
scraperapi/scraperapi-mcp
MCP server for ScraperAPI web scraping with JavaScript rendering, geotargeting, premium proxies, and auto-parsing support.
talonicdev/talonic-mcp
Schema-validated document extraction with searchable workspace memory. Extract structured fields from PDFs, scans, images, and forms; AI agents can also search, filter, and query past extractions.
keenableai/keenable-mcp
Live web search and clean-markdown page fetch over the Keenable web index. Two tools: `search_web_pages`, `fetch_page_content`. Keyless by default (1,000 req/hour); an optional API key lifts the cap. Hosted Streamable HTTP at `https://api.keenable.ai/mcp`, or run `npx -y @keenable/mcp`.
pepabo/muumuu-domain-mcp
Official remote MCP server for Muumuu Domain (GMO Pepabo). Search and register domains, manage owned domains and contracts, and configure DNS records via natural language.
mrslbt/rippr
YouTube transcript extraction for AI agents. Clean text, timestamps, or structured JSON from any video. No API keys required. Install via `npx rippr-mcp`.
lionkiii/rss-feeds-mcp
RSS feeds MCP server with 8 tools — fetch, filter, search, and manage RSS feeds by category or source. Zero config, no API keys required.
zlatkoc/youtube-summarize
MCP server that fetches YouTube video transcripts and optionally summarizes them. Supports multiple transcript formats (text, JSON, SRT, WebVTT), multi-language retrieval, and flexible YouTube URL parsing.
AIMLPM/markcrawl
Crawl websites into clean Markdown, search pages, and extract structured data with LLMs. Built-in MCP server for web research and RAG pipelines.
Khamel83/argus
Multi-provider search broker with automatic fallback, RRF ranking, content extraction, and budget enforcement.
comparedge/mcp-server-comparedge
Verified SaaS, AI, and LLM pricing for 490+ tools: plans, hidden costs, alternatives, and comparisons. Free, no API key.
ashlrai/webfetch
License-first federated image search across 25 providers. Returns open/platform/editorial license tags, attribution strings, dimensions, and download-ready URLs via `npx -y getwebfetch-mcp`.
MKirovBG/scribefy-mcp
Extract timestamped YouTube transcripts, plus search, video metadata, and related-video tools for Claude, Cursor, Windsurf, and AI agents.
sifter-ai/sifter
Structure any document, query it like a database. Open-source extraction engine that turns any document into typed, schema-defined records, queryable in natural language from Claude, ChatGPT, Gemini, or any MCP client.
wd041216-bit/free-web-search-ultimate
Zero-cost, privacy-first universal web search MCP server. Enforces a **Search-First** paradigm — instructs LLMs to retrieve real-time information before answering factual questions. Supports 10+ search engines (DuckDuckGo, Bing, Google, Brave, Wikipedia, Arxiv, YouTube, Reddit) and deep page browsing. No API key required.
whw23/searxng-http-mcp
Self-contained SearXNG MCP server in Docker. 200+ search engines, 30+ categories, multi-page fanout, autocomplete, and engine discovery. Dual transport (HTTP + stdio), API key auth, and built-in SearXNG Web UI reverse proxy. Zero-install deploy.
linxule/mineru-mcp
MCP server for MinerU document parsing API. Parse PDFs, images, DOCX, and PPTX with OCR (109 languages), batch processing (200 docs), page ranges, and local file upload. 73% token reduction with structured output.
paulieb89/govuk-mcp
Search GOV.UK content, retrieve full government pages, look up organisations, and resolve UK postcodes to local authorities. 5 read-only tools, no API keys required.
serkan-ozal/driflyte-mcp-server
The Driflyte MCP Server exposes tools that allow AI assistants to query and retrieve topic-specific knowledge from recursively crawled and indexed web pages.
vectorize-io/vectorize-mcp-server
[Vectorize](https://vectorize.io) MCP server for advanced retrieval, Private Deep Research, Anything-to-Markdown file extraction and text chunking.
echology-io/decompose
Decompose text into classified semantic units with authority, risk, attention scores, and entity extraction. No LLM. Deterministic. Works as MCP server or CLI.
searchcraft-inc/searchcraft-mcp-server
Official MCP server for managing Searchcraft clusters, creating a search index, generating an index dynamically given a data file and for easily importing data into a search index given a feed or local json file.
AutomateLab-tech/citation-intelligence
What LLMs cite, for agents. Check which URLs Perplexity, Claude, ChatGPT, Gemini, Bing, and Google AI Overviews cite for any query. Self-hosted, BYO API key. Install via `npx @automatelab/citation-intelligence`.
pgalyen1987/gate402-mcp
Pay-per-call agent APIs over x402 (USDC on Base): per-token LLM inference (Llama 3.1/Qwen/Mistral), per-second GPU/CPU compute, on-chain/DeFi + SEC-EDGAR + news data, and clean-Markdown/Cloudflare-stealth web scraping. Signed receipts, no signup — auto-claims a free-tier key. Install via `npx -y gate402-mcp`.
AceDataCloud/MCPSerp
Google SERP search including web, images, news, maps, places, videos, and knowledge graph results via Ace Data Cloud API.
capad-xyz/searchts
Keyless web access for AI agents: an escalating open-source unlocker (browser-fingerprint fetch → JS-render relay → stealth browser) reads bot-walled pages as clean Markdown, plus multi-provider web search with rank fusion, subtitles-first video transcripts, and page asset grabbing. No API keys; ships a reproducible benchmark.
Crawlora-org/crawlora-mcp
Hosted MCP for structured public web data — 319 tools across search, maps, commerce, social, and finance, each returning clean JSON. Free 2,000 credits/mo.
divyanshu-iitian/SearchForge
Free capability-routed search and web reading for agents: GitHub, Crossref, Hacker News, Wikipedia, private SearXNG, and URL-to-Markdown, with live health diagnostics and no telemetry. Exposes `web_search`, `read_url`, and `search_status`.
goofrey/zoom-search
MCP search and evidence tool for AI agents. Rewrites queries, zooms into source domains, and returns sourced answers with metrics.
hanoak/unsplash-mcp-server
Unsplash API server exposing 21 tools across photos, search, users, collections, topics, and stats. Ships Unsplash-guideline compliance built in — ready-to-use attribution with UTM parameters, a download-tracking tool, rate-limit surfacing, and `content_filter=high` by default. Install via `npx -y @hanoak/unsplash-mcp-server`.
Opedd/opedd-mcp
Licensed, rights-cleared content for AI agents — discover, purchase, verify, and retrieve expert analysis with a verifiable license key per article, on-chain proof, and EU AI Act Article 53 attestation. The alternative to unlicensed scraping for RAG and AI search. `npx opedd-mcp`
chefcohen/corroborate-mcp
Tells AI agents how independently a claim is being reported: syndication-aware source counting (a wire story reprinted by 40 outlets counts as one origin), 0-1 confidence scores, and published error rates in-repo. Measures corroboration, not truth — no stance detection. Keyless, no accounts. Install: `npx -y corroborate-mcp`.
kimdonghwi94/Web-Analyzer-MCP
Extracts clean web content for RAG and provides Q&A about web pages.
pranciskus/newsmcp
Real-time world news for AI agents — events clustered from hundreds of sources, classified by topic and geography, ranked by importance. Free, no API key. `npx -y @newsmcp/server`
robbyczgw-cla/web-search-plus-mcp
Multi-provider web search with intelligent auto-routing (Serper, Tavily, Exa). Available via `uvx web-search-plus-mcp`.
VoxellInc/forge-mcp
Official MCP server for [Forge](https://voxell.ai), Voxell's hosted text-embedding API. Generate vector embeddings (turbo 1024d, pro 2560d, ultra 4096d; Matryoshka truncation) for semantic search and RAG. `npx -y @voxell/forge-mcp`
andybrandt/mcp-simple-pubmed
MCP to search and read medical / life sciences papers from PubMed.
imprvhub/mcp-domain-availability
A Model Context Protocol (MCP) server that enables Claude Desktop to check domain availability across 50+ TLDs. Features DNS/WHOIS verification, bulk checking, and smart suggestions. Zero-clone installation via uvx.
boikot-xyz/boikot
Model Context Protocol Server for looking up company ethics information. Learn about the ethical and unethical actions of major companies.