Search: multimodal
Search MCP servers and agent skills by name, description, category or topic — 8 results.
stabgan/openrouter-mcp-multimodal
All-in-one multimodal MCP for 300+ OpenRouter models: text chat, image / audio / video analysis, and image / audio / video generation (Veo 3.1, Sora 2 Pro, Seedance, Wan). Structured `_meta.code` error taxonomy, IPv4+IPv6 SSRF guards, path-sandbox for disk writes, retry-after-aware backoff, multi-arch Docker.
video-db/skills
Realtime and batch video workflows: capture screen/audio, ingest URLs/YouTube/RTSP, transcribe, index, search, generate subtitles, edit timelines, and stream HLS output
Zacccck/Claude-MCP-Read-Email-Attachments
Remote HTTP MCP server that reads Outlook email attachments via Microsoft Graph. Parses PDF, Word (with embedded image extraction for multimodal analysis), Excel, and text files in-memory and returns structured content directly to Claude.
tan-yong-sheng/ai-vision-mcp
Multimodal AI vision MCP server for image, video, and object detection analysis. Enables UI/UX evaluation, visual regression testing, and interface understanding using Google Gemini and Vertex AI.
jj-cheng25/weixin-articles-mcp
Read WeChat (微信) Official Account articles with native multimodal output — body, images, and video keyframes as MCP content blocks. Handles all three embed types: Tencent Video (yt-dlp keyframes), WeChat-native (mp4 keyframes), Channels/视频号 (metadata + cover via public API).
unixlamadev-spec/lightningprox-mcp
MCP server for LightningProx — pay-per-request AI access via Bitcoin Lightning. Supports vision/multimodal.