Design, Media & Creative MCP servers connect AI assistants to design platforms, video platforms, and audio tools. These servers let models inspect UI specifications in tools like Zeplin, extract subtitles via YouTube Video Summarizer MCP and yt-dlp-mcp, manage multimedia records, and generate spoken audio through voice synthesis servers.
Design, Media & Creative MCP servers let AI assistants create, inspect, and manipulate visual, audio, and video assets directly from chat or coding environments. Instead of relying on manual file uploads or static descriptions, these Model Context Protocol integrations give assistants programmatic access to video streams, subtitle tracks, design specifications, media metadata, and voice synthesis engines.
When evaluating servers in this category, focus on three technical criteria:
Top picks include Zeplin, which exposes interface specifications and design assets directly to coding assistants; yt-dlp-mcp, which downloads video and audio streams across web platforms; YouTube Video Summarizer MCP, which pulls video context for quick conceptual synthesis; and Zundamon Voice Synthesis, which provides speech generation capabilities to assign distinct spoken output to AI workflows.
| MCP server | What it connects to | Type | Repository |
|---|---|---|---|
| YouTube Video Summarizer MCP | The YouTube Video Summarizer MCP acts like a smart bridge between video content and AI assistants, making it easy to | Open Source | |
| yt-dlp-mcp | The yt-dlp-mcp tool serves as a bridge that allows AI assistants to fetch video and audio content directly from the | Open Source | |
| YouTube Video Summarizer | The YouTube Video Summarizer MCP acts as a bridge between video content and artificial intelligence, allowing models like Claude to | Open Source | |
| Zeplin | The Zeplin MCP server acts as a professional bridge between design files and coding environments, allowing AI coding assistants to | Unknown | |
| YouTube Translate MCP | The YouTube Translate MCP is a specialized tool that allows AI assistants to "read" and understand YouTube videos by accessing | Unknown | |
| Voice Mcp | MCP Server | Unknown | |
| yt-dlp | The yt-dlp-mcp tool acts as a bridge between AI assistants and the vast world of online video content. It allows | Open Source | |
| MCP YouTube Transcript Server | Retrieves transcripts from YouTube videos for content analysis and processing. | Unknown | |
| YTTranscipterMultilingualMCP | YTTranscipterMultilingualMCP serves as a digital ear for AI assistants, allowing them to "listen" to and read the contents of YouTube | Open Source | |
| Zundamon Voice Synthesis | The Zundamon Voice Synthesis MCP server is a specialized tool designed to give AI-driven applications a unique and recognizable personality. | Open Source |
The YouTube Video Summarizer MCP and Zeplin are strong choices for Claude Desktop. Zeplin connects assistants directly to UI components and design specs, while YouTube Video Summarizer MCP provides video transcripts and summaries without leaving the conversational interface.
Add the server configuration to your Cursor MCP settings file using the stdio command. For example, configure Zeplin to supply UI design specs to your coding agent, or register yt-dlp-mcp to allow your assistant to fetch media assets during development.
MCP YouTube Transcript Server, YTTranscipterMultilingualMCP, and YouTube Translate MCP specialize in fetching YouTube transcripts. They retrieve multilingual dialogue tracks and pass raw text or timestamps directly into the model context for translation or analysis.
Yes, Zundamon Voice Synthesis and Voice Mcp offer audio capabilities. Zundamon Voice Synthesis connects AI agents to speech synthesis systems, letting assistants generate spoken voice files alongside standard text responses.
Yes, Movies MCP Server provides structured database tools for creative projects. It runs against PostgreSQL to deliver advanced search, CRUD operations, and image asset metadata management for movie and media catalogs.