Qdrant Retrieve is an MCP server that enables AI assistants to execute semantic searches across vector collections stored in a Qdrant database instance. Developed for machine learning engineers, knowledge managers, and developers building retrieval-augmented generation pipelines, this server connects clients like Claude Desktop to local or remote Qdrant deployments. By integrating local embedding generation through models such as Xenova/all-MiniLM-L6-v2, it automatically converts search queries into vectors and retrieves relevant document chunks without requiring external embedding APIs. The server exposes a dedicated retrieval tool that accepts multiple text queries and searches across multiple collection names simultaneously, returning matching documents accompanied by similarity scores and collection identifiers. This functionality enables language models to locate domain knowledge, ground answers in enterprise reference materials, and retrieve contextual information dynamically during user conversations. It supports both standard input/output and HTTP transports, allowing flexible integration into diverse agentic workflows, local developer environments, and production retrieval architectures.
Category: AI Memory & Context
Tags: embeddings, qdrant, rag, semantic search, vector-database
claude_desktop_config.json). 2. Add the server under the mcpServers object using npx: json { "mcpServers": { "qdrant": { "command": "npx", "args": ["-y", "@gergelyszerovay/mcp-server-qdrant-retrive"], "env": { "QDRANT_API_KEY": "your_api_key_here" } } } } 3. If connecting to a remote or non-default Qdrant instance, append --qdrantUrl=<url> to args (default is http://localhost:6333). 4. Save the configuration file and restart Claude Desktop. The initial search may take longer while the default embedding model downloads.Part of MCP Servers
Qdrant Retrieve is an open-source MCP server that provides semantic search capabilities across Qdrant vector database collections. It allows compatible AI clients to run vector similarity searches using local embedding models like Xenova/all-MiniLM-L6-v2, returning matched document text, source collection names, and similarity scores.
You can configure Qdrant Retrieve in Claude Desktop by editing your claude_desktop_config.json file. Add an entry under mcpServers running npx with the package @gergelyszerovay/mcp-server-qdrant-retrive. Optionally include your QDRANT_API_KEY in the env block and specify your database location using the --qdrantUrl argument.
The qdrant_retrieve tool accepts three primary parameters: collectionNames, which takes an array of Qdrant collection names to query; query, which takes an array of search strings; and topK, an optional number specifying how many top similar documents to return per query, defaulting to three.
No, Qdrant Retrieve downloads and runs a local embedding model, defaulting to Xenova/all-MiniLM-L6-v2. Because the embeddings are generated locally by the server, an external API key for embedding providers is not required, though the initial search may take slightly longer while the model files download.
Yes, Qdrant Retrieve is an open-source project hosted on GitHub under the repository gergelyszerovay/mcp-server-qdrant-retrive. It can be executed directly using npx or cloned and customized locally to adjust embedding models, transport protocols, and database connection settings.