Skip to content
Stackship documentation Svenska

Vector stores (Qdrant)Users

Vector stores (Qdrant)

Run a Qdrant vector database in a resource group for embeddings and similarity search, and understand how it is secured, exposed and kept.

A vector store is a Qdrant cluster that the platform runs for you in a resource group. Applications store embeddings in it and search them by similarity — the storage behind semantic search and retrieval-augmented generation. In the portal, vector stores are under Vector Store in the sidebar.

What you get

  • One or more Qdrant servers (the replicas), each with its own persistent volume, sized by a compute plan.
  • Endpoints for Qdrant's REST API (port 6333) and gRPC API (port 6334) inside the cluster, and optionally from outside it — see Connect to a vector store.
  • An API key that clients send with every request.

You work with Qdrant itself — collections, points, searches — through Qdrant's own API and client libraries. The platform manages the servers, not what is stored in them.

The API key is the only credential

Qdrant has no users and no per-collection permissions: whoever presents the API key can read, change and delete everything in the cluster. The portal creates every vector store with API key authentication turned on. You can turn it off, but then any workload that can reach the cluster can do the same without a key, and the platform refuses to expose a cluster that has no API key outside the Kubernetes cluster.

Showing the key in the portal is recorded in the vector store's activity log. Platform roles decide who may see or rotate the key, not what a client holding it may do.

Replicas and sharding

Each replica is one Qdrant server, and the servers of a vector store form one distributed Qdrant cluster. How data is spread over them is decided per collection, when you create it with Qdrant's API:

  • shard_number — how many shards the collection is split into. When you leave it out, Qdrant uses the number of servers at the time the collection is created.
  • replication_factor — how many copies of each shard are kept. It is 1 unless you set it, which means every shard exists once: a vector store with three replicas then holds each part of a collection on one server only, and loses access to that part while that server is down.

To make a collection survive the loss of a server, create it with replication_factor of 2 or more on a vector store with at least that many replicas. The platform does not ask Kubernetes to place the servers on different nodes.

Adding replicas later does not move existing shards onto the new servers, and removing replicas stops servers without moving their shards away first. See Change the replica count.

Exposure

Inside the Kubernetes cluster, a vector store is reachable by the workloads the boundary's network rules let through. Two independent options open it to the outside:

  • HTTPS ingress — the REST API and Qdrant's dashboard over HTTPS on a hostname, with a managed certificate. The create wizard turns this on by default.
  • Load balancer — an external address serving the REST and gRPC ports directly, without TLS.

Both require API key authentication, and both can be narrowed to a list of allowed networks. See Networking and the API key.

The HTTPS endpoint is also reachable by workloads in other boundaries on the same cluster, whose network rules allow outbound HTTPS to any address — see Connect from a workload.

Data and backups

The platform takes no backups or snapshots of a vector store and cannot restore one. The data lives on the servers' volumes. To keep a copy, take collection snapshots with Qdrant's own API and store them elsewhere — see Back up a collection.

Pages