Vector Database
A vector database stores embeddings alongside their source data and metadata and quickly retrieves the vectors most similar to a query vector, using approximate nearest-neighbour indexes. It is a core component of retrieval-augmented generation.
Exact nearest-neighbour search over millions of vectors is slow, so vector databases use approximate nearest-neighbour (ANN) indexes such as HNSW graphs, inverted file indexes and product quantisation, trading a small loss of recall for large gains in speed. Queries usually combine vector similarity with metadata filters, and many systems support hybrid search that blends keyword ranking, such as BM25, with semantic similarity.
Vector search underpins RAG chatbots that answer from manuals, procedures and tickets, semantic search across engineering documents, recommendation, image similarity search and deduplication. Options include dedicated vector databases and vector extensions to existing databases and search engines, such as pgvector for PostgreSQL.
Design choices include the embedding model, chunk size, index type and parameters, filtering strategy and how updates and deletions are handled. Retrieval quality should be measured with recall and relevance metrics on real queries. Access control must be enforced at retrieval time so that users cannot retrieve documents they are not permitted to see.
Key points
- Stores embeddings and finds the vectors most similar to a query
- Uses approximate nearest-neighbour indexes such as HNSW
- Hybrid search combines keyword and semantic matching
- Document access control must be enforced at retrieval time
Where AiVibe comes in
AiVibe delivers AI and machine learning services, chatbots and virtual assistants with RAG, MCP tools and voice, AI quality management including bias detection and model validation, and the AIMURUGA AI agent, and builds Intel-based edge AI devices using the Intel Distribution of OpenVINO toolkit.