pgvector-filterbench: Making Vector Search Measurable by Niko Minadzepgvector-filterbench: Making Vector Search Measurable by Niko Minadze

pgvector-filterbench: Making Vector Search Measurable

Niko Minadze

Niko Minadze

pgvector-filterbench: Making Vector Search Measurable

A vector search can return plausible results while silently missing neighbors that exact search would return. Add a selective filter, and an approximate search may return too few rows. Even an excellent recall score can mislead if PostgreSQL answered with a sequential scan instead of the HNSW index being tested.
I built pgvector-filterbench to make those failure modes measurable. It is an open-source Go command-line tool for comparing filtered approximate retrieval with exact PostgreSQL results.

What I built

The benchmark uses seeded database samples or supplied query vectors, then runs the same workload across bounded HNSW search, strict iterative scanning and relaxed iterative scanning. Configurable sweeps cover search depth and scan limits.
Every measured configuration records recall@K, latency and returned-row counts. The tool also checks the actual EXPLAIN plan against the configured index, table and schema. This connects the quality numbers to the execution path that produced them.
Conceptual editorial artwork representing exact and approximate search comparison, filtered results and query-plan verification.
Conceptual editorial artwork representing exact and approximate search comparison, filtered results and query-plan verification.

Engineering the checks

A low recall score alone does not explain whether a query returned a full set of weak matches or simply ran out of candidates. I added explicit short-result tracking and an optional completeness gate so those cases remain distinguishable.
For repeatable comparisons, the benchmark supports warmup runs and counterbalanced exact/approximate query ordering. Strict configuration validation catches malformed settings before a database connection is opened. Queries run in read-only, timeout-bounded transactions.
I also built privacy controls into the reporting: raw vectors and row identifiers are omitted by default, connection strings are redacted, and database identifiers can be replaced with a neutral label. Reports are available as JSON and Markdown for automated checks and human review.

Result

The result is a reusable benchmark and CI gate for deciding whether a retrieval configuration meets its own recall, completeness and query-plan requirements. Teams can compare settings, investigate regressions and retain the evidence behind an index-tuning decision.
The repository documents a completed Docker benchmark matrix for PostgreSQL 16, 17 and 18. Its seeded demo makes the workflow reproducible; measurements from a team's own dataset remain the basis for production decisions.
Like this project

Posted Sep 27, 2026

A Go toolkit that checks filtered pgvector recall, latency and HNSW query plans against exact PostgreSQL results, with repeatable reports and CI gates.