Lakebase Vector Revolutionizes Postgres Search with Scalable Hybrid Vector-BM25 Integration
September 28, 2026
Lakebase Vector introduces scalable vector search built on hierarchical IVF clustering and binary quantization (RaBitQ) that offloads index construction to distributed engines like Spark, separating storage from compute to enable fast queries with caching and object storage.
The architecture separates storage from compute, uses RaBitQ to shrink vectors and prune candidates, and is stateless with zero-replication scaling, charging for storage at rest.
Lakebase Search brings semantic, keyword, and hybrid search directly into Postgres alongside operational data, delivering low-latency, high-accuracy retrieval without external ETL pipelines.
A customer example from Conexiom shows BM25 hybrid search over more than 100 million rows with half the compute footprint versus a prior pgvector setup, delivering a threefold cost reduction and a fivefold throughput increase.
Overall, Conexiom achieved about 3x lower infrastructure costs and 5x higher throughput with Lakebase Vector by enabling hybrid search and unifying OLTP and search workloads.
Databricks identified three major pain points with pgvector: high scaling costs due to in-memory indexing, long index build times and write bottlenecks, and lack of parallelism for single queries.
Lakebase Vector reports strong performance, including 71 ms P99 latency at 97% recall on 100 million vectors, with higher throughput and lower cost compared with pgvector-based deployments.
In VectorDBBench’s 100M benchmark (LAION dataset), Lakebase Vector delivers roughly twice the throughput of the next-best system and about four times cheaper than a cloud Postgres vendor using pgvector, noting pgvector and DiskANN were tested on a single large instance.
Hybrid queries combine semantic vector scoring with BM25 relevance, enabling filtering, joins, and live data access within a single SQL query, scalable from one to billions of vectors and from one to thousands of queries per second.
Hybrid search enables combining vector and BM25 scoring with standard SQL predicates, letting queries run against live operational tables without multi-system pipelines.
BM25 text search (lakebase_text) improves over tsvector by using global IDF and upper-bound pruning for faster top-K results, enabling efficient hybrid queries.
lakebase_text adds native BM25 search to Postgres for corpus-wide relevance scoring, improves speed over tsvector with GIN indexes, and supports seamless hybrid search by fusing semantic vector search with BM25 in a single query.
Summary based on 2 sources
Get a daily email with more AI stories
Sources

Unite.AI • Sep 28, 2026
Databricks Brings Full-Text and Vector Search to Lakebase Postgres
Databricks • Sep 28, 2026
Lakebase Search: State-of-the-art full text and vector search for Postgres