Lakebase Vector Revolutionizes Postgres Search with Scalable Hybrid Vector-BM25 Integration

September 28, 2026
Lakebase Vector Revolutionizes Postgres Search with Scalable Hybrid Vector-BM25 Integration
  • Databricks unveils Lakebase Search, a built-in search engine for Lakebase Postgres, delivered as two extensions: lakebasevector for approximate nearest-neighbor search and lakebasetext for BM25 full-text search, generally available on AWS and Azure.

  • Lakebase Search includes BM25 text search (lakebase_text) that improves over traditional tsvector by weighting rare terms with global IDF and speeding top-K results through upper-bound pruning, enabling efficient hybrid queries.

  • Databricks positions Lakebase AI Search as a managed alternative for high-quality retrieval, while Lakebase Search aims to consolidate operational data and search within a single database.

  • Databricks highlights three vector-search pain points: high costs at scale due to in-memory indices, long index builds and write bottlenecks, and limited parallelism for single queries.

  • Benchmark claims tout twice the throughput of the next best system and four times lower cost than a cloud Postgres with pgvector, though results are company-generated without external validation.

  • The vector-search architecture can serve up to 100 million vectors per Lakebase Compute Unit by caching hot data and decoupling storage from compute, enabling scale-to-zero when idle.

  • Index construction is distributed across cores via LTAP to offload builds from the primary database to distributed engines like Spark, cutting build times to minutes.

  • Conexiom reports BM25 hybrid search over over 100 million rows with half the compute footprint of its prior pgvector setup, reducing costs threefold and tripling throughput.

  • As a concrete example, Conexiom achieved roughly 3x lower infrastructure costs and 5x higher throughput with Lakebase Vector versus pgvector by enabling hybrid search and unifying OLTP and search workloads.

  • Lakebase uses a storage-backed index approach with hierarchical inverted-file clustering and RaBitQ quantization, claiming 100 million vectors at 97% recall with about 71 ms P99 latency in benchmarks.

  • Post-GA signals to monitor include reproducible performance against tuned pgvector and DiskANN, adoption by teams consolidating search pipelines, and competitive responses from pgvector and others.

  • Lakebase_vector reports strong performance, with 71 ms P99 latency at 97% recall on 100 million vectors, claiming higher throughput and lower cost than pgvector-based setups.

Summary based on 3 sources


Get a daily email with more AI stories

More Stories