PAVEDB

RETRIEVAL YOU CAN OWN

The inspectable retrieval database.

Small enough to embed. Complete enough to operate.

Start inside Python. Serve the same engine over HTTP. Self-host it or let Flowlexi operate it. Open-source semantic search with sources, query history and replay.

Quick start

Python · embedded

pip install pavedb

from pavesdk.client import connect

db = connect("./data")  # in-process; no server, no config
books = db.create_collection("books")
books.add("Captain Nemo commands the Nautilus.", docid="note-1")
hits = books.search("submarine captain", k=3)

Search locally, with no server to run.

When you deploy a core, the same connect() takes a URL.

Start building →

Semantic search and RAG, from your own documents

From Python to a service. Same retrieval engine.

Start in a local directory, then serve it over HTTP. Choose who operates the same open-source engine.

  1. Prototype connect() Ephemeral store in a temp directory; nothing left behind.
  2. Persist connect("./data") Point at a directory. Same code, now durable.
  3. Serve pavesrv --data-dir ./data The same directory over HTTP — staging speaks the production wire protocol.
  4. Deploy docker run … pavedb:latest-cpu The prebuilt image on your server; clients switch with one line.
  5. Migrate pavecli dump-archive Snapshot every tenant and collection, restore it into the remote instance.

Managed hosting · a FLOWLEXI. service

Embedded or managed. Same database.

Cloud changes who operates PaveDB, not your retrieval stack. Keep the same engine and client API while Flowlexi handles provisioning, updates and monitoring. PaveDB is also the evidence layer in Flowlexi review workflows.

Collection export/import gives your data an exit path. Move a collection between compatible PaveDB instances, including one you run yourself.

Explore PaveDB Cloud →

Inspect a result. Replay a query. See what changed.

Trace each match to its source and inspect the logged query, parameters and results. After updating your collection, replay the query against its current state and compare with the original record to understand retrieval drift.

Python
from pavesdk.client import connect

db = connect("./data")
books = db.create_collection("books")
books.add("Captain Nemo commands the Nautilus.", docid="note-1")

# Search, then inspect the stored request and replay it.
hits = books.search("submarine captain", k=3)
query = books.queries(limit=1)[0]
record = books.get_query(query["query_id"])
again = books.replay(query["query_id"])
Elixir
client =
  PaveDBClient.connect("http://localhost:8086",
    tenant: "tenant",
    api_key: "secret"
  )

{:ok, books} =
  PaveDBClient.create_collection(client, "books", display_name: "Books")

{:ok, _doc} =
  PaveDBClient.Collection.add(books, "Captain Nemo commands the Nautilus.",
    docid: "note-1",
    metadata: %{"kind" => "note"}
  )

{:ok, response} =
  PaveDBClient.Collection.search(books, "captain", k: 3)

matches = response["matches"]

Connect with Python, Elixir, JavaScript or HTTP

Pick a client, then point it at a PaveDB core you run — installed directly or in Docker.

Choose your path →

Why PaveDB

Document ingestion included

Hand PaveDB a PDF, CSV, or TXT file; it chunks, embeds, and indexes the content without a separate preprocessing pipeline.

Full provenance on text-backed hits

Trace a match to its document, page or character offset, and exact snippet. Vector-only records may not carry text.

Query history and replay

Keep a record of text queries, parameters, timing and ranked results. Replay by query ID and compare records as your collection changes.

Tenant boundaries and limits

Configure tenant-scoped API keys, concurrency and storage limits. Keep workloads separate without building these controls yourself.

Controls for operating a service

Health and readiness endpoints, Prometheus metrics, request limits and archives come with the engine.

Your choice of embeddings

Choose a local or hosted embedding model per collection, or bring your own vectors. Text-backed records keep their sources available for inspection.

Where PaveDB fits

Retrieval inside your application

Document search, internal knowledge tools and domain-specific assistants that need source evidence and a self-contained deployment. Start embedded and run the same engine as a service when needed.

A deliberate single-process design

One process owns one data directory. Size an instance for its workload or run separate instances for separate workloads. If you need a distributed index or multi-node high availability today, choose a database designed for that.

The book · public draft coming soon

Inspectable Retrieval with PaveDB

Building RAG and semantic search systems you can deploy, operate, observe, and trust.

From embedded Python to an operated service: inspect sources, compare query replays and keep retrieval under your control. The full draft will be available by email.

About the book →
Ready to build? Start building →