Skip to content

MAX_CONTENT_STEMS_SIZE truncates FTS on SQLite; make backend-conditional or configurable #1657

Description

@dougvann

Summary

MAX_CONTENT_STEMS_SIZE = 6000 in src/basic_memory/services/search_service.py is applied on every
backend, but its own comment states the reason is Postgres-specific:

# Maximum size for content_stems field to stay under Postgres's 8KB index row limit.
# We use 6000 characters to leave headroom for other indexed columns and overhead.
MAX_CONTENT_STEMS_SIZE = 6000

SQLite FTS5 has no equivalent per-row index size limit, so on the SQLite backends this silently
truncates the searchable text of every file larger than the cap.

Why it is easy to miss

The failure mode is a false negative, not an error. Search keeps working and keeps returning
results; what disappears is the ability to match anything that appears only past the 6000th
character of a note. A query for such a token returns an empty set, which is indistinguishable
from the token genuinely not existing.

That matters because an empty search result is normally read as evidence of absence.

Measurements

From a 411-file corpus of long-form notes (v0.23.2, SQLite, FTS5, no embeddings):

  • 70 search_index entity rows have length(content_stems) exactly 6000 — i.e. truncated:
    SELECT COUNT(*) FROM search_index WHERE type='entity' AND length(content_stems)=6000;
  • 67 source files exceed 6000 bytes.
  • Largest file is 246 KB, so roughly 2.4% of it is searchable. The next is 119 KB (~5%).
  • Concretely: tokens that exist only in the tail of a large note (an env var name, an identifier)
    return nothing, while the note itself is clearly in the index and returned by other queries.

Suggested fix

Make the cap backend-conditional rather than removing it — Postgres is a supported backend here
(DatabaseBackend.POSTGRES, _create_postgres_engine), so the 6000 limit is load-bearing there and
should stay. Something like:

MAX_CONTENT_STEMS_SIZE = 6000           # Postgres 8KB index row limit
MAX_CONTENT_STEMS_SIZE_SQLITE = 10_000_000

def _max_content_stems_size() -> int:
    try:
        from basic_memory.config import ConfigManager, DatabaseBackend
        if ConfigManager().config.database_backend == DatabaseBackend.POSTGRES:
            return MAX_CONTENT_STEMS_SIZE
    except Exception:
        return MAX_CONTENT_STEMS_SIZE
    return MAX_CONTENT_STEMS_SIZE_SQLITE

applied at the two truncation sites (entity stems and observation stems). Making it a config value
would work equally well and would be more flexible.

Happy to open a PR if the approach looks right. I am running this as a one-line fork locally and
would rather it live upstream.

Environment

  • basic-memory v0.23.2, installed via uvx
  • SQLite backend, FTS5, no vector/semantic search enabled
  • macOS, Python 3.13

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions