Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 9 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -7,3 +7,12 @@ VOYAGE_API_KEY=
# Optional only for a portable/custom Cursor profile. The standard macOS/Linux
# database path is detected automatically.
SESSION_RECALL_CURSOR_DB=

# Optional Qwen inference gateway; see .project-docs/processes/inference-api-setup.md.
# SESSION_RECALL_EMBED=inference-api
# SESSION_RECALL_EMBED_BASE_URL=https://inference.example/v1
# SESSION_RECALL_EMBED_DIM=4096
# SESSION_RECALL_INFERENCE_MAX_TOKENS=8192
# SESSION_RECALL_EMBED_REVISION=deployment-revision
# SESSION_RECALL_DB_PATH=/path/to/separate/index-qwen.db
INFERENCE_API_KEY=
33 changes: 33 additions & 0 deletions .project-docs/agent-context.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
<!-- Synchronized project guidance. Canonical editable path: .project-docs/agent-context.md. -->

# Project guidance

Project knowledge lives under `.project-docs/`.

Before making a non-trivial project-specific claim:

1. read `.project-docs/manifest.yaml`;
2. follow `.project-docs/navigation.md`;
3. search the routed locations and read the canonical record plus material
related links;
4. check status, freshness, and sources.

Use `services/` for current systems, `processes/` for exact actions,
`decisions/` for rationale, `reactions/` for conditional first actions,
`bugs/` for known failures, and `timeline/` for chronology.

Do not treat observations, stale claims, conflicts, recommendations, or
unknowns as confirmed facts. Never store, quote, partially reproduce,
transform, or echo secret values.

Change project guidance only through the `project-documentation` workflow at
the canonical editable source, `.project-docs/agent-context.md`. Do not edit
`AGENTS.md` or `CLAUDE.md` directly; both generated targets are byte-identical
to the canonical source.

After every documentation mutation, run
`python3 <skill-directory>/scripts/validate_project_documentation.py
--project-root <project-root> --tracked-documentation`. Add `--include` for a
new untracked documentation file. If canonical guidance changed, also run the
guidance synchronization script with `--diff`, `--write`, and `--check`.
Documentation work is incomplete until validation and synchronization pass.
31 changes: 31 additions & 0 deletions .project-docs/manifest.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,31 @@
schema_version: 1
entrypoints:
project: project.md
navigation: navigation.md
routes:
current_state:
- project.md
- services/
how_to:
- processes/
why:
- decisions/
incident_first_action:
- reactions/
- processes/
- services/
known_problem:
- bugs/
history:
- timeline/
- changelog/
unresolved:
- open-questions.md
- observations/
search_fields:
- id
- title
- summary
- tags
- canonical_for
- related
28 changes: 28 additions & 0 deletions .project-docs/navigation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
# Documentation navigation

## Search order

1. Classify the question using `manifest.yaml` routes.
2. Search only routed paths with task terms, identifiers, tags, and synonyms.
3. Read the best canonical match.
4. Follow only material `related` links.
5. Check status, freshness, and sources before using a claim.

Example:

```bash
rg -n -i 'trace|tracing|observability' \
.project-docs/services .project-docs/processes .project-docs/decisions
```

## Stopping rules

- No canonical match means unknown.
- Observations remain unconfirmed.
- Stale records require re-verification.
- Conflicts preserve every sourced version.
- Recommendations are not current behavior.
- Missing sources invalidate confirmed claims.

Record missing knowledge in `open-questions.md`; do not invent a project
default.
6 changes: 6 additions & 0 deletions .project-docs/open-questions.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
# Open questions

- Gateway deployments must supply their own endpoint, credentials, dimensions,
context limit, and embedding revision. These are not repository defaults.
- The gateway registry does not provide a required immutable embedding revision;
operators must update the configured revision when the backend changes.
66 changes: 66 additions & 0 deletions .project-docs/processes/inference-api-setup.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,66 @@
---
id: inference-api-setup
type: process
title: Connect a Qwen inference gateway
summary: Configure the gateway and build a separate index before switching retrieval.
status: confirmed
tags: [qwen, setup, migration, backup]
canonical_for: [inference-api-setup]
verified_at: 2026-09-14
sources:
- type: repository
reference: src/session_recall/config.py
confirmed_at: 2026-09-14
- type: repository
reference: src/session_recall/inference.py
confirmed_at: 2026-09-14
- type: repository
reference: src/session_recall/cli.py
confirmed_at: 2026-09-14
related: [../project.md, ../services/inference-api.md]
---

# Connect a Qwen inference gateway

Check the [required gateway contract](../services/inference-api.md) first.
Obtain the endpoint, embedding dimension and context limit from the deployment
owner or authenticated model registry. Supply `INFERENCE_API_KEY` through your
secret manager or environment; never commit its value.

Example configuration for a gateway with 4096-dimensional vectors and an
8192-token context window:

```sh
export SESSION_RECALL_EMBED=inference-api
export SESSION_RECALL_EMBED_BASE_URL=https://inference.example/v1
export SESSION_RECALL_EMBED_DIM=4096
export SESSION_RECALL_INFERENCE_MAX_TOKENS=8192
export SESSION_RECALL_EMBED_REVISION=deployment-revision
export SESSION_RECALL_DB_PATH="$HOME/.local/share/session-recall/index-qwen.db"
session-recall index
session-recall health
session-recall search "why did we choose"
```

The preset uses model aliases `embedder` and `reranker`. Override them with
`SESSION_RECALL_EMBED_MODEL` and `SESSION_RECALL_RERANK_MODEL` if your gateway
uses different public names. Existing provider overrides also take precedence
over the preset, so remove or update stale overrides when migrating.

Before migration, preserve the old provider settings and make a consistent
SQLite backup of the old index. Build the new embedding space in a separate
file with `SESSION_RECALL_DB_PATH`; do not mix Voyage and Qwen vectors.
Review indexing errors, corpus coverage and health before switching clients.
An unavailable source database cannot supply new history; retain its old
index until its historical records have been accounted for.

Configure CLI invocations, background indexers, and MCP launchers with the same
provider settings and database path. Restart existing MCP processes after
switching, because they retain their configuration and open database.
To roll back, restore the previous launcher configuration and use the old
index with its original embedding provider.

When the backend behind the public alias changes incompatibly, update
`SESSION_RECALL_EMBED_REVISION` and rebuild. A stable model alias alone does
not identify an immutable embedding space. Session Recall does not auto-load
`.env` files; export these settings or load them through your launcher.
25 changes: 25 additions & 0 deletions .project-docs/project.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
---
id: session-recall-project
type: project
title: Session Recall
summary: Semantic retrieval over local Claude Code, Codex, and Cursor history.
status: confirmed
tags: [recall, indexing, embeddings]
canonical_for: [project-overview]
verified_at: 2026-09-14
sources:
- type: repository
reference: README.md
confirmed_at: 2026-09-14
related: [services/inference-api.md, processes/inference-api-setup.md]
---

# Session Recall

Session Recall extracts conversation text, stores text and vectors in SQLite,
and exposes retrieval through its CLI and MCP server. Provider selection and
index identity are configured outside the repository.

See the [Inference API contract](services/inference-api.md) and
[connection procedure](processes/inference-api-setup.md) for hosted Qwen gateways.
The root README covers existing providers and the remaining product workflows.
56 changes: 56 additions & 0 deletions .project-docs/services/inference-api.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,56 @@
---
id: inference-api-provider
type: service
title: Qwen inference gateway provider
summary: Embeddings and reranking through an authenticated capability-based HTTP API.
status: confirmed
tags: [qwen, inference, embeddings, reranker, codex]
canonical_for: [inference-api-provider]
verified_at: 2026-09-14
sources:
- type: repository
reference: src/session_recall/inference.py
confirmed_at: 2026-09-14
- type: repository
reference: src/session_recall/config.py
confirmed_at: 2026-09-14
- type: repository
reference: tests/test_inference.py
confirmed_at: 2026-09-14
related: [../project.md, ../processes/inference-api-setup.md]
---

# Inference API provider

The `inference-api` preset selects the public `embedder` and `reranker` model
aliases. It defaults to 4096 embedding dimensions; deployments can override
that value. The base URL and credential have no default.

The client sends `POST embeddings` beneath the configured `/v1` base URL,
with `input_type=document` for indexing and `input_type=query` for retrieval.
It sends batches of up to 64 strings and requests base64 float32 vectors.
Response indices determine ordering. Both float arrays and base64 responses
are validated for dimension and finite values before storage.

Reranking uses `POST rerank` with `query`, `documents`, and `top_n`. Results
must contain distinct valid document indices and finite `relevance_score`
values. The client returns scores in descending order.

Both endpoints receive `truncate_prompt_tokens`, defaulting to 8192. This
limits model input; the full extracted text remains in the local index.
The gateway must support these request fields, including the query/document
instruction distinction; a generic embeddings-only API is not sufficient.

The client uses `INFERENCE_API_KEY`, verifies HTTPS for remote endpoints,
does not follow redirects, and ignores proxy environment variables. Local
loopback HTTP is supported. Requests use a 120-second timeout with a
10-second connection timeout and up to three attempts for transport errors,
429, and selected temporary server errors. HTTP error messages omit response
bodies so echoed transcripts do not enter indexing logs.

Index fingerprints include the provider, model alias, dimension, endpoint,
operator-supplied revision, token limit, and query/document preprocessing
version. Changing any of these requires compatible reindexing. Encoding and
batch size do not change the embedding space.

See the [setup procedure](../processes/inference-api-setup.md).
Loading
Loading