Skip to content

fix(gemma4): preserve sliding-window KV history during verification - #824

Draft
pculaf wants to merge 2 commits into
Luce-Org:mainfrom
pculaf:fix/gemma4-swa-cache-headroom
Draft

pculaf wants to merge 2 commits into
Luce-Org:mainfrom
pculaf:fix/gemma4-swa-cache-headroom

Conversation

@pculaf

@pculaf pculaf commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

This PR fixes the KV cache allocation for the sliding-window attention layers and adds checks on forward lengths and limits on layer-split chunk sizes to ensure they fit the allocated cache. Once the context reached the sliding-window boundary, verification could overwrite older K/V entries that earlier positions in the verification sequence still needed. This corrupted the attention inputs and could change the model’s predictions. Also, the values from one head could overflow and be written in the memory allocated for the next head. These overflows are now prevented.

Cache capacity is increased to reserve space for both the attention window and the maximum forward sequence length, following a similar approach to the Laguna implementation in Lucebox.

Forward calls are checked against the allocated capacity before writing to the cache. Layer-split execution also limits each shard’s forward length to what its cache can safely accommodate.

I added a CPU unit test covering cache sizing, range checks, history preservation across ring wraparound and rejection, and layer-split chunk limits.

View guided diff

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant