Repository navigation
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR fixes the KV cache allocation for the sliding-window attention layers and adds checks on forward lengths and limits on layer-split chunk sizes to ensure they fit the allocated cache. Once the context reached the sliding-window boundary, verification could overwrite older K/V entries that earlier positions in the verification sequence still needed. This corrupted the attention inputs and could change the model’s predictions. Also, the values from one head could overflow and be written in the memory allocated for the next head. These overflows are now prevented.
Cache capacity is increased to reserve space for both the attention window and the maximum forward sequence length, following a similar approach to the Laguna implementation in Lucebox.
Forward calls are checked against the allocated capacity before writing to the cache. Layer-split execution also limits each shard’s forward length to what its cache can safely accommodate.
I added a CPU unit test covering cache sizing, range checks, history preservation across ring wraparound and rejection, and layer-split chunk limits.