Repository navigation
fix(extraction): put the extraction instructions in the fork's prompt text - #52
Merged
Merged
Conversation
kuitos
added this pull request to stack #55
October 7, 2026 14:53
kuitos
force-pushed
the
fix/frame-extraction-transcript
branch
from
October 7, 2026 15:19
4ecdc41 to
8f64339
Compare
… text The extraction fork's only user message was the raw `### User / ### Assistant` transcript; the task lived in the system prompt alone. Fast models (Claude Haiku 4.5, GPT-5.4 Mini/Nano) read it as a conversation waiting for a reply and answered its last question instead of extracting, in 9 of 9 reported V2 runs. The fork's prompt text now states the task, wraps the transcript in a <transcript> block (a quoted closing tag is escaped), and ends with a reminder to only record memories and not answer or continue the conversation. The system prompt is unchanged. The auto-dream fork's message restates its task the same way. Both hosts send the same text. Refs #48
kuitos
force-pushed
the
fix/frame-extraction-transcript
branch
from
October 7, 2026 15:23
8f64339 to
227b6d6
Compare
kuitos
marked this pull request as ready for review
October 7, 2026 15:24
A deadline timer can fire a millisecond before Date.now() reaches the fork's budget. remaining() then still read > 0, so the TimeoutError was rethrown as a plain error: the fork was removed without being interrupted. This made 'a hung wait times out, interrupts, then removes' and the V2 setup timeout test flaky on ubuntu CI. A deadline that ran on the fork's whole remaining budget now always counts as the fork's timeout.
|
🎉 This PR is included in version 2.1.3 🎉 The release is available on: Your semantic-release bot 📦🚀 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
On V2, the extraction fork's only user message was the raw
### User / ### Assistanttranscript (ExtractionCoordinator:system: buildExtractionSystemPrompt(...),text: conversation). The task itself lived only in the system prompt. #48 (point 1) reports that Claude Haiku 4.5, GPT-5.4 Mini and GPT-5.4 Nano read the transcript as a conversation waiting for a reply and answered its last question instead of extracting, in 9/9 runs. One fork explained Python decorators; after a coding turn, another re-ranpytest. Recall already restates its task in the prompt text (buildGeneratePrompt); extraction did not.Fix
buildExtractionUserMessage()(src/extraction/prompts.ts) builds the fork's prompt text:<transcript>block, with any quoted</transcript>escaped so it cannot end the block early;The system prompt is unchanged. The auto-dream fork's message restates its task the same way. Both hosts send the same text, through
MemoryHost.runFork.(The V2 fork-sandbox half of #48 is handled in a separate PR.)
Tests
test/extraction/ExtractionCoordinator.test.ts:test/v2/plugin.test.ts: the V2 fork prompt starts with the task and contains the closing reminder.bun run lint,bun run typecheck,bun testandbun run buildpass.Real-environment verification (OpenCode 2.0.22, isolated DB/config)
Setup:
opencode/big-pickle) was asked a question about Python decorators, withmemory_save,memory_deleteandexternal_directorydenied so that only the fork could save anything.opencode/gpt-6-luna.Results:
memory_list/memory_saveand ended with a one-line summary.Limitation, stated plainly: the models that failed in #48 (Haiku 4.5, GPT-5.4 Mini/Nano) could not be run here. Zen's free models return 403 for sandboxed forks, and gpt-6-luna may well have extracted correctly before this change too. So this run shows no regression, not the fix itself. The fix follows the issue's diagnosis: the same framing recall already uses.
Refs #48
Also: flaky V2 fork timeout (CI fix)
runV2Forkcheckedremaining() <= 0after aTimeoutError, but a deadline timer can fire a millisecond beforeDate.now()reaches the budget. The timeout was then rethrown as a plain error and the fork was removed without being interrupted, which maderunV2Fork > a hung wait times out…andV2 setup: extraction > ⑧fail intermittently on ubuntu. A deadline that ran on the fork's whole remaining budget now always counts as the fork timeout. New regression test freezesDate.now()withsetSystemTimeand fails without the fix.🤖 Generated with Claude Code