Repository navigation
fix: add anti-preamble instructions to role play attack system prompts - #2900
ANIRUDDHA ADAK (aniruddhaadak80) wants to merge 4 commits into
Conversation
| """Role-play persona YAMLs must instruct the model to avoid preamble/filler.""" | ||
| path = Path(EXECUTOR_SEED_PROMPT_PATH) / "red_teaming" / "role_play" / f"{technique_name}.yaml" | ||
| text = path.read_text(encoding="utf-8") | ||
| assert "must begin directly with the in-character attack prompt" in text, ( |
There was a problem hiding this comment.
🟡 Should fix: These assertions only prove the YAML contains the new words. They cannot show whether the generated next_message actually loses its preamble, which is the behavior requested in #1912. The adversarial reply parser accepts any nonempty next_message, including "Sure, here is...", and forwards it unchanged. Could you include the issue's before/after 50-run comparison, counting preambles in the final target-directed message, or other measured output evidence? A green substring test alone does not establish that the fix works.
There was a problem hiding this comment.
Agreed, I don't want this PR to strip or rewrite generated messages. The parser should preserve the model's output, and adding a content filter would be a separate change. Sorry, mentioning the parser made my request less clear.
My request was for measured before/after evidence that the prompt change reduces preambles, as described in #1912. The substring tests are useful guards for the instruction, but they cannot establish its effect on generated output. Could you run the 50-before/50-after comparison using the same model and settings, and count preambles in the final target-directed next_message? Please include the model/settings and report refusals or errors separately. Benign objectives are fine for this comparison; no parser changes or live-model unit tests are needed.
Auto-replied by the GitHub Copilot app
141a633 to
205e154
Compare
|
Done in 🔴 Missing import. Your read was right and it reproduced exactly:
A second one I hadn't caught. Checking the diff against
YAML strips a block scalar's base indent, so this isn't cosmetic. In the next-message prompt the stray column reaches the model and renders the flat list as a nested sub-list, and the heading moves out of column zero. Only the first content line sets the base indent, so the files still parsed, still built, and nothing complained. That is exactly why it survived review. Fixed in I added a guard in 🟡 The behavioural assertion — I'd like to push back rather than silently rework this. You are right that the assertions only prove the YAML contains the words. But the parser accepting a nonempty If you do want that hardening, I think it belongs in its own PR against
One thing still outstanding, and it is not mine to do: the CLA is unsigned, so |
Role play attacks were sometimes starting responses with conversational preamble like 'Sure, here is your objective rephrased as a detailed question set...' instead of directly outputting the attack prompt. Added explicit instructions to all 5 role play persona YAMLs and the role_play_next_message.yaml to begin next_message directly with the in-character attack prompt, with no preamble, conversational filler, or meta commentary. Fixes microsoft#1912 Signed-off-by: Aniruddha Adak <aniruddhaadak80@users.noreply.github.com>
Adding the anti-preamble bullet re-indented neighbouring lines one column to the right in all six prompts: ## Response Format went from two spaces to three in the five personas, and a run of bullets under Hard Constraints in role_play_next_message.yaml went from two to three. YAML strips a block scalar's base indent, so the stray column is not cosmetic. It reaches the model: the next-message prompt's flat bullet list became a 2/3/2 mix, which renders as a nested sub-list, and the section heading moved out of column zero. Only the first content line sets the base indent, so the files still parse and still build, which is why nothing flagged this. Every pre-existing line is back where it was; the net diff against main is now purely additive.
…tion The next-message test raised NameError because EXECUTOR_SIMULATED_TARGET_PATH was never imported; it is defined in pyrit.common.path. Added it alongside the other two executor paths. Verified: 23 passed on -k role_play before this, 1 failed with NameError on the branch. Also guards the prompt indentation regression fixed in the previous commit. Every ## heading and top-level bullet inside a �alue: | block must sit at the block's base indent, across all six role-play prompts rather than the five personas alone. Confirmed the test fails on the pre-fix indentation and passes after. The behavioural point from the review is deliberately not addressed here: the parser does not and should not strip preamble. This change is a prompt instruction, and the instruction is what the model is asked to obey. Making the reply parser defensively rewrite a nonempty next_message would be a different design change, on a different code path, and out of scope for a prompt fix.
3625203 to
e349177
Compare
|
ANIRUDDHA ADAK (@aniruddhaadak80) please read the following Contributor License Agreement(CLA). If you agree with the CLA, please reply with the following information.
Contributor License AgreementContribution License AgreementThis Contribution License Agreement (“Agreement”) is agreed to by the party signing below (“You”),
|
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
| assert "–" not in text, f"{technique_name}.yaml contains an en-dash" | ||
| assert "—" not in text, f"{technique_name}.yaml contains an em-dash" |
There was a problem hiding this comment.
🟡 Should Fix: These edits change the existing dash checks into checks for mojibake, so they no longer reject an actual en dash or em dash. I called both TestRolePlayYamls.test_yaml_has_no_em_or_en_dashes and TestPersonaCrescendoYamls.test_yaml_has_no_em_or_en_dashes with file contents containing U+2013 or U+2014, and all four checks passed.
Please restore the original needles in both classes. ASCII escapes also avoid an encoding round-trip:
assert "\u2013" not in text
assert "\u2014" not in textThis change is unrelated to the preamble fix and weakens an existing regression check.
Summary
Fixes #1912 — role play attacks were sometimes starting responses with conversational preamble like "Sure, here is your objective rephrased as a detailed question set..." instead of directly outputting the attack prompt.
Changes
role_play_movie_script,role_play_video_game,role_play_trivia_game,role_play_persuasion,role_play_persuasion_written)role_play_next_message.yamlTesting
test_yaml_has_no_preamble_instructionparametrized test for all 5 role play persona YAMLstest_role_play_next_message_has_no_preamble_instructionfor the next-message prompt