Repository navigation
FIX: keep Prompt Shield metadata in a shape the Score model accepts - #3045
Open
Chen Yufeiyang (feiiiiii5) wants to merge 1 commit into
Open
Chen Yufeiyang (feiiiiii5) wants to merge 1 commit into
Chen Yufeiyang (feiiiiii5) wants to merge 1 commit into
Conversation
PromptShieldScorer stored the parsed endpoint body in score_metadata, but Score only accepts flat str/int/float values, so every real Prompt Shield response raised a ValidationError that score_async re-raised as a RuntimeError. Store the response as the JSON text it arrived as, which is what the scorer's docstring already says the metadata attribute holds.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
FIX:
PromptShieldScorerraised aValidationErrorfor every real Prompt Shield response, so the scorer could not be used against the endpoint at all._score_piece_asyncparses the endpoint body and puts it inscore_metadata:Score.score_metadatais a flatdict[str, str | int | float](pyrit/models/score/score.py:99), and the real body is nested —{"userPromptAnalysis": {...}, "documentsAnalysis": [{...}]}, which is the shape the scorer's own test fixture uses. SoScore(...)fails with six validation errors, andscore_asyncwraps that intoRuntimeError: Error in scorer PromptShieldScorer: ...:The line dates from
7b5a7481(#1104), whenScorewas still a plain dataclass; it became a validating pydantic model in0899b214(#2491) and this site was not updated. The symmetric construction path in the same package already does it right:_build_unvalidated_scorenormalises metadata to primitives and drops anything else ("Unrecognized metadata shape; drop to avoid downstream errors",pyrit/score/response_handler.py:154-162).What changed: the response is stored as the JSON text it arrived as —
score_metadata={"response": response}— which is also what_parse_response_to_boolean_list's docstring already tells callers the metadata attribute holds ("you can just access the metadata attribute to get the original Prompt Shield endpoint response, and then just calljson.loads()on it"). The# type: ignoreis gone because the value now satisfies the model.Related: #2925 fixed the same class of problem for
SelfAskLikertScorer's metadata; this is thePromptShieldScorerinstance of it.Tests and Documentation
tests/unit/score/test_prompt_shield_scorer.py::test_prompt_shield_scorer_metadata_is_the_response_textdrivesscore_text_asyncagainst a mockedPromptTargetthat returns the real-shaped body, and asserts the score and the metadata survive theScoremodel, and thatjson.loadson the stored value round-trips the body. It fails onmainwith theValidationErrorabove and passes on this branch.The 42 failures are identical on
main(3d279a84) with the same names — they come from optional dependencies this environment does not have (transformers, and similar), not from this change.ruff checkandruff format --check(0.16.10) are clean on both changed files. I did not run JupyText or the doc notebooks, and no documentation changes were needed: the scorer's documented behaviour is what the fix restores.