Skip to content

FIX: keep Prompt Shield metadata in a shape the Score model accepts - #3045

Open
Chen Yufeiyang (feiiiiii5) wants to merge 1 commit into
microsoft:mainfrom
feiiiiii5:fix/prompt-shield-metadata
Open

Chen Yufeiyang (feiiiiii5) wants to merge 1 commit into
microsoft:mainfrom
feiiiiii5:fix/prompt-shield-metadata

Conversation

@feiiiiii5

Copy link
Copy Markdown
Contributor

Description

FIX: PromptShieldScorer raised a ValidationError for every real Prompt Shield response, so the scorer could not be used against the endpoint at all.

_score_piece_async parses the endpoint body and puts it in score_metadata:

# pyrit/score/true_false/prompt_shield_scorer.py:88
meta = json.loads(response)
...
score_metadata=meta,  # type: ignore[ty:invalid-argument-type]

Score.score_metadata is a flat dict[str, str | int | float] (pyrit/models/score/score.py:99), and the real body is nested — {"userPromptAnalysis": {...}, "documentsAnalysis": [{...}]}, which is the shape the scorer's own test fixture uses. So Score(...) fails with six validation errors, and score_async wraps that into RuntimeError: Error in scorer PromptShieldScorer: ...:

>>> await scorer.score_text_async("hello")
RuntimeError: Error in scorer PromptShieldScorer: 6 validation errors for Score
score_metadata.userPromptAnalysis.str  Input should be a valid string
...

The line dates from 7b5a7481 (#1104), when Score was still a plain dataclass; it became a validating pydantic model in 0899b214 (#2491) and this site was not updated. The symmetric construction path in the same package already does it right: _build_unvalidated_score normalises metadata to primitives and drops anything else ("Unrecognized metadata shape; drop to avoid downstream errors", pyrit/score/response_handler.py:154-162).

What changed: the response is stored as the JSON text it arrived as — score_metadata={"response": response} — which is also what _parse_response_to_boolean_list's docstring already tells callers the metadata attribute holds ("you can just access the metadata attribute to get the original Prompt Shield endpoint response, and then just call json.loads() on it"). The # type: ignore is gone because the value now satisfies the model.

Related: #2925 fixed the same class of problem for SelfAskLikertScorer's metadata; this is the PromptShieldScorer instance of it.

Tests and Documentation

tests/unit/score/test_prompt_shield_scorer.py::test_prompt_shield_scorer_metadata_is_the_response_text drives score_text_async against a mocked PromptTarget that returns the real-shaped body, and asserts the score and the metadata survive the Score model, and that json.loads on the stored value round-trips the body. It fails on main with the ValidationError above and passes on this branch.

pytest tests/unit/score/test_prompt_shield_scorer.py -q      3 passed
pytest tests/unit/score -q                                   42 failed, 3041 passed, 51 skipped

The 42 failures are identical on main (3d279a84) with the same names — they come from optional dependencies this environment does not have (transformers, and similar), not from this change. ruff check and ruff format --check (0.16.10) are clean on both changed files. I did not run JupyText or the doc notebooks, and no documentation changes were needed: the scorer's documented behaviour is what the fix restores.

PromptShieldScorer stored the parsed endpoint body in score_metadata, but Score
only accepts flat str/int/float values, so every real Prompt Shield response
raised a ValidationError that score_async re-raised as a RuntimeError. Store the
response as the JSON text it arrived as, which is what the scorer's docstring
already says the metadata attribute holds.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant