Repository navigation
FEAT Add multi-label true/false scoring with explicit label selection - #2858
biefan (biefan) wants to merge 1 commit into
Conversation
|
Thanks for this contribution. We are interested in support for multiple labeled true/false verdicts and want to keep this PR open. I will return to it after we merge #2899, which changes the shared scorer contract and target-backed scoring path. That will give us a stable base to align this implementation with the broader scorer design. |
| request_prompt = render_wildguard_prompt( | ||
| response=message_piece.converted_value, user_prompt=user_prompt, prompt_template=self._prompt_template | ||
| ) | ||
| parsed = await _run_llm_scoring_async( |
There was a problem hiding this comment.
Please rebase this implementation onto the scorer contract from #2899 and use the new contracts directly. The merged _run_llm_scoring_async accepts a JudgmentRequest, so the separate arguments here cause normal scoring to fail with an unexpected-keyword TypeError.
The concrete scorer should construct TargetJudge(target=chat_target, requirements=type(self).TARGET_REQUIREMENTS) and implement _score_piece_with_expectation_async(..., expectation=...). Render the prompt here, put the full effective expectation in a JudgmentRequest, call _capture_judgment_evidence to attach the scored evidence, and pass that request and the response handler to self._judge.judge_async. Construct the labeled scores with scored_expectation=expectation, rather than the compatibility objective input.
Please also remove chat_target from the new MessageMultiLabelTrueFalseScorer constructor and its forwarding to the base. That parameter is now deprecated validation-only wiring. The concrete judge owns target requirements; the family base owns labels, message policy, and aggregation. Keep target discovery through get_chat_target().
We want this new code to use only the new contracts, not an adapter for the old transport signature or objective-only piece hook. Please migrate the synthetic test scorer and its patched hooks as well, and cover retained observations and full expectations through direct and selector scoring.
| fallback = self._build_neutral_fallback_score(message=message, objective=objective, neutral_value="false") | ||
| return self._label_undetermined_score(fallback[0]) if fallback else [] | ||
|
|
||
| def _finalize_message_scores( |
There was a problem hiding this comment.
This override no longer participates in the message pipeline after #2899. The merged base calls _finalize_message_scores_async, not _finalize_message_scores. If a judge blocks and raise_if_scorer_blocks=False, the shared handler returns one unlabeled undetermined score. Without this expansion, validate_return_scores rejects it instead of returning an undetermined verdict for each label.
Please move this behavior to the current async finalization hook and await the base implementation after expanding the scores. Preserve the shared blocked-judge policy, observation IDs, evidence anchors, and expectation stamping. Do not add a compatibility dispatch to the removed synchronous hook.
Please update the blocked-judge regression to exercise the typed piece hook on the rebased code. It should verify one undetermined score per declared label, distinct score IDs, the retained error observation when available, and the full scored expectation. Also cover this through a selector, where one public projection is returned and the labeled child results are retained as intermediate scores.
| return user_prompt if user_prompt.strip() else None | ||
| if not message_piece.conversation_id or message_piece.sequence < 1: | ||
| return None | ||
| conversation = await asyncio.to_thread(memory.get_message_pieces, conversation_id=message_piece.conversation_id) |
There was a problem hiding this comment.
Please extract the current async context-resolution implementation from main, rather than restoring the deprecated synchronous memory path. WildGuard now reads this history with await memory.get_message_pieces_async(conversation_id=...). Wrapping get_message_pieces in asyncio.to_thread still uses the legacy API. We do not want new code to depend on that compatibility layer.
Keep the existing semantics: an explicit blank override must not fall back to history, choose the latest earlier user turn before filtering modalities, and use converted text rather than the original seed text.
The new tests and paired notebook also call synchronous memory methods. Please migrate their setup, score and observation queries, and cleanup to the async APIs, including add_message_to_memory_async, get_scores_async, get_observations_async, and dispose_engine_async. Make their storage helpers async. The category-membership fix should remain in the shared query implementation used by get_scores_async, with coverage through that API. Do not implement a separate fix only in the deprecated sync entry point.
| ) | ||
| return [ | ||
| score.model_copy(deep=True, update={"id": uuid.uuid4(), "scorer_class_identifier": self.get_identifier()}) | ||
| for score in scores |
There was a problem hiding this comment.
Please use the current _create_wrapper_score(score) helper for the selected result instead of duplicating the copy logic here. It creates a wrapper-owned result with a new ID, timestamp, and scorer identity while preserving the child judgment and evidence links.
The persistence contract also changed with #2916: _score_nested_async collects all labeled child results as intermediate scores, and the public root commits them with the selected projection. The statement above and in the notebook that only the projection is persisted is therefore no longer correct. Default queries return the public projection, but include_intermediate=True exposes the retained child labels.
Please update the documentation and tests to show this distinction. For one message with three labels, a selector should make one classifier call, return one public score, retain all three child verdicts, and give the projection a distinct ID and the selector identity. Querying intermediate scores should expose those labels without another inference. This is the useful attack path: select one objective verdict and retain the other judgments from the same call, rather than telling users to score the root again to save them.
Separate selector calls can still make separate inferences. This does not require adding a cache.
Description
Related to #2565.
A classifier can answer several independent questions in one inference, but a single
MessageTrueFalseScoreraggregates them into one boolean. WildGuard currently exposes one selected verdict and keeps the other judgments in metadata. Users cannot query or evaluate those secondary judgments as normal scores without making separate classifier calls.This adds an opt-in multi-label contract and a working WildGuard implementation. For example, one successful classification of a harmful request that receives a refusal can now produce:
score_categoryharmful_requestresponse_refusalharmful_responseThe three scores have distinct IDs and reference the same retained classifier observation. Choosing
harmful_responseas the objective therefore yields False, regardless of the other labels' True values.API and design choices
MultiLabelTrueFalseScorerdeclares the output labels and validates exactly one category per verdict.MessageMultiLabelTrueFalseScoreradds the existing message pipeline and aggregates supported pieces within each label. Unknown, duplicate, or missing labels are rejected before persistence. Non-applicable evidence still returns[].Score.score_category=[label], so scores retain their existing IDs, evidence anchors, expectations, serialization, and observation links. There is no schema migration. The existing category query is fixed to match complete JSON-array elements, case-insensitively; comparing the entire array with a string previously missed these stored scores.TrueFalseScoreSelectoradapts a named label toTrueFalseScorerfor attacks, boolean wrappers, and objective evaluation. Its identity includes the source configuration and selected label, and its batch path retains the source target's rate-limit policy. Raw multi-label scorers cannot silently enter single-verdict consumers or float thresholding.WildGuardMultiLabelScoreruses the shared transport, three-label parser, prompt template and context resolution. It makes one classifier call per supported text piece, excluding retries for malformed responses. Any label may returnN/Awithout discarding the other judgments or triggering a retry; it becomesUNDETERMINED. Fully blocked/unreadable evidence also remains undetermined for every label.Persistence and compatibility: a selector follows the existing wrapper contract and persists only its selected projection. Call the multi-label root directly to save every label. Separate selector calls do not share cached inference. Existing
WildGuardScorer(label=...)and single-verdict scorers retain their APIs and verdict behavior; only WildGuard's unchanged context-resolution logic is shared with the new implementation.Tests and Documentation
make unit-test(Python 3.11, default dependencies): 20,464 passed, 146 skipped, 1 failed. The sole failure is the pre-existingtest_get_seed_dataset_summaries_follows_a_trailing_blank_insensitive_collationintests/unit/memory/memory_interface/test_interface_seed_prompts.py, independently reproduced at unmodified basef65263eusing the same environment.The new offline notebook was executed with the checkout's virtual-environment kernel. Its retained output shows 1 classifier call, 3 persisted scores, 1 shared observation, and no additional call when reading saved categories. The paired Python source matches the notebook. The framework documentation and scoring navigation are updated.
Validation is local and offline; no live WildGuard endpoint, model weights or paid model calls were used.