Skip to content

FEAT Add multi-label true/false scoring with explicit label selection - #2858

Open
biefan (biefan) wants to merge 1 commit into
microsoft:mainfrom
biefan:feat/labeled-true-false-scores
Open

biefan (biefan) wants to merge 1 commit into
microsoft:mainfrom
biefan:feat/labeled-true-false-scores

Conversation

@biefan

Copy link
Copy Markdown
Contributor

Description

Related to #2565.

A classifier can answer several independent questions in one inference, but a single MessageTrueFalseScorer aggregates them into one boolean. WildGuard currently exposes one selected verdict and keeps the other judgments in metadata. Users cannot query or evaluate those secondary judgments as normal scores without making separate classifier calls.

This adds an opt-in multi-label contract and a working WildGuard implementation. For example, one successful classification of a harmful request that receives a refusal can now produce:

score_category Verdict Persisted independently
harmful_request True Yes
response_refusal True Yes
harmful_response False Yes

The three scores have distinct IDs and reference the same retained classifier observation. Choosing harmful_response as the objective therefore yields False, regardless of the other labels' True values.

API and design choices

  • Separate result family. MultiLabelTrueFalseScorer declares the output labels and validates exactly one category per verdict. MessageMultiLabelTrueFalseScorer adds the existing message pipeline and aggregates supported pieces within each label. Unknown, duplicate, or missing labels are rejected before persistence. Non-applicable evidence still returns [].
  • Existing score representation. Labels use Score.score_category=[label], so scores retain their existing IDs, evidence anchors, expectations, serialization, and observation links. There is no schema migration. The existing category query is fixed to match complete JSON-array elements, case-insensitively; comparing the entire array with a string previously missed these stored scores.
  • Explicit single-label projection. TrueFalseScoreSelector adapts a named label to TrueFalseScorer for attacks, boolean wrappers, and objective evaluation. Its identity includes the source configuration and selected label, and its batch path retains the source target's rate-limit policy. Raw multi-label scorers cannot silently enter single-verdict consumers or float thresholding.
  • WildGuard implementation. WildGuardMultiLabelScorer uses the shared transport, three-label parser, prompt template and context resolution. It makes one classifier call per supported text piece, excluding retries for malformed responses. Any label may return N/A without discarding the other judgments or triggering a retry; it becomes UNDETERMINED. Fully blocked/unreadable evidence also remains undetermined for every label.
classifier = WildGuardMultiLabelScorer(chat_target=wildguard_target)
scores = await classifier.score_async(scorable=evidence)  # Saves all three labels.

objective_scorer = TrueFalseScoreSelector(
    scorer=classifier,
    label="harmful_response",
)

Persistence and compatibility: a selector follows the existing wrapper contract and persists only its selected projection. Call the multi-label root directly to save every label. Separate selector calls do not share cached inference. Existing WildGuardScorer(label=...) and single-verdict scorers retain their APIs and verdict behavior; only WildGuard's unchanged context-resolution logic is shared with the new implementation.

Tests and Documentation

  • 45 new cases cover independent persistence/querying, JSON round trips, per-label AND/OR aggregation, invalid output contracts, abstentions, unavailable evidence, shared observations, loose-content anchors, batch context isolation, explicit selection, wrapper composition, rate-limit enforcement, and label-specific evaluation. An attack-level test confirms the selected label controls the actual attack outcome rather than the first classifier result.
  • WildGuard tests exercise the real completion target, normalizer, parser and SQLite memory, mocking only the external SDK response. They verify one successful response produces three scores and one shared observation, and malformed-response retries do not duplicate scores.
  • Category-query regressions fail against unmodified upstream and pass here. SQLite behavior and SQL Server query construction/bound parameters are covered; no live SQL Server was used.
  • Focused scorer, memory, registry and lazy-import suite: 589 passed.
  • make unit-test (Python 3.11, default dependencies): 20,464 passed, 146 skipped, 1 failed. The sole failure is the pre-existing test_get_seed_dataset_summaries_follows_a_trailing_blank_insensitive_collation in tests/unit/memory/memory_interface/test_interface_seed_prompts.py, independently reproduced at unmodified base f65263e using the same environment.
  • All applicable pre-commit hooks passed with all optional dependencies installed, including repository-wide type checking and documentation validation.

The new offline notebook was executed with the checkout's virtual-environment kernel. Its retained output shows 1 classifier call, 3 persisted scores, 1 shared observation, and no additional call when reading saved categories. The paired Python source matches the notebook. The framework documentation and scoring navigation are updated.

uv run pytest tests/unit/score/test_multi_label_true_false_scorer.py tests/unit/score/test_wildguard_multi_label_scorer.py -q
uv run jupytext --to ipynb --execute --set-kernel python3 doc/code/scoring/6_multi_label_true_false_scorers.py

Validation is local and offline; no live WildGuard endpoint, model weights or paid model calls were used.

@richlundeen

Copy link
Copy Markdown
Contributor

Thanks for this contribution. We are interested in support for multiple labeled true/false verdicts and want to keep this PR open. I will return to it after we merge #2899, which changes the shared scorer contract and target-backed scoring path. That will give us a stable base to align this implementation with the broader scorer design.

request_prompt = render_wildguard_prompt(
response=message_piece.converted_value, user_prompt=user_prompt, prompt_template=self._prompt_template
)
parsed = await _run_llm_scoring_async(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please rebase this implementation onto the scorer contract from #2899 and use the new contracts directly. The merged _run_llm_scoring_async accepts a JudgmentRequest, so the separate arguments here cause normal scoring to fail with an unexpected-keyword TypeError.

The concrete scorer should construct TargetJudge(target=chat_target, requirements=type(self).TARGET_REQUIREMENTS) and implement _score_piece_with_expectation_async(..., expectation=...). Render the prompt here, put the full effective expectation in a JudgmentRequest, call _capture_judgment_evidence to attach the scored evidence, and pass that request and the response handler to self._judge.judge_async. Construct the labeled scores with scored_expectation=expectation, rather than the compatibility objective input.

Please also remove chat_target from the new MessageMultiLabelTrueFalseScorer constructor and its forwarding to the base. That parameter is now deprecated validation-only wiring. The concrete judge owns target requirements; the family base owns labels, message policy, and aggregation. Keep target discovery through get_chat_target().

We want this new code to use only the new contracts, not an adapter for the old transport signature or objective-only piece hook. Please migrate the synthetic test scorer and its patched hooks as well, and cover retained observations and full expectations through direct and selector scoring.

fallback = self._build_neutral_fallback_score(message=message, objective=objective, neutral_value="false")
return self._label_undetermined_score(fallback[0]) if fallback else []

def _finalize_message_scores(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This override no longer participates in the message pipeline after #2899. The merged base calls _finalize_message_scores_async, not _finalize_message_scores. If a judge blocks and raise_if_scorer_blocks=False, the shared handler returns one unlabeled undetermined score. Without this expansion, validate_return_scores rejects it instead of returning an undetermined verdict for each label.

Please move this behavior to the current async finalization hook and await the base implementation after expanding the scores. Preserve the shared blocked-judge policy, observation IDs, evidence anchors, and expectation stamping. Do not add a compatibility dispatch to the removed synchronous hook.

Please update the blocked-judge regression to exercise the typed piece hook on the rebased code. It should verify one undetermined score per declared label, distinct score IDs, the retained error observation when available, and the full scored expectation. Also cover this through a selector, where one public projection is returned and the labeled child results are retained as intermediate scores.

return user_prompt if user_prompt.strip() else None
if not message_piece.conversation_id or message_piece.sequence < 1:
return None
conversation = await asyncio.to_thread(memory.get_message_pieces, conversation_id=message_piece.conversation_id)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please extract the current async context-resolution implementation from main, rather than restoring the deprecated synchronous memory path. WildGuard now reads this history with await memory.get_message_pieces_async(conversation_id=...). Wrapping get_message_pieces in asyncio.to_thread still uses the legacy API. We do not want new code to depend on that compatibility layer.

Keep the existing semantics: an explicit blank override must not fall back to history, choose the latest earlier user turn before filtering modalities, and use converted text rather than the original seed text.

The new tests and paired notebook also call synchronous memory methods. Please migrate their setup, score and observation queries, and cleanup to the async APIs, including add_message_to_memory_async, get_scores_async, get_observations_async, and dispose_engine_async. Make their storage helpers async. The category-membership fix should remain in the shared query implementation used by get_scores_async, with coverage through that API. Do not implement a separate fix only in the deprecated sync entry point.

)
return [
score.model_copy(deep=True, update={"id": uuid.uuid4(), "scorer_class_identifier": self.get_identifier()})
for score in scores

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please use the current _create_wrapper_score(score) helper for the selected result instead of duplicating the copy logic here. It creates a wrapper-owned result with a new ID, timestamp, and scorer identity while preserving the child judgment and evidence links.

The persistence contract also changed with #2916: _score_nested_async collects all labeled child results as intermediate scores, and the public root commits them with the selected projection. The statement above and in the notebook that only the projection is persisted is therefore no longer correct. Default queries return the public projection, but include_intermediate=True exposes the retained child labels.

Please update the documentation and tests to show this distinction. For one message with three labels, a selector should make one classifier call, return one public score, retain all three child verdicts, and give the projection a distinct ID and the selector identity. Querying intermediate scores should expose those labels without another inference. This is the useful attack path: select one objective verdict and retain the other judgments from the same call, rather than telling users to score the root again to save them.

Separate selector calls can still make separate inferences. This does not require adding a cache.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants