Skip to content

[BREAKING] [FIX]: Evaluate attack verdicts over final traces - #150

Draft
Spencer Schoenberg (spencrr) wants to merge 9 commits into
microsoft:mainfrom
spencrr:dev/spencrr/trace-xpia-stopping
Draft

Spencer Schoenberg (spencrr) wants to merge 9 commits into
microsoft:mainfrom
spencrr:dev/spencrr/trace-xpia-stopping

Conversation

@spencrr

@spencrr Spencer Schoenberg (spencrr) commented Aug 4, 2026 •

Copy link
Copy Markdown
Contributor

Description

XPIA now evaluates the attack objective once over the final trace, separately from online stopping. stop_when accepts an evaluator, None, or the default StopWhen.AUTO. StopWhen.AUTO reuses the verdict evaluator as the stop condition only when detection stays true as turns are appended: ToolCalled, SideEffectOccurred, and ResponseContains with ResponseScope.ANY_TURN. StopWhen is exported from rampart.attacks and rampart.

When the stop and verdict evaluators are the same object, the latest online evaluation is reused if it already covers the final trace, so no duplicate judge call is made. Final verdict evidence is stored on Result.final_trace_evaluation, trace completion uses TraceEndReason, and online stop evidence uses EvaluationPurpose.STOP_CHECK. Final evaluation runs inside the active session and injection stack. Observability downgrades are recorded without mutating response metadata, and cleanup failures discard otherwise successful verdict evidence and return ERROR.

The list-based resolve_as_attack and per-turn evaluate_turn_async helpers are removed because no built-in strategy uses them. The trace-contract declaration records a compatible change against main.

Depends on #149.

Breaking changes

  • XPIA verdicts are computed once over the final trace. Single-trigger attacks with deterministic evaluators keep their verdicts.
  • With the default StopWhen.AUTO, other evaluators, including LLM judges, no longer stop early. They are called once on the final trace, and adaptive drivers can run up to max_turns. Pass the same evaluator as stop_when to restore per-turn early stopping without a duplicate final call.
  • With the default, stochastic evaluators are sampled once per run instead of once per turn, so trial pass rates can shift.
  • Auto-stopped attacks expose online evaluator feedback to adaptive drivers.
  • Removed helpers: replace resolve_as_attack(eval_results=...) with resolve_attack_verdict(evaluation=...), and replace evaluate_turn_async with run_trace_async plus evaluate_final_trace_async.

The upgrade note in docs/attacks/xpia.md covers these changes.

Checklist

  • pre-commit run --all-files passes
  • Tests added or updated for changes — automatic, explicit, and disabled stopping; StopWhen values; exact evaluator call counts; stable-evaluator classification and composition; cleanup ordering; zero and max turns; observability; metadata isolation; and summaries
  • Documentation updated

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
There may be pipelines that require an authorized user to comment /azp run to run.

@spencrr
Spencer Schoenberg (spencrr) force-pushed the dev/spencrr/trace-xpia-stopping branch from d80c55b to 84d0197 Compare August 8, 2026 02:31
@spencrr
Spencer Schoenberg (spencrr) force-pushed the dev/spencrr/trace-xpia-stopping branch 2 times, most recently from cdcad35 to bedbf6b Compare August 27, 2026 17:30
@spencrr
Spencer Schoenberg (spencrr) force-pushed the dev/spencrr/trace-xpia-stopping branch 2 times, most recently from 3051684 to 4074ae9 Compare September 8, 2026 22:46
@spencrr
Spencer Schoenberg (spencrr) force-pushed the dev/spencrr/trace-xpia-stopping branch 4 times, most recently from 3ee9e25 to bcc3fd6 Compare October 1, 2026 00:59
@spencrr Spencer Schoenberg (spencrr) changed the title [FIX]: Evaluate attack verdicts over final traces [FIX] [BREAKING]: Evaluate attack verdicts over final traces Oct 1, 2026
@spencrr
Spencer Schoenberg (spencrr) force-pushed the dev/spencrr/trace-xpia-stopping branch from bcc3fd6 to bb58b64 Compare October 8, 2026 02:14
@spencrr Spencer Schoenberg (spencrr) changed the title [FIX] [BREAKING]: Evaluate attack verdicts over final traces [BREAKING] [FIX]: Evaluate attack verdicts over final traces Oct 8, 2026
Spencer Schoenberg (spencrr) added a commit that referenced this pull request Oct 9, 2026
<!-- This repository loosely follows the "Conventional Commits"
specification for commit messages. See
https://www.conventionalcommits.org/ for more information. -->
<!-- Squash-merge commit messages must match:
^\[(FEAT|FIX|REFACTOR|STYLE|TEST|DOCS|CI|MAINT|META|REVERT)\](
\[BREAKING\])?:\s.+\(#\d+\) -->
<!-- GitHub appends (#N) automatically during squash-merge; just use the
[TAG]: description format for your PR title. -->
<!-- If your PR is not yet ready for review, please mark it as [DRAFT].
-->

## Description

<!-- What does this PR do? Provide a brief summary of the changes. -->
<!-- Mention any relevant issues or pull requests with #<issue_number>
-->
<!-- Tag any relevant reviewers or teams using @<username> -->

Adds the linear trace runner shared by attack and probe strategies.
`run_trace_async` drives a conversation with an optional online
`stop_when` check, keeps separate raw and annotated turn histories, and
records `TraceEndReason`. `evaluate_final_trace_async` evaluates the
final trace once. It reuses the latest online evaluation only when the
evaluator, raw turns, manifest, and observability level are all
identical.

The runner does not own session lifetime, polarity, cleanup, or
exception conversion; those remain strategy and `BaseExecution`
responsibilities. Driver history receives a shallow copy, evaluator
contexts contain annotation-free turns, and reused evaluator evidence is
copied defensively.

`EvaluationRecord`, `TraceRun`, `run_trace_async`, and
`evaluate_final_trace_async` are exported from `rampart.core` for custom
strategies and documented in the API reference. Built-in probes adopt
them in #149 and XPIA in #150.

## Breaking changes
<!-- If none, write "None". If breaking, describe the impact and
migration path. -->

None. This PR adds APIs and does not change existing strategies.

## Checklist

- [x] `pre-commit run --all-files` passes
- [x] Tests added or updated for changes <!-- Please describe what tests
were added or updated --> — termination reasons, raw and annotated
histories, stop behavior, exact-context reuse, changed observability and
manifest, optional-evidence copying, post-run mutation, exceptions, and
turn budgets
- [x] Documentation updated
Retire the list reducer after probe execution adopts terminal evaluation. Keep explicit response scopes, online evidence separation, zero-turn errors and xdist v3 semantics consistent across tests, exports and extension guidance.
Call evaluate_final_trace_async and describe probe verdict evidence as final-trace evaluation. Record a compatible trace-contract decision against the merged v2 base; Result fields and schemas are unchanged.
Explain the per-turn to final-trace verdict change, its single-prompt blast radius, trial sampling impact, and resolver replacement. Remove private xdist envelope details already covered by schema-drift rejection and the mixed-version limitation.
Remove the old reducers and helper now that both built-in strategies evaluate terminal traces. Keep explicit-scope stopping classification, provenance and extension documentation aligned with the shared runner.
Call evaluate_final_trace_async and describe XPIA verdict evidence as final-trace evaluation in code, tests, and extension guidance.
Explain the per-turn to final-trace verdict change, its single-trigger blast radius, automatic stopping defaults, LLM-judge cost trade-off, trial sampling impact, and replacements for removed helpers.
Type Attacks.xpia(stop_when=...) as Evaluator | StopWhen | None and default to StopWhen.AUTO, following the enum-over-Literal standard. StopWhen is a string enum, so its equal string value remains accepted at runtime. Export it from rampart.attacks and the top-level package.
@spencrr
Spencer Schoenberg (spencrr) force-pushed the dev/spencrr/trace-xpia-stopping branch from bb58b64 to 8bbadc9 Compare October 9, 2026 00:48

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant