Skip to content

feat: Add role-aware Scenario target-attempt accounting (#3043) - #3062

Open
Nimit Jain (Nimit3418) wants to merge 10 commits into
microsoft:mainfrom
Nimit3418:feat-role-aware-accounting
Open

Nimit Jain (Nimit3418) wants to merge 10 commits into
microsoft:mainfrom
Nimit3418:feat-role-aware-accounting

Conversation

@Nimit3418

Copy link
Copy Markdown

Description

This PR implements the role-aware target-attempt accounting projection requested in #3043.

Key Changes:

  • Database-Side Aggregations: Updated memory_interface.py to extract result_role and aggregate attempt metrics by role (TARGET_FACING, ORCHESTRATION, and UNKNOWN_ROLE) directly via SQL, ensuring SQLite and Azure SQL parity.
  • Conservative Role Decoding: Unrecognized or legacy roles explicitly default to UNKNOWN_ROLE to protect legacy runs and prevent silent redefinitions.
  • Separation of Raw vs. Logical Progress: Introduced ScenarioProducerCounts. Target-facing attempts now strictly represent physical requests, while orchestration envelopes complete logical units without inflating raw attempt metrics.
  • API Mappings: ScenarioProducerCounts is now properly typed and mapped seamlessly into ScenarioRunListItem and ScenarioRunSummary.

Closes #3043

Testing

  • Ran make unit-test locally and verified that all 24,000+ unit tests pass 100% green.
  • Verified that legacy scenarios with no result_role correctly fallback to UNKNOWN_ROLE without crashing.

Roman Lutz (romanlutz) and others added 8 commits October 8, 2026 00:21
Interpret saved outcome counts through an async SDK with shared loop-bound admission, cooperative cancellation, and deterministic cleanup. Match bounded Unicode profile aggregation to the complete SQL fallback, with exact drilldowns, focused compatibility coverage, and SDK lifecycle documentation.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Preserve cancellation when admission races with a slot grant on Python 3.11. Release unscheduled operation capacity and close rejected coroutines, keep shutdown retryable, and log failures that race with caller cancellation. Add deterministic lifecycle regressions and verify raw-result analytics remains distinct from scenario-unit statistics and explicit result-role policy.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Expose decided-only and all-outcome success rates through one calculator and shared model. Keep raw saved-result selection distinct from scenario latest-unit selection while reusing count validation, rates, shares, and percentage formatting. Preserve legacy defaults and carry both rates through scenario progress and JSON projections.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Reject contradictory scenario totals and outcome breakdowns, including mutated inputs at aggregation boundaries. Expose an explicit read-only success_rate_decided alias in shared statistics and JSON. Add typed include_outcome_statistics opt-ins to maintained result and async technique analytics while preserving the six-field default and existing deprecation schedules.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@Nimit3418

Copy link
Copy Markdown
Author

@microsoft-github-policy-service agree

@Nimit3418

Copy link
Copy Markdown
Author

Hi Roman Lutz (@romanlutz), thanks for the guidance and the heads-up on the overlapping changes!

I've updated the PR to coordinate with the shared OutcomeStatistics calculation from #3059 and fully addressed the role-aware accounting constraints you outlined.

Here’s a breakdown of how it was implemented in this update:

  1. Integrated with FEAT: Attack analytics SDK #3059's OutcomeStatistics:
    Reused compute_outcome_statistics directly rather than building a custom success-rate calculation.
    We cleanly append producer_counts to ScenarioProgressCounts while leaving the success_rate_decided and success_rate_all logic (and the base denominators) completely untouched.

  2. Analytics / SQL Counting Parity:
    Analytics owns the policy: scenario_statistics.py now explicitly loops over attempts and buckets them into target_facing, orchestration, and unknown (tracking attempts, errors, and retries for each). We rely on AttackResultMetadata as the role contract and explicitly preserve UNKNOWN roles rather than discarding them.
    SQL Aggregate Match: I updated _build_scenario_history_aggregate_statement to mirror the SDK's selection logic exactly. The SQL case() statements explicitly filter on is_planned == 1 when tallying producer counts, ensuring we correctly ignore unplanned/legacy attempts just like the SDK does. This guarantees 100% mathematical parity between the database rollups and the in-memory calculations.

  3. Preserved Existing Base Metrics:
    As requested, I added separate role-aware rollups without redefining or mutating any existing counts. unit_count, completed_units, successful_units, and the latest-attempt selection logic are all preserved.
    The edge cases work as intended: for example, 2 target-facing child results + 1 Sequential envelope will successfully show as 2 target-facing attempts and 1 orchestration attempt inside producer_counts, while still correctly rolling up as exactly 1 outer planned unit in the core metrics. Preparation failures also cleanly map to their respective producer categories.

  4. Testing & Validation:
    Expanded the SQLite parity tests (test_scenario_statistics_parity.py) with role_aware_accounting scenarios to cover mixed-role histories and validate the fallback/unplanned grouping behavior.
    Added Azure SQL query-compilation checks (test_azure_sql_memory.py) to ensure the JsonScalar role-extraction compiles and executes safely.
    Ran the full suite locally—all 24,000+ unit tests are passing cleanly.

Let me know if there's anything else you'd like adjusted. Otherwise, it should be good for another look whenever you have time!

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

FEAT Add role-aware Scenario target-attempt accounting

2 participants