Repository navigation
[BREAKING] FEAT: Make scenario dataset sources and limits explicit - #2956
Richard Lundeen (richlundeen) wants to merge 5 commits into
Conversation
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Resolve latent-injection test coverage with an explicit total limit while preserving default and all-limit behavior. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Use explicit dataset sources and max_total=all for the complete latent-injection population check. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Preserve upstream request bounds alongside explicit dataset limits. Increment the benchmark version to 7 for independent named-source sampling. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Remove sampling_scope, retain empty dataset keys, and preserve Psychosocial sub-harm coverage within the total cap. Update scenario guidance and adjust the merged adaptive run-plan test. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
| if not isinstance(max_dataset_size, _Unset): | ||
| if not isinstance(max_total, _Unset): | ||
| raise ValueError("Use only one of 'max_dataset_size' and 'max_total'.") |
There was a problem hiding this comment.
max_dataset_size=10 with one dataset gives 10 on main but 5 here, bc max_per_dataset still defaults to 5. None went from unlimited to 5 too. the warning just says rename, so no one will catch it. can we set max_per_dataset="all" here
| raise DatasetConstraintError( | ||
| f"Psychosocial max_total ({cap}) must cover every selected sub-harm " | ||
| f"({len(self._selected_sub_harms())}); use max_per_dataset=1 for one objective per sub-harm." |
There was a problem hiding this comment.
the check is cap < len(sub_harms), so max_per_dataset=1 doesn't change it. and the backend resolver only takes max_dataset_size, so you can't even set it there. maybe say bump max_total or pick fewer sub-harms?
hannahwestra25
left a comment
There was a problem hiding this comment.
ghcp found :
doc/scanner/garak.py:602 uses dataset_names=, which is deprecated now. and the line below still has max_dataset_size. those are pre-existing, but since we're already in this file could we migrate them? otherwise the docs teach the old way.
so double check that all the docs have been updated
Description
Scenario dataset preparation and limits need clear, consistent rules so overrides do not change fetch policies or remove unrelated caps. This implements the next step in the dataset-generation proposal.
Adds named dataset sources, explicit preparation, and separate source and total limits while keeping dataset reads free of fetching. Uses one limit contract across Python, CLI, API, and GUI: omitted,
None, empty, and"default"use defaults;"all"removes the specified cap. Migrates built-in scenarios and preserves retained source settings during overrides; changingNonefrom unlimited to default is a breaking change.Tests and Documentation
Added regression coverage for preparation, selection, overrides, limit normalization, request persistence, and GUI controls. Updated scenario guidance and paired scenario/scanner notebooks.
Validation commands and results
$env:UV_NO_SYNC = '1'; uv run --no-sync pre-commit run --all-files— passed all hooks, including full Python type checking.uv run --no-sync pytest tests\unit\scenario tests\unit\models\test_scenario_request.py tests\unit\models\test_import_boundary.py tests\unit\backend\test_scenario_configuration_resolver.py tests\unit\backend\test_scenario_service.py tests\unit\backend\test_scenario_run_service.py tests\unit\backend\test_scenario_resume.py tests\unit\cli\test_cli_args.py tests\unit\cli\test_pyrit_scan.py tests\unit\cli\test_pyrit_shell.py tests\unit\cli\test_api_client.py -q -n 4 --tb=short --disable-warnings— 2,252 passed.frontend:npm run type-check -- --pretty false— passed.frontend:npm test -- --runInBand --silent src\components\Scenarios\ScenarioDetail.test.tsx— 65 passed.Jupytext cell comparison confirmed that both changed
.py/.ipynbpairs match. Notebooks were not executed.