Repository navigation
FEAT: paginated seed browsing - #2989
Roman Lutz (romanlutz) merged 13 commits into
Conversation
|
@microsoft-github-policy-service agree |
- Return domain Seed objects from memory; detail members are SeedUnion. - Remove the is_jinja_template column, its migration, and the legacy routes. - Remove the is_template and is_configuration summary labels. - Exclude simulated-conversation JSON from text search. - Replace the seed browsing tests with focused memory, route, and integration tests. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
|
Jackson Severino da Rocha (@jacksonsr451) Roman Lutz (@romanlutz) I pushed a commit (343b7a7) that makes the seed browsing API smaller. Most of the PR follows Roman's contract in #2748. These are the places where it is different, and why: 1. Text search does not include simulated-conversation configurations. 2. There are no 3. The detail returns domain seed objects. 4. We do not have special code for legacy simulated-conversation rows. Other changes in this push (not from the contract):
This comment was drafted with GitHub Copilot and reviewed by me. |
Return stored seed records without loading legacy template files or dropping group members. Mask standalone text references in previews and use canonical UUID ordering for cross-database paging. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> # Conflicts: # pyrit/memory/memory_interface.py
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Recognize HTTP(S) and data URI schemes case-insensitively after leading whitespace. Derive media labels only from URL paths while preserving stored values and local filenames. Cover the shared formatter and dataset list/detail routes. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Description
Fixes #2748. This PR adds a read-only REST API to browse the seeds that are stored in memory, one page at a time. It uses the dataset
selection_keyfrom #2762 and follows the contract from Roman Lutz (@romanlutz) in #2748. This comment lists the places where it is different from that contract, and why.Routes
GET /api/datasets/seeds?selection_key=<key>returns one page of logical examples.GET /api/datasets/seeds/{example_id}?selection_key=<key>returns all seeds of one example.A logical example is the set of seeds that have the same
prompt_group_id. A seed without a group is an example by itself, and its ID is the seed ID.Behavior
selection_keyand filters. A bad or mismatched cursor returns 400.date_addedof all their members, newest first, then by example ID.modality,seed_type, andharm_category. Values of one filter use OR, and different filters use AND. Different members of one example can match different filters, and the API returns the complete example. Harm categories match complete values and ignore case.searchfinds literal text in text prompts and objectives, and ignores case.%,_,[, and\are not patterns. Simulated-conversation configurations are not searched; useseed_type=simulated_conversationto find them.preview_truncatedflag. It also has the seed types, modalities, piece and objective counts, harm categories, andhas_unlabeled_harm. Media members show only the file name. No list item shows an absolute path or a URL query string.membersis a list of domain seeds (SeedPrompt,SeedObjective, orSeedSimulatedConversation), objectives first, then by sequence. These are the same objects asget_seeds()returns, with IDs, group IDs, provenance, hashes,parameters, andconditions.Code
MemoryInterface.get_seed_examples_asyncandget_seed_example_asyncselect the example IDs for one page in SQL, then get the members of only those examples. SQLite and Azure SQL have their own harm-category predicates.DatasetServicemakes the summaries and previews. The REST models are inpyrit/backend/models/datasets.py.get_seeds()behavior does not change. This PR does not change the schema and adds no migrations.Tests and Documentation
Documentation: There is a new "Browse stored seeds" section in
doc/code/datasets/0_dataset.md.Tests:
tests/unit/memory/memory_interface/test_interface_seed_examples.py: order, tied timestamps, order when a filter matches only some members, complete groups, filter matches across members, named and unnamed datasets, literal search characters, search that ignores media and simulated-conversation JSON, and skipped unreadable seeds.tests/unit/backend/test_seed_example_routes.py: pages and cursors, bad requests, safe previews, full text in the detail when the preview is truncated, objective conditions, domain seed members, and empty pages.tests/integration/memory/test_seed_examples_azure_sql_integration.py: filters, literal search, and order on Azure SQL.Results:
uv run --no-sync pytest tests/unit/memory/memory_interface/test_interface_seed_examples.py tests/unit/backend/test_seed_example_routes.py tests/unit/memory/test_memory_models.py tests/unit/memory/test_migration.py -q: 180 passeduv run --no-sync pytest tests/unit/backend -q: 1995 passed, 5 skipped, 1 failed. The failed test,test_scenario_run_routes.py::TestResumeScenarioRunRoute::test_resume_has_no_get_preflight, passes when it runs alone. It does not use the code in this PR.pre-commit run --all-files: passed, exceptenforce_alembic_revision_immutability. Locally, that hook flags the deletion of this PR's own earlier migration, which is not onmain. The CI check compares againstorigin/main, so it does not see that file.