Skip to content
Open
16 changes: 16 additions & 0 deletions .pyrit_conf_example
Original file line number Diff line number Diff line change
Expand Up @@ -134,6 +134,22 @@ enable_live_reinitialization: false
# Default: false
allow_custom_initializers: false

# When true, the backend downloads http(s) media URLs that API callers ask it to
# import (`import_url` on message pieces and converter previews, and URLs given for
# converter file parameters) once into managed storage, with time, size, and redirect
# limits, and passes only the stored copy on. Media URLs that are not imported stay
# references. When false, import requests are rejected and media must be uploaded.
#
# Default: true
allow_media_url_import: true

# Directory that targets created through the backend API may upload local files from
# (HTTPXAPITarget). The server passes it to the target; API callers cannot choose it.
# When unset, such targets cannot be created through the API.
#
# Default: unset
# target_upload_directory: /path/to/upload/files

# Optional storage for custom initializer Python scripts. This may be a local
# directory or an Azure Blob container URI with an optional blob prefix.
# Container URIs may include a SAS; otherwise DefaultAzureCredential is used. Defaults to
Expand Down
46 changes: 46 additions & 0 deletions pyrit/backend/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -200,3 +200,49 @@ Environment variables:
- `PYRIT_API_HOST` - Host to bind to (default: localhost)
- `PYRIT_API_PORT` - Port to listen on (default: 8000)
- `PYRIT_API_RELOAD` - Enable auto-reload (default: false)

## Input Validation

The backend is the part of PyRIT that accepts requests from other machines, so it checks
request values before using them:

- Media values in messages, previews, and prepended conversations, and file parameters of
converters, must be uploaded content, a media URL, or a reference into this server's media
storage: the `prompt-memory-entries` and `seed-prompt-entries` folders under the memory
results path. Other file paths are rejected.
- Media URLs are kept as references by default, and `url` pieces always pass through
unchanged. Azure Blob URLs outside the configured results container are rejected as media
references, because the storage layer would read them from this server's own container;
blob URLs inside it are kept without their query string.
- A caller imports a media URL by setting `import_url` on a message piece or a converter
preview with an `image_path`, `audio_path`, `video_path`, or `binary_path` type. The server
downloads it once into managed storage (10 second connect, 30 second read, and 60 second
total limits, 100 MiB limit, at most 3 redirects, no request credentials forwarded) and
stores it under the declared type without format conversion. The format extension comes
from the response MIME type, the caller MIME type if the response is missing or generic,
or the URL suffix. Unknown formats use `.bin`, not a modality default such as `.wav`.
Converters and targets then see only the stored copy,
and the piece's prompt metadata records the source URL, without credentials or query
string, and the resolved content type. A preview that imports returns the stored copy and
that metadata, so sending them reuses the same bytes. Set `allow_media_url_import: false`
in `.pyrit_conf` to turn imports off. Converter file parameters given a URL are downloaded
the same way.
- Target types that load model code (`HuggingFaceChatTarget`) cannot be created through the
API, and target parameters that name server paths cannot be set through it; the target type
catalog leaves both out. Targets that upload local files (`HTTPXAPITarget`) can be created
through the API only when `target_upload_directory` is set in `.pyrit_conf`; the server passes
that directory to the target, which uploads files only from inside it. Register such targets
in Python or with an initializer for other settings.

Intentional exceptions:

- Target endpoints, raw HTTP requests, media URLs, and their redirects are chosen by the
operator and are not restricted to particular hosts. Limit outbound network access in the
deployment instead.
- Prompt content is not filtered. It is adversarial test data by design.
- Any file type can be stored as a payload. `GET /api/media` only renders known image,
audio, and video types inline; everything else downloads as a file.
- `GET /api/media` does not require authentication so the browser can load media. It only
serves files from the media folders above.
- Custom initializer scripts are trusted Python. Uploading them requires an administrator
and `allow_custom_initializers: true`.
26 changes: 26 additions & 0 deletions pyrit/backend/models/attacks.py
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,7 @@
TextStr,
)
from pyrit.models import (
MEDIA_PATH_DATA_TYPES,
AttackResult,
ChatMessageRole,
ConversationReference,
Expand Down Expand Up @@ -398,6 +399,31 @@ class MessagePieceRequest(BaseModel):
source_piece_id: uuid.UUID | None = Field(
None, description="Source piece for a complete conversation save; verified against its source conversation."
)
import_url: bool = Field(
False,
description="Download this piece's http(s) media values once into managed storage instead of keeping the "
"URLs as references. The stored copies keep the declared image_path, audio_path, video_path, or "
"binary_path type.",
)

@model_validator(mode="after")
def _validate_import_url(self) -> "MessagePieceRequest":
"""
Validate that a URL import names the media type to store.

Returns:
The validated request piece.

Raises:
ValueError: If import_url is set on a piece without a media path type.
"""
converted_type = self.converted_value_data_type or self.data_type
if self.import_url and not {self.data_type, converted_type} & MEDIA_PATH_DATA_TYPES:
raise ValueError(
"import_url needs an image_path, audio_path, video_path, or binary_path value; "
"declare the media type to store the URL as."
)
return self

@model_validator(mode="after")
def _validate_converted_value_data_type(self) -> "MessagePieceRequest":
Expand Down
36 changes: 33 additions & 3 deletions pyrit/backend/models/converters.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,10 +9,10 @@

from typing import Any

from pydantic import BaseModel, Field
from pydantic import BaseModel, Field, model_validator

from pyrit.backend.models.common import MAX_ITEMS, REGISTRY_INSTANCE_NAME_PATTERN, IdentifierStr
from pyrit.models import ConverterIdentifier, Parameter, PromptDataType
from pyrit.models import MEDIA_PATH_DATA_TYPES, ConverterIdentifier, Parameter, PromptDataType

__all__ = [
"ConverterInstance",
Expand Down Expand Up @@ -121,13 +121,43 @@ class ConverterPreviewRequest(BaseModel):
converter_ids: list[IdentifierStr] = Field(..., max_length=MAX_ITEMS, description="Converter instance IDs to apply")
start_token: str = Field(default="⟪", min_length=1, description="Opening marker for selected text regions")
end_token: str = Field(default="⟫", min_length=1, description="Closing marker for selected text regions")
import_url: bool = Field(
False,
description="Download an http(s) original_value once into managed storage instead of passing the URL on. "
"The stored copy keeps the declared image_path, audio_path, video_path, or binary_path type.",
)

@model_validator(mode="after")
def _validate_import_url(self) -> "ConverterPreviewRequest":
"""
Validate that a URL import names the media type to store.

Returns:
The validated preview request.

Raises:
ValueError: If import_url is set without a media path type.
"""
if self.import_url and self.original_value_data_type not in MEDIA_PATH_DATA_TYPES:
raise ValueError(
"import_url needs an image_path, audio_path, video_path, or binary_path original_value_data_type; "
"declare the media type to store the URL as."
)
return self


class ConverterPreviewResponse(BaseModel):
"""Response from converter preview."""

original_value: str = Field(..., description="Original input text")
original_value: str = Field(
..., description="Original input, or the stored copy that conversion used when an http(s) URL was imported"
)
original_value_data_type: PromptDataType = Field(..., description="Data type of original value")
converted_value: str = Field(..., description="Final converted text")
converted_value_data_type: PromptDataType = Field(..., description="Data type of converted value")
steps: list[PreviewStep] = Field(..., description="Step-by-step conversion results")
prompt_metadata: dict[str, str] = Field(
default_factory=dict,
description="Source information for an imported original value. Send it as the piece's prompt_metadata "
"with the stored copy to keep it in the conversation.",
)
7 changes: 3 additions & 4 deletions pyrit/backend/routes/attacks.py
Original file line number Diff line number Diff line change
Expand Up @@ -251,10 +251,9 @@ async def create_attack(request: CreateAttackRequest) -> CreateAttackResponse:
try:
return await service.create_attack_async(request=request)
except ValueError as e:
raise HTTPException(
status_code=status.HTTP_404_NOT_FOUND,
detail=str(e),
) from e
error_msg = str(e)
error_status = status.HTTP_404_NOT_FOUND if "not found" in error_msg.lower() else status.HTTP_400_BAD_REQUEST
raise HTTPException(status_code=error_status, detail=error_msg) from e


@router.get(
Expand Down
25 changes: 4 additions & 21 deletions pyrit/backend/routes/media.py
Original file line number Diff line number Diff line change
Expand Up @@ -23,15 +23,13 @@
from fastapi import APIRouter, HTTPException, Query
from fastapi.responses import FileResponse

from pyrit.backend.services.media_persistence import resolve_managed_media_path
from pyrit.memory import CentralMemory

logger = logging.getLogger(__name__)

router = APIRouter()

# Only serve files from known media subdirectories under results_path.
_ALLOWED_SUBDIRECTORIES = {"prompt-memory-entries", "seed-prompt-entries"}

# Only these known-safe media types render inline. Every other extension is
# served as an application/octet-stream attachment.
_INLINE_EXTENSIONS = {
Expand Down Expand Up @@ -62,12 +60,7 @@

def _validate_media_path(*, path: str, allowed_root: Path) -> Path:
"""
Validate and sanitize a user-provided file path against an allowed root directory.

Uses ``Path.resolve()`` to resolve symlinks and ``..`` components, then
verifies the canonical path is under the allowed root. This is the standard
sanitization pattern recognized by static analysis tools (e.g. CodeQL
``py/path-injection``).
Validate a user-provided file path against the allowed results directory.

Args:
path: The user-provided file path to validate.
Expand All @@ -79,20 +72,10 @@ def _validate_media_path(*, path: str, allowed_root: Path) -> Path:
Raises:
HTTPException 403: If the path fails any validation check.
"""
real_path = Path(path).resolve(strict=False)

try:
relative_parts = real_path.relative_to(allowed_root).parts
return resolve_managed_media_path(path=path, allowed_root=allowed_root)
except ValueError as exc:
raise HTTPException(
status_code=403, detail="Access denied: path is outside the allowed results directory."
) from exc

# Restrict to known media subdirectories (e.g. prompt-memory-entries/)
if not relative_parts or relative_parts[0] not in _ALLOWED_SUBDIRECTORIES:
raise HTTPException(status_code=403, detail="Access denied: path is not in a media subdirectory.")

return real_path
raise HTTPException(status_code=403, detail=f"Access denied: {exc}") from exc


@router.get("/media")
Expand Down
3 changes: 2 additions & 1 deletion pyrit/backend/services/attack_service.py
Original file line number Diff line number Diff line change
Expand Up @@ -54,7 +54,7 @@
)
from pyrit.backend.models.common import PaginationInfo
from pyrit.backend.models.message_sends import MessageSendRequest, MessageSendStatus
from pyrit.backend.services.media_persistence import persist_message_pieces_async
from pyrit.backend.services.media_persistence import media_source_entries, persist_message_pieces_async
from pyrit.backend.services.message_send_service import (
MessageSendService,
get_message_send_service,
Expand Down Expand Up @@ -787,6 +787,7 @@ async def _prepare_message_pieces_async(
saved.original_value = request_piece.original_value
converted_value = request_piece.converted_value
saved.converted_value = converted_value if converted_value is not None else saved.original_value
saved.prompt_metadata.update(media_source_entries(request_piece.prompt_metadata))
await set_message_piece_sha256_async(saved)
return [saved for saved, _ in prepared_pieces]

Expand Down
45 changes: 32 additions & 13 deletions pyrit/backend/services/converter_service.py
Original file line number Diff line number Diff line change
Expand Up @@ -37,8 +37,13 @@
CreateConverterRequest,
PreviewStep,
)
from pyrit.backend.services.media_persistence import persist_media_value_async
from pyrit.common.azure_storage import is_azure_blob_uri
from pyrit.backend.services.media_persistence import (
is_managed_blob_url,
media_source_metadata,
persist_media_value_async,
)
from pyrit.backend.services.media_url_import import download_media_url_async, media_extension
from pyrit.common.azure_storage import redact_url_credentials
from pyrit.memory import data_serializer_factory
from pyrit.models import MessagePiece, PromptDataType
from pyrit.prompt_normalizer import ConverterConfiguration, PromptNormalizer
Expand Down Expand Up @@ -216,15 +221,18 @@ async def preview_conversion_async(self, *, request: ConverterPreviewRequest) ->

For non-text data types (image_path, audio_path, etc.), persists base64 data
to a temporary file so converters can operate on file paths. Marked text
regions use the request's delimiter settings for every stage.
regions use the request's delimiter settings for every stage. When the
request imports an http(s) URL, the response returns the stored copy and
its source metadata, so a later send can reuse the same bytes.

Returns:
ConverterPreviewResponse with step-by-step conversion results.
"""
original_value = request.original_value
data_type = request.original_value_data_type
source_metadata: dict[str, str] = {}

# For path-based data types, resolve references or persist base64/data URIs.
# For path-based data types, resolve references, import URLs on request, or persist base64/data URIs.
if str(data_type).endswith("_path"):
result = await persist_media_value_async(
value=original_value,
Expand All @@ -234,9 +242,11 @@ async def preview_conversion_async(self, *, request: ConverterPreviewRequest) ->
# explicit/data-URI MIME metadata.
use_data_uri_mime_type=False,
require_valid_base64_after_path_error=True,
import_url=request.import_url,
serializer_factory=data_serializer_factory,
)
original_value = result.value
source_metadata = media_source_metadata(result)

converters = self._gather_converters(converter_ids=request.converter_ids)
steps, final_value, final_type = await self._apply_converters_async(
Expand All @@ -248,11 +258,12 @@ async def preview_conversion_async(self, *, request: ConverterPreviewRequest) ->
)

return ConverterPreviewResponse(
original_value=request.original_value,
original_value=original_value if source_metadata else request.original_value,
original_value_data_type=request.original_value_data_type,
converted_value=final_value,
converted_value_data_type=final_type,
steps=steps,
prompt_metadata=source_metadata,
)

def get_converter_objects_for_ids(self, *, converter_ids: list[str]) -> list[Any]:
Expand Down Expand Up @@ -289,8 +300,9 @@ async def _persist_data_uri_params_async(
directory this service owns, and the client never names a server path. Every
``Path`` parameter is handled the same way, so a converter opts in simply by
declaring the type; there is no per-converter or per-parameter table.
``Path | str`` parameters also accept Azure Blob URLs, which pass through
unchanged. Their data-URI uploads use the same local storage.
An http(s) URL is downloaded once into the same local storage. ``Path | str``
parameters also accept Azure Blob URLs inside this server's result storage,
which pass through unchanged.

Inputs remain local until converter deletion or backend shutdown, even with
Azure-backed memory. Converter outputs still use the configured result storage.
Expand All @@ -308,7 +320,8 @@ async def _persist_data_uri_params_async(
set of request-created files owned by the future registry entry.

Raises:
ValueError: If a ``Path`` value is not a valid data URI.
ValueError: If a ``Path`` value is not a data URI or an http(s) URL, or a URL
cannot be downloaded.
"""
metadata = self._registry.get_registered_class_metadata(converter_type)
path_params = (
Expand All @@ -330,13 +343,19 @@ async def _persist_data_uri_params_async(
if value is None:
continue
parameter = path_params[name]
if not isinstance(value, str) or not value.startswith("data:"):
if parameter.is_path_or_str and isinstance(value, str) and is_azure_blob_uri(value):
if isinstance(value, str) and value.startswith(("http://", "https://")):
if parameter.is_path_or_str and is_managed_blob_url(value):
result[name] = redact_url_credentials(value)
continue
alternative = " or supplied as an Azure Blob URL" if parameter.is_path_or_str else ""
raise ValueError(f"Path parameter '{name}' must be uploaded as a data URI{alternative}")
download = await download_media_url_async(url=value)
content, extension = download.content, media_extension(download, default="")
elif isinstance(value, str) and value.startswith("data:"):
content, extension = self._decode_data_uri(parameter_name=name, data_uri=value)
else:
raise ValueError(
f"Path parameter '{name}' must be uploaded as a data URI or given as an http(s) URL"
)

content, extension = self._decode_data_uri(parameter_name=name, data_uri=value)
file_path = self._upload_path / f"{uuid.uuid4().hex}{extension}"
async with aiofiles.open(file_path, "xb") as file:
owned_paths.append(file_path)
Expand Down
Loading
Loading