Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/security-regressions.yml
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@ jobs:
python-version: '3.12'
- run: python -m pip install -r backend/requirements-ai.txt pillow numpy uv pip-audit
- run: python -m pip check
- run: python -m pytest backend/tests/test_security_regressions.py backend/tests/test_biomedparse_demo.py backend/tests/test_clinical_platform.py backend/tests/test_ai_live.py backend/tests/test_ai_text.py backend/tests/test_ai_codex.py backend/tests/test_ai_research.py backend/tests/test_ai_providers.py backend/tests/test_ai_openai.py backend/tests/test_ai_connection_races.py backend/tests/test_ai_screen_awareness.py backend/tests/test_ai_credentials.py backend/tests/test_ai_research_settings.py backend/tests/test_ai_evidence_review.py backend/tests/test_ai_evidence_routes.py backend/tests/test_ai_evidence_provenance.py backend/tests/evidence_review -q
- run: python -m pytest backend/tests/test_security_regressions.py backend/tests/test_biomedparse_demo.py backend/tests/test_clinical_platform.py backend/tests/test_ai_live.py backend/tests/test_ai_text.py backend/tests/test_ai_codex.py backend/tests/test_ai_codex_exploration.py backend/tests/test_ai_actions.py backend/tests/test_ai_exploration_contracts.py backend/tests/test_ai_exploration_coverage.py backend/tests/test_ai_exploration_routes.py backend/tests/test_ai_exploration_lifecycle.py backend/tests/test_ai_exploration_tools.py backend/tests/test_ai_research.py backend/tests/test_ai_providers.py backend/tests/test_ai_openai.py backend/tests/test_ai_connection_races.py backend/tests/test_ai_screen_awareness.py backend/tests/test_ai_credentials.py backend/tests/test_ai_research_settings.py backend/tests/test_ai_evidence_review.py backend/tests/test_ai_evidence_routes.py backend/tests/test_ai_evidence_provenance.py backend/tests/evidence_review -q
- name: Resolve and audit research dependencies
run: |
uv pip compile --python 3.12 backend/requirements.txt -o /tmp/radsysx-research.txt
Expand Down
4 changes: 4 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -350,3 +350,7 @@ The standalone [evidence-review runbook](backend/evidence_review/README.md) docu

- The desktop sidebar supports a separately managed ChatGPT/Codex subscription login for typed chat and public PubMed research. Read `roadmap/ai-backend/CODEX_SUBSCRIPTION.md` for its account, execution and validation boundaries. Realtime voice still uses its own API-key billing. Gemini/NVIDIA retain DeepAgents/LangGraph; subscription execution uses pinned Codex App Server. Never reuse subscription tokens as API keys or read/copy the user's existing Codex credentials.
- Root npm pins `@openai/codex` 0.154.0. Desktop enables `RADSYSX_CODEX_ENABLED`; other backends default off. Private `.ai-codex/` directories and OS-keyring entries are per RadSysX actor and must never be committed. Clinical mode remains disabled.

## Scoped study exploration

- Codex text/research supports explicit current-image, reading-view and entire-series sharing, plus separately permitted native viewer tools. It does not require a Realtime connection. Backend grants, renderer claims, acknowledged frame coverage, reversible action receipts and Stop/Take over own this path. See `roadmap/ai-backend/CODEX_STUDY_EXPLORATION.md`. Do not equate image delivery with diagnosis or claim unverified specialized modality parity.
8 changes: 7 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -561,4 +561,10 @@ NVIDIA NIM is available for explicit evidence evaluation and opt-in PubMed resea

### Attach an image to Codex Chat or Research

With a ChatGPT/Codex model selected, confirm synthetic/deidentified data and choose **Attach current view**. Expand **Preview image** to inspect the exact snapshot and visible measurement overlays, then type your question and choose **Send** or **Research**. The selected model must advertise image input; `gpt-6-astra` was verified on 2026-09-23. This works without Realtime. Each request receives only that captured viewport, not the whole series or continuous screen access. Attach again for another view. Saved image receipts record what was submitted; pixels are not saved in RadSysX history. See the [subscription runbook](roadmap/ai-backend/CODEX_SUBSCRIPTION.md).
With a ChatGPT/Codex model selected, confirm synthetic/deidentified data and choose **Current image** under **Share images with AI** (or **Attach current view** in a client without study sharing). Expand **Preview image** to inspect the exact snapshot and visible measurement overlays, then type your question and choose **Send** or **Research**. The selected model must advertise image input; single-image input on `gpt-6-astra` was verified on 2026-09-23. This works without Realtime. This option shares only that captured viewport. Saved image receipts record what was submitted; pixels are not saved in RadSysX history. See the [subscription runbook](roadmap/ai-backend/CODEX_SUBSCRIPTION.md).

### Share a study with your Codex model

In the AI sidebar, select your ChatGPT/Codex model in Settings, confirm synthetic/deidentified data, then open **Share images with AI**. Choose the active image, the reading view, or the entire active series. **Allow viewer tools** separately permits native navigation and reversible edits. Prepare the scope and send your question with Send or Research; a voice connection is optional.

The study card shows acknowledged frame delivery, actions, Stop and Take over. Partial series coverage remains visible and requires explicit Continue review. Saving a report still requires review. See the [implementation and acceptance record](roadmap/ai-backend/CODEX_STUDY_EXPLORATION.md).
21 changes: 20 additions & 1 deletion backend/clinical/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -98,7 +98,26 @@

- `ai_codex.py` and `ai_codex_routes.py` own local per-actor Codex App Server stdio, browser login, status, model discovery and bounded text/PubMed execution. Require signed ai.run, enabled pilot/research, explicit write Origin, strict empty write bodies and private no-store errors. No generic RPC or browser tokens. Account mutations stop owned jobs. Clinical logout closes its process/pending login; subscription Sign out also clears its isolated keyring login.
- Workspace-pinned Codex 0.154.0 receives a private 0700 database-adjacent `.ai-codex/<actor-hash>` home, 077 umask, allowlisted environment and keyring-only forced ChatGPT auth. Never inherit API keys or the user's Codex config/auth. Fail closed without private storage/keyring.
- Every ephemeral thread/turn has no execution environments; shell, filesystem-image tools, plugins, hooks, memory and subagents are disabled. Verify exact model/provider, empty loaded instructions and read-only/network-disabled sandbox before sending text or an explicit inline image. Only bounded public PubMed dynamic calls are handled; deny all other requests. Eight tool calls, owned job/expiry deadline, final text only and ledger-backed citations. Unconfirmed cancellation terminates the child, including lost turn-start acknowledgements. Do not replay.
- Every ephemeral thread/turn has no execution environments; shell, filesystem-image tools, plugins, hooks, memory and subagents are disabled. Verify exact model/provider, empty loaded instructions and read-only/network-disabled sandbox before sending text or an explicit inline image. Only bounded public PubMed calls and explicitly granted study tools are handled; deny all other requests. Eight tool calls, owned job/expiry deadline, final text only and ledger-backed citations. Unconfirmed cancellation terminates the child, including lost turn-start acknowledgements. Do not replay.
- The codex research preference is authenticated catalog-backed and independent of voice. Gemini/NVIDIA research remains DeepAgents/LangGraph. Subscription uses the Codex harness; never mislabel its orchestration or bill it as an API key.

- `ai_view_image.py` owns explicit Codex viewport input: strict JPEG/base64, at most 512 KiB decoded, 768 px per edge, matching dimensions/target/context and capture age up to five minutes. Require the exact authenticated model to advertise image input before creating a job. Forward an inline App Server `image` data URL; never enable filesystem image, screen/computer or other autonomous tools. Pixels stay in memory for this one request and are omitted from persisted tool args, history and logs. Save only scope/dimensions/time/hash/model submission receipts. Prior chat images are historical text context and their pixels are never replayed. Keep visual observations separate from PubMed claims, and public queries free of identifiers. Voice session routes reject typed image attachments; clinical mode remains disabled.

## Scoped study exploration foundations

- `ai_exploration_contracts.py` and `ai_exploration_coverage.py` define strict observation records, bounded task budgets and acknowledged distinct-frame accounting. Raw JPEG data is excluded from serialization and repr; validate base64, hash, dimensions and full Pillow decoding before forwarding. The private `/sessions/{sessionId}/explorations` routes use these records for owned preparation, command polling/claiming and completion.
- A thumbnail or overview never establishes full-series delivery. Transport submission remains unconfirmed until the matching tool acknowledgment; repeated frame deliveries consume budget without increasing distinct coverage. Continuation restores safe coverage metadata, never image bytes or executable pending operations.

- `ai_actions.py` owns transport-independent tool validation, immutable call identity, reviewed decisions, report-save permission/study checks and actor-scoped mutation locks. Live uses this broker for native edits and saves. A granted text exploration can exclusively own the mutation channel while voice continues conversation/research. Existing voice proposals keep their five-minute lifetime; new exploration proposals use two minutes. Lost delivery of a committed completion cannot rewrite that action as failed or unknown.

- `ai_exploration.py`, its repository and routes own voice-independent renderer grants: prepared scopes last 60 seconds, active scopes at most ten minutes, one task per owner/two globally, and a 15-second heartbeat/claim deadline. Bind every command to a renderer epoch, context and revision; history reads never execute. Metadata and manifest pages are capped at 64 KiB; only observation results accept 8 MiB encoded images plus 64 KiB metadata. Manifest pages use the same claim channel with a dedicated metadata-only completion route. Strict decisions reject extra fields. Recovered tasks are interrupted; claimed mutations without completion remain unknown. Context/account/model changes, Stop, takeover and deletion revoke work. Persist only metadata and receipts; completed image futures release their pixel references and late cancellation cannot recreate deleted rows.

- Native reading tool declarations in `ai_tools.py` use explicit finite numeric bounds and fixed enums for orientation, cine, synchronization, rendering and panels. Capability summaries retain only known semantic names/availability, never renderer-provided schemas or executable native command IDs. Renderer post-state is independently sanitized.

- Measurement/segmentation/region schemas forbid extra fields, nonfinite/degenerate geometry, arbitrary tools and unknown coordinate spaces. New geometry operations carry frame/revision pairs; the exploration bridge additionally requires acknowledgment of that frame before mutation. Calibration requires broker approval for the exact physical reference. Native statistic summaries omit free text and use explicit calculation/geometry completeness states.

## Scoped Codex study tools

- `ai_codex_tools.py` connects owned study grants to pinned App Server dynamic tools. Every request binds thread, turn and call IDs; unknown namespaces/general execution remain disabled. At most four request handlers, 64 dynamic calls, eight PubMed searches, 128 image deliveries and ten minutes per exploration. Stdio reads/writes are bounded to 12 MiB; writes serialize and time out within 15 seconds. An uncertain process is terminated without replay.
- Exact `dynamicToolCall` completion acknowledges matching submitted observations. Persist requested/captured/unconfirmed/delivered coverage separately; duplicate image calls return receipt-only `pixelsUnavailable`, and repeated mutations never execute again. Geometry requires an acknowledged current pane/frame/revision. Unknown mutations revoke mutation permission immediately.
- Text turns may carry `explorationId` or the legacy single `image`, never both. Preparation inventories the explicitly shared series; Send activates it. Completed/failed/cancelled runs retain metadata only. Continue creates a new explicit run; model prose cannot establish complete coverage. PubMed and Jev remain separate public-evidence workflows.
97 changes: 97 additions & 0 deletions backend/clinical/ai_actions.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,97 @@
"""Transport-independent authority for reviewed, idempotent native viewer actions."""
from __future__ import annotations
import asyncio
import re
from fastapi import HTTPException
from .ai_repository import TERMINAL_TOOLS
from .ai_tools import validate_tool, requires_approval, bound_json, safe_state
from .contracts import ReportDraftRequest, parse_iso_z, utc_now


class ActionBroker:
def __init__(self, live):
self.live, self.repo = live, live.repository
self.locks = {}
self.mutation_owners = {}

def lock(self, actor):
return self.locks.setdefault(actor.sub, asyncio.Lock())

def claim_viewer(self, actor, grant_id):
if self.lock(actor).locked() or self.mutation_owners.get(actor.sub) not in (None, grant_id):
raise HTTPException(409, 'Another task is changing this viewer.')
self.mutation_owners[actor.sub] = grant_id

def release_viewer(self, actor, grant_id):
if self.mutation_owners.get(actor.sub) == grant_id: self.mutation_owners.pop(actor.sub, None)

def require(self, session_id, actor, context_version=None):
self.live.require_research_settings(actor)
row = self.repo.owned(session_id, actor, active=True)
if not row['attestation'] or row['status'] in {'closed','interrupted'} or (context_version is not None and row['contextVersion'] != context_version):
raise HTTPException(409, 'AI context is no longer authorized.')
return row

def prepare(self, session_id, call_id, name, args, actor, *, context_version, grant=None):
row = self.require(session_id, actor, context_version)
if not isinstance(call_id,str) or not re.fullmatch(r'[A-Za-z0-9_.:-]{1,128}',call_id):
raise ValueError('Invalid tool call identity')
args = validate_tool(name, args)
bound_json(args)
tool, fresh = self.repo.add_tool(session_id,call_id,name,args,context_version,requires_approval(name,args),
approval_seconds=120 if grant else 300)
if not fresh and (tool['contextVersion'] != context_version or tool['name'] != name or tool['args'] != args):
raise HTTPException(409, 'Tool call identity belongs to another request.')
if fresh: self.live.audit(row, actor, f'{session_id}:{call_id}')
return tool, fresh

def decide(self, session_id, call_id, decision, actor, *, grant=None):
self.require(session_id, actor, decision.context_version)
tool = self.repo.tool(session_id,call_id)
if tool['contextVersion'] != decision.context_version or tool['status'] != 'awaiting_approval' or parse_iso_z(tool['expiresAt']) <= utc_now():
raise HTTPException(409, 'Approval expired or belongs to another context.')
return self.repo.set_tool(session_id,call_id,'pending' if decision.approved else 'denied')

async def execute(self, session_id, call_id, actor, *, check, dispatch, grant=None):
self.require(session_id, actor)
async with self.lock(actor):
tool = self.repo.tool(session_id,call_id)
if tool['status'] in TERMINAL_TOOLS: return tool['result'] or {'status':tool['status']}
if tool['status'] == 'awaiting_approval': raise HTTPException(409,'Review this exact proposal before execution.')
if tool['status'] not in {'pending','running'}: raise HTTPException(409,'Action unavailable.')
dispatched = False
try:
row = check()
self.require(session_id,actor,tool['contextVersion'])
if self.mutation_owners.get(actor.sub) not in (None, grant):
raise HTTPException(409,'Another task owns viewer tools. Take over or finish it first.')
self.repo.set_tool(session_id,call_id,'running')
if tool['name'] == 'report_save':
if 'report.write' not in actor.scopes: raise HTTPException(403,'Report write permission required.')
uid = row['viewerContext'].get('studyInstanceUID')
if not uid or not self.live.clinical_repository.get_worklist_row(uid):
raise HTTPException(409,'Import or associate this local study through the worklist before saving a report.')
record = self.live.clinical.save_report(ReportDraftRequest(studyInstanceUID=uid,
findingsSummary=tool['args']['findings'], impression=tool['args']['impression']),actor=actor,source_ip='ai-sidebar')
result = {'reportId':record.report_id,'status':'draft_saved'}
else:
dispatched = True
raw = bound_json(await dispatch(call_id,tool['name'],tool['args']),32768)
check()
result = safe_state(raw)
if 'result' in raw: result['result'] = safe_state(raw['result'])
if raw.get('error'): result = {'status':'failed','message':'The native tool could not complete.'}
current = self.repo.tool(session_id,call_id)
if current['status'] in TERMINAL_TOOLS: return current['result'] or {'status':current['status']}
status = 'failed' if result.get('status') == 'failed' else 'outcome_unknown' if result.get('status') == 'outcome_unknown' else 'completed'
self.repo.set_tool(session_id,call_id,status,result)
return result
except asyncio.CancelledError:
self.repo.set_tool(session_id,call_id,'outcome_unknown' if dispatched else 'cancelled')
raise
except Exception as error:
status = 'outcome_unknown' if dispatched else 'failed'
message = 'No valid completion receipt. The action was not retried.' if dispatched else error.detail if isinstance(error,HTTPException) else 'The native tool could not complete.'
result = {'status':status,'message':message}
self.repo.set_tool(session_id,call_id,status,result)
return result
Loading
Loading