Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .pre-commit-config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ repos:
- id: check-toml
- id: detect-private-key
- repo: https://github.com/astral-sh/ruff-pre-commit
rev: v0.13.0
rev: v0.16.9
hooks:
- id: ruff-check
args: [ --fix ]
Expand Down
30 changes: 17 additions & 13 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co

A FastAPI service that provides AI-powered explanations of compiler assembly output for the Compiler Explorer
website, using Anthropic's Claude API. Runs locally for development or as an AWS Lambda function (Mangum adapter)
behind an API Gateway HTTP API. The production explainer model is Sonnet 5 (see `app/prompt.yaml`); the
behind an API Gateway HTTP API. The production explainer model is Sonnet 5.5 (see `app/prompt.yaml`); the
prompt-testing framework's correctness reviewer is Opus 5.

Request pipeline: input validation, then smart assembly filtering plus hard character caps, then the Claude API
Expand Down Expand Up @@ -44,17 +44,20 @@ Pre-commit hooks may modify files (e.g. ruff format); re-`git add` if a hook rep
- **`max_tokens` includes thinking tokens.** When thinking is enabled it counts against `max_tokens`, and can
starve the visible text on complex cases. `Prompt.__init__` refuses to load a thinking-enabled config with
`max_tokens < 4096` (production uses 4096).
- **Neither production model accepts `temperature`.** Opus 5 (reviewer) and Sonnet 5 (explainer) reject
- **Neither production model accepts `temperature`.** Opus 5 (reviewer) and Sonnet 5.5 (explainer) reject
non-default sampling parameters with a 400, so neither sets one. Only pre-5 Sonnet models accept
`temperature`; restore it in the YAML if you ever pin one of those.
- **Sonnet 5 runs adaptive thinking by default when `thinking` is omitted.** `app/prompt.yaml` therefore sets
`thinking: {type: disabled}` explicitly; dropping that line silently turns thinking on and eats the
`max_tokens` budget. The same trap applies to the reviewer, which always sends an explicit thinking config so
`--reviewer-thinking off` really means off.
- **Sonnet 5's tokenizer produces ~30% more tokens than 4.6 for the same text.** Don't reuse token counts or
`temperature`; restore it in the YAML if you ever pin one of those. anthropic SDK 1.x has no `temperature`
kwarg, so `build_api_payload` sends a configured one via `extra_body`.
- **Sonnet 5.5 runs adaptive thinking by default when `thinking` is omitted, and 400s on `disabled`.**
`app/prompt.yaml` therefore sets `thinking: {type: between_tools}`, the lowest setting (no thinking when no
tools are sent; only valid at effort `high` or below). Dropping that line silently turns thinking on and eats
the `max_tokens` budget. Sonnet 5 and Opus 5 still take `disabled`; the reviewer (Opus 5) always sends an
explicit thinking config so `--reviewer-thinking off` really means off.
- **Sonnet 5/5.5's tokenizer produces ~30% more tokens than 4.6 for the same text.** Don't reuse token counts or
cost baselines measured on 4.6-era models.
- **`model.effort` is plumbed but a no-op with thinking disabled.** The 2026-07 sweep (low/medium/high, 21
cases) showed identical latency and cost across levels with thinking off: effort mostly modulates thinking
cases) showed identical latency and cost across levels with thinking off, and the 2026-09 Sonnet 5.5 eval
found medium and high indistinguishable under `between_tools`: effort mostly modulates thinking
depth, so there is nothing to modulate. Production leaves it unset (API default `high`). It becomes meaningful
on the `useThinking` path or if adaptive thinking is ever made the default.
- **Production explainer thinking is opt-in per request** (`useThinking: true`; default off). The 2026-07 eval
Expand All @@ -72,14 +75,15 @@ Pre-commit hooks may modify files (e.g. ruff format); re-`git add` if a hook rep
without rerunning that eval.
- **Prompt caching: evaluated 2026-07 and rejected at current traffic.** ~104 fresh Claude calls/day
(CloudWatch, 14-day window), only ~35 hours/fortnight above 12 calls/hour, against a 5-minute cache TTL and a
prefix fragmented by language/arch/audience/type. Generous math: ~$0.40 saved per fortnight of ~$22 spend,
before counting the restructuring needed to clear Sonnet 5's 1024-token minimum cacheable prefix (the system
prompt is only ~620 tokens; the per-audience guidance lives in the user prompt). Revisit if traffic grows
prefix fragmented by language/arch/audience/type. Generous math: ~$0.40 saved per fortnight of ~$22 spend.
(That analysis also counted restructuring to clear Sonnet 5's 1024-token minimum; Sonnet 5.5's is 512, which
the ~620-token system prompt already clears, but the savings math still doesn't pay.) Revisit if traffic grows
~50x, or if sustained >3 same-combo requests/hour makes the 1-hour TTL viable. Rerun the analysis with
`aws cloudwatch get-metric-statistics` on `CompilerExplorer/ClaudeExplainFreshResponse`.
- **Safety refusals are handled before the empty-response path.** Claude 5-family classifiers can decline a
request (HTTP 200 with `stop_reason: "refusal"`, empty or partial content), plausible here since CE users
compile arbitrary, sometimes exploit-adjacent code. `app/explain.py` returns a distinct user-facing message,
compile arbitrary, sometimes exploit-adjacent code. Sonnet 5.5 declines in more categories than Sonnet 5
(adds `bio`, `reasoning_extraction`, `general_harms`); watch the metric after model bumps. `app/explain.py` returns a distinct user-facing message,
discards any partial output, and emits `ClaudeExplainRefusal`.
- **Multi-block responses.** With thinking enabled the API returns thinking blocks before the text block; both
`app/explain.py` and `prompt_testing/runner.py` pick the last text block via
Expand Down
2 changes: 1 addition & 1 deletion app/explain.py
Original file line number Diff line number Diff line change
Expand Up @@ -138,7 +138,7 @@ async def _call_anthropic_api(
LOGGER.info(
"Using Anthropic client with model: %s (thinking=%s)",
prompt_data["model"],
bool(prompt_data.get("thinking")),
(prompt_data.get("thinking") or {}).get("type"),
)
# Bound the call to a wall-clock budget below the API Gateway HTTP API
# integration timeout (a hard 30s ceiling). Without this, a slow generation
Expand Down
6 changes: 3 additions & 3 deletions app/explanation_types.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,17 +5,17 @@
prompt configuration and accessed via the Prompt class.
"""

from enum import Enum
from enum import StrEnum


class AudienceLevel(str, Enum):
class AudienceLevel(StrEnum):
"""Target audience for the explanation."""

BEGINNER = "beginner"
EXPERIENCED = "experienced"


class ExplanationType(str, Enum):
class ExplanationType(StrEnum):
"""Type of explanation to generate."""

ASSEMBLY = "assembly"
Expand Down
12 changes: 6 additions & 6 deletions app/model_costs.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@ class ModelCost(NamedTuple):


# Model family costs in USD per million tokens
# Updated: 2026-05-06 based on https://platform.claude.com/docs/en/about-claude/pricing
# Updated: 2026-09-29 based on https://platform.claude.com/docs/en/about-claude/pricing
#
# Notes:
# - Opus 4.5+ moved to a new lower price tier ($5/$25) and now bundle the 1M
Expand All @@ -30,11 +30,11 @@ class ModelCost(NamedTuple):
# from the lookup; the regex normaliser still parses their names so callers
# get a clear "not found" error rather than a parse failure.
MODEL_FAMILIES = {
# Claude 5 family
# Sonnet 5 sticker price is $3/$15; an introductory $2/$10 applies through
# 2026-08-31. We use the sticker price so estimates don't silently rot
# when the intro window closes.
"sonnet-5": ModelCost(3.0, 15.0),
# Claude 5 family. Sonnet 5's launch $2/$10 became its standard price;
# the planned 2026-09-01 rise to $3/$15 was cancelled.
"sonnet-5.5": ModelCost(2.0, 10.0),
"sonnet-5": ModelCost(2.0, 10.0),
"opus-5.5": ModelCost(4.0, 20.0),
"opus-5": ModelCost(5.0, 25.0),
# Opus 4.5+: new pricing tier, 1M context bundled
"opus-4.8": ModelCost(5.0, 25.0),
Expand Down
13 changes: 7 additions & 6 deletions app/prompt.py
Original file line number Diff line number Diff line change
Expand Up @@ -48,10 +48,10 @@ def __init__(self, config: dict[str, Any] | Path):
# Extract model configuration
self.model = self.config["model"]["name"]
self.max_tokens = self.config["model"]["max_tokens"]
self.temperature = self.config["model"].get("temperature", 0.0)
# Optional extended-thinking config, e.g. {"type": "adaptive"} or
# {"type": "enabled", "budget_tokens": 2000}. When set, callers
# should drop `temperature` (the API requires it to be unset/1).
# Only pre-5 models accept temperature; unset means don't send one.
self.temperature = self.config["model"].get("temperature")
# Optional thinking config, e.g. {"type": "adaptive"} or
# {"type": "between_tools"}. When set, `temperature` is never sent.
self.thinking = self.config["model"].get("thinking")
# Optional effort level ("low" | "medium" | "high" | "xhigh" | "max").
# Controls the model's reasoning/token spend; unset means the API
Expand Down Expand Up @@ -322,8 +322,9 @@ def build_api_payload(self, request: ExplainRequest) -> dict[str, Any]:
# rejects temperature when thinking is set). Floor max_tokens
# so adaptive thinking can't starve the visible text.
payload["max_tokens"] = max(payload["max_tokens"], MIN_MAX_TOKENS_WITH_THINKING)
else:
payload["temperature"] = base["temperature"]
elif base["temperature"] is not None:
# anthropic 1.x dropped the `temperature` kwarg; the API still takes it on older models.
payload["extra_body"] = {"temperature": base["temperature"]}
return payload

def generate_messages(self, request: ExplainRequest) -> dict[str, Any]:
Expand Down
16 changes: 8 additions & 8 deletions app/prompt.yaml
Original file line number Diff line number Diff line change
@@ -1,14 +1,14 @@
name: Production Sonnet 5
description: Sonnet 5 with conciseness tuning; thinking off by default (per-request opt-in unchanged)
name: Production Sonnet 5.5
description: Sonnet 5.5 with conciseness tuning; no up-front thinking by default (per-request opt-in unchanged)
model:
name: claude-sonnet-5
# Sonnet 5 runs adaptive thinking by default when `thinking` is omitted,
# which would starve a small max_tokens budget; keep it explicitly disabled
# (the useThinking request flag still switches to adaptive per-request).
# Sonnet 5 also rejects non-default temperature, so none is set here.
name: claude-sonnet-5-5
# Sonnet 5.5 runs adaptive thinking when `thinking` is omitted and rejects
# `disabled` with a 400; `between_tools` is its lowest setting and, with no
# tools, means no thinking. The useThinking request flag still switches to
# adaptive per-request. No temperature: Claude 5 models reject it.
max_tokens: 4096
thinking:
type: disabled
type: between_tools
audience_levels:
beginner:
description: For beginners learning assembly language. Uses simple language and explains technical terms.
Expand Down
20 changes: 20 additions & 0 deletions app/test_explain.py
Original file line number Diff line number Diff line change
Expand Up @@ -223,6 +223,26 @@ def test_no_effort_means_no_output_config(self, sample_request):
payload = Prompt(Path("app/prompt.yaml")).build_api_payload(sample_request)
assert "output_config" not in payload

def test_temperature_sent_via_extra_body_without_thinking(self, sample_request):
"""anthropic 1.x has no `temperature` kwarg; pre-5 models still get it via extra_body."""
yaml = YAML(typ="safe")
with Path("app/prompt.yaml").open(encoding="utf-8") as f:
config = yaml.load(f)
del config["model"]["thinking"]
config["model"]["temperature"] = 0.2
payload = Prompt(config).build_api_payload(sample_request)
assert "temperature" not in payload
assert payload["extra_body"] == {"temperature": 0.2}

def test_no_temperature_unless_configured(self, sample_request):
yaml = YAML(typ="safe")
with Path("app/prompt.yaml").open(encoding="utf-8") as f:
config = yaml.load(f)
del config["model"]["thinking"]
payload = Prompt(config).build_api_payload(sample_request)
assert "temperature" not in payload
assert "extra_body" not in payload

def test_invalid_effort_rejected_at_load(self):
"""A typo'd effort level fails loudly at config load, not at request time."""
config = {
Expand Down
10 changes: 10 additions & 0 deletions app/test_model_costs.py
Original file line number Diff line number Diff line change
Expand Up @@ -83,6 +83,16 @@ def test_sonnet_4_6_cost(self):
assert input_cost == 3.0 / 1_000_000 # $3 per million
assert output_cost == 15.0 / 1_000_000 # $15 per million

def test_sonnet_5_cost(self):
input_cost, output_cost = get_model_cost("claude-sonnet-5")
assert input_cost == 2.0 / 1_000_000
assert output_cost == 10.0 / 1_000_000

def test_sonnet_5_5_cost(self):
input_cost, output_cost = get_model_cost("claude-sonnet-5-5")
assert input_cost == 2.0 / 1_000_000
assert output_cost == 10.0 / 1_000_000

def test_opus_4_cost(self):
"""Test Claude 4 Opus costs (legacy $15/$75 pricing)."""
input_cost, output_cost = get_model_cost("claude-opus-4-0")
Expand Down
24 changes: 12 additions & 12 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -5,27 +5,27 @@ description = "Explain Compiler Explorer output using AI"
readme = "README.md"
requires-python = ">=3.13"
dependencies = [
"anthropic>=0.100.0",
"anthropic>=1.9.0",
"aws-embedded-metrics>=3.5.0",
"boto3>=1.43.4",
"click>=8.3.3",
"fastapi>=0.136.1",
"boto3>=1.43.104",
"click>=8.5.0",
"fastapi>=0.141.1",
"humanfriendly>=10.0",
"mangum>=0.21.0",
"pydantic-settings>=2.14.0",
"python-dotenv>=1.2.2",
"requests>=2.33.1",
"mangum>=0.22.0",
"pydantic-settings>=2.15.0",
"python-dotenv>=1.2.3",
"requests>=2.34.2",
"ruamel.yaml>=0.19.1",
]

[dependency-groups]
dev = [
# In development mode, include the FastAPI development server.
"fastapi[standard]>=0.141.1",
"pre-commit>=4.6.0",
"pytest>=9.0.3",
"pytest-asyncio>=1.3.0",
"ruff>=0.15.12",
"pre-commit>=4.6.2",
"pytest>=9.1.1",
"pytest-asyncio>=1.4.0",
"ruff>=0.16.9",
]

[tool.ruff]
Expand Down
Loading
Loading