Repository navigation
Report Durable Object failures without secrets instead of an opaque 500 - #36
Draft
RhysSullivan wants to merge 5 commits into
Draft
RhysSullivan wants to merge 5 commits into
RhysSullivan wants to merge 5 commits into
Conversation
RhysSullivan
added this pull request to stack #38
October 8, 2026 04:35
The Worker no longer retries a failed stub call: a retryable flag does not prove the object never ran the request, and replaying a reset or an OAuth authorize changes state twice. Reports carry an instance hash, the route template, the error class and Cloudflare's flags, never a message, path or stack. The core router and control plane pass flagged platform failures through, and the Durable Object rethrows unexpected router errors so they are reported the same way. Invocation logs, which record request URLs, are off.
…port Neither boundary lets an error escape to Cloudflare's exception logging. The Durable Object reports Cloudflare's flagged failures itself instead of rethrowing them, so core no longer passes them through its router. Report fields come from fixed lists: error class names, HTTP methods, registered services and declared route templates. Seed and credential requests answer 400 only for typed rejections; other errors go to the error handler.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Executor's CI loses about one e2e job in 20 on busy days to an HTTP 500 from emulators.dev on
/_emulate/credentials,/_emulate/resetor/_emulate/seed. From 30 Sep to 8 Oct there were 11 such failures; 9 of them were on 7 Oct, and 3 of 250 main CI runs went red from them. The client gets Cloudflare's opaque "Worker threw exception" page, so neither side could say what failed.emulate-hostshas no Workers Logs.Cause
Cloudflare analytics for
emulate-hostsover the same window show two kinds of failure:stub.fetchfrom the Worker to the instance's Durable Object throws after about 12 s, and the Durable Object records no invocation at all. Every one of the 11 CI failures falls in one of these minutes (chart below). They come from LAX, SJC and PHX; 7 Oct 20:38 to 23:26 was all PHX, which is where GitHub's runners enter Cloudflare. The rest of the Worker's traffic in that window succeeded (PHX: about 10k requests). This is a transient fault on Cloudflare's side between the Worker and the object. Without logs its exact message is not knowable yet.scriptThrewException(52, including 36 in one minute on one Autumn instance on 6 Oct) andexceededMemory(3, one Resend instance). None of them coincides with a CI failure.The Worker didn't catch either kind, so both reached the client as a bare 500.
Solution
diagnostics.ts) and never throw:durable-object.ts). Its own exceptions becomeemulator_error(500). Cloudflare's flagged failures, such ascannot access storage because object has moved to a different machine, becomeemulator_unavailablewithretryable/overloadedset and a 503. They used to be rethrown so the Worker could see their flags, and that rethrow is what Cloudflare logged. The Worker now passes the object's report through unchanged.worker.ts). Reading the body,idFromNameandstub.fetchare all inside one catch. A failure there becomesemulator_unavailable: 503 when flagged, 500 otherwise (for example aremotememory kill). Any other failure in the Worker's router becomesworker_error(500).retryableflag does not prove the object never ran the request. Replaying a reset changes Google's signing key and PostHog's token and erases a seed written in between. Replaying WorkOS'sGET /oauth2/authorizeissues a second code. Each failure is answered once. A retry would need request-id dedupe inside the object first.error:emulator_unavailable,emulator_errororworker_error;service: a key of the service registry, elseunknown;instanceId: the first 12 hex digits of SHA-256(service:instance), ornullfor an unknown service. Generated instance names end in 96 random bits, so their hashes can't be enumerated. The hash does not hide a predictable or low-entropy name, such as one a caller chose or a legacy fixed name: hashing guesses confirms it.method: one of the 7 standard methods, elseOTHER;route: a template the service's router declares (such as/domains/:id) or one of the object's own control paths, elseunmatchedorunknown;errorClass: an allowlisted name (the built-in JS error classes and the runtime's DOMException names), elseother;retryable,overloadedandremoteflags;ray, when thecf-rayheader has a ray's format.400 invalid_seed/400 unsupportedonly for a typedControlPlaneRejection, which the emulators throw for their own refusals (an unsupported credential type, an invalid PlanetScale client). Any other error goes to the app's error handler. On Cloudflare that means the redacted 500 report; before, a storage failure came back as a 400 with its raw message.createServergainsrethrowUnexpectedErrors. The Durable Object sets it, so a route error without an HTTP status reaches the object's report instead of a 500 that carries the raw message. API errors such as 404s are answered as before.emulate. The core router no longer has a flagged-error passthrough, so local routing is unchanged. One local change remains: an unexpected (untyped) error from a seed or credential request is now a 500 from the error handler, not a 400. Its message is still shown locally.console.errorwith the invocation's URL, whether or not Workers Logs is on (Cloudflare: "Issues detection does not require Workers Logs or tracing"). For an instance, that URL contains the instance name. These settings are not versioned, sowrangler rollbackleaves them as they are. So:wrangler.jsoncsetsobservability.enabled,logs,tracesandissuesall tofalse. That makes each deploy also turn off anything enabled from the dashboard. Production hasobservability: nulltoday, so this deploy changes no setting.The failure path in the Worker and the Durable Object (
worker.ts,durable-object.ts,diagnostics.ts) writes nothing to the console. Bundled provider handlers still can: the GitHub and Vercel OAuth authorize routesconsole.warnthe caller'sredirect_urion a mismatch (github/src/routes/oauth.ts:112,vercel/src/routes/oauth.ts:69), and core'sdebug()logs whenDEBUGis set, which it isn't in the Worker. With telemetry off, Cloudflare stores none of these lines. Those handlers are local-dev diagnostics shared with the CLI, so this PR leaves them alone.Each failure is written once, by whichever boundary answered it, to the
emulate_hosts_failuresAnalytics Engine dataset (FAILURESbinding). A data point stores only what is written to it. Every row reads back with all 20 blob and 20 double slots, plusindex1,timestamp,datasetand_sample_interval. The slots a point doesn't write read back as''and0. The written slots are:service;error,service,method,route,errorClass,ray;status,retryable,overloaded,remote.The instance hash is left out, and the ray joins a data point to the client's copy of the report.
Analytics Engine samples writes, and samples again at query time, so
SUM(_sample_interval)is an estimate, and no single record is guaranteed. Rows are kept for 3 months. Reading them through SQL needs an account-scoped token with Account Analytics Read.A Worker-side
emulator_unavailablereport says the call to the object failed. It does not prove the object never ran the request, which is also why there is no retry.Not changed: why Cloudflare sometimes cannot reach the object. That is outside this code; the logs will show its flags the next time it happens. Deploying this stack is a production change to emulators.dev and has not been done.
Runtime probe (workerd)
src/__tests__/runtime.test.tsin #37 bundles the real Worker and Durable Object, runs them in Miniflare/workerd, and captures every tail event (the source of Workers Logs, exception events included) plus workerd's raw stdout and stderr. It injects 7 faults, each with a synthetic secret in the error message, the errornameand the instance URL:/_emulate/seedand/_emulate/credentials;idFromNamethrowing;outcome: exceptionname: "SYNTHETICtoken482913", message with the full instance URL, full stackThe probe waits for each fault's exact set of tail events before checking: a Worker event, plus an object event when the fault is inside the object (11 in all). It checks the complete events, with nothing stripped. The instance name appears in exactly three fields of the invocations' requests: the client's
event.request.url, and thex-emulator-instance/x-emulator-base-urlheaders. Our Worker sets those headers on its request to the object; they are not platform metadata. Cloudflare can store an invocation's request with each telemetry record. The URL and these headers stay out of Cloudflare's stores only because telemetry is off.Saved-record proof (isolated deployment, synthetic data)
A throwaway Worker,
d040-probe-r3on workers.dev, ran this code with the same fault injection. Everything was read back through the Telemetry API (real-time-issuesandcloudflare-workersdatasets), the Issues API (/workers/observability/issuesand each issue's occurrences) and Analytics Engine SQL. The Worker has since been deleted. Each run used a fresh synthetic instance and sent 7 faults plus 2 ordinary requests.error-log, 4http-status)invocation.url/pathin every Worker occurrenceobservability: null''or0invocation.urlwith the instance path, for botherror-log(console.error) andhttp-status(a handled 5xx). Query strings andx-emulator-*headers were not kept; the request headers kept werecf-ipcountry,cf-ray,content-type,content-lengthanduser-agent.wrangler deployreturned, while that version was still rolling out.Post-deploy checks
These checks roll back only on evidence of disclosure. Ingestion is reported as advice. They claim only what the data can decide:
_sample_intervalis a statistical weight that names no request (sampling). So a ray with a row is recorded, and a ray without one is not found: compatible with sampling, completeness unverified. Nearby rows never confirm or excuse it.SUM(_sample_interval)is reported as an estimate.The scripts are in the D-040 findings data, outside this repo:
checkpoint.py,positive_control.sh,report_body.py,harness_failures.py,ae_check_sql.pyandae_query.sh. Cloudflare reads need a token with Workers Scripts Read and Account Analytics Read.Deploy.
GET /accounts/{account}/workers/scripts/emulate-hosts/settingsreadsobservability: null, and 87d64cf7 is at 100%.packages/@emulators/cloudflarewithnpx wrangler@4.148.0 deploy. Wrangler 4.134 or later is needed, because older versions don't send theissueskey.wrangler deployreturned.At each checkpoint (T0, 24 h, 7 days):
Settings.
GET …/settings, saved asS.json. It gives the observability settings and the bindings.Positive control.
positive_control.sh COMMIT C.json 3 120(see below).Smoke checks, at T0 only. Use a fresh instance with a synthetic name and synthetic secrets.
checkpoint.pyrequires one result for each of these names (SMOKE_CHECKS):create-instance;credentials-api-keyandcredentials-oauth-client(/_emulate/credentials);seed(/_emulate/seed);resend-domain-create, thenresend-domain-read:POST /domainswith a synthetic name, thenGET /domains/{id}, which returns 200;reset(/_emulate/reset), thenreset-cleared: the sameGET /domains/{id}returns 404. Reset reapplies the instance's seed, so only state created through the API shows that reset ran;github-providerandgoogle-provider: a provider route on each;oauth-authorize,oauth-consentandoauth-token;workos-authorize-ledger: after one WorkOSGET /oauth2/authorize, the instance ledger shows exactly one entry. The ledger skips/_emulaterequests and reset clears it, so it isn't checked for reset;openid-configuration(/.well-known/openid-configuration);unsupported-credential: 400 with its message;unknown-serviceandunmatched-path: 404.Record each result in
SMOKE.json: the expected status, then the first run's status, whether it was readable, and whether the body was right (bodyOk). Opaque means a 5xx without emulate's report, or no response. If a readable run is wrong, repeat that check on a second fresh instance and record it asrerun; otherwise setrerunto null.report_body.py check BODY --routes routes-2cedc9b.json --forbid SECRETS, whereSECRETSlists the synthetic names and secrets. Save each result asB<n>.json. (2xx bodies return the instance's own data by design.)alljob.Harness, at 24 h and 7 days. Run
harness_failures.py ZIPDIR T0 H.jsonover the CI, Cloud tests on main and Deploy job logs since T0. It reads the executor-next#2199 lines.Queries. Run these at least W = 15 min after the last control:
ae_check_sql.py sqlwritesviolations.sql(the allowlist query, below) androutes.sql.checkpoint.py sql --dataset emulate_hosts_failures --deploy T0 Q/ C*.json H.json, given every control file so far, writes:lookup.sql: one export of the rows whose ray is an observed ray;natural.sql: the rows without any control's ray, withSUM(_sample_interval).Rays are compared by their 16-hex id on both sides (
substring(blob6, 1, 16)). The Worker storescf-rayas it arrives, with or without the colo suffix.Run each query with
ae_query.sh.Decide. Run
checkpoint.py check --name T0|24h|7d --now … --settings S.json --violations V.out --routes routes-2cedc9b.json --routes-result R.out --lookup L.out --natural N.out --control C.json [--harness H.json] [--body-check B<n>.json …] [--smoke SMOKE.json](--smokeis required at T0). It refuses to run before the last control is 15 minutes old, so a checkpoint is never left pending.Rollback triggers (only these)
Telemetry on. The settings read-back has any
trueunderobservability: logs, invocation logs, traces or Issues.A stored row outside the allowlist. The violations query returns a row, or
ae_check_sql.py routesrejects a (service, route) pair or a ray. This query must return 0 rows:AE SQL has no regular expressions and rejects queries over 10,000 characters. So the exact route (one of the 1,051 templates the routers declare at 2cedc9b, or
/__seed,/__reset,/__token,unmatchedorunknown) and the ray format (^[0-9a-f]{16}(-[A-Z]{3})?$) are checked byae_check_sql.py routes. Regenerate the template list if the deployed commit isn't 2cedc9b.A failure report with a field it shouldn't have. A report received by a control or a smoke check has an
erroroutside its three values, a key outside its 10 fields, a value outside its list or format, or a synthetic secret or instance name. A body counts as a report by its structure: a JSON object whoseerroris one of the three, or that has at least 5 of the 10 report keys. So a changederrorcan't hide a report. Secrets are searched in the raw bytes and in every decoded JSON key and string, so a JSON escape can't hide one. Reports are this PR's surface. A secret in any other body, such as an API error unchanged by this PR, is marked investigate, not a trigger.A reproducible smoke-check regression. A T0 smoke check returns the wrong status or body with a readable (non-opaque) response, and the same check is wrong again, readably, on a second fresh instance. Emulators broken for every e2e job is a regression, not an ingestion question.
checkpoint.pydecides it fromSMOKE.json. A failure that passes on the rerun, a failure that wasn't rerun, an opaque answer on either run, and a missing or malformed result all make the checkpoint unverified.Never a trigger:
The CI harness (#2199) logs only values it vouches for, so it can't show trigger 3. A logged value outside the lists would be a #2199 bug.
The positive control
d040-control, generated suffix) and posts{"domains":1}to/_emulate/seedthree times, 2 minutes apart.seedFromConfigiteratesdomainsand throws aTypeErrorbefore any insert. That gives a 500emulator_errorand one row, with nothing stored. In workerd, the unmodified Worker and Durable Object gave exactly that response and row. On production today (87d64cf7) the same request is a 400.ok: 500emulator_error,resend,POST,/_emulate/seed,TypeError, no flags, the instance's own hash, and a cf-ray header and a report ray with the same 16-hex id;unavailable: 400invalid_seed;opaque: a 5xx without a report, or no response within the 30 s curl deadline. This is inconclusive: the script retries it at most twice, 30 s apart. If it persists, check Workers analytics for that minute (exceptions and Durable Object invocations).invalid: anything else;disclosure: the report failed the body check.ok.Decision table
At each checkpoint:
• a safety read or query didn't run;
• the control batch isn't complete (unavailable, opaque after retries, invalid, or a slot missing);
• a control has no row with its exact ray, or its row has other fields;
• the lookup failed;
• the binding isn't listed;
• a body is marked investigate;
• a smoke check is missing, failed once but passed on a second fresh instance, failed without a rerun, or had an opaque response.
• A read that didn't run: rerun it.
• Binding missing: redeploy this commit.
• Opaque control: check Workers analytics for that minute.
• Control unavailable: check the deployments list and commit.
• Smoke-check failure that didn't reproduce: record it and escalate to Rhys.
At 7 days,
checkpoint.py finalgives the verdict. Pass it every checkpoint result, reruns included. A checkpoint given more than once keeps its worst outcome and the reasons from every run, so a later pass can't erase an earlier trigger or unverified result.It also reports these estimates:
SUM(_sample_interval)without the control rays.Compare with the baseline:
scriptThrewException: 52;exceededMemory: 3.The rows'
retryable/overloadedflags decide whether a dedupe-backed retry is worth building.Tested (synthetic data and read-only queries only):
-SJCrays on every side;error, with and without an extra key;SYNTHETIC-secret-1hidden by JSON escapes;lookup.sqlandnatural.sqlran against the probe's dataset. 17 of its 17 rows match by exact ray. The 4 failures with no row are not found: run B's rollout-window miss, run B2's two sampled points and a fake ray. The earlier reconciler counted B2's two points assampled.positive_control.shran against a local mock. It retried opaque slots, applied the 30 s deadline, stopped on a 400 and on each disclosure (an instance name in a field, the same name fully JSON-escaped, an out-of-listerror), refused another commit, and wrote no instance name.Why not change the code.
writeDataPointthrow can't be reached. A point has 1 index of at most 11 bytes, 6 blobs from fixed lists (the longest route is 87 bytes) and 4 doubles. The documented limits are 1 index of 96 bytes, 20 blobs totalling 16 KB, 20 doubles and 250 points per invocation.writeDataPointreturns nothing, so no counter in the Worker can see these. The binding is read directly in step 1, and Test redacted Durable Object failure reports and flags through the router #37 pins it inwrangler.jsonc.Rollback (drilled on the probe)
Observability settings are not part of a version:
wrangler rollbackto a version deployed with it off left logs and Issues on.So a rollback is two steps:
PATCH /accounts/{account}/workers/scripts/emulate-hosts/script-settingswith{"observability":{"enabled":false,"logs":{"enabled":false,"invocation_logs":false},"traces":{"enabled":false},"issues":{"enabled":false}}}. Read it back and confirmobservabilityisnullor disabled.wrangler rollback 87d64cf7-aee9-4ba4-b2ee-b20d0a0d2d26. This also removes theFAILURESbinding, since bindings are part of a version.Before / after
Before: the 11 CI failures carry no body; the harness shows
External emulator failed: /_emulate/reset (HTTP 500).After:
{"error":"emulator_unavailable","service":"resend","instanceId":"781f9f1e9db7","method":"GET","route":"/emails","errorClass":"Error","retryable":true,"overloaded":false,"remote":false,"ray":null}executor-next#2199 prints these fields. Tests are in #37.
How measured
Cloudflare GraphQL Analytics (
workersInvocationsAdaptiveanddurableObjectsInvocationsAdaptiveGroups, filtered toemulate-hosts, by minute, colo and object name), 30 Sep 00:00 to 8 Oct 05:00 UTC, compared with the failure timestamps in 748 executor-next CI job logs. The only deployment in the window is 87d64cf7 (29 Sep, this repo'smain).