Skip to content

Sync planning scans the full pending change set before yielding; repeated Durable Object CPU-limit failures #225

Description

@lumpinif

Describe the bug

The sync coalescer performs work for the entire pending change set before yielding its first entry. A downstream block limit therefore does not bound this initial planning work.

We investigated repeated container-exec failures with @cloudflare/computer@0.4.1. Six invocations on the same SQLite-backed Durable Object reported exceededCpu at 32,500 ms and an instance reset. Each invocation included the same large batch of VFS path-resolution queries before failure.

Evidence boundary: the eager planning behavior is confirmed in source. The telemetry makes it a strong candidate contributor to these CPU failures, but we do not have per-stage CPU profiling or a standalone reproducer that proves it accounts for the entire CPU budget. We also cannot establish from these traces whether the shell command had started. A Durable Object reset is not evidence that the container restarted.

Source path

Incident package: @cloudflare/computer@0.4.1. We also checked the following upstream sources while preparing this report:

  • coalesce.ts: coalesceChanges calls db.all for the full touched-inode revision range, resolves all touched inodes through pathsOfMany, builds the candidate map, merges tombstones and sorts the candidates before the first yield.
  • paths.ts: pathsOfMany batches 512 inodes per recursive SQL query, but synchronously processes every batch before returning.

In the inspected 0.4.1 bundle, the exec path pushes changes before spawnShell. Its mode-selection probe and block entry/byte limits consume the coalescer after this eager work. Smaller output blocks therefore do not by themselves bound the work before the first block.

Observed behavior

Sanitized aggregate measurements from six failed invocations:

Measurement Observed value
Invocation active CPU 32,500 ms each
Invocation wall time Approximately 37.7–40.2 seconds
Path-resolution query batches 449 per invocation
Billed rows read for those path batches Approximately 10.34 million per invocation
Earlier touched-range query Approximately 229,500 billed rows read

The SQL templates were matched to coalesceChanges → pathsOfMany → collectBatch. The large enumeration repeated on subsequent attempts. Other lightweight operations on the same object could succeed.

rows_read is a platform billing metric, not a unique-node count or a CPU measurement. We did not recover the incident cursor values and are not claiming cursor corruption. We have omitted workspace contents, paths, object IDs and private trace identifiers.

Expected behavior

Preparing the first sync block should have a bounded work budget that can make durable progress on a large change set under the Durable Object CPU limit. Reconnection should not force unbounded planning to restart before any progress can be acknowledged.

Reducing the wire block size should either bound the preparatory work too, or the API should clearly expose the separate planning cost and a supported way to bound it.

Steps to reproduce / current reproduction limits

The observed trigger was a container-backed workspace with a large pending change set, followed by a minimal shell-command request through workspace.runtime.exec. Repeated attempts encountered the measurements above.

We do not yet have a portable, self-contained reproduction of the exact deployed CPU failure. The source-level eager-work boundary can be isolated as follows:

  1. Use the existing real SQLite test harness to create a large nested tree and retain a cursor preceding those writes.
  2. Instrument database calls, then request only the first entry from coalesceChanges(db, cursor).
  3. Observe that the full touched range and all path-resolution batches are processed before that first entry is returned.
  4. For a platform reproduction, prepare equivalent data across bounded setup requests in a disposable SQLite Durable Object with the container backend, then issue a minimal exec and record platform CPU plus sync-planning/spawn stage markers.

These are suggested isolation steps, not a claim that this exact standalone fixture has already reproduced the CPU reset. Tree depth, pending revision range and peer cursor state should be reported with any reproduction.

Environment

  • Incident SDK and computerd image version: 0.4.1.
  • Cloudflare Workers, SQLite-backed Durable Object, container execution backend.
  • Deployed compatibility date: 2026-07-15.
  • No explicit limits.cpu_ms override.
  • Source inspection during this report: upstream commit 61f5aeae56e8e7ed75dddf235a9e8983508a0b39. We have not reproduced the platform failure on that newer commit.
  • The consumer had local patches for stream decoding, cancellation and read-handle behavior; the incident patch did not alter sync planning or schema.

Possible direction

Could the existing sync planner bound candidate reading, path reconstruction and sorting before the first output block, using its existing cursor/target mechanism? Any change needs to preserve same-revision path ordering, hard links, tombstones, concurrent changes and lost-acknowledgement recovery; simply limiting the initial SQL result could silently skip changes.

Related: #67 concerns retained tombstones. This report concerns eager planning over the current pending live-node set as well, so pruning old tombstones alone would not address the boundary described here.

Submitting an issue as requested by CONTRIBUTING.md. Maintainer guidance on a supported bounded-planning approach would be appreciated.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions