Skip to content

test(playground): proxy the Studio suite to isolated local apps with Endform - #2

Open
ostenbom wants to merge 8 commits into
codex/mastra-playwright-baselinefrom
codex/mastra-endform
Open

ostenbom wants to merge 8 commits into
codex/mastra-playwright-baselinefrom
codex/mastra-endform

Conversation

@ostenbom

@ostenbom ostenbom commented Oct 1, 2026 •

Copy link
Copy Markdown

Run the existing Studio browser suite once on Endform against applications on the CLI host. The runner-local application prototype is replaced with a local pool: build the supported mastra dev entrypoint once, lease an isolated application port per attempt, and proxy browser/API/SSE traffic back through Endform. Remote workers receive test code and fixtures, without a Mastra server or Studio asset deployment.

Ordinary successful cases can reuse a warm application after the native storage-reset endpoint. Failed, filesystem, workflow, agent-builder and CMS-agent cases restart the process and restore source/database state. Each app serves only one attempt at a time. The effective Endform limit is capped at application capacity, with streaming capped at 12. Native Playwright remains available; all 335 Chromium cases, 35 existing skips, original assertions and CI retries of two are preserved.

The control server allocates/reset leases only. Application traffic goes directly to its assigned local port through proxyNetworkHosts: ['<loopback>']. ARIA snapshots and the playground tsconfig are transferred for runtime reads and alias resolution. GitHub authentication remains job-scoped OIDC, with no API key added.

Completed experiments

Setup Test-stage wall time Passed / flaky / failed / skipped
Native passing baseline, four CI shards 10m17s 297 / 3 / 0 / 35
Initial direct proxy, 8 apps, four-core CI 21m41s 236 / 25 / 39 / 35
Restart every attempt, 16 apps, Mac about 11m11s 212 / 27 / 61 / 35
Restart every attempt, 32 apps, Mac 11m47s 156 passed including retries / 144 failed / 35 skipped
Warm reuse + HTTP-only, 32 apps, Mac 4m54s 241 / 28 / 31 / 35
Final default: 8 warm apps, four-core CI 18m40s 242 / 21 / 37 / 35
32 warm apps, four-core CI 11m55s 184 / 33 / 83 / 35

The local warm run is a provisional timing result, not a passing sub-five-minute migration. The CI target remains unmet: warm32 took 11m55s for tests and 18m54s for the workflow; warm8 took 18m40s and 26m09s. The conservative CI default stays at eight apps because 32 raised failures to 83. Failed assertions mean some journeys were not fully exercised. It is one sample, only 5.73 seconds under the target. Mac runs use 10 logical CPUs, 32 GiB and Node 25.8.1; GitHub uses four cores, approximately 15 GiB reported memory and Node 24.18.0. Local timing is not equivalent CI proof. Some jobs overlap and may compete for Endform organization capacity.

Increasing local capacity alone made things worse: the strict 32-app run failed 144 cases, used 346.8 Endform runner minutes, and averaged about 7.4 seconds per app reset. The 16-app run averaged about 2.4 seconds. Warm reuse reduced reset overhead to about one second and the local run used 135.3 runner minutes. Transport, helper-origin handling and reuse changed together, so the full improvement cannot be attributed to one variable.

A native Playwright check against the same pool passed eight of nine selected journeys: history cases took 1.8–2.3 seconds and text streaming 9.9 seconds. The filesystem version-count check failed with expected 1/actual 2. The generated-entrypoint pool lacks the complete native CLI watcher, so this does not establish an upstream product bug. A first native attempt was blocked by missing Chromium; it was installed and rerun, and its launch errors are excluded from product results.

Asset-cache and raw-TCP-only probes did not establish an improvement and were removed. The final setup uses local applications with HTTP proxying. Warm32 CI also reported at least one application-reset HTTP 500 (viewer-role.spec.ts:354), whose underlying cause is unresolved. Remaining failures involve missing UI elements, history, scorer/workflow journey timeouts and filesystem behavior; they remain visible and were not quarantined or weakened.

The next capacity experiment would use roughly 10–16 CPU cores and 32 GiB memory with 32 warm apps. This is inferred from the Mac result, not a CI guarantee. The fork has no configured self-hosted runners, and the larger-hosted-runner API reports that hosted runners are not supported for the organization. No settings, billing or runner provisioning were changed.

Evidence

  • Native baseline commit acd0ca604383413038e22ba87dfa6088d879a878: passing rerun.
  • Initial direct local proxy commit c7c31d24e9d3342f415453198acaa3b6282bc620: completed CI, dashboard.
  • Local strict 16 apps: dashboard.
  • Local strict 32 apps: dashboard.
  • Final tested head: fee14a1aafcf413201cd35e08b8d05c705fa7ee9. Completed default warm8 PR CI, dashboard; completed manual warm32 CI, dashboard.
  • Local warm 32 apps: dashboard; /usr/bin/time -p measured 294.27 seconds around the entire command, including preparation and collection. This local run started with uncommitted warm/proxy changes while HEAD still recorded c7c31d24e9; those changes became fee14a1aaf during preparation, with reset instrumentation changes. CI above verifies the exact committed implementation.

Validation: exact 335-case discovery; all 57 spec bodies/assertions compared to baseline; playground and targeted fixture/config typechecks; formatting and diff checks; real database/source isolation, CLI package metadata and warm-process reuse checks. The local commit hook could not find corepack; manual checks passed. CLI 0.81.3, config dependency 0.81.1, pnpm 11.21.0. No changeset for this private package.

This remains stacked on baseline PR #1 in the fork. Leave both unmerged; the historical baseline and October 1 Endform experiments remain reproducible. The earlier criticism of the .mastra project-root value was incorrect: tracing the CLI call shows that value matches its current startup convention. Package metadata and generated-entrypoint differences were real.

Keep four isolated application shards and cap each Endform run at one concurrent test because fixture cleanup clears shared storage. Preserve suite coverage and retries, transfer ARIA snapshots, and retain failed-attempt traces.
@ostenbom ostenbom changed the title ci: run the existing Studio browser suite on Endform test(playground): run the full Studio suite once on Endform with isolated apps Oct 1, 2026
@ostenbom ostenbom changed the title test(playground): run the full Studio suite once on Endform with isolated apps test(playground): proxy the Studio suite to isolated local apps with Endform Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant