Skip to content

Fix the iOS Maestro E2E suite on Linux runners against remote simulators - #3

Merged
jwajgelt merged 8 commits into
mainfrom
jwajgelt/skip-foreground-remote-sim
Sep 2, 2026
Merged

jwajgelt merged 8 commits into
mainfrom
jwajgelt/skip-foreground-remote-sim

Conversation

@jwajgelt

@jwajgelt jwajgelt commented Aug 27, 2026 •

Copy link
Copy Markdown
Member

Summary:

Follow-up to #2, which moved the iOS Maestro E2E tests onto Linux runners driving remote fleet simulators via sim-remote. That path had never actually completed a run — this PR fixes everything that stood between it and a green suite.

The whole workflow now passes end to end (run 33160192651, completed/success, no re-runs):

Suite Debug Release
rntester 38 flows, 0 failures 38 flows, 0 failures
templateapp passed passed

No macOS runner is involved any more except for building the apps.

What changed

2f2a187 — skip simulator foregrounding for remote simulators. maestro-ios.js ran open -a simulator, which fell through the sim-remote open shim (it only no-ops on *Simulator.app*, Maestro's own auto-boot call) to xdg-open, which rejects -a. A remote simulator has no local GUI to foreground, so the step is skipped when SIM_REMOTE_BIN is set.

75c1a42 — skip Maestro helper fragments when walking the flow directory. The runner recurses into .maestro/ and executed helpers/search.yml as a standalone test. That fragment has no launchApp; it only passes if a previous flow left the app foregrounded. On macOS that state survives between maestro test invocations, which is why upstream is unaffected — on the remote fleet it does not, and the flow failed against the iOS home screen. Helper directories are now skipped. No coverage is lost: every real flow either starts with launchApp or delegates to helpers/launch-app-and-search.yml, which does, and search.yml's steps still run as part of every flow that includes it.

040bac1 — retry fleet acquisition. Four jobs raced for fleet machines and three died with no simulator is available right now. The router also clamps a single acquire wait to 300s server-side (DEFAULT_ACQUIRE_MAX_WAIT_SECS), so a longer --timeout is silently truncated. Now: login --no-acquire, then retry acquire --timeout 300 up to 8 times. (This commit also serialized the flavors; 8ccc27e below lifts that again.)

04247fb — fail fast when sim-remote lacks the Maestro shims. The softu release is a rolling tag and was rebuilt mid-run from a tree without install-shims, surfacing only later as an opaque unrecognized subcommand during the test step. A preflight check now fails immediately with the reason.

572b987 — raise the scrollUntilVisible timeout. ScrollViewMaintainVisibleContentPositionExample sits far down the Components list, and per-interaction latency against a remote simulator pushed the scroll past Maestro's default timeout (upstream passes the same flow in 56s; ours failed at 58s). Pinned to 90s in both scrollview-minindex-maintainvisible and scrollview-threshold-maintainvisible — both now pass in ~2m.

fe7be20 — reap the video recorder instead of leaking a zombie per flow. stopVideoRecording sent SIGINT and returned immediately. The flows run in a synchronous loop, so the event loop never turned and libuv could not collect the exited children — one zombie accumulated per flow (37 over a full rntester suite). Returning early also raced the recorder's finalization, which is why video artifacts came out missing (No files were found with the provided path: video_record_1.mov). Now the child's exit is awaited, with a 30s SIGKILL escalation so a hung recorder cannot stall the suite, and the flow-walking callers are async to match. Verified by an A/B run of the script against stubs: 6 flows leave 4 zombies before, 0 after — and in CI all three video artifacts now upload (10.4 MB / 6.1 MB / 5.9 MB).

8ccc27e — run every flavor concurrently. Fleet capacity is being increased, so the flavors are no longer serialized; all four E2E jobs start together. A job that cannot get a machine immediately still rides the acquire retry rather than failing.

Notes for reviewers

  • Requires the SIM_ROUTER_USERNAME Actions variable and SIM_ROUTER_API_KEY secret (SIM_ROUTER_URL optional — a default is baked into the binary).
  • A dispatch now wants four fleet machines at once. The acquire retry is what keeps that safe if capacity is short — jobs wait rather than fail.
  • The only test files touched are the two scrollview flows, and the change is a timeout: value, not an assertion.
  • 04247fb assumes the softu rolling tag carries install-shims (software-mansion/radon-cloud#143, still unmerged). Once that lands, pinning to an immutable release would be better than a mutable tag that can change under a running job.
  • In Debug, the 30s recorder shutdown timeout fires on ~5 of 38 flows (bigger recordings, Metro over the reverse tunnel); those videos are truncated and each costs 30s. Release hits it 0 times. Raising it to ~60s would capture them if the videos matter.
  • Unrelated flakiness, not introduced here: long rntester jobs were killed mid-run three times across earlier runs (~27, ~27, ~35 min) with The runner has received a shutdown signal or a bare SIGTERM, always with zero failing flows, and always passing on re-run. A memory leak was ruled out by instrumenting a full 38-flow suite on a Linux VM: memory is flat (355 MB at start, 390 MB after 38 flows, oscillating with each Maestro JVM). This looks like GitHub-side runner reclamation. If it becomes disruptive, sharding the 38 flows across 2–3 jobs (~25 min each) is the durable fix.

Changelog:

[INTERNAL] [FIXED] - Fix the iOS Maestro E2E suite running on Linux runners against remote simulators

Test Plan:

Dispatched the E2E Maestro workflow on this branch: run 33160192651 — completed/success with no re-runs.

  • rntester Debug and Release: 38 [Passed], 0 [Failed] each, ending in Test finished; Report status is skipped in both, confirming the E2E step's outcome was success rather than being masked by continue-on-error.
  • templateapp Debug and Release: both pass; both report jobs green.
  • Previously failing flows all pass: scrollview-minindex-maintainvisible (2m 4s), scrollview-threshold-maintainvisible (2m 11s), and text (1m 47s) — the last of which no earlier run had reached.
  • Two earlier runs on 572b987 (33083393983, 33144746540) also had all four E2E jobs pass.
  • Not yet exercised in CI: 8ccc27e (concurrency) landed after the green run above. It only deletes two max-parallel: 1 lines, but a dispatch would confirm it and test four simultaneous fleet acquires.

A remote simulator has no local GUI; on Linux `open -a simulator`
fell through the sim-remote open shim (which only no-ops on
*Simulator.app*) to xdg-open, which rejects -a and failed the run.
Three of four E2E jobs failed with 'no simulator is available right
now': they all raced for a fleet machine and only one was free. The
router also clamps a single acquire wait to 300s server-side, so the
--timeout 900 was silently truncated.

Log in without acquiring, then retry 'acquire --timeout 300' up to 8
times, and run at most one fleet-backed job at a time (max-parallel: 1
per matrix, plus templateapp sequenced after rntester).
helpers/ holds flow fragments that other flows pull in with runFlow.
They have no launchApp and assume the app is already on a given
screen, so running them standalone fails on leftover state — and
after MAX_ATTEMPTS the whole suite aborts, skipping the flows that
sort after helpers/.
The 'softu' release is a rolling tag and was rebuilt mid-run from a
tree without install-shims, which surfaced only later as an opaque
'unrecognized subcommand' during the test step.
The ScrollView maintainVisibleContentPosition example sits far down the
Components list; remote-simulator latency pushes the scroll past
Maestro's default timeout, so pin it to the 90s that worked before.

The fleet can serve two machines, so drop the sequencing between the
rntester and templateapp suites — each still takes one machine at a
time via max-parallel: 1 on its matrix.
@jwajgelt
jwajgelt marked this pull request as ready for review August 28, 2026 05:13
@jwajgelt jwajgelt changed the title Skip simulator foregrounding when driving a remote simulator Fix the iOS Maestro E2E suite on Linux runners against remote simulators Aug 28, 2026
stopVideoRecording sent SIGINT and returned immediately. The flows run
in a synchronous loop, so the event loop never turned and libuv could
not collect the exited children — one zombie accumulated per flow (37
over a full rntester suite). Returning early also raced the recorder's
finalization, which is why video artifacts sometimes came out missing.

Await the child's exit (SIGKILL after 30s so a hung recorder cannot
stall the suite), and make the flow-walking callers async to match.

Verified with a stubbed A/B run of the script: 6 flows leave 4 zombies
before, 0 after.
Fleet capacity is being increased, so the flavors no longer need to be
serialized. All four jobs now start together; a job that cannot get a
machine immediately still rides the acquire retry rather than failing.
The slow part is shipping the whole accessibility tree over the wire on
each scroll step, not scroll latency.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant