Repository navigation
Fix the iOS Maestro E2E suite on Linux runners against remote simulators - #3
Merged
Merged
Conversation
A remote simulator has no local GUI; on Linux `open -a simulator` fell through the sim-remote open shim (which only no-ops on *Simulator.app*) to xdg-open, which rejects -a and failed the run.
Three of four E2E jobs failed with 'no simulator is available right now': they all raced for a fleet machine and only one was free. The router also clamps a single acquire wait to 300s server-side, so the --timeout 900 was silently truncated. Log in without acquiring, then retry 'acquire --timeout 300' up to 8 times, and run at most one fleet-backed job at a time (max-parallel: 1 per matrix, plus templateapp sequenced after rntester).
helpers/ holds flow fragments that other flows pull in with runFlow. They have no launchApp and assume the app is already on a given screen, so running them standalone fails on leftover state — and after MAX_ATTEMPTS the whole suite aborts, skipping the flows that sort after helpers/.
The 'softu' release is a rolling tag and was rebuilt mid-run from a tree without install-shims, which surfaced only later as an opaque 'unrecognized subcommand' during the test step.
The ScrollView maintainVisibleContentPosition example sits far down the Components list; remote-simulator latency pushes the scroll past Maestro's default timeout, so pin it to the 90s that worked before. The fleet can serve two machines, so drop the sequencing between the rntester and templateapp suites — each still takes one machine at a time via max-parallel: 1 on its matrix.
jwajgelt
marked this pull request as ready for review
August 28, 2026 05:13
stopVideoRecording sent SIGINT and returned immediately. The flows run in a synchronous loop, so the event loop never turned and libuv could not collect the exited children — one zombie accumulated per flow (37 over a full rntester suite). Returning early also raced the recorder's finalization, which is why video artifacts sometimes came out missing. Await the child's exit (SIGKILL after 30s so a hung recorder cannot stall the suite), and make the flow-walking callers async to match. Verified with a stubbed A/B run of the script: 6 flows leave 4 zombies before, 0 after.
Fleet capacity is being increased, so the flavors no longer need to be serialized. All four jobs now start together; a job that cannot get a machine immediately still rides the acquire retry rather than failing.
The slow part is shipping the whole accessibility tree over the wire on each scroll step, not scroll latency.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary:
Follow-up to #2, which moved the iOS Maestro E2E tests onto Linux runners driving remote fleet simulators via
sim-remote. That path had never actually completed a run — this PR fixes everything that stood between it and a green suite.The whole workflow now passes end to end (run 33160192651,
completed/success, no re-runs):No macOS runner is involved any more except for building the apps.
What changed
2f2a187— skip simulator foregrounding for remote simulators.maestro-ios.jsranopen -a simulator, which fell through the sim-remoteopenshim (it only no-ops on*Simulator.app*, Maestro's own auto-boot call) toxdg-open, which rejects-a. A remote simulator has no local GUI to foreground, so the step is skipped whenSIM_REMOTE_BINis set.75c1a42— skip Maestro helper fragments when walking the flow directory. The runner recurses into.maestro/and executedhelpers/search.ymlas a standalone test. That fragment has nolaunchApp; it only passes if a previous flow left the app foregrounded. On macOS that state survives betweenmaestro testinvocations, which is why upstream is unaffected — on the remote fleet it does not, and the flow failed against the iOS home screen. Helper directories are now skipped. No coverage is lost: every real flow either starts withlaunchAppor delegates tohelpers/launch-app-and-search.yml, which does, andsearch.yml's steps still run as part of every flow that includes it.040bac1— retry fleet acquisition. Four jobs raced for fleet machines and three died withno simulator is available right now. The router also clamps a single acquire wait to 300s server-side (DEFAULT_ACQUIRE_MAX_WAIT_SECS), so a longer--timeoutis silently truncated. Now:login --no-acquire, then retryacquire --timeout 300up to 8 times. (This commit also serialized the flavors;8ccc27ebelow lifts that again.)04247fb— fail fast when sim-remote lacks the Maestro shims. Thesofturelease is a rolling tag and was rebuilt mid-run from a tree withoutinstall-shims, surfacing only later as an opaqueunrecognized subcommandduring the test step. A preflight check now fails immediately with the reason.572b987— raise thescrollUntilVisibletimeout.ScrollViewMaintainVisibleContentPositionExamplesits far down the Components list, and per-interaction latency against a remote simulator pushed the scroll past Maestro's default timeout (upstream passes the same flow in 56s; ours failed at 58s). Pinned to 90s in bothscrollview-minindex-maintainvisibleandscrollview-threshold-maintainvisible— both now pass in ~2m.fe7be20— reap the video recorder instead of leaking a zombie per flow.stopVideoRecordingsent SIGINT and returned immediately. The flows run in a synchronous loop, so the event loop never turned and libuv could not collect the exited children — one zombie accumulated per flow (37 over a full rntester suite). Returning early also raced the recorder's finalization, which is why video artifacts came out missing (No files were found with the provided path: video_record_1.mov). Now the child's exit is awaited, with a 30s SIGKILL escalation so a hung recorder cannot stall the suite, and the flow-walking callers areasyncto match. Verified by an A/B run of the script against stubs: 6 flows leave 4 zombies before, 0 after — and in CI all three video artifacts now upload (10.4 MB / 6.1 MB / 5.9 MB).8ccc27e— run every flavor concurrently. Fleet capacity is being increased, so the flavors are no longer serialized; all four E2E jobs start together. A job that cannot get a machine immediately still rides the acquire retry rather than failing.Notes for reviewers
SIM_ROUTER_USERNAMEActions variable andSIM_ROUTER_API_KEYsecret (SIM_ROUTER_URLoptional — a default is baked into the binary).timeout:value, not an assertion.04247fbassumes thesofturolling tag carriesinstall-shims(software-mansion/radon-cloud#143, still unmerged). Once that lands, pinning to an immutable release would be better than a mutable tag that can change under a running job.The runner has received a shutdown signalor a bare SIGTERM, always with zero failing flows, and always passing on re-run. A memory leak was ruled out by instrumenting a full 38-flow suite on a Linux VM: memory is flat (355 MB at start, 390 MB after 38 flows, oscillating with each Maestro JVM). This looks like GitHub-side runner reclamation. If it becomes disruptive, sharding the 38 flows across 2–3 jobs (~25 min each) is the durable fix.Changelog:
[INTERNAL] [FIXED] - Fix the iOS Maestro E2E suite running on Linux runners against remote simulators
Test Plan:
Dispatched the
E2E Maestroworkflow on this branch: run 33160192651 —completed/successwith no re-runs.[Passed], 0[Failed]each, ending inTest finished;Report statusis skipped in both, confirming the E2E step's outcome was success rather than being masked bycontinue-on-error.scrollview-minindex-maintainvisible(2m 4s),scrollview-threshold-maintainvisible(2m 11s), andtext(1m 47s) — the last of which no earlier run had reached.572b987(33083393983, 33144746540) also had all four E2E jobs pass.8ccc27e(concurrency) landed after the green run above. It only deletes twomax-parallel: 1lines, but a dispatch would confirm it and test four simultaneous fleet acquires.