Repository navigation
Conversation
Contributor
|
@vzaidman has exported this pull request. If you are a Meta employee, you can view the originating Diff in D122997558. |
robhogan
reviewed
Oct 7, 2026
| every run's bundle is byte-identical to unpatched. | ||
|
|
||
| <entry>:<platform> Entry point relative to xplat/js, and platform. E.g. | ||
| RKJSModules/EntryPoints/TwilightBundle.js:android |
Collaborator
There was a problem hiding this comment.
Might want to use a different example path in OSS ;)
robhogan
reviewed
Oct 7, 2026
| '%e %U %S', | ||
| '-o', | ||
| timeFile, | ||
| 'scripts/metro/fb-metro', |
meta-codesync Bot
pushed a commit
that referenced
this pull request
Oct 8, 2026
) Summary: Adds `scripts/benchmark-patches/benchmark-patches.js`, a script that runs a named benchmark unpatched and with each of a set of patches applied, and compares the results. The first benchmark, `bundling_speed`, measures how the patches affect Metro build speed and output on real bundles. ## Why Performance changes in Metro are small (a few percent) and noisy (builds vary by ±2–7s run to run), so ad-hoc measurements easily produce wrong conclusions. While evaluating `cloneInputAst: false` in `metro-transform-worker`, two separate one-off investigations reached contradicting results: one measured the eventual fix as +3.4% slower on Meta real app 1, the other as no wall-time change. Both were artifacts of the method: a baseline measured hours earlier than the variants, no warmup build, 1–3 runs per variant, no variance or significance reported, a "prod" build without `--minify`, and an unbounded worker count. This script gives a standardised way to benchmark patches that can be reviewed and improved once, instead of bespoke tests that each need to be validated. Different people can run the same benchmark on the same patches, locally or on a remote machine, and validate each other's results the same way: a diff's numbers are reproduced with one command. ## How For every benchmark, `benchmark-patches.js`: - measures the baseline in the same batch as the patches, and interleaves runs in a rotating order, so machine drift and position effects don't land on one variant; - does one unmeasured warmup run per target; - compares any number of candidate patches in one run: each is a diff, relative to the repository root, saved as `*.patch` in `patches/`; - applies and reverts each patch around its run, including on Ctrl-C, so the working copy is left clean; - keeps the last run's outputs and the results table in `dist/` for inspection. Each benchmark is a module in `benchmarks/<name>.js` that exports `HELP`, `parseTargets`, `measure`, `describe` and `formatTable`, and is chosen explicitly by name on the command line. Other benchmarks, such as source map generation speed, can be added the same way and reuse the patch handling, run scheduling and output layout. `bundling_speed` (`benchmarks/bundling_speed.js`): - cold builds only (`--reset-cache --no-use-saved-state`), a fixed `--max-workers=16`, and real prod flags (`--minify`); - every bundle is built in both prod and dev; - CPU (user + sys, including workers) is reported next to wall time, as mean ± sd with Welch's t vs unpatched, so effects within noise are visible as such; - every run's bundle is hashed and compared with the unpatched one. A speed-up that changes output (raw `cloneInputAst: false` silently deleted a live function in the Meta real app 1 prod bundle) shows up as `NO` instead of a win; - the bundles and build logs of the last run stay in `dist/<variant>/`. `dist/` and `patches/*.patch` are ignored by `.gitignore` files. Run `node scripts/benchmark-patches/benchmark-patches.js` for usage. ## Example: measuring D123636983 with `bundling_speed` D123636983 (`cloneInputAst: false` in the worker's main Babel pass, plus a Babel scope-cache clear) was chosen with this script. Each candidate fix was saved as a patch: ``` scripts/benchmark-patches/patches/clear-before-constant-folding.patch scripts/benchmark-patches/patches/clear-before-transform.patch scripts/benchmark-patches/patches/clearscope-before-constant-folding.patch scripts/benchmark-patches/patches/clearscope-before-transform.patch ``` and compared against the unpatched tree on three Meta production apps, referred to as Meta real apps 1–3 (`app1`–`app3` in the table): ``` node scripts/benchmark-patches/benchmark-patches.js bundling_speed --runs=5 <app1 entry>:android <app2 entry>:ios ``` That is 5 variants × 4 bundles (app1 android, app2 ios, each in prod and dev) × 5 runs = 100 measured builds, plus 4 warmup builds, about 5 hours. The same patches were then run on app3 android (46,534 modules; 50 measured builds, about 3 hours). The analysis of these results is in D123636983. Variants (all set `cloneInputAst: false` in the worker's main Babel pass): - `clear-before-transform`: `traverse.cache.clear()` before the main pass. - `clear-before-constant-folding`: `traverse.cache.clear()` before the prod-only constant-folding pass. - `clearscope-before-transform`: `traverse.cache.clearScope()` before the main pass. - `clearscope-before-constant-folding`: `traverse.cache.clearScope()` before the prod-only constant-folding pass. This is D123636983's change. | Bundle | Variant | Wall (s) | Wall Δ (t) | CPU (s) | CPU Δ (t) | Size (bytes) | Same as unpatched | |---|---|--:|--:|--:|--:|--:|---| | app1-android-prod | unpatched | 207.5 ± 4.6 | | 2956.1 ± 74.0 | | 37,357,607 | yes | | | clear-before-constant-folding | 204.4 ± 5.0 | -1.48% (-0.91) | 2854.2 ± 41.9 | -3.45% (-2.40) | 37,357,607 | yes | | | clear-before-transform | 203.5 ± 4.1 | -1.91% (-1.29) | 2856.4 ± 42.4 | -3.37% (-2.34) | 37,357,607 | yes | | | clearscope-before-constant-folding | 204.4 ± 3.9 | -1.48% (-1.02) | 2853.2 ± 23.5 | -3.48% (-2.65) | 37,357,607 | yes | | | clearscope-before-transform | 210.0 ± 6.6 | +1.20% (+0.61) | 2917.8 ± 87.5 | -1.30% (-0.67) | 37,357,607 | yes | | app1-android-dev | unpatched | 183.4 ± 6.4 | | 2454.7 ± 66.2 | | 95,688,124 | yes | | | clear-before-constant-folding | 178.2 ± 5.6 | -2.83% (-1.21) | 2304.4 ± 82.6 | -6.12% (-2.84) | 95,688,124 | yes | | | clear-before-transform | 182.6 ± 4.1 | -0.42% (-0.19) | 2391.7 ± 66.4 | -2.57% (-1.34) | 95,688,124 | yes | | | clearscope-before-constant-folding | 177.2 ± 3.8 | -3.37% (-1.65) | 2320.7 ± 59.9 | -5.46% (-3.00) | 95,688,124 | yes | | | clearscope-before-transform | 182.0 ± 4.1 | -0.73% (-0.34) | 2367.3 ± 70.9 | -3.56% (-1.80) | 95,688,124 | yes | | app2-ios-prod | unpatched | 153.5 ± 2.4 | | 1502.7 ± 34.0 | | 18,768,596 | yes | | | clear-before-constant-folding | 148.1 ± 3.2 | -3.53% (-2.72) | 1456.3 ± 43.6 | -3.09% (-1.67) | 18,768,596 | yes | | | clear-before-transform | 148.5 ± 2.9 | -3.23% (-2.61) | 1430.8 ± 20.5 | -4.79% (-3.62) | 18,768,596 | yes | | | clearscope-before-constant-folding | 143.8 ± 3.5 | -6.33% (-4.56) | 1402.6 ± 26.1 | -6.66% (-4.67) | 18,768,596 | yes | | | clearscope-before-transform | 144.9 ± 2.5 | -5.57% (-4.88) | 1431.0 ± 16.4 | -4.77% (-3.80) | 18,768,596 | yes | | app2-ios-dev | unpatched | 132.5 ± 1.9 | | 1135.5 ± 14.8 | | 48,693,569 | yes | | | clear-before-constant-folding | 126.3 ± 2.0 | -4.68% (-4.51) | 1073.1 ± 10.8 | -5.50% (-6.83) | 48,693,569 | yes | | | clear-before-transform | 127.4 ± 2.8 | -3.92% (-3.06) | 1084.7 ± 19.8 | -4.47% (-4.10) | 48,693,569 | yes | | | clearscope-before-constant-folding | 124.7 ± 4.3 | -5.89% (-3.33) | 1057.6 ± 37.0 | -6.86% (-3.91) | 48,693,569 | yes | | | clearscope-before-transform | 131.0 ± 4.9 | -1.20% (-0.61) | 1114.8 ± 50.1 | -1.82% (-0.79) | 48,693,569 | yes | | app3-android-prod | unpatched | 215.3 ± 4.2 | | 3479.3 ± 35.7 | | 114,719,789 | yes | | | clear-before-constant-folding | 210.9 ± 3.2 | -2.05% (-1.9) | 3371.2 ± 61.2 | -3.11% (-3.4) | 114,719,789 | yes | | | clear-before-transform | 212.1 ± 5.3 | -1.48% (-1.1) | 3415.7 ± 91.2 | -1.83% (-1.5) | 114,719,789 | yes | | | clearscope-before-constant-folding | 211.2 ± 4.6 | -1.91% (-1.5) | 3382.7 ± 71.0 | -2.78% (-2.7) | 114,719,789 | yes | | | clearscope-before-transform | 226.0 ± 20.2 | +4.99% (+1.2) | 3633.0 ± 339.7 | +4.42% (+1.0) | 114,719,789 | yes | | app3-android-dev | unpatched | 192.3 ± 7.0 | | 2953.4 ± 130.4 | | 289,411,642 | yes | | | clear-before-constant-folding | 184.4 ± 4.2 | -4.12% (-2.2) | 2777.0 ± 78.3 | -5.97% (-2.6) | 289,411,642 | yes | | | clear-before-transform | 185.8 ± 6.7 | -3.39% (-1.5) | 2826.0 ± 102.1 | -4.31% (-1.7) | 289,411,642 | yes | | | clearscope-before-constant-folding | 184.4 ± 8.1 | -4.09% (-1.6) | 2803.5 ± 113.8 | -5.08% (-1.9) | 289,411,642 | yes | | | clearscope-before-transform | 185.6 ± 5.7 | -3.46% (-1.7) | 2823.2 ± 92.1 | -4.41% (-1.8) | 289,411,642 | yes | 5 measured runs per variant, each a cold build (`--reset-cache --no-use-saved-state`, `--max-workers=16`) on a 72-core Linux machine. CPU is user + sys of the build and its workers. Values are mean ± sd; t is Welch's t vs unpatched (at 5 runs, |t| < 2.3 is not significant). "Same as unpatched" compares bundle hashes. app1 and app2 were measured with an earlier revision of the script that ran variants in a fixed order (unpatched first) and printed population sd; their t values were computed afterwards from the same per-run timings. app3 was measured with this revision: rotating variant order, sample sd, and output identity checked on every run. Changelog: [Internal] Differential Revision: D122997558
meta-codesync
Bot
force-pushed
the
export-D122997558
branch
from
October 8, 2026 08:56
228ff9c to
72af3bf
Compare
meta-codesync Bot
pushed a commit
that referenced
this pull request
Oct 8, 2026
) Summary: Adds `scripts/benchmark-patches/benchmark-patches.js`, a script that runs a named benchmark unpatched and with each of a set of patches applied, and compares the results. The first benchmark, `bundling_speed`, measures how the patches affect Metro build speed and output on real bundles. ## Why Performance changes in Metro are small (a few percent) and noisy (builds vary by ±2–7s run to run), so ad-hoc measurements easily produce wrong conclusions. While evaluating `cloneInputAst: false` in `metro-transform-worker`, two separate one-off investigations reached contradicting results: one measured the eventual fix as +3.4% slower on Meta real app 1, the other as no wall-time change. Both were artifacts of the method: a baseline measured hours earlier than the variants, no warmup build, 1–3 runs per variant, no variance or significance reported, a "prod" build without `--minify`, and an unbounded worker count. This script gives a standardised way to benchmark patches that can be reviewed and improved once, instead of bespoke tests that each need to be validated. Different people can run the same benchmark on the same patches, locally or on a remote machine, and validate each other's results the same way: a diff's numbers are reproduced with one command. ## How For every benchmark, `benchmark-patches.js`: - measures the baseline in the same batch as the patches, and interleaves runs in a rotating order, so machine drift and position effects don't land on one variant; - does one unmeasured warmup run per target; - compares any number of candidate patches in one run: each is a diff, relative to the repository root, saved as `*.patch` in `patches/`; - applies and reverts each patch around its run, including on Ctrl-C, so the working copy is left clean; - keeps the last run's outputs and the results table in `dist/` for inspection. Each benchmark is a module in `benchmarks/<name>.js` that exports `HELP`, `parseTargets`, `measure`, `describe` and `formatTable`, and is chosen explicitly by name on the command line. Other benchmarks, such as source map generation speed, can be added the same way and reuse the patch handling, run scheduling and output layout. `bundling_speed` (`benchmarks/bundling_speed.js`): - cold builds only (`--reset-cache --no-use-saved-state`), a fixed `--max-workers=16`, and real prod flags (`--minify`); - every bundle is built in both prod and dev; - CPU (user + sys, including workers) is reported next to wall time, as mean ± sd with Welch's t vs unpatched, so effects within noise are visible as such; - every run's bundle is hashed and compared with the unpatched one. A speed-up that changes output (raw `cloneInputAst: false` silently deleted a live function in the Meta real app 1 prod bundle) shows up as `NO` instead of a win; - the bundles and build logs of the last run stay in `dist/<variant>/`. `dist/` and `patches/*.patch` are ignored by `.gitignore` files. Run `node scripts/benchmark-patches/benchmark-patches.js` for usage. ## Example: measuring D123636983 with `bundling_speed` D123636983 (`cloneInputAst: false` in the worker's main Babel pass, plus a Babel scope-cache clear) was chosen with this script. Each candidate fix was saved as a patch: ``` scripts/benchmark-patches/patches/clear-before-constant-folding.patch scripts/benchmark-patches/patches/clear-before-transform.patch scripts/benchmark-patches/patches/clearscope-before-constant-folding.patch scripts/benchmark-patches/patches/clearscope-before-transform.patch ``` and compared against the unpatched tree on three Meta production apps, referred to as Meta real apps 1–3 (`app1`–`app3` in the table): ``` node scripts/benchmark-patches/benchmark-patches.js bundling_speed --runs=5 <app1 entry>:android <app2 entry>:ios ``` That is 5 variants × 4 bundles (app1 android, app2 ios, each in prod and dev) × 5 runs = 100 measured builds, plus 4 warmup builds, about 5 hours. The same patches were then run on app3 android (46,534 modules; 50 measured builds, about 3 hours). The analysis of these results is in D123636983. Variants (all set `cloneInputAst: false` in the worker's main Babel pass): - `clear-before-transform`: `traverse.cache.clear()` before the main pass. - `clear-before-constant-folding`: `traverse.cache.clear()` before the prod-only constant-folding pass. - `clearscope-before-transform`: `traverse.cache.clearScope()` before the main pass. - `clearscope-before-constant-folding`: `traverse.cache.clearScope()` before the prod-only constant-folding pass. This is D123636983's change. | Bundle | Variant | Wall (s) | Wall Δ (t) | CPU (s) | CPU Δ (t) | Size (bytes) | Same as unpatched | |---|---|--:|--:|--:|--:|--:|---| | app1-android-prod | unpatched | 207.5 ± 4.6 | | 2956.1 ± 74.0 | | 37,357,607 | yes | | | clear-before-constant-folding | 204.4 ± 5.0 | -1.48% (-0.91) | 2854.2 ± 41.9 | -3.45% (-2.40) | 37,357,607 | yes | | | clear-before-transform | 203.5 ± 4.1 | -1.91% (-1.29) | 2856.4 ± 42.4 | -3.37% (-2.34) | 37,357,607 | yes | | | clearscope-before-constant-folding | 204.4 ± 3.9 | -1.48% (-1.02) | 2853.2 ± 23.5 | -3.48% (-2.65) | 37,357,607 | yes | | | clearscope-before-transform | 210.0 ± 6.6 | +1.20% (+0.61) | 2917.8 ± 87.5 | -1.30% (-0.67) | 37,357,607 | yes | | app1-android-dev | unpatched | 183.4 ± 6.4 | | 2454.7 ± 66.2 | | 95,688,124 | yes | | | clear-before-constant-folding | 178.2 ± 5.6 | -2.83% (-1.21) | 2304.4 ± 82.6 | -6.12% (-2.84) | 95,688,124 | yes | | | clear-before-transform | 182.6 ± 4.1 | -0.42% (-0.19) | 2391.7 ± 66.4 | -2.57% (-1.34) | 95,688,124 | yes | | | clearscope-before-constant-folding | 177.2 ± 3.8 | -3.37% (-1.65) | 2320.7 ± 59.9 | -5.46% (-3.00) | 95,688,124 | yes | | | clearscope-before-transform | 182.0 ± 4.1 | -0.73% (-0.34) | 2367.3 ± 70.9 | -3.56% (-1.80) | 95,688,124 | yes | | app2-ios-prod | unpatched | 153.5 ± 2.4 | | 1502.7 ± 34.0 | | 18,768,596 | yes | | | clear-before-constant-folding | 148.1 ± 3.2 | -3.53% (-2.72) | 1456.3 ± 43.6 | -3.09% (-1.67) | 18,768,596 | yes | | | clear-before-transform | 148.5 ± 2.9 | -3.23% (-2.61) | 1430.8 ± 20.5 | -4.79% (-3.62) | 18,768,596 | yes | | | clearscope-before-constant-folding | 143.8 ± 3.5 | -6.33% (-4.56) | 1402.6 ± 26.1 | -6.66% (-4.67) | 18,768,596 | yes | | | clearscope-before-transform | 144.9 ± 2.5 | -5.57% (-4.88) | 1431.0 ± 16.4 | -4.77% (-3.80) | 18,768,596 | yes | | app2-ios-dev | unpatched | 132.5 ± 1.9 | | 1135.5 ± 14.8 | | 48,693,569 | yes | | | clear-before-constant-folding | 126.3 ± 2.0 | -4.68% (-4.51) | 1073.1 ± 10.8 | -5.50% (-6.83) | 48,693,569 | yes | | | clear-before-transform | 127.4 ± 2.8 | -3.92% (-3.06) | 1084.7 ± 19.8 | -4.47% (-4.10) | 48,693,569 | yes | | | clearscope-before-constant-folding | 124.7 ± 4.3 | -5.89% (-3.33) | 1057.6 ± 37.0 | -6.86% (-3.91) | 48,693,569 | yes | | | clearscope-before-transform | 131.0 ± 4.9 | -1.20% (-0.61) | 1114.8 ± 50.1 | -1.82% (-0.79) | 48,693,569 | yes | | app3-android-prod | unpatched | 215.3 ± 4.2 | | 3479.3 ± 35.7 | | 114,719,789 | yes | | | clear-before-constant-folding | 210.9 ± 3.2 | -2.05% (-1.9) | 3371.2 ± 61.2 | -3.11% (-3.4) | 114,719,789 | yes | | | clear-before-transform | 212.1 ± 5.3 | -1.48% (-1.1) | 3415.7 ± 91.2 | -1.83% (-1.5) | 114,719,789 | yes | | | clearscope-before-constant-folding | 211.2 ± 4.6 | -1.91% (-1.5) | 3382.7 ± 71.0 | -2.78% (-2.7) | 114,719,789 | yes | | | clearscope-before-transform | 226.0 ± 20.2 | +4.99% (+1.2) | 3633.0 ± 339.7 | +4.42% (+1.0) | 114,719,789 | yes | | app3-android-dev | unpatched | 192.3 ± 7.0 | | 2953.4 ± 130.4 | | 289,411,642 | yes | | | clear-before-constant-folding | 184.4 ± 4.2 | -4.12% (-2.2) | 2777.0 ± 78.3 | -5.97% (-2.6) | 289,411,642 | yes | | | clear-before-transform | 185.8 ± 6.7 | -3.39% (-1.5) | 2826.0 ± 102.1 | -4.31% (-1.7) | 289,411,642 | yes | | | clearscope-before-constant-folding | 184.4 ± 8.1 | -4.09% (-1.6) | 2803.5 ± 113.8 | -5.08% (-1.9) | 289,411,642 | yes | | | clearscope-before-transform | 185.6 ± 5.7 | -3.46% (-1.7) | 2823.2 ± 92.1 | -4.41% (-1.8) | 289,411,642 | yes | 5 measured runs per variant, each a cold build (`--reset-cache --no-use-saved-state`, `--max-workers=16`) on a 72-core Linux machine. CPU is user + sys of the build and its workers. Values are mean ± sd; t is Welch's t vs unpatched (at 5 runs, |t| < 2.3 is not significant). "Same as unpatched" compares bundle hashes. app1 and app2 were measured with an earlier revision of the script that ran variants in a fixed order (unpatched first) and printed population sd; their t values were computed afterwards from the same per-run timings. app3 was measured with this revision: rotating variant order, sample sd, and output identity checked on every run. Changelog: [Internal] Differential Revision: D122997558
meta-codesync Bot
pushed a commit
that referenced
this pull request
Oct 8, 2026
) Summary: Adds `scripts/benchmark-patches/benchmark-patches.js`, a script that runs a named benchmark unpatched and with each of a set of patches applied, and compares the results. The first benchmark, `bundling_speed`, measures how the patches affect Metro build speed and output on real bundles. ## Why Performance changes in Metro are small (a few percent) and noisy (builds vary by ±2–7s run to run), so ad-hoc measurements easily produce wrong conclusions. While evaluating `cloneInputAst: false` in `metro-transform-worker`, two separate one-off investigations reached contradicting results: one measured the eventual fix as +3.4% slower on Meta real app 1, the other as no wall-time change. Both were artifacts of the method: a baseline measured hours earlier than the variants, no warmup build, 1–3 runs per variant, no variance or significance reported, a "prod" build without `--minify`, and an unbounded worker count. This script gives a standardised way to benchmark patches that can be reviewed and improved once, instead of bespoke tests that each need to be validated. Different people can run the same benchmark on the same patches, locally or on a remote machine, and validate each other's results the same way: a diff's numbers are reproduced with one command. ## How For every benchmark, `benchmark-patches.js`: - measures the baseline in the same batch as the patches, and interleaves runs in a rotating order, so machine drift and position effects don't land on one variant; - does one unmeasured warmup run per target; - compares any number of candidate patches in one run: each is a diff, relative to the repository root, saved as `*.patch` in `patches/`; - applies and reverts each patch around its run, including on Ctrl-C, so the working copy is left clean; - keeps the last run's outputs and the results table in `dist/` for inspection. Each benchmark is a module in `benchmarks/<name>.js` that exports `HELP`, `parseTargets`, `measure`, `describe` and `formatTable`, and is chosen explicitly by name on the command line. Other benchmarks, such as source map generation speed, can be added the same way and reuse the patch handling, run scheduling and output layout. `bundling_speed` (`benchmarks/bundling_speed.js`): - cold builds only (`--reset-cache --no-use-saved-state`), a fixed `--max-workers=16`, and real prod flags (`--minify`); - every bundle is built in both prod and dev; - CPU (user + sys, including workers) is reported next to wall time, as mean ± sd with Welch's t vs unpatched, so effects within noise are visible as such; - every run's bundle is hashed and compared with the unpatched one. A speed-up that changes output (raw `cloneInputAst: false` silently deleted a live function in the Meta real app 1 prod bundle) shows up as `NO` instead of a win; - the bundles and build logs of the last run stay in `dist/<variant>/`. `dist/` and `patches/*.patch` are ignored by `.gitignore` files. Run `node scripts/benchmark-patches/benchmark-patches.js` for usage. ## Example: measuring D123636983 with `bundling_speed` D123636983 (`cloneInputAst: false` in the worker's main Babel pass, plus a Babel scope-cache clear) was chosen with this script. Each candidate fix was saved as a patch: ``` scripts/benchmark-patches/patches/clear-before-constant-folding.patch scripts/benchmark-patches/patches/clear-before-transform.patch scripts/benchmark-patches/patches/clearscope-before-constant-folding.patch scripts/benchmark-patches/patches/clearscope-before-transform.patch ``` and compared against the unpatched tree on three Meta production apps, referred to as Meta real apps 1–3 (`app1`–`app3` in the table): ``` node scripts/benchmark-patches/benchmark-patches.js bundling_speed --runs=5 <app1 entry>:android <app2 entry>:ios ``` That is 5 variants × 4 bundles (app1 android, app2 ios, each in prod and dev) × 5 runs = 100 measured builds, plus 4 warmup builds, about 5 hours. The same patches were then run on app3 android (46,534 modules; 50 measured builds, about 3 hours). The analysis of these results is in D123636983. Variants (all set `cloneInputAst: false` in the worker's main Babel pass): - `clear-before-transform`: `traverse.cache.clear()` before the main pass. - `clear-before-constant-folding`: `traverse.cache.clear()` before the prod-only constant-folding pass. - `clearscope-before-transform`: `traverse.cache.clearScope()` before the main pass. - `clearscope-before-constant-folding`: `traverse.cache.clearScope()` before the prod-only constant-folding pass. This is D123636983's change. | Bundle | Variant | Wall (s) | Wall Δ (t) | CPU (s) | CPU Δ (t) | Size (bytes) | Same as unpatched | |---|---|--:|--:|--:|--:|--:|---| | app1-android-prod | unpatched | 207.5 ± 4.6 | | 2956.1 ± 74.0 | | 37,357,607 | yes | | | clear-before-constant-folding | 204.4 ± 5.0 | -1.48% (-0.91) | 2854.2 ± 41.9 | -3.45% (-2.40) | 37,357,607 | yes | | | clear-before-transform | 203.5 ± 4.1 | -1.91% (-1.29) | 2856.4 ± 42.4 | -3.37% (-2.34) | 37,357,607 | yes | | | clearscope-before-constant-folding | 204.4 ± 3.9 | -1.48% (-1.02) | 2853.2 ± 23.5 | -3.48% (-2.65) | 37,357,607 | yes | | | clearscope-before-transform | 210.0 ± 6.6 | +1.20% (+0.61) | 2917.8 ± 87.5 | -1.30% (-0.67) | 37,357,607 | yes | | app1-android-dev | unpatched | 183.4 ± 6.4 | | 2454.7 ± 66.2 | | 95,688,124 | yes | | | clear-before-constant-folding | 178.2 ± 5.6 | -2.83% (-1.21) | 2304.4 ± 82.6 | -6.12% (-2.84) | 95,688,124 | yes | | | clear-before-transform | 182.6 ± 4.1 | -0.42% (-0.19) | 2391.7 ± 66.4 | -2.57% (-1.34) | 95,688,124 | yes | | | clearscope-before-constant-folding | 177.2 ± 3.8 | -3.37% (-1.65) | 2320.7 ± 59.9 | -5.46% (-3.00) | 95,688,124 | yes | | | clearscope-before-transform | 182.0 ± 4.1 | -0.73% (-0.34) | 2367.3 ± 70.9 | -3.56% (-1.80) | 95,688,124 | yes | | app2-ios-prod | unpatched | 153.5 ± 2.4 | | 1502.7 ± 34.0 | | 18,768,596 | yes | | | clear-before-constant-folding | 148.1 ± 3.2 | -3.53% (-2.72) | 1456.3 ± 43.6 | -3.09% (-1.67) | 18,768,596 | yes | | | clear-before-transform | 148.5 ± 2.9 | -3.23% (-2.61) | 1430.8 ± 20.5 | -4.79% (-3.62) | 18,768,596 | yes | | | clearscope-before-constant-folding | 143.8 ± 3.5 | -6.33% (-4.56) | 1402.6 ± 26.1 | -6.66% (-4.67) | 18,768,596 | yes | | | clearscope-before-transform | 144.9 ± 2.5 | -5.57% (-4.88) | 1431.0 ± 16.4 | -4.77% (-3.80) | 18,768,596 | yes | | app2-ios-dev | unpatched | 132.5 ± 1.9 | | 1135.5 ± 14.8 | | 48,693,569 | yes | | | clear-before-constant-folding | 126.3 ± 2.0 | -4.68% (-4.51) | 1073.1 ± 10.8 | -5.50% (-6.83) | 48,693,569 | yes | | | clear-before-transform | 127.4 ± 2.8 | -3.92% (-3.06) | 1084.7 ± 19.8 | -4.47% (-4.10) | 48,693,569 | yes | | | clearscope-before-constant-folding | 124.7 ± 4.3 | -5.89% (-3.33) | 1057.6 ± 37.0 | -6.86% (-3.91) | 48,693,569 | yes | | | clearscope-before-transform | 131.0 ± 4.9 | -1.20% (-0.61) | 1114.8 ± 50.1 | -1.82% (-0.79) | 48,693,569 | yes | | app3-android-prod | unpatched | 215.3 ± 4.2 | | 3479.3 ± 35.7 | | 114,719,789 | yes | | | clear-before-constant-folding | 210.9 ± 3.2 | -2.05% (-1.9) | 3371.2 ± 61.2 | -3.11% (-3.4) | 114,719,789 | yes | | | clear-before-transform | 212.1 ± 5.3 | -1.48% (-1.1) | 3415.7 ± 91.2 | -1.83% (-1.5) | 114,719,789 | yes | | | clearscope-before-constant-folding | 211.2 ± 4.6 | -1.91% (-1.5) | 3382.7 ± 71.0 | -2.78% (-2.7) | 114,719,789 | yes | | | clearscope-before-transform | 226.0 ± 20.2 | +4.99% (+1.2) | 3633.0 ± 339.7 | +4.42% (+1.0) | 114,719,789 | yes | | app3-android-dev | unpatched | 192.3 ± 7.0 | | 2953.4 ± 130.4 | | 289,411,642 | yes | | | clear-before-constant-folding | 184.4 ± 4.2 | -4.12% (-2.2) | 2777.0 ± 78.3 | -5.97% (-2.6) | 289,411,642 | yes | | | clear-before-transform | 185.8 ± 6.7 | -3.39% (-1.5) | 2826.0 ± 102.1 | -4.31% (-1.7) | 289,411,642 | yes | | | clearscope-before-constant-folding | 184.4 ± 8.1 | -4.09% (-1.6) | 2803.5 ± 113.8 | -5.08% (-1.9) | 289,411,642 | yes | | | clearscope-before-transform | 185.6 ± 5.7 | -3.46% (-1.7) | 2823.2 ± 92.1 | -4.41% (-1.8) | 289,411,642 | yes | 5 measured runs per variant, each a cold build (`--reset-cache --no-use-saved-state`, `--max-workers=16`) on a 72-core Linux machine. CPU is user + sys of the build and its workers. Values are mean ± sd; t is Welch's t vs unpatched (at 5 runs, |t| < 2.3 is not significant). "Same as unpatched" compares bundle hashes. app1 and app2 were measured with an earlier revision of the script that ran variants in a fixed order (unpatched first) and printed population sd; their t values were computed afterwards from the same per-run timings. app3 was measured with this revision: rotating variant order, sample sd, and output identity checked on every run. Changelog: [Internal] Differential Revision: D122997558
) Summary: Pull Request resolved: #2032 Adds `scripts/benchmark-patches/benchmark-patches.js`, a script that runs a named benchmark unpatched and with each of a set of patches applied, and compares the results. The first benchmark, `bundling_speed`, measures how the patches affect Metro build speed and output on real bundles. ## Why Performance changes in Metro are small (a few percent) and noisy (builds vary by ±2–7s run to run), so ad-hoc measurements easily produce wrong conclusions. While evaluating `cloneInputAst: false` in `metro-transform-worker`, two separate one-off investigations reached contradicting results: one measured the eventual fix as +3.4% slower on Meta real app 1, the other as no wall-time change. Both were artifacts of the method: a baseline measured hours earlier than the variants, no warmup build, 1–3 runs per variant, no variance or significance reported, a "prod" build without `--minify`, and an unbounded worker count. This script gives a standardised way to benchmark patches that can be reviewed and improved once, instead of bespoke tests that each need to be validated. Different people can run the same benchmark on the same patches, locally or on a remote machine, and validate each other's results the same way: a diff's numbers are reproduced with one command. ## How For every benchmark, `benchmark-patches.js`: - measures the baseline in the same batch as the patches, and interleaves runs in a rotating order, so machine drift and position effects don't land on one variant; - does one unmeasured warmup run per target; - compares any number of candidate patches in one run: each is a diff, relative to the repository root, saved as `*.patch` in `patches/`; - applies and reverts each patch around its run, including on Ctrl-C, so the working copy is left clean; - keeps the last run's outputs and the results table in `dist/` for inspection. Each benchmark is a module in `benchmarks/<name>.js` that exports `HELP`, `parseTargets`, `measure`, `describe` and `formatTable`, and is chosen explicitly by name on the command line. Other benchmarks, such as source map generation speed, can be added the same way and reuse the patch handling, run scheduling and output layout. `bundling_speed` (`benchmarks/bundling_speed.js`): - cold builds only (`--reset-cache --no-use-saved-state`), a fixed `--max-workers=16`, and real prod flags (`--minify`); - every bundle is built in both prod and dev; - CPU (user + sys, including workers) is reported next to wall time, as mean ± sd with Welch's t vs unpatched, so effects within noise are visible as such; - every run's bundle is hashed and compared with the unpatched one. A speed-up that changes output (raw `cloneInputAst: false` silently deleted a live function in the Meta real app 1 prod bundle) shows up as `NO` instead of a win; - the bundles and build logs of the last run stay in `dist/<variant>/`. `dist/` and `patches/*.patch` are ignored by `.gitignore` files. Run `node scripts/benchmark-patches/benchmark-patches.js` for usage. ## Example: measuring D123636983 with `bundling_speed` D123636983 (`cloneInputAst: false` in the worker's main Babel pass, plus a Babel scope-cache clear) was chosen with this script. Each candidate fix was saved as a patch: ``` scripts/benchmark-patches/patches/clear-before-constant-folding.patch scripts/benchmark-patches/patches/clear-before-transform.patch scripts/benchmark-patches/patches/clearscope-before-constant-folding.patch scripts/benchmark-patches/patches/clearscope-before-transform.patch ``` and compared against the unpatched tree on three Meta production apps, referred to as Meta real apps 1–3 (`app1`–`app3` in the table): ``` node scripts/benchmark-patches/benchmark-patches.js bundling_speed --runs=5 <app1 entry>:android <app2 entry>:ios ``` That is 5 variants × 4 bundles (app1 android, app2 ios, each in prod and dev) × 5 runs = 100 measured builds, plus 4 warmup builds, about 5 hours. The same patches were then run on app3 android (46,534 modules; 50 measured builds, about 3 hours). The analysis of these results is in D123636983. Variants (all set `cloneInputAst: false` in the worker's main Babel pass): - `clear-before-transform`: `traverse.cache.clear()` before the main pass. - `clear-before-constant-folding`: `traverse.cache.clear()` before the prod-only constant-folding pass. - `clearscope-before-transform`: `traverse.cache.clearScope()` before the main pass. - `clearscope-before-constant-folding`: `traverse.cache.clearScope()` before the prod-only constant-folding pass. This is D123636983's change. | Bundle | Variant | Wall (s) | Wall Δ (t) | CPU (s) | CPU Δ (t) | Size (bytes) | Same as unpatched | |---|---|--:|--:|--:|--:|--:|---| | app1-android-prod | unpatched | 207.5 ± 4.6 | | 2956.1 ± 74.0 | | 37,357,607 | yes | | | clear-before-constant-folding | 204.4 ± 5.0 | -1.48% (-0.91) | 2854.2 ± 41.9 | -3.45% (-2.40) | 37,357,607 | yes | | | clear-before-transform | 203.5 ± 4.1 | -1.91% (-1.29) | 2856.4 ± 42.4 | -3.37% (-2.34) | 37,357,607 | yes | | | clearscope-before-constant-folding | 204.4 ± 3.9 | -1.48% (-1.02) | 2853.2 ± 23.5 | -3.48% (-2.65) | 37,357,607 | yes | | | clearscope-before-transform | 210.0 ± 6.6 | +1.20% (+0.61) | 2917.8 ± 87.5 | -1.30% (-0.67) | 37,357,607 | yes | | app1-android-dev | unpatched | 183.4 ± 6.4 | | 2454.7 ± 66.2 | | 95,688,124 | yes | | | clear-before-constant-folding | 178.2 ± 5.6 | -2.83% (-1.21) | 2304.4 ± 82.6 | -6.12% (-2.84) | 95,688,124 | yes | | | clear-before-transform | 182.6 ± 4.1 | -0.42% (-0.19) | 2391.7 ± 66.4 | -2.57% (-1.34) | 95,688,124 | yes | | | clearscope-before-constant-folding | 177.2 ± 3.8 | -3.37% (-1.65) | 2320.7 ± 59.9 | -5.46% (-3.00) | 95,688,124 | yes | | | clearscope-before-transform | 182.0 ± 4.1 | -0.73% (-0.34) | 2367.3 ± 70.9 | -3.56% (-1.80) | 95,688,124 | yes | | app2-ios-prod | unpatched | 153.5 ± 2.4 | | 1502.7 ± 34.0 | | 18,768,596 | yes | | | clear-before-constant-folding | 148.1 ± 3.2 | -3.53% (-2.72) | 1456.3 ± 43.6 | -3.09% (-1.67) | 18,768,596 | yes | | | clear-before-transform | 148.5 ± 2.9 | -3.23% (-2.61) | 1430.8 ± 20.5 | -4.79% (-3.62) | 18,768,596 | yes | | | clearscope-before-constant-folding | 143.8 ± 3.5 | -6.33% (-4.56) | 1402.6 ± 26.1 | -6.66% (-4.67) | 18,768,596 | yes | | | clearscope-before-transform | 144.9 ± 2.5 | -5.57% (-4.88) | 1431.0 ± 16.4 | -4.77% (-3.80) | 18,768,596 | yes | | app2-ios-dev | unpatched | 132.5 ± 1.9 | | 1135.5 ± 14.8 | | 48,693,569 | yes | | | clear-before-constant-folding | 126.3 ± 2.0 | -4.68% (-4.51) | 1073.1 ± 10.8 | -5.50% (-6.83) | 48,693,569 | yes | | | clear-before-transform | 127.4 ± 2.8 | -3.92% (-3.06) | 1084.7 ± 19.8 | -4.47% (-4.10) | 48,693,569 | yes | | | clearscope-before-constant-folding | 124.7 ± 4.3 | -5.89% (-3.33) | 1057.6 ± 37.0 | -6.86% (-3.91) | 48,693,569 | yes | | | clearscope-before-transform | 131.0 ± 4.9 | -1.20% (-0.61) | 1114.8 ± 50.1 | -1.82% (-0.79) | 48,693,569 | yes | | app3-android-prod | unpatched | 215.3 ± 4.2 | | 3479.3 ± 35.7 | | 114,719,789 | yes | | | clear-before-constant-folding | 210.9 ± 3.2 | -2.05% (-1.9) | 3371.2 ± 61.2 | -3.11% (-3.4) | 114,719,789 | yes | | | clear-before-transform | 212.1 ± 5.3 | -1.48% (-1.1) | 3415.7 ± 91.2 | -1.83% (-1.5) | 114,719,789 | yes | | | clearscope-before-constant-folding | 211.2 ± 4.6 | -1.91% (-1.5) | 3382.7 ± 71.0 | -2.78% (-2.7) | 114,719,789 | yes | | | clearscope-before-transform | 226.0 ± 20.2 | +4.99% (+1.2) | 3633.0 ± 339.7 | +4.42% (+1.0) | 114,719,789 | yes | | app3-android-dev | unpatched | 192.3 ± 7.0 | | 2953.4 ± 130.4 | | 289,411,642 | yes | | | clear-before-constant-folding | 184.4 ± 4.2 | -4.12% (-2.2) | 2777.0 ± 78.3 | -5.97% (-2.6) | 289,411,642 | yes | | | clear-before-transform | 185.8 ± 6.7 | -3.39% (-1.5) | 2826.0 ± 102.1 | -4.31% (-1.7) | 289,411,642 | yes | | | clearscope-before-constant-folding | 184.4 ± 8.1 | -4.09% (-1.6) | 2803.5 ± 113.8 | -5.08% (-1.9) | 289,411,642 | yes | | | clearscope-before-transform | 185.6 ± 5.7 | -3.46% (-1.7) | 2823.2 ± 92.1 | -4.41% (-1.8) | 289,411,642 | yes | 5 measured runs per variant, each a cold build (`--reset-cache --no-use-saved-state`, `--max-workers=16`) on a 72-core Linux machine. CPU is user + sys of the build and its workers. Values are mean ± sd; t is Welch's t vs unpatched (at 5 runs, |t| < 2.3 is not significant). "Same as unpatched" compares bundle hashes. app1 and app2 were measured with an earlier revision of the script that ran variants in a fixed order (unpatched first) and printed population sd; their t values were computed afterwards from the same per-run timings. app3 was measured with this revision: rotating variant order, sample sd, and output identity checked on every run. Changelog: [Internal] Differential Revision: D122997558
meta-codesync
Bot
force-pushed
the
export-D122997558
branch
from
October 8, 2026 08:58
72af3bf to
812c4cc
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary:
Adds
scripts/benchmark-patches/benchmark-patches.js, a script that runs a named benchmark unpatched and with each of a set of patches applied, and compares the results. The first benchmark,bundling_speed, measures how the patches affect Metro build speed and output on real bundles.Why
Performance changes in Metro are small (a few percent) and noisy (builds vary by ±2–7s run to run), so ad-hoc measurements easily produce wrong conclusions. While evaluating
cloneInputAst: falseinmetro-transform-worker, two separate one-off investigations reached contradicting results: one measured the eventual fix as +3.4% slower on Meta real app 1, the other as no wall-time change. Both were artifacts of the method: a baseline measured hours earlier than the variants, no warmup build, 1–3 runs per variant, no variance or significance reported, a "prod" build without--minify, and an unbounded worker count.This script gives a standardised way to benchmark patches that can be reviewed and improved once, instead of bespoke tests that each need to be validated. Different people can run the same benchmark on the same patches, locally or on a remote machine, and validate each other's results the same way: a diff's numbers are reproduced with one command.
How
For every benchmark,
benchmark-patches.js:*.patchinpatches/;dist/for inspection.Each benchmark is a module in
benchmarks/<name>.jsthat exportsHELP,parseTargets,measure,describeandformatTable, and is chosen explicitly by name on the command line. Other benchmarks, such as source map generation speed, can be added the same way and reuse the patch handling, run scheduling and output layout.bundling_speed(benchmarks/bundling_speed.js):--reset-cache --no-use-saved-state), a fixed--max-workers=16, and real prod flags (--minify);cloneInputAst: falsesilently deleted a live function in the Meta real app 1 prod bundle) shows up asNOinstead of a win;dist/<variant>/.dist/andpatches/*.patchare ignored by.gitignorefiles. Runnode scripts/benchmark-patches/benchmark-patches.jsfor usage.Example: measuring D123636983 with
bundling_speedD123636983 (
cloneInputAst: falsein the worker's main Babel pass, plus a Babel scope-cache clear) was chosen with this script. Each candidate fix was saved as a patch:and compared against the unpatched tree on three Meta production apps, referred to as Meta real apps 1–3 (
app1–app3in the table):That is 5 variants × 4 bundles (app1 android, app2 ios, each in prod and dev) × 5 runs = 100 measured builds, plus 4 warmup builds, about 5 hours. The same patches were then run on app3 android (46,534 modules; 50 measured builds, about 3 hours). The analysis of these results is in D123636983.
Variants (all set
cloneInputAst: falsein the worker's main Babel pass):clear-before-transform:traverse.cache.clear()before the main pass.clear-before-constant-folding:traverse.cache.clear()before the prod-only constant-folding pass.clearscope-before-transform:traverse.cache.clearScope()before the main pass.clearscope-before-constant-folding:traverse.cache.clearScope()before the prod-only constant-folding pass. This is D123636983's change.5 measured runs per variant, each a cold build (
--reset-cache --no-use-saved-state,--max-workers=16) on a 72-core Linux machine. CPU is user + sys of the build and its workers. Values are mean ± sd; t is Welch's t vs unpatched (at 5 runs, |t| < 2.3 is not significant). "Same as unpatched" compares bundle hashes.app1 and app2 were measured with an earlier revision of the script that ran variants in a fixed order (unpatched first) and printed population sd; their t values were computed afterwards from the same per-run timings. app3 was measured with this revision: rotating variant order, sample sd, and output identity checked on every run.
Changelog: [Internal]
Differential Revision: D122997558