Benchmarking Methods
How to benchmark MVT code without fooling yourself, and how to run this repo's benchmarks. Per-frame costs in a game loop are measured in microseconds, and at that scale the benchmark itself can distort the result more than the code under test does.
Related: Why Performance Matters · Performance Measurements · Hot Paths · Reactivity: Why MVT Uses Polling
About the figures
The timings on this page are examples from one 2025 machine, chosen to show how flawed methods can distort a result. What matters in each is the difference between the two numbers, not the numbers themselves.
Hot Paths says what to avoid, and Performance Measurements says what things cost. This page says how those costs were measured, and how to measure your own.
The Rules at a Glance
| Rule | Why |
|---|---|
| One approach per process | Approaches timed side by side distort each other's timings |
| Bundle to plain JavaScript first | A TypeScript loader adds its own noise to the timings |
| Warm up, time in batches, repeat across processes | One run of one process is not a measurement |
| Check that the code under test does its work | A reactive system that never reacts looks free |
| Measure allocation without a collection in the way | A garbage collection mid-measurement hides what was allocated |
| Vary one thing at a time | Dividing by the wrong unit hides part of the cost |
| Read profiles with inlining in mind | Inlining puts almost every sample on the outermost function |
| Measure the whole frame, and say what is excluded | A cost moved elsewhere is still a cost |
One Approach per Process
V8 shares inline caches and optimisation decisions across all the code in a process. When two approaches are timed side by side, whichever runs first shapes how the others are compiled. A benchmark suite that declares several cases in one file, as vitest bench suites do, has this problem by default.
In this repo, the same two approaches gave these numbers (1000 Pixi containers, each with three dynamic properties, nothing changed):
| In one Vitest process | One approach per process, bundled | |
|---|---|---|
| JSX runtime (an earlier version) | 26.5-29.2 µs | 9.75-10.11 µs |
| Hand-written refresh methods | 10.9-11.6 µs | 5.55-5.70 µs |
| Ratio | 2.3-2.7x | about 1.8x |
The in-process numbers also showed a per-container overhead for refreshView that does not exist, and made a faster version of its loop look slower. An idea was rejected on that evidence.
Rule: run each approach in a fresh process. A small driver script that starts one Node process per case and collects the results is enough; that is what packages/benchmarks/run.ts does.
Bundle First
Running TypeScript through a loader such as tsx puts the loader's own work in the timings. Under tsx, 12 identical runs of one case ranged from 8.2 to 18.0 µs, and the loader accounted for about 40% of the samples in a CPU profile. Bundled to plain JavaScript with esbuild, 8 runs ranged from 7.80 to 8.05 µs.
Rule: bundle the code under test to plain JavaScript, then time and profile that. The driver may still run under tsx, because it does no timing itself.
npx esbuild bench.ts --bundle --platform=node --format=esm --outfile=bench.mjs
node bench.mjsWarm Up, Batch and Repeat
A single timing is dominated by whatever else the machine and the engine were doing at that moment. The pattern used in this repo:
- Warm up. Run at least 1000 frames and 300 ms untimed, so the code under test is optimised before timing starts. A frame slow enough to take 2 s over it stops at 100 frames: a frame that slow loops over tens of thousands of objects, so its code is optimised within the first few, and a thousand would take tens of seconds.
- Time in batches. Time batches of many frames rather than single frames, because a frame of a few microseconds is close to the timer's resolution. Each batch here lasts about 30 ms, however long a frame takes.
- Report the median batch, not the mean. Garbage collection and other interruptions produce outliers that drag a mean upwards.
- Repeat across processes. Run each case in several fresh processes and report the median with the range. Here, each case runs twice, and a third time only when the two disagree by more than 5%, the point at which the tables mark a result as noisy.
Expect noise of about 5-10% between identical processes. Two approaches whose ranges overlap are not measurably different. When a comparison depends on a small difference, check it by running two identical builds side by side first: if they differ by as much, the comparison means nothing.
Check That the Work Happens
A benchmark is only as good as its evidence that the code under test did what it claims. Two ways this went wrong while measuring Solid's signals:
- The wrong build. Under Node,
solid-jsresolves to its server build, in which effects never run. Signals then look almost free. Import the browser build (solid-js/dist/solid.js) by path, or alias it in the bundler. - Deferred work. Inside the body of Solid's
createRoot, every effect is deferred until the body returns. Frames timed there ran zero effects across 18,000 frames, and reported signals costing well under half their real cost. Build inside the root, but time outside it.
The same applies to any variant built by patching code: check that the patch applied. A patched build that silently matched nothing once produced a 10% "difference" between two identical bundles.
Rule: assert that the work happened. Count effect runs, or read back a value the work should have written, and fail the run if it did not. The Solid cases here write a signal and check that its effect ran before timing anything.
Measuring Allocation
Allocation per frame shows how much garbage a frame leaves for the collector, which is what the hot path rules are about. Node exposes the heap's size (process.memoryUsage().heapUsed) but no running total of bytes allocated, so the measurement works by making sure nothing is collected while it watches:
- Start Node with
--expose-gcand a large young generation (--max-semi-space-size=128). - Force a full collection, then note the heap size.
- Run 10,000 frames, and note the heap size again. With no collection in between, the growth is exactly what the frames allocated. Keep the window the same length however long a frame takes: a game does not allocate the same in every frame, and windows sized to half a second reported 2.5 times as much for one demo.
- Watch for collections with a
PerformanceObserverforgcentries. If one ran, the window is thrown away and retried with fewer frames. - Subtract the same measurement of an empty frame, which is the measurement's own cost.
Two checks keep it honest:
- A deliberately wasteful case. One view builds a template string and an
array.map()result every frame. If the measurement reports zero for it, the measurement is broken. - The benchmark's own variables. A module-level
letholding a non-integer number allocates a new heap number on every+=. A benchmark that adds its results to such a variable reports that allocation against every case. Keep results in aFloat64Array, which stores numbers in place.
For garbage collection over time, count the gc entries and their durations over a fixed number of frames with Node's default heap settings. For memory kept alive, force a collection before and after building many items and keep a reference to them.
Vary One Thing at a Time
A frame's cost usually has more than one part: something per container, and something per dynamic property within each container. Dividing the frame time by one of them hides the other. An early result here reported "nanoseconds per property" this way. Varying the dynamic properties per container separately from the number of containers showed that each container costs a fixed amount, plus about 1-2 ns for each dynamic property (a static property, set once at construction, costs nothing per frame).
Rule: vary one thing (the number of containers, the dynamic properties per container, the share changed per frame) while holding the others fixed, and only then choose the unit to report.
Read Profiles With Inlining in Mind
V8 inlines small functions into their callers. With inlining on, almost all samples in a CPU profile land on the outermost function of the frame, so the profile says little about where time goes inside it.
node --cpu-prof --cpu-prof-interval 25 bench.mjs
node --cpu-prof --max-inlined-bytecode-size=0 bench.mjsThe second form disables inlining, which restores per-function attribution but slows everything down.
Rule: use profiles taken with inlining disabled for relative shares only, never for absolute costs. --trace-deopt confirms whether anything is deoptimised after warm-up.
Measure the Whole Frame
Time everything a frame does to keep the view in step with the model: the model's changes plus the refresh, not just the refresh. Signals and events do their work when the model changes, so timing only the view side makes them look free.
State plainly what the measurement excludes. The measurements in this repo exclude rendering (Pixi's transform updates and draw calls), which is likely to be larger than anything measured when values change.
Running the Repo's Benchmarks
The benchmarks live in the repo's packages/benchmarks/ directory, one suite per topic. Each suite's cases run under plain Node, each in its own process: two processes per case, or three when the first two disagree by more than 5%. Timed cases run one process at a time. Cases that only count (bytes allocated, memory kept alive) run several processes at once, since sharing the machine cannot change a count. A full run of every suite takes about 25 minutes.
npm run bench # list the suites
npm run bench -- reactivity # run one suite
npm run bench -- reactivity approach=solid # only the matching cases
npm run bench -- reactivity --runs=5 # always this many processes per case
npm run bench -- all --save # every suite whose inputs changed, saving the results
npm run bench -- all --save --force # every suite, changed or not
npm run bench -- all --save --extended # every suite, its slow extended cases includedA save skips a suite whose inputs are unchanged since its saved results: the bundled code it measures, its cases, the harness, and the Node, Pixi and Solid versions. Some slow cases that answer a settled question are in an extended tier, run only with --extended; a save without it keeps their previous results, marked † with the date they were measured.
--save writes packages/benchmarks/results/<suite>.json (every run's numbers, the machine and versions they were measured on, and a hash of the inputs) and <suite>.md (the tables). Performance Measurements includes those tables directly, so re-running with --save updates the page. The suites are described in packages/benchmarks/README.md.