Files
Eric Allam 60d71da90e perf(webapp,run-engine): cut CPU on the engine-facing worker-action routes (#4746)
Cuts CPU on the `engine/v1/worker-actions/*` routes a managed supervisor
calls, and adds the benchmark harness the numbers come from.

Measured on a local stack: **on-CPU per completed run 9.07ms → 6.59ms
(−27%)**, busy fraction 45.6% → 33.8%, with every worker-action p50 down
23–27%. Load was 5,000 runs / 24 virtual supervisors / 90s window /
30,120 requests / 0 errors.

Query-count work from the same investigation is deliberately **not**
here — it will follow as a separate PR.

## The three changes

**1. Split the event-loop monitor in two (~14% of on-CPU, plus ~5pp of
GC).**

`eventLoopMonitor.server.ts` installs a global `async_hooks` hook:
`init` writes a `Map` entry for *every* async resource the process
creates, `before` calls `process.hrtime()` and `context.active()` on
every one. Enabling any async hook also puts V8 on the slow path for
promise instrumentation process-wide. `EVENT_LOOP_MONITOR_ENABLED`
defaulted to `"1"`, so this was the shipping configuration.

The blocked-loop detector is now opt-in (`EVENT_LOOP_MONITOR_ENABLED`,
default `0`). The event-loop *utilization* gauge — a single interval
timer with no per-request cost — moves to its own flag
(`EVENT_LOOP_UTILIZATION_MONITOR_ENABLED`, default `1`) and stays on, so
the useful half survives without the expensive half.

A/B under identical load:

| | monitor on | monitor off | change |
|---|---|---|---|
| on-CPU per run | 9.08ms | 7.25ms | −20% |
| GC self time | 9.80% | 5.05% | −4.75pp |
| dequeue p50 | 76.6ms | 62.8ms | −18% |
| attempts/start p50 | 56.3ms | 43.5ms | −23% |

**2. Bucket route matching by first static path segment (10.4% → 3.9% of
on-CPU).**

`patches/@remix-run__router@1.23.3.patch` already memoized flattened
branches and compiled path regexes. What remained was the linear scan:
`matchRouteBranch` walked the ranked branch list calling `matchPath` per
branch across 521 route files, so every worker-action request paid a
scan proportional to the whole route table.

Branches are now indexed by their lowercased leading segment, with one
always-considered list for branches whose leading segment is dynamic,
splat or optional (and for root/pathless paths). A request walks only
its own bucket merged with that list. Route-matching self time dropped
64% (3.6s → 1.3s over a 90s window).

Ordering is preserved exactly: both lists hold indexes into the already
rank-sorted branch array and are walked in ascending-index order, so the
first match found is the same branch the full scan would have found.
Bucketing lowercases on both sides, so case-insensitive matching still
resolves and `caseSensitive: true` routes are still rejected by
`matchPath` itself. A pathname whose own leading segment can't be
bucketed falls back to the full scan.

Verified equivalent to the unpatched matcher over 20,050 pathnames
(literal, dynamic, splat, optional, case variants, basenames,
percent-encoded) with zero mismatches.
`apps/webapp/test/routeMatchingPatch.test.ts` pins the matching
semantics rather than the optimisation, so it still passes without the
patch.

**3. Demote per-heartbeat and per-dequeue `info` logs to `debug`.**

These are the two highest-rate engine calls and each wrote a synchronous
structured log line on every request. Synchronous `console` writes can
block the loop when stdout backs up, which costs more than the ~1.3% CPU
share suggests.

## The harness

Two benchmarks, neither in the default suite (they run for minutes,
attach the V8 profiler, and report numbers rather than assert on them).
See `apps/webapp/test/bench/README.md`.

- `apps/webapp/test/bench/engineHttp.bench.test.ts` — spawns a real
webapp against throwaway Postgres/Redis containers, seeds a production
environment with a promoted managed deployment, and drives a closed-loop
supervisor pool through the full lifecycle. Profiling runs over CDP
rather than `--cpu-prof` so it covers only the measured window instead
of being swamped by boot, and `performance.eventLoopUtilization()` is
sampled *inside* the webapp process.
-
`internal-packages/run-engine/src/engine/bench/runEngineLifecycle.bench.test.ts`
— drives `RunEngine` directly, profiling enqueue and lifecycle
separately so engine cost isn't mixed with request-stack overhead.
- `apps/webapp/test/bench/analyzeProfile.ts` — dependency-free
`.cpuprofile` analyzer that symbolicates through the build's source maps
and ranks CPU by package, self time and total time. Percentages are
shares of on-CPU time (V8's `(idle)`/`(program)` excluded).

`startWebapp` gains `overrideEnv`, applied after the worker-disable
defaults, so the HTTP bench can re-enable the run engine worker that
drains the master queue into the worker queues a supervisor dequeues
from.

The local OTel collector gains a traces pipeline. It only defined a
metrics pipeline, so pointing `INTERNAL_OTEL_TRACE_EXPORTER_URL` at it
locally failed and the webapp silently fell back to the console span
logger.

## Configuration

For operators upgrading:

- `EVENT_LOOP_MONITOR_ENABLED` (now defaults to `0`) — the
per-async-resource blocked-loop detector. Set to `1` to restore the
previous behaviour and keep emitting `event-loop-blocked` spans.
- `EVENT_LOOP_UTILIZATION_MONITOR_ENABLED` (new, defaults to `1`) — the
`nodejs.event_loop.utilization` gauge. Unchanged in behaviour; it just
has its own flag now so it survives turning the detector off.

## Notes for review

- `pnpm-lock.yaml` changes only because the router patch content
changed, which changes its patch hash.
- One thing the profile ruled out: with a real OTLP collector receiving
spans, tracing costs ~1.7% of on-CPU at 100% sampling and ~0.8% at the
production rate. Span shipping is not a hidden cost, so nothing here
touches it.
- Caveats on the numbers: a laptop, not production hardware, so DB and
Redis *latency* are unrepresentative (client-side CPU is what's ranked);
single webapp process; throughput varies ~5% run to run, which is why
the claims rest on on-CPU per run rather than req/s.

## Verification

- 20,050-pathname router equivalence check vs the unpatched matcher,
zero mismatches
- `apps/webapp/test/routeMatchingPatch.test.ts` (12 cases) passes
- webapp e2e smoke suite (68 tests) passes through the patched router
- run-engine suites covering the snapshot/attempt paths pass
- `typecheck`, `format`, `lint`, `knip` clean
2026-08-21 11:53:16 +01:00

126 lines
7.0 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Patches
This directory holds [pnpm patches](https://pnpm.io/cli/patch) applied on install via
`pnpm.patchedDependencies` in the root `package.json`. Each `.patch` is a diff against the
published package. Most are small and self-explanatory from the diff; the non-obvious ones
are documented below.
---
## `@remix-run/router@1.23.2` — route-matching memoization
**File:** `patches/@remix-run__router@1.23.2.patch` (patches `dist/router.cjs.js`)
### What it does
Four changes to `matchRoutesImpl` / `compilePath`, all derived from work that depends only
on the **static** route manifest:
1. **Cache flattened + ranked branches per route-tree** (`WeakMap` keyed by the `routes`
ref). `flattenRoutes()` + `rankRouteBranches()` were recomputed on *every* `matchRoutes`
call across all ~436 webapp routes.
2. **Hoist `decodePath(pathname)` out of the branch-match loop** — it's loop-invariant but
was recomputed once per branch.
3. **Memoize `compilePath` compiled regexes** by `path|caseSensitive|end` (bounded `Map`,
cap 2000). The matcher RegExp was rebuilt on every `matchPath` call.
4. **Bucket ranked branches by first static path segment** (`WeakMap` keyed by the branch
array). Even with (1)(3), matching was still a linear scan calling `matchPath` on every
branch until one matched — O(route table) per request, now across 521 route files.
Branches are indexed by their lowercased leading segment, with one always-considered list
for branches whose leading segment is dynamic, splat or optional (and for root/pathless
paths). A request walks only its own bucket merged with that list.
Ordering is preserved exactly: both lists hold indexes into the already rank-sorted
branch array and are walked in ascending-index order, so the first match found is the
same branch the full scan would have found. Bucketing lowercases on both sides, so
case-insensitive matching still resolves and `caseSensitive: true` routes are still
rejected by `matchPath` itself. A pathname whose own leading segment can't be bucketed
falls back to the full scan.
Verified equivalent to the unpatched matcher over 20,050 pathnames (literal, dynamic,
splat, optional, case variants, basenames, percent-encoded). The semantics it depends on
are pinned by `apps/webapp/test/routeMatchingPatch.test.ts`, which asserts matching
behaviour rather than the optimisation, so it still passes without the patch.
### Why
Profiling the realtime runs feed under load (100 concurrent tag feeds, ~425 req/s) found
**~68% of webapp CPU was spent in react-router's `matchRoutes`** — re-flattening,
re-ranking, and re-compiling the entire route table on every request. It is **not** a dev
artifact: there is no `NODE_ENV` gate, and a `NODE_ENV=production` profile was identical
(67.9% vs 68.3%). The realtime feed's high request rate (each long-poll returns fast and
immediately re-polls) just amplifies a latent per-request cost that large route tables pay
everywhere.
Measured on a single instance, same load, before vs after this patch:
| | before | after |
|---|---|---|
| active CPU (self-time / window) | 28.3s | 18.5s (**34%**) |
| route-matching self-time | 19.2s | 7.5s (**61%**) |
| event-loop lag p99 | 322ms | 113ms (**65%**) |
| idle headroom | 26% | 52% |
The realtime machinery itself (router/hydrate/serialize/diff) was ~0% — the bottleneck was
entirely generic Remix request overhead.
Change (4) was added later, from a CPU audit of the engine-facing worker-action routes
(`apps/webapp/test/bench`). With (1)(3) already in place, route matching was still 10.4%
of on-CPU time on that path — the residual linear scan rather than any recompilation.
Bucketing took route-matching self-time from 3.6s to 1.3s (**64%**) over a 90s window at
~300 req/s.
### Upstream status (why we patch instead of upgrade)
This is a known, acknowledged inefficiency, and it is **only partially fixed in React
Router v7** — which we can't adopt without a full Remix 2 → RR7 framework migration.
- [Issue #8653 "Performance issues"](https://github.com/remix-run/react-router/issues/8653)
reported it (a user with 12k routes, ~67ms per match) and was closed as a dup of the
route-ranking discussion [remix#4786](https://github.com/remix-run/remix/discussions/4786).
- [PR #14866 "Optimize route matching performance with caching"](https://github.com/remix-run/react-router/pull/14866)
implemented *exactly this patch* (hoist `decodePath`, cache `compilePath`, cache
flatten/rank), claiming **~80% route-matching CPU reduction on a 400+ route app**. It was
**closed, not merged.**
- [PR #14967 "perf: cache flattened/ranked route branches"](https://github.com/remix-run/react-router/pull/14967)
is the partial fix that *did* ship (in v7): it caches only the branches, threaded via a
`precomputedBranches` param through the framework's server-runtime (~15% SSR gain). It
does **not** cache `compilePath` — that regex rebuild remains even on `main`.
([PR #14971](https://github.com/remix-run/react-router/pull/14971) added client-side wins.)
The maintainer's reasoning for closing the fuller PR (#14866), verbatim:
> "This is great as a `patch-package` optimization for those who want it, but we are
> actively working on integrating the more performant route-pattern library from Remix 3 so
> we'd rather just do the right 'fix' and ship the new algorithm instead of trying to
> band-aide perf improvements to the existing algorithm which was written with a very
> different set of constraints. Those constraints come from early v6 when it was only
> declarative mode so route trees were defined at render time and thus had to be
> re-flattened/re-ranked/re-compiled every time."
So: the re-compute-everything design is a holdover from early React Router v6 declarative
mode (route trees defined at render time, so recomputing was correct then). The maintainer
**explicitly endorsed patch-package as the interim approach** and is betting on the Remix 3
route-pattern rewrite for the real fix. This patch is that sanctioned stopgap — and it also
includes the `compilePath` cache the merged PR left on the table.
### Safety
Pure memoization of deterministic, internal-only values:
- `flattenRoutes`/`rankRouteBranches` and the compiled regexes depend solely on the static
route manifest; the cached values are never returned to or mutated by the framework.
- The compiled `RegExp` has no `/g` flag, so `.exec()` carries no cross-call state — safe to
share under concurrency.
- The branch cache is a `WeakMap` (collected with its route tree); the compile cache is
bounded at 2000 entries (route patterns are a static set; the cap only guards any dynamic
`matchPath()` use).
- Targets the **CJS** build (`dist/router.cjs.js`), which the webapp server loads at runtime
(`@remix-run/router` is not bundled into the server build).
### When to remove
Drop this patch if/when the webapp moves to React Router v7+ (which threads
`precomputedBranches` itself) or the Remix 3 route-pattern matcher lands. Re-profile at that
point — the `compilePath` cache may still be worth keeping since upstream never added it.