Files
Matt Aitken 1a033b665b fix(webapp,core): retry run resume through transient database outages (#4161)
## Summary

When the platform database is briefly unreachable while a run is
resuming from a wait, the run no longer fails with
`TASK_EXECUTION_ABORTED`. The worker now retries the resume through the
outage instead of aborting on the first blip.

## Root cause

Resuming a run calls the engine's `continue` worker-action endpoint.
That route caught every error and returned a `422`, which the worker's
HTTP client treats as non-retryable. So a transient Prisma
infrastructure error (for example `P1001` "Can't reach database server")
was flattened into a permanent failure: the worker gave up, force-killed
the run process, and completed it with `TASK_EXECUTION_ABORTED`.

## Fix

- The `continue` route now lets infrastructure errors propagate to the
generic 500 handler (message scrubbed, and retryable by the worker's
HTTP client), the same treatment the trigger path already gives them via
`isInfrastructureError`. Genuine validation errors (snapshot mismatch,
invalid state) still return `422`, so a stale retry stays non-retryable.
Resuming is idempotent server-side (guarded by the snapshot id), so
retrying is safe.
- The worker's `continueRunExecution` calls (both the
runner-to-supervisor and supervisor-to-engine hops) retry with a longer,
jittered backoff so they can ride out an outage lasting tens of seconds,
and the jitter keeps a fleet of resuming runs from stampeding the
database the moment it recovers.

Builds on #3960, which scrubbed the leaked message on these routes but
left the status non-retryable.

No changeset: this is a server-side behaviour fix recorded via
`.server-changes`. The `@trigger.dev/core` edits are internal run-engine
worker plumbing, not a public API change.
2026-07-07 11:51:59 +01:00
..

Server Changes

This directory tracks changes to server-only components (webapp, supervisor, coordinator, etc.) that are not captured by changesets. Changesets only track published npm packages — server changes would otherwise go undocumented.

When to add a file

Server-only PRs: If your PR only changes apps/webapp/, apps/supervisor/, apps/coordinator/, or other server components (and does NOT change anything in packages/), add a .server-changes/ file.

Mixed PRs (both packages and server): Just add a changeset as usual. No .server-changes/ file needed — the changeset covers it.

Package-only PRs: Just add a changeset as usual.

File format

Create a markdown file with a descriptive name:

.server-changes/fix-batch-queue-stalls.md

With this format:

---
area: webapp
type: fix
---

Speed up batch queue processing by removing stalls and fixing retry race

Fields

  • area (required): webapp | supervisor | coordinator | kubernetes-provider | docker-provider
  • type (required): feature | fix | improvement | breaking

Description

The body text (below the frontmatter) is a one-line description of the change. Keep it concise — it will appear in release notes.

Writing guidance

These entries are public-facing - they ship verbatim in user-visible release notes. A few rules to keep them clean:

  • One sentence is usually enough. The body is the bullet in the changelog. If you need a paragraph, you're probably describing the implementation rather than the change.
  • Describe behavior, not implementation. Skip internal scopes, middleware names, library specifics, framework internals. Users care about what's different for them, not how it's wired.
  • Never name internal tools or infra. Observability stacks, internal services, infra components, monitoring backends, CI surfaces, AWS specifics - none of these belong in user-facing notes.

Lifecycle

  1. Engineer adds a .server-changes/ file in their PR
  2. Files accumulate on main as PRs merge
  3. The changeset release PR includes these in its summary
  4. After the release merges, CI cleans up the consumed files

Examples

New feature:

---
area: webapp
type: feature
---

TRQL query language and the Query page

Bug fix:

---
area: webapp
type: fix
---

Fix schedule limit counting for orgs with custom limits

Improvement:

---
area: webapp
type: improvement
---

Use the replica for API auth queries to reduce primary load