Commit Graph

4050 Commits

Author SHA1 Message Date
Daniel Sutton 4c2c25511b test(webapp): split triggerTask engine test into per-concern files (#4167)
The engine `triggerTask` suite was a single 2447-line file with 23
`containerTest` cases, each spinning its own Postgres + Redis. vitest
shards by whole file, so all 23 container setups landed on one shard and
dominated its wall-clock. The recorded entry in `test-timings.json`
badly under-counts the real cost (it does not capture the
per-`containerTest` container startup that dominates on CI), so the
duration-sharding sequencer treated the file as light and stacked it,
producing one ~21 minute shard.

Splitting does not reduce the number of container setups; it lets those
23 cases distribute across shards instead of stacking on one. The webapp
unit-test stage is gated by its slowest shard, so this cuts the stage's
wall-clock roughly in half.

## CI timing (before vs after)

Real CI wall-clock of the `Unit Tests: Webapp` shards (`--shard=i/10`).
"Before" is sampled from recent runs on other branches (unsplit file,
from `main`); "after" is this PR.

| Shard | Before (s) | After (s) |
|------:|-----------:|----------:|
| 1  | 250 | 359 |
| 2  | 444 | 411 |
| 3  | 497 | 659 |
| 4  | **1257** | 284 |
| 5  | 545 | 641 |
| 6  | 284 | 644 |
| 7  | 244 | 214 |
| 8  | 340 | 445 |
| 9  | 188 | 395 |
| 10 | 234 | 567 |
| **Slowest shard (gates the stage)** | **~1247s (≈21m)** | **659s
(≈11m)** |
| Sum of all shards | 4283 | 4619 |

Before: shard 4 is the long pole at 1237s / 1247s / 1257s across three
sampled runs (the `triggerTask` file plus whatever else the packer put
with it). After: the six pieces spread across shards, the slowest drops
to 659s. The small rise in summed time is the extra per-file container
startup, paid in parallel across shards, so the gating number still
falls by about 10 minutes.

## Change

Split into six per-concern files that share a `triggerTaskTestHelpers`
module (the `vi.mock` calls stay per-file, since vitest hoists them):

- `triggerTask.test.ts` (3): trigger + concurrencyKey coercion
- `triggerTask.idempotency.test.ts` (4): idempotency + queue resolution
- `triggerTask.debounce.test.ts` (4): retries + debounce validation
- `triggerTask.mollifier.test.ts` (4): mollifier call-site behaviour
- `triggerTask.metadataCache.test.ts` (4): DefaultQueueManager task
metadata cache
- `triggerTask.residency.test.ts` (4): child run residency inheritance

All 23 cases are preserved. The file's `test-timings.json` entry is
split across the new files so bin-packing stays balanced.

While rewriting these files, cleanup was moved to `onTestFinished(() =>
engine.quit())` so an `engine`/`Redis` leaked on a failing assertion no
longer persists on the worker-scoped Redis and cascades into later cases
(`hookTimeout` raised to 60s so the after-cleanup gets the full budget).
Prisma lookups switched from `findUnique` to `findFirst` to match the
repo convention.

Verified: all six files run green locally (23/23), oxlint and oxfmt
clean.
2026-07-06 13:03:27 +01:00
Daniel Sutton e4ae8cbcd4 fix(run-engine,run-store,webapp): stop split-mode waits hanging on resume (#4164)
## Summary

On the run-ops database split, a run that waits (`triggerAndWait`,
`batchTriggerAndWait`, `wait.forToken`) could hang forever after its
wait had already completed. The runner reads a resume from
`/snapshots/since` exactly once: if that read returned the resume
snapshot without its completed-waitpoints, the runner logged "executing
without completed waitpoints", advanced its cursor, and never re-read
it, so the awaiting run never continued.

## Root cause

The resume snapshot and its completed-waitpoint rows were written as two
separate commits. This regressed when the split replaced Prisma's atomic
nested `connect` with an FK-free insert (in
[#4163](https://github.com/triggerdotdev/trigger.dev/pull/4163)), and
`/snapshots/since` is served from a read replica. A fetch landing in the
sub-millisecond gap between the two commits, or a multi-reader replica
serving the snapshot from a different point in time than its join rows,
delivered an empty resume. Because the runner consumes each snapshot
once and treats an empty resume as terminal, a single stale read was
fatal and produced a permanent, nondeterministic hang.

## Fixes

- Commit a snapshot and its completed-waitpoint links in one
transaction, restoring the atomicity the split removed.
- Repair the completed-waitpoints from the owning primary when a
multi-reader replica serves the snapshot without its join rows. This
covers single-waitpoint resumes, which carry no
`completedWaitpointOrder` and so were missed by the count-based repair.
- Read the primary in the checkpoint `WAIT_FOR_BATCH` pre-check, so a
batch that already resumed is not re-suspended into a stall.
- Fall back to the primary when a waitpoint token misses both read
replicas, so a token completed immediately after it was minted no longer
returns a spurious 404.
- Route batch-item creation by `batchTaskRunId`, consistent with the
batch-completion count and the row's foreign key.
- Reject control-plane-only relation selects on the dedicated schema
with a clear error instead of an opaque Prisma failure, and stop
`createDateTimeWaitpoint` bypassing residency routing through a caller
transaction.

Verified against the deployed split topology: a resume snapshot and its
completed-waitpoints are now always delivered together, so the runner
can no longer drop a resume.
2026-07-06 10:55:02 +00:00
Oskar Otwinowski de65370fb9 feat(webapp): Directory Sync (SCIM) for Identity & Access (#4148)
Extend the SSO plugin contract for directory sync and apply membership
effects
from the accounts webhook worker: provision users in mapped groups (role
from
group mapping, else the org default role), deprovision on removal, and
keep a
sticky-removal tombstone so JIT never silently re-adds a removed user.
JIT and
Directory Sync coexist; roles default to Developer (the JIT default-role
picker
has no 'None'). Changing a group's role in the dashboard re-applies it
to that
group's current members immediately. The Directory Sync settings section
(group→role mapping, external-domain + manual-membership policy,
deferred Save)
appears once a domain is verified — independent of SSO — gated by the
hasSso
flag. The settings page polls the whole page while entitled with
override-aware
drafts so in-progress edits are never clobbered.
2026-07-06 09:31:14 +02:00
Daniel Sutton f101983a70 fix(run-store,run-engine): fix run-ops split hangs from wrong-store reads on the resume path (#4163)
## Summary

On the run-ops split, NEW-residency runs could hang. Time-based waits
(`wait.for`, `wait.until`, `delay`, waitpoint tokens),
`batchTriggerAndWait`, and attempt starts stalled and never resumed.
Each was a run-ops read or update that hit the wrong database: either
the owning store's read replica when it needed read-your-writes, or the
wrong store entirely because it routed by an id that does not encode
residency.

## Fixes

**Waitpoint resume (the main hang).** The managed resume path reads a
run's completed waitpoints by snapshot id
(`findSnapshotCompletedWaitpointIds`). Snapshot ids are cuids, which
always classify to the legacy store, so a NEW run's join rows (which
live on the new store) were never found. The resumed run saw zero
completed waitpoints and hung. It now fans out across both stores and
merges, like its sibling readers.

**Batch completion.** Batch item completion
(`updateManyBatchTaskRunItems`) routed by the item id, which is also a
cuid, so a NEW batch's items were updated on the wrong store, matched
zero rows, and the batch was treated as already complete (its parent's
`batchTriggerAndWait` then hung). It now routes by the batch id, which
does encode residency, matching the sibling `countBatchTaskRunItems`.

**Read-your-writes on the resume path.** The block-time
pending-waitpoint check (`countPendingWaitpoints`) and the attempt-start
lock check (`findRun` in `startRunAttempt`) both read the owning store's
replica with no read-your-writes guarantee, so a just-committed
waitpoint completion or dequeue lock could be missed under replica lag
and strand the run. Both now read the owning primary.

Each fix ships with a two-database store or engine test that reproduces
the hang and passes with the fix.
2026-07-05 22:39:40 +00:00
Daniel Sutton 712c7c3b1a feat(webapp): add RUN_REPLICATION_RUN_OPS_DATABASE_URL for the runs-replication source (#4160)
## Summary

The run-ops runs-replication source now takes its connection URL from
`RUN_REPLICATION_RUN_OPS_DATABASE_URL`, required whenever the run-ops
split is enabled.

The runs replicator speaks the Postgres streaming replication protocol,
which cannot run through a transaction pooler, so it needs its own
direct endpoint separate from the app's `RUN_OPS_DATABASE_URL` (which
may point at a pooler). When the split is on and this is unset, boot
fails via `SplitReplicationMisconfiguredError` rather than silently
falling back to a wrong endpoint.
2026-07-05 13:29:22 +00:00
Daniel Sutton 092b9ef07a fix(run-ops): DNS-safe, sortable base32hex run id (replace base62 KSUID) (#4154)
## Problem

The run-ops split mints NEW-store run ids as **27-char base62 KSUIDs**.
The supervisor writes the run id into the Kubernetes pod name
(`runner-<id>`), and pod names must be DNS-1123 labels (lowercase
`[a-z0-9-]`) — so uppercase base62 ids make k8s reject the pod (422) and
**those runs never launch** (they loop in `PENDING_EXECUTING` until the
heartbeat-stall handler nacks them, forever). `.toLowerCase()` can't fix
it: base62 has both `A`(10) and `a`(36) as distinct symbols, so folding
collides distinct ids and destroys sort order.

## Fix: change the encoding, not the structure

Mint a **26-char lowercase base32hex** run id:

```
run_<24-char base32hex core><region char><version char>
      [ 6-byte ms timestamp ][ 9 CSPRNG bytes ]
```

- **base32hex** (RFC 4648 §7, alphabet `0-9a-v`): lowercase,
order-preserving, DNS-safe; 15 bytes → exactly 24 chars, no padding.
Hand-rolled encode/decode (no new dependency).
- **48-bit ms timestamp** in the leading bytes → plain string sort ==
creation order at millisecond resolution.
- **72 bits CSPRNG** entropy; PK unique constraint is the backstop (no
retry loop).
- **region / version** are raw positional chars (read via one `charAt`
before decoding/routing), version = `"1"`.

DNS-safe from birth and hyphen-free, so **firekeeper is unchanged** —
`runner-<id>-attempt-N` → strip `runner-`, cut at first hyphen still
recovers the exact id incl. region+version.

## Residency discriminator: length → version char

`classifyKind`/`classifyResidency` (`runOpsResidency.ts`) previously
distinguished NEW vs LEGACY by **id length**. That gets ambiguous with a
third format. It now discriminates on the **version char at a fixed
position** (`isRunOpsIdBody`: 26 chars, `[25] === "1"`, base32hex
alphabet) → NEW; everything else → LEGACY. Total, never throws. The
`Residency` (NEW/LEGACY) contract the routing store consumes is
unchanged; the `"ksuid"` `ResidencyKind` label is retained only because
it's the persisted `runOpsMintKsuid` feature-flag value.

## Scope / verification

- Generator + discriminator in `@trigger.dev/core` isomorphic; mint path
+ all id-shape call sites swept (~40 webapp files); changeset added
(`@trigger.dev/core` patch).
- Core unit tests (encode/decode round-trip + property, generator shape,
ms sort-order incl. intra-second, parse partitioned-vs-legacy,
firekeeper round-trip): **24 pass**. `@trigger.dev/core` builds; webapp
typechecks; format/lint clean.

## Open decisions (flagged, not silently chosen)

1. **Backward-compat**: existing 27-char base62 KSUID runs now classify
LEGACY. On test cloud these are the broken/looping runs that never
completed, so this is acceptable — but worth a conscious call before
prod. No transitional length-recognition added (keeps the discriminator
clean).
2. **Storage collation**: the sort guarantee is byte-order — if the
run-ops id column is `TEXT` with default locale collation it's silently
not honored. Confirm whether `COLLATE "C"` / `BYTEA` is needed on the
run-ops schema.
3. **Region sourcing** wiring — see `regionCharForRegion` /
`REGION_CODES`.


---

## ⚠️ Required migration — deploy in lockstep

This PR renames a persisted feature-flag key/value and an env var. These
are **not** changed by the code alone and must be migrated when this
deploys, or affected orgs silently fall back to `cuid` minting (no crash
— `defaultValue: "cuid"`):

1. **Env var** (terraform): `RUN_OPS_MINT_KSUID_ENABLED` →
`RUN_OPS_MINT_ENABLED` (carry the value over).
2. **DB** `organization.featureFlags`: migrate both the key and value
together:
   - key `runOpsMintKsuid` → `runOpsMintKind`
   - value `"ksuid"` → `"runOpsId"`

Until an org's flag row is migrated, its `runOpsMintKind` lookup misses
and it mints `cuid` (legacy) — so no NEW-store ids for that org until
the data lands.

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 10:05:54 +01:00
Daniel Sutton d977691219 fix(run-store): route caller-passed read clients to the owning store's primary (#4153) 2026-07-05 00:17:24 +01:00
Daniel Sutton 119189f31a fix(webapp): accept all valid run and batch IDs in dashboard filters (#4152)
## Summary

The **Run ID** and **Batch ID** filters on the runs list, batches list,
and logs view rejected valid IDs. The input showed an error and the
**Apply** button stayed disabled, so filtering by an affected run or
batch ID from the dashboard was impossible.

The filter validators hard-coded exact friendly-id character lengths.
Friendly IDs come in three generations that all still exist in the data
(`<prefix>_` plus a 21-char nanoid, a 25-char cuid, or a 27-char ksuid),
and the hard-coded lengths never covered all three at once.

## Fix

All the ID filter validators (run, batch, waitpoint, schedule) now share
one helper, `makeFriendlyIdValidator`
(`apps/webapp/app/utils/friendlyId.ts`), which validates by prefix plus
a base62 body of any known generator length (21 / 25 / 27). The cuid and
ksuid lengths are sourced from core so the helper tracks any future
change to those formats. Unit tests assert it accepts the output of the
real id generators and rejects malformed input.

Downstream was already unaffected: run/batch route params and
URL-applied filters use unconstrained validation, so only the manual
filter inputs needed the fix.
2026-07-04 19:57:32 +01:00
Daniel Sutton 962bc48738 feat(run-ops): automatically migrate the dedicated run-ops database (#4150)
## What

Adds the ability to **automatically migrate the dedicated run-ops
database** (the NEW DB in the run-ops split), matching how every other
database in the system is migrated. Follow-up to the run-ops split
activation.

## Changes

- **Migrate runner** — new
`internal-packages/run-ops-database/scripts/migrate.mjs`, exposed as
`db:migrate:deploy` / `db:migrate:status`. Connects via
`RUN_OPS_DATABASE_URL` (the same var the app uses) and expands `${VAR}`
refs like Prisma's dotenv.
- **Self-host** — `docker/scripts/entrypoint.sh` runs the run-ops
migration on boot when the DB is configured, gated by
`SKIP_RUN_OPS_MIGRATIONS`. Single-DB installs never set the URL, so it's
a clean no-op.
- **Single env-var family** — the run-ops DB is now addressed by one
canonical `RUN_OPS_*` family, connect path and migrations resolving the
identical URL:
  - `RUN_OPS_DATABASE_URL` (writer) — replaces `TASK_RUN_DATABASE_URL`
- `RUN_OPS_LEGACY_DATABASE_URL` — replaces
`TASK_RUN_LEGACY_DATABASE_URL`
- `RUN_OPS_DATABASE_READ_REPLICA_URL` — replaces
`TASK_RUN_DATABASE_READ_REPLICA_URL`
- the old `TASK_RUN_*` aliases, the `??` coalesce, the
`runOpsNewDatabaseUrl` indirection, and the migrate-only `directUrl` are
all removed (consumers read `env.RUN_OPS_DATABASE_URL` directly).

`directUrl` was dropped because it was only ever used by `prisma
migrate` (never the app runtime) to bypass a pooler for advisory locks —
premature here since the run-ops connection isn't wired to the app yet.
If a pooler is later introduced for the app, a direct URL can be
reintroduced then.

## Safety

- **Pure rename** — nothing deployed sets any `TASK_RUN_*` var (the
split isn't activated anywhere yet; `.env.example`, docker-compose, and
cloud already use `RUN_OPS_*`), so there is no config migration.
- **Single-DB / self-host** — no new required env var; entrypoint and
migrate are no-ops when `RUN_OPS_DATABASE_URL` is unset.
- **Cloud** — runs migrations as pre-deploy ECS tasks (companion cloud
PR), calling these same `db:migrate:deploy` / `db:migrate:status`
commands.

## Verification

- Live migration against a fresh scratch DB with only
`RUN_OPS_DATABASE_URL` set: both migrations applied, no `P1012`/`P1013`;
`${VAR}` expansion, idempotent re-run, `status`, and no-op skip all
pass.
- Schema parity 4/4; `typecheck --filter webapp` 18/18; affected
split/replication tests 34/34.

## Scope

This delivers automatic migrations only. Enabling the app to *use* the
new DB (setting `RUN_OPS_DATABASE_URL` + `RUN_OPS_SPLIT_ENABLED` on the
service) is a separate activation step.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 15:44:47 +00:00
Daniel Sutton c9f427e21f fix(replication): key logical replication leader lock on slot name (#4151)
## Problem

`LogicalReplicationClient` uses a Redlock leader lock to guarantee a
single active consumer per Postgres logical replication slot. The lock
resource was keyed on the client `name`:

```
logical-replication-client:${this.options.name}
```

A slot permits exactly one consumer, so the lock's job is to serialize
consumers **of a given slot**. Keying it on `name` breaks that whenever
two clients target the same slot with different names — most notably
across a rolling deploy where the client `name` changes but `slotName`
does not. Both acquire *distinct* locks, both consider themselves
leader, and the second to reach `START_REPLICATION` hits `replication
slot "<slot>" is active for PID <n>`. Because that query was
fire-and-forget and its failure was only logged (no retry), the consumer
stopped and replication stalled until the process was restarted.

## Fix

**1. Key the leader lock on `slotName`** — the actual single-consumer
resource:

```
logical-replication-client:${this.options.slotName}
```

Consumers of the same slot now contend on the same lock and hand off
cleanly across restarts/deploys; different slots stay independent.
`name` is kept for logging and the pg `application_name`.

**2. Self-healing resubscribe** (`resubscribeOnFailure`, opt-in) —
instead of logging-and-dying, a client re-subscribes with exponential
backoff after a lost election or a failed `START_REPLICATION`, so a
rolling deploy self-heals: the incoming pod retries until the draining
pod releases the slot, then takes over. Safety:
- `#cleanupAttempt()` unconditionally ends the pg client (freeing the
walsender) and releases the leader lock before rescheduling — retries
never leak connections/locks.
- `shutdown()` sets an intentional-stop latch re-checked after every
`await` in `subscribe()` (and aborts the lock-acquire spin), so a
resubscribe can never race or outlive an intentional shutdown.
- Backoff resets only on genuine stream start, so a permanently stuck
slot backs off to the ceiling and logs loudly rather than tight-looping;
an epoch guard neutralises stale `START_REPLICATION` catches.

Runs- and sessions-replication opt in and use `shutdown()` for all
intentional stops.

**3. Observability** — the admin runs-replication status route probed
the old name-keyed Redis key (would report `leader:false` for every
source after fix #1); now probes the slot-keyed key.

## Tests

`internal-packages/replication/src/client.test.ts` (real Postgres +
Redis containers):
- same-slot/different-name → second client must not double-lead or race
into "slot is active" (the regression)
- a failing `START_REPLICATION` retry loop must not leak connections or
locks
- `shutdown()` during an in-flight `subscribe()` must not leave a zombie
leader
- `subscribe()` after `shutdown()` re-arms `resubscribeOnFailure`
- self-heals once the leader releases the slot

Plus the multi-source wiring test updated to the slot-keyed lock keys.

## Rollout

With the self-healing resubscribe, this ships as a **plain rolling
deploy** — the incoming pods retry across the one-time lock-key
transition and take over once the old pods drain (a brief replication
stall that the durable slot replays on reconnect — no data loss). No
stop-before-start required.
2026-07-04 15:02:05 +00:00
Daniel Sutton 70bca82d84 feat(run-ops): activation — drop cross-DB FKs, provision run-ops DB, enable split (#4124) 2026-07-04 07:02:28 +01:00
Daniel Sutton 618b921a51 feat(run-ops): webapp routes — friendlyId reads, cross-seam token resolution, co-location writes (#4123) 2026-07-03 21:10:42 +01:00
Daniel Sutton 0a51341347 feat(run-ops): read presenters — de-join control-plane relations + read-through hydration (#4122) 2026-07-03 20:37:59 +01:00
Daniel Sutton 5be6a4fe38 feat(run-ops): ClickHouse multi-source replication fan-in + admin ops (#4119)
## What

Extends the ClickHouse runs-replication service to fan in from multiple
Postgres sources (the control-plane DB and the run-ops DB) instead of a
single source, plus the admin operations to run and observe it.

- **Multi-source fan-in** (`services/runsReplicationService.server.ts`,
new `runsReplicationInstance.server.ts`,
`runsReplicationGlobal.server.ts`): factors the replication service into
per-source instances and a coordinator so a single ClickHouse target is
fed from more than one Postgres source.
- **Admin ops** (`routes/admin.api.v1.runs-replication.status.ts`,
`admin.api.v1.runs-replication.backfill.ts`,
`v3/services/adminWorker.server.ts`): adds a status endpoint reporting
per-source replication state and updates the backfill entrypoint for the
multi-source shape.

## Why

PR7 of the run-ops split stack, and the final piece: once run state can
live in a separate run-ops DB (earlier PRs), the analytics replication
into ClickHouse has to consume both sources so runs remain queryable
regardless of residency. Behavior-changing for the replication service
internals; the ClickHouse-facing output is unchanged (still one runs
stream), and single-source operation is preserved when the split is not
enabled.

## Tests

New vitest coverage: `runsReplicationInstance.test.ts` (per-source
instance behavior) and `runsReplicationService.part8`/`part9` suites
exercising the multi-source coordinator. Testcontainers-backed
(ClickHouse + Postgres); no mocks.

## Notes

Draft, **stacked on #4118** (`runops/pr06-write-path`). Review that
first; this diff is against it.

Server-change / changeset note to be added at stack-assembly time.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-03 20:01:56 +01:00
Daniel Sutton 84f3e1b39c feat(run-ops): webapp write path — trigger/batch minting, idempotency routing, run lifecycle (#4118)
## What

Routes the webapp write path through the run-ops split seam:
trigger/batch minting, idempotency-key resolution, and the run-lifecycle
services now determine residency and dispatch writes to the correct
store.

- **Trigger & batch** (`runEngine/services/triggerTask.server.ts`,
`batchTrigger.server.ts`, `createBatch.server.ts`,
`streamBatchItems.server.ts`, `v3/services/batchTriggerV3.server.ts`):
mint ids with the run-ops-aware minting and route creation/streaming
through the store; batch children inherit the parent's residency.
- **Idempotency** (`runEngine/concerns/idempotencyKeys.server.ts` + new
`idempotencyResidency.server.ts`): idempotency-key lookup/dedup is
residency-aware so a keyed retrigger resolves against the store that
owns the original run.
- **Run lifecycle services** (`createCheckpoint`,
`createTaskRunAttempt`, `enqueueDelayedRun`, `expireEnqueuedRun`,
`finalizeTaskRun`, `resumeBatchRun`, `cancelDevSessionRuns`,
`executeTasksWaitingForDeploy`, `triggerFailedTask`): resolve their
target run through the store rather than a fixed client.
- **Reads that fan out from writes** (`runsRepository` +
`clickhouseRunsRepository`, `BulkActionV2` + batch read-through,
realtime `sessions`/`runReader`, alerts
`deliverAlert`/`performTaskRunAlerts`): route through the read-through
resolver.
- `9535ae63d` — resolves the parent run through an injectable run store
in `TriggerFailedTaskService`.
- `bf8f7c881` — drops the "known-migrated" concept from write-path and
read repos; residency is id-shape only.
- `515b897ea` — self-defaults `resolveWaitpointThroughReadThrough` to
the safe run-ops clients.

## Why

PR6 of the run-ops split stack. This is the write-path counterpart to
the read foundation in the previous PRs: with it in place, both reads
and writes route through the seam. Additive when the split is disabled
(id-shape resolution collapses to the control-plane client);
behavior-changing on the minting, idempotency, and lifecycle paths when
enabled.

## Tests

Large new/expanded vitest suite under `apps/webapp/test/` and colocated
service tests: trigger-task and batch-trigger store routing, residency
inheritance, idempotency dedup residency + legacy-authority, bulk-action
read routing, cancel-dev-session routing, alerts store routing,
runs-repository read-through, realtime session/run-reader read-through
and stream-registration routing, and the waitpoint read-through default.
Testcontainers-backed; no mocks.

## Notes

Draft, **stacked on #4117** (`runops/pr05-webapp-foundation`). Review
that first; this diff is against it.

Server-change / changeset note to be added at stack-assembly time.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-03 18:52:08 +01:00
Daniel Sutton 8465ac5ac3 feat(run-ops): webapp db topology, flags, and split-mode resolver wiring (#4117)
## What

Wires the run-ops split into the webapp: database topology, environment
flags, split-mode gating, and the control-plane resolver/cache layer
that the run-store and run-engine seams from the previous PR plug into.

- **DB topology & env** (`apps/webapp/app/db.server.ts`,
`env.server.ts`, `entry.server.tsx`): adds the run-ops database
clients/topology and the environment variables that configure and gate
the split.
- **runOpsMigration module** (new
`apps/webapp/app/v3/runOpsMigration/`): the webapp-side machinery —
`splitMode.server.ts`, `controlPlaneResolver.server.ts` +
`controlPlaneCache.server.ts`, `readThrough.server.ts`,
`crossSeamGuard.server.ts`, `distinctDbSentinel.server.ts`, id-minting
helpers (`mintBatchFriendlyId`, `runOpsMintKind`,
`resolveInheritedMintKind`), `runOpsCascadeCleanup.server.ts`, the split
read gate, and route/unblock catalogs.
- **Store/engine wiring** (`app/v3/runStore.server.ts`,
`runEngine.server.ts`, `runEngineHandlers.server.ts` + new
`runEngineHandlersShared.server.ts`): points the webapp's store/engine
construction at the resolver, and factors shared handler logic out so
both seams use one path.
- **Read-path touch-ups**: `runtimeEnvironment.server.ts`,
`eventRepository/index.server.ts`, `taskRunHeartbeatFailed.server.ts`,
`engineVersion.server.ts` route their run/environment lookups
read-through the resolver.
- `413a94511` — interlocks split mode against the native realtime
backend so the two aren't enabled in an incompatible combination (see
`.server-changes/run-ops-split-realtime-interlock.md`).
- `dc74c57fd` — drops the earlier "known-migrated" read layer; residency
is determined by id-shape only.

## Why

PR5 of the run-ops split stack. This is the webapp foundation layer: it
stands up the DB topology, flags, and resolver/cache the rest of the
stack depends on, and repoints webapp read paths through the resolver.
Additive when the split is not enabled (existing single-DB behavior
preserved behind flags); behavior-changing on the read-through paths and
the realtime interlock.

## Tests

New vitest coverage across `apps/webapp/test/` and colocated
`*.server.test.ts` files: db topology, split mode, split read gate,
cross-seam guard, mint cutover / flip latency, control-plane cache,
control-plane resolver, distinct-db sentinel, read-through loaders
(route loaders, run-detail loaders, `findEnvironmentFromRun`), and the
run-engine handlers. Testcontainers-backed; no mocks. `pnpm-lock.yaml`
synced for the two new webapp deps.

## Notes

Draft, **stacked on #4116** (`runops/pr04-store-engine`). Review that
first; this diff is against it.

Server-change / changeset note to be added at stack-assembly time.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-03 18:02:22 +01:00
Daniel Sutton c266e96c87 feat(run-ops): run-store routing seam + run-engine read seams (#4116)
## What

Introduces the run-store routing seam and the run-engine read seams that
let run lifecycle operations be dispatched to either the control-plane
database or a separately-generated run-ops database, depending on where
a run/batch resides.

- **run-store** (`internal-packages/run-store`): adds `runOpsStore.ts`
and substantially expands `PostgresRunStore.ts` so the store can resolve
residency and route reads/writes to the correct backing client.
`types.ts` grows the routing/residency types; `NoopRunStore.ts` is
removed.
- **run-engine** (`internal-packages/run-engine`): adds
`engine/controlPlaneResolver.ts` and routes the per-system read paths
(dequeue, enqueue, waitpoint, checkpoint, run-attempt, ttl, delayed-run,
execution-snapshot, pending-version, debounce, batch) through the
resolver/store instead of talking to a single Prisma client directly.
`engine/errors.ts`, `engine/types.ts`, and `engine/index.ts` are
extended to support injecting the store/resolver.

Three fixes are included on top of the seam work:

- `c6cadd85f` — routes read-your-writes to the owning store's
**writer**, not its lagging replica, so an operation immediately reading
back what it just wrote sees a consistent result.
- `05c912e05` — normalizes run-ops-generation Prisma errors to the
control-plane error class at the store **write boundary**, so
`instanceof` checks and the `P2002` → 422 handling continue to work
across the separately-generated run-ops Prisma client.
- `88d12907f` — resolves NEW-resident batches in
`ApiBatchResultsPresenter` by routing the batch read through the store,
so a dedicated-DB batch resolves instead of returning 404.

The change is heavily test-first: the bulk of the diff is new
unit/integration coverage for the store routing, residency, and each
run-engine system's control-plane resolver path.

## Why

PR4 of the run-ops split stack (PR1–PR3 land the ClickHouse
test-container and earlier plumbing). This PR is the read-path
foundation: it adds the seam and read-routing but leaves the write path
to route through the same seam in a later PR. Behavior-changing where
the three fixes above touch existing read-your-writes /
error-normalization / batch-resolution paths; otherwise additive (new
store module, new resolver, injectable dependencies with existing
single-client behavior preserved when no dedicated store is configured).

## Tests

Extensive new vitest coverage under `run-store/src/*.test.ts` (routing,
residency, dual-schema select, cross-generation error normalization,
read-after-write, idempotency dedup, mixed residency, waitpoint
co-location) and `run-engine/src/engine/**/*.test.ts` (per-system
`controlPlaneResolver` tests, injectability, block-edge residency,
waitpoint read residency, trigger-create routing, lifecycle router).
Testcontainers-backed; no mocks.

## Notes

Draft, **stacked on #4114** (`runops/pr03-clickhouse-tc`). Review that
first; this diff is against it.

Server-change / changeset note to be added at stack-assembly time.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-03 17:42:44 +01:00
Chris Arderne 86235e5cc7 feat(webapp): add Abort button to bulk actions list rows (#4144) 2026-07-03 15:42:13 +01:00
Wes Mason 58b114f4c9 feat(webapp): seed a local CLI personal access token (#4135)
## Summary

Getting the CLI talking to a local instance meant the browser magic-link
login, which is no good when you're driving things headlessly (an agent,
a container, or just no browser to hand). The seed already prints dev
secret keys for the batch-limit orgs, so it now also mints a personal
access token for the seeded `local@trigger.dev` user and prints a
ready-to-run `export TRIGGER_ACCESS_TOKEN=...` next to them.

Re-seeding stays idempotent: it decrypts and reprints the existing
`local-dev-cli` token rather than piling up a new one on every run.

<!-- GitButler Footer Boundary Top -->
---
This is **part 2 of 2 in a stack** made with GitButler:
- <kbd>&nbsp;2&nbsp;</kbd> #4135 👈 
- <kbd>&nbsp;1&nbsp;</kbd> #4137
<!-- GitButler Footer Boundary Bottom -->

---------

Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2026-07-03 14:10:15 +01:00
Chris Arderne 20f9c781f6 fix(webapp): db seed script (#4137)
## Summary

`pnpm run db:seed` failed with `SyntaxError: The requested module
'./app/models/organization.server' does not provide an export named
'createOrganization'`, even though that export exists. Renaming the seed
entry
point from `seed.mts` to `seed.ts` runs it as CommonJS and fixes the
failure.
No seed logic changes.

## Root cause

`seed.mts` is an ES module, but the server modules it imports (`.ts`
files, and
no package declares `"type": "module"`) resolve as CommonJS. tsx
compiles those
to CommonJS using esbuild's getter-based export shape
(`Object.defineProperty(exports, name, { get })`), which Node's
`cjs-module-lexer` does not detect when it links the ESM importer. The
named
exports look absent, so linking throws before any code runs.

The seed only ever needed the `.mts` extension: it has no top-level
`await` and
no `import.meta`. Running it as `.ts` keeps the whole import chain
CommonJS to
CommonJS and never crosses the lexer boundary. Verified the seed runs to
completion after the change.

Possibly caused by Node 22 upgrade?
2026-07-03 12:44:59 +01:00
Eric Allam ea9d6f7430 feat(webapp): pin the dashboard agent to a deployed version (#4128)
## Summary

The dashboard agent runs as a `chat.agent` in its own Trigger.dev
project, deployed separately from the app that starts its sessions. This
adds independent, non-disruptive deploys for it: a new workflow deploys
the agent with `--skip-promotion` (so a deploy never becomes the current
version on its own), and the app pins its sessions to a chosen version.

## How it works

`DASHBOARD_AGENT_VERSION` (unset by default) pins agent sessions to a
specific deployed version; when unset, sessions run on the project
environment's current version. The pin is passed on session start and
head start and is forwarded to every continuation run, so a pinned
session stays on its version for its whole life. Cutting over to a new
build becomes a config change (set the version) rather than a redeploy,
and rollback is flipping it back.

The workflow (`dashboard-agent-deploy.yml`) runs a leg per environment
(staging and prod), each gated by its own environment, and triggers on
pushes to `main` that touch the agent or its store (also available via
manual dispatch). Deploy versions are per-environment, so each
environment pins to its own leg's version.
2026-07-03 10:19:38 +01:00
nicktrn 0f349dd636 test(webapp): regression tests for resuming manually paused environments (#4127)
Follow-up to #4120, adding regression coverage for the pause/resume path
in `PauseEnvironmentService`.

Three tests against the real service with testcontainers Postgres (no
mocks), seeded org/project/environment rows, and the same
`AuthenticatedEnvironment` coercion production uses:

1. **resumes a manually paused env** - the actual regression: pause
leaves `pauseSource` null, resume must succeed. Verified this fails with
the exact pre-#4120 symptom (false billing-limit error) when the fix is
reverted locally.
2. **rejects resume of a billing-limit paused env** - the guard still
blocks manual resume while `pauseSource = BILLING_LIMIT`, and the env
stays paused.
3. **manual pause while billing-limit paused is a no-op** - returns
success without overwriting `pauseSource`, so billing-limit converge can
still find and unpause the environment.

Tests-only change, no runtime behavior touched.
2026-07-02 22:47:03 +01:00
Oskar Otwinowski c7f6ed501c feat(cli,core): add opt-in dev-only telnet log streaming (#4110)
Stream dev logs over a local telnet/TCP socket. `trigger dev` mirrors
its terminal output on port 6767 by default (override with
--telnet-logs-port or TRIGGER_DEV_TELNET_LOGS_PORT, 0 disables). webapp,
supervisor, and coordinator each expose an opt-in stream gated on a
per-service *_TELNET_LOGS_PORT env var. New
@trigger.dev/core/v3/telnetLogServer module (localhost-only,
backpressure-safe, plain-text) plus optional static Logger.onLog /
SimpleStructuredLogger.onLog sinks.

Then you (or your agent) can use `nc` to connect and filter out the
stream.
<img width="1103" height="239" alt="image"
src="https://github.com/user-attachments/assets/b4d47efc-8a57-4185-a159-10f2806627ae"
/>
2026-07-02 19:16:14 +00:00
James Ritchie 22a7e33df7 fix(webapp): render error page full screen and use Enter shortcut (#4121)
Makes the app error page always render full screen — it previously
inherited the width/offset of whatever container the error boundary was
mounted in (e.g. the centered `max-w-xs` column in the root boundary),
so `min-h-screen` alone couldn't fill the viewport. The root container
now uses `fixed inset-0 z-50` to break out and cover the full screen
regardless of nesting.

Also changes the "Go to homepage" shortcut from `Cmd/Ctrl+G` (which
collides with the browser's native "Find Again") to `Enter`.
2026-07-02 19:09:56 +01:00
nicktrn 8947c07a0c fix(webapp): allow resuming manually paused environments (#4120)
## Bug

Manually pausing an environment works, but resuming it always fails
with:

> This environment is paused because your organization reached its
billing limit. Resolve the limit on the billing limits settings page to
resume.

even when no billing limit is in effect. Once paused by a user, an
environment cannot be resumed at all.

## Root cause

A manual pause leaves `RuntimeEnvironment.pauseSource` as `NULL` (only
billing-limit enforcement sets `BILLING_LIMIT`). The resume path in
`PauseEnvironmentService` guards its `updateMany` with:

```ts
NOT: { pauseSource: EnvironmentPauseSource.BILLING_LIMIT }
```

Prisma's `NOT` on a nullable field translates to SQL `!=`, which
excludes `NULL` rows. So the update matches zero rows for every
user-paused environment, and the zero-count branch (meant to catch a
race with billing-limit pausing) returns the misleading billing-limit
error.

Introduced in #3996 (the guard is correct for `BILLING_LIMIT` rows; it
just also swallows `NULL`).

## Fix

Explicitly include `pauseSource: null` rows:

```ts
OR: [
  { pauseSource: null },
  { NOT: { pauseSource: EnvironmentPauseSource.BILLING_LIMIT } },
]
```

Billing-limit-paused environments are still blocked from manual resume,
both by the `getManualPauseEnvironmentResult` guard and by this clause.

## Verification

Reproduced locally: paused an environment via `PauseEnvironmentService`
(DB shows `paused = true`, `pauseSource = NULL`), resume returned the
billing-limit error with `updateMany` matching 0 rows. With the fix,
resume succeeds and the environment unpauses. Billing-paused rows remain
excluded by the same clause.
2026-07-02 18:58:35 +01:00
James Ritchie 478adb5d74 fix(webapp): truncate task title in landing page side menu (#4115)
## Summary

Long task names in the task landing page side menu pushed the **Test**
button off the edge of the panel instead of truncating. The heading now
truncates with an ellipsis so the Test button always stays in view, on
the standard and agent task pages.

## Root cause

The side menu lives in a fixed-width resizable panel with `overflow:
hidden`. Its header row is a grid item, and a grid item's default
`min-width: auto` lets it grow to its content's width. The title
`<span>` uses `truncate` (`white-space: nowrap`), whose min-content is
the full, untruncated name, so the header row expanded past the panel
and the Test button was clipped off the edge.

Adding `min-w-0` to the header container lets it shrink back to the
panel width so the title truncates. The scheduled task page already had
this class; the standard and agent pages did not.
2026-07-02 18:56:19 +01:00
Matt Aitken 6563793132 fix(webapp): dashboard timezone display + preference persistence (#4104)
## Summary

Two related timezone bugs in the dashboard.

1. The date/time tooltip could show a UTC offset label that contradicted
the time it displayed. A viewer whose machine clock differs from their
saved timezone (or when a date falls in the other DST phase) would see
something like `Local (UTC +0)` next to a value that isn't at +0.
2. A user's timezone preference silently failed to save whenever their
browser reported a zone like `UTC`, `Etc/UTC`, or `Asia/Kolkata`,
leaving their timestamps stuck in a previously-saved timezone.

## Offset label

The "Local" row formatted its time using the viewer's configured
timezone but computed the `(UTC +n)` label from `new
Date().getTimezoneOffset()`, the browser's offset at the current moment.
Those are two independent sources, so they disagreed when the configured
timezone differed from the machine, and also when the displayed date was
in the opposite DST phase. The label is now derived from the same date
and timezone used to render the row (via `Intl.DateTimeFormat` with
`timeZoneName: "longOffset"`), so it always matches the displayed time.

## Preference persistence

`/resources/timezone` validated the incoming zone against
`Intl.supportedValuesOf("timeZone")`, which lists only canonical zone
ids. Browsers report zones that aren't in that list via
`resolvedOptions().timeZone`, notably `UTC` (and `Etc/UTC`,
`Asia/Kolkata`, `GMT`), so those requests returned 400 and the
preference was never stored. Validation now checks whether the runtime
can resolve the zone at all, which accepts every real zone and still
rejects invalid input.

Added unit tests for both.
2026-07-02 15:04:09 +00:00
Matt Aitken fd4f02b2f8 fix(webapp): onboard new cloud orgs via plan selection; allow Free plan without GitHub verification (#4109)
## What & why

Two related fixes to how new cloud organizations get onboarded onto the
Free plan.

### 1. Route new cloud orgs through plan selection

New cloud organizations were created already activated, so they skipped
the plan-selection step and went straight to creating projects — which
meant their plan and usage limits were never set up. They're now created
deactivated and routed through plan selection, which activates them once
a plan is chosen. Self-hosted installs have no plan-selection step, so
they're activated immediately on creation and are unaffected.

The `Organization.v3Enabled` field is renamed to `isActivated` to better
describe what it now gates. It's mapped to the existing `v3Enabled`
column, so there's no data migration — only a schema/code rename.

### 2. Allow selecting the Free plan without GitHub verification

Choosing the Free plan no longer requires connecting and verifying a
GitHub account. The plan is applied immediately when selected. This
removes:

- the "Connect to GitHub" dialog and the GitHub-verified badge from the
plan picker
- the account-rejected state
- the now-unreachable GitHub-connect return routes

## Notes

- These changes pair with the corresponding change in the billing
service that applies the Free plan directly; they should be released
together.

## Testing

Verified locally end to end: a new cloud org is routed to plan
selection, the Free plan applies in one click with no GitHub step, the
org is activated, its usage allowance is provisioned, and it lands on
the new-project page.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-02 16:34:41 +02:00
Chris Arderne e49f74e0c6 chore(webapp): enable dev branches for everyone (#4106)
Remove the `devBranchesEnabled` feature flag and enable behaviour by
default.
2026-07-02 12:22:09 +01:00
Chris Arderne c7861be520 chore: activate no-unused-vars and import linters (#4096)
Once this is merged, oxlint is at a pretty sensible baseline.

**Enable `no-unused-vars`, `typescript/consistent-type-imports`, and
`import/no-duplicates` lint rules**

Turns on three previously-disabled oxlint rules across the monorepo and
fixes all violations:

- **`no-unused-vars`** – enabled as an error with standard ignore
patterns: unused function arguments are ignored by default (`args:
"none"`), variables/caught errors/destructured array elements prefixed
with `_` are allowed, and rest siblings are permitted.
- **`typescript/consistent-type-imports`** – enforced as an error; all
type-only imports now use the `import type` syntax.
- **`import/no-duplicates`** – enforced as an error; duplicate import
statements from the same module have been merged.

The remaining commits clean up the violations found across the codebase:
removing unused variables/imports/type aliases, adding `_` prefixes to
intentionally unused bindings, fixing duplicate imports, and converting
value imports to `import type` where appropriate.
2026-07-02 11:37:05 +01:00
nicktrn 6cf7677f18 chore(deps): bump @remix-run/* to 2.17.5 (#4102)
Bumps the `@remix-run/*` family in the webapp from 2.17.4 to 2.17.5 to
keep dependencies current.

2.17.5 pulls `@remix-run/router` 1.23.2 → 1.23.3, so the local
route-matching perf patch was rebased onto 1.23.3 (regenerated via `pnpm
patch`). It is functionally identical to the previous one - the only
difference is that 1.23.3 already hoists `decodePath` out of the match
loop upstream, so that hunk is dropped; the per-route-tree branch cache
and the compiled-path cache are unchanged. Also updated the
`@remix-run/dev>tar-fs` override key to track the new dev version.

Verified locally against latest main: `typecheck --filter webapp`
passes, `--frozen-lockfile` is consistent, and the dev server boots and
server-renders pages cleanly (route matching exercised via the patched
router).
2026-07-01 18:06:36 +00:00
Chris Arderne bfa902bd18 chore: enable more linters (#4080)
Re-enables ~15 oxlint rules that were blanket-disabled before.
2026-07-01 08:43:12 +01:00
Oskar Otwinowski 93755cbdbd feat(webapp): Identity & Access setup flow improvements (#4082)
- auto-reloading for setting up identity & access
- ui/ux improvements around the process of setting up connections for
the identity & access
2026-06-30 17:46:15 +01:00
Chris Arderne 5b01ebdb61 fix: conform error snuck through (#4089) 2026-06-30 16:45:20 +01:00
Chris Arderne 1e6cc19088 chore(webapp): upgrade @conform-to to v1 (#4044)
## Summary

Upgrades the dashboard form layer from `@conform-to` 0.9 to 1.x. No
behaviour change is intended; this is the conform API migration only.

conform 1.x peer-depends on `zod` `^3.21 || ^4`, so it runs on the
current zod 3 and is a prerequisite for upgrading the repo to zod 4:
conform 0.9 imports `ZodNativeEnum`, `ZodEffects`, and `ZodPipeline`,
all removed in zod 4, so the webapp cannot build against zod 4 until
conform is on 1.x. Landing this first (on zod 3) keeps the zod 4 PR
focused on zod alone.

## Testing
Did a bunch of local smoke tests and E2E playwright tests, several
rounds of different reviewers, all clean.
2026-06-30 16:51:15 +02:00
Katia Bulatova cd9d20cc99 fix(webapp): accept invites for orgs with many projects (#4043)
Invite acceptance could fail for cloud organizations with many projects
because the whole flow ran inside a single transaction and did too much
work before it completed. In larger orgs, that pushed the transaction
past its timeout and blocked the invite from being accepted.

This PR moves the expensive parts of invite acceptance out of the
transaction, excludes deleted projects from environment setup, fixes
error handling on /invites, and adds regression coverage for the failure
cases.
2026-06-30 14:35:35 +02:00
Katia Bulatova cc752cc375 fix(webapp): fetch run-scoped trace subtrees for large traces (#4024)
Fixed trace rendering for child and nested runs in large traces.
Dashboard and trace API responses now load the requested run's trace
subtree instead of depending on the run span appearing in the initial
trace slice.
2026-06-30 13:35:31 +02:00
Eric Allam c8d085a541 fix(webapp): clamp oversized toast messages and quiet expected pause errors (#4077)
## Summary

Two robustness fixes in the dashboard's error handling, found while
testing the billing-limit pause/resume flow.

## Toast cookie overflow

Toast messages are flashed into the `__message` session cookie, which
the session store rejects once the serialized cookie passes the
browser's ~4KB limit. Any call site that flashes a raw caught error (a
verbose database or validation message, for example) could turn a toast
into a failed request. `setErrorMessage` / `setSuccessMessage` now clamp
the message length, so a toast can never overflow the cookie. This
protects every toast helper at once.

## Pause/resume reporting

Resuming an environment that is paused by a billing limit is an
expected, user-actionable state, but `PauseEnvironmentService` threw it,
and the service's catch reports every throw at error level. It now
returns that case as a failure result, so callers still surface the
message to the user while genuine errors keep reporting.
2026-06-30 09:14:25 +01:00
Eric Allam aae1c6a512 fix(schedule-engine,webapp): log out-of-entitlements scheduled triggers as warnings (#4067)
## Summary

When a scheduled task fires for an organization that is out of
entitlements, the trigger can't proceed. That's an expected outcome, but
it was being logged at error level and surfaced as a failure.

## Fix

The trigger callback now classifies an out-of-entitlements result as its
own error type (`OUT_OF_ENTITLEMENTS`), and the schedule engine logs
both that and the existing queue-limit result as warnings rather than
errors. The run still doesn't fire and the `schedule_execution_failure`
metric still records the outcome (now tagged `out_of_entitlements`), so
nothing about observability or behavior changes beyond the log level.
2026-06-29 19:10:10 +01:00
Matt Aitken 729ed9d107 feat(webapp): gate billing limits on a dedicated permission (#4073)
## Summary

The billing limits page and the usage-limit banners now require a
dedicated `manage:billing-limits` permission instead of the broader
billing permission. This lets a role be granted control over billing
limits independently of subscription and payment management.

## Details

Both the loader and action of the billing limits settings page check
`manage:billing-limits`. The "Configure billing limit" and "Resolve"
actions in the limit banners (the no-limit-configured, grace, and
rejected states) gate on the same permission. The subscription and
billing pages keep using the existing billing permission.
2026-06-29 16:48:50 +01:00
Matt Aitken d720690073 feat(webapp): add region override to the bulk replay action (#4022)
## Summary

When replaying runs in bulk from a deployed environment, you can now
choose which region the replayed runs run in. The bulk action inspector
shows an "Override region" dropdown that defaults to "Don't override",
which keeps each run in its original region, so replaying a selection
that spans multiple regions doesn't silently re-route anything. Pick a
region and every matched run is replayed there instead.

The dropdown only appears for the replay action in a deployed
environment with more than one region available; cancel actions and
development environments don't show it.

## Design

The selected region is carried through the bulk action as a dedicated
`replayRegion` param, kept separate from the run-list selection filters
so it can't be confused with a region selection filter. When the action
runs, each replay passes it through to the existing region override on
the replay service, which already falls back to each run's original
region when no override is set. "Don't override" is a sentinel value
that the action normalizes away so the service only ever sees a real
region or nothing.

---------

Co-authored-by: Eric Allam <eric@trigger.dev>
2026-06-29 10:54:14 +01:00
James Ritchie 9f29eb2090 fix(webapp): update task filter icons and the scheduled run icon (#4065)
## Summary

Fixes the Tasks icon on the Runs, Logs, Errors, and metrics pages (and
the task-derived queue indicators) to use the correct "T" tasks icon,
and shows a clock icon for scheduled runs in the run inspector header.

## Before
<img width="1440" height="704" alt="CleanShot 2026-06-28 at 21 39 57@2x"
src="https://github.com/user-attachments/assets/d00d5cbf-4726-4695-825b-59f5bdd84529"
/>

## After
<img width="1370" height="604" alt="CleanShot 2026-06-28 at 21 38 59@2x"
src="https://github.com/user-attachments/assets/fb709073-0b9f-409d-a852-72979bf14da6"
/>
2026-06-28 21:50:14 +01:00
James Ritchie b2b4c510e2 feat(webapp): improve task and dashboard activity charts (#4064)
## Summary

Improves and unifies the run-activity charts by extracting a shared set
of chart primitives and adopting them on the three task landing pages
(agent, standard, scheduled), with the density and label fixes also
carried over to the dashboard and custom query charts.

Main changes a reviewer should know about:

- **Shared primitives (DRY).** New `ChartCard` (title +
maximize/fullscreen), `ChartSyncContext` (cross-chart hover + zoom
state), `useXAxisTicks` (width-aware tick selection),
`activityTimeAxis`, and `statusColors`, plus a server-side
`activitySeries.server.ts` holding `chooseBucketSeconds`, status
grouping, and the zero-fill helpers. The three task routes and both
presenters were refactored onto these, removing roughly 3x duplicated
tick logic, status-color tables, and bucket-ladder code.
- **Denser bars on short ranges.** Server-side bucketing now uses
`chooseBucketSeconds` (nice-interval ladder, ~72 target, capped at 120
buckets) instead of the hardcoded 1h/6h/1d ladder, so a 5m or 1h range
no longer collapses into a single bar.
- **Width-aware x-axis labels.** Labels are selected to fit the measured
plot width (always first + last, evenly spaced, de-duplicated by
rendered text), stay horizontal, and reflow on panel/window resize.
Y-axis values default to compact form (8K, 1.2M) in
`ChartBar`/`ChartLine`.
- **Synced hover line.** Hovering one of the agent page's three charts
draws a dashed vertical line at the same bucket on the *other* two, and
suppresses it on the hovered chart. It is opt-in via
`ChartSyncProvider`, so single-chart pages are unaffected.
- **Maximize button.** Each chart gets a fullscreen dialog toggle
(reuses the existing dashboard-widget pattern, `v` shortcut while
hovered).
- **Drag-to-zoom on task pages.** Dragging across a task chart sets the
Time/Date filter (`from`/`to` URL params, clearing `period`/`cursor`),
with a From/To tooltip shown during the drag.
- **Custom query charts.** Long categorical x labels (run IDs, task
names) middle-truncate and auto-rotate only when needed, and label
thinning is now width-aware for both bar and line variants. Dashboard
line-chart label density is also width-aware, tuned by a
`TIME_AXIS_LABEL_SPACING_PX` constant.
- **Tests.** 46 new unit tests for the pure logic (bucket selection,
tick spacing, time-axis formatting, zoom range, truncation).

## Intentionally unchanged

- **No click/drag zoom on the dashboard or custom query charts.**
Drag-to-zoom is wired up on the task landing pages only; zooming the
dashboard and custom charts is deliberately deferred to a separate
follow-up PR. A plain click (without a drag) on a task chart is a no-op.
- **The 25 mini activity charts on the Task list (`_index`) page are
untouched.** They are hand-rolled raw-Recharts sparklines kept
deliberately lightweight and do not use these primitives.
- **Other raw-Recharts sparklines are untouched** (the usage sparkline,
errors and prompts pages).
- **No ClickHouse query semantics changed** beyond the bucket-interval
parameter (same filters, same FINAL / `_is_deleted` handling).
- **Webapp-only.** No public package (`packages/*`) changes, so there is
no changeset; the `.server-changes/` entries cover it.

---

## Testing

Added 46 unit tests covering server-side bucket selection, width-aware
tick spacing, time-axis formatting, zoom-range math, and categorical
label truncation (`pnpm --filter webapp run test`), and `pnpm run
typecheck --filter webapp` passes. Manually exercised each task landing
page (agent, standard, scheduled) plus the dashboard and custom query
charts, stepping the Date/Time filter through 5m, 1h, 24h, 7d, and 30d
to confirm dense bars on short ranges, non-overlapping labels that
reflow on resize, the synced hover line across the agent charts, the
maximize button, and drag-to-zoom updating the filter.

---

## Changelog

The activity charts on the task landing pages and the dashboard and
custom query charts now share one set of reusable primitives. X-axis
labels are width-aware so they never overlap and reflow when a panel
resizes, y-axis values are abbreviated (8K, 1.2M), and short time ranges
render dense bars instead of collapsing into a single bar. Hovering any
agent chart mirrors a vertical line on the others, every chart gains a
maximize button, and dragging across a task chart zooms the Time/Date
filter. Long categorical labels such as run IDs and task names
middle-truncate and auto-rotate only when needed.

---



https://github.com/user-attachments/assets/6be09e38-3a0e-4947-b6e3-4839daa2fbe0
2026-06-28 17:26:43 +01:00
Eric Allam f1bd11a7ef feat(webapp): gracefully shut down the v3 engine behind a flag (#4017)
## Summary

Adds a single env flag, `DEPRECATE_V3_ENABLED` (default off), that
gracefully winds down the v3 engine (`RunEngineVersion.V1`). While it's
off nothing changes, so self-hosted instances still on v3 keep working.
When it's on:

- Triggers that resolve to v3 are rejected with a clear, actionable
error pointing at the [v4 migration
guide](https://trigger.dev/docs/migrating-from-v3), instead of silently
creating runs that never execute. This covers single triggers, batches,
scheduled fires, replays, and `triggerAndWait`, which all funnel through
one place.
- The legacy `trigger dev` websocket used by v3 CLIs is closed with an
upgrade message (v4 CLIs use a different dev transport).
- The v3 shared-queue consumer refuses to start, so no deployed v3 runs
are dequeued.
- The v3 run-lifecycle background jobs (heartbeat timeout, TTL expiry,
retry, resume batch/dependency, delayed-run enqueue, and scheduled
fires) become no-ops, so abandoned v3 runs stop generating database
load.

This builds on the existing deploy deprecation flag, which already
rejects v3 CLI deploys.

## Design

Enforcement is read through one helper, `isV3Disabled()`. Every gate
combines it with a per-run or per-project engine check (`isV3Disabled()
&& engine === "V1"`), so a v4 run that happens to reach a shared service
behaves exactly as before. v4 (V2) is never affected.

The flag is a hard switch, not a drain: when it's on, in-flight v3 runs
are abandoned in place rather than failed or expired, which is the
intended behaviour for the final shutdown.
2026-06-27 15:51:18 +01:00
Katia Bulatova b1987dc090 feat(webapp): billing limits — pause, reject, recovery, and settings UI (#3996)
## Summary

Adds Billing Limits to the webapp.

Customers can set a monthly spend cap. When usage crosses the limit,
billable environments enter a grace period. If the limit is not resolved
before grace expires, new triggers are rejected until the organization
increases or removes the limit.
2026-06-26 17:12:53 +02:00
Eric Allam 5d994577d6 fix(webapp): stop hydration errors on the Tasks page (#4058)
The Tasks page logged a burst of React hydration errors (#421, "this
Suspense boundary received an update before it finished hydrating") on
every load, one per task row for the Running and Activity cells. The
page still worked, but it spammed the console.

The Running and Activity (24h) cells stream in via Remix `defer()` +
`<Suspense>`/`<Await>`, two boundaries per row. A streamed boundary
stays in React's "hydrating" state until its data arrives; if the
backing queries are slow enough that the data is still in flight after
the page loads, a normal background re-render (a server-sent-events
update, a panel layout effect, a revalidation) hits the boundary and
React bails it to client rendering and throws #421. With N rows that is
2N errors. It never reproduced locally because those queries return
instantly there.

Fix: wrap the two cells in `ClientOnly` so they mount after hydration.
The stats still load asynchronously (the task list renders immediately),
but there is no longer an SSR boundary to bail. In the slow-query case
those cells already client-rendered (that was the bail); this just makes
it explicit and silent.

Verified by simulating slow stat queries against a local build: the
errors go from 2-per-row to zero, and the cells render correctly once
the data resolves.
2026-06-26 15:11:28 +01:00
Chris Arderne 4fde283e76 chore: format and lint webapp also (#4056)
#3977 added formatting and linting everywhere else.

This extends it to the webapp.
2026-06-26 13:02:53 +01:00
Chris Arderne b54201f986 chore: switch to oxfmt, oxlint - add ci checks (#3977) 2026-06-26 12:19:29 +01:00
Eric Allam 63bbea0d96 feat(webapp): move the in-dashboard agent launcher into the page header (#4053)
<img width="2400" height="1794" alt="chat-ui-closed"
src="https://github.com/user-attachments/assets/35016a72-c6d2-4b6a-8760-b1da5b4a166c"
/>
<img width="2400" height="1794" alt="chat-ui-open"
src="https://github.com/user-attachments/assets/25b3df60-9aa2-40fd-a1ba-1447eeee4c52"
/>

## Summary

The in-dashboard agent was launched from a button pinned to the
bottom-right of every page, which floated over page controls (for
example the run inspector's action bar). It now opens from a compact
"Chat" button on the far right of the page header, and the same button
toggles to "Collapse" while the panel is open. The panel and launcher
are labelled "Chat" in the UI.

The launcher only renders on env-scoped pages where the agent is enabled
(same feature flag gating), so it stays hidden for everyone who doesn't
have it. The existing "Ask AI" support button is untouched; it stays in
place until the agent is turned on by default.

## How it works

`DashboardAgent` (env layout) shares the open/close state through a
small context, and `NavBar` renders a launcher that self-hides whenever
that context is absent. No floating overlay, and the launcher can't
appear on pages where the agent can't open.
2026-06-26 11:09:26 +01:00
Eric Allam 9ef5cf0055 fix(webapp): stop showing the in-dashboard agent to admins by default (#4050)
## Summary

The in-dashboard agent button was rendered for all admins and
impersonators regardless of the `hasDashboardAgentAccess` flag, so it
appeared even where the agent is disabled (for example, floating over
the run inspector controls). It is now gated by the flag for everyone,
so it stays hidden until the flag is turned on.

## Rollout

Both levers default off, so nothing changes for users until deliberately
enabled:

- **Per-org:** set `hasDashboardAgentAccess` on an org's feature flags
to enable the agent for just that org.
- **All admins:** set `DASHBOARD_AGENT_ADMIN_PREVIEW=1` to give admins
and impersonators an everywhere-preview, independent of the per-org
flag.

Previously admins bypassed the flag unconditionally, which is why the
button showed up before the agent was ready to ship.
2026-06-26 11:09:14 +01:00