Commit Graph

7588 Commits

Author SHA1 Message Date
Matt Aitken ef8661375e feat(webapp): pass database writer and reader config to the RBAC plugin
📦 Preview packages (pkg.pr.new) / Build and publish previews (push) Has been cancelled
The host now resolves writer and read-replica URLs from its env (the same
fallback chain its own Prisma clients use) and hands them to an installed
RBAC plugin at create time, with separate connection limits for writes
(default 2) and reads (default 5). A plugin that owns its own database
client can then route per-request auth reads to the read replica instead
of holding them on the primary.
2026-07-10 16:33:42 +01:00
Iss 983bd03131 feat: support isSecret in syncEnvVars (#4203)
## What

Adds per-variable secret support to the `syncEnvVars` build extension.
Return `{ name, value, isSecret: true }` and the variable is stored as a
secret (redacted in the dashboard, value non-revealable), just like a
manually created secret env var. Secret and non-secret variables can be
mixed in one callback.

```ts
syncEnvVars(async () => [
  { name: "PUBLIC_API_URL", value: "https://api.example.com" },
  { name: "DATABASE_URL", value: "postgres://...", isSecret: true },
]);
```

## How

Env vars flow through the build pipeline as a flat name→value map, and
the import API's `isSecret` is per-call. So secret vars are carried
through the layer + manifest in parallel `secretEnv` / `secretParentEnv`
maps, and at deploy time they go up in a second `importEnvVars` call
with `isSecret: true` (the plain vars in the first call). The record
form (`{ KEY: "value" }`) is unchanged and stays non-secret.

## Commits

- `feat(core)`: carry secret env vars through the build layer + manifest
schema
- `feat(build)`: partition `isSecret` vars in `syncEnvVars`
- `feat(cli)`: merge secret layers and import them with `isSecret: true`
at deploy
- `test(build)`: cover the partitioning + document `isSecret`

## Testing

- vitest covers the partitioning (secret/non-secret × child/parent) and
that the record form stays non-secret.
- Verified against a local webapp that the deploy's import contract
stores the secret var redacted (`isSecret: true`) and the plain var
visible.

Closes TRI-11099
2026-07-10 11:03:28 -04:00
Chris Arderne 7faa52597d chore: format prisma schemas (#4224)
Creating a Prisma migration now formats its schema first, keeping
migration-related schema edits consistently formatted without adding
work to the repository-wide format command. Run `pnpm run format:prisma`
to format either schema on demand.
2026-07-10 14:01:11 +01:00
Eric Allam 02cf9c81ad fix(tsql): make JSON functions work on the output and error columns (#4221)
## Summary

A Query page (TRQL) query that pulls fields out of a run's `output` with
JSON functions (`JSONExtractString`, `JSONExtractInt`, `JSONHas`, and
the rest of the family) failed with "The first argument of function ...
should be a string containing JSON, illegal type: JSON". Those queries
now work.

## Root cause and fix

`output` is a native ClickHouse `JSON` column, but `JSONExtract*`,
`JSONHas`, `JSONLength`, and `JSONType` all expect a String containing
JSON text. The compiler already swaps in the column's String companion
(`output_text`) when a JSON column is selected or compared, but not
inside function-call arguments, so it emitted `JSONExtractInt(output,
'x')` against the native column.

The fix prints the companion column for the first argument of these
functions when it resolves to a bare JSON field, keeping the table alias
when qualified (so it works in JOINs):

JSONExtractInt(output, 'x') -> JSONExtractInt(output_text, 'x')
JSONExtractArrayRaw(assumeNotNull(output), 'y') ->
JSONExtractArrayRaw(assumeNotNull(output_text), 'y')

It also reaches through value-preserving passthrough wrappers like
`assumeNotNull(...)`, while leaving value-changing wrappers like
`toJSONString(output)` on the native column (that argument is already a
String). The swap is also semantically correct, not just a type fix:
`output_text` is the unwrapped data JSON that the TRQL `output` model
already represents, so field paths line up.

Covered by printer unit tests and a ClickHouse integration test that
runs the whole family (plus the wrapped and `toJSONString` cases)
against a real native-JSON column. Both new cases fail with the exact
"illegal type: JSON" error without the fix.
2026-07-10 12:44:52 +01:00
claude[bot] de536622c8 Add oxlint rule to catch thrown un-awaited redirect helpers (#4222)
##  Checklist

- [x] I have followed every step in the [contributing
guide](https://github.com/triggerdotdev/trigger.dev/blob/main/CONTRIBUTING.md)
- [x] The PR title follows the convention.
- [x] I ran and tested the code works

---

## Testing

- Verified the rule emits exactly five errors for un-awaited throws of
the known
async redirect helpers while ignoring awaited throws, returned promises,
and
  synchronous `redirect(...)`.
- Verified `--fix` inserts `await` in async functions and produces a
clean
  second lint run.
- Verified synchronous functions remain diagnostic-only so autofix
cannot
  introduce invalid syntax.
- Ran `pnpm run format`, `pnpm run lint`,
  `pnpm run typecheck --filter webapp`, and `git diff --check`.

---

## Changelog

Adds an Oxlint rule that prevents async redirect helpers from being
thrown
without awaiting their `Response`. Existing violations are fixed, the
autofix
is limited to async functions, and the plugin uses an explicit ESM
extension.

---

## Screenshots

See the test-results comment for CLI evidence.

💯


Link to Devin session:
https://app.devin.ai/sessions/e60ad7610773401da3d3040cf1252337
Requested by: @ericallam

---------

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Eric Allam <eric@trigger.dev>
2026-07-10 13:19:13 +02:00
Chris Arderne b4866f0184 docs: improve bulk actions docs (#4211)
- Combine SDK and dashboard bulk actions docs
- Fix API reference pages for bulk actions
- Fix weird rendering on bulk actions page

## Todo
- [ ] not sure about having the SDK+dashboard combined and under "Using
the dashboard"... need to find the right place
2026-07-10 11:49:45 +01:00
Wes Mason e6e8aeb993 docs(limits): document automatic payload offloading for triggers and batches (#4217)
## Summary

The limits page didn't spell out that large payloads offload to object
storage automatically, and its single "512KB" note conflated two
different thresholds. This clarifies the behaviour.

On the way in, the SDK uploads any trigger or batch-item payload over
128KB to object storage before sending, so large triggers and batches
don't hit the request body limit (`trigger` / `triggerAndWait` since
4.5.0, `batchTrigger` / `batchTriggerAndWait` since 4.5.2). On
retrieval, payloads and outputs over 512KB are stored in object storage
and returned as a presigned URL from `runs.retrieve`.
2026-07-10 11:31:42 +01:00
James Ritchie 32e5edbd04 fix(webapp): restore magic link login on the login page (#4220)
## Summary

Magic link login could appear completely broken: submitting your email
on the login page showed a stale "This email is unauthorized" error
instead of the "we've sent you a magic link" confirmation, even when the
address was fine.

This PR reverts
[#4215](https://github.com/triggerdotdev/trigger.dev/pull/4215) (whose
diagnosis and fix turned out to be wrong) and fixes the actual bug,
which was in how login errors are stored and consumed.

## Root cause

Two session bugs compounded on the login page:

- The `/login` loader read the flashed `auth:error` without committing
the session. A Remix flash is only consumed when the session is
committed after the read, so once any attempt flashed an error (for
example an address rejected on an instance with `WHITELISTED_EMAILS`
set), it stayed in the session cookie and reappeared on every later
`/login` visit, making successful attempts look like failures.
- The `/login/magic` action stored its validation and rate limit errors
with `session.set`, which survives every later read and commit, so those
errors stuck permanently.

[#4215](https://github.com/triggerdotdev/trigger.dev/pull/4215) had
instead diagnosed a server-only module leaking into the client bundle
and crashing navigation. Checking the shipped images' client bundles via
their sourcemaps shows `.server` modules were always stubbed out, so
that change fixed nothing and is reverted here.

## Fix

- `/login` reads the flashed error and commits the session when one was
present, so an error renders once and clears. The `redirectTo` branch
now surfaces the error too instead of leaving it in the cookie.
- `/login/magic` flashes its errors instead of `set`ting them.

Verified end-to-end on a live preview environment: a rejected address
shows the error once and a reload clears it; a valid address lands on
the confirmation screen with the address named; GitHub, Google, and SSO
login paths are untouched by this diff.
2026-07-10 10:59:10 +01:00
github-actions[bot] 9f76c92021 chore: release v4.5.3 (#4219)
🚀 Publish Trigger.dev Docker / typecheck (push) Failing after 3s
🚀 Publish Trigger.dev Docker / scan-webapp (push) Has been skipped
🚀 Publish Trigger.dev Docker / publish-webapp (push) Has been skipped
🚀 Publish Trigger.dev Docker / publish-worker-v4 (push) Has been skipped
🚀 Publish Trigger.dev Docker / units (push) Failing after 4s
🧭 Helm Chart Release / lint-and-test (push) Has been cancelled
🧭 Helm Chart Release / release (push) Has been cancelled
🚀 Publish Trigger.dev Docker / 📣 Dispatch main image (push) Has been cancelled
## Summary
1 improvement, 2 bug fixes.

## Breaking changes
- Removed support for the end-of-life v3 `trigger dev` CLI. Starting a
dev session with an old v3 CLI now returns an upgrade message instead of
connecting - upgrade to the v4 CLI to continue using `trigger dev`.
([#4198](https://github.com/triggerdotdev/trigger.dev/pull/4198))

## Bug fixes
- Fix TS2742 ("inferred type cannot be named") when exporting a
`chat.agent` from a project with declaration emit: `ChatTaskWirePayload`
and `ChatInputChunk` are now declared in the public
`@trigger.dev/sdk/chat` subpath, so inferred agent types emit portable
declarations and the wire types are directly importable.
([#4218](https://github.com/triggerdotdev/trigger.dev/pull/4218))

## Server changes

These changes affect the self-hosted Docker image and Trigger.dev Cloud:

- Reduce primary database load on the runs page by serving its
empty-state check from ClickHouse instead of Postgres.
([#4202](https://github.com/triggerdotdev/trigger.dev/pull/4202))
- Fixed submitting your email on the login page reloading back to an
empty form instead of showing the magic link confirmation screen.
([#4215](https://github.com/triggerdotdev/trigger.dev/pull/4215))

<details>
<summary>Raw changeset output</summary>

# Releases
## @trigger.dev/build@4.5.3

### Patch Changes

- Updated dependencies:
  - `@trigger.dev/core@4.5.3`
## trigger.dev@4.5.3

### Patch Changes

- Updated dependencies:
  - `@trigger.dev/build@4.5.3`
  - `@trigger.dev/core@4.5.3`
  - `@trigger.dev/schema-to-json@4.5.3`
## @trigger.dev/python@4.5.3

### Patch Changes

- Updated dependencies:
  - `@trigger.dev/sdk@4.5.3`
  - `@trigger.dev/build@4.5.3`
  - `@trigger.dev/core@4.5.3`
## @trigger.dev/react-hooks@4.5.3

### Patch Changes

- Updated dependencies:
  - `@trigger.dev/core@4.5.3`
## @trigger.dev/redis-worker@4.5.3

### Patch Changes

- Updated dependencies:
  - `@trigger.dev/core@4.5.3`
## @trigger.dev/rsc@4.5.3

### Patch Changes

- Updated dependencies:
  - `@trigger.dev/core@4.5.3`
## @trigger.dev/schema-to-json@4.5.3

### Patch Changes

- Updated dependencies:
  - `@trigger.dev/core@4.5.3`
## @trigger.dev/sdk@4.5.3

### Patch Changes

- Fix TS2742 ("inferred type cannot be named") when exporting a
`chat.agent` from a project with declaration emit: `ChatTaskWirePayload`
and `ChatInputChunk` are now declared in the public
`@trigger.dev/sdk/chat` subpath, so inferred agent types emit portable
declarations and the wire types are directly importable.
([#4218](https://github.com/triggerdotdev/trigger.dev/pull/4218))
- Updated dependencies:
  - `@trigger.dev/core@4.5.3`
## @trigger.dev/core@4.5.3

</details>

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
helm-v4.5.3 v4.5.3 v.docker.4.5.3
2026-07-10 08:42:34 +01:00
Eric Allam 25254d0201 fix(sdk): make inferred chat agent types portable for declaration emit (#4218)
## Summary

Exporting a `chat.agent` from a project with `declaration: true` failed
with TS2742: the inferred type of the agent references
`ChatTaskWirePayload`, which was declared in an internal module not
reachable through the package exports map, so tsc could only name it via
a file path into `node_modules` and refused to emit. Consumers had to
hand-mirror the wire type and annotate their export.

## Fix

`ChatTaskWirePayload` and `ChatInputChunk` are now declared in
`@trigger.dev/sdk/chat` (a public subpath) and re-exported type-only
from the internal shared module, so every internal import is unchanged
and the browser/server module split is untouched. Declaration emit for
an inferred agent type now produces a portable specifier:

```ts
export declare const chatAgent: Task<"chat-agent", import("@trigger.dev/sdk/chat").ChatTaskWirePayload<MyUIMessage, MyClientData>, unknown>;
```

As a side effect the wire types are now directly importable, which is
what affected users were reconstructing by hand.

## Verification

Reproduced against the built 4.5.2-equivalent package: a consumer
fixture with declaration emit produced `import("<file
path>/ai-shared.js")` in its declaration (the TS2742 trigger); after the
fix the same fixture emits the public specifier with zero diagnostics. A
regression test now builds that consumer simulation in a temp directory
on every test run: it copies the built package into a fake node_modules
(copied, not symlinked, because tsc only applies exports-map naming to
real node_modules paths), compiles the fixture with the TypeScript API,
and asserts no errors, no relative-path imports, and no internal module
references in the emit.
2026-07-10 07:35:03 +01:00
Matt Aitken 6b0588bef1 chore: vouch brentshulman-silkline (#4216)
Adds `brentshulman-silkline` to the list of vouched outside contributors
so their PRs aren't auto-closed by the vouch check.
2026-07-09 21:02:09 +01:00
James Ritchie afc8f9e210 fix(webapp): show magic link confirmation instead of reloading login (#4215)
## Summary

Submitting your email on the login page could reload back to an empty
login form instead of showing the "we've sent you a magic link"
confirmation. The magic link email was still sent, so it looked like
nothing happened.

## Root cause

The `/login/magic` route imported a server-only cookie module
(`magicLinkEmailCookie.server.ts`) whose top-level `env.NODE_ENV` read
got bundled into the route's client JS. On the client `env` is
undefined, so the module threw a `TypeError` at module eval, which
aborted Remix's client-side navigation to the confirmation and
hard-reloaded back to `/login`. It only surfaced in production builds
(local dev auto-logs-in, and local prod builds happen to tree-shake the
module out), which is why it slipped through.

## Fix

The email-link strategy already stores the submitted address in the
session (`auth:email`), so the separate cookie was redundant. Deleted
the cookie module and read the address from the session in the loader.
With the module gone, nothing server-only can leak into the client
bundle regardless of tree-shaking.

Verified the confirmation renders with the email address, the SSO
domain-policy redirect (with the email prefilled) still works, and a
production build no longer bundles the module.
2026-07-09 20:20:06 +01:00
Oskar Otwinowski dc6c98af5e chore(webapp): trim comments in directorySyncEffects (#4207)
Condense the kept rationale comments (logLevel/warn, last-Owner dedup,
role
overwrite) and drop the obvious function-header comments that just
restated
the code. No behavior change.
2026-07-09 20:13:21 +01:00
claude[bot] 105f48927d fix(release): populate changelog and server-changes on release/dispatch (#4204) 2026-07-09 17:46:39 +01:00
Chris Arderne 580f94a955 chore: ignore plugins package in changesets (#4210)
## Summary

Excludes the non-published plugins workspace from Changesets release
planning so it cannot drive public package version bumps.

## Verification

Ran `pnpm run changeset:version` with temporary changesets for
`@trigger.dev/plugins` and `@trigger.dev/core`; the ignored workspace
produced no release-driver updates, and the public package changeset
versioned normally.
2026-07-09 17:30:19 +01:00
Oskar Otwinowski e57fd9ce90 fix(webapp): downgrade retryable directory-sync effect failures to warn (#4200)
Directory-sync effects are idempotent and the accounts-webhook worker
retries
the whole event, so a single failed attempt (typically a role assignment
losing a serializable race during a backfill burst) is self-healing
rather
than alert-worthy. Tag those thrown errors with logLevel "warn" so the
worker
logs at warn instead of error, keeping them visible for triage without
paging.
2026-07-09 17:57:40 +02:00
Eric Allam 1a0198cc5e perf(webapp,clickhouse): move runs empty-state check to ClickHouse (#4202)
## Summary

The runs page's empty-state check (whether an environment has ever had a
run, which decides between the "getting started" and "no runs match your
filters" states) ran a `findFirst` against the Postgres `TaskRun` table.
This moves it to ClickHouse, the same store the runs list itself reads
from, so the check no longer queries `TaskRun`.

## Design

Only the runs list triggers the check now (via an `includeHasAnyRuns`
flag); the other presenters that reuse `NextRunListPresenter` (API,
schedule detail, waitpoint detail, error group) no longer issue it. When
the list is empty it runs `SELECT 1 FROM task_runs_v2 ... LIMIT 1`
filtered on the full `(organization_id, project_id, environment_id)`
sort-key prefix with a configurable `created_at` lower bound
(`RUN_LIST_HAS_RUNS_LOOKBACK_DAYS`, default 30), so it hits the primary
index and reads minimal granules.

Results are cached in a tiered memory + Redis SWR cache. Only positive
("has runs") results are cached, so an environment with no runs is
always re-checked and its first run shows up immediately.
2026-07-09 15:24:31 +01:00
nicktrn 0631c8373c chore: retire legacy v3 dev websocket + delete legacy self-hosting docs (#4198)
Follow-up to #4194 (v3 execution app + core-helper removal). The v3
(engine V1) is end-of-lifed and enforced off in prod, so this removes a
self-contained slice of the remaining dead v3 code while **keeping every
user-facing deprecation message** - a user still on v3 must still be
told to upgrade.

## Legacy dev websocket

`app/v3/handleWebsockets.server.ts` backs the `/ws` transport used
**only** by the legacy v3 `trigger dev` CLI (v4 dev uses a different
transport). It's now authenticate-then-close with
`V3_DEV_DEPRECATION_MESSAGE`, so an old CLI is still told what to do -
only the legacy `AuthenticatedSocketConnection` / `DevQueueConsumer`
execution behind it (which can no longer run) is removed.

- Deleted `app/v3/authenticatedSocketConnection.server.ts` (its only
consumer).
- `engineDeprecation.server.ts` and the deprecation message constants
are untouched.

## Docs

Deleted the intentionally-legacy "Docker (legacy)" self-hosting page
(`open-source-self-hosting.mdx`) and redirected
`/open-source-self-hosting` (+ the existing
`/v3/open-source-self-hosting` alias) to `/self-hosting/overview`;
repointed the two inbound links. The current `self-hosting/*` docs
already describe the v4 (single supervisor) setup.

## Deliberately out of scope

Despite the branch name, this PR does **not** touch MarQS or the
socket.io coordinator/provider namespaces. Investigation found MarQS is
entangled with **live v2** queue/metrics/concurrency/project-cleanup
code (`runQueue`, `queueSizeLimits`, `taskRunConcurrencyTracker`,
`EnvironmentQueuePresenter`, `registerProjectMetrics`, `deleteProject`),
so it needs a per-file reviewed pass, not a bulk delete. That remainder
stays on TRI-11883.

refs TRI-11883
2026-07-09 14:40:44 +01:00
Chris Arderne a3dca98d43 fix: run npm release jobs on ubuntu-latest (#4201)
🚀 Publish Trigger.dev Docker / typecheck (push) Failing after 4s
🚀 Publish Trigger.dev Docker / scan-webapp (push) Has been skipped
🚀 Publish Trigger.dev Docker / publish-webapp (push) Has been skipped
🚀 Publish Trigger.dev Docker / publish-worker-v4 (push) Has been skipped
🚀 Publish Trigger.dev Docker / units (push) Failing after 20s
🧭 Helm Chart Release / lint-and-test (push) Has been cancelled
🧭 Helm Chart Release / release (push) Has been cancelled
🚀 Publish Trigger.dev Docker / 📣 Dispatch main image (push) Has been cancelled
Failed trying to run trust npm publish on warp runner:

https://github.com/triggerdotdev/trigger.dev/actions/runs/29016820615/job/86113758922
helm-v4.5.2 v.docker.4.5.2 v4.5.2
2026-07-09 13:23:39 +00:00
github-actions[bot] 188f008715 chore: release v4.5.2 (#4180)
## Summary
4 improvements, 5 bug fixes.

## Improvements
- Add SDK and API client helpers for run bulk actions.
([#4105](https://github.com/triggerdotdev/trigger.dev/pull/4105))
- Large batch payloads now offload to object storage instead of riding
inline in the trigger request. `batchTrigger` and `batchTriggerAndWait`
(and the by-id and by-task variants) offload any per-item payload over
128KB before sending, the same way single `trigger` and `triggerAndWait`
already do, so a big batch no longer blows past the API body limit.
([#4165](https://github.com/triggerdotdev/trigger.dev/pull/4165))
- Removed internal helpers that were only used by the end-of-life v3
self-hosted compute providers.
([#4194](https://github.com/triggerdotdev/trigger.dev/pull/4194))
- Add an `onEvent` callback to `TriggerChatTransport` /
`useTriggerChatTransport` that emits typed lifecycle events for sends,
stream connects, first chunk, and turn completion. Send-success metrics,
time-to-first-token, and "sent but never answered" watchdogs become a
few lines of client code.
([#4187](https://github.com/triggerdotdev/trigger.dev/pull/4187))
  
  ```ts
  onEvent: (event) => {
if (event.type === "message-sent") metrics.timing("chat.send_ms",
event.durationMs);
if (event.type === "first-chunk") metrics.timing("chat.ttft_ms",
event.sinceSendMs ?? 0);
  },
  ```

## Bug fixes
- fix(cli): honor the MCP server's `--dev-only` flag
([#4199](https://github.com/triggerdotdev/trigger.dev/pull/4199))
- Fix chat turns that throw (for example from an `onTurnStart` hook)
leaking their message listener, which lost or duplicated messages sent
during later turns.
([#4176](https://github.com/triggerdotdev/trigger.dev/pull/4176))
- Fix `chat.agent` and `chat.createSession` permanently dropping user
messages when several arrived during a single turn: every buffered
message is now dispatched as its own turn instead of only the first.
([#4176](https://github.com/triggerdotdev/trigger.dev/pull/4176))
- Fix chat continuation runs replaying already-answered messages: turns
delivered while the run was suspended now advance the session.in resume
cursor, so a new run picks up exactly where the previous one left off.
([#4176](https://github.com/triggerdotdev/trigger.dev/pull/4176))
- Fix `chat.createSession` swallowing a message sent shortly after
stopping a turn: the turn's message listener now detaches when the
stream settles, so those messages run as the next turn.
([#4176](https://github.com/triggerdotdev/trigger.dev/pull/4176))

<details>
<summary>Raw changeset output</summary>

# Releases
## @trigger.dev/build@4.5.2

### Patch Changes

- Updated dependencies:
  - `@trigger.dev/core@4.5.2`
## trigger.dev@4.5.2

### Patch Changes

- fix(cli): honor the MCP server's `--dev-only` flag
([#4199](https://github.com/triggerdotdev/trigger.dev/pull/4199))
- Updated dependencies:
  - `@trigger.dev/core@4.5.2`
  - `@trigger.dev/build@4.5.2`
  - `@trigger.dev/schema-to-json@4.5.2`
## @trigger.dev/core@4.5.2

### Patch Changes

- Add SDK and API client helpers for run bulk actions.
([#4105](https://github.com/triggerdotdev/trigger.dev/pull/4105))
- Large batch payloads now offload to object storage instead of riding
inline in the trigger request. `batchTrigger` and `batchTriggerAndWait`
(and the by-id and by-task variants) offload any per-item payload over
128KB before sending, the same way single `trigger` and `triggerAndWait`
already do, so a big batch no longer blows past the API body limit.
([#4165](https://github.com/triggerdotdev/trigger.dev/pull/4165))
- Removed internal helpers that were only used by the end-of-life v3
self-hosted compute providers.
([#4194](https://github.com/triggerdotdev/trigger.dev/pull/4194))
## @trigger.dev/python@4.5.2

### Patch Changes

- Updated dependencies:
  - `@trigger.dev/core@4.5.2`
  - `@trigger.dev/sdk@4.5.2`
  - `@trigger.dev/build@4.5.2`
## @trigger.dev/react-hooks@4.5.2

### Patch Changes

- Updated dependencies:
  - `@trigger.dev/core@4.5.2`
## @trigger.dev/redis-worker@4.5.2

### Patch Changes

- Updated dependencies:
  - `@trigger.dev/core@4.5.2`
## @trigger.dev/rsc@4.5.2

### Patch Changes

- Updated dependencies:
  - `@trigger.dev/core@4.5.2`
## @trigger.dev/schema-to-json@4.5.2

### Patch Changes

- Updated dependencies:
  - `@trigger.dev/core@4.5.2`
## @trigger.dev/sdk@4.5.2

### Patch Changes

- Add SDK and API client helpers for run bulk actions.
([#4105](https://github.com/triggerdotdev/trigger.dev/pull/4105))
- Fix chat turns that throw (for example from an `onTurnStart` hook)
leaking their message listener, which lost or duplicated messages sent
during later turns.
([#4176](https://github.com/triggerdotdev/trigger.dev/pull/4176))
- Fix `chat.agent` and `chat.createSession` permanently dropping user
messages when several arrived during a single turn: every buffered
message is now dispatched as its own turn instead of only the first.
([#4176](https://github.com/triggerdotdev/trigger.dev/pull/4176))
- Fix chat continuation runs replaying already-answered messages: turns
delivered while the run was suspended now advance the session.in resume
cursor, so a new run picks up exactly where the previous one left off.
([#4176](https://github.com/triggerdotdev/trigger.dev/pull/4176))
- Fix `chat.createSession` swallowing a message sent shortly after
stopping a turn: the turn's message listener now detaches when the
stream settles, so those messages run as the next turn.
([#4176](https://github.com/triggerdotdev/trigger.dev/pull/4176))
- Add an `onEvent` callback to `TriggerChatTransport` /
`useTriggerChatTransport` that emits typed lifecycle events for sends,
stream connects, first chunk, and turn completion. Send-success metrics,
time-to-first-token, and "sent but never answered" watchdogs become a
few lines of client code.
([#4187](https://github.com/triggerdotdev/trigger.dev/pull/4187))

  ```ts
  onEvent: (event) => {
if (event.type === "message-sent") metrics.timing("chat.send_ms",
event.durationMs);
if (event.type === "first-chunk") metrics.timing("chat.ttft_ms",
event.sinceSendMs ?? 0);
  },
  ```

- Large batch payloads now offload to object storage instead of riding
inline in the trigger request. `batchTrigger` and `batchTriggerAndWait`
(and the by-id and by-task variants) offload any per-item payload over
128KB before sending, the same way single `trigger` and `triggerAndWait`
already do, so a big batch no longer blows past the API body limit.
([#4165](https://github.com/triggerdotdev/trigger.dev/pull/4165))
- Updated dependencies:
  - `@trigger.dev/core@4.5.2`
## @trigger.dev/plugins@4.5.2

### Patch Changes

- Updated dependencies:
  - `@trigger.dev/core@4.5.2`

</details>

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-07-09 12:02:34 +00:00
Chris Arderne 34b1a181c2 fix: security release 2026-07-06 (#4199) 2026-07-09 11:58:33 +00:00
James Ritchie bb450e608d feat(webapp): SSO & Directory Sync settings UI improvements (#4196)
📚 Publish docs / publish (push) Has been cancelled
## Summary

UI/layout/copy pass over the org **SSO & Directory Sync** settings page
(formerly "Identity & Access"). No logic, gates, flags, or data flow
changed — server-side auth (`manage:sso`), Enterprise entitlement,
action validation, and data loading are all untouched.

- Renamed the nav item, page title, and meta from "Identity & Access" to
"SSO & Directory Sync".
- Added a reusable `SettingsLayout` component system (container,
section, header, row, block, actions) modeled on `/account/security`,
and refactored the SSO page onto it (section titles, dividers, left
title/subtitle + right action rows).
- Tightened all UI copy: concise, active voice, consistent labels, no
em-dashes.
- `Select` primitive: additive `wrap`, `popoverClassName`, and
`placement` props (all default to prior behavior) so role options show a
bright title with a wrapping description, right-aligned popover, and no
horizontal overflow.
- Removed the external-link arrow icon from buttons that open a modal;
kept it only on genuinely external actions (Contact us, Open in new
tab).
- Polished the admin portal link dialog: smaller description, tighter
spacing, `ClipboardField` with a permanent copy button, removed the
redundant Copy link button, and a provider-aware Open label (e.g. "Open
in WorkOS") derived from the link host with a safe fallback.

### SSO page UI
<img width="3568" height="2550" alt="CleanShot 2026-07-08 at 18 52
11@2x"
src="https://github.com/user-attachments/assets/009d2437-7552-4ff0-a457-64744a9fcd88"
/>

### Login with SSO and normal email test (local)


https://github.com/user-attachments/assets/b33a4ce9-c1fa-45c9-bd3c-077cb6fc9473



## Test plan

- [ ] Non-Enterprise org: SSO page shows the upsell state
- [ ] Enterprise org, non-Owner without `manage:sso`: 403
- [ ] Enterprise Owner: verify domains, configure SSO, connect
directory, JIT/default/group role selects, and enforcement toggle all
work
- [ ] Role select popovers: bright title + wrapping description,
right-aligned, no horizontal scroll
- [ ] Admin portal dialog: copy button works, "Open in WorkOS" opens the
portal in a new tab

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
docs-release-2026-07-09
2026-07-09 12:09:45 +01:00
Matt Aitken 71e4b00880 docs: add ClickHouse chat agent example project page (#4195)
## What

Adds a new Example projects page: **ClickHouse chat agent** — a
`chat.agent()` that answers questions about your data by writing and
running SQL against ClickHouse Cloud via the official Node.js ClickHouse
client.

The page follows the existing example-project format (overview,
features, GitHub repo card, how-it-works with code excerpts, relevant
code links) and is registered in `docs.json` in alphabetical order.

## Note

The GitHub repo card links to
`triggerdotdev/examples/tree/main/clickhouse-chat-agent`, which lands in
a companion examples PR — merge that one first.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-09 11:31:21 +01:00
nicktrn a6bd370e42 chore: remove end-of-life v3 execution components (#4194)
v3 (engine V1) is end-of-lifed and the v3 clusters are gone, so this
removes the dead v3 execution code from the monorepo. It's the first
pass of TRI-11824 - the webapp v3 code paths are deliberately left
untouched and gated for a follow-up.

## Apps

Deletes the three v3-only execution apps and their build wiring:

- `apps/coordinator`, `apps/kubernetes-provider`, `apps/docker-provider`
- `.github/workflows/publish-worker.yml` - it built only those three;
the v4 worker publish is a separate workflow
- Their references in `.changeset/config.json`, `.cursorignore`,
`CHANGESETS.md`, `CONTRIBUTING.md`, `.server-changes/README.md`
- `pnpm-lock.yaml` regenerated to prune the apps and their app-only
dependencies (`socket.io`, `@kubernetes/client-node`, `p-queue`,
`execa`, `prom-client`, `tinyexec`)

## Core

Removes the helpers in `@trigger.dev/core` that only those apps used -
`ProviderShell`, `SimpleLogger`, the `Exec`/process helpers,
`isExecaChildProcess`, `getTextBody`, and `testDockerCheckpoint`. Each
was verified to have no remaining consumers anywhere in the repo.

Kept the helpers still used elsewhere: `ExponentialBackoff` (warm-start
client), `HttpReply`/`getJsonBody` (serverOnly http server),
`SimpleStructuredLogger` (widely used), and
`ZodNamespace`/`ZodSocketConnection` (still referenced by legacy v3
webapp code, hence the follow-up pass).

The `./v3/apps` and `./v3/serverOnly` export subpaths remain - only dead
members were trimmed from their barrels, so no `package.json` exports
changed.

## Verification

`@trigger.dev/core` builds, and `typecheck` passes for core, supervisor,
cli-v3, run-engine, redis-worker, and webapp.

refs TRI-11824
2026-07-08 19:47:20 +02:00
Iss e0208f3a27 fix(webapp): keep playground chat requests same-origin (#4193)
### Problem

The agent playground chat builds its realtime transport baseURL from
apiOrigin, but points it at a same-origin /resources/... dashboard
route. When API_ORIGIN differs from APP_ORIGIN, the in/append POST goes
cross-origin, fails the CORS preflight, and messages never reach the
agent ("Failed to fetch").

It only reproduces where the two origins differ — not locally, where
both default to localhost:3030.

Fixes #4149.

### Fix

Build the base URL from window.location.origin (falling back to
apiOrigin on SSR), so realtime traffic stays same-origin — the same
approach AgentView.tsx already uses.

### Testing

Typecheck passes. The CORS path only manifests when API_ORIGIN !=
APP_ORIGIN, so verify on test-cloud (can't reproduce locally).
2026-07-08 13:18:08 -04:00
Chris Arderne 7a19bb4cbb chore: add security docs (#4192)
Adds SECURITY.md and security page to the docs.
2026-07-08 17:23:18 +01:00
Daniel Sutton 80d4819a03 fix(webapp): stop slow database cleanup on project deletion (#4191)
## Summary

Deleting a project triggered an unbounded database cleanup that scanned
the project's entire run history, so deleting a project with many runs
could be very slow. Project deletion is a soft delete again: run data is
retained and the deletion completes quickly.

## Fix

Project deletion ran a cascade hard-delete whose `BulkActionItem` step
filtered through a relation to `TaskRun` scoped by `projectId`. Prisma
compiles that to an `EXISTS`-join over the project's entire `TaskRun`
set (a large, hot table with no `projectId` index), and it ran on every
project deletion unconditionally.

Removing the cascade-cleanup call restores the prior soft-delete
behaviour: queues are removed, the project is marked deleted, and run
data is retained. The cascade-cleanup service (added in
[#4117](https://github.com/triggerdotdev/trigger.dev/pull/4117)) had no
other callers, so it and its test are deleted.
2026-07-08 16:53:24 +01:00
Matt Aitken a682f1d171 feat(hosting): default self-hosted realtime streams to v2 (s2-lite) (#4185)
## Summary

Realtime streams (AI-agent token streaming and run streams) now default
to v2 for self-hosters, backed by a bundled [s2-lite](https://s2.dev)
service. Self-hosting previously shipped no S2 configuration, so streams
ran on the Redis-backed v1 path and there were no docs for wiring up v2.
Both the Docker Compose stack and the Helm chart now provision s2-lite
with persistent storage and set the stream env vars out of the box.

## What's included

- **Docker Compose**: a persistent `s2` service (s2-lite), a basin init
spec, and the `REALTIME_STREAMS_S2_*` plus
`REALTIME_STREAMS_DEFAULT_VERSION=v2` env on the webapp. `.env.example`
documents the v1 fallback and hosted-S2 options.
- **Helm**: an `s2` StatefulSet, PVC, Service and ConfigMap (runs as the
non-root image user via `fsGroup`), an `s2` values block, and webapp env
wiring with an existing-secret path for hosted S2.
- **Docs**: the `REALTIME_STREAMS_S2_*` and
`REALTIME_STREAMS_DEFAULT_VERSION` vars in the webapp env reference,
plus a "Realtime streams" section in the Docker and Kubernetes
self-hosting guides.

## Notes

- The OSS code default stays `v1`; v2 becomes the default purely through
the self-hosting artifacts, so non-self-host deployments are unaffected.
Disabling s2, or setting the version back to `v1`, cleanly reverts to
Redis-backed v1.
- With v2 enabled, the bundled s2 service is a required dependency for
streaming: if it is down, streams error while the task itself still
runs. That is the intended trade for the better v2 path.
- You can point at a hosted S2 at s2.dev instead of the bundled server.
2026-07-08 14:58:37 +01:00
Katia Bulatova 6e827f1da3 chore: Tailwind CSS v4 migration (#4139)
Migrates the webapp from Tailwind CSS 3.4 to 4.x.
2026-07-08 15:11:40 +02:00
Katia Bulatova e0bf74bfae docs: billing limits and alerts page (#4132)
New `/billing-limits` page covering the full [billing limits
feature](https://trigger.dev/changelog/billing-limits) : the three limit
options (plan / custom / no limit), billing alerts (% of limit or dollar
thresholds), what happens when the limit is reached, the recovery flow,
the soft-limits caveat, and the billing limit marker on the Usage page.
2026-07-08 13:46:42 +02:00
DKP 00ee0751ec feat(webapp): proxy PostHog through a same-origin /ph path (#4183)
## Summary

posthog-js sent product analytics to PostHog Cloud directly from the
browser. This points `api_host` at a same-origin `/ph` path that
forwards to PostHog Cloud EU server-side, following PostHog's standard
first-party reverse-proxy setup.

## How it works

A resource route forwards each request server-side, splitting by path:
`/ph/static/*` and `/ph/array/*` go to the asset host, everything else
(analytics events, feature flags) goes to the ingest host. It rewrites
the `Host` header, strips the `/ph` prefix, and streams the response
back. Only PostHog's own cookies are forwarded, so the app session
cookie stays first-party. Upstream hosts default to PostHog Cloud EU,
overridable via `POSTHOG_INGEST_HOST` / `POSTHOG_ASSETS_HOST`.

It also sets `cross_subdomain_cookie` so a single PostHog session is
shared across the marketing site and app.

Verified locally: static assets return 200 from the EU asset host, and
analytics events return 200 through the ingest host.
2026-07-08 11:35:04 +01:00
Eric Allam fbd86b6ee9 feat(sdk): onEvent observability callback on the chat transport (#4187)
## Summary

`sendMessage` from `useChat` gives no feedback about whether a message
actually reached the backend, and the `fetch` override is wire-level: it
requires knowing endpoint semantics, cannot attribute requests to
messages, and misses the headStart first-turn POST entirely. This adds a
typed `onEvent` observability callback to `TriggerChatTransport` /
`useTriggerChatTransport` so send-success metrics, time-to-first-token,
and "sent but never answered" watchdogs become a few lines of client
code.

## Example

```ts
const transport = useTriggerChatTransport({
  task: "my-chat",
  accessToken: ({ chatId }) => mintChatAccessToken(chatId),
  onEvent: (event) => {
    switch (event.type) {
      case "message-sent":
        // Durably acknowledged by the session's input stream, not just "request accepted".
        metrics.increment("chat.message_sent", { source: event.source });
        metrics.timing("chat.send_duration_ms", event.durationMs);
        break;
      case "message-send-failed":
        metrics.increment("chat.message_send_failed", { status: event.status });
        break;
      case "first-chunk":
        metrics.timing("chat.ttft_ms", event.sinceSendMs ?? 0);
        break;
      case "turn-completed":
        metrics.timing("chat.turn_duration_ms", event.sinceSendMs ?? 0);
        break;
    }
  },
});
```

## Design

One callback, one discriminated union (`ChatTransportEvent`):

- `message-sent` / `message-send-failed`: terminal send outcomes with
`messageId`, a `source` discriminator (submit, regenerate, steer,
action, stop, head-start), `durationMs`, `bodyBytes`, the append's
idempotency key (`partId`, also stored on the server-side record), and
error + HTTP status on failure. `message-sent` means the append was
durably acknowledged, after any internal token-refresh retries.
- `stream-connected` (with a `resumed` flag and the cursor it connected
from), `first-chunk` (chunk type plus `sinceSendMs` for
time-to-first-token), `turn-completed` (`sinceSendMs` full-turn latency
and the agent's committed input cursor), and `stream-error` follow the
response side, so a send can be paired with the answer that should
follow it. `messageId` on response events is client-side attribution
from the last turn-producing send on that chat.

Emissions sit at the transport's existing choke points, covering every
send path uniformly (including steering and headStart, which the fetch
override cannot observe). Exceptions thrown by the callback are
swallowed: observability can never break the chat. The React hook keeps
the callback live across renders instead of freezing the first-render
closure.

## Verification

Unit tests drive the transport directly with the `fetch` override as the
network stub (send success/failure per source, stream lifecycle, resumed
flag, field enrichment, callback exceptions swallowed). Verified
end-to-end against a realistic metrics setup in the ai-chat reference
app (counters, send-duration and TTFT histograms, and both watchdogs
built purely on these events): a healthy two-turn chat produces exactly
the expected event sequence and TTFT values; an oversized append records
`message_send_failed` with status 413; and killing the worker after a
durable send fires both `sent_but_no_stream` and `sent_but_unanswered`,
reproducing and detecting the "message disappeared" failure mode that
motivated this feature.
2026-07-08 11:01:40 +01:00
Eric Allam fe07de4a2c fix(webapp): use provider-reported cost for AI generations when present (#4186)
## Summary

The run page could show an AI generation cost well above what the
provider actually charged, most visibly for OpenRouter and Vercel AI
Gateway requests where a heavily cache-read prompt was priced at the
full input rate. When the provider reports an exact per-request cost, we
now use that instead of catalog pricing.

## Fix

Gateway and OpenRouter include the exact per-request cost in
`ai.response.providerMetadata` (`openrouter.usage.cost` /
`gateway.cost`). That figure already reflects the cache-read discount
and the real per-provider rate, which the catalog cannot reconstruct:
cache-read counts do not arrive in `gen_ai.usage.*`, and per-model
catalog prices drift from what the provider billed, in either direction.
So provider-reported cost is now preferred, and the catalog is used only
when no provider cost is present.

Fallback routing is covered by the same change: when OpenRouter routes
to a different model, `gen_ai.response.model` already carries the served
model, so the cost follows the served model and the provider's own
figure makes it exact.

`extractProviderCost` now runs on every AI span, so it gets a cheap
`"cost"` substring guard to skip the JSON parse on reasoning-model spans
whose provider metadata carries large reasoning text and no cost field.

Regression tests cover the cache-discount overcharge, fallback
served-model pricing, gateway cost, and the catalog fallback path.
2026-07-07 22:47:34 +01:00
James Ritchie 7c5f089d3d feat(webapp): rework login page and SSO sign-in UI (#4182) 2026-07-07 22:34:16 +01:00
Wes Mason 76c37ecd24 feat(sdk,core,webapp): offload large batch payloads to object storage (#4165)
## Summary

`batchTrigger` and `batchTriggerAndWait` (and the by-id and by-task
variants) now offload any per-item payload over 128KB to object storage
before sending, the same way single `trigger`/`triggerAndWait` already
do since
[#3785](https://github.com/triggerdotdev/trigger.dev/pull/3785). A batch
of large items no longer inflates the request body past the API limit.

## Demo

A live local run: `batchTriggerAndWait` of 5 items × 300KB (1.5MB
total). Each item offloads to object storage, so the receiver run rows
hold a 65-byte `application/store` pointer instead of the 300KB body,
and every item round-trips (received == sent).

<img width="1000" height="494" alt="batch large-payload offload demo"
src="https://github.com/user-attachments/assets/77ae3958-97d6-4b5c-ab25-39b217caefbc"
/>


## Design

Both the array and streaming batch paths funnel through
`executeBatchTwoPhase`, so offloading happens once there: each item is
measured, then offloaded through the existing
`conditionallyExportPacket` when it crosses 128KB, with bounded
concurrency so a big batch doesn't fire an unbounded number of presigned
PUTs.

Because items are offloaded before the request, SDK batches arrive as
small `application/store` references, so the server-side inline offload
during item ingest (parallelised in
[#3777](https://github.com/triggerdotdev/trigger.dev/pull/3777)) mostly
no longer fires for them.

Every trigger and item also carries its pre-offload serialised size as
`options.payloadSize`. The trigger span records that value, so an
offloaded payload shows its real size instead of the size of the small
object-store reference (previously the span measured the reference).
2026-07-07 16:21:48 +00:00
nicktrn 94b30fc1a6 fix(webapp): reject deploy images with runtime-incompatible zstd layers (#4184)
Container runtimes (cri-o / containerd / podman) can't pull
zstd-compressed layers carried in a Docker v2s2 manifest
(`application/vnd.docker.image.rootfs.diff.tar.zstd`). A deploy built
with an outdated CLI can produce exactly that combination - and today
it's promoted to current and then fails every run at image-pull time.

This extends the pre-promotion image check (#4049) to also inspect the
manifest's layer media types. If any layer uses the unpullable zstd/v2s2
media type, the deploy is rejected at finalize with a clear message to
upgrade the CLI and re-deploy, instead of silently shipping a version
that can't start.

The manifest is already returned by the existing ECR `BatchGetImage`
call, so there's no extra registry request for single-arch images.
Parsing is a lenient Zod schema and **fails open** - a manifest we can't
read never blocks a deploy. Manifest lists / OCI indexes (no top-level
`layers[]`) and OCI zstd (`...tar+zstd`, which runtimes support) pass
unaffected.

Also clarifies in the contributor docs that changesets and
`.server-changes/` notes are user-facing and should be written for
users, not maintainers.

refs TRI-11702
2026-07-07 17:20:07 +01:00
claude[bot] 8bf5879b60 test(webapp): poll for replicated rows instead of fixed sleeps in runs replication tests (#4181)
<!-- ccr-slack-attribution -->
_Requested by **Matt Aitken** · [Slack
thread](https://triggerdotdev.slack.com/archives/C032WA2S43F/p1783430373189849?thread_ts=1783430373.189849&cid=C032WA2S43F)_

##  Checklist

- [x] The PR title follows the convention.
- [x] I ran and tested the code works (typecheck of the edited files is
clean; see Testing)

---

## Testing

**Before:** the webapp run-replication test shard failed on nearly every
PR because assertions waited a fixed 1s for rows to replicate from
Postgres → ClickHouse and intermittently checked before the row arrived
under CI load.

**After:** those assertions poll (up to 30s, 250ms interval) until the
rows land, so they pass as soon as replication completes and stop
flaking, without slowing the happy path.

These tests are testcontainers-backed (need Docker + Postgres +
ClickHouse), so the full suite is exercised in CI. Locally I confirmed
the edited `runsReplicationService.part1..part8.test.ts` files
type-check with no new errors.

---

## Changelog

**How:** wrapped the ~21 present-row assertions across
`runsReplicationService.part1..part8.test.ts` in `vi.waitFor`, matching
the existing poll pattern in `part9.test.ts`. Left absence assertions
(expecting 0 rows / no spans) on a fixed settle delay since there is
nothing to poll for. Tests only — no production code changed.

Note: this does NOT touch the `subscribe()` startup race in
`internal-packages/replication/src/client.ts` (a riskier, separate
follow-up).

💯

---
_Generated by [Claude
Code](https://claude.ai/code/session_01KtUdSLKrK17eFVuRYXT6uj)_

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-07 16:15:15 +01:00
Chris Arderne aa74e68c71 feat(sdk): add bulk replay to api and sdk (#4105)
## Summary

Adds SDK and API support for run bulk actions. You can now create bulk
cancel or replay actions from `@trigger.dev/sdk` using run IDs or the
same filters as `runs.list()`, then retrieve, list, poll, or abort the
action by its `bulk_` handle.

Tests, docs, changesets added.

## Design

The dashboard bulk action service now accepts structured filters instead
of reading directly from a dashboard request, so the dashboard and API
share the same creation path. Replay actions created through the API are
attributed with the existing `api` trigger source, while
dashboard-created actions keep `dashboard`.

The SDK exposes the new surface under `runs.bulk.*`, including
`targetRegion` for replay region overrides and cursor pagination for
listing bulk actions.

## Filters and runIds

Nuance on filters. If `filter` is provided, it MUST have at least one
key. This is to remove the footgun of passing no filter and selecting
all runs.

```typescript
   { action: "cancel", runIds: ["run_1"] } // valid
   { action: "cancel", runIds: [] } // invalid, min(1)
   { action: "cancel", filter: { status: "FAILED" } } // valid
   { action: "cancel", filter: {} } // invalid
   { action: "cancel", filter: {}, runIds: ["run_1"] } // invalid
```
2026-07-07 15:43:30 +01:00
Daniel Sutton d59743bd35 fix(webapp,run-ops-database): keep run-ops batch items co-resident with their batch (#4178)
## Summary

Three fixes to the run-ops database split (the Cloud-only mode where
run-lifecycle rows live on a dedicated Postgres). All are inert in the
default single-database deployment.

The main fix: on the batch trigger paths, a parentless batch's item runs
chose their physical store from a fresh per-org mint-flag read at
processing time, so flipping an org's flag mid-batch could land an item
in a different store than its batch, breaking the `TaskRun.batchId`
foreign key (or silently orphaning the item). The other two harden the
split's safety nets: the schema-parity test now actually compares
columns, and the read fan-out gate now signals when it has been silently
disabled.

## Batch item residency

`RunEngineBatchTriggerService` (api.v2) and the BatchQueue item callback
(api.v3) now anchor each item's id mint on the batch's own friendlyId,
mirroring the already-safe `BatchTriggerV3Service`. Residency is a pure
id-shape check, so an item can no longer diverge from its batch across a
mid-batch flag flip. The pre-failed-run fallback is anchored the same
way (it also sets `batchId`), and the shared mint branch is consolidated
into one helper so every mint path stays in lockstep. No new database
queries; single-database mode is unchanged (a cuid-shaped batch
friendlyId yields a cuid item).

## Schema parity test

The parity test previously read only the dedicated schema and matched
model headers with regexes, so it never compared columns and could not
catch a run-subgraph column that diverged between the two physical
schemas. It now parses both schemas and asserts bidirectional
scalar-column parity (type, nullability, array-ness, default) across the
run-subgraph models, and fails on any field line it can't parse. Scoped
to the run-subgraph models so unrelated control-plane edits don't break
it.

## Read fan-out signal

The split read fan-out gate is decided by the object identity of the NEW
vs control-plane clients. It now warns when both run-ops URLs are set
but the NEW client isn't a distinct instance (fan-out silently off), and
a new test exercises the real topology-into-gate wiring so a future
refactor that aliases the clients can't disable fan-out unnoticed.

## Verification

New unit and glue tests cover all three changes; the DB-backed
residency, store-routing, and topology suites pass against real
Postgres; `typecheck` is clean for both packages.
2026-07-07 14:17:23 +01:00
Eric Allam add0a7da0a fix(sdk,core): stop chat sessions dropping messages that arrive during a turn (#4176)
## Summary

Sending a message to a chat whose run had ended could make the message
vanish: the continuation run replayed already-answered messages, never
processed the new one, and a page refresh lost it entirely. Chasing that
report surfaced four composing message-loss bugs in the chat session
runtime; this PR fixes all of them, each with a regression test.

## The fixes

1. **Stale resume cursor.** Records delivered while a run was suspended
(the waitpoint path) advanced the SSE resume counter but not the
committed-consume cursor, so the `session-in-event-id` header stamped on
turn-completes went stale by one record per suspended turn. Continuation
boots seed from that header, which is what made them replay
already-processed messages. `session.in.wait()` now advances both
cursors.

2. **Only the first buffered message dispatched.** Messages arriving
during a turn are consumed into a buffer whose end-of-turn pickup
dispatched only the first entry; the buffer was recreated each turn, so
the rest were discarded, and since consuming a record commits the cursor
the loss was permanent. A continuation boot's replay delivers several
records back-to-back, which put the user's new message at index 1 or
later. The buffer now outlives the turn and drains one message per turn
in both `chat.agent` and `chat.createSession` (whose equivalent buffer
was never read at all).

3. **Post-stop window in `chat.createSession`.** The turn's message
listener stayed attached through the stopped turn's post-stream work, so
a message sent shortly after stopping a turn was consumed into the dead
steering queue and lost. The listener now detaches when the stream
settles, matching the `chat.agent` loop.

4. **Handler leak on errored turns.** A turn that threw outside the
streaming section (for example from an `onTurnStart` hook) leaked its
message listener. Previously that silently lost mid-turn messages; with
the loop-level buffer it would have duplicated them instead. The
subscription handle is now detached by the turn's catch/finally, and
`chat.createSession` defensively detaches its prior turn's listener when
user code exits a turn without `complete()`/`done()`.

## Verification

Reproduced end-to-end with the ai-chat reference project before the fix
(message consumed but never answered, two replayed turns, gone on
refresh) and verified after (single clean turn, survives refresh,
turn-complete cursors strictly advancing). Regression tests in
`packages/trigger-sdk/test/pending-message-drain.test.ts` cover all
four, each verified red against the unfixed behavior. A smoke sweep of
the standard chat scenarios (basic send, multi-turn, suspend/resume,
mid-stream refresh, stop, steering, cancel + continue, and the
`createSession` variant) passes on the final branch state.
2026-07-07 14:10:46 +01:00
Matt Aitken 1a033b665b fix(webapp,core): retry run resume through transient database outages (#4161)
## Summary

When the platform database is briefly unreachable while a run is
resuming from a wait, the run no longer fails with
`TASK_EXECUTION_ABORTED`. The worker now retries the resume through the
outage instead of aborting on the first blip.

## Root cause

Resuming a run calls the engine's `continue` worker-action endpoint.
That route caught every error and returned a `422`, which the worker's
HTTP client treats as non-retryable. So a transient Prisma
infrastructure error (for example `P1001` "Can't reach database server")
was flattened into a permanent failure: the worker gave up, force-killed
the run process, and completed it with `TASK_EXECUTION_ABORTED`.

## Fix

- The `continue` route now lets infrastructure errors propagate to the
generic 500 handler (message scrubbed, and retryable by the worker's
HTTP client), the same treatment the trigger path already gives them via
`isInfrastructureError`. Genuine validation errors (snapshot mismatch,
invalid state) still return `422`, so a stale retry stays non-retryable.
Resuming is idempotent server-side (guarded by the snapshot id), so
retrying is safe.
- The worker's `continueRunExecution` calls (both the
runner-to-supervisor and supervisor-to-engine hops) retry with a longer,
jittered backoff so they can ride out an outage lasting tens of seconds,
and the jitter keeps a fleet of resuming runs from stampeding the
database the moment it recovers.

Builds on #3960, which scrubbed the leaked message on these routes but
left the status non-retryable.

No changeset: this is a server-side behaviour fix recorded via
`.server-changes`. The `@trigger.dev/core` edits are internal run-engine
worker plumbing, not a public API change.
2026-07-07 11:51:59 +01:00
Chris Arderne 018f445467 chore: switch runners to warpbuild (#4175) 2026-07-07 11:24:17 +01:00
Chris Arderne f4af13981a feat: auto-update dev and preview branch selector (#4171) 2026-07-07 10:45:06 +01:00
Wes Mason ce123c13e1 fix(llm-model-catalog): stop generated files showing as modified (#4170)
📚 Publish docs / publish (push) Has been cancelled
## Summary

Running `pnpm run db:migrate` locally left `defaultPrices.ts` and
`modelCatalog.ts` in `llm-model-catalog` showing as modified every time,
a ~10k-line diff that only ever changed formatting. This stops the
churn.

## Root cause

The root `db:migrate` script ends in `&& turbo run generate`, which runs
the `generate` script in every package that has one, including this one.
The generator writes its output with `JSON.stringify` (quoted keys, no
trailing commas), but the checked-in copies had been reformatted by
oxfmt (unquoted keys, trailing commas). So the generator output never
matched what was committed, even though the parsed data was identical.

## Fix

Add the two generated files to `.oxfmtrc.json`'s ignore list and commit
the raw generator output, matching how other codegen files in the repo
are already handled (e.g. the tsql grammar). Generation is
deterministic, so `generate` and `format` are both no-ops on a clean
tree now.

No changeset: internal package, dev tooling only, no runtime or public
API change.
docs-release-20260707-1007
2026-07-07 08:35:06 +01:00
Chris Arderne 5158ee8ec8 chore: use latest self-host compose images (#4140)
## Summary

Depends on #4136.

Docker Compose self-hosting now uses the maintained `latest` image tag
by default instead of the frozen prerelease tag. The version-locking
docs keep pointing production users at explicit versioned tags when they
want pinned upgrades.
2026-07-06 22:03:16 +01:00
github-actions[bot] c584236937 chore: release v4.5.1 (#4126)
🚀 Publish Trigger.dev Docker / units (push) Failing after 22s
🚀 Publish Trigger.dev Docker / typecheck (push) Failing after 22s
🚀 Publish Trigger.dev Docker / publish-worker-v4 (push) Has been skipped
🚀 Publish Trigger.dev Docker / scan-webapp (push) Has been skipped
🚀 Publish Trigger.dev Docker / publish-webapp (push) Has been skipped
🚀 Publish Trigger.dev Docker / publish-worker (push) Has been skipped
🧭 Helm Chart Release / lint-and-test (push) Has been cancelled
🧭 Helm Chart Release / release (push) Has been cancelled
🚀 Publish Trigger.dev Docker / 📣 Dispatch main image (push) Has been cancelled
## Summary
1 improvement.

## Improvements
- Extend the SSO plugin contract with WorkOS Directory Sync (SCIM)
support.
([#4148](https://github.com/triggerdotdev/trigger.dev/pull/4148))

<details>
<summary>Raw changeset output</summary>

# Releases
## @trigger.dev/build@4.5.1

### Patch Changes

- Updated dependencies:
  - `@trigger.dev/core@4.5.1`
## trigger.dev@4.5.1

### Patch Changes

- Updated dependencies:
  - `@trigger.dev/build@4.5.1`
  - `@trigger.dev/core@4.5.1`
  - `@trigger.dev/schema-to-json@4.5.1`
## @trigger.dev/python@4.5.1

### Patch Changes

- Updated dependencies:
  - `@trigger.dev/build@4.5.1`
  - `@trigger.dev/core@4.5.1`
  - `@trigger.dev/sdk@4.5.1`
## @trigger.dev/react-hooks@4.5.1

### Patch Changes

- Updated dependencies:
  - `@trigger.dev/core@4.5.1`
## @trigger.dev/redis-worker@4.5.1

### Patch Changes

- Updated dependencies:
  - `@trigger.dev/core@4.5.1`
## @trigger.dev/rsc@4.5.1

### Patch Changes

- Updated dependencies:
  - `@trigger.dev/core@4.5.1`
## @trigger.dev/schema-to-json@4.5.1

### Patch Changes

- Updated dependencies:
  - `@trigger.dev/core@4.5.1`
## @trigger.dev/sdk@4.5.1

### Patch Changes

- Updated dependencies:
  - `@trigger.dev/core@4.5.1`
## @trigger.dev/core@4.5.1


## @trigger.dev/plugins@4.5.1

### Patch Changes

- Extend the SSO plugin contract with WorkOS Directory Sync (SCIM)
support.
([#4148](https://github.com/triggerdotdev/trigger.dev/pull/4148))
- Updated dependencies:
  - `@trigger.dev/core@4.5.1`

</details>

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
helm-v4.5.1 v4.5.1 v.docker.4.5.1
2026-07-06 19:17:36 +01:00
Chris Arderne ef3eadab3c chore: back to github runners (#4168) 2026-07-06 16:22:31 +01:00
Daniel Sutton 4c2c25511b test(webapp): split triggerTask engine test into per-concern files (#4167)
The engine `triggerTask` suite was a single 2447-line file with 23
`containerTest` cases, each spinning its own Postgres + Redis. vitest
shards by whole file, so all 23 container setups landed on one shard and
dominated its wall-clock. The recorded entry in `test-timings.json`
badly under-counts the real cost (it does not capture the
per-`containerTest` container startup that dominates on CI), so the
duration-sharding sequencer treated the file as light and stacked it,
producing one ~21 minute shard.

Splitting does not reduce the number of container setups; it lets those
23 cases distribute across shards instead of stacking on one. The webapp
unit-test stage is gated by its slowest shard, so this cuts the stage's
wall-clock roughly in half.

## CI timing (before vs after)

Real CI wall-clock of the `Unit Tests: Webapp` shards (`--shard=i/10`).
"Before" is sampled from recent runs on other branches (unsplit file,
from `main`); "after" is this PR.

| Shard | Before (s) | After (s) |
|------:|-----------:|----------:|
| 1  | 250 | 359 |
| 2  | 444 | 411 |
| 3  | 497 | 659 |
| 4  | **1257** | 284 |
| 5  | 545 | 641 |
| 6  | 284 | 644 |
| 7  | 244 | 214 |
| 8  | 340 | 445 |
| 9  | 188 | 395 |
| 10 | 234 | 567 |
| **Slowest shard (gates the stage)** | **~1247s (≈21m)** | **659s
(≈11m)** |
| Sum of all shards | 4283 | 4619 |

Before: shard 4 is the long pole at 1237s / 1247s / 1257s across three
sampled runs (the `triggerTask` file plus whatever else the packer put
with it). After: the six pieces spread across shards, the slowest drops
to 659s. The small rise in summed time is the extra per-file container
startup, paid in parallel across shards, so the gating number still
falls by about 10 minutes.

## Change

Split into six per-concern files that share a `triggerTaskTestHelpers`
module (the `vi.mock` calls stay per-file, since vitest hoists them):

- `triggerTask.test.ts` (3): trigger + concurrencyKey coercion
- `triggerTask.idempotency.test.ts` (4): idempotency + queue resolution
- `triggerTask.debounce.test.ts` (4): retries + debounce validation
- `triggerTask.mollifier.test.ts` (4): mollifier call-site behaviour
- `triggerTask.metadataCache.test.ts` (4): DefaultQueueManager task
metadata cache
- `triggerTask.residency.test.ts` (4): child run residency inheritance

All 23 cases are preserved. The file's `test-timings.json` entry is
split across the new files so bin-packing stays balanced.

While rewriting these files, cleanup was moved to `onTestFinished(() =>
engine.quit())` so an `engine`/`Redis` leaked on a failing assertion no
longer persists on the worker-scoped Redis and cascades into later cases
(`hookTimeout` raised to 60s so the after-cleanup gets the full budget).
Prisma lookups switched from `findUnique` to `findFirst` to match the
repo convention.

Verified: all six files run green locally (23/23), oxlint and oxfmt
clean.
2026-07-06 13:03:27 +01:00
Daniel Sutton e4ae8cbcd4 fix(run-engine,run-store,webapp): stop split-mode waits hanging on resume (#4164)
## Summary

On the run-ops database split, a run that waits (`triggerAndWait`,
`batchTriggerAndWait`, `wait.forToken`) could hang forever after its
wait had already completed. The runner reads a resume from
`/snapshots/since` exactly once: if that read returned the resume
snapshot without its completed-waitpoints, the runner logged "executing
without completed waitpoints", advanced its cursor, and never re-read
it, so the awaiting run never continued.

## Root cause

The resume snapshot and its completed-waitpoint rows were written as two
separate commits. This regressed when the split replaced Prisma's atomic
nested `connect` with an FK-free insert (in
[#4163](https://github.com/triggerdotdev/trigger.dev/pull/4163)), and
`/snapshots/since` is served from a read replica. A fetch landing in the
sub-millisecond gap between the two commits, or a multi-reader replica
serving the snapshot from a different point in time than its join rows,
delivered an empty resume. Because the runner consumes each snapshot
once and treats an empty resume as terminal, a single stale read was
fatal and produced a permanent, nondeterministic hang.

## Fixes

- Commit a snapshot and its completed-waitpoint links in one
transaction, restoring the atomicity the split removed.
- Repair the completed-waitpoints from the owning primary when a
multi-reader replica serves the snapshot without its join rows. This
covers single-waitpoint resumes, which carry no
`completedWaitpointOrder` and so were missed by the count-based repair.
- Read the primary in the checkpoint `WAIT_FOR_BATCH` pre-check, so a
batch that already resumed is not re-suspended into a stall.
- Fall back to the primary when a waitpoint token misses both read
replicas, so a token completed immediately after it was minted no longer
returns a spurious 404.
- Route batch-item creation by `batchTaskRunId`, consistent with the
batch-completion count and the row's foreign key.
- Reject control-plane-only relation selects on the dedicated schema
with a clear error instead of an opaque Prisma failure, and stop
`createDateTimeWaitpoint` bypassing residency routing through a caller
transaction.

Verified against the deployed split topology: a resume snapshot and its
completed-waitpoints are now always delivered together, so the runner
can no longer drop a resume.
2026-07-06 10:55:02 +00:00
Oskar Otwinowski de65370fb9 feat(webapp): Directory Sync (SCIM) for Identity & Access (#4148)
Extend the SSO plugin contract for directory sync and apply membership
effects
from the accounts webhook worker: provision users in mapped groups (role
from
group mapping, else the org default role), deprovision on removal, and
keep a
sticky-removal tombstone so JIT never silently re-adds a removed user.
JIT and
Directory Sync coexist; roles default to Developer (the JIT default-role
picker
has no 'None'). Changing a group's role in the dashboard re-applies it
to that
group's current members immediately. The Directory Sync settings section
(group→role mapping, external-domain + manual-membership policy,
deferred Save)
appears once a domain is verified — independent of SSO — gated by the
hasSso
flag. The settings page polls the whole page while entitled with
override-aware
drafts so in-progress edits are never clobbered.
2026-07-06 09:31:14 +02:00