Files
triggerdotdev--trigger.dev/apps/webapp/test/dashboardAgentCreateChatOrdering.test.ts
T
Katia Bulatova 0b750d00dd feat(webapp): dashboard agent — Watch (#4525)
Watch is the agent noticing something later: you ask it to tell you when
a condition holds, and it answers when it does — or when it can't any
more.

A watch is a **durable one-shot promise**. The condition is checked on a
schedule by deterministic code (no LLM in the checks), the answer lands
in the chat once, and then the watch is over. Ten kinds: three on a run,
five on a queue, error recurrence, health recovery.

## Stack

Stacked on **#4529** (UI), which is stacked on **#4418** (chat, reports,
investigate). Merge those first. **#4516** (storybook gallery) sits on
top of this branch.

## How to review


[**GUIDEBOOK.md**](https://github.com/triggerdotdev/trigger.dev/blob/feat/dashboard-agent-flows-watch/internal-packages/dashboard-agent/GUIDEBOOK.md)
on this branch is the behaviour reference — it states the conditions
rather than the code, so you can predict what happens without running
anything. "The ten watch kinds, and what makes each fire" and "Creating
a watch" describe exactly this PR, and the tables there are the spec the
code is written against.

## What's inside

- **Ten watch kinds**, one deterministic check each
(`dashboardAgentWatch*Checks.ts`), with the spec union in
`dashboard-agent-contracts/src/watch.ts`.
- **Scheduling** — each watch schedules its own next check; due watches
of one `(environment, cadence)` group can be checked together in one
batch pass, with a sweep as the backstop for expiry, redelivery and
retention.
- **Delivery** — the in-chat wake and card, an optional email alert (new
`DASHBOARD_AGENT_WATCH` alert channel, so it shows on the project's
Alerts page with one-click unsubscribe), and an optional investigation
when the outcome needs attention.
- **Submission ledger** — `watch_submissions`, keyed `(chat_id,
client_request_id)`, so a retried card submission replays the recorded
outcome instead of creating a second watch.
- **Watch token** — a delayed-execution credential accepted only by the
watch endpoints, re-checked against the user's live access on every
tick.
- **Unread work** — the panel polls for wakes that landed while it was
closed, so a chat can go unread and light the launcher dot.

## Key decisions

**A check result is a 4-way, and only two of them are verdicts.**
`satisfied` / `terminal_unsatisfied` are answers; `pending` and
`unavailable` are not. Any exception inside any check is caught in one
place and becomes `unavailable` with an unverified observation — a check
that failed is never evidence.

**A completed window is an answer, and whether it is good or bad news is
declared per kind, never inferred.** There is a table for that in the
guidebook: `run_failed` completing its window is *good* news ("hasn't
failed"), `backlog_drain` completing it is not. One rule overrides the
table: a window that completed on an unverified observation is neutral
and says only that the watch ended without a confirmed answer. **An
unreadable source is never a negative answer** — and, because
investigations only open on `attention`, it never starts one either.

**Identity is `(chat, project, environment)` plus the condition,**
enforced by a partial unique index over active rows
(`watches_chat_active_identity_key`), not by the read-then-insert check.
Cadence, window, note and `ticks` are deliberately not part of it. Two
different chats may watch the same thing — a watch is a promise to a
chat.

**The server resolves the target's name, whatever the model calls it.**
The model can't tell a task queue (`task/<id>`) from a custom queue, so
both spellings are tried and the stored one wins — and the rewrite
happens **before** identity and before the row is written, so the
identity, the checks, the link and the wording all see one spelling.

**Freshness fences.** Depth falls back from the live counter to the
newest 60 s ClickHouse bucket, which only counts as current within 60 s
of now. A non-current reading at or below the *quiet line* is refused as
`unavailable` rather than believed, so a stale empty bucket is never
read as "drained". The stall streak is the one piece of carried state:
it lives in the previous check's facts and *freezes* on an unreadable
reading rather than breaking.

**Chain reliability.** There is no shared cron — each watch (or batch
group) schedules its own next tick, so the failure mode to review is the
chain dying. A failed batch check is caught, the next tick is scheduled
anyway and the run resolves rather than failing, so the chain survives a
check that couldn't run; the sweep re-arms groups and finalizes anything
still active past its deadline, even when delivery isn't configured.
Wake redelivery is id-deduped rather than conditional, because the sweep
can't know whether the user was already told. Access is re-authorized on
**every** check against the primary — replica lag would extend access
the user has already lost.

**Wording lives in one place.** `watch-wording.ts` is read by the card,
banner, toast, email and the agent's own narration, and the numbers come
from the frozen observation rather than a fresh read, so a retry
produces the same sentence. Replay reproduces the **recorded** decision
instead of deciding again — the transcript is append-once, so a second
decision would contradict it forever.

**Cancellation is the ending without an answer** — no resolution, no
wake. One exception, decided during testing: a watch the *user*
cancelled leaves a single neutral transcript line ("Stopped watching
…"), keyed off the watch id so a retry can't repeat it. The other four
reasons stay silent.

**Email is opt-in and only a fired watch emails.** An expiry is narrated
in the chat and nowhere else. Both gates (agent access, a configured
email transport) are checked at subscribe time *and* again at delivery,
and the subscription outcome is frozen on the ledger row so a retry
replays it. Neither gate is a plan check.

**One watch offer per turn.** The prompt and the renderer guard this
independently — if the turn already proposed a watch card, the action
button is dropped, because the card is the better affordance. Two eval
cases pin the prompt side: exactly one offer with the line last and the
button after it, and zero offers when the rendered card already carries
one — deterministic assertions, over a real-model run.

## Testing

Unit tests (vitest, testcontainers, no mocks) under
`apps/webapp/test/dashboardAgentWatch*.test.ts` and
`internal-packages/dashboard-agent/src/watch-*.test.ts` cover the
invariants above: the 4-way check results and the freshness fences,
identity/dedup and the submission ledger, queue-name resolution, the
batch chain surviving a failed check, sweep boundaries and alert-once,
tenancy and the watch token's scope, and the wording snapshot. The
load-bearing ones were verified by control-breaking the guard first and
checking the test goes red.

Live-tested end to end against a local stack, following the guidebook:
all ten watch kinds firing and expiring, cancellation, the email pair (a
fired watch mails, an expired one does not), and watch recovery from a
health report.
2026-08-12 09:51:40 +02:00

210 lines
8.3 KiB
TypeScript

import { beforeEach, describe, expect, it, vi } from "vitest";
const mocks = vi.hoisted(() => ({
createChat: vi.fn(),
findEnvironmentBySlug: vi.fn(),
mintUserActorToken: vi.fn(),
mintPublicToken: vi.fn(),
headStart: vi.fn(),
startSession: vi.fn(),
softDeleteChat: vi.fn(),
logger: { debug: vi.fn(), error: vi.fn(), warn: vi.fn(), info: vi.fn() },
// Mutable so a test can take the head start away and drive the cold path.
env: { SESSION_SECRET: "test-session-secret", ANTHROPIC_API_KEY: "sk-test" } as Record<
string,
string | undefined
>,
}));
vi.mock("~/db.server", () => ({ $replica: {}, prisma: {} }));
vi.mock("~/env.server", () => ({ env: mocks.env }));
vi.mock("~/services/session.server", () => ({
requireUser: async () => ({ id: "usr_real", admin: false, isImpersonating: false }),
}));
vi.mock("~/v3/canAccessDashboardAgent.server", () => ({
canAccessDashboardAgent: async () => true,
}));
vi.mock("~/models/project.server", () => ({
findProjectBySlug: async () => ({
id: "proj_real",
organizationId: "org_real",
externalRef: "proj_ref_real",
}),
}));
vi.mock("~/models/runtimeEnvironment.server", () => ({
findEnvironmentBySlug: mocks.findEnvironmentBySlug,
}));
vi.mock("~/services/dashboardAgent.server", () => ({
dashboardAgentApiOrigin: () => "https://api.trigger.dev",
isDashboardAgentConfigured: () => true,
mintDashboardAgentToken: mocks.mintPublicToken,
mintDashboardAgentUserActorToken: mocks.mintUserActorToken,
resolveDashboardAgentRepoSnapshot: async () => null,
startDashboardAgentSession: mocks.startSession,
dashboardAgentWakeFeedCounter: { inc: vi.fn() },
}));
vi.mock("~/services/dashboardAgentHeadStart.server", () => ({
startDashboardAgentHeadStart: mocks.headStart,
}));
// The chat route reaches the ClickHouse factory through the watch services, and the factory
// builds its client at import time from an env var no test sets.
vi.mock("~/services/clickhouse/clickhouseFactoryInstance.server", () => ({
clickhouseFactory: { getClickhouseForOrganization: async () => ({}) },
}));
vi.mock("~/services/dashboardAgentDb.server", () => ({ dashboardAgentDb: {} }));
vi.mock("~/services/resolveTriggerUri.server", () => ({ resolveTriggerUri: () => null }));
// Spread the real module so this doesn't have to track every query the route imports.
vi.mock("@internal/dashboard-agent-db", async (importOriginal) => ({
...((await importOriginal()) as Record<string, unknown>),
createChat: mocks.createChat,
softDeleteChat: mocks.softDeleteChat,
}));
vi.mock("~/services/logger.server", () => ({ logger: mocks.logger }));
import { action } from "~/routes/resources.orgs.$organizationSlug.projects.$projectParam.env.$envParam.dashboard-agent";
function createChatRequest() {
const form = new URLSearchParams({
intent: "create",
message: JSON.stringify({ id: "m1", role: "user", parts: [{ type: "text", text: "hi" }] }),
});
return action({
request: new Request(
"https://app.trigger.dev/resources/orgs/acme/projects/api/env/dev/dashboard-agent",
{
method: "POST",
headers: { "content-type": "application/x-www-form-urlencoded" },
body: form.toString(),
}
),
params: { organizationSlug: "acme", projectParam: "api", envParam: "dev" },
context: {},
} as any);
}
describe("dashboard agent chat creation — nothing fallible after the row exists", () => {
beforeEach(() => {
mocks.createChat.mockReset().mockResolvedValue(undefined);
mocks.headStart.mockReset().mockResolvedValue(undefined);
mocks.findEnvironmentBySlug
.mockReset()
.mockResolvedValue({ id: "env_real", type: "DEVELOPMENT" });
mocks.mintUserActorToken.mockReset().mockResolvedValue("tr_uat_real");
mocks.mintPublicToken.mockReset().mockResolvedValue("pat_public");
mocks.startSession.mockReset().mockResolvedValue(undefined);
mocks.softDeleteChat.mockReset().mockResolvedValue({ deleted: true, cancelledWatches: [] });
mocks.env.ANTHROPIC_API_KEY = "sk-test";
});
it("creates no chat when the environment slug resolves to nothing", async () => {
mocks.findEnvironmentBySlug.mockResolvedValue(null);
const response = await createChatRequest();
expect(response.status).toBe(404);
expect(mocks.createChat).not.toHaveBeenCalled();
});
it("creates no chat when the delegated token mint fails", async () => {
mocks.mintUserActorToken.mockRejectedValue(new Error("signing key unavailable"));
const response = await createChatRequest();
expect(response.status).toBe(500);
expect(mocks.createChat).not.toHaveBeenCalled();
});
it("still creates the chat and head-starts it on the happy path", async () => {
const response = await createChatRequest();
expect(response.status).toBe(200);
expect(await response.json()).toMatchObject({ headStarted: true });
expect(mocks.createChat).toHaveBeenCalledTimes(1);
expect(mocks.headStart).toHaveBeenCalledTimes(1);
expect(mocks.headStart.mock.calls[0][0].metadata).toMatchObject({
userActorToken: "tr_uat_real",
environmentId: "env_real",
environmentName: "dev",
});
expect(mocks.softDeleteChat).not.toHaveBeenCalled();
});
});
// A failed start means no handover was dispatched and no message was sent, so any session it
// did create idles out having done nothing — the chat row is safe to take back. Once the start
// has resolved the session is live, and removing the chat would hide a running agent.
describe("dashboard agent chat creation — a start that fails part way", () => {
beforeEach(() => {
mocks.createChat.mockReset().mockResolvedValue(undefined);
mocks.headStart.mockReset().mockResolvedValue(undefined);
mocks.findEnvironmentBySlug
.mockReset()
.mockResolvedValue({ id: "env_real", type: "DEVELOPMENT" });
mocks.mintUserActorToken.mockReset().mockResolvedValue("tr_uat_real");
mocks.mintPublicToken.mockReset().mockResolvedValue("pat_public");
mocks.startSession.mockReset().mockResolvedValue(undefined);
mocks.softDeleteChat.mockReset().mockResolvedValue({ deleted: true, cancelledWatches: [] });
mocks.env.ANTHROPIC_API_KEY = "sk-test";
mocks.logger.error.mockReset();
});
it("takes the chat back when the head start fails", async () => {
mocks.headStart.mockRejectedValue(new Error("session create failed"));
const response = await createChatRequest();
expect(response.status).toBe(500);
expect(mocks.createChat).toHaveBeenCalledTimes(1);
expect(mocks.softDeleteChat).toHaveBeenCalledTimes(1);
expect(mocks.softDeleteChat.mock.calls[0][1]).toMatchObject({
chatId: mocks.createChat.mock.calls[0][1].id,
userId: "usr_real",
});
});
it("takes the chat back when the cold start fails", async () => {
mocks.env.ANTHROPIC_API_KEY = undefined;
mocks.startSession.mockRejectedValue(new Error("session create failed"));
const response = await createChatRequest();
expect(response.status).toBe(500);
expect(mocks.createChat).toHaveBeenCalledTimes(1);
expect(mocks.softDeleteChat).toHaveBeenCalledTimes(1);
});
it("keeps the chat when the session is live and only its access token failed", async () => {
mocks.mintPublicToken.mockRejectedValue(new Error("token mint failed"));
const response = await createChatRequest();
expect(response.status).toBe(500);
expect(mocks.headStart).toHaveBeenCalledTimes(1);
expect(mocks.createChat).toHaveBeenCalledTimes(1);
expect(mocks.softDeleteChat).not.toHaveBeenCalled();
});
it("keeps the chat when a cold-started session's access token failed", async () => {
mocks.env.ANTHROPIC_API_KEY = undefined;
mocks.mintPublicToken.mockRejectedValue(new Error("token mint failed"));
const response = await createChatRequest();
expect(response.status).toBe(500);
expect(mocks.startSession).toHaveBeenCalledTimes(1);
expect(mocks.softDeleteChat).not.toHaveBeenCalled();
});
it("surfaces the start's own failure when taking the chat back also fails", async () => {
mocks.headStart.mockRejectedValue(new Error("session create failed"));
mocks.softDeleteChat.mockRejectedValue(new Error("chat store unavailable"));
const response = await createChatRequest();
expect(response.status).toBe(500);
const logged = mocks.logger.error.mock.calls.map((call: any[]) => call[1]?.error?.message);
expect(logged).toContain("session create failed");
});
});