0b750d00dd
Watch is the agent noticing something later: you ask it to tell you when a condition holds, and it answers when it does — or when it can't any more. A watch is a **durable one-shot promise**. The condition is checked on a schedule by deterministic code (no LLM in the checks), the answer lands in the chat once, and then the watch is over. Ten kinds: three on a run, five on a queue, error recurrence, health recovery. ## Stack Stacked on **#4529** (UI), which is stacked on **#4418** (chat, reports, investigate). Merge those first. **#4516** (storybook gallery) sits on top of this branch. ## How to review [**GUIDEBOOK.md**](https://github.com/triggerdotdev/trigger.dev/blob/feat/dashboard-agent-flows-watch/internal-packages/dashboard-agent/GUIDEBOOK.md) on this branch is the behaviour reference — it states the conditions rather than the code, so you can predict what happens without running anything. "The ten watch kinds, and what makes each fire" and "Creating a watch" describe exactly this PR, and the tables there are the spec the code is written against. ## What's inside - **Ten watch kinds**, one deterministic check each (`dashboardAgentWatch*Checks.ts`), with the spec union in `dashboard-agent-contracts/src/watch.ts`. - **Scheduling** — each watch schedules its own next check; due watches of one `(environment, cadence)` group can be checked together in one batch pass, with a sweep as the backstop for expiry, redelivery and retention. - **Delivery** — the in-chat wake and card, an optional email alert (new `DASHBOARD_AGENT_WATCH` alert channel, so it shows on the project's Alerts page with one-click unsubscribe), and an optional investigation when the outcome needs attention. - **Submission ledger** — `watch_submissions`, keyed `(chat_id, client_request_id)`, so a retried card submission replays the recorded outcome instead of creating a second watch. - **Watch token** — a delayed-execution credential accepted only by the watch endpoints, re-checked against the user's live access on every tick. - **Unread work** — the panel polls for wakes that landed while it was closed, so a chat can go unread and light the launcher dot. ## Key decisions **A check result is a 4-way, and only two of them are verdicts.** `satisfied` / `terminal_unsatisfied` are answers; `pending` and `unavailable` are not. Any exception inside any check is caught in one place and becomes `unavailable` with an unverified observation — a check that failed is never evidence. **A completed window is an answer, and whether it is good or bad news is declared per kind, never inferred.** There is a table for that in the guidebook: `run_failed` completing its window is *good* news ("hasn't failed"), `backlog_drain` completing it is not. One rule overrides the table: a window that completed on an unverified observation is neutral and says only that the watch ended without a confirmed answer. **An unreadable source is never a negative answer** — and, because investigations only open on `attention`, it never starts one either. **Identity is `(chat, project, environment)` plus the condition,** enforced by a partial unique index over active rows (`watches_chat_active_identity_key`), not by the read-then-insert check. Cadence, window, note and `ticks` are deliberately not part of it. Two different chats may watch the same thing — a watch is a promise to a chat. **The server resolves the target's name, whatever the model calls it.** The model can't tell a task queue (`task/<id>`) from a custom queue, so both spellings are tried and the stored one wins — and the rewrite happens **before** identity and before the row is written, so the identity, the checks, the link and the wording all see one spelling. **Freshness fences.** Depth falls back from the live counter to the newest 60 s ClickHouse bucket, which only counts as current within 60 s of now. A non-current reading at or below the *quiet line* is refused as `unavailable` rather than believed, so a stale empty bucket is never read as "drained". The stall streak is the one piece of carried state: it lives in the previous check's facts and *freezes* on an unreadable reading rather than breaking. **Chain reliability.** There is no shared cron — each watch (or batch group) schedules its own next tick, so the failure mode to review is the chain dying. A failed batch check is caught, the next tick is scheduled anyway and the run resolves rather than failing, so the chain survives a check that couldn't run; the sweep re-arms groups and finalizes anything still active past its deadline, even when delivery isn't configured. Wake redelivery is id-deduped rather than conditional, because the sweep can't know whether the user was already told. Access is re-authorized on **every** check against the primary — replica lag would extend access the user has already lost. **Wording lives in one place.** `watch-wording.ts` is read by the card, banner, toast, email and the agent's own narration, and the numbers come from the frozen observation rather than a fresh read, so a retry produces the same sentence. Replay reproduces the **recorded** decision instead of deciding again — the transcript is append-once, so a second decision would contradict it forever. **Cancellation is the ending without an answer** — no resolution, no wake. One exception, decided during testing: a watch the *user* cancelled leaves a single neutral transcript line ("Stopped watching …"), keyed off the watch id so a retry can't repeat it. The other four reasons stay silent. **Email is opt-in and only a fired watch emails.** An expiry is narrated in the chat and nowhere else. Both gates (agent access, a configured email transport) are checked at subscribe time *and* again at delivery, and the subscription outcome is frozen on the ledger row so a retry replays it. Neither gate is a plan check. **One watch offer per turn.** The prompt and the renderer guard this independently — if the turn already proposed a watch card, the action button is dropped, because the card is the better affordance. Two eval cases pin the prompt side: exactly one offer with the line last and the button after it, and zero offers when the rendered card already carries one — deterministic assertions, over a real-model run. ## Testing Unit tests (vitest, testcontainers, no mocks) under `apps/webapp/test/dashboardAgentWatch*.test.ts` and `internal-packages/dashboard-agent/src/watch-*.test.ts` cover the invariants above: the 4-way check results and the freshness fences, identity/dedup and the submission ledger, queue-name resolution, the batch chain surviving a failed check, sweep boundaries and alert-once, tenancy and the watch token's scope, and the wording snapshot. The load-bearing ones were verified by control-breaking the guard first and checking the test goes red. Live-tested end to end against a local stack, following the guidebook: all ten watch kinds firing and expiring, cancellation, the email pair (a fired watch mails, an expired one does not), and watch recovery from a health report.
135 lines
4.4 KiB
TypeScript
135 lines
4.4 KiB
TypeScript
import { generateJWT } from "@trigger.dev/core/v3/jwt";
|
|
import { beforeEach, describe, expect, it, vi } from "vitest";
|
|
|
|
/**
|
|
* The agent reads a queue's live row — paused, depth, limit — through the environment JWT it
|
|
* exchanges its delegated token for. Metrics already answer that JWT; without the same on the
|
|
* retrieve route the agent got a 401, which reaches the model as absent data and had it
|
|
* telling users a queue of thousands of runs did not exist.
|
|
*
|
|
* These drive the real loader with a real signed environment JWT: the route builder
|
|
* authenticates it, compiles its scopes into an ability, and gates on `read:queues`.
|
|
*/
|
|
|
|
const ENVIRONMENT_ID = "env_1234";
|
|
const API_KEY = "tr_dev_abcdefghijklmnop";
|
|
|
|
const environment = {
|
|
id: ENVIRONMENT_ID,
|
|
type: "DEVELOPMENT",
|
|
slug: "dev",
|
|
branchName: null,
|
|
apiKey: API_KEY,
|
|
organizationId: "org_1",
|
|
projectId: "proj_1",
|
|
archivedAt: null,
|
|
concurrencyLimitBurstFactor: 1,
|
|
maximumConcurrencyLimit: 10,
|
|
project: { id: "proj_1", externalRef: "proj_ref", deletedAt: null },
|
|
organization: { id: "org_1" },
|
|
orgMember: null,
|
|
parentEnvironment: null,
|
|
};
|
|
|
|
const queueRow = {
|
|
id: "tq_1",
|
|
friendlyId: "queue_1234",
|
|
name: "task/my-task",
|
|
type: "VIRTUAL",
|
|
runtimeEnvironmentId: ENVIRONMENT_ID,
|
|
paused: true,
|
|
concurrencyLimit: 5,
|
|
concurrencyLimitBase: 5,
|
|
concurrencyLimitOverriddenAt: null,
|
|
concurrencyLimitOverriddenBy: null,
|
|
concurrencyLimitOverridePercent: null,
|
|
};
|
|
|
|
const mocks = vi.hoisted(() => ({
|
|
runtimeEnvironmentFindFirst: vi.fn(),
|
|
taskQueueFindFirst: vi.fn(),
|
|
revokedApiKeyFindMany: vi.fn(),
|
|
}));
|
|
|
|
vi.mock("~/db.server", () => {
|
|
const client = {
|
|
runtimeEnvironment: { findFirst: mocks.runtimeEnvironmentFindFirst },
|
|
taskQueue: { findFirst: mocks.taskQueueFindFirst },
|
|
revokedApiKey: { findMany: mocks.revokedApiKeyFindMany, findFirst: async () => null },
|
|
};
|
|
return { prisma: client, $replica: client };
|
|
});
|
|
vi.mock("~/env.server", () => ({ env: { SESSION_SECRET: "test-session-secret" } }));
|
|
vi.mock("~/v3/engineVersion.server", () => ({ determineEngineVersion: async () => "V2" }));
|
|
vi.mock("~/v3/runEngine.server", () => ({
|
|
engine: {
|
|
lengthOfQueues: async () => ({ "task/my-task": 1234 }),
|
|
currentConcurrencyOfQueues: async () => ({ "task/my-task": 2 }),
|
|
},
|
|
}));
|
|
vi.mock("~/services/logger.server", () => ({
|
|
logger: { debug: vi.fn(), error: vi.fn(), warn: vi.fn(), info: vi.fn() },
|
|
}));
|
|
vi.mock("~/v3/services/worker/workerGroupTokenService.server", () => ({
|
|
WorkerGroupTokenService: class {},
|
|
}));
|
|
|
|
import { loader } from "~/routes/api.v1.queues.$queueParam";
|
|
|
|
/** The claims the env-JWT exchange mints (api.v1.projects.$projectRef.$env.jwt.ts). */
|
|
function mintEnvJwt(scopes: string[]) {
|
|
return generateJWT({
|
|
secretKey: API_KEY,
|
|
payload: {
|
|
sub: ENVIRONMENT_ID,
|
|
pub: true,
|
|
scopes,
|
|
act: { sub: "usr_1", client: "dashboard-agent" },
|
|
},
|
|
expirationTime: "1h",
|
|
});
|
|
}
|
|
|
|
async function retrieveQueue(token: string) {
|
|
const response = await loader({
|
|
request: new Request("https://api.trigger.dev/api/v1/queues/my-task?type=task", {
|
|
headers: { Authorization: `Bearer ${token}` },
|
|
}),
|
|
params: { queueParam: "my-task" },
|
|
context: {},
|
|
} as any);
|
|
return { status: response.status, body: await response.json() };
|
|
}
|
|
|
|
describe("queue retrieve through an environment JWT", () => {
|
|
beforeEach(() => {
|
|
// Only the JWT's own `sub` lookup resolves — a bearer read as an API key finds nothing.
|
|
mocks.runtimeEnvironmentFindFirst
|
|
.mockReset()
|
|
.mockImplementation(async ({ where }: any) =>
|
|
where?.id === ENVIRONMENT_ID ? environment : null
|
|
);
|
|
mocks.taskQueueFindFirst.mockReset().mockResolvedValue(queueRow);
|
|
mocks.revokedApiKeyFindMany.mockReset().mockResolvedValue([]);
|
|
});
|
|
|
|
it("answers a JWT carrying read:queues with the queue's live row", async () => {
|
|
const result = await retrieveQueue(await mintEnvJwt(["read:runs", "read:queues"]));
|
|
|
|
expect(result.status).toBe(200);
|
|
expect(result.body).toMatchObject({
|
|
id: "queue_1234",
|
|
name: "my-task",
|
|
paused: true,
|
|
queued: 1234,
|
|
});
|
|
});
|
|
|
|
it("refuses a JWT without it — widening who may ask must not widen what they may read", async () => {
|
|
const result = await retrieveQueue(await mintEnvJwt(["read:runs", "read:query"]));
|
|
|
|
expect(result.status).toBe(403);
|
|
expect(mocks.taskQueueFindFirst).not.toHaveBeenCalled();
|
|
});
|
|
});
|