* Auto-yield run execution to help prevent duplicate task executions
* Add auto-yield config to endpoints
* Refactor run execution with buffer and limits
Introduced constants RUN_CHUNK_EXECUTION_BUFFER and MAX_RUN_CHUNK_EXECUTION_LIMIT. Adjusted PerformRunExecutionV2Service to use the new constants to fine-tune execution timings and buffers.
* Add endpoint probing functionality
Added new `RESPONSE_TIMEOUT_STATUS_CODES` in `consts.ts` to manage timeout responses. Additional functions `detectResponseIsTimeout(response: Response)` was added in `endpoint.server.ts` to detect if a response was a timeout based on the status codes from `RESPONSE_TIMEOUT_STATUS_CODES`.
Update actions to use new endpoint probing endpoint service. This allows for the early probing of endpoints to determine if they're up and running.
A new class `ProbeEndpointService` was created in `probeEndpoint.server.ts` which makes HTTP requests to a given endpoint and updates its properties based on the result.
Finally, `detectResponseIsTimeout(response)` is used in `performRunExecutionV2.server.ts` for marking the execution as succeeded when facing a timeout.
* Refactored probe method in EndpointApi class
The probe method of the EndpointApi class has been refactored to remove the error handling part and it now takes a timeout sent from the client directly. The corresponding changes were also made in the ProbeEndpointService and TriggerClient objects to reflect the alterations in the probe method.
The error handling related to the timeout has been removed and the responsibility of handling the timeout has been shifted to the client. Thus, the probe method has been greatly simplified. The `probeEndpoint.server.ts` file was also changed to accommodate the change in behavior of the probe result.
In the `triggerClient.ts` the timeout for probe is now read from the incoming request object. For backward compatibility, if no timeout is provided in the request, the default value of 15 minutes is used.
* Remove performRunExecution v1 enqueue function
* Better document limits and add docs on increasing function timeouts
* Upgrade webapp docker container to use 18.18.2
* force clients to yield when a run is executing in a gracefully shutting down worker
* Renamed task `key` to `cacheKey` and added more task documentation
* Index the `@trigger.dev/sdk` version on Endpoints
* Added EndpointIndex status. Default is PENDING, existing rows are SUCCESS
* Made it easier to create a migration SQL file
* Created a job-catalog file for misconfigured Jobs that should error when running the CLI
* EndpointIndex data and state are now optional
* The Environments page now shows the status of the last refresh
* Improved the UI about endpoints
* Added EndpointIndex error column
* WIP making the endpoint indexing more robust
* Indexing errors are now surfaced
* Improved the error show it shows the job id
* Use “performEndpointIndexing” when you create your first endpoint from the UI
* Use “performEndpointIndexing” for the recurring endpoint checker
* Staging is now auto-indexed every 10 mins too
* Use “performEndpointIndexing” for the webhook
* Moved the throttling to a util
* Removed instructional comments
* Created a reusable retry system with exponential backoff
* Use p-retry for retrying with backoff
* The CLI gets indexing results and displays errors
* Improved the indexing error messages and display in the console
* Use a pre so the Indexing error is correctly split over multiple lines
* Use a db transaction for webhook that triggers endpoint indexing
* Tidied up imports
* Support older versions of the server
* Improved the comment on the misconfigured job
* Changeset: When indexing user's jobs errors are now stored and displayed
* feat: add internal flag to job-run table
* add backfill SQL
* Update internal backfill to only update internal jobs
---------
Co-authored-by: Eric Allam <eallam@icloud.com>
* Improves the perform of run resuming
When runs resume, we try and make sure that tasks that have already been completed are cached and reused. Worst case scenario the client needs to hit the API server once for a non-cached task that is indeed completed on the server, but this can get pretty expensive when there are a larger number of tasks.
This commit does 2 different things to help:
- noop tasks are no longer “cached” using the cachedTasks strategy, instead their idempotency keys are shoved into a bloom filter and the client tests for their inclusion in the bloom filter before running them (since they don’t have any concept of output, this works)
- Additional cached tasks are lazy loaded when a task is run. This allows us to progressively fetch additional tasks to be cached on the client, which will cut down on cache misses by a decent amount
* Create warm-carrots-float.md
* Make io.yield backwards compat with older platform versions
* Better support old clients connecting to server versions that support lazy loading cached tasks
* Fixed type errors when settings headers with unknown value
* Better yield not support error message
* Rename _version to _serverVersion to be more clear
* feat: BYO Auth
Define client-side auth resolvers to be able to supply custom authentication credentials for integrations before a run is performed
- Added new defineAuthResolver
- Update all integrations to support the new auth resolvers
- Strip internal symbols from .d.ts in integrations and trigger-sdk
- Added BYO Auth docs
- Update Dynamic Schedule to support associated account IDs
- Create external accounts just-in-time
- Added Account ID field to test job when there are external auth integrations
- Show Account ID on run dashboard
- Added new Run error state called “Unresolved auth”
* Added changeset
* Remove @internal from TriggerIntegration public methods
* Add void to the result union
* DynamicTriggers now work with the new BYO auth system, and added a bunch of docs and docs changes
* Add additional key material for registering dynamic trigger task
* Add new define* instance methods to the overview
* Added stripInternal to SDK tsconfig
* Statuses can now be set from a run, and are stored in the database
* Added the key to the returned status
* Made the test job have an extra step and only pass in some of the options
* client.getRunStatuses() and the corresponding endpoint
* client.getRun() now includes status info
* Fixed circular dependency schema
* Translate null to undefined
* Added the react package to the nextjs-reference tsconfig
* Removed unused OpenAI integration from nextjs-reference project
* New hooks for getting the statuses
* Disabled most of the nextjs-reference jobs
* Updated the hooks UI
* Updated the endpoints to deal with null statuses values
* The hook is working, with an example
* Changeset: “You can create statuses in your Jobs that can then be read using React hooks”
* Changeset config is back to the old changelog style
* WIP on new React hooks guide
* Guide docs for the new hooks
* Added the status hooks to the React hooks guide
* Removed the links to the status hooks reference for now
* Re-ordered the hooks
* Fix for an error in the docs
* Set a default of a blank array for the GetRunSchema
* CLI create-integration command now accepts an Open AI api key
* Create integration docs separated into multiple pages
* Initial Airtable integration commit, with OpenAI generated code
* OAuth page coming soon
* Export DisplayProperty from the SDK
* TSConfig made to match GitHub’s with paths
* getRecords
* Removed duplicate Stripe job from the catalog
* Renamed Airtable apiKey option to token
* First Airtable job
* Export Collaborator and Attachment field types
* A typesafe example that uses runTask
* WIP on new integration tasks… not working yet
* Attempt with class
* Revert "Attempt with class"
This reverts commit 93a48330019f754c3216c5b49964fa4b0218bd3f.
* WIP changing how tasks work
* Mock of async local storage
* Moved client creation from constructor
* New approach with a clone method on TriggerIntegration
* Added runTask to Airtable which is used by integration tasks
* Added the Airtable icon and connection when using runTask
* base().table() is working
* runTask options moved to the 3rd param, made optional with optional name
* Added some generic arguments
* Added generic type to table
* Removed old comment
* We don’t need to repeat the icon
* getRecords and getRecord now returning the right data and types
* Creating records
* Update records
* Delete records
* The internal properties of integrations are now hidden by the TypeScript types
* Sprinkled a Prettify in there
* Improved the types
* Added Airtable to the integration catalog
* Early work on Airtable webhook registration
* More progress with webhooks
* connectionKey needs to be cloned for webhooks to work
* connectionKey needs to be cloned for webhooks to work
* It was unclear that the ActivateSourceService was using a graphileJob id
* ActivateSourceService optionally takes a jobId, if missing it generate a unique id
* When retrying trigger registration, don’t pass an id so it is generated
* Removed Airtable webhooks tasks from the job-catalog example
* Added TriggerSourceOption, removed TriggerSourceEvent
* WIP with new ExternalSource options
* ExternalSourceTrigger setup
* DynamicTrigger changed to options, will need some more work
* filter gets options passed to it
* SourceMetadata v2 renamed to SourceMetadataV2, kept original
* Started versioning the backend
* Moved param order on io.getEvent and io.cancelEvent
* The runTask stuff that allows unknown to work is back
* Indexing for v1 and v2, with version on “activateSource” schema
* Added todos, to deal with Airtable SDK calls inside the webhook handler
* “deliverHttpSourceRequest” queueName changed to the source id so they process in order
* ActivateSource changes to deal with old and new data formats
* Update existing TriggerSources to v2
* Fix for dynamic.ts typescript errors, need to revisit this later
* UpdateSourceService v1 and v2, with new v2 API endpoint
* Removed unused imports
* More progress on v1 and v2
* Airtable webhooks are now triggering a job
* Moved webhooks to a new file
* You can do API calls in the webhook handler now, Airtable webhook data is being processed
* Airtable events coming through
* Defined the Airtable table payload type
* TriggerSource metadata is being stored and used
* Removed some logs
* Added filtering and don’t allow any webhooks that use automated sources
* Resend switched to new integration
* Moved Resend test jobs to the catalog, and tested it worked
* WIP on Slack, there are compile errors
* Created a generic type that strips out indexes
* Slack updated to use new integration
* SendGrid migrated over
* Integration runTask is now allowing regular types
* Changed io.runTask types so it only allows Json-able types
* OpenAI models tasks working
* Added Airtable changes to runTask
* Removed the index signature crap from the Slack integration
* Don’t need to cast the callback result
* Updated Resend
* Re-ordered runTask params
* WIP on openai
* onAccountUpdated is Connect only
* Removed RunTaskResult
* Handle Resend errors, the official SDK doesn’t expose them properly at the moment
* Removed OmitIndexSignature
* OpenAI converted to new integration, with backwards compatible functions
* Put the openai catalog back to what it was originally
* Export a standard retry with backoff, to be used
* Use the standard exponential backoff in the integrations
* Retry options moved earlier so they can be overriden by a task
* GitHub tasks migrated
* Added sources, fixed one bundling issue
* Added GitHub jobs to catalog
* Remove duplicate options
* Deduplicate events
* Removed duplicate Job
* Switched Plain over
* Set the Plain icon
* Converted Stripe over
* Supabase adapted
* Typeform working
* Added dynamic-schedule to catalog
* Added background-fetch job catalog
* Created dynamic-triggers catalog file
* Fixed old general file with runTask param order
* Dynamic triggers working
* SendGrid updated to use the same tsconfig as other integrations
* Removed Airtable webhook, until we have batch support
* Added OAuth airtable auth example
* Created beta changeset tag
* Beta changesets for most packages
---------
Co-authored-by: Eric Allam <eallam@icloud.com>
This commit fixes various issues with run executions, including a pretty gnarly memory bloat issue when resuming a run that had a decent number of completed tasks (e.g. anything over a few). Other issues fixed:
- Run executions no longer are bound to a queue, which will allow more parallel runs in a single job (instead of 1).
- Serverless function timeouts (504) errors are now handled better, and no longer are retried using the graphile worker failure/retry mechanism (causing massive delays).
- Fixed the job_key design of the performRunExecutionV2 task, which will ensure resumed runs are executed
- Added a mechanism to measure the amount of execution time a given run has accrued, and added a maximum execution duration on the org to be able to limit total execution time for a single run
* Don’t show the Ready To Run Job prompt if you have an Integration that needs attention
* Don’t highlight the row red or green
* Improvements to the contrast of the app sections and dividers
* Removed the Jobs page title as it’s duped info
* using the new border variable
* Sticky last table cell
* Improved the sticky last table cell
* Make any last cell in a table sticky by adding isSticky to it
* Added a dropdown menu to the menu table cell
* table rows can be marked as disabled by adding disabled
* Using the jobTestPath function for the test path
* tidy up imports
* Removed un-used props
* Removed the green badge variant
* Added a new status badge to the Runs table
* Clicking the gradient clicks the row
* Delete Job triggers a modal popup
* Added a large danger button type
* Added some modal styling and started adding data
* Added more styling and data to the delete job modal
* A table can now be given a full width prop
* Large danger button added to Storybook
* Danger button disabled state looks disabled now
* New active badge component to display in the table and logic for showing the env data
* Style updates to the dialog component
* active and job status badges can now have a small size
* Added a new named icon
* The Job page shows the Job status in the PageInfoRow
* Small badge style update
* Runs table has a sticky right cell
* Created a JobStatusTable component
* Added some placeholder help panel content for disabling a Job
* WIP creating a Settings page
* Added a delete button that triggers the delete modal – just need data hooking up
* Implemented deleting jobs from the dashboard
---------
Co-authored-by: Eric Allam <eallam@icloud.com>
* WIP job run performance improvements
- Added a `perf` tool to better measure job run performance under heavy load
- Removed `runFinished` job (not really needed)
- startQueuedRuns now uses a jobKey with replace
- Fixed an issue with ZodWorker when using jobKey
* Publish improvement docker images
* fixed the improvement docker publishing
* Downgrade back to prisma 4.16.0 because 5.1.x broke docker builds
* Changes to how queued runs work
- Split the worker into two different workers, one dedicated to performRunExecution
- Schedule performRunExecution in a single place, with a queue and using a round robin manually controlled concurrency
- Remove startQueuedRuns
- All runs are queued before they are started
- Setting the worker maxPoolSize to the same as the worker concurrency
- Starting to be able to split the docker image
* Remove queue name from startRun graphile job
* Make the prisma connection pool stuff configurable through env vars
* Hardcode (for now) the max concurrent runs limit
* Rewrite performRunExecution to be more performant
PerformRunExecutionV2:
- Does not create and manage jobRunExecution records
- Does not reimplement retrying, uses graphile worker retrying instead
I’ve kept around PerformRunExecutionV1 so this works when deploying. Definitely needs LOTS of testing
* Fix issues with cached tasks
- Limit the size of the cached tasks sent when executing a run, using the knapsack problem dynamic programming approach
- Actually USE the cached tasks in IO by using the idempotencyKey instead of the task ID
- Remove output from all logs
- Added a stress test job catalog
* Forgot to commit the logger updates
* Never log connectionString
* Login to docker hub to get around rate limits
* Add additional logging to the graphile workers
* Fix the *_ENABLED env vars
* Allow adding and removing jobs to be done from the webapp
* Don’t set the job to failed if it’s being retried
* Deprecated queue options in the job and removed startPosition. Now using the job/env combo as the job queue name
* Dequeung jobs doesn’t check if the runner is initialized
* Fixed issues with retrying a run getting stuck on a cancelled task, and errors from parsing the results of dequeing a job
* Remove queued round robin thing that isn’t used anymore
* Added slack to job catalog
* Better forwards compat
* Added long delay
* Fixed lock file
* chore: update prisma from version 4.16.0 to 5.1.0
* fixed the ts-errors due to the prisma update
---------
Co-authored-by: Matt Aitken <matt@mattaitken.com>
* Setup project-wide prettier
* Remove old workspace file
* Remove old debugging directives
* New top-level .prettierignore
* Updated Prettier config settings
* Contrubuting guide: Fix for some bad code blocks
* Added more ignores
* Improved the format script command
* printWidth set to 100
* Formatted entire repo (pnpm run format)
* WIP supabase integration
* supabase oauth working
* Supabase database triggers
* Specify postgres:14
* Limit refreshOAuthToken jobs to 10 attempts
* Better displaying types and removing onChange for now
* WIP on the supabase db client
* Finishing the supabase-js integration
* Adding changeset
* Added supabase to the integration catalogs, and added an optional icon to Integrations
* Reworking how we handle types for the triggers (wip)
* Update fully over to the new way to define supabase triggers
* Go back to using the type for the event name
* Add back in the icon to the JobListPresenter since it was moved from the ProjectPresenter
* Remove unused import