From f1becff0b8e6785934d29de2e63418d58dcb93a2 Mon Sep 17 00:00:00 2001 From: Kazuhiro Sera Date: Tue, 28 Jul 2026 15:08:39 +0900 Subject: [PATCH] docs: update missing info --- docs/models/index.md | 4 ++++ docs/release.md | 2 +- docs/running_agents.md | 2 ++ mkdocs.yml | 1 + 4 files changed, 8 insertions(+), 1 deletion(-) diff --git a/docs/models/index.md b/docs/models/index.md index 2c274ae1..819e0258 100644 --- a/docs/models/index.md +++ b/docs/models/index.md @@ -232,6 +232,8 @@ If you use a custom OpenAI-compatible endpoint or proxy, websocket transport als - You can use [`Runner.run_streamed()`][agents.run.Runner.run_streamed] directly after enabling websocket transport. For multi-turn workflows where you want to reuse the same websocket connection across turns (and nested agent-as-tool calls), the [`responses_websocket_session()`][agents.responses_websocket_session] helper is recommended. See the [Running agents](../running_agents.md) guide and [`examples/basic/stream_ws.py`](https://github.com/openai/openai-agents-python/tree/main/examples/basic/stream_ws.py). - For long reasoning turns or networks with latency spikes, customize websocket keepalive behavior with `responses_websocket_options`. Increase `ping_timeout` to tolerate delayed pong frames, or set `ping_timeout=None` to disable heartbeat timeouts while keeping pings enabled. Prefer HTTP/SSE transport when reliability is more important than websocket latency. - By default the SDK disables the incoming message-size limit (`max_size=None`). For long-lived agent processes behind proxies or in memory-constrained containers, set `responses_websocket_options={"max_size": 8 * 1024 * 1024}` to bound per-message memory usage. +- The [Responses API WebSocket service](https://developers.openai.com/api/docs/guides/websocket-mode) processes one response at a time on each connection and limits each connection to 60 minutes. Open a new connection after that limit; use multiple connections when you need parallel runs. +- The service keeps only the most recent response in connection-local memory. A failed `4xx` or `5xx` turn evicts the referenced `previous_response_id`. After reconnecting, a stored response can still be continued when available, but `store=False` and ZDR flows have no persisted fallback. Start a new chain with `previous_response_id=None` and send the full input context, or rebuild that context from locally managed session state. ### Hosted multi-agent (experimental) @@ -496,6 +498,8 @@ english_agent = Agent( Retries are runtime-only and opt in. The SDK does not retry general model requests unless you set `ModelSettings(retry=...)` and your retry policy chooses to retry. +On the Responses websocket transport, `retry_policies.provider_suggested()` recognizes pre-response overload frames and code-less `server_error` frames as retry suggestions. This does not enable retries by itself: you still need `ModelRetrySettings`, and the normal replay-safety checks still apply. If any response event has already arrived, the SDK does not replay the request. + ```python from agents import Agent, ModelRetrySettings, ModelSettings, retry_policies diff --git a/docs/release.md b/docs/release.md index 472bf5db..468575e6 100644 --- a/docs/release.md +++ b/docs/release.md @@ -29,7 +29,7 @@ Highlights: - Added the public `agents.decorators` module and the shorter `@tool` alias alongside the existing function and guardrail decorators. Function tools now also support async callable objects. - SDK configuration now consistently accepts either typed settings objects or dictionaries across agents, runs, models, sessions, sandboxes, and voice pipelines, with validation for unknown settings. - Hardened error and diagnostic logging across models, tools, MCP, Realtime, sessions, sandboxes, and tracing to avoid exposing raw sensitive payloads while preserving useful debugging context. -- Improved AnyLLM, LiteLLM, and Chat Completions compatibility, preserved session history across model retries, and added retries for WebSocket overloads that occur before a response starts. +- Improved AnyLLM, LiteLLM, and Chat Completions compatibility, preserved session history across model retries, and added provider retry guidance for WebSocket overloads that occur before a response starts so opt-in Runner retry policies can act when replay is permitted. - Added [create-time-only S3 mounts for Vercel sandboxes](sandbox/clients.md#mounts-and-remote-storage) through `VercelCloudBucketMountStrategy`. Mounted sessions exclude bucket contents from workspace persistence and intentionally do not support dynamic mount changes or session resume. ### 0.18.0 diff --git a/docs/running_agents.md b/docs/running_agents.md index 866c50be..593d73e7 100644 --- a/docs/running_agents.md +++ b/docs/running_agents.md @@ -117,6 +117,8 @@ asyncio.run(main()) Finish consuming streamed results before the context exits. Exiting the context while a websocket request is still in flight may force-close the shared connection. +The service processes one response at a time on each websocket connection and limits a connection to 60 minutes. The helper reuses the connection but does not remove those constraints. After a reconnect, `store=False` and ZDR flows cannot recover an uncached `previous_response_id`; start a new chain with full input context or rebuild it from locally managed session state. See the [Responses WebSocket transport notes](models/index.md#responses-websocket-transport) for the full recovery behavior. + If long reasoning turns hit websocket keepalive timeouts, increase `ping_timeout` or set `ping_timeout=None` to disable heartbeat timeouts. Use HTTP/SSE transport for runs where reliability matters more than websocket latency. ### Run config diff --git a/mkdocs.yml b/mkdocs.yml index c38e7476..dd4aa2f3 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -122,6 +122,7 @@ plugins: - Memory: ref/memory.md - REPL: ref/repl.md - Tools: ref/tool.md + - Decorators: ref/decorators.md - Tool context: ref/tool_context.md - Results: ref/result.md - Streaming events: ref/stream_events.md