Compare commits
312 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 668fabd6fe | |||
| 4fde807f1f | |||
| b4606c8f4c | |||
| 909c5c434a | |||
| 4a7dc08d31 | |||
| 94c1f6095f | |||
| 7a828f6db2 | |||
| 0ffb9e062c | |||
| 5d4a752c24 | |||
| 3d07ff1687 | |||
| 1643dbe617 | |||
| e7f835b6c6 | |||
| 849241498f | |||
| ec8fbb875e | |||
| 70d01a9eca | |||
| a898061dfa | |||
| 106ecb6bfe | |||
| fe7d10625d | |||
| 2afb6e1a24 | |||
| 343c06fae1 | |||
| 56921c89fc | |||
| f13a3f78f1 | |||
| 8d78731826 | |||
| ecfa2e0db6 | |||
| 829e19a7b7 | |||
| 655c805c92 | |||
| 7713342bed | |||
| fea17ef542 | |||
| 97bc99101a | |||
| 81fe6ab348 | |||
| 8c688b96ac | |||
| ab38ee5d45 | |||
| 3321967587 | |||
| 92137d222b | |||
| 8a0e34caaa | |||
| 30ff4a5239 | |||
| 16d16ec2be | |||
| f706d65124 | |||
| 2df8adcae7 | |||
| 7192bc968a | |||
| a9b7b7ef4a | |||
| 28d2a6c838 | |||
| 3f6324c2d2 | |||
| f115dd6015 | |||
| e605bd20a2 | |||
| 9168c387eb | |||
| 593331a0a6 | |||
| 0bd3d1b865 | |||
| 7328f790be | |||
| ce4cc62d86 | |||
| 50c8b74c50 | |||
| 4ebf880953 | |||
| b20cb7f897 | |||
| 9f93077a74 | |||
| 83e8a1b6e2 | |||
| 13ef230468 | |||
| 869e869c70 | |||
| 306d426930 | |||
| f918c6a4f0 | |||
| 510b7de2ce | |||
| 09f0ed45e7 | |||
| 7fc489eb0d | |||
| a4ba63f704 | |||
| 288455911c | |||
| 12e268c8e2 | |||
| c359c20647 | |||
| 131c3aef76 | |||
| 1b0076dda0 | |||
| d8068d96a9 | |||
| 5a3028f37c | |||
| b855c76a56 | |||
| a4d17da8ae | |||
| 6927622331 | |||
| 8eaa4b9872 | |||
| 5121f302a5 | |||
| f4e5d96f9f | |||
| b051a70393 | |||
| 0e9807b6c4 | |||
| bf0d025dc9 | |||
| 9470af01ed | |||
| 9520a5c558 | |||
| 464043a5e6 | |||
| bd8b1b7541 | |||
| d87fec8589 | |||
| 9411cee8e2 | |||
| 921135630b | |||
| 6d5b888f6f | |||
| 9a92ea4c8e | |||
| fedc809afd | |||
| 573530febb | |||
| 813a00324e | |||
| eda7b51aea | |||
| 45effb4b7d | |||
| 9303389d84 | |||
| 9514473c23 | |||
| 445ad2273c | |||
| 17f7caeaa3 | |||
| f6ba91b120 | |||
| 65774ce88d | |||
| 13d1815218 | |||
| 41fe113dd0 | |||
| 4970fd6ce4 | |||
| 4bd459e403 | |||
| e37b3d266b | |||
| 94523f946e | |||
| cb0461d177 | |||
| 5e71d09554 | |||
| cb7211b599 | |||
| 43a389d462 | |||
| 656baf8ead | |||
| eca794c0bc | |||
| 7865ec9819 | |||
| f7e8196484 | |||
| bdd228b4db | |||
| 900cab6212 | |||
| 8dc30e4807 | |||
| f9eb3a4f33 | |||
| ec016cde9a | |||
| 564584513f | |||
| a466860fdc | |||
| 03ff10ef6c | |||
| 55b859d05c | |||
| 76f42d8ed7 | |||
| 55e4a850fc | |||
| 5594fb464d | |||
| 64de1cb211 | |||
| 1d3df7fa04 | |||
| ae3263b840 | |||
| e3c99d9964 | |||
| b2bd8218c6 | |||
| 8715a0585a | |||
| 4a95659892 | |||
| 5635fab9d3 | |||
| 92e014a4cc | |||
| 852a8dbad4 | |||
| eada610734 | |||
| 028abc6a74 | |||
| 4c3de2df65 | |||
| cf151f91ff | |||
| 970db97acb | |||
| 4bc942af3c | |||
| 52656a51cb | |||
| 479640e214 | |||
| 650212ef62 | |||
| d7417fb2e7 | |||
| a6a4cc5143 | |||
| 94ba76de0f | |||
| 62560996c2 | |||
| 94fa9e21ba | |||
| 610c2742dc | |||
| 4b8744903a | |||
| fa75ffc404 | |||
| bc3607bae4 | |||
| 9154d836a2 | |||
| 1289fb0bef | |||
| 73bc963398 | |||
| 43e051e725 | |||
| 38bef807f5 | |||
| fc1abe0999 | |||
| dbb3181386 | |||
| 3031b13eaa | |||
| 8f40dde4a4 | |||
| 4bb388c955 | |||
| 29b2acffb1 | |||
| f3d434cdfd | |||
| 33b9a2c9fd | |||
| 272dd18b9f | |||
| cc2998845b | |||
| 035271db60 | |||
| 55eb3dec30 | |||
| f8ba94d416 | |||
| 7aeb39e150 | |||
| 7e9089f270 | |||
| 960e979f74 | |||
| e6a3ed887a | |||
| e86dff2907 | |||
| 12ad112fec | |||
| 250fb97fef | |||
| 62f9416ead | |||
| 76702572ef | |||
| 1b6876a8c2 | |||
| 89f368df03 | |||
| 730ee55e14 | |||
| b4741e077d | |||
| 4a3d33e44f | |||
| 9242902e1b | |||
| b4be17586f | |||
| ec5523d620 | |||
| 1d38492b44 | |||
| f32f613a7f | |||
| de60d05b0c | |||
| cc5a392583 | |||
| 8b04ee0939 | |||
| e19ac4e44d | |||
| c7bcdd4a4a | |||
| 58a89c810f | |||
| 9a0c07f9c7 | |||
| 8619dfda75 | |||
| 683b6e79e5 | |||
| e3746c52d9 | |||
| 0fea7e8347 | |||
| 0a76dd03ce | |||
| f47d486985 | |||
| 1660d306b5 | |||
| ee36d43584 | |||
| a91f630f79 | |||
| bd84b65258 | |||
| 28de3652d3 | |||
| a67d95f58a | |||
| 3a11cf5225 | |||
| 170ee73f94 | |||
| 6f5fbf6dbb | |||
| 516aa0c9d9 | |||
| 0fb2e0944c | |||
| eed9100777 | |||
| f185dfa2a0 | |||
| c5ebf809ea | |||
| 1683357483 | |||
| 0ed4ee6e18 | |||
| 9f361ba7dd | |||
| e8856de1b4 | |||
| 5a10e46f11 | |||
| 8c8a2eb32e | |||
| ff8e3db4b2 | |||
| 8526723b49 | |||
| 0466636b77 | |||
| f903926394 | |||
| 1a1b35d4fb | |||
| fc2d208f30 | |||
| b9cbab149f | |||
| bed924b45d | |||
| e1cb2be4ca | |||
| b1722a7459 | |||
| 0370fd3527 | |||
| 9a2b4a9b85 | |||
| e22f25a431 | |||
| 6e691ee2ea | |||
| ce462354fd | |||
| 02a6b21151 | |||
| 75da8e0200 | |||
| 3d1231e0fe | |||
| 601ecf5503 | |||
| 574a598fae | |||
| 7a5d32bc83 | |||
| 613b8f39a6 | |||
| 9b57f057b4 | |||
| ae60947451 | |||
| 1b7d878b7c | |||
| b80d541946 | |||
| 54ec5f0091 | |||
| 3854c124cb | |||
| 63ebf5ada1 | |||
| fbc5a44045 | |||
| 4bb4400731 | |||
| 4b5a0b89cd | |||
| 9d24382d0e | |||
| 044d44ce0d | |||
| f2fb9ffb66 | |||
| 60b7bee807 | |||
| e9a3e3610c | |||
| ceb238fd1b | |||
| 41c646d898 | |||
| 4b2881c7c4 | |||
| a47b7ea7ec | |||
| 42c3015518 | |||
| ae224b4449 | |||
| 756fa431a7 | |||
| 841f72f296 | |||
| 611d080ff0 | |||
| 48d7e9cae7 | |||
| f7410c8e96 | |||
| da3f15708e | |||
| 498390a741 | |||
| 7833715ac2 | |||
| d996707c0a | |||
| ec99da6375 | |||
| 2d40c09c88 | |||
| 3a3f34f18d | |||
| 7029ea8fff | |||
| 572c7bf1b5 | |||
| 0661c9e9ca | |||
| 05004336a2 | |||
| 8d7f05b6b2 | |||
| ebdb0f2ee1 | |||
| b3688db750 | |||
| 8025ed0b42 | |||
| ba889de480 | |||
| 5df41d3013 | |||
| 2eb8713b53 | |||
| 9a207b6938 | |||
| 9af6ad111c | |||
| 1821bf8094 | |||
| 071e2b68f6 | |||
| c88f339d32 | |||
| 5ffc1ecee4 | |||
| c2cb031461 | |||
| fe3a5e6c27 | |||
| 16e040929b | |||
| 5be06a16a9 | |||
| 81c57c5e7e | |||
| 638388ad17 | |||
| 734d42490a | |||
| 4e43cbaf09 | |||
| fdf2d009a6 | |||
| 333721d72f | |||
| 4c68780ad3 | |||
| 3aad7eba85 | |||
| 4c5112cbf4 | |||
| 106ef05317 | |||
| 2a515f0eb4 | |||
| 9e228fc959 | |||
| bf3e9d178c |
@@ -0,0 +1,40 @@
|
||||
# SDK Maintainer References
|
||||
|
||||
This directory captures long-lived implementation contracts of the OpenAI Agents Python SDK that are not replaceable by OpenAI API or platform facts from the Developer Docs MCP. The repo's `docs/` remain an SDK-specific behavioral contract; these references distill the ownership, compatibility, ordering, and failure semantics that maintainers need to preserve that contract.
|
||||
|
||||
## Usage
|
||||
|
||||
Read the reference map before changing or reviewing an affected runtime boundary, then open only the files relevant to that boundary. During issue and PR review, treat this directory as read-only background: use it to identify expected invariants, adjacent surfaces, and regression risks, but verify the current claim against the remote issue or PR, current code, tests, docs, release boundary, and focused runtime evidence. Do not edit references as a side effect of a review or treat them as proof of current issue status, PR behavior, or repository readiness.
|
||||
|
||||
When implementation or dedicated repository-maintenance work establishes a reusable invariant that remains valid beyond one issue or PR, update the narrowest owning reference separately. Preserve the generalized contract, not the case history or decision outcome that revealed it.
|
||||
|
||||
## Inclusion Criteria
|
||||
|
||||
Add or retain a reference when the knowledge is SDK-specific, stable across multiple releases, easy to violate from one local code path, and expensive to reconstruct from source, tests, and repo docs during every review. Treat `docs/` as the SDK's user-facing behavioral contract; use these references to preserve the implementation constraints behind that contract. Prefer invariants and ownership rules over summaries of individual issues, PRs, or recent fixes.
|
||||
|
||||
Do not store current issue or PR status, generic maintainer-review workflow, release notes, OpenAI API or platform behavior available through `$openai-knowledge`, or one-off implementation details in this directory. Put review methodology under `.agents/skills/`, released migration notes in `docs/release.md`, and API or platform facts behind `$openai-knowledge`.
|
||||
|
||||
## Reference Map
|
||||
|
||||
| Reference | Read before changing or reviewing |
|
||||
|---|---|
|
||||
| [Agent definition and run context](agent-definition-and-run-context.md) | Agent fields, cloning, dynamic instructions, enabled tools or handoffs, context wrappers, usage, or public agent identity |
|
||||
| [Runner lifecycle](runner-lifecycle.md) | Turn accounting, guardrails, handoffs, interruptions, cancellation, or streaming parity |
|
||||
| [Run item lifecycle](run-item-lifecycle.md) | Model output processing, new item types, stream events, replay conversion, session persistence, or RunState serialization |
|
||||
| [Function and output schema](function-and-output-schema.md) | Function-tool signatures and metadata, strict JSON schema conversion, or structured output types |
|
||||
| [Conversation state ownership](conversation-state-ownership.md) | Sessions versus server-managed continuation, input deltas, retries, compaction, or conversation resume |
|
||||
| [Session persistence](session-persistence.md) | Session input callbacks, per-turn saves, retry rewind, atomicity, or compaction replacement |
|
||||
| [RunState schema and resume boundary](runstate-schema.md) | Serialized state, schema versions, approvals, agent identity, or durable resume data |
|
||||
| [Tool identity and routing](tool-identity.md) | Tool names, namespaces, lookup, approvals, MCP naming, handoffs, or call IDs |
|
||||
| [Tool execution lifecycle](tool-execution-lifecycle.md) | Function-tool planning, approvals, guardrails, concurrency, cancellation, timeouts, or failure conversion |
|
||||
| [Local MCP server lifecycle](local-mcp-server-lifecycle.md) | Local MCP connection ownership, manager state, request serialization, caching, filtering, retries, or cleanup |
|
||||
| [Model and provider boundaries](model-provider-boundaries.md) | Model resolution, provider adapters, feature capability, request conversion, terminal events, or retries |
|
||||
| [Tracing lifecycle](tracing-lifecycle.md) | Trace and span context, processors, export, flush, shutdown, resume, or sensitive data |
|
||||
| [Realtime session lifecycle](realtime-session-lifecycle.md) | Realtime listeners, connections, background tasks, handoffs, event iteration, or cleanup |
|
||||
| [Realtime tracing architecture](realtime-tracing.md) | Realtime API server traces versus Agents SDK client traces |
|
||||
| [Voice pipeline lifecycle](voice-pipeline-lifecycle.md) | VoicePipeline STT/workflow/TTS ownership, event and audio ordering, stream cleanup, PCM framing, or tracing |
|
||||
| [Sandbox runtime boundary](sandbox-runtime-boundary.md) | Sandbox session ownership, preparation, resume state, manifests, materialization, or cleanup |
|
||||
|
||||
## Maintenance Rules
|
||||
|
||||
Keep each rule in the narrowest reference that owns it. Cross-link instead of copying detailed rules between files. Describe current architecture and compatibility boundaries, not the chronology of how a bug was found. Use source paths and durable public contracts as anchors, and remove or rewrite guidance when ownership moves.
|
||||
@@ -0,0 +1,57 @@
|
||||
# Agent Definition and Run Context
|
||||
|
||||
Use this reference for changes to public `Agent` fields, cloning, dynamic instructions, enabled tools or handoffs, output schemas, `RunContextWrapper`, `ToolContext`, usage aggregation, or the distinction between a public agent and an internal prepared clone.
|
||||
|
||||
## Public Definition and Cloning
|
||||
|
||||
- Exported `Agent` and `AgentBase` dataclass field order is a positional compatibility boundary. Append optional fields where possible and test old positional construction when the order changes.
|
||||
- `Agent.__post_init__()` is the eager boundary for invalid field categories such as names, tools, handoffs, hooks, model settings, output types, and tool-use behavior. Dynamic callbacks are validated when invoked because their result depends on the current run.
|
||||
- `Agent.clone()` uses `dataclasses.replace()` and is shallow. Mutable fields and contained tool, handoff, hook, and provider objects remain shared unless the caller explicitly supplies replacements.
|
||||
- When `clone(model=...)` replaces a model whose settings still equal the old model's implicit defaults, recompute the new model's implicit defaults. Preserve explicitly customized `model_settings` instead of silently resetting them.
|
||||
|
||||
## Per-Turn Resolution
|
||||
|
||||
- Resolve dynamic instructions with the current context and public agent for each model turn. Enforce the documented two-argument callable shape and await async results.
|
||||
- Evaluate callable `FunctionTool.is_enabled` and `Handoff.is_enabled` against the current run context. Do not cache a prior run's enabled set on the reusable agent.
|
||||
- Use one resolved tool and handoff view for model exposure, reserved-name and collision checks, local dispatch, tracing, and Realtime session updates. Re-resolving independently at those surfaces can expose one set and execute another.
|
||||
- An internal prepared agent may add bound tools, instructions, or sampling settings, but hooks, `ToolContext.agent`, handoff callbacks, and public results should identify the public agent unless an internal identity is explicitly part of the contract.
|
||||
- The effective output schema belongs to the agent and model call that produced the candidate output. A handoff can change the final output type, so do not assume the starting agent's schema when parsing or typing the final result.
|
||||
|
||||
## Context Ownership
|
||||
|
||||
- Every agent, tool, handoff, guardrail, and lifecycle hook in one run must agree on the same application context type. The context object is local runtime state and is never added to model input automatically.
|
||||
- A normal `ToolContext.from_agent_context()` shares the underlying application object, usage accumulator, and approval mapping with its parent while adding call-scoped fields such as call ID, namespace, arguments, and conversation history.
|
||||
- Nested `Agent.as_tool()` execution has a separate run loop, approval scope, and resumable tool state. On the normal function-tool path it still shares the application object and usage accumulator, while `tool_input` belongs to the nested wrapper and must not overwrite the parent's scoped value.
|
||||
- Sharing the application object is not the same as sharing every wrapper field. Add explicit application-level isolation when nested mutation is unsafe, and do not reuse parent approval decisions for nested calls merely because the tool name or call ID looks similar.
|
||||
- Context serialization is a separate durability decision. Read [RunState schema and resume boundary](runstate-schema.md) before persisting custom context objects, approvals, usage, or nested tool input.
|
||||
|
||||
## Usage Accounting
|
||||
|
||||
- `RunContextWrapper.usage` is the run-wide mutable accumulator. Add each model response exactly once across streaming, non-streaming, retries, nested runs, handoffs, and resume paths.
|
||||
- Preserve authoritative `request_usage_entries` when combining usage. Do not synthesize a second per-request entry from aggregate totals when the provider or retry layer already supplied request-level records.
|
||||
- Retry accounting may include failed attempts with no token totals. Keep request count, aggregate tokens, request-level entries, and trace span usage internally consistent without inventing provider token data.
|
||||
- Streamed usage remains incomplete until terminal chunks and the stream driver finish. Do not finalize billing, result summaries, or usage-bearing spans from the last visible text delta alone.
|
||||
|
||||
## Review Checklist
|
||||
|
||||
1. Test direct construction and clone behavior without mutating shared caller-owned objects.
|
||||
2. Resolve dynamic instructions, tools, and handoffs through the same public agent and current context used for dispatch.
|
||||
3. Verify handoff and internal prepared-agent paths expose the intended public identity and effective output schema.
|
||||
4. Test nested agent tools for shared application state and isolated scoped metadata.
|
||||
5. Compare aggregate and per-request usage after streaming, retries, handoffs, interruption resume, and nested runs.
|
||||
|
||||
## Sources
|
||||
|
||||
- `docs/agents.md`
|
||||
- `docs/context.md`
|
||||
- `docs/results.md`
|
||||
- `src/agents/agent.py`
|
||||
- `src/agents/run_context.py`
|
||||
- `src/agents/tool_context.py`
|
||||
- `src/agents/usage.py`
|
||||
- `src/agents/run_internal/turn_preparation.py`
|
||||
- `src/agents/run_internal/run_loop.py`
|
||||
- `tests/test_agent_config.py`
|
||||
- `tests/test_agent_clone_shallow_copy.py`
|
||||
- `tests/test_agent_as_tool.py`
|
||||
- `tests/test_usage.py`
|
||||
@@ -0,0 +1,67 @@
|
||||
# Conversation State Ownership
|
||||
|
||||
Use this reference for changes involving multi-turn input, sessions, `conversation_id`, `previous_response_id`, `auto_previous_response_id`, compaction, retries, `call_model_input_filter`, or `RunState` resume.
|
||||
|
||||
## Choose One Conversation Strategy
|
||||
|
||||
The state owner determines what the next model request should contain.
|
||||
|
||||
| Strategy | State owner | Next-turn input |
|
||||
|---|---|---|
|
||||
| Explicit replay with `result.to_input_list()` | Application | Replay-ready history plus the new turn |
|
||||
| SDK session | Application storage plus the SDK | The same session plus the new turn |
|
||||
| `conversation_id` | OpenAI Conversations API | The same conversation ID plus only the new turn |
|
||||
| `previous_response_id` or `auto_previous_response_id` | OpenAI Responses API | The previous response ID plus only the new turn |
|
||||
| `RunState` resume | Serialized Agents SDK run | Resume the same interrupted run; this is not a new conversation strategy |
|
||||
|
||||
In normal use, select one conversation strategy. Mixing client-managed replay or sessions with server-managed continuation can duplicate context unless the implementation explicitly reconciles both owners. Read [Session persistence](session-persistence.md) for the client-managed storage contract.
|
||||
|
||||
## Server-Managed Continuation
|
||||
|
||||
- `OpenAIServerConversationTracker` in `src/agents/run_internal/oai_conversation.py` owns delta calculation for `conversation_id`, `previous_response_id`, and `auto_previous_response_id`.
|
||||
- Send only items that the server has not already acknowledged. Object identity is useful only within one process; resume and retry paths also require stable item IDs, tool call IDs, and content fingerprints.
|
||||
- Update `previous_response_id` from the most recent response that actually has an ID. Do not erase a valid chain because an adjacent provider response lacks one.
|
||||
- Session persistence cannot be combined with server-managed continuation. `validate_session_conversation_settings()` rejects a session with `conversation_id`, `previous_response_id`, or `auto_previous_response_id`; do not add a second history writer without defining reconciliation and dedupe semantics.
|
||||
- Treat `conversation_id` and `previous_response_id` / `auto_previous_response_id` chaining as mutually exclusive state owners.
|
||||
|
||||
## Filters, Retries, and Resume
|
||||
|
||||
- `call_model_input_filter` runs on the prepared model payload. With server-managed continuation, that payload may already be a new-turn delta rather than full history.
|
||||
- The filter must return `ModelInputData` with list input. Mark exactly the returned list as sent immediately before the request so nested preparation cannot add unsent items, rewind that tracking before retrying a failed request, and preserve it after success.
|
||||
- Keep streaming and non-streaming tracker updates aligned. Both paths must preserve the same delta, retry, and response-ID semantics.
|
||||
- Stateful retries require replay-safety evidence. Do not blindly resend a request that may already have advanced server state.
|
||||
- `RunState` persists conversation identifiers and reconstructs tracker knowledge for resumed runs. Resume must not replay acknowledged input, lose unsent tool outputs, or increment the turn count without a model call.
|
||||
- Conversation continuation carries context into a new turn. `RunState` resume continues a paused run. Do not substitute one mechanism for the other.
|
||||
|
||||
## Compaction
|
||||
|
||||
- `compaction_mode="previous_response_id"` depends on a usable stored response chain.
|
||||
- `compaction_mode="input"` rebuilds from client-held items and is the fallback when the server chain is unavailable or `store=False` prevents later response lookup.
|
||||
- Compaction must preserve the chosen state owner. Do not compact from local history and then also replay that history through server-managed continuation.
|
||||
|
||||
## Handoffs
|
||||
|
||||
- Server-managed conversations send deltas, so handoff input filters are not supported. `Handoff.input_filter` and `RunConfig.handoff_input_filter` should raise instead of rewriting a history the server already owns.
|
||||
- `nest_handoff_history` is a client-history transformation. When server-managed continuation is active, disable it with a warning and continue with delta-only input.
|
||||
- Keep generated items and session items distinct during handoff processing. The next model input may be filtered, but session history needs the full unfiltered item sequence when client-managed sessions are active.
|
||||
|
||||
## Review Checklist
|
||||
|
||||
1. Name the state owner before changing request construction.
|
||||
2. Specify whether the model receives full history or a delta on every affected path.
|
||||
3. Verify first turn, follow-up turn, retry, interruption, serialized resume, and streaming behavior.
|
||||
4. Test tool calls and outputs separately; call IDs and output fingerprints have different dedupe roles.
|
||||
5. Confirm that filtering, compaction, and session persistence do not introduce a second source of truth.
|
||||
|
||||
## Sources
|
||||
|
||||
- [OpenAI conversation state guide](https://developers.openai.com/api/docs/guides/conversation-state)
|
||||
- [OpenAI running agents guide](https://developers.openai.com/api/docs/guides/agents/running-agents#choose-one-conversation-strategy)
|
||||
- `src/agents/run_internal/oai_conversation.py`
|
||||
- `src/agents/run_internal/run_loop.py`
|
||||
- `src/agents/run_internal/session_persistence.py`
|
||||
- `src/agents/run_state.py`
|
||||
- `docs/running_agents.md`
|
||||
- `docs/sessions/index.md`
|
||||
|
||||
Recheck the official API reference with `$openai-knowledge` before changing server-managed continuation behavior.
|
||||
@@ -0,0 +1,50 @@
|
||||
# Function and Output Schema
|
||||
|
||||
Use this reference for changes to function-tool signature inspection, parameter metadata, strict JSON schema conversion, tool argument reconstruction, or structured agent output schemas.
|
||||
|
||||
Schema behavior is a compatibility boundary shared by Python callables, Pydantic, model providers, and runtime validation. Keep schema generation and invocation aligned rather than fixing one representation in isolation.
|
||||
|
||||
## Function Schema Ownership
|
||||
|
||||
- Explicit decorator arguments for a function name or description override values inferred from the callable and its docstring.
|
||||
- Parameter descriptions from parsed docstrings take precedence over description strings carried by `Annotated`. Preserve `Field` constraints, aliases, and defaults when merging `Annotated` metadata.
|
||||
- A run context parameter is special only in the first parameter position. Exclude it from the model-visible schema while still supplying it during invocation; do not silently treat later context-typed parameters as injected context.
|
||||
- Keep the inspected signature, generated Pydantic model, JSON schema, and `to_call_args()` reconstruction consistent. Cover positional-only parameters, keyword-only parameters, `*args`, and `**kwargs` when changing this path.
|
||||
- Reject unsupported callable shapes or invalid schemas when the tool is constructed so failures do not depend on whether a particular model later selects the tool.
|
||||
|
||||
## Strict JSON Schema Conversion
|
||||
|
||||
- Strict conversion closes object schemas with `additionalProperties: false` and marks their declared properties required. Reject an explicit `additionalProperties: true` instead of silently changing its meaning.
|
||||
- Preserve the meaning of unions, intersections, definitions, and references. Normalize `oneOf` where required, process `allOf`, retain chained references, and merge a referenced schema with sibling keys without discarding the siblings.
|
||||
- Remove defaults that only encode Python `None`; a nullable type must remain represented by its type schema rather than by an unsupported default.
|
||||
- `ensure_strict_json_schema()` may mutate a non-empty input dictionary. Copy caller-owned schemas at public boundaries before conversion. Empty-schema conversion must return a fresh object rather than shared mutable state.
|
||||
- Keep strictness explicit. If a tool or output schema opts out of strict mode, preserve that choice through provider conversion instead of partially applying strict normalization.
|
||||
|
||||
## Structured Output Schemas
|
||||
|
||||
- Plain `str` output and no declared output type use the plain-text path. Pydantic models and dictionary-shaped outputs expose their object schema directly; other Python types use the SDK's wrapper object with the `response` key.
|
||||
- Keep generated output names stable and descriptive for nested generics, unions, and `Literal` types. These names are observable in provider requests and diagnostics.
|
||||
- Parse model output as JSON and validate it through the output type adapter. Convert JSON or validation failures to the SDK's model-behavior error boundary rather than leaking provider- or Pydantic-specific exceptions.
|
||||
- Streaming and non-streaming adapters must carry the same schema, strictness flag, wrapper behavior, and validation result.
|
||||
|
||||
## Review Checklist
|
||||
|
||||
1. Test precedence among explicit metadata, docstrings, `Annotated`, and `Field` values.
|
||||
2. Test invocation reconstruction for positional-only, keyword-only, variadic, and context-bearing callables.
|
||||
3. Test nested objects, unions, intersections, sibling and chained references, nullable fields, and caller-owned schema mutation.
|
||||
4. Test plain text, direct object output, wrapped scalar or generic output, invalid JSON, and validation failure.
|
||||
5. Verify every provider adapter receives the same normalized schema and strictness decision.
|
||||
|
||||
## Sources
|
||||
|
||||
- `src/agents/function_schema.py`
|
||||
- `src/agents/strict_schema.py`
|
||||
- `src/agents/tool.py`
|
||||
- `src/agents/agent_output.py`
|
||||
- `src/agents/models/`
|
||||
- `tests/test_function_schema.py`
|
||||
- `tests/test_function_tool_decorator.py`
|
||||
- `tests/test_strict_schema.py`
|
||||
- `tests/test_strict_schema_oneof.py`
|
||||
- `tests/test_output_tool.py`
|
||||
- `tests/models/`
|
||||
@@ -0,0 +1,56 @@
|
||||
# Local MCP Server Lifecycle
|
||||
|
||||
Use this reference for changes to Python-managed MCP servers, `MCPServerManager`, client-session request ordering, tool caching or filtering, local MCP retries, cancellation, or cleanup. Hosted MCP is a provider tool and follows the OpenAI API contract; use `$openai-knowledge` for that protocol surface. Read [Tool identity and routing](tool-identity.md) for server-prefixed names and [Tool execution lifecycle](tool-execution-lifecycle.md) for approval and invocation behavior after an MCP tool is converted to a `FunctionTool`.
|
||||
|
||||
## Connection Ownership and Task Affinity
|
||||
|
||||
- A local `MCPServer` owns its transport, `ClientSession`, and `AsyncExitStack` from `connect()` through `cleanup()`. Partial connection failure still requires closing every context already entered.
|
||||
- Some MCP transports use AnyIO cancel scopes that require connection and cleanup in the same task. Do not wrap either operation in a helper that silently creates another task.
|
||||
- `MCPServerManager` preserves task affinity in sequential mode and uses one long-lived worker task per server in parallel mode. Timeouts must run inside that owning task; on Python versions without `asyncio.timeout()`, cancel the current worker task and translate only timer-originated cancellation to `TimeoutError`.
|
||||
- Cleanup runs servers in reverse order and continues across ordinary cleanup failures. Cancellation suppression is an explicit manager policy; do not accidentally convert unrelated `BaseException` failures into recoverable connection errors.
|
||||
- Server cleanup must clear session and transport-visible state even when exit-stack cleanup raises, so the same server object can reconnect without exposing stale session handles or workers.
|
||||
|
||||
## Manager State
|
||||
|
||||
- Keep configured servers, connected servers, failed servers, active servers, and per-server errors as distinct views. `active_servers` is the agent-facing list; with `drop_failed_servers=True` it excludes failed connections while preserving configured order.
|
||||
- Non-strict connection records failures and continues with the connected subset. Strict connection cleans up work started by the failed attempt and restores the previous coherent active state before raising.
|
||||
- `reconnect(failed_only=True)` retries the deduplicated failed set without disturbing healthy connections. A full reconnect cleans up all servers first and rebuilds manager state.
|
||||
- Parallel connection still needs deterministic per-server state and complete cleanup after cancellation or one hard failure. Do not let completion order decide `active_servers`, `failed_servers`, or which workers remain registered.
|
||||
|
||||
## Shared Session Requests and Retries
|
||||
|
||||
- Streamable HTTP can require requests on one shared MCP session to be serialized. The same lock must cover tool calls, tool listing, prompts, and resource operations that share that session; serializing only `call_tool()` still permits sibling cancellation and protocol races.
|
||||
- Preserve outer cancellation. A cancelled shared request may qualify for an isolated-session retry only when the transport identifies it as an inner or transient session failure and retry budget remains.
|
||||
- Isolated-session retries are transport-specific recovery. Count isolated session setup and execution against the same retry budget, retry only the supported transient failure shapes, and never replay mixed exception groups or ordinary 4xx failures as if they were safe.
|
||||
- Generic `list_tools()` and `call_tool()` retries use the configured attempt count and backoff. Validate required arguments locally before starting retries so deterministic input errors never reach the server or consume retry budget.
|
||||
- MCP tool failure conversion follows the effective server or agent `failure_error_function`. Explicit `None` means propagate; cancellation of the parent run must not become model-visible tool failure output.
|
||||
|
||||
## Tool Discovery, Cache, and Filtering
|
||||
|
||||
- The unfiltered server tool list is the cacheable value. Apply static or dynamic filters to a copy for each requesting agent and run context; never let one request's filtered or merged metadata mutate the shared cache.
|
||||
- `cache_tools_list=True` assumes server schemas are stable until `invalidate_tools_cache()` marks them dirty. Connection or filter changes must not accidentally make a stale filtered list authoritative.
|
||||
- Dynamic filters require both `run_context` and agent. A filter exception excludes that tool and logs the failure rather than exposing it by default.
|
||||
- Schema conversion to strict form is best effort and must not mutate the MCP server's original input schema. If strict conversion fails, preserve the original schema and keep metadata isolated per converted `FunctionTool`.
|
||||
- Tool list collision errors, prefixed-name generation, and approval policy validation must be deterministic regardless of server response or connection completion order.
|
||||
|
||||
## Review Checklist
|
||||
|
||||
1. Identify the task that owns connect, every request, timeout cancellation, and cleanup for each transport.
|
||||
2. Test partial connect failure, strict and non-strict manager modes, reconnect, repeated cleanup, and manager cancellation.
|
||||
3. Test overlapping tool, prompt, and resource requests when shared-session serialization is enabled.
|
||||
4. Prove retries preserve outer cancellation, consume one budget, and do not replay deterministic or unsupported failures.
|
||||
5. Test cache invalidation, context-dependent filters, schema immutability, duplicate names, and reconnect with the public runner path.
|
||||
|
||||
## Sources
|
||||
|
||||
- `docs/mcp.md`
|
||||
- `src/agents/mcp/server.py`
|
||||
- `src/agents/mcp/manager.py`
|
||||
- `src/agents/mcp/util.py`
|
||||
- `tests/mcp/test_mcp_server_manager.py`
|
||||
- `tests/mcp/test_connect_disconnect.py`
|
||||
- `tests/mcp/test_client_session_retries.py`
|
||||
- `tests/mcp/test_caching.py`
|
||||
- `tests/mcp/test_tool_filtering.py`
|
||||
- `tests/mcp/test_server_errors.py`
|
||||
- `tests/mcp/test_runner_calls_mcp.py`
|
||||
@@ -0,0 +1,75 @@
|
||||
# Model and Provider Boundaries
|
||||
|
||||
Use this reference for changes to model resolution, `ModelSettings`, provider adapters, Responses versus Chat Completions behavior, request conversion, streaming terminal events, transport reuse, or model retries.
|
||||
|
||||
## Core Boundary
|
||||
|
||||
The run loop depends on the `Model` interface, not on one provider's request or response schema.
|
||||
|
||||
- `Model.get_response()` returns a normalized `ModelResponse`.
|
||||
- `Model.stream_response()` yields normalized response stream events while preserving provider payloads needed by public raw-event consumers.
|
||||
- `ModelProvider.get_model()` resolves names to model implementations and owns provider-level caches or connections.
|
||||
- `Model.close()` and `ModelProvider.aclose()` release persistent transport resources when an implementation owns them.
|
||||
|
||||
Provider adapters own request construction, provider feature validation, terminal event interpretation, usage conversion, and translation into SDK item shapes. Keep provider-specific branching out of the core run loop unless it represents a shared SDK contract.
|
||||
|
||||
## Model and Settings Resolution
|
||||
|
||||
- An explicit `RunConfig.model` overrides the agent model. A model instance is used directly; a model name is resolved through the configured `ModelProvider`.
|
||||
- Implicit default settings must follow the resolved model name, including when a run-level model name replaces the agent default.
|
||||
- Resolve agent settings with run-level settings by overlaying non-`None` values. Preserve the documented merge behavior for structured fields such as `extra_args` and retry settings.
|
||||
- Do not pass provider request extras into tracing by default. `ModelSettings.to_traceable_dict()` is the boundary for settings considered safe and meaningful in traces.
|
||||
|
||||
## Capability Ownership
|
||||
|
||||
Do not infer that a feature available in one adapter is supported by every `Model` implementation.
|
||||
|
||||
- Responses-specific features include server-managed response chaining, conversation-aware request fields, tool namespaces, deferred tool loading, tool search, response includes, compaction, and Responses websocket transport.
|
||||
- Chat Completions generally requires client-managed replay and adapter conversion of Responses-compatible SDK items. Unsupported server-state or tool features should be rejected or explicitly ignored according to the adapter's documented validation mode.
|
||||
- Realtime has its own session protocol, event model, and server tracing. Do not route Realtime behavior through the standard Responses or Chat Completions assumptions.
|
||||
- Third-party model adapters may preserve only the shared `Model` contract. New provider-specific fields need an explicit conversion and fallback policy.
|
||||
|
||||
Validate capabilities at the adapter boundary where the resolved model and complete request are known. Avoid public flags that appear accepted by the SDK but are silently dropped before the provider request.
|
||||
|
||||
## Provider Data and Terminal Semantics
|
||||
|
||||
- Preserve provider-supplied string IDs, request IDs, usage, and opaque provider data when the public SDK contract exposes them.
|
||||
- Normalize provider objects and mapping payloads without relying on truthiness for valid empty or zero values.
|
||||
- A transport stream ending is not automatically a successful model response. Responses `failed` and `incomplete` terminals, explicit error events, and a missing terminal payload must produce the documented failure behavior in both HTTP and websocket paths.
|
||||
- Keep semantically equivalent HTTP, websocket, streaming, and non-streaming paths aligned on final `ModelResponse`, errors, request IDs, and usage.
|
||||
|
||||
## Transport Resource Ownership
|
||||
|
||||
- Persistent Responses websocket models are loop-bound resources. Cache reusable websocket model instances by running event loop and model name; do not share one connection or `asyncio.Lock` across loops.
|
||||
- Use weak loop ownership so an unused cache does not keep a closed event loop alive. When a live connection itself pins a closed loop, prune it with synchronous abort and state clearing rather than awaiting work on that closed loop.
|
||||
- A provider that caches persistent models must make `aclose()` close every unique cached model and clear its caches. Close on a still-running owner loop when possible; do not drive an inactive foreign loop inside `asyncio.to_thread()`.
|
||||
- A model used without a running loop cannot safely join the loop-scoped websocket cache. Preserve the non-reuse fallback rather than attaching it to an arbitrary global loop.
|
||||
- Connection reuse ends after protocol errors, pre-terminal disconnects, cancellation that invalidates framing, or explicit close. Clear connection and loop-bound lock state together so a later request cannot reuse half-closed transport state.
|
||||
|
||||
## Retry and Replay Safety
|
||||
|
||||
- Provider retry advice can describe retryability, delay, and replay safety; the runner must not replace provider-specific evidence with a generic status-code assumption.
|
||||
- Requests that use server-managed conversation state or may have produced side effects are not automatically replay-safe. A retry policy must account for whether the provider could have accepted the previous attempt.
|
||||
- Retry conversion and error handlers must preserve the original exception semantics and avoid leaking sensitive request payloads through chaining, logs, traces, or provider error objects.
|
||||
|
||||
## Review Checklist
|
||||
|
||||
1. Identify which adapter owns the feature and how unsupported adapters behave.
|
||||
2. Verify model and implicit-settings resolution when run config overrides the agent.
|
||||
3. Compare HTTP/websocket and streaming/non-streaming terminal behavior when applicable.
|
||||
4. Preserve request IDs, usage, provider data, and error semantics through normalization.
|
||||
5. Prove retries are safe for the request's state ownership and side effects.
|
||||
6. Test transport reuse, cross-loop access, closed-loop pruning, and provider shutdown when persistent connections are involved.
|
||||
|
||||
## Sources
|
||||
|
||||
- `src/agents/models/interface.py`
|
||||
- `src/agents/model_settings.py`
|
||||
- `src/agents/run_internal/turn_preparation.py`
|
||||
- `src/agents/models/openai_responses.py`
|
||||
- `src/agents/models/openai_chatcompletions.py`
|
||||
- `src/agents/models/multi_provider.py`
|
||||
- `src/agents/models/_response_terminal.py`
|
||||
- `src/agents/run_internal/model_retry.py`
|
||||
- `tests/models/`
|
||||
- `tests/test_config.py`
|
||||
@@ -0,0 +1,73 @@
|
||||
# Realtime Session Lifecycle
|
||||
|
||||
Use this reference for `RealtimeSession` changes involving entry, exit, listeners, connections, background tasks, approvals, handoffs, event iteration, tracing context, or cleanup.
|
||||
|
||||
## Resource Ownership
|
||||
|
||||
Treat the session as the owner of these resources once they are acquired:
|
||||
|
||||
| Resource | Acquisition | Required release or terminal state |
|
||||
|---|---|---|
|
||||
| Model listener | `add_listener()` during entry | `remove_listener()` |
|
||||
| Model connection | `model.connect()` | `model.close()` |
|
||||
| Event iterators | Waiting on the event queue | Wake or terminate every waiter on close |
|
||||
| Guardrail tasks | Created during output processing | Complete, or cancel and account for completion |
|
||||
| Tool-call tasks | Created when `async_tool_calls=True` | Complete, or cancel and account for completion |
|
||||
| Pending approvals and outputs | Added during tool execution | Resolve, retain for retry, or clear during terminal cleanup |
|
||||
| Agent and model settings | Updated on handoff or `update_agent()` | Keep runtime state and model configuration aligned |
|
||||
|
||||
Do not add a new side effect before a failure point without defining who releases it.
|
||||
|
||||
## Entry and Exit
|
||||
|
||||
- Python does not call `__aexit__` when `__aenter__` raises. Any listener, connection, task, tracing scope, or other resource acquired before the exception needs explicit failure cleanup.
|
||||
- Keep construction free of external side effects. Acquire listeners and connections during entry where failures can be handled coherently.
|
||||
- `close()` and internal cleanup must be idempotent. Repeated close paths should still wake event iterators without closing the model twice.
|
||||
- Mark the session closed only after the cleanup state is coherent. If model close fails, decide deliberately whether retry is possible and which resources remain owned.
|
||||
|
||||
## Async Task and Context Rules
|
||||
|
||||
- `asyncio` tasks inherit a snapshot of the creator's context. A background task cannot update the caller task's `ContextVar` state.
|
||||
- A `ContextVar` token must be reset in the same context that created it. Never pass a token to a different task and assume cleanup can reset it safely.
|
||||
- Shared session fields can be mutated by the listener path, tool-call tasks, `close()`, `update_agent()`, and handoff handling. Review ordering and races whenever one of those paths changes.
|
||||
- Calling `task.cancel()` requests cancellation; it does not prove the task has finished its `finally` blocks or released resources. Await cancelled tasks when completion matters, or document and test why dropping them is safe.
|
||||
- Background-task exceptions must reach a deterministic owner. They must not silently disappear or leave event consumers blocked.
|
||||
|
||||
## Agent Transitions
|
||||
|
||||
- Handoffs and the public `update_agent()` API are equivalent agent-transition surfaces. Keep their model settings, tool and handoff resolution, emitted events, and tracing metadata aligned unless a difference is intentional and documented.
|
||||
- Resolve dynamic tools and enabled handoffs once per transition when possible, then reuse the exact resolved values for model settings and metadata.
|
||||
- With concurrent tool calls, capture the agent snapshot associated with each call. Do not route a call through whichever agent happens to be current when the task eventually runs.
|
||||
|
||||
## Guardrails and Response Ordering
|
||||
|
||||
- Realtime output guardrails inspect accumulated transcript text at configured debounce thresholds, not each token and not a final `Runner` output object. They emit `guardrail_tripped` instead of raising a normal Runner tripwire exception.
|
||||
- A tripped output guardrail marks the response interrupted before awaiting transport work, emits one trip event per response, forces response cancellation, and sends safe follow-up input naming the guardrail. Concurrent guardrail tasks must not interrupt or message the same response twice.
|
||||
- Guardrail callbacks can run after audio has already been buffered or played. Consumers must treat `audio_interrupted` as the signal to stop local playback; text rejection alone cannot retract audio already delivered.
|
||||
- An exception from one output guardrail is logged and skipped so it does not silently terminate the live session. Exceptions that escape the background guardrail task must become a `RealtimeError` event rather than disappearing.
|
||||
- Realtime function-tool input guardrails follow the same optional pre-approval and mandatory post-approval ordering as standard function tools, but their rejection is returned through Realtime tool output and events.
|
||||
- Follow-up `response.create` work triggered by tools, handoffs, or guardrails must respect the active response lifecycle. Wait for `response.done` or the model layer's equivalent gate before starting a conflicting response.
|
||||
|
||||
## Failure-Path Tests
|
||||
|
||||
Add focused tests for affected phases:
|
||||
|
||||
1. Instruction, tool, or handoff resolution fails during entry.
|
||||
2. Model connection fails after listener registration.
|
||||
3. A background tool or guardrail task raises or is cancelled.
|
||||
4. Cleanup runs while event iterators are waiting.
|
||||
5. `close()` is called repeatedly or from another task.
|
||||
6. A handoff or `update_agent()` fails partway through model-settings application.
|
||||
7. Tool output sending fails after local execution and must be retried without running the tool twice.
|
||||
8. Concurrent guardrail tasks trip once, cancel playback, and do not overlap follow-up responses.
|
||||
|
||||
Verify lifecycle changes with the real public path where feasible; helper-only tests are insufficient when task ownership or context propagation determines the result.
|
||||
|
||||
## Sources
|
||||
|
||||
- `src/agents/realtime/session.py`
|
||||
- `src/agents/realtime/model.py`
|
||||
- `src/agents/realtime/openai_realtime.py`
|
||||
- `tests/realtime/test_session.py`
|
||||
- `tests/realtime/test_session_exceptions.py`
|
||||
- `docs/realtime/guide.md`
|
||||
@@ -0,0 +1,42 @@
|
||||
# Realtime Tracing Architecture
|
||||
|
||||
Use this reference when reviewing or implementing Realtime tracing behavior, especially claims that `RealtimeSession` should emit the same trace hierarchy as `Runner`.
|
||||
|
||||
## Two Separate Tracing Systems
|
||||
|
||||
Realtime integrations involve two independent tracing paths:
|
||||
|
||||
| Path | Owner | Configuration | Result |
|
||||
|---|---|---|---|
|
||||
| Realtime API server tracing | Realtime API | `"auto"`, `workflow_name`, `group_id`, and `metadata` | The server creates a Realtime session trace in the Traces Dashboard. |
|
||||
| Agents SDK client tracing | Agents SDK tracing provider | `trace()`, `agent_span()`, and other SDK span factories | The SDK exports locally created traces and spans through its tracing processor. |
|
||||
|
||||
The current Python SDK has no mapping from an Agents SDK client `trace_id`, `span_id`, or parent context into `RealtimeModelTracingConfig` or the model's `session.update`. A server-created Realtime trace is therefore not attached as a child of an SDK-created trace or span by this implementation. Likewise, adding an SDK `agent_span()` around `RealtimeSession` does not make server-side trace contents children of that span.
|
||||
|
||||
If both paths are enabled, the dashboard can contain two separate traces. A shared `group_id` can make them easier to filter and correlate, but it does not merge them or create a parent-child relationship.
|
||||
|
||||
## Current Python SDK Behavior
|
||||
|
||||
- `RealtimeModelTracingConfig` exposes only `workflow_name`, `group_id`, and `metadata` in `src/agents/realtime/config.py`.
|
||||
- `OpenAIRealtimeWebSocketModel` defaults the Realtime tracing configuration to `"auto"` when the caller does not provide one.
|
||||
- After receiving `session.created`, the model sends the tracing configuration through a `session.update` event.
|
||||
- `RealtimeRunConfig.tracing_disabled` prevents the SDK from enabling Realtime tracing for that session.
|
||||
|
||||
Verify these paths in `src/agents/realtime/openai_realtime.py` and `src/agents/realtime/session.py`; do not rely on old issue descriptions because Realtime tracing support has changed over time.
|
||||
|
||||
## Maintainer Constraints
|
||||
|
||||
1. Identify whether the behavior belongs to the Realtime API's server trace or an Agents SDK client trace created with `trace()`.
|
||||
2. A client-side agent span does not repair missing server tracing and does not create the unified hierarchy produced by `Runner`.
|
||||
3. The current Python SDK cannot place server-created Realtime spans under an SDK-created trace or span because it does not carry client trace parentage through the Realtime tracing configuration. Recheck the live protocol with `$openai-knowledge` before treating that implementation gap as permanent.
|
||||
4. Use the server trace for Realtime model activity. Use a shared `group_id` or metadata when correlation with a client trace is required.
|
||||
5. Parallel SDK spans need an explicit product and maintenance contract covering the dual-trace user experience, async task context, handoff parenting, failure cleanup, and the client-only operations represented by those spans.
|
||||
6. A client trace becoming non-empty is not evidence that Realtime server activity has been captured or parented correctly.
|
||||
|
||||
## Sources
|
||||
|
||||
- `src/agents/realtime/config.py`
|
||||
- `src/agents/realtime/openai_realtime.py`
|
||||
- `src/agents/realtime/session.py`
|
||||
|
||||
Recheck the official API reference with `$openai-knowledge` before changing this guidance or implementing new protocol behavior.
|
||||
@@ -0,0 +1,79 @@
|
||||
# Run Item Lifecycle
|
||||
|
||||
Use this reference for changes to model output processing, `RunItem` types, tool call and output items, stream events, replay conversion, session history, or serialized run state.
|
||||
|
||||
## Item Flow
|
||||
|
||||
The runtime carries one semantic item through several representations:
|
||||
|
||||
1. A model adapter returns provider output in `ModelResponse.output`.
|
||||
2. `process_model_response()` converts recognized output into public `RunItem` objects and internal executable tool-run records in `ProcessedResponse`.
|
||||
3. Tool execution and handoffs add output items and choose a `SingleStepResult.next_step`.
|
||||
4. The resulting items feed `RunResult`, semantic stream events, session persistence, tracing, and `RunState` serialization.
|
||||
5. Replayable items convert back to model input through `RunItem.to_input_item()` or `run_item_to_input_item()` after SDK-only metadata is handled.
|
||||
|
||||
Keep provider payloads, public run items, and internal execution records distinct. A provider item may be observable without requiring local execution, while a local tool-run record may need to preserve the selected SDK tool object and routing identity.
|
||||
|
||||
## Generated, Session, and Model Input Views
|
||||
|
||||
- `new_step_items` describes items generated by the current step.
|
||||
- `session_step_items` preserves the full unfiltered sequence when session history must retain items that a handoff or input filter omitted from the next model request.
|
||||
- `generated_items` is the public observability view and prefers `session_step_items` when present.
|
||||
- Model input is a replay view, not the canonical storage view. Approval placeholders, SDK-only metadata, unsupported IDs, and orphaned calls may need filtering or normalization before an API request.
|
||||
|
||||
Do not force these views into one list. History persistence, user-visible results, and the next provider request have different correctness requirements.
|
||||
|
||||
## Adding or Changing an Item Type
|
||||
|
||||
Update every applicable surface together:
|
||||
|
||||
- `src/agents/items.py` for the public `RunItem` type, accessors, and replay conversion.
|
||||
- `src/agents/run_internal/run_steps.py` for processed response and executable tool-run records.
|
||||
- `src/agents/run_internal/turn_resolution.py` for provider output recognition, item creation, side effects, and next-step selection.
|
||||
- `src/agents/run_internal/tool_execution.py`, `tool_actions.py`, or `tool_planning.py` for execution, dedupe, approvals, and outputs.
|
||||
- `src/agents/run_internal/items.py` for normalization, replay conversion, fingerprints, dedupe, and provider-boundary metadata stripping.
|
||||
- `src/agents/stream_events.py` and streaming queue helpers for public semantic events.
|
||||
- `src/agents/run_state.py` for serialization and deserialization when the item can survive interruption.
|
||||
- `src/agents/run_internal/session_persistence.py` for session conversion, sanitization, and retry accounting.
|
||||
- Tracing and usage conversion when the item contributes observable tool or model work.
|
||||
|
||||
## Compatibility Rules
|
||||
|
||||
- Public stream event names are compatibility-sensitive. Do not rename an existing event, even to fix spelling, without an explicit breaking-change plan.
|
||||
- Preserve provider-supplied IDs and opaque provider data until the owning boundary deliberately removes them. Do not invent IDs or coerce malformed values to make replay appear valid.
|
||||
- Preserve SDK-only metadata needed for display, routing, approvals, tool origin, or resume, but strip it before sending payloads to a provider that does not accept it.
|
||||
- Tool call and output pairs must retain the same string call ID across execution, replay, session persistence, and resume.
|
||||
- Empty, falsey, structured, image, file, and custom tool outputs are valid values unless the public tool contract explicitly rejects them; do not use broad truthiness checks to decide whether output exists.
|
||||
|
||||
## Replay Integrity
|
||||
|
||||
- Prune orphan calls only from runner-generated or resumed history where the SDK owns call/output pairing. Preserve caller-supplied initial input unless an explicit public normalization contract says otherwise.
|
||||
- When dropping an orphan tool call, also drop reasoning items tied to that removed call so the provider does not receive a reasoning item without its required following item. Do not drop a lone reasoning item merely because its following item is absent locally; server-managed conversation state may own that item.
|
||||
- `reasoning_item_id_policy="omit"` strips IDs only from SDK-generated follow-up reasoning items. It does not rewrite initial caller input, must survive `RunState` resume, and can be superseded by a later `call_model_input_filter` that deliberately returns IDs.
|
||||
- Pair anonymous tool-search outputs with the latest compatible anonymous call and never pair a named call with an anonymous output. A missing call ID does not justify inventing a persistent provider identity.
|
||||
- `provider_data` and provider IDs have boundary-specific ownership. Preserve them for raw results and provider requests that accept them, but strip private or replay-unsafe metadata from session and server-conversation history where the SDK contract requires sanitized items.
|
||||
|
||||
## Review Checklist
|
||||
|
||||
1. Follow the item from provider response through result, stream, session, replay, and `RunState`.
|
||||
2. Test both typed provider objects and mapping payloads when adapters support both.
|
||||
3. Verify IDs, metadata, and output values survive every required round-trip.
|
||||
4. Test filtering and dedupe without losing the latest valid call/output pair.
|
||||
5. Compare streaming event order with the non-streaming item sequence.
|
||||
6. Test orphan pruning and reasoning pairing with client-managed replay and server-managed continuation separately.
|
||||
|
||||
## Sources
|
||||
|
||||
- `src/agents/items.py`
|
||||
- `src/agents/stream_events.py`
|
||||
- `src/agents/run_internal/items.py`
|
||||
- `src/agents/run_internal/run_steps.py`
|
||||
- `src/agents/run_internal/turn_resolution.py`
|
||||
- `src/agents/run_internal/session_persistence.py`
|
||||
- `src/agents/run_state.py`
|
||||
- `tests/test_items_helpers.py`
|
||||
- `tests/test_run_internal_items.py`
|
||||
- `tests/test_stream_events.py`
|
||||
- `tests/test_run_state.py`
|
||||
- `docs/running_agents.md`
|
||||
- `docs/results.md`
|
||||
@@ -0,0 +1,67 @@
|
||||
# Runner Lifecycle
|
||||
|
||||
Use this reference for changes to `Runner`, turn accounting, guardrails, hooks, handoffs, interruptions, cancellation, or streaming and non-streaming behavior.
|
||||
|
||||
## Turn Boundary
|
||||
|
||||
A turn is one logical model invocation plus processing of that response. Tool execution, handoff resolution, session persistence, interruption resume, and retries inside that logical invocation do not independently consume turns.
|
||||
|
||||
- Increment the turn counter exactly once when the run loop starts a logical model turn. Transport or provider retries inside `get_new_response()` remain part of that turn.
|
||||
- A handoff changes the current agent, but the next turn begins only when the new agent invokes a model.
|
||||
- Resuming `NextStepInterruption` continues the paused turn. Resolve stored approvals and tool work before deciding whether another model call is needed.
|
||||
- Preserve `max_turns` and the current turn in `RunState`; resume must not reset the budget or charge a turn twice.
|
||||
|
||||
## Guardrail Ordering
|
||||
|
||||
- Input guardrails belong to the starting agent and run only for the initial user input. Do not rerun them after handoffs or when resuming an interruption.
|
||||
- Sequential input guardrails must finish before model-side effects begin. Parallel input guardrails may overlap the model call, so a tripwire or exception must cancel and await the in-flight model task and sibling guardrail tasks.
|
||||
- Tool input guardrails run before the approved tool side effect. Tool output guardrails run after local execution and before the output is accepted into the next step.
|
||||
- Output guardrails run only after a candidate final output exists. Streaming must await them and preserve the same tripwire and exception behavior as non-streaming execution before declaring completion.
|
||||
- Guardrail results are observable run state. Preserve them across handoffs, error handlers, streamed completion, and `RunState` round-trips.
|
||||
|
||||
## Step State Machine
|
||||
|
||||
`SingleStepResult.next_step` is the control boundary after one model response and its local side effects:
|
||||
|
||||
| Step | Meaning |
|
||||
|---|---|
|
||||
| `NextStepRunAgain` | Continue with the current agent and make another model call |
|
||||
| `NextStepHandoff` | Switch the current agent, emit the transition, then continue |
|
||||
| `NextStepFinalOutput` | A final candidate exists; finish terminal hooks, output guardrails, persistence, and result construction |
|
||||
| `NextStepInterruption` | Persist enough processed state to resume pending approvals without rerunning completed work |
|
||||
|
||||
Do not bypass this state machine with path-local completion logic. New terminal or pausable behavior must define non-streaming, streaming, session, tracing, and serialized-resume semantics.
|
||||
|
||||
## Streaming Parity and Cancellation
|
||||
|
||||
- Streaming and non-streaming paths must produce equivalent final output, generated items, current agent, usage, guardrail results, session history, and interruption state for the same model behavior.
|
||||
- Raw transport events may differ, but semantic `RunItemStreamEvent` and `AgentUpdatedStreamEvent` emission must follow the same processed items and agent transitions used by the non-streaming result.
|
||||
- `stream_events()` is the stream driver's cleanup boundary. Keep consuming it until exhaustion after normal completion or `cancel()`, or explicitly close the async iterator; merely breaking after the last visible token does not prove session writes, guardrails, compaction, sandbox cleanup, usage, or terminal errors have settled.
|
||||
- Immediate cancellation marks the result complete and requests task cancellation. `after_turn` cancellation leaves the current model/tool turn running so it can persist state and usage before the next turn. Preserve this distinction instead of treating both modes as queue shutdown.
|
||||
- Terminal run-loop, guardrail, and max-turn errors must be surfaced from `stream_events()` after the required queued events are handled. Preserve `run_loop_exception` as a diagnostic view of the background task, not as a replacement completion primitive.
|
||||
- `task.cancel()` is a request, not cleanup completion. Await cancelled tasks when their `finally` blocks, exceptions, or owned resources affect run correctness.
|
||||
- Keep lifecycle hooks aligned across both paths, especially model start/end, handoff, tool start/end, and final-output hooks.
|
||||
|
||||
## Review Checklist
|
||||
|
||||
1. Identify which turn and which agent own the behavior.
|
||||
2. Trace every `NextStep` outcome, including interruption resume.
|
||||
3. Compare streaming and non-streaming side effects and terminal ordering.
|
||||
4. Test guardrail tripwires and exceptions in sequential and parallel modes when relevant.
|
||||
5. Verify normal exhaustion, explicit iterator close, immediate cancellation, and after-turn cancellation leave the documented result and owned resources in a coherent state.
|
||||
|
||||
## Sources
|
||||
|
||||
- `src/agents/run.py`
|
||||
- `src/agents/run_internal/run_loop.py`
|
||||
- `src/agents/run_internal/run_steps.py`
|
||||
- `src/agents/run_internal/turn_preparation.py`
|
||||
- `src/agents/run_internal/turn_resolution.py`
|
||||
- `src/agents/run_internal/guardrails.py`
|
||||
- `tests/test_agent_runner.py`
|
||||
- `tests/test_agent_runner_streamed.py`
|
||||
- `tests/test_cancel_streaming.py`
|
||||
- `tests/test_guardrails.py`
|
||||
- `tests/test_run_state.py`
|
||||
- `docs/streaming.md`
|
||||
- `docs/results.md`
|
||||
@@ -0,0 +1,64 @@
|
||||
# RunState Schema and Resume Boundary
|
||||
|
||||
Use this reference for changes involving `RunState` serialization, deserialization, approvals, trace state, sandbox state, agent identity, tool output payloads, or any persisted resume data.
|
||||
|
||||
## Compatibility Boundary
|
||||
|
||||
`RunState` is the durable SDK pause/resume boundary. Treat the serialized JSON shape as compatibility-sensitive once a schema version has shipped in a release.
|
||||
|
||||
- `to_json()` always emits `CURRENT_SCHEMA_VERSION`.
|
||||
- `from_json()` must continue reading every version in `SUPPORTED_SCHEMA_VERSIONS`.
|
||||
- Older SDKs intentionally reject newer or unsupported versions rather than attempting forward compatibility.
|
||||
- Unreleased schema versions may be renumbered or squashed before release when intermediate snapshots are intentionally unsupported.
|
||||
- Every supported version must have a non-empty one-line entry in `SCHEMA_VERSION_SUMMARIES`.
|
||||
|
||||
## When to Bump the Schema
|
||||
|
||||
Bump `CURRENT_SCHEMA_VERSION` when a serialized `RunState` snapshot changes in a way that affects resume correctness or would silently lose data when read under an older schema label.
|
||||
|
||||
Examples include:
|
||||
|
||||
- New persisted fields on `RunState`, `ModelResponse`, `ProcessedResponse`, interruptions, approvals, tool outputs, sandbox state, trace state, or agent-owned state.
|
||||
- New run item, tool call, approval, or output item variants that can appear in serialized state.
|
||||
- New SDK-only metadata needed to route, dedupe, approve, retry, or resume a tool call.
|
||||
- A changed meaning for an existing serialized field.
|
||||
|
||||
Do not rely on current-reader tests alone. Add a regression that rewrites `$schemaVersion` to an older supported label when appropriate and proves the old label is accepted, rejected, or migrated deliberately.
|
||||
|
||||
## Identity and Routing State
|
||||
|
||||
Serialized state must preserve enough identity to resume without changing behavior:
|
||||
|
||||
- Agent identity must distinguish duplicate agent names in the same graph.
|
||||
- Function tools should persist canonical lookup keys, including `bare`, `namespaced`, and `deferred_top_level`.
|
||||
- Tool call IDs must remain provider-supplied strings; do not coerce arbitrary values into IDs.
|
||||
- Approval decisions and rejection messages must restore against the same tool identity and call ID they originally targeted.
|
||||
- The per-agent tool-use tracker must preserve stable duplicate-agent identity so tool-choice reset behaves the same after resume.
|
||||
- Server-managed conversation identifiers must restore into `OpenAIServerConversationTracker` without replaying acknowledged input.
|
||||
|
||||
## Context and Secrets
|
||||
|
||||
Context serialization is intentionally conservative.
|
||||
|
||||
- Mapping contexts can round-trip directly.
|
||||
- Custom contexts need explicit serializers and deserializers when exact restoration matters.
|
||||
- Without a safe serializer, snapshots may record metadata and warnings rather than the raw object.
|
||||
- Do not persist secrets in `RunContextWrapper.context`, trace data, tool outputs, or custom data unless the caller explicitly chose that durability boundary.
|
||||
|
||||
## Review Checklist
|
||||
|
||||
1. Identify every serialized field whose shape or meaning changes.
|
||||
2. Decide whether the affected schema version is released or unreleased.
|
||||
3. Update `CURRENT_SCHEMA_VERSION` and `SCHEMA_VERSION_SUMMARIES` when resume compatibility requires it.
|
||||
4. Keep released schema versions readable, or fail with an explicit compatibility error if the old label cannot safely represent the new data.
|
||||
5. Test `to_json()` output, `from_json()` restoration, string round-trips, and resumed execution through the public `Runner.run(...)` or `Runner.run_streamed(...)` path.
|
||||
|
||||
## Sources
|
||||
|
||||
- `src/agents/run_state.py`
|
||||
- `src/agents/result.py`
|
||||
- `src/agents/run_internal/agent_runner_helpers.py`
|
||||
- `src/agents/run_internal/oai_conversation.py`
|
||||
- `src/agents/run_internal/run_steps.py`
|
||||
- `src/agents/run_internal/tool_execution.py`
|
||||
- `tests/test_run_state.py`
|
||||
@@ -0,0 +1,70 @@
|
||||
# Sandbox Runtime Boundary
|
||||
|
||||
Use this reference for changes to sandbox session ownership, `SandboxAgent` preparation, manifests, capabilities, host-path materialization, snapshots, resume state, agent transitions, or cleanup.
|
||||
|
||||
## Runtime Ownership
|
||||
|
||||
The outer `Runner` owns agent turns, approvals, handoffs, tracing, session history, and `RunState`. A sandbox session owns the execution environment, workspace, processes, mounts, and provider-specific connection state. Do not move one layer's lifecycle into the other without defining resume and cleanup behavior for both.
|
||||
|
||||
- A live `SandboxRunConfig.session` is caller-owned. The runner may configure and use it but must not delete or fully tear it down.
|
||||
- A session created or resumed through `SandboxRunConfig.client` is runner-owned. Cleanup runs pre-stop hooks, persists snapshot-backed workspace state, stops and shuts down the session, deletes provider resources when required, and closes dependencies.
|
||||
- Session cleanup must be idempotent and release acquired `SandboxAgent` concurrency guards even when persistence or provider cleanup fails.
|
||||
- A `SandboxAgent` instance cannot be reused concurrently across runs because prepared capability tools and session state are bound to one live run. Clone or construct separate agents for concurrent work.
|
||||
|
||||
## Session Source and Saved State
|
||||
|
||||
Resolve the session source in this order: injected live session, resumable sandbox state carried by `RunState`, explicit `SandboxRunConfig.session_state`, then a newly created session. Manifest and snapshot inputs seed only a fresh session; they do not overwrite an injected or resumed workspace.
|
||||
|
||||
- `RunState` sandbox data and explicit `session_state` represent provider connection or session state used to reconnect to existing work.
|
||||
- A snapshot represents saved workspace contents used to seed a new session. It is not interchangeable with provider session state.
|
||||
- Preserve stable per-agent resume identity across handoffs, including graphs with duplicate agent names. Object identity is process-local, so serialized state needs stable keys and explicit current-agent selection.
|
||||
- Serialize runner-owned sessions after stop-time persistence has completed so a later resume can reattach when the backend survives or reconstruct the workspace from the saved snapshot when it does not.
|
||||
|
||||
## Agent Preparation
|
||||
|
||||
- Clone capability instances per run before binding them to a live session. Reusing mutable capability objects can leak tools, sampling settings, or session references across runs.
|
||||
- Validate capability dependencies before exposing tools. Capability tool construction, instruction fragments, input processing, and sampling adjustments must use the same effective capability set.
|
||||
- Build instructions in the documented order: SDK sandbox base prompt or explicit replacement, agent instructions, capability instructions, remote-mount policy, then the rendered filesystem description.
|
||||
- Bind capability tools to the live session and preserve a link from the prepared clone to the public `SandboxAgent`. Dynamic instructions and hooks should observe the public agent rather than an internal clone with implementation-only state.
|
||||
- Handoffs stay in the outer run loop and select another agent-bound sandbox session. A nested `Agent.as_tool()` run owns its own nested runner and sandbox lifecycle.
|
||||
|
||||
## Filesystem Trust Boundary
|
||||
|
||||
- Manifest entry destinations are workspace-relative and must not escape the workspace. The workspace root itself must be absolute where the backend requires an absolute runtime root.
|
||||
- `LocalFile` and `LocalDir` sources are host-side inputs. Resolve them against a trusted base directory, require explicit application-controlled `extra_path_grants` outside that base, and reject untrusted manifests that try to authorize their own host access.
|
||||
- Validate local sources at use time, not only when parsing the manifest. Defend against symlinked sources, parent-directory swaps, platform path aliases, and archive members that change meaning between validation and extraction.
|
||||
- Archive extraction must reject traversal, unsafe links, and unsupported member types before writing, and enforce entry, byte, and expansion limits without materializing an unbounded member list.
|
||||
- Extra path grants are runtime access, not durable workspace content. Snapshots and `persist_workspace()` include the workspace root, not arbitrary granted paths.
|
||||
- Credentials for mounts or providers must remain in the owning adapter and must not appear in generated shell commands, model-visible errors, logs, or serialized sandbox state.
|
||||
|
||||
## Provider and Error Boundary
|
||||
|
||||
- Normalize backend failures to sandbox errors without discarding provider details needed for diagnosis. Preserve explicit retryability instead of inferring it later from a message string.
|
||||
- Keep portable sandbox paths separate from host filesystem paths and provider identifiers. Conversion belongs in the backend or materialization boundary, not in agent-facing tools.
|
||||
- Temporary clones, mounts, sinks, and dependency resources need failure cleanup during partial startup as well as normal shutdown.
|
||||
- Capability tools should report bounded output and preserve provider exit status or structured error data without exposing private runtime metadata to the model.
|
||||
|
||||
## Review Checklist
|
||||
|
||||
1. Name the owner of every live session, provider client, mount, process, capability, and temporary resource.
|
||||
2. Test injected, resumed, explicit-state, snapshot-seeded, and fresh-session paths separately.
|
||||
3. Verify handoffs, duplicate agent names, interruption resume, and cleanup failure preserve the intended session mapping.
|
||||
4. Test host-path, symlink, traversal, archive-limit, and credential-redaction boundaries on applicable platforms.
|
||||
5. Exercise the public `Runner` path so agent preparation, capability binding, persistence, and cleanup run together.
|
||||
|
||||
## Sources
|
||||
|
||||
- `docs/sandbox/guide.md`
|
||||
- `docs/sandbox/clients.md`
|
||||
- `src/agents/sandbox/runtime.py`
|
||||
- `src/agents/sandbox/runtime_session_manager.py`
|
||||
- `src/agents/sandbox/runtime_agent_preparation.py`
|
||||
- `src/agents/sandbox/manifest.py`
|
||||
- `src/agents/sandbox/materialization.py`
|
||||
- `src/agents/sandbox/workspace_paths.py`
|
||||
- `src/agents/sandbox/session/archive_extraction.py`
|
||||
- `tests/sandbox/test_runtime.py`
|
||||
- `tests/sandbox/test_runtime_agent_preparation.py`
|
||||
- `tests/sandbox/test_session_state_roundtrip.py`
|
||||
- `tests/sandbox/test_materialization.py`
|
||||
- `tests/sandbox/test_extract.py`
|
||||
@@ -0,0 +1,80 @@
|
||||
# Session Persistence
|
||||
|
||||
Use this reference for changes to client-managed sessions, session input callbacks, per-turn persistence, retries, rewind, compaction replacement, or session backend implementations.
|
||||
|
||||
Read [Conversation state ownership](conversation-state-ownership.md) first when server-managed continuation is also involved. A client-managed session is a history store; it is not a second owner for a server-managed conversation.
|
||||
|
||||
## Session Contract
|
||||
|
||||
- `get_items(limit=N)` returns the latest `N` items in chronological order.
|
||||
- `add_items()` appends one logical batch. Backends should make the batch atomic so partial turns are not visible after failure.
|
||||
- `pop_item()` removes the current tail item and is used only for guarded rollback of items the current run can prove it owns.
|
||||
- `clear_session()` clears the session boundary; compaction decorators that replace history must provide stronger restore behavior around destructive replacement.
|
||||
|
||||
Third-party implementations target the `Session` protocol. Internal base classes and backend-specific metadata are not the compatibility contract unless explicitly documented.
|
||||
|
||||
## Backend Consistency
|
||||
|
||||
- An explicit `get_items(limit=N)` argument overrides the backend's default session limit. Return the latest `N` items in chronological order, with a deterministic tie-breaker when timestamps can collide.
|
||||
- Preserve caller batch order. Persist the items and any indexes or structural metadata required to read them as one atomic operation. A failed batch must leave earlier history unchanged, and any backend-internal retry must not create duplicates.
|
||||
- Serialize initialization and conflicting writes at the backend's actual consistency boundary. Concurrent first writers must not race, and cancellation or failure must not strand locks or transactions.
|
||||
- Apply configured table or collection names and session settings consistently across reads, writes, deletes, metadata updates, and wrapper operations.
|
||||
- For backends that deserialize stored records, a corrupt record must not hide valid history or cause unrelated records to be deleted. Define consistent `get_items()` and `pop_item()` behavior that isolates the bad record and continues safely.
|
||||
- Preserve creation timestamps and advance update timestamps deliberately. Backend-only identifiers and metadata must not leak into model-facing session items.
|
||||
- Close only resources the backend owns. An injected engine, client, or connection remains caller-owned unless the public contract explicitly transfers ownership.
|
||||
|
||||
## Preparing Input Versus Persisting Input
|
||||
|
||||
`prepare_input_with_session()` returns two different values: the normalized input for the next model request and the subset of new-turn items that should be appended to the session.
|
||||
|
||||
- Existing history must not be re-appended as new input, even when `session_input_callback` deep-copies, reorders, filters, duplicates, or reconstructs items.
|
||||
- A callback may change the model view without rewriting already stored history.
|
||||
- Handoff and model-input filters may omit items from the next request while `session_step_items` retains the complete unfiltered sequence for history and observability.
|
||||
- Normalize and deduplicate the model request and persistence candidates through the same canonical item helpers, then apply boundary-specific sanitization.
|
||||
|
||||
## Per-Turn Save and Resume
|
||||
|
||||
- Persist each completed turn, not only the final run result. Tool outputs and handoff items must survive a later error or interruption.
|
||||
- `_current_turn_persisted_item_count` tracks which generated items have already been saved during streaming, retry, or resume. Count items after conversion and persistence filtering, not from the unsanitized source list.
|
||||
- Resuming an interruption must save newly produced approval and tool output items without duplicating inputs or previously persisted outputs.
|
||||
- Preserve full session items separately from filtered model input when updating `RunState` after resume.
|
||||
- A guardrail trip must preserve the accepted user input while excluding speculative assistant or tool work that the tripwire invalidated. Test sequential and parallel guardrails in streaming and non-streaming modes because their persistence timing differs even though the resulting history must remain coherent.
|
||||
|
||||
## Retry Rewind
|
||||
|
||||
Retry cleanup is ownership-sensitive and best effort.
|
||||
|
||||
- Rewind only an exact serialized suffix that belongs to the failed attempt. Never scan backward and delete merely similar historical items.
|
||||
- Verify the complete suffix before popping. If a pop fails or returns an unexpected item, restore already popped items in chronological order.
|
||||
- Wait for backends with asynchronous cleanup semantics before starting the next retry when stale tail items could be observed.
|
||||
- Do not forward live `RunContextWrapper` objects through retry rewind or compaction storage paths unless the session API explicitly owns that runtime context.
|
||||
|
||||
## Compaction Replacement
|
||||
|
||||
- Treat history replacement as a transaction: capture the prior state, apply the compacted state, and restore the prior state if clear or replacement fails.
|
||||
- Defer response-based compaction while local tool outputs still need to be associated with the response chain.
|
||||
- Choose input-based or previous-response-based compaction according to the actual state owner and `store` behavior; do not combine a local replay with a server-owned history chain.
|
||||
- Compaction output is a run item and must follow the item lifecycle, session sanitization, and `RunState` rules rather than bypassing them as backend-only data.
|
||||
|
||||
## Review Checklist
|
||||
|
||||
1. Distinguish model input, new-turn persistence candidates, and full session history.
|
||||
2. Test atomic failure, duplicate content, reordered callbacks, and filtered handoff input.
|
||||
3. Test save behavior after tool execution, handoff, guardrail trip, interruption, and resume.
|
||||
4. Prove retry rewind removes only the attempt-owned suffix and restores on partial failure.
|
||||
5. Test compaction replacement failures without losing the previous history.
|
||||
6. Test backend ordering, atomic batches, concurrent first writes, configured names and limits, corrupt records, and resource ownership.
|
||||
|
||||
## Sources
|
||||
|
||||
- `src/agents/memory/session.py`
|
||||
- `src/agents/memory/session_settings.py`
|
||||
- `src/agents/memory/sqlite_session.py`
|
||||
- `src/agents/extensions/memory/`
|
||||
- `src/agents/run_internal/session_persistence.py`
|
||||
- `src/agents/run_internal/items.py`
|
||||
- `src/agents/run_internal/run_steps.py`
|
||||
- `tests/memory/`
|
||||
- `tests/extensions/memory/`
|
||||
- `tests/test_agent_runner.py`
|
||||
- `tests/test_agent_runner_streamed.py`
|
||||
@@ -0,0 +1,64 @@
|
||||
# Tool Execution Lifecycle
|
||||
|
||||
Use this reference for changes to function-tool planning, approvals, tool guardrails, concurrency, cancellation, timeouts, hooks, error conversion, or resumed execution. Read [Tool identity and routing](tool-identity.md) when names, namespaces, lookup keys, or call IDs also change.
|
||||
|
||||
## Plan Before Side Effects
|
||||
|
||||
`process_model_response()` discovers executable work, but `tool_planning.py` decides which work may run now. Keep discovery, approval partitioning, and invocation as separate phases.
|
||||
|
||||
- Fresh and resumed turns need different plans. A resumed interruption must execute unresolved or newly approved work without rediscovering or rerunning completed calls.
|
||||
- Approval state is authoritative once resolved. Do not call a dynamic `needs_approval` checker again for a call whose status is already approved or rejected.
|
||||
- Deduplicate by invocation identity before execution while preserving model order for public call and output items. A repeated tool definition is not a repeated call, and a repeated call ID must not execute twice.
|
||||
- Validate enabled tools and canonical lookup before side effects. A tool disabled after model output or absent from the resolved tool set must follow the configured missing-tool behavior rather than reaching a stale callable.
|
||||
|
||||
## Approval and Guardrail Ordering
|
||||
|
||||
- Pre-approval input guardrails are an early rejection optimization. They may run before an approval interruption, but input guardrails must run again immediately before invocation because state, policy, or arguments may have changed while approval was pending.
|
||||
- Rechecking guardrails does not mean rechecking approval. Persisted approval decisions and rejection messages must remain attached to the same tool identity and call ID across `RunState` resume.
|
||||
- Tool input guardrails finish before the local side effect. Tool output guardrails finish before output becomes accepted run state, model input, or persisted session history.
|
||||
- The tool guardrail pipeline applies to `FunctionTool` invocation. Handoffs, hosted tools, built-in provider tools, and nested `Agent.as_tool()` runs have separate execution boundaries unless they explicitly opt into equivalent checks.
|
||||
|
||||
## Concurrency and Failure Semantics
|
||||
|
||||
SDK-side function-tool concurrency is independent of provider-side parallel tool-call generation. The provider controls how many calls appear in one response; `RunConfig.tool_execution.max_function_tool_concurrency` controls how many local function handlers run at once.
|
||||
|
||||
- Preserve model order in emitted outputs even when handlers complete out of order.
|
||||
- Isolate sibling results. A cancelled or failed call must not discard outputs already produced by successful siblings.
|
||||
- Distinguish cancellation of one tool handler from cancellation of the parent run. Tool-local cancellation can follow the configured tool failure policy; parent cancellation must propagate promptly instead of becoming model-visible tool output.
|
||||
- `task.cancel()` is not terminal cleanup. On sibling failure, drain cancelled handlers and wait for post-invocation work within the bounded cleanup policy. On parent cancellation, cancel remaining tasks and attach result callbacks so late exceptions are observed without delaying cancellation indefinitely.
|
||||
- Select and raise failures deterministically when several tasks fail, while still observing secondary failures. Do not let task-set iteration order or eager task execution change the public result.
|
||||
|
||||
## Invocation Boundary
|
||||
|
||||
- Decorated synchronous Python functions run through `asyncio.to_thread()` so they do not block the event loop. Async function tools run in the event loop and are the only decorated handlers that support SDK timeouts.
|
||||
- Timeout handling and ordinary exception handling are distinct policies. `timeout_behavior` and `timeout_error_function` own timeout conversion; `failure_error_function=None` means ordinary exceptions propagate instead of becoming model-visible output.
|
||||
- Tool start/end hooks and function spans surround the actual invocation once per call, including failure and cancellation paths. Do not emit a successful end state before output guardrails complete.
|
||||
- Per-run resources such as resolved `Computer` implementations must be initialized and disposed by the run that acquired them.
|
||||
- Nested `Agent.as_tool()` execution owns a nested run loop and nested resumable state. Scope cached nested state by the parent `RunState` and call identity, not only by the reusable agent or tool object.
|
||||
- `AgentToolUseTracker` records tool use per agent identity. When `reset_tool_choice=True`, reset the effective next-turn tool choice after that agent uses a tool so `required` or a named choice cannot force an accidental loop; do not mutate the agent's declared settings across independent runs.
|
||||
- Persist and restore the tool-use tracker across interruption and sandbox resume, including graphs with duplicate agent names, so resumed tool-choice behavior matches uninterrupted execution.
|
||||
|
||||
## Review Checklist
|
||||
|
||||
1. Trace fresh execution, approval interruption, approval rejection, and serialized resume separately.
|
||||
2. Verify guardrail, approval, hook, trace, invocation, output, and persistence order.
|
||||
3. Test sequential, bounded-concurrency, sibling failure, tool-local cancellation, and parent cancellation paths.
|
||||
4. Test default, custom, and disabled failure conversion plus timeout behavior where applicable.
|
||||
5. Confirm every started task and per-run resource reaches a deterministic terminal state.
|
||||
|
||||
## Sources
|
||||
|
||||
- `docs/running_agents.md`
|
||||
- `docs/tools.md`
|
||||
- `docs/guardrails.md`
|
||||
- `docs/human_in_the_loop.md`
|
||||
- `src/agents/run_internal/tool_planning.py`
|
||||
- `src/agents/run_internal/tool_execution.py`
|
||||
- `src/agents/tool.py`
|
||||
- `tests/test_agent_runner.py`
|
||||
- `tests/test_agent_runner_streamed.py`
|
||||
- `tests/test_function_tool.py`
|
||||
- `tests/test_tool_guardrails.py`
|
||||
- `tests/test_tool_choice_reset.py`
|
||||
- `tests/test_tool_use_tracker.py`
|
||||
- `tests/test_run_state.py`
|
||||
@@ -0,0 +1,71 @@
|
||||
# Tool Identity and Routing
|
||||
|
||||
Use this reference for changes involving function-tool names, namespaces, provider wire names, lookup, approvals, tracing, MCP exposure, handoffs, or tool call IDs.
|
||||
|
||||
## Identity Layers
|
||||
|
||||
One tool can have several related identifiers. They are not interchangeable.
|
||||
|
||||
| Layer | Purpose | Canonical source |
|
||||
|---|---|---|
|
||||
| Public name | User- and model-facing tool name | `tool.name` |
|
||||
| Explicit namespace | Distinguishes tools with the same public name | Tool namespace metadata |
|
||||
| Qualified or dispatch name | Routes a model call to the intended tool | `namespace.name` when a namespace exists |
|
||||
| Lookup key | Collision-free internal identity | `bare`, `namespaced`, or `deferred_top_level` tuple |
|
||||
| Approval keys | Matches approval decisions to the intended tool | Canonical qualified and permitted alias keys |
|
||||
| Trace name | Human-readable tracing label | Explicit trace name or public name |
|
||||
| Call ID | Identifies one invocation, not the tool definition | Provider-supplied string |
|
||||
|
||||
Do not collapse these layers into one string or introduce local rules that only one caller uses.
|
||||
|
||||
## Canonical Helpers
|
||||
|
||||
Use `src/agents/_tool_identity.py` as the single implementation layer. Important helpers include:
|
||||
|
||||
- `get_function_tool_lookup_key_for_tool()` and `get_function_tool_lookup_key_for_call()` for canonical lookup identity.
|
||||
- `get_function_tool_dispatch_name()` and `get_function_tool_qualified_name()` for routing and display surfaces that require qualification.
|
||||
- `get_function_tool_approval_keys()` for approval matching.
|
||||
- `get_function_tool_trace_name()` and `get_tool_call_trace_name()` for trace labels.
|
||||
- `validate_function_tool_lookup_configuration()` and `build_function_tool_lookup_map()` for collision detection and dispatch maps.
|
||||
- `normalize_tool_call_for_function_tool()` when provider payloads must be normalized for a selected tool.
|
||||
|
||||
If a proposed change bypasses these helpers, first prove that the target surface has intentionally different semantics.
|
||||
|
||||
## MCP and Handoff Rules
|
||||
|
||||
- `include_server_in_tool_names` is opt-in. Server-prefixed MCP names affect the model-exposed collision-safe name; they do not rename the original tool on the MCP server.
|
||||
- Reserved names and enabled handoff names participate in collision avoidance only on the paths that expose generated model-facing names.
|
||||
- `Handoff.default_tool_name()` is the source of default handoff tool names. Keep Realtime and non-Realtime handoff conversion aligned with it.
|
||||
- Do not forward a naming option through a path where the downstream helper does not consult it and then describe the change as runtime behavior. Trace the complete caller-to-dispatch path first.
|
||||
|
||||
## Deferred Tool Search Rules
|
||||
|
||||
- A top-level `FunctionTool` with `defer_loading=True` and no explicit namespace uses the synthetic lookup key `("deferred_top_level", tool.name)`.
|
||||
- The Responses wire shape for a loaded deferred top-level tool can look like `namespace == name`. Treat that namespace as reserved for the synthetic deferred tool-search path, not as a normal explicit namespace.
|
||||
- `tool_namespace()` must reject an explicit namespace that equals the inner tool name. Otherwise a normal namespaced tool and a deferred top-level tool would have the same wire shape.
|
||||
- Preserve the synthetic namespace on approval, interruption, tracing, and `ToolContext` surfaces when it identifies the model call, but dispatch the actual local tool through the deferred lookup key and strip the synthetic namespace before invoking the tool.
|
||||
- Permanent approvals for deferred top-level tools should key by `deferred_top_level:<name>`. A bare-name approval alias is allowed only when no visible bare sibling can make that alias ambiguous.
|
||||
|
||||
## Tool Call ID Rules
|
||||
|
||||
- Preserve provider-supplied string call IDs across call items, approvals, outputs, retries, and serialized state.
|
||||
- Do not coerce arbitrary values with `str(...)`. Canonical extractors return a call ID only when the source value is already a string.
|
||||
- Do not use a call ID as a tool-definition identity or a tool name as an invocation identity.
|
||||
- When a provider omits a stable identifier, use an existing fingerprint or dedupe policy for that item type instead of inventing a cross-provider ID contract.
|
||||
|
||||
## Review Checklist
|
||||
|
||||
1. Identify every identifier layer affected by the change.
|
||||
2. Trace the actual runtime path from model-visible name to lookup, approval, invocation, output, and trace metadata.
|
||||
3. Compare adjacent canonical helpers before adding conversion or fallback behavior.
|
||||
4. Test collisions between bare, namespaced, deferred, MCP, local function, and handoff tools when applicable.
|
||||
5. Require a regression test that fails on the base and proves the model-visible or dispatch behavior, not only an intermediate argument value.
|
||||
|
||||
## Sources
|
||||
|
||||
- `src/agents/_tool_identity.py`
|
||||
- `src/agents/agent.py`
|
||||
- `src/agents/mcp/`
|
||||
- `src/agents/handoffs/__init__.py`
|
||||
- `src/agents/run_internal/tool_execution.py`
|
||||
- `src/agents/run_state.py`
|
||||
@@ -0,0 +1,58 @@
|
||||
# Tracing Lifecycle
|
||||
|
||||
Use this reference for changes to SDK trace or span context, processors, export, flush, shutdown, resumed trace state, or sensitive-data handling. Read [Realtime tracing architecture](realtime-tracing.md) before applying these client-side rules to Realtime server traces.
|
||||
|
||||
## Context and Parenting
|
||||
|
||||
- The current trace and span are held in `ContextVar` state. Async tasks inherit a snapshot when created; later changes in a child task do not rewrite the parent task's context.
|
||||
- A context token must be reset in the context that created it. Start and finish ownership cannot be transferred between tasks without an explicit context boundary.
|
||||
- A no-op trace or span cannot be a real parent. Propagate no-op behavior instead of exporting children with the sentinel `no-op` trace or span ID.
|
||||
- Span factories should inherit trace metadata needed by processors, but they must not mutate the trace's caller-owned metadata mapping.
|
||||
|
||||
## Run and Resume Ownership
|
||||
|
||||
- A runner-created trace encloses run-loop-owned guardrails, model calls, tool execution, handoffs, session persistence, and error handling. Do not assume every completion callback or resource cleanup runs before trace finish; place newly traced cleanup explicitly inside the trace lifetime or create a deliberate separate trace/span context.
|
||||
- An existing caller trace remains caller-owned. `Runner` may create child spans but must not finish or flush the caller's trace.
|
||||
- `RunState` stores enough trace metadata to continue an interrupted run. Resume may reattach only when the trace ID was previously started in the process and the effective workflow name, group ID, metadata, and tracing key identity still match.
|
||||
- Reattachment must not emit a duplicate trace-start event. If the saved state cannot prove a compatible live trace, create a normal trace according to the current run configuration instead of pretending to resume the old context.
|
||||
- Tracing API keys are omitted from serialized `RunState` by default. A hash can verify that the caller supplied the same explicit key without persisting the secret; raw key persistence is opt-in.
|
||||
|
||||
## Processor and Export Isolation
|
||||
|
||||
- Trace processors are observability extensions and must not change application success. Catch processor callback, exporter, flush, and shutdown failures and report them as non-fatal.
|
||||
- The default batch worker starts lazily on first queued item to avoid import-time thread and fork hazards. Keep top-level imports free of worker creation and shutdown-handler duplication.
|
||||
- An exporter exception must not kill the batch worker and strand future traces. Drop or report the failed batch according to policy, then keep the worker usable.
|
||||
- `flush_traces()` waits for queued and in-flight export work, so callers should invoke it after the trace closes when they require immediate delivery. It is not a substitute for finishing a partially built trace.
|
||||
- Shutdown is best effort and deadline-aware. It should request exporter shutdown, interrupt retry backoff, drain within the remaining deadline, and return without changing the process exit code when an exporter blocks or a backend remains unavailable.
|
||||
- Keep `TraceProvider.force_flush()` and `shutdown()` defaulting to no-ops for compatibility with custom providers that predate these lifecycle methods.
|
||||
|
||||
## Data Boundaries
|
||||
|
||||
- `trace_include_sensitive_data=False` controls captured span payload fields; it does not automatically sanitize exception objects, chaining, tracebacks, logs, or telemetry created elsewhere.
|
||||
- Redaction must cover `__cause__`, `__context__`, formatter failures, and model-visible error conversion when an original exception carries tool arguments or provider payloads. `raise ... from None` changes display, not object retention.
|
||||
- The OpenAI trace exporter owns ingest-specific payload sanitization such as field-size limits and supported usage keys. Custom processors should continue receiving the SDK's normal trace data unless their contract says otherwise.
|
||||
- Per-run tracing keys, organization, and project routing must stay attached to the trace or exported item that selected them; do not let mutable global exporter state reroute an already-created trace.
|
||||
|
||||
## Review Checklist
|
||||
|
||||
1. Identify which task and context own each trace and span start, finish, and token reset.
|
||||
2. Test success, exception, cancellation, interruption, serialized resume, full stream exhaustion, and explicit stream close.
|
||||
3. Verify processor and exporter failures remain non-fatal and do not kill later export work.
|
||||
4. Test flush and shutdown with queued work, in-flight export, retry backoff, and a blocking exporter.
|
||||
5. Audit sensitive data through span payloads, exception chains, logs, and serialized state.
|
||||
|
||||
## Sources
|
||||
|
||||
- `docs/tracing.md`
|
||||
- `src/agents/tracing/context.py`
|
||||
- `src/agents/tracing/scope.py`
|
||||
- `src/agents/tracing/traces.py`
|
||||
- `src/agents/tracing/spans.py`
|
||||
- `src/agents/tracing/provider.py`
|
||||
- `src/agents/tracing/processors.py`
|
||||
- `src/agents/tracing/setup.py`
|
||||
- `src/agents/run_state.py`
|
||||
- `tests/test_trace_processor.py`
|
||||
- `tests/test_tracing.py`
|
||||
- `tests/test_run_state.py`
|
||||
- `tests/tracing/test_import_side_effects.py`
|
||||
@@ -0,0 +1,56 @@
|
||||
# Voice Pipeline Lifecycle
|
||||
|
||||
Use this reference for changes to `VoicePipeline`, `AudioInput`, `StreamedAudioInput`, STT sessions, TTS task ordering, voice lifecycle events, PCM framing, result streaming, or voice tracing. Realtime agents use a different live-session architecture; read [Realtime session lifecycle](realtime-session-lifecycle.md) for that path.
|
||||
|
||||
## Pipeline Ownership
|
||||
|
||||
`VoicePipeline` owns an STT-to-workflow-to-TTS producer task and returns a `StreamedAudioResult` that drives its observable completion.
|
||||
|
||||
- Static `AudioInput` produces one transcription and one workflow turn. `StreamedAudioInput` creates a long-lived transcription session and runs one workflow turn for each emitted transcript until the input or session ends.
|
||||
- The multi-turn pipeline owns the transcription session and closes it in `finally` before marking output complete. Partial setup and workflow failure must not strand the STT connection or producer task.
|
||||
- `workflow.on_start()` applies only to the streamed multi-turn path. Its failure is logged and skipped so the transcription session can still start; normal per-turn workflow failures are terminal and surface through the result stream.
|
||||
- The SDK does not provide application-level interruption handling for `StreamedAudioInput`. Lifecycle events expose turn boundaries, but microphone muting, playback interruption, and barge-in policy remain application-owned.
|
||||
|
||||
## Text, Audio, and Event Ordering
|
||||
|
||||
- A workflow can yield multiple text fragments. The text splitter returns ready-to-synthesize text plus a remainder; synthesize non-empty ready text even when it is shorter than a default sentence threshold, and retain the remainder for the turn's final flush.
|
||||
- TTS segment tasks may run concurrently, but `_ordered_tasks` and the dispatcher must emit their audio and lifecycle events in workflow text order rather than completion order.
|
||||
- `turn_started` precedes audio for that turn. `turn_ended` is emitted only after the turn's final text remainder has been synthesized and its audio dispatched. `session_ended` follows all ordered segment queues and all turns.
|
||||
- A `VoiceStreamEventError` terminates result streaming and the stored exception is raised after task cleanup. `session_ended` is a lifecycle marker, not proof of success; consumers must still observe the terminal exception from `stream()`.
|
||||
- Consuming `StreamedAudioResult.stream()` is the public completion and error boundary. On normal `session_ended`, let the producer finish before cleanup so session close and trace end are not cancelled by result teardown.
|
||||
|
||||
## PCM and Caller Data
|
||||
|
||||
- PCM16 samples span two bytes. Preserve a trailing half-sample across TTS chunks, combine it with the next chunk, and pad only the final unmatched byte at end of segment.
|
||||
- Apply `buffer_size` to TTS source chunks without changing sample order. Convert to float32 only after PCM16 framing is complete, then apply caller-provided `transform_data` to each emitted array.
|
||||
- `AudioInput.to_base64()` and audio-file conversion must not mutate the caller's NumPy buffer when converting float input to PCM16.
|
||||
- Empty input and empty text-splitter output are valid boundaries. They must not cause NumPy reduction errors, phantom TTS calls, or missing turn/session lifecycle events.
|
||||
|
||||
## Trace Lifetime and Data
|
||||
|
||||
- The pipeline trace stays active for the full asynchronous producer lifecycle, not only until `VoicePipeline.run()` returns its result object.
|
||||
- Each output turn owns a speech-group span and each synthesized segment owns a child speech span. Finish the turn span after ordered audio dispatch and finish the pipeline trace after STT session close and output completion.
|
||||
- Text and audio sensitivity are independent controls. `trace_include_sensitive_data` governs transcript and TTS text, while `trace_include_sensitive_audio_data` governs encoded audio payloads.
|
||||
- Error paths must finish active speech spans and the enclosing trace without replacing the original pipeline exception.
|
||||
|
||||
## Review Checklist
|
||||
|
||||
1. Test static and streamed input, including STT setup failure, workflow failure, TTS failure, and transcription-session close.
|
||||
2. Verify fragment concurrency never changes audio, turn, or session event order.
|
||||
3. Test short splitter output, empty output, odd-byte chunks, cross-chunk sample boundaries, int16, and float32 conversion.
|
||||
4. Consume the public result stream and verify terminal errors, task cleanup, session close, and trace-end order.
|
||||
5. Confirm sensitive text and audio are independently omitted from trace payloads.
|
||||
|
||||
## Sources
|
||||
|
||||
- `docs/voice/pipeline.md`
|
||||
- `docs/voice/tracing.md`
|
||||
- `src/agents/voice/pipeline.py`
|
||||
- `src/agents/voice/result.py`
|
||||
- `src/agents/voice/input.py`
|
||||
- `src/agents/voice/model.py`
|
||||
- `src/agents/voice/models/openai_stt.py`
|
||||
- `tests/voice/test_pipeline.py`
|
||||
- `tests/voice/test_input.py`
|
||||
- `tests/voice/test_openai_stt.py`
|
||||
- `tests/voice/test_openai_tts.py`
|
||||
@@ -19,9 +19,15 @@ Ensure work is only marked complete after formatting, linting, type checking, an
|
||||
6. If any command fails, fix the issue, rerun the script, and report the failing output.
|
||||
7. Confirm completion only when all commands succeed with no remaining issues.
|
||||
|
||||
## Environment setup
|
||||
|
||||
The verification scripts assume repository dependencies are already installed. Do not run `make sync` as part of every verification pass; use it for a fresh checkout, after dependency files change, or when dependency resolution fails before the checks start.
|
||||
|
||||
On Linux, some Python packages with native extensions may require system packages such as `libffi-dev`, Python development headers, or build tools. If verification cannot start because one of these packages is missing, treat it as a local environment setup issue. Install the missing dependency when possible, or report the failing command and missing dependency in the PR test plan before rerunning verification in a prepared environment.
|
||||
|
||||
## Manual workflow
|
||||
|
||||
- If dependencies are not installed or have changed, run `make sync` first to install dev requirements via `uv`.
|
||||
- For a fresh checkout, or if dependencies are not installed or have changed, run `make sync` first to install dev requirements via `uv`.
|
||||
- Run from the repository root with `make format` first, then `make lint`, `make typecheck`, and `make tests`.
|
||||
- Do not skip steps; stop and fix issues immediately when a command fails.
|
||||
- If you run the steps manually, you may parallelize `make lint`, `make typecheck`, and `make tests` after `make format` completes, but you must stop the remaining steps as soon as one fails.
|
||||
|
||||
@@ -286,13 +286,21 @@ check_for_missing_reporters() {
|
||||
log_file="${STEP_LOGS[$idx]}"
|
||||
start_time="${STEP_STARTS[$idx]}"
|
||||
now=$(date +%s)
|
||||
wait "${pid}" 2>/dev/null || true
|
||||
set +e
|
||||
wait "${pid}" 2>/dev/null
|
||||
step_status=$?
|
||||
set -e
|
||||
|
||||
if [ "${step_status}" -eq 0 ]; then
|
||||
finish_step "${name}" 0
|
||||
return 0
|
||||
fi
|
||||
|
||||
echo "code-change-verification: make ${name} exited before reporting completion status after $((now - start_time))s." >&2
|
||||
echo "--- ${name} log (last 80 lines) ---" >&2
|
||||
tail -n 80 "${log_file}" >&2 || true
|
||||
stop_running_steps
|
||||
return 1
|
||||
return "${step_status}"
|
||||
done
|
||||
|
||||
return 0
|
||||
|
||||
@@ -8,6 +8,8 @@ description: Run python examples in auto mode with logging, rerun helpers, and b
|
||||
## What it does
|
||||
|
||||
- Runs `uv run examples/run_examples.py` with:
|
||||
- Optional dependency extras enabled by default:
|
||||
`litellm`, `any-llm`, `sqlalchemy`, `redis`, `blaxel`, `modal`, `runloop`, and `temporal`.
|
||||
- `EXAMPLES_INTERACTIVE_MODE=auto` (auto-input/auto-approve).
|
||||
- Per-example logs under `.tmp/examples-start-logs/`.
|
||||
- Main summary log path passed via `--main-log` (also under `.tmp/examples-start-logs/`).
|
||||
@@ -44,6 +46,11 @@ description: Run python examples in auto mode with logging, rerun helpers, and b
|
||||
.agents/skills/examples-auto-run/scripts/run.sh rerun
|
||||
```
|
||||
|
||||
## Codex execution environment
|
||||
|
||||
- When Codex invokes `start` or `rerun`, run the command outside the Codex sandbox by default (`sandbox_permissions=require_escalated`). Many examples start nested sandboxes, browsers, npm helpers, or cloud/local service processes; running from inside the Codex sandbox can produce environment-only failures such as `sandbox-exec: sandbox_apply: Operation not permitted`, Playwright cache permission errors, or npm cache permission errors.
|
||||
- Use sandboxed execution only when the user explicitly asks for it or when running a narrow dry-run / log inspection command that does not execute examples.
|
||||
|
||||
## Defaults (overridable via env)
|
||||
|
||||
- `EXAMPLES_INTERACTIVE_MODE=auto`
|
||||
@@ -51,6 +58,7 @@ description: Run python examples in auto mode with logging, rerun helpers, and b
|
||||
- `EXAMPLES_INCLUDE_SERVER=0`
|
||||
- `EXAMPLES_INCLUDE_AUDIO=0`
|
||||
- `EXAMPLES_INCLUDE_EXTERNAL=0`
|
||||
- `EXAMPLES_UV_EXTRAS="litellm any-llm sqlalchemy redis blaxel modal runloop temporal"` (set to an empty string to disable extras)
|
||||
- Auto-approvals in auto mode: `APPLY_PATCH_AUTO_APPROVE=1`, `SHELL_AUTO_APPROVE=1`, `AUTO_APPROVE_MCP=1`
|
||||
|
||||
## Log locations
|
||||
@@ -62,7 +70,8 @@ description: Run python examples in auto mode with logging, rerun helpers, and b
|
||||
|
||||
## Notes
|
||||
|
||||
- The runner delegates to `uv run examples/run_examples.py`, which already writes per-example logs and supports `--collect`, `--rerun-file`, and `--print-auto-skip`.
|
||||
- The runner delegates to `uv run --extra ... examples/run_examples.py`, which already writes per-example logs and supports `--collect`, `--rerun-file`, and `--print-auto-skip`.
|
||||
- `examples/sandbox/extensions/vercel_runner.py` is temporarily excluded from auto runs due to credential issues. Do not force-run it until the credential setup is fixed.
|
||||
- `start` uses `--write-rerun` so failures are captured automatically.
|
||||
- If `.tmp/examples-rerun.txt` exists and is non-empty, invoking the skill with no args runs `rerun` by default.
|
||||
|
||||
|
||||
@@ -5,6 +5,23 @@ ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/../../../.." && pwd)"
|
||||
PID_FILE="$ROOT/.tmp/examples-auto-run.pid"
|
||||
LOG_DIR="$ROOT/.tmp/examples-start-logs"
|
||||
RERUN_FILE="$ROOT/.tmp/examples-rerun.txt"
|
||||
DEFAULT_UV_EXTRAS="litellm any-llm sqlalchemy redis blaxel modal runloop temporal"
|
||||
|
||||
build_uv_prefix() {
|
||||
UV_RUN=(uv run)
|
||||
local extras_value
|
||||
if [[ -n "${EXAMPLES_UV_EXTRAS+x}" ]]; then
|
||||
extras_value="$EXAMPLES_UV_EXTRAS"
|
||||
else
|
||||
extras_value="$DEFAULT_UV_EXTRAS"
|
||||
fi
|
||||
|
||||
local extra
|
||||
for extra in $extras_value; do
|
||||
UV_RUN+=(--extra "$extra")
|
||||
done
|
||||
export EXAMPLES_UV_EXTRAS="$extras_value"
|
||||
}
|
||||
|
||||
ensure_dirs() {
|
||||
mkdir -p "$LOG_DIR" "$ROOT/.tmp"
|
||||
@@ -28,8 +45,9 @@ cmd_start() {
|
||||
main_log="$LOG_DIR/main_${ts}.log"
|
||||
stdout_log="$LOG_DIR/stdout_${ts}.log"
|
||||
|
||||
build_uv_prefix
|
||||
local run_cmd=(
|
||||
uv run examples/run_examples.py
|
||||
"${UV_RUN[@]}" examples/run_examples.py
|
||||
--auto-mode
|
||||
--write-rerun
|
||||
--main-log "$main_log"
|
||||
@@ -152,7 +170,8 @@ collect_rerun() {
|
||||
exit 1
|
||||
fi
|
||||
cd "$ROOT"
|
||||
uv run examples/run_examples.py --collect "$log_file" --output "$RERUN_FILE"
|
||||
build_uv_prefix
|
||||
"${UV_RUN[@]}" examples/run_examples.py --collect "$log_file" --output "$RERUN_FILE"
|
||||
}
|
||||
|
||||
cmd_rerun() {
|
||||
@@ -171,8 +190,9 @@ cmd_rerun() {
|
||||
export APPLY_PATCH_AUTO_APPROVE="${APPLY_PATCH_AUTO_APPROVE:-1}"
|
||||
export SHELL_AUTO_APPROVE="${SHELL_AUTO_APPROVE:-1}"
|
||||
export AUTO_APPROVE_MCP="${AUTO_APPROVE_MCP:-1}"
|
||||
build_uv_prefix
|
||||
set +e
|
||||
uv run examples/run_examples.py --auto-mode --rerun-file "$file" --write-rerun --main-log "$main_log" --logs-dir "$LOG_DIR" 2>&1 | tee "$stdout_log"
|
||||
"${UV_RUN[@]}" examples/run_examples.py --auto-mode --rerun-file "$file" --write-rerun --main-log "$main_log" --logs-dir "$LOG_DIR" 2>&1 | tee "$stdout_log"
|
||||
local run_status=${PIPESTATUS[0]}
|
||||
set -e
|
||||
return "$run_status"
|
||||
@@ -194,6 +214,7 @@ Commands:
|
||||
Environment overrides:
|
||||
EXAMPLES_INTERACTIVE_MODE (default auto)
|
||||
EXAMPLES_INCLUDE_SERVER/INTERACTIVE/AUDIO/EXTERNAL (defaults: 0/1/0/0)
|
||||
EXAMPLES_UV_EXTRAS (default: litellm any-llm sqlalchemy redis blaxel modal runloop; set empty to disable)
|
||||
APPLY_PATCH_AUTO_APPROVE, SHELL_AUTO_APPROVE, AUTO_APPROVE_MCP (default 1 in auto mode)
|
||||
EOF
|
||||
}
|
||||
|
||||
@@ -117,8 +117,7 @@ https://github.com/openai/openai-agents-python/compare/<tag>...<target-commit>
|
||||
- <working tree status, tag/target assumptions, or re-run guidance>
|
||||
```
|
||||
|
||||
If no risks are found, include a “No material risks identified” line under Risk assessment and still provide a ship call. If you did not run local verification, do not add a verification status section or use it as a release blocker; note any assumptions briefly in Notes.
|
||||
If the report is not blocked, omit the `Unblock checklist` section.
|
||||
If no risks are found, include a "No material risks identified" line under Risk assessment and still provide a ship call. If you did not run local verification, do not add a verification status section or use it as a release blocker; note any assumptions briefly in Notes. If the report is not blocked, omit the `Unblock checklist` section.
|
||||
|
||||
### Resources
|
||||
|
||||
|
||||
@@ -38,6 +38,15 @@ Use this skill before editing code when the task changes runtime behavior or any
|
||||
- If review feedback claims a change is breaking, verify it against the latest release tag and actual external impact before accepting the feedback.
|
||||
- If a change truly crosses the latest released contract boundary, call that out explicitly in the ExecPlan, release notes context, and user-facing summary.
|
||||
|
||||
## SDK-specific decision rules
|
||||
|
||||
- When unsupported OpenAI API or provider-adapter behavior already has a released default path, avoid turning it into a default hard error unless the latest release boundary justifies that break. Prefer an opt-in strict mode such as `strict_feature_validation=True`, while keeping the default path compatible through warning, ignoring unsupported data, or a clearly non-empty placeholder.
|
||||
- For OpenAI API feature gaps, evaluate streaming and non-streaming paths together. Custom tool calls, multi-choice Chat Completions chunks, non-text tool outputs, and similar provider payload differences must not be strict in one path and permissive or malformed in the other.
|
||||
- When a change creates new public SDK behavior, do not expose it only through hard-coded module globals. Prefer an explicit public configuration object or parameter, preserve the existing default behavior when compatibility-sensitive, and make opt-in SDK defaults explicit.
|
||||
- Append new optional fields or constructor parameters to public dataclasses and constructors. Do not insert them before existing public fields unless you also provide a compatibility layer and regression coverage for the old positional call shape.
|
||||
- Treat threshold and quota values as part of the API design when they affect runtime behavior. Distinguish OpenAI platform quota-derived values from defensive SDK defaults; if the value is not anchored in a documented platform limit, avoid making it an unconditional default-on behavior.
|
||||
- Define `None` semantics deliberately for public configuration. For example, use separate meanings for "feature disabled or no SDK limit", "use SDK default limits", and "disable only this specific limit" rather than relying on implicit truthiness checks.
|
||||
|
||||
## When to stop and confirm
|
||||
|
||||
- The change would alter behavior shipped in the latest release tag.
|
||||
|
||||
@@ -0,0 +1,211 @@
|
||||
---
|
||||
name: maintainer-review
|
||||
description: Review a GitHub issue or pull request URL as an openai-agents-python maintainer, with a staged assessment of whether the claim is real, practically important, already solvable with supported functionality, correctly scoped, better served by another design, and worth maintainer and contributor effort. Use when assessing issue validity or severity, deciding whether an issue should be prioritized or closed, determining whether a requested feature represents an unmet need rather than a discoverability or usage gap, judging whether a PR is worth bringing to mergeable quality, comparing open PRs or alternative designs, separating code quality from repository readiness, or drafting a concise maintainer assessment. When closure, additional evidence, or code changes should be requested, also produce a polite, concise, complete, copy-paste-ready maintainer comment.
|
||||
---
|
||||
|
||||
# Maintainer Review
|
||||
|
||||
## Objective
|
||||
|
||||
Make a maintainer decision, not a generic code-review summary. Separate these questions:
|
||||
|
||||
1. Is the claimed behavior real?
|
||||
2. What user outcome or constraint exists independently of the reporter's proposed API or fix?
|
||||
3. Can supported functionality already achieve that outcome with reasonable composition or configuration?
|
||||
4. If a gap remains, is the proposed solution the best design and implementation layer?
|
||||
5. Can normal users plausibly reach the gap, and what happens when they do?
|
||||
6. Is it important enough to act on now?
|
||||
7. If this PR did not already exist, would maintainers choose to open and implement the same work?
|
||||
8. For a PR, is this solution worth merging and maintaining?
|
||||
9. Can overlapping or stale operations corrupt shared state or clean up resources owned by surviving work?
|
||||
10. If competing PRs exist, which single implementation path should maintainers pursue?
|
||||
11. Which ambiguous scope or semantic choices are maintainer-owned product/API decisions, and what concrete direction should the contributor implement?
|
||||
12. What concise maintainer message should communicate a closure or change request clearly and politely?
|
||||
|
||||
Treat an issue's requested field, callback, flag, class, or implementation strategy as a proposed mechanism, not as the accepted requirement. Do not begin by asking how to implement it. First prove that a concrete user outcome is not already supported and that the proposed mechanism is better than the available alternatives.
|
||||
|
||||
Lead with the current review state. Use `Preliminary assessment` while runtime approval or evidence is pending, and `Maintainer decision` only when the review can be concluded. Use the diff, issue narrative, or contributor effort as evidence, not as a proxy for impact.
|
||||
|
||||
## Workflow
|
||||
|
||||
### 1. Establish the exact target
|
||||
|
||||
- Accept a GitHub issue or PR URL as the primary input. Resolve its owner, repository, item type, and number before reviewing it.
|
||||
- For an issue, read the full report, comments, reproduction, environment, linked material, and maintainer responses.
|
||||
- For a PR, inspect the current remote base and head, full patch, commit history when relevant, tests, linked issue, and review discussion. Do not substitute the current local checkout for the remote change under review.
|
||||
- State the claim in one falsifiable sentence. Distinguish the reported symptom from the reporter's proposed cause or fix.
|
||||
- Identify the released behavior boundary when compatibility or regression claims matter.
|
||||
- Verify whether linked evidence matches the PR's exact runtime variant, provider or tool type, triggering condition, and user outcome. A generic issue title, conceptual similarity, or wording such as `Related to` does not transfer evidence of need to an adjacent extension. If the reported scenario has already been fixed, treat additional variants as new needs requiring their own evidence.
|
||||
|
||||
Respect repository instructions for remote access and mutation. A review does not authorize comments, labels, branch changes, pushes, or other remote writes.
|
||||
|
||||
### 2. Establish the unmet need and challenge the proposed solution
|
||||
|
||||
Complete this pass before deeply evaluating a proposed implementation and before any positive issue or PR assessment.
|
||||
|
||||
First assign one `Need evidence` status:
|
||||
|
||||
- **Demonstrated**: The exact scope has a concrete supported scenario, a real-path reproduction, a released compatibility requirement, repeated demand, or a broad invariant with a meaningful consequence.
|
||||
- **Plausible but unproven**: The path can exist, but realistic provider behavior, user reach, frequency, consequence, or demand is not established.
|
||||
- **Already covered**: A reasonable supported workflow already satisfies the outcome.
|
||||
- **Unsupported**: The outcome belongs outside the SDK contract or at a provider, adapter, or caller-owned layer.
|
||||
|
||||
Only `Demonstrated` need may receive `Merge-worthy as-is` or `Merge-worthy after focused changes`. For `Plausible but unproven`, prefer `Needs evidence` or `Not worth completing`; for `Already covered` or `Unsupported`, prefer closure or the relevant simpler alternative.
|
||||
|
||||
1. Restate the desired user outcome without naming the requested API, class, file, option, or implementation. Separate the actual constraint from the reporter's preferred mechanism.
|
||||
2. Trace the closest supported ways to achieve that outcome in the current release and current target. Inspect the owning code path, public API, tests, and relevant docs rather than assuming that an unfamiliar capability is missing. Consider configuration, composition, cloning, callbacks, extension points, provider adapters, and doing the work at a caller-owned layer.
|
||||
3. Determine whether the report shows a capability gap, an ergonomics or discoverability problem, an unsupported use case, or no demonstrated problem. A more convenient spelling is not automatically a missing capability.
|
||||
4. Compare the proposed solution against the strongest existing approach and at least one better-design candidate: no code change, clearer documentation or validation, a narrower fix, reuse of an existing abstraction, or enforcement at a more coherent shared boundary.
|
||||
5. For each viable approach, compare whether it satisfies the concrete scenario, what new public or internal contract it creates, cross-path consistency, compatibility, and permanent maintenance cost.
|
||||
|
||||
Do not treat a test proving that new code can work as evidence that the feature is needed. A `FakeModel` response, manually constructed provider item, mock, or new regression test can establish code-path reachability and implementation correctness; it does not by itself establish realistic provider behavior, user reach, frequency, practical consequence, or demand.
|
||||
|
||||
API symmetry, naming consistency, and parity with an adjacent tool, provider, or output type are design arguments, not evidence of need. Parity may justify work when it removes existing complexity or enforces a broad demonstrated invariant, but adding branches, tests, documentation, or public behavior requires independent practical justification.
|
||||
|
||||
If the need is not `Demonstrated`, inspect the patch only far enough to understand its contract, risk, and maintenance cost. Do not turn implementation defects, missing tests, or documentation gaps into a request-changes recommendation, because those questions become merge-blocking only after the need gate passes. If the report provides no concrete scenario, the existing functionality appears sufficient, or the requested mechanism solves only a hypothetical convenience problem, prefer `Needs evidence`, `Close`, `Supersede with a simpler alternative`, or `Not worth completing` over designing the requested feature on the reporter's behalf.
|
||||
|
||||
### 3. Discover competing open PRs proportionally
|
||||
|
||||
Do this before deeply evaluating a specified PR. A PR URL selects the starting point, not necessarily the entire comparison set.
|
||||
|
||||
- Determine the primary issue from explicit closing keywords, linked issues, issue timeline or development links, PR body and comments, and the reproduced symptom. If the association is inferred rather than explicit, state the evidence.
|
||||
- When an issue is explicitly linked, enumerate all open PRs that address it through the issue timeline, development links, cross-references, closing keywords, and ordinary references. Include draft PRs but label them as drafts.
|
||||
- When no issue is linked, run a bounded duplicate search using the strongest two or three signals from the title, reproduction, violated invariant, and runtime path. Stop when additional queries are unlikely to produce a credible competing implementation.
|
||||
- Exclude closed or merged PRs from the active comparison set, while using them as history when relevant.
|
||||
- Do not group PRs merely because they mention the same subsystem. Require a shared issue, symptom, violated invariant, or materially overlapping fix.
|
||||
- Record the search methods and candidate set internally. If repository access cannot establish completeness, say so instead of claiming that every open PR was found. Do not list unrelated search hits in the final report.
|
||||
|
||||
When multiple candidates exist, compare them on need coverage, runtime correctness, scope, implementation layer, tests, compatibility, complexity, readiness, remaining maintainer work, and whether useful parts can be combined. Prefer the best maintainable solution, not the first submission or the smallest diff by default.
|
||||
|
||||
### 4. Use a two-stage evidence flow
|
||||
|
||||
Always begin with a desk review. Inspect the concrete runtime path before judging a small change as either trivial or meaningful. Check callers, adjacent helpers, validation layers, fallback paths, and existing tests. Search history or documentation only when it changes the decision. Inspecting test code is part of the desk review; executing tests, imports, examples, reproductions, benchmarks, or service calls is a runtime probe.
|
||||
|
||||
For repository-specific runtime invariants, start with `.agents/references/README.md` and open only the references that match the affected boundary. Treat `.agents/references/` as read-only during issue and PR review: use it to identify expected invariants, adjacent surfaces, and regression risks, then verify the current claim against the remote change, current code, tests, docs, release boundary, and focused runtime evidence. Do not edit references as a side effect of the review, infer current issue or PR status from them, or treat old issue or PR outcomes as current evidence. If the review reveals a reusable invariant that should be captured, recommend a separate repository-maintenance update unless the user explicitly asks to update references in the same task.
|
||||
|
||||
Use this evidence order across the two stages:
|
||||
|
||||
1. Trace the closest existing supported capabilities and determine whether they already satisfy the underlying user outcome.
|
||||
2. Inspect existing tests and complete the code-path trace, including the mandatory interleaving and ownership pass when triggered, without executing code.
|
||||
3. With explicit user approval, run a focused local reproduction of the exact claim when the desk-review rules below require it.
|
||||
4. A comparison with the released version, base branch, or known-good control.
|
||||
5. A broader runtime matrix only when the maintainer decision remains uncertain and the user approves it.
|
||||
|
||||
#### Stage 1: desk review
|
||||
|
||||
Produce an initial result from static evidence before running code:
|
||||
|
||||
##### Mandatory unmet-need and design pass
|
||||
|
||||
Before a positive assessment, complete the pass in step 2 and be able to state all of the following from concrete evidence:
|
||||
|
||||
1. The user outcome that current supported behavior cannot achieve.
|
||||
2. The closest existing API or composition path and the exact reason it is insufficient.
|
||||
3. Why the proposed behavior belongs at the chosen abstraction layer instead of a caller, adapter, validation, documentation, or existing extension point.
|
||||
4. Why the proposed permanent contract is better than no code change and the strongest narrower alternative.
|
||||
5. What real scenario, compatibility requirement, or repeated demand justifies the new maintenance surface.
|
||||
6. Whether maintainers would choose to pursue the same work if no contributor had already supplied a patch.
|
||||
|
||||
If any answer is missing and could change whether code should exist at all, do not call the issue actionable or the PR merge-worthy. Request only the evidence needed to distinguish a genuine capability gap from a usage, discoverability, or solution-design problem. This is a product and architecture evidence gap, not a runtime-probe trigger by itself.
|
||||
|
||||
##### Mandatory interleaving and ownership pass
|
||||
|
||||
Run this pass before any positive PR assessment when a patch adds, removes, or reorders cleanup, retry, reconnect, cancellation, listeners, shared futures or tasks, connections or streams, state flags, or mutable state across an `await`, callback, event, or deferred completion.
|
||||
|
||||
1. Name each shared resource or state value and the operation that owns it. Include listeners, futures, tasks, connections, streams, locks, caches, state flags, persistence, and telemetry.
|
||||
2. Trace at least two overlapping operations, `A` and `B`, across every suspension or re-entry point. Check `A pending -> B starts -> A fails -> B succeeds`, `A pending -> B starts -> B fails -> A succeeds`, close or cancellation between setup and completion, and a stale completion arriving after newer work.
|
||||
3. For every cleanup or rollback, identify the exact attempt and resource generation it is allowed to dispose. Treat unconditional cleanup after a suspension point as a regression candidate until the code proves it cannot tear down newer or surviving work.
|
||||
4. Compare base and head for the survivor invariant. Replacing duplicated work with missing handlers, a closed shared resource, reverted state, or a failed surviving task is a regression, not successful cleanup.
|
||||
5. Inspect tests for controlled interleavings using deferred futures, callbacks, or events. Require assertions about the surviving operation's observable behavior and final resource state, not only listener counts or individual exception results.
|
||||
|
||||
Do not mark a concurrency-sensitive patch `Merge-worthy as-is` merely because sequential reconnect, retry, failure, and close tests pass. If the code trace proves an unsafe interleaving, conclude from static evidence and request a focused fix and regression test. If ownership remains ambiguous, keep the result preliminary and request approval for the smallest decisive runtime probe.
|
||||
|
||||
- If the claim or PR is decisively negative from a complete reachable code-path trace, conclude the review without a runtime probe. Examples include an impossible or unsupported path, duplicated existing handling, a demonstrated no-op, a direct compatibility break, or a clearly wrong abstraction. Do not call an ambiguous result negative merely to avoid a probe.
|
||||
- If the initial result is positive and there is no unresolved runtime concern, and any triggered interleaving and ownership pass is complete, the desk review may be sufficient for a final maintainer decision. Do not run a probe only to restate evidence that cannot plausibly change the decision.
|
||||
- If the initial result is positive but there is any unresolved runtime concern that could plausibly change claim validity, severity, merge-worthiness, required changes, or the preferred competing PR, stop before executing code. Report a `Preliminary assessment`, name the concern, propose the smallest decisive probe and control, and ask the user for approval to run it.
|
||||
- A purely stylistic, documentation, CI-status, or repository-readiness concern does not trigger a runtime probe unless it masks a runtime question.
|
||||
|
||||
Do not issue a definitive positive maintainer decision while a decision-relevant runtime concern remains unresolved. If the user declines the probe, keep the result preliminary and state the exact confidence limitation.
|
||||
|
||||
#### Stage 2: approved runtime probe
|
||||
|
||||
After explicit approval, run only the smallest probe needed to resolve the stated concern. Exercise the real public or internal path and include a base, release, or known-good control when relevant. Do not stop at a happy-path smoke check when failure behavior determines the decision. Return to the user for separate approval before expanding materially beyond the approved probe.
|
||||
|
||||
For latency, timeout, buffering, backpressure, or cleanup claims, measure at least one observable elapsed-time or state-transition path when feasible. Do not assume that a mocked unit test exercises real scheduling or provider behavior. Prefer a local probe first; use an approval-gated live-service probe only when local evidence cannot settle the decision.
|
||||
|
||||
Use `$runtime-behavior-probe` only when the user explicitly invokes it and the skill is available, or when the user explicitly approves using it for the proposed runtime work. Preserve its environment-variable approval, live-service, cost, cleanup, and reporting gates. Do not make ordinary maintainer review depend on that skill being available.
|
||||
|
||||
For changes involving validation, fail-fast behavior, cleanup, retries, interruption, or concurrency, trace lifecycle ordering in addition to the main behavior:
|
||||
|
||||
- Identify listeners, tasks, connections, files, locks, state mutations, and other resources acquired before the new check or failure point.
|
||||
- Verify cleanup when construction, context-manager entry, validation, connection, or execution raises before normal teardown runs.
|
||||
- Require a negative-path test when a failure can leave observable state or resources behind.
|
||||
|
||||
Do not over-investigate. Stop when additional evidence is unlikely to change validity, severity, or the maintainer recommendation.
|
||||
|
||||
### 5. Calibrate validity and impact
|
||||
|
||||
Use `references/evaluation-framework.md` to assess claim validity, realistic reach, consequence, breadth, frequency, recoverability, compatibility, and severity. Keep observed facts separate from inference and state any missing evidence that could change the decision.
|
||||
|
||||
Report the `Need evidence` status before classifying the need as a capability gap, ergonomics or discoverability gap, unsupported use case, or no demonstrated gap. Do not assign practical impact to the absence of the requested mechanism when an existing supported workflow already produces the requested outcome. Do not infer practical importance merely from reachability, API asymmetry, or a technically successful patch.
|
||||
|
||||
For a PR, make `Severity` describe the underlying issue or user need only. Do not combine it with the risk created by the proposed patch. Report a meaningful patch-induced regression, compatibility, lifecycle, or maintenance risk separately as `Patch risk`.
|
||||
|
||||
Do not infer that a report is low-value merely because an AI may have found or written it. Do not speculate about authorship or motive. Identify contribution-shaped reports through objective signals: no reproducible behavior, unrealistic inputs, an impossible call path, duplicated existing handling, tests that do not exercise the claim, or a fix whose runtime result is a no-op.
|
||||
|
||||
### 6. Apply the maintainer-effort test
|
||||
|
||||
Use the framework's issue dispositions and PR checks to decide whether the outcome justifies permanent code, tests, documentation, and maintainer attention. Classify code quality separately from repository readiness.
|
||||
|
||||
Use one code recommendation:
|
||||
|
||||
- **Merge-worthy as-is**: real need, sound implementation, proportionate scope, adequate tests.
|
||||
- **Merge-worthy after focused changes**: real need and viable direction, with bounded corrections.
|
||||
- **Supersede with a simpler alternative**: real need, but a smaller or more coherent fix is preferable.
|
||||
- **Not worth completing**: negligible or unsupported impact, no-op behavior, wrong abstraction, or excessive completion cost.
|
||||
|
||||
`Merge-worthy as-is` and `Merge-worthy after focused changes` are invalid unless `Need evidence` is `Demonstrated`. A bounded set of implementation fixes cannot promote a `Plausible but unproven` need into a merge-worthy recommendation.
|
||||
|
||||
For `Merge-worthy as-is` and `Merge-worthy after focused changes`, use one repository-readiness status when it helps communicate the integration state:
|
||||
|
||||
- **Ready**: current head is reviewable and required checks are green.
|
||||
- **CI or review pending**: code recommendation is stable, but required external gates are incomplete.
|
||||
- **Rebase or conflict resolution required**: the head cannot merge cleanly or is materially stale.
|
||||
- **Blocked**: a concrete external or repository condition prevents a reliable merge decision.
|
||||
|
||||
Omit repository readiness for `Supersede with a simpler alternative` and `Not worth completing`; CI, review, mergeability, or branch freshness does not change those dispositions. Put any validation limitation that materially affects confidence in the evidence instead. When readiness is included, use exactly one of the four statuses above and do not invent variants such as `ready mechanically` or use rebase status for semantic staleness.
|
||||
|
||||
Do not downgrade an otherwise sound code recommendation solely because CI is pending. Do not call a PR ready when semantic conflict resolution or material code changes remain.
|
||||
|
||||
When multiple open PRs address the same issue, make one portfolio-level recommendation: select the strongest PR, request focused changes in one candidate, combine specific ideas into one PR, supersede all candidates with a simpler approach, or close duplicates. Explain why the recommended path is better than each alternative without turning the report into line-by-line review.
|
||||
|
||||
Always compare the proposed patch with the strongest existing supported approach and at least one alternative: no code change, validation or documentation, a narrower fix, reuse of an existing helper, or a different layer that enforces the invariant consistently. A review is incomplete if it establishes only that the patch works without establishing why the current product cannot meet the underlying need and why this design is preferable.
|
||||
|
||||
When multiple plausible semantic scopes, compatibility boundaries, or public API contracts remain, do not ask the contributor to choose among maintainer-owned options. Decide the preferred scope from the evidence, compatibility contract, and product/API design principles, then request that specific change. If the evidence is insufficient to choose, mark the review preliminary or request maintainer input; do not present an open-ended implementation fork as the contributor's decision.
|
||||
|
||||
### 7. Report findings and maintainer action
|
||||
|
||||
Choose the assessment language using this precedence:
|
||||
|
||||
1. Follow an explicit language request in the current conversation.
|
||||
2. Follow an applicable language instruction from `~/.codex/AGENTS.md`, the repository's `AGENTS.md`, or another governing instruction file.
|
||||
3. If recent conversation turns are consistently in one language, use that language.
|
||||
4. Otherwise, default to English.
|
||||
|
||||
Do not infer the assessment language from the GitHub URL, contributor, code, or browser locale. Maintainer comment drafts remain English regardless of the assessment language. Keep the report decision-oriented and compact. Use no more than five evidence bullets by default; add more only when the decision genuinely depends on them.
|
||||
|
||||
Use the matching compact report variant in `references/evaluation-framework.md`. While runtime approval is pending, use its preliminary-assessment variant and end with the approval request instead of presenting a final recommendation. Collapse sections for simple cases rather than padding the answer. Put unexpected or negative runtime findings first, and name the preferred PR or approach explicitly when candidates compete.
|
||||
|
||||
For PRs, put `Need evidence` before code recommendation. When the need is not `Demonstrated`, lead with that result, omit repository readiness, and avoid presenting patch fixes as the primary maintainer action.
|
||||
|
||||
When existing functionality or a better alternative materially affects the decision, state it explicitly in the evidence and recommendation. Name the exact supported path, what it does and does not cover, and why it is preferable. Do not bury a `Not worth completing` or `Supersede with a simpler alternative` conclusion beneath praise for implementation quality.
|
||||
|
||||
When recommending closure, requesting more evidence, requesting code changes, or superseding a PR, append the English, copy-paste-ready maintainer comment defined by the framework. If multiple PRs need different actions, label one draft for each affected PR. Include only merge-blocking requests in the main action paragraph; keep optional documentation or polish clearly non-blocking or omit it.
|
||||
|
||||
For request-changes comments, phrase maintainer-owned semantic decisions as a directive, not as a menu. It is fine to mention the rejected alternative briefly in the rationale, but the requested action must identify the chosen behavior, scope, or compatibility boundary. Use "please do X because..." instead of "either do X or Y" when X versus Y changes the SDK contract or user-visible semantics.
|
||||
|
||||
Do not produce a line-by-line review unless requested. Do not equate passing tests with merge-worthiness, or a logically correct patch with practical value.
|
||||
|
||||
## Resource
|
||||
|
||||
- `references/evaluation-framework.md` contains the severity rubric, evidence checks, lifecycle review, issue dispositions, PR quality checks, maintainer-comment guidance, and report variants.
|
||||
@@ -0,0 +1,4 @@
|
||||
interface:
|
||||
display_name: "Maintainer Review"
|
||||
short_description: "Gate PR value on demonstrated user need"
|
||||
default_prompt: "Use $maintainer-review with this GitHub issue or PR URL. Before evaluating implementation quality, verify that linked evidence matches the exact runtime variant and assign Need evidence as Demonstrated, Plausible but unproven, Already covered, or Unsupported. Only a Demonstrated need may receive a merge-worthy recommendation; synthetic tests, API parity, and contributor effort do not establish need. Then compare existing and alternative approaches, complete the desk review and required lifecycle ownership checks, request approval before any decision-relevant runtime probe, compare credible competing PRs, recommend the best maintainer action, and include an English comment draft when closure or changes are needed."
|
||||
@@ -0,0 +1,366 @@
|
||||
# Maintainer Evaluation Framework
|
||||
|
||||
Use this reference when a claim is ambiguous, severity is disputed, or a PR is technically correct but may not justify merge effort.
|
||||
|
||||
## Contents
|
||||
|
||||
- [Decision model](#decision-model)
|
||||
- [Severity rubric](#severity-rubric)
|
||||
- [Evidence-strength checks](#evidence-strength-checks)
|
||||
- [Unmet need and alternative design gate](#unmet-need-and-alternative-design-gate)
|
||||
- [Issue disposition](#issue-disposition)
|
||||
- [PR quality and value](#pr-quality-and-value)
|
||||
- [Documentation threshold](#documentation-threshold)
|
||||
- [Lifecycle and failure-path review](#lifecycle-and-failure-path-review)
|
||||
- [Concurrency and cleanup ownership](#concurrency-and-cleanup-ownership)
|
||||
- [Better-alternative prompts](#better-alternative-prompts)
|
||||
- [Competing PR comparison](#competing-pr-comparison)
|
||||
- [Maintainer comment drafts](#maintainer-comment-drafts)
|
||||
- [Compact report variants](#compact-report-variants)
|
||||
|
||||
## Decision model
|
||||
|
||||
Treat validity, severity, and merge-worthiness as separate results. Also distinguish a `Preliminary assessment`, which may still require approved runtime evidence, from a final `Maintainer decision`. Do not label a provisional positive result as a verdict or final decision.
|
||||
|
||||
| Dimension | Questions | Strong evidence |
|
||||
|---|---|---|
|
||||
| Claim validity | Does the exact reported behavior occur? Is the proposed cause correct? | Reproduction, failing focused test, or complete reachable code path |
|
||||
| Reachability | Can supported, realistic inputs reach it? | Public API trace, real configuration, linked user report, or release comparison |
|
||||
| Consequence | What fails, and is the result silent or recoverable? | Observed output/error/state plus downstream effect |
|
||||
| Breadth | Who is affected? | Supported providers, platforms, versions, and configurations identified precisely |
|
||||
| Frequency | Is this normal, intermittent, or pathological? | Repeat runs, telemetry or reports when available, deterministic preconditions |
|
||||
| Need evidence | Is the exact scope demonstrated, merely plausible, already covered, or unsupported? | Same-scope user scenario, real-path reproduction, released compatibility requirement, repeated demand, or broad consequential invariant |
|
||||
| Unmet need | What user outcome cannot be achieved through supported behavior today? | Concrete scenario plus a trace showing why the closest existing path is insufficient |
|
||||
| Existing capability | Can configuration, composition, cloning, callbacks, extension points, or a caller-owned layer already satisfy the outcome? | Current release code, tests, docs, and an exact supported workflow |
|
||||
| Compatibility | Is released behavior or durable state changed? | Latest release comparison and explicit contract inspection |
|
||||
| Solution fit | Is the requested mechanism the best design and implementation layer? | Proposed solution compared with the strongest existing path and at least one narrower or more coherent alternative |
|
||||
| Maintainer-owned scope | When several plausible semantics remain, which behavior should the SDK own? | A concrete maintainer decision grounded in compatibility, user outcome, and API design, not an open-ended contributor choice |
|
||||
| Resource ownership | Can stale, failed, cancelled, or overlapping work mutate or clean up resources owned by surviving work? | Interleaving trace, attempt or generation ownership, and survivor assertions |
|
||||
| Maintenance cost | What permanent complexity and review burden does it add? | Changed surface, new branches/configuration, test burden, remaining work |
|
||||
|
||||
## Severity rubric
|
||||
|
||||
- **Negligible**: No runtime difference, unreachable or unsupported input, cosmetic inconsistency, or a fully harmless edge case. Usually close, document, or decline code complexity.
|
||||
- **Low**: Real but narrow and recoverable behavior with a simple workaround and no data, security, or compatibility risk. Merge only when the fix is small and clearly improves an invariant.
|
||||
- **Moderate**: Plausible supported use fails or produces incorrect behavior for a meaningful subset of users. Prioritize a bounded fix and regression test.
|
||||
- **High**: Common or important supported use is broken, causes serious compatibility problems, leaks sensitive data, or risks persistent corruption. Treat as urgent and require strong validation.
|
||||
- **Critical**: Broadly exploitable security impact, severe data loss, or systemic failure requiring immediate coordinated action. Use only with concrete evidence.
|
||||
|
||||
Severity is approximately consequence multiplied by realistic reach and frequency, reduced by recoverability. Do not raise severity because a report sounds alarming or lower it because a patch is small.
|
||||
|
||||
## Evidence-strength checks
|
||||
|
||||
Before calling a claim confirmed, answer:
|
||||
|
||||
- Does the reproduction exercise the same public or internal path named in the report?
|
||||
- Does the failure still occur on the relevant base, release, or current target?
|
||||
- Does the test fail without the patch and pass with it?
|
||||
- Are setup failures, stale builds, environment leakage, proxies, caches, or unsupported options excluded?
|
||||
- Does an adjacent helper or equivalent path follow different semantics?
|
||||
- Is the observed behavior prohibited by an actual contract, or merely surprising?
|
||||
- For latency, timeout, buffering, backpressure, or cleanup claims, was observable elapsed time or a real state transition measured when feasible rather than inferred only from mocks?
|
||||
- For shared asynchronous state, do tests control completion order and prove that stale failure or cleanup cannot affect the surviving operation?
|
||||
|
||||
Use `partially confirmed` when the symptom is real but the cause, reach, or claimed scope is wrong. Use `unproven` when decisive evidence is missing. Use `contradicted` only when evidence directly disproves the claim.
|
||||
|
||||
## Unmet need and alternative design gate
|
||||
|
||||
Issue reports often combine a desired outcome with a proposed API or implementation. Treat the proposed mechanism as a hypothesis. Confirm the unmet outcome before evaluating how well the patch implements that mechanism.
|
||||
|
||||
### Linked-evidence scope
|
||||
|
||||
Evidence from a linked issue applies only when the issue and PR share the same runtime variant, provider or tool type, trigger, supported configuration, and user outcome. A broad title, ordinary reference, `Related to` statement, or conceptual similarity is not enough. If an earlier change already resolved the concrete reported scenario, an adjacent extension starts with no inherited evidence of need.
|
||||
|
||||
### Need evidence status
|
||||
|
||||
Assign one status before deep implementation review:
|
||||
|
||||
- **Demonstrated**: The exact scope has a concrete supported scenario, real-path reproduction, released compatibility requirement, repeated demand, or broad invariant with meaningful consequence.
|
||||
- **Plausible but unproven**: The code path is possible, but realistic reach, frequency, consequence, provider behavior, or demand is missing.
|
||||
- **Already covered**: A reasonable supported workflow already satisfies the outcome.
|
||||
- **Unsupported**: The outcome is outside the SDK contract or belongs at a provider, adapter, or caller-owned layer.
|
||||
|
||||
Only `Demonstrated` need can support a merge-worthy code recommendation. `Plausible but unproven` maps to `Needs evidence` or `Not worth completing`, even when the patch is technically correct and its remaining fixes are bounded. `Already covered` and `Unsupported` normally map to closure or a simpler non-core alternative.
|
||||
|
||||
Before accepting an issue or recommending a PR, record:
|
||||
|
||||
| Question | Required evidence |
|
||||
|---|---|
|
||||
| What outcome is needed? | A concrete supported scenario stated without the proposed API or fix |
|
||||
| What exists today? | The closest current-release API, configuration, composition, extension point, or caller-owned solution |
|
||||
| Why is it insufficient? | An exact behavioral, compatibility, lifecycle, or operational constraint, not preference alone |
|
||||
| What are the alternatives? | The proposed patch, the strongest existing path, and at least one no-code, narrower, or better-layer design |
|
||||
| Why add a contract? | Practical benefit sufficient to justify public surface, runtime branches, cross-path tests, documentation, and long-term maintenance |
|
||||
|
||||
Classify the result:
|
||||
|
||||
- **Capability gap**: a supported, realistic outcome cannot be achieved with current functionality. Code may be warranted.
|
||||
- **Ergonomics or discoverability gap**: the outcome is already possible, but the supported route is confusing or unnecessarily difficult. Prefer documentation, validation, or a narrowly justified convenience improvement.
|
||||
- **Unsupported use case**: the desired outcome lies outside the SDK contract or belongs at a provider, adapter, application, or other caller-owned layer. Do not expand the core API merely to make it possible.
|
||||
- **No demonstrated gap**: no concrete scenario proves that existing functionality is insufficient. Request evidence or close rather than designing from the proposed mechanism.
|
||||
|
||||
Passing tests for a new implementation establish feasibility and correctness, not need. A `FakeModel` response, manually constructed provider item, mock, or synthetic fixture does not establish realistic provider behavior, user reach, frequency, consequence, or demand. API symmetry and parity with an adjacent runtime are design arguments, not need evidence. A technically coherent patch can still be `Not worth completing` when the motivating scenario is hypothetical, already supported, or better solved elsewhere.
|
||||
|
||||
Use the counterfactual maintainer test: if the PR did not already exist, would maintainers choose to file and implement the same work from the available evidence? Contributor effort lowers implementation cost but does not create product need or remove permanent maintenance cost.
|
||||
|
||||
When the need is not `Demonstrated`, inspect implementation only far enough to estimate contract, risk, and maintenance cost. Do not convert patch defects, missing tests, or documentation gaps into a request-changes disposition; those become merge blockers only after the need gate passes.
|
||||
|
||||
## Issue disposition
|
||||
|
||||
Choose one primary action:
|
||||
|
||||
- **Prioritize**: confirmed moderate-or-higher impact or an important invariant with no safe workaround.
|
||||
- **Accept, low priority**: confirmed low impact, existing supported functionality is insufficient for the demonstrated scenario, and a proportionate fix appears possible.
|
||||
- **Narrow scope**: a valid core exists, but the report overstates affected paths or expected behavior.
|
||||
- **Needs evidence**: plausible claim, but no minimal reproduction, supported setup, contract basis, or concrete scenario showing why existing functionality is insufficient.
|
||||
- **Close**: duplicate, unsupported, unreachable, contradicted, no-op, already addressed by a reasonable supported path, or not worth permanent complexity.
|
||||
|
||||
When requesting evidence, ask only for information that could change the disposition.
|
||||
|
||||
## PR quality and value
|
||||
|
||||
Assess these independently:
|
||||
|
||||
1. **Need**: Same-scope issue or runtime evidence demonstrates a concrete unmet user outcome that the closest supported capability cannot reasonably satisfy. Do not inherit evidence from an adjacent variant or already-fixed scenario.
|
||||
2. **Correctness**: The fix works for the reported case and meaningful boundaries.
|
||||
3. **Placement**: The invariant is enforced once at the right layer instead of duplicating existing functionality, patching locally, or moving caller- or provider-owned policy into the core SDK.
|
||||
4. **Consistency**: Equivalent sync/async, streaming/non-streaming, provider, serialization, and resume paths remain aligned where applicable.
|
||||
5. **Tests**: A regression test fails on the base, passes on the head, and tests the exact non-happy-path value or state. When shared state crosses an asynchronous boundary, tests control relevant completion orders and assert the surviving operation's behavior and final resource state.
|
||||
6. **Compatibility**: Released positional APIs, wire formats, persisted schemas, and established error behavior are preserved or intentionally migrated.
|
||||
7. **Proportionality**: Complexity and public surface are justified by impact.
|
||||
8. **Completion cost**: Remaining fixes, docs, tests, and design work are bounded enough to justify maintainer attention.
|
||||
|
||||
A PR can be correct but not merge-worthy. Typical reasons include a nonexistent or negligible need, an outcome already supported through a reasonable existing mechanism, a no-op on the actual runtime path, incomplete cross-path semantics, an abstraction cost larger than the benefit, or a simpler design at another layer.
|
||||
|
||||
Do not use implementation correctness, bounded remaining work, CI status, or contributor effort to upgrade a need that is only `Plausible but unproven`. Merge-worthiness is gated by demonstrated need, not by how close the patch is to completion.
|
||||
|
||||
Keep issue impact and patch risk separate. `Severity` describes the underlying issue or user need. A regression, compatibility break, lifecycle leak, or maintenance hazard introduced by the proposed patch belongs under `Patch risk` and must not inflate or obscure the issue severity.
|
||||
|
||||
When a PR exposes an ambiguous semantic boundary, decide whether that boundary belongs to maintainers before drafting requests. If the choice affects SDK contract, compatibility, persistence, error semantics, public API meaning, or cross-path behavior, the review should pick one direction or explicitly block on maintainer input. Do not delegate that choice to the contributor as "either X or Y"; ask for the chosen behavior and the tests or docs needed to lock it down.
|
||||
|
||||
## Documentation threshold
|
||||
|
||||
Do not treat documentation as automatically required for every public option, constructor parameter, provider setting, or behavior change. Make docs merge-blocking only when at least one of these is true:
|
||||
|
||||
- Existing user-facing docs become materially false, unsafe, or misleading.
|
||||
- Correct or safe use depends on a non-obvious constraint, migration step, compatibility boundary, or operational warning.
|
||||
- Repository policy, the accepted issue scope, or an explicit maintainer decision requires documentation in the same PR.
|
||||
- The intended feature would be practically unusable or undiscoverable by its target users without a documented entry point, and generated API reference or clear code-level discovery is insufficient.
|
||||
|
||||
If docs would merely improve discoverability or completeness, keep them non-blocking. Do not change `Merge-worthy as-is` to `Merge-worthy after focused changes` solely for optional docs, and do not include optional docs in the maintainer comment's required-action paragraph. Respect an explicit maintainer choice to omit docs or defer them to a separate follow-up.
|
||||
|
||||
## Lifecycle and failure-path review
|
||||
|
||||
Apply this section when a change adds validation, fail-fast behavior, cleanup, retries, interruption, background work, or concurrency.
|
||||
|
||||
- Identify the earliest point where all dynamic inputs needed for a correct decision are available.
|
||||
- List side effects before and after that point: listeners, tasks, connections, files, locks, caches, state mutations, and telemetry.
|
||||
- Exercise failure during construction, context-manager entry, validation, connection, and execution when those phases exist.
|
||||
- Confirm that normal teardown is actually entered. If an enter or constructor fails, verify cleanup explicitly rather than assuming an exit hook runs.
|
||||
- Prefer validation after dynamic configuration is resolved but before avoidable side effects begin.
|
||||
- Require a regression test for any listener, task, connection, or state that could remain after failure.
|
||||
|
||||
## Concurrency and cleanup ownership
|
||||
|
||||
Apply this section before a positive assessment whenever lifecycle work crosses an `await`, callback, event, deferred completion, retry, reconnect, cancellation, or shared resource boundary. Sequential correctness is insufficient because a patch can improve isolated cleanup while introducing cross-attempt teardown.
|
||||
|
||||
Use a two-operation interleaving matrix during desk review:
|
||||
|
||||
| Ordering | Required question |
|
||||
|---|---|
|
||||
| `A pending -> B starts -> A fails -> B succeeds` | Can A's cleanup remove or revert anything B needs? |
|
||||
| `A pending -> B starts -> B fails -> A succeeds` | Can B's cleanup leave A successful but non-functional? |
|
||||
| `A succeeds -> B starts -> stale A completion` | Can stale A overwrite B's newer state or generation? |
|
||||
| setup -> close/cancel -> late completion | Can late work resurrect listeners, state, tasks, or connections after teardown? |
|
||||
|
||||
For each ordering:
|
||||
|
||||
- Identify the resource owner before and after every suspension point.
|
||||
- Distinguish per-attempt resources from shared runner, session, transport, cache, or listener state.
|
||||
- Require cleanup to carry an ownership token, generation, identity check, serialization guarantee, or another invariant that prevents cross-attempt disposal.
|
||||
- Compare base and head on the survivor invariant. Fewer duplicates do not justify losing the only active handler, connection, task, or state update.
|
||||
- Require a controlled interleaving test when the ordering is reachable. The test must assert both the failing operation and the surviving operation's observable behavior after all completions settle.
|
||||
|
||||
An unscoped `finally`, `except`, close handler, cancellation callback, or rollback that mutates shared state after a suspension point is merge-blocking when another operation can still own or use that state.
|
||||
|
||||
## Better-alternative prompts
|
||||
|
||||
Start with the strongest existing supported path, then test at least one additional alternative against the proposed patch. Do not complete a positive review without this comparison.
|
||||
|
||||
- Can the requested outcome already be achieved through configuration, composition, cloning, callbacks, extension points, a custom provider or adapter, or caller-owned code?
|
||||
- If the existing route is awkward, is the problem discoverability or ergonomics rather than missing capability?
|
||||
- What happens if maintainers make no code change?
|
||||
- Can input validation or an existing helper enforce the invariant earlier?
|
||||
- Can the fix be limited to the one supported path that fails?
|
||||
- Would documentation or a clearer error prevent misuse without runtime complexity?
|
||||
- Can the test be added first to reveal the smallest correct change?
|
||||
- Is the proposed public option compensating for an internal design issue?
|
||||
- Is the proposed core behavior actually provider- or application-specific policy that belongs at another layer?
|
||||
|
||||
## Competing PR comparison
|
||||
|
||||
When two or more open PRs address the same issue, first verify that they belong in one comparison set. Accept an explicit issue link, the same minimal reproduction, the same violated invariant, or materially overlapping runtime paths as association evidence. Do not treat a shared label or subsystem as sufficient.
|
||||
|
||||
Compare each candidate on the same evidence basis:
|
||||
|
||||
| Criterion | Question |
|
||||
|---|---|
|
||||
| Need | Does a concrete user outcome remain unmet after tracing existing supported functionality? |
|
||||
| Existing capability | Could every candidate be avoided by configuration, composition, an extension point, or a better caller- or provider-owned solution? |
|
||||
| Coverage | Does it solve the whole confirmed issue, a useful subset, or an adjacent problem? |
|
||||
| Correctness | Does the fix work on the real path and meaningful boundaries? |
|
||||
| Placement | Does it enforce the invariant at the correct shared layer? |
|
||||
| Tests | Does it reproduce the base failure and distinguish the candidate approaches? |
|
||||
| Compatibility | Does it preserve released APIs, state, protocol, providers, and established behavior? |
|
||||
| Complexity | What permanent branches, abstractions, configuration, or coupling does it add? |
|
||||
| Readiness | Is it mergeable now, or how much focused work remains? |
|
||||
| Reuse | Are there valuable tests or implementation pieces that should be combined into another candidate? |
|
||||
|
||||
Choose one portfolio-level disposition:
|
||||
|
||||
- **Prefer one PR**: identify the strongest candidate and close or supersede duplicates.
|
||||
- **Prefer one after focused changes**: keep one candidate active and state bounded changes required before merge.
|
||||
- **Combine selectively**: identify the destination PR and the exact ideas or tests worth transferring; avoid asking maintainers to reconcile entire competing implementations.
|
||||
- **Replace all**: explain the simpler or more coherent implementation that should supersede every candidate.
|
||||
- **Merge none**: the issue is invalid, negligible, unsupported, or none of the approaches justify completion cost.
|
||||
|
||||
Do not split the decision into independent approvals. Competing PRs consume overlapping review and maintenance budgets, so recommend one path for the issue as a whole.
|
||||
|
||||
## Maintainer comment drafts
|
||||
|
||||
Always write maintainer comments in English, regardless of the assessment language. Produce a draft when the recommendation is to close, request evidence, request focused code changes, supersede a PR, or choose one competing PR over another.
|
||||
|
||||
Keep each draft polite, direct, and copy-paste-ready. Usually use 60-160 words in one to three short paragraphs:
|
||||
|
||||
1. Acknowledge the contribution or report.
|
||||
2. Explain the decision with the smallest amount of decisive technical evidence.
|
||||
3. Give the exact next action or the condition for reconsideration.
|
||||
|
||||
Do not include internal labels such as `severity: low`, speculate about AI authorship or contributor intent, repeat the full review, or soften the message until the requested action becomes unclear.
|
||||
|
||||
Do not ask contributors to choose maintainer-owned semantics. If two implementations are technically possible but one changes the SDK contract, decide the contract in the review and make the comment actionable. Use a short rationale such as "This keeps the new handler scoped to the existing raise site" or "This makes the handler name match all invalid final messages", then request the exact code and tests for that decision.
|
||||
|
||||
### Close
|
||||
|
||||
```text
|
||||
Thanks for taking the time to investigate this. I traced the reported case through <path or behavior>, and <decisive finding>. In the supported path, <practical result>, so the added complexity is not justified by the demonstrated impact.
|
||||
|
||||
I am going to close this <issue/PR>. If you can provide <specific reproduction or evidence that would change the decision>, we can revisit the underlying problem with that narrower scope.
|
||||
```
|
||||
|
||||
### Request changes
|
||||
|
||||
```text
|
||||
Thanks for the contribution. The underlying issue is valid, and this approach is directionally reasonable. Before we can merge it, please address the following points: <bounded list of required changes>.
|
||||
|
||||
These changes are needed because <concise contract, lifecycle, compatibility, or test reason>. Once they are covered with a regression test that fails on the base and passes on the updated branch, the PR should be ready for another review.
|
||||
```
|
||||
|
||||
Adapt the wording to the actual evidence. Do not use these templates as generic filler.
|
||||
|
||||
### Existing capability or better alternative
|
||||
|
||||
```text
|
||||
Thanks for the contribution. I traced the underlying use case through <existing API or workflow>, which already supports <desired outcome and relevant limits>. The proposed change adds <new contract or complexity>, but the issue does not demonstrate a concrete supported case that the existing approach cannot handle.
|
||||
|
||||
I am going to close this <issue/PR> for now. If you can provide <specific scenario showing the existing approach is insufficient>, we can revisit the unmet need and choose the narrowest appropriate design from that evidence.
|
||||
```
|
||||
|
||||
## Compact report variants
|
||||
|
||||
Use `Maintainer decision` for a concluded review. Use `Preliminary assessment` when a desk review is tentatively positive but a decision-relevant runtime concern remains. `Verdict` is intentionally avoided in the report headings because it does not communicate whether the result is provisional or final.
|
||||
|
||||
### Runtime approval gate
|
||||
|
||||
```markdown
|
||||
## Preliminary assessment
|
||||
<Tentative issue or PR assessment based on desk review only.>
|
||||
|
||||
## Static evidence
|
||||
- <decisive code-path or test-inspection evidence>
|
||||
- <what remains uncertain at runtime>
|
||||
|
||||
## Proposed runtime probe
|
||||
- Concern: <the uncertainty that could change the decision>
|
||||
- Probe: <smallest exact execution path>
|
||||
- Control: <base, release, or known-good comparison when relevant>
|
||||
- Scope: <local-only or any live-service, cost, mutation, or cleanup implications>
|
||||
|
||||
## Approval request
|
||||
<Ask whether to run this exact probe. Do not present a final positive recommendation yet.>
|
||||
```
|
||||
|
||||
### Issue
|
||||
|
||||
```markdown
|
||||
## Maintainer decision
|
||||
<Real/partial/unproven/contradicted, severity, and disposition.>
|
||||
|
||||
## Evidence
|
||||
- <decisive evidence>
|
||||
- <scope or uncertainty>
|
||||
|
||||
## Existing capability and alternatives
|
||||
<Closest supported path, why it is or is not sufficient, and the preferred design alternative.>
|
||||
|
||||
## Recommendation
|
||||
<Prioritize, accept low priority, narrow, request evidence, or close.>
|
||||
|
||||
## Maintainer comment draft
|
||||
<Include when closure or additional evidence should be requested.>
|
||||
```
|
||||
|
||||
### Pull request
|
||||
|
||||
```markdown
|
||||
## Maintainer decision
|
||||
<Need, practical impact, and merge-worthiness.>
|
||||
- Need evidence: <Demonstrated / Plausible but unproven / Already covered / Unsupported>
|
||||
- Code recommendation: <code disposition>
|
||||
- Repository readiness: <integration status; include only for a merge-worthy recommendation when material>
|
||||
|
||||
## Evidence
|
||||
- <runtime or code-path result>
|
||||
- <test and compatibility result>
|
||||
|
||||
## Existing capability and alternatives
|
||||
<Closest supported path, why the demonstrated scenario cannot use it, and why this patch is preferable to no code change or a narrower design.>
|
||||
|
||||
## Issue impact
|
||||
- Validity: <claim validity>
|
||||
- Severity: <severity of the underlying issue or need>
|
||||
- Reach: <realistic reach>
|
||||
|
||||
## Patch risk
|
||||
<Include only when the proposed patch introduces a meaningful regression, compatibility, lifecycle, or maintenance risk.>
|
||||
|
||||
## PR quality
|
||||
- Solution fit: <assessment>
|
||||
- Tests: <assessment>
|
||||
- Remaining effort: <bounded/unbounded and why>
|
||||
|
||||
## Recommendation
|
||||
<Merge, focused changes, simpler replacement, or close.>
|
||||
|
||||
## Maintainer comment draft
|
||||
<Include only when closure, evidence, or changes should be requested.>
|
||||
```
|
||||
|
||||
### Competing pull requests
|
||||
|
||||
```markdown
|
||||
## Maintainer decision
|
||||
<Issue validity, practical severity, and preferred implementation path.>
|
||||
|
||||
## Open PR comparison
|
||||
| PR | Approach | Correctness | Tests | Compatibility/complexity | Readiness |
|
||||
|---|---|---|---|---|---|
|
||||
| #... | ... | ... | ... | ... | ... |
|
||||
|
||||
## Recommendation
|
||||
<Select one, request focused changes, combine specific parts, replace all, or merge none.>
|
||||
<State what should happen to every other open candidate.>
|
||||
|
||||
## Maintainer comment drafts
|
||||
<One copy-paste-ready draft for each PR that should be closed, changed, or superseded.>
|
||||
```
|
||||
@@ -1,16 +1,17 @@
|
||||
---
|
||||
name: pr-draft-summary
|
||||
description: Create the required PR-ready summary block, branch suggestion, title, and draft description for openai-agents-python. Use in the final handoff after moderate-or-larger changes to runtime code, tests, examples, build/test configuration, or docs with behavior impact; skip only for trivial or conversation-only tasks, repo-meta/doc-only tasks without behavior impact, or when the user explicitly says not to include the PR draft block.
|
||||
description: Create the required PR-ready summary block, branch suggestion, title, and draft description for openai-agents-python. Use before the final response whenever the current task changed runtime code, tests, examples, build/test configuration, or docs with behavior impact, regardless of perceived change size and including local-only or uncommitted work. Skip only for trivial or conversation-only tasks, repo-meta/doc-only tasks without behavior impact, or when the user explicitly says not to include the PR draft block.
|
||||
---
|
||||
|
||||
# PR Draft Summary
|
||||
|
||||
## Purpose
|
||||
Produce the PR-ready summary required in this repository after substantive code work is complete: a concise summary plus a PR-ready title and draft description that begins with "This pull request <verb> ...". The block should be ready to paste into a PR for openai-agents-python.
|
||||
Produce the PR-ready summary required in this repository after eligible code work is complete: a concise summary plus a PR-ready title and draft description that begins with "This pull request <verb> ...". The block should be ready to paste into a PR for openai-agents-python.
|
||||
|
||||
## When to Trigger
|
||||
- The task for this repo is finished (or ready for review) and it touched runtime code, tests, examples, docs with behavior impact, or build/test configuration.
|
||||
- Treat this as the default final handoff step for substantive code work. Run it after any required verification or changeset work and before sending the "work complete" response.
|
||||
- Before every final response, check whether the current task changed runtime code (`src/agents/`), tests (`tests/`), examples (`examples/`), build/test configuration, or docs with behavior impact.
|
||||
- If it did, run this skill after required verification and before sending the final response. Do not use perceived change size to decide whether to run it.
|
||||
- Run it for eligible local-only and uncommitted work even when the user did not ask to create a pull request. Producing this text does not authorize creating a branch, committing, pushing, or opening a pull request.
|
||||
- Skip only for trivial or conversation-only tasks, repo-meta/doc-only tasks without behavior impact, or when the user explicitly says not to include the PR draft block.
|
||||
|
||||
## Inputs to Collect Automatically (do not ask the user)
|
||||
|
||||
@@ -24,6 +24,7 @@ Use this skill to investigate real runtime behavior, not to restate code or docu
|
||||
- Use temporary files or a temporary directory for one-off probe scripts.
|
||||
- Keep temporary artifacts until the final response is drafted. Then delete them by default unless the user asked to keep them or they are needed for follow-up. Even when artifacts are deleted, keep a short run summary of the command shape, runtime context, and artifact status in the report.
|
||||
- Before executing a live probe that will read environment variables, tell the user the exact variable names you plan to use and why, then wait for explicit approval. Examples include `OPENAI_API_KEY` and other expected default names for the system under test.
|
||||
- When the environment-variable approval gate is required and the `request_user_input` tool is available, use that tool instead of a plain-text approval question. Ask one concise question with mutually exclusive choices such as `Allow once (Recommended)` and `Do not allow`, omit `autoResolutionMs`, and make the approval single-probe and limited to the exact named variables and destination. If the tool is unavailable, fall back to a concise plain-text approval question and do not proceed until the user explicitly approves.
|
||||
- Never print secrets, even when they come from standard environment variables that this skill may use.
|
||||
- For OpenAI API or OpenAI platform probes in this repository, use [$openai-knowledge](../openai-knowledge/SKILL.md) early to confirm contract-sensitive details such as supported parameters, field names, and limits. Use runtime probing to validate or challenge the documented behavior, not to skip the documentation pass entirely. If the docs MCP is unavailable, fall back to the official OpenAI docs and say that you used the fallback in the report.
|
||||
- For benchmark or comparison probes, make parity explicit before execution. Record what is held constant, what variable is under test, which response-shape constraints keep the comparison fair, and any usage or token counters that matter for interpreting latency or cost.
|
||||
@@ -51,7 +52,7 @@ Use this skill to investigate real runtime behavior, not to restate code or docu
|
||||
7. For comparative probes, define parity before execution. Record prompt or input shape, tool-choice setup, model-settings parity, state reuse rules, and any response-shape constraint that keeps the comparison fair. If materially different output length could bias the result, record usage or token notes too.
|
||||
8. If the question asks whether one option has the same intelligence or quality as another, decide whether the matrix supports only example-pattern parity or a broader quality claim. For broader claims, add at least one harder or more open-ended case. Otherwise say explicitly that the result is limited to the covered patterns.
|
||||
9. Plan state controls before execution when hidden state could affect the result. Record whether each case uses fresh or reused state, how cache reuse or cache busting is handled, what unique IDs isolate repeated runs, and how cleanup is verified.
|
||||
10. If any live case will read environment variables, list the exact variable names and purpose for each case, then ask the user for approval before execution. Keep the approval ask short and include destination, read-only versus mutating or costly risk, exact variable names, and cleanup or rollback if relevant.
|
||||
10. If any live case will read environment variables, list the exact variable names and purpose for each case, then ask the user for approval before execution. Prefer `request_user_input` for this gate when it is available, with no auto-resolution and choices that grant or deny only this specific probe. Keep the approval ask short and include destination, read-only versus mutating or costly risk, exact variable names, and cleanup or rollback if relevant.
|
||||
11. Build task-specific probe scripts in a temporary location. Keep the script small, observable, and easy to discard.
|
||||
12. In `openai-agents-python`, make the runtime context explicit:
|
||||
- Run Python probes from the repository root with `uv run python` when practical.
|
||||
|
||||
@@ -27,6 +27,17 @@ Do not read these variables automatically. Before a live probe uses any of them,
|
||||
|
||||
If the task targets another standard integration, use that integration's expected default variable names under the same rule.
|
||||
|
||||
## Environment False Signals
|
||||
|
||||
Before attributing a failure to the patch under review, exclude environment and source-selection problems with a control run.
|
||||
|
||||
- Confirm the commit and worktree under test. When editable installs, shared environments, `PYTHONPATH`, or generated artifacts can select stale code, verify the imported package path and rebuild before probing.
|
||||
- Run base and head controls with the same interpreter, dependencies, environment variables, and command shape.
|
||||
- Treat proxy initialization, sandbox denials, unavailable containers, expired snapshots, authentication, quotas, rate limits, service outages, and stale caches as environment conditions until a controlled rerun ties them to the patch.
|
||||
- Never print proxy URLs or credentials. Change only the minimum in-scope environment or disposable state needed for the control run, and record which variable names or constraints changed.
|
||||
|
||||
In the final report, distinguish code failures, unsupported configurations, environment blockers, and inconclusive probes. Do not combine them into one failed-test count.
|
||||
|
||||
## Responses API Probe Patterns
|
||||
|
||||
For Responses API work, start from the uncertainty instead of from the full feature surface.
|
||||
|
||||
@@ -5,6 +5,7 @@
|
||||
### Test plan
|
||||
|
||||
<!-- Please explain how this was tested -->
|
||||
<!-- If verification could not complete because of local environment setup, include the failing command, the missing dependency, and why it is unrelated to this PR. Leave the pass checkbox below unchecked until all verification steps pass. -->
|
||||
|
||||
### Issue number
|
||||
|
||||
@@ -12,7 +13,7 @@
|
||||
|
||||
### Checks
|
||||
|
||||
- [ ] I've added new tests (if relevant)
|
||||
- [ ] I've added/updated the relevant documentation
|
||||
- [ ] I've run `make lint` and `make format`
|
||||
- [ ] I've made sure tests pass
|
||||
- [ ] I've added new tests, if relevant
|
||||
- [ ] I've run `.agents/skills/code-change-verification/scripts/run.sh`
|
||||
- [ ] I've confirmed all verification steps pass
|
||||
- [ ] If using Codex, I've run `/review` before submitting this PR
|
||||
|
||||
@@ -1,73 +0,0 @@
|
||||
# PR auto-labeling
|
||||
|
||||
You are Codex running in CI to propose labels for a pull request in the openai-agents-python repository.
|
||||
|
||||
Inputs:
|
||||
- PR context: .tmp/pr-labels/pr-context.json
|
||||
- PR diff: .tmp/pr-labels/changes.diff
|
||||
- Changed files: .tmp/pr-labels/changed-files.txt
|
||||
|
||||
Task:
|
||||
- Inspect the PR context, diff, and changed files.
|
||||
- Output JSON with a single top-level key: "labels" (array of strings).
|
||||
- Only use labels from the allowed list.
|
||||
- Prefer false negatives over false positives. If you are unsure, leave the label out.
|
||||
- Return the smallest accurate set of labels for the PR's primary intent and primary surface area.
|
||||
|
||||
Allowed labels:
|
||||
- documentation
|
||||
- project
|
||||
- bug
|
||||
- enhancement
|
||||
- dependencies
|
||||
- feature:chat-completions
|
||||
- feature:core
|
||||
- feature:extensions
|
||||
- feature:mcp
|
||||
- feature:realtime
|
||||
- feature:sandboxes
|
||||
- feature:sessions
|
||||
- feature:tracing
|
||||
- feature:voice
|
||||
|
||||
Important guidance:
|
||||
- `documentation`, `project`, and `dependencies` are also derived deterministically elsewhere in the workflow. You may include them when the evidence is explicit, but do not stretch to infer them from weak signals.
|
||||
- Use direct evidence from changed implementation files and the dominant intent of the diff. Do not add labels based only on tests, examples, comments, docstrings, imports, type plumbing, or shared helpers.
|
||||
- Cross-cutting features often touch many adapters and support layers. Only add a `feature:*` label when that area is itself a primary user-facing surface of the PR, not when it receives incidental compatibility or parity updates.
|
||||
- Mentions of a feature area in helper names, comments, tests, or trace metadata are not enough by themselves.
|
||||
- Prefer the most general accurate feature label over a larger set of narrower labels. For broad runtime work, this usually means `feature:core`.
|
||||
- A secondary `feature:*` label needs two things: a non-test implementation/docs change in that area, and evidence that the area is a user-facing outcome of the PR rather than support work for another feature.
|
||||
|
||||
Label rules:
|
||||
- documentation: Documentation changes (docs/), or src/ changes that only modify comments/docstrings without behavior changes. If only comments/docstrings change in src/, do not add bug/enhancement.
|
||||
- project: Any change to pyproject.toml.
|
||||
- dependencies: Dependencies are added/removed/updated (pyproject.toml dependency sections or uv.lock changes).
|
||||
- bug: The PR's primary intent is to correct existing incorrect behavior. Use only with strong evidence such as the title/body/tests clearly describing a fix, regression, crash, incorrect output, or restore/preserve behavior. Do not add `bug` for incidental hardening that accompanies a new feature.
|
||||
- enhancement: The PR's primary intent is to add or expand functionality. Prefer `enhancement` for feature work even if the diff also contains some fixes or guardrails needed to support that feature.
|
||||
- bug vs enhancement: Prefer exactly one of these. Include both only when the PR clearly contains two separate substantial changes and both are first-order outcomes.
|
||||
- feature:chat-completions: Chat Completions support or conversion is a primary deliverable of the PR. Do not add it for a small compatibility guard or parity update in `chatcmpl_converter.py`.
|
||||
- feature:core: Core agent loop, tool calls, run pipeline, or other central runtime behavior is a primary surface of the PR. For cross-cutting runtime changes, this is usually the single best feature label.
|
||||
- feature:extensions: `src/agents/extensions/` surfaces are a primary deliverable of the PR, including extension models/providers such as Any-LLM and LiteLLM. Changes under `src/agents/extensions/sandbox/` can warrant this label alongside `feature:sandboxes`.
|
||||
- feature:mcp: MCP-specific behavior or APIs are a primary deliverable of the PR. Do not add it for incidental hosted/deferred tool plumbing touched by broader runtime work.
|
||||
- feature:realtime: Realtime-specific behavior, API shape, or session semantics are a primary deliverable of the PR. Do not add it for small parity updates in realtime adapters.
|
||||
- feature:sandboxes: Sandbox runtime or sandbox extension behavior is a primary deliverable of the PR, including changes under `src/agents/sandbox/` and `src/agents/extensions/sandbox/`. Prefer this over `feature:core` for sandbox-focused work; for `src/agents/extensions/sandbox/`, `feature:extensions` may also be appropriate.
|
||||
- feature:sessions: Session or memory behavior is a primary deliverable of the PR. Do not add it for persistence updates that merely support a broader feature.
|
||||
- feature:tracing: Tracing is a primary deliverable of the PR. Do not add it for trace naming or metadata changes that accompany another feature.
|
||||
- feature:voice: Voice pipeline behavior is a primary deliverable of the PR.
|
||||
|
||||
Decision process:
|
||||
1. Determine the PR's primary intent in one sentence from the PR title/body and dominant runtime diff.
|
||||
2. Start with zero labels.
|
||||
3. Add `bug` or `enhancement` conservatively.
|
||||
4. Add only the minimum `feature:*` labels needed to describe the primary surface area.
|
||||
5. Treat extra `feature:*` labels as guilty until proven necessary. Keep them only when the PR would feel mislabeled without them.
|
||||
6. Re-check every label. Drop any label that is supported only by secondary edits, parity work, or touched files outside the PR's main focus.
|
||||
|
||||
Examples:
|
||||
- If a new cross-cutting runtime feature touches Chat Completions, Realtime, Sessions, MCP, and tracing support code for parity, prefer `["enhancement","feature:core"]` over labeling every touched area.
|
||||
- If a PR mainly adds a Responses/core capability and touches realtime or sessions files only to keep shared serialization, replay, or adapters in sync, do not add `feature:realtime` or `feature:sessions`.
|
||||
- If a PR mainly fixes realtime transport behavior and also updates tests/docs, prefer `["bug","feature:realtime"]`.
|
||||
|
||||
Output:
|
||||
- JSON only (no code fences, no extra text).
|
||||
- Example: {"labels":["enhancement","feature:core"]}
|
||||
@@ -1,23 +0,0 @@
|
||||
# Release readiness review
|
||||
|
||||
You are Codex running in CI. Produce a release readiness report for this repository.
|
||||
|
||||
Steps:
|
||||
1. Determine the latest release tag (use local tags only):
|
||||
- `git tag -l 'v*' --sort=-v:refname | head -n1`
|
||||
2. Set TARGET to the current commit SHA: `git rev-parse HEAD`.
|
||||
3. Collect diff context for BASE_TAG...TARGET:
|
||||
- `git diff --stat BASE_TAG...TARGET`
|
||||
- `git diff --dirstat=files,0 BASE_TAG...TARGET`
|
||||
- `git diff --name-status BASE_TAG...TARGET`
|
||||
- `git log --oneline --reverse BASE_TAG..TARGET`
|
||||
4. Review `.agents/skills/final-release-review/references/review-checklist.md` and analyze the diff.
|
||||
|
||||
Output:
|
||||
- Write the report in the exact format used by `$final-release-review` (see `.agents/skills/final-release-review/SKILL.md`).
|
||||
- Use the compare URL: `https://github.com/${GITHUB_REPOSITORY}/compare/BASE_TAG...TARGET`.
|
||||
- Include clear ship/block call and risk levels.
|
||||
- If no risks are found, include "No material risks identified".
|
||||
|
||||
Constraints:
|
||||
- Output only the report (no code fences, no extra commentary).
|
||||
@@ -1,29 +0,0 @@
|
||||
{
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["labels"],
|
||||
"properties": {
|
||||
"labels": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "string",
|
||||
"enum": [
|
||||
"documentation",
|
||||
"project",
|
||||
"bug",
|
||||
"enhancement",
|
||||
"dependencies",
|
||||
"feature:chat-completions",
|
||||
"feature:core",
|
||||
"feature:extensions",
|
||||
"feature:mcp",
|
||||
"feature:realtime",
|
||||
"feature:sandboxes",
|
||||
"feature:sessions",
|
||||
"feature:tracing",
|
||||
"feature:voice"
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1,442 +0,0 @@
|
||||
#!/usr/bin/env python3
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import pathlib
|
||||
import subprocess
|
||||
import sys
|
||||
from collections.abc import Sequence
|
||||
from dataclasses import dataclass
|
||||
from typing import Any, Final
|
||||
|
||||
ALLOWED_LABELS: Final[set[str]] = {
|
||||
"documentation",
|
||||
"project",
|
||||
"bug",
|
||||
"enhancement",
|
||||
"dependencies",
|
||||
"feature:chat-completions",
|
||||
"feature:core",
|
||||
"feature:extensions",
|
||||
"feature:mcp",
|
||||
"feature:realtime",
|
||||
"feature:sandboxes",
|
||||
"feature:sessions",
|
||||
"feature:tracing",
|
||||
"feature:voice",
|
||||
}
|
||||
|
||||
DETERMINISTIC_LABELS: Final[set[str]] = {
|
||||
"documentation",
|
||||
"project",
|
||||
"dependencies",
|
||||
}
|
||||
|
||||
MODEL_ONLY_LABELS: Final[set[str]] = {
|
||||
"bug",
|
||||
"enhancement",
|
||||
}
|
||||
|
||||
FEATURE_LABELS: Final[set[str]] = ALLOWED_LABELS - DETERMINISTIC_LABELS - MODEL_ONLY_LABELS
|
||||
|
||||
SOURCE_FEATURE_PREFIXES: Final[dict[str, tuple[str, ...]]] = {
|
||||
"feature:realtime": ("src/agents/realtime/",),
|
||||
"feature:sandboxes": ("src/agents/sandbox/", "src/agents/extensions/sandbox/"),
|
||||
"feature:voice": ("src/agents/voice/",),
|
||||
"feature:mcp": ("src/agents/mcp/",),
|
||||
"feature:tracing": ("src/agents/tracing/",),
|
||||
"feature:sessions": ("src/agents/memory/",),
|
||||
}
|
||||
|
||||
CORE_EXCLUDED_PREFIXES: Final[tuple[str, ...]] = (
|
||||
"src/agents/realtime/",
|
||||
"src/agents/voice/",
|
||||
"src/agents/mcp/",
|
||||
"src/agents/tracing/",
|
||||
"src/agents/memory/",
|
||||
"src/agents/extensions/",
|
||||
"src/agents/models/",
|
||||
)
|
||||
|
||||
PR_CONTEXT_DEFAULT_PATH = ".tmp/pr-labels/pr-context.json"
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class PRContext:
|
||||
title: str = ""
|
||||
body: str = ""
|
||||
|
||||
|
||||
def read_file_at(commit: str | None, path: str) -> str | None:
|
||||
if not commit:
|
||||
return None
|
||||
try:
|
||||
return subprocess.check_output(["git", "show", f"{commit}:{path}"], text=True)
|
||||
except subprocess.CalledProcessError:
|
||||
return None
|
||||
|
||||
|
||||
def dependency_lines_for_pyproject(text: str) -> set[int]:
|
||||
dependency_lines: set[int] = set()
|
||||
current_section: str | None = None
|
||||
in_project_dependencies = False
|
||||
|
||||
for line_number, raw_line in enumerate(text.splitlines(), start=1):
|
||||
stripped = raw_line.strip()
|
||||
if stripped.startswith("[") and stripped.endswith("]"):
|
||||
if stripped.startswith("[[") and stripped.endswith("]]"):
|
||||
current_section = stripped[2:-2].strip()
|
||||
else:
|
||||
current_section = stripped[1:-1].strip()
|
||||
in_project_dependencies = False
|
||||
if current_section in ("project.optional-dependencies", "dependency-groups"):
|
||||
dependency_lines.add(line_number)
|
||||
continue
|
||||
|
||||
if current_section in ("project.optional-dependencies", "dependency-groups"):
|
||||
dependency_lines.add(line_number)
|
||||
continue
|
||||
|
||||
if current_section != "project":
|
||||
continue
|
||||
|
||||
if in_project_dependencies:
|
||||
dependency_lines.add(line_number)
|
||||
if "]" in stripped:
|
||||
in_project_dependencies = False
|
||||
continue
|
||||
|
||||
if stripped.startswith("dependencies") and "=" in stripped:
|
||||
dependency_lines.add(line_number)
|
||||
if "[" in stripped and "]" not in stripped:
|
||||
in_project_dependencies = True
|
||||
|
||||
return dependency_lines
|
||||
|
||||
|
||||
def pyproject_dependency_changed(
|
||||
diff_text: str,
|
||||
*,
|
||||
base_sha: str | None,
|
||||
head_sha: str | None,
|
||||
) -> bool:
|
||||
import re
|
||||
|
||||
base_text = read_file_at(base_sha, "pyproject.toml")
|
||||
head_text = read_file_at(head_sha, "pyproject.toml")
|
||||
if base_text is None and head_text is None:
|
||||
return False
|
||||
|
||||
base_dependency_lines = dependency_lines_for_pyproject(base_text) if base_text else set()
|
||||
head_dependency_lines = dependency_lines_for_pyproject(head_text) if head_text else set()
|
||||
|
||||
in_pyproject = False
|
||||
base_line: int | None = None
|
||||
head_line: int | None = None
|
||||
hunk_re = re.compile(r"@@ -(\d+)(?:,\d+)? \+(\d+)(?:,\d+)? @@")
|
||||
|
||||
for line in diff_text.splitlines():
|
||||
if line.startswith("+++ b/"):
|
||||
current_file = line[len("+++ b/") :].strip()
|
||||
in_pyproject = current_file == "pyproject.toml"
|
||||
base_line = None
|
||||
head_line = None
|
||||
continue
|
||||
|
||||
if not in_pyproject:
|
||||
continue
|
||||
|
||||
if line.startswith("@@ "):
|
||||
match = hunk_re.match(line)
|
||||
if not match:
|
||||
continue
|
||||
base_line = int(match.group(1))
|
||||
head_line = int(match.group(2))
|
||||
continue
|
||||
|
||||
if base_line is None or head_line is None:
|
||||
continue
|
||||
|
||||
if line.startswith(" "):
|
||||
base_line += 1
|
||||
head_line += 1
|
||||
continue
|
||||
|
||||
if line.startswith("-"):
|
||||
if base_line in base_dependency_lines:
|
||||
return True
|
||||
base_line += 1
|
||||
continue
|
||||
|
||||
if line.startswith("+"):
|
||||
if head_line in head_dependency_lines:
|
||||
return True
|
||||
head_line += 1
|
||||
continue
|
||||
|
||||
return False
|
||||
|
||||
|
||||
def infer_specific_feature_labels(changed_files: Sequence[str]) -> set[str]:
|
||||
source_files = [path for path in changed_files if path.startswith("src/")]
|
||||
labels: set[str] = set()
|
||||
|
||||
for label, prefixes in SOURCE_FEATURE_PREFIXES.items():
|
||||
if any(path.startswith(prefix) for path in source_files for prefix in prefixes):
|
||||
labels.add(label)
|
||||
|
||||
if any(path.startswith("src/agents/extensions/") for path in source_files):
|
||||
labels.add("feature:extensions")
|
||||
|
||||
if any(
|
||||
path.startswith(("src/agents/models/", "src/agents/extensions/models/"))
|
||||
and ("chatcmpl" in path or "chatcompletions" in path)
|
||||
for path in source_files
|
||||
):
|
||||
labels.add("feature:chat-completions")
|
||||
|
||||
return labels
|
||||
|
||||
|
||||
def infer_feature_labels(changed_files: Sequence[str]) -> set[str]:
|
||||
source_files = [path for path in changed_files if path.startswith("src/")]
|
||||
specific_labels = infer_specific_feature_labels(source_files)
|
||||
core_touched = any(
|
||||
path.startswith("src/agents/") and not path.startswith(CORE_EXCLUDED_PREFIXES)
|
||||
for path in source_files
|
||||
)
|
||||
|
||||
if core_touched and len(specific_labels) != 1:
|
||||
return {"feature:core"}
|
||||
return specific_labels
|
||||
|
||||
|
||||
def infer_fallback_labels(changed_files: Sequence[str]) -> set[str]:
|
||||
return infer_feature_labels(changed_files)
|
||||
|
||||
|
||||
def load_json(path: pathlib.Path) -> Any:
|
||||
return json.loads(path.read_text())
|
||||
|
||||
|
||||
def load_pr_context(path: pathlib.Path) -> PRContext:
|
||||
if not path.exists():
|
||||
return PRContext()
|
||||
|
||||
try:
|
||||
payload = load_json(path)
|
||||
except json.JSONDecodeError:
|
||||
return PRContext()
|
||||
|
||||
if not isinstance(payload, dict):
|
||||
return PRContext()
|
||||
|
||||
title = payload.get("title", "")
|
||||
body = payload.get("body", "")
|
||||
if not isinstance(title, str):
|
||||
title = ""
|
||||
if not isinstance(body, str):
|
||||
body = ""
|
||||
|
||||
return PRContext(title=title, body=body)
|
||||
|
||||
|
||||
def load_codex_labels(path: pathlib.Path) -> tuple[list[str], bool]:
|
||||
if not path.exists():
|
||||
return [], False
|
||||
|
||||
raw = path.read_text().strip()
|
||||
if not raw:
|
||||
return [], False
|
||||
|
||||
try:
|
||||
payload = load_json(path)
|
||||
except json.JSONDecodeError:
|
||||
return [], False
|
||||
|
||||
if not isinstance(payload, dict):
|
||||
return [], False
|
||||
|
||||
labels = payload.get("labels")
|
||||
if not isinstance(labels, list):
|
||||
return [], False
|
||||
|
||||
if not all(isinstance(label, str) for label in labels):
|
||||
return [], False
|
||||
|
||||
return list(labels), True
|
||||
|
||||
|
||||
def fetch_existing_labels(pr_number: str) -> set[str]:
|
||||
result = subprocess.check_output(
|
||||
["gh", "pr", "view", pr_number, "--json", "labels", "--jq", ".labels[].name"],
|
||||
text=True,
|
||||
).strip()
|
||||
return {label for label in result.splitlines() if label}
|
||||
|
||||
|
||||
def infer_title_intent_labels(pr_context: PRContext) -> set[str]:
|
||||
normalized_title = pr_context.title.strip().lower()
|
||||
|
||||
bug_prefixes = ("fix:", "fix(", "bug:", "bugfix:", "hotfix:", "regression:")
|
||||
enhancement_prefixes = ("feat:", "feat(", "feature:", "enhancement:")
|
||||
|
||||
if normalized_title.startswith(bug_prefixes):
|
||||
return {"bug"}
|
||||
if normalized_title.startswith(enhancement_prefixes):
|
||||
return {"enhancement"}
|
||||
return set()
|
||||
|
||||
|
||||
def compute_desired_labels(
|
||||
*,
|
||||
pr_context: PRContext,
|
||||
changed_files: Sequence[str],
|
||||
diff_text: str,
|
||||
codex_ran: bool,
|
||||
codex_output_valid: bool,
|
||||
codex_labels: Sequence[str],
|
||||
base_sha: str | None,
|
||||
head_sha: str | None,
|
||||
) -> set[str]:
|
||||
desired: set[str] = set()
|
||||
codex_label_set = {label for label in codex_labels if label in ALLOWED_LABELS}
|
||||
codex_feature_labels = codex_label_set & FEATURE_LABELS
|
||||
codex_model_only_labels = codex_label_set & MODEL_ONLY_LABELS
|
||||
fallback_feature_labels = infer_fallback_labels(changed_files)
|
||||
title_intent_labels = infer_title_intent_labels(pr_context)
|
||||
|
||||
if "pyproject.toml" in changed_files:
|
||||
desired.add("project")
|
||||
|
||||
if any(path.startswith("docs/") for path in changed_files):
|
||||
desired.add("documentation")
|
||||
|
||||
dependencies_allowed = "uv.lock" in changed_files
|
||||
if "pyproject.toml" in changed_files and pyproject_dependency_changed(
|
||||
diff_text, base_sha=base_sha, head_sha=head_sha
|
||||
):
|
||||
dependencies_allowed = True
|
||||
if dependencies_allowed:
|
||||
desired.add("dependencies")
|
||||
|
||||
if codex_ran and codex_output_valid and codex_feature_labels:
|
||||
desired.update(codex_feature_labels)
|
||||
else:
|
||||
desired.update(fallback_feature_labels)
|
||||
|
||||
if title_intent_labels:
|
||||
desired.update(title_intent_labels)
|
||||
elif codex_ran and codex_output_valid:
|
||||
desired.update(codex_model_only_labels)
|
||||
|
||||
if any(path.startswith("src/agents/extensions/sandbox/") for path in changed_files):
|
||||
desired.update({"feature:extensions", "feature:sandboxes"})
|
||||
|
||||
return desired
|
||||
|
||||
|
||||
def compute_managed_labels(
|
||||
*,
|
||||
pr_context: PRContext,
|
||||
codex_ran: bool,
|
||||
codex_output_valid: bool,
|
||||
codex_labels: Sequence[str],
|
||||
) -> set[str]:
|
||||
managed = DETERMINISTIC_LABELS | FEATURE_LABELS
|
||||
title_intent_labels = infer_title_intent_labels(pr_context)
|
||||
codex_label_set = {label for label in codex_labels if label in MODEL_ONLY_LABELS}
|
||||
if title_intent_labels or (codex_ran and codex_output_valid and codex_label_set):
|
||||
managed |= MODEL_ONLY_LABELS
|
||||
return managed
|
||||
|
||||
|
||||
def parse_args(argv: Sequence[str] | None = None) -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("--pr-number", default=os.environ.get("PR_NUMBER", ""))
|
||||
parser.add_argument("--base-sha", default=os.environ.get("PR_BASE_SHA", ""))
|
||||
parser.add_argument("--head-sha", default=os.environ.get("PR_HEAD_SHA", ""))
|
||||
parser.add_argument(
|
||||
"--codex-output-path",
|
||||
default=os.environ.get("CODEX_OUTPUT_PATH", ".tmp/codex/outputs/pr-labels.json"),
|
||||
)
|
||||
parser.add_argument("--codex-conclusion", default=os.environ.get("CODEX_CONCLUSION", ""))
|
||||
parser.add_argument(
|
||||
"--pr-context-path",
|
||||
default=os.environ.get("PR_CONTEXT_PATH", PR_CONTEXT_DEFAULT_PATH),
|
||||
)
|
||||
parser.add_argument(
|
||||
"--changed-files-path",
|
||||
default=os.environ.get("CHANGED_FILES_PATH", ".tmp/pr-labels/changed-files.txt"),
|
||||
)
|
||||
parser.add_argument(
|
||||
"--changes-diff-path",
|
||||
default=os.environ.get("CHANGES_DIFF_PATH", ".tmp/pr-labels/changes.diff"),
|
||||
)
|
||||
return parser.parse_args(argv)
|
||||
|
||||
|
||||
def main(argv: Sequence[str] | None = None) -> int:
|
||||
args = parse_args(argv)
|
||||
if not args.pr_number:
|
||||
raise SystemExit("Missing PR number.")
|
||||
|
||||
changed_files_path = pathlib.Path(args.changed_files_path)
|
||||
changes_diff_path = pathlib.Path(args.changes_diff_path)
|
||||
codex_output_path = pathlib.Path(args.codex_output_path)
|
||||
pr_context_path = pathlib.Path(args.pr_context_path)
|
||||
codex_conclusion = args.codex_conclusion.strip().lower()
|
||||
codex_ran = bool(codex_conclusion) and codex_conclusion != "skipped"
|
||||
pr_context = load_pr_context(pr_context_path)
|
||||
|
||||
changed_files = []
|
||||
if changed_files_path.exists():
|
||||
changed_files = [
|
||||
line.strip() for line in changed_files_path.read_text().splitlines() if line.strip()
|
||||
]
|
||||
|
||||
diff_text = changes_diff_path.read_text() if changes_diff_path.exists() else ""
|
||||
codex_labels, codex_output_valid = load_codex_labels(codex_output_path)
|
||||
if codex_ran and not codex_output_valid:
|
||||
print(
|
||||
"Codex output missing or invalid; using fallback feature labels and preserving "
|
||||
"model-only labels."
|
||||
)
|
||||
desired = compute_desired_labels(
|
||||
pr_context=pr_context,
|
||||
changed_files=changed_files,
|
||||
diff_text=diff_text,
|
||||
codex_ran=codex_ran,
|
||||
codex_output_valid=codex_output_valid,
|
||||
codex_labels=codex_labels,
|
||||
base_sha=args.base_sha or None,
|
||||
head_sha=args.head_sha or None,
|
||||
)
|
||||
|
||||
existing = fetch_existing_labels(args.pr_number)
|
||||
managed_labels = compute_managed_labels(
|
||||
pr_context=pr_context,
|
||||
codex_ran=codex_ran,
|
||||
codex_output_valid=codex_output_valid,
|
||||
codex_labels=codex_labels,
|
||||
)
|
||||
to_add = sorted(desired - existing)
|
||||
to_remove = sorted((existing & managed_labels) - desired)
|
||||
|
||||
if not to_add and not to_remove:
|
||||
print("Labels already up to date.")
|
||||
return 0
|
||||
|
||||
cmd = ["gh", "pr", "edit", args.pr_number]
|
||||
if to_add:
|
||||
cmd += ["--add-label", ",".join(to_add)]
|
||||
if to_remove:
|
||||
cmd += ["--remove-label", ",".join(to_remove)]
|
||||
subprocess.check_call(cmd)
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -16,7 +16,7 @@ jobs:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
|
||||
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0
|
||||
- name: Determine docs-only push
|
||||
id: docs-only
|
||||
run: |
|
||||
@@ -36,8 +36,9 @@ jobs:
|
||||
fi
|
||||
- name: Setup uv
|
||||
if: steps.docs-only.outputs.skip != 'true'
|
||||
uses: astral-sh/setup-uv@cec208311dfd045dd5311c1add060b2062131d57
|
||||
uses: astral-sh/setup-uv@fac544c07dec837d0ccb6301d7b5580bf5edae39 # setup-uv v8.1.0; uv 0.11.14
|
||||
with:
|
||||
version: "0.11.14"
|
||||
enable-cache: true
|
||||
- name: Install dependencies
|
||||
if: steps.docs-only.outputs.skip != 'true'
|
||||
|
||||
@@ -10,7 +10,7 @@ jobs:
|
||||
issues: write
|
||||
pull-requests: write
|
||||
steps:
|
||||
- uses: actions/stale@b5d41d4e1d5dceea10e7104786b73624c18a190f
|
||||
- uses: actions/stale@eb5cf3af3ac0a1aa4c9c45633dd1ae542a27a899
|
||||
with:
|
||||
days-before-issue-stale: 7
|
||||
days-before-issue-close: 3
|
||||
|
||||
@@ -1,204 +0,0 @@
|
||||
name: Auto label PRs
|
||||
|
||||
on:
|
||||
pull_request_target:
|
||||
types:
|
||||
- opened
|
||||
- reopened
|
||||
- synchronize
|
||||
- ready_for_review
|
||||
workflow_dispatch:
|
||||
inputs:
|
||||
pr_number:
|
||||
description: "PR number to label."
|
||||
required: true
|
||||
type: number
|
||||
|
||||
permissions:
|
||||
contents: read
|
||||
issues: write
|
||||
pull-requests: write
|
||||
|
||||
jobs:
|
||||
label:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- name: Ensure main workflow
|
||||
if: ${{ github.event_name == 'workflow_dispatch' && github.ref != 'refs/heads/main' }}
|
||||
run: |
|
||||
echo "This workflow must be dispatched from main."
|
||||
exit 1
|
||||
|
||||
- name: Resolve PR context
|
||||
id: pr
|
||||
uses: actions/github-script@ed597411d8f924073f98dfc5c65a23a2325f34cd
|
||||
env:
|
||||
MANUAL_PR_NUMBER: ${{ inputs.pr_number || '' }}
|
||||
with:
|
||||
github-token: ${{ secrets.GITHUB_TOKEN }}
|
||||
script: |
|
||||
const isManual = context.eventName === 'workflow_dispatch';
|
||||
let pr;
|
||||
if (isManual) {
|
||||
const prNumber = Number(process.env.MANUAL_PR_NUMBER);
|
||||
if (!prNumber) {
|
||||
core.setFailed('workflow_dispatch requires pr_number input.');
|
||||
return;
|
||||
}
|
||||
const { data } = await github.rest.pulls.get({
|
||||
owner: context.repo.owner,
|
||||
repo: context.repo.repo,
|
||||
pull_number: prNumber,
|
||||
});
|
||||
pr = data;
|
||||
} else {
|
||||
pr = context.payload.pull_request;
|
||||
}
|
||||
if (!pr) {
|
||||
core.setFailed('Missing pull request context.');
|
||||
return;
|
||||
}
|
||||
const headRepo = pr.head.repo.full_name;
|
||||
const repoFullName = `${context.repo.owner}/${context.repo.repo}`;
|
||||
core.setOutput('pr_number', pr.number);
|
||||
core.setOutput('base_sha', pr.base.sha);
|
||||
core.setOutput('head_sha', pr.head.sha);
|
||||
core.setOutput('head_repo', headRepo);
|
||||
core.setOutput('is_fork', headRepo !== repoFullName);
|
||||
core.setOutput('title', pr.title || '');
|
||||
core.setOutput('body', pr.body || '');
|
||||
|
||||
- name: Checkout base
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
|
||||
with:
|
||||
fetch-depth: 0
|
||||
ref: ${{ steps.pr.outputs.base_sha }}
|
||||
- name: Fetch PR head
|
||||
env:
|
||||
PR_HEAD_REPO: ${{ steps.pr.outputs.head_repo }}
|
||||
PR_HEAD_SHA: ${{ steps.pr.outputs.head_sha }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
git fetch --no-tags --prune --recurse-submodules=no \
|
||||
"https://github.com/${PR_HEAD_REPO}.git" \
|
||||
"${PR_HEAD_SHA}"
|
||||
- name: Collect PR diff
|
||||
id: diff
|
||||
env:
|
||||
PR_BASE_SHA: ${{ steps.pr.outputs.base_sha }}
|
||||
PR_HEAD_SHA: ${{ steps.pr.outputs.head_sha }}
|
||||
PR_TITLE: ${{ steps.pr.outputs.title }}
|
||||
PR_BODY: ${{ steps.pr.outputs.body }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
mkdir -p .tmp/pr-labels
|
||||
diff_base_sha="$(git merge-base "$PR_BASE_SHA" "$PR_HEAD_SHA")"
|
||||
echo "diff_base_sha=${diff_base_sha}" >> "$GITHUB_OUTPUT"
|
||||
git diff --name-only "$diff_base_sha" "$PR_HEAD_SHA" > .tmp/pr-labels/changed-files.txt
|
||||
git diff "$diff_base_sha" "$PR_HEAD_SHA" > .tmp/pr-labels/changes.diff
|
||||
python - <<'PY'
|
||||
import json
|
||||
import os
|
||||
import pathlib
|
||||
|
||||
pathlib.Path(".tmp/pr-labels/pr-context.json").write_text(
|
||||
json.dumps(
|
||||
{
|
||||
"title": os.environ.get("PR_TITLE", ""),
|
||||
"body": os.environ.get("PR_BODY", ""),
|
||||
},
|
||||
ensure_ascii=False,
|
||||
indent=2,
|
||||
)
|
||||
+ "\n"
|
||||
)
|
||||
PY
|
||||
- name: Prepare Codex output
|
||||
id: codex-output
|
||||
run: |
|
||||
set -euo pipefail
|
||||
output_dir=".tmp/codex/outputs"
|
||||
output_file="${output_dir}/pr-labels.json"
|
||||
mkdir -p "$output_dir"
|
||||
echo "output_file=${output_file}" >> "$GITHUB_OUTPUT"
|
||||
- name: Run Codex labeling
|
||||
id: run_codex
|
||||
if: ${{ (github.event_name == 'workflow_dispatch' || steps.pr.outputs.is_fork != 'true') && github.actor != 'dependabot[bot]' }}
|
||||
uses: openai/codex-action@c25d10f3f498316d4b2496cc4c6dd58057a7b031
|
||||
with:
|
||||
openai-api-key: ${{ secrets.PROD_OPENAI_API_KEY }}
|
||||
prompt-file: .github/codex/prompts/pr-labels.md
|
||||
output-file: ${{ steps.codex-output.outputs.output_file }}
|
||||
output-schema-file: .github/codex/schemas/pr-labels.json
|
||||
# Keep the legacy Linux sandbox path until the default bubblewrap path
|
||||
# works reliably on GitHub-hosted Ubuntu runners.
|
||||
codex-args: '["--enable","use_legacy_landlock"]'
|
||||
safety-strategy: drop-sudo
|
||||
sandbox: read-only
|
||||
- name: Apply labels
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
PR_NUMBER: ${{ steps.pr.outputs.pr_number }}
|
||||
PR_BASE_SHA: ${{ steps.diff.outputs.diff_base_sha }}
|
||||
PR_HEAD_SHA: ${{ steps.pr.outputs.head_sha }}
|
||||
CODEX_OUTPUT_PATH: ${{ steps.codex-output.outputs.output_file }}
|
||||
CODEX_CONCLUSION: ${{ steps.run_codex.conclusion }}
|
||||
run: |
|
||||
python .github/scripts/pr_labels.py
|
||||
|
||||
- name: Comment on manual run failure
|
||||
if: ${{ github.event_name == 'workflow_dispatch' && always() }}
|
||||
uses: actions/github-script@ed597411d8f924073f98dfc5c65a23a2325f34cd
|
||||
env:
|
||||
PR_NUMBER: ${{ steps.pr.outputs.pr_number }}
|
||||
JOB_STATUS: ${{ job.status }}
|
||||
RUN_URL: ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }}
|
||||
CODEX_CONCLUSION: ${{ steps.run_codex.conclusion }}
|
||||
with:
|
||||
github-token: ${{ secrets.GITHUB_TOKEN }}
|
||||
script: |
|
||||
const marker = '<!-- pr-labels-manual-run -->';
|
||||
const jobStatus = process.env.JOB_STATUS;
|
||||
if (jobStatus === 'success') {
|
||||
return;
|
||||
}
|
||||
const prNumber = Number(process.env.PR_NUMBER);
|
||||
if (!prNumber) {
|
||||
core.setFailed('Missing PR number for manual run comment.');
|
||||
return;
|
||||
}
|
||||
const body = [
|
||||
marker,
|
||||
'Manual PR labeling failed.',
|
||||
`Job status: ${jobStatus}.`,
|
||||
`Run: ${process.env.RUN_URL}.`,
|
||||
`Codex labeling: ${process.env.CODEX_CONCLUSION}.`,
|
||||
].join('\n');
|
||||
const { data: comments } = await github.rest.issues.listComments({
|
||||
owner: context.repo.owner,
|
||||
repo: context.repo.repo,
|
||||
issue_number: prNumber,
|
||||
per_page: 100,
|
||||
});
|
||||
const existing = comments.find(
|
||||
(comment) =>
|
||||
comment.user?.login === 'github-actions[bot]' &&
|
||||
comment.body?.includes(marker),
|
||||
);
|
||||
if (existing) {
|
||||
await github.rest.issues.updateComment({
|
||||
owner: context.repo.owner,
|
||||
repo: context.repo.repo,
|
||||
comment_id: existing.id,
|
||||
body,
|
||||
});
|
||||
core.info(`Updated existing comment ${existing.id}`);
|
||||
return;
|
||||
}
|
||||
const { data: created } = await github.rest.issues.createComment({
|
||||
owner: context.repo.owner,
|
||||
repo: context.repo.repo,
|
||||
issue_number: prNumber,
|
||||
body,
|
||||
});
|
||||
core.info(`Created comment ${created.id}`);
|
||||
@@ -21,14 +21,15 @@ jobs:
|
||||
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
|
||||
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0
|
||||
- name: Setup uv
|
||||
uses: astral-sh/setup-uv@cec208311dfd045dd5311c1add060b2062131d57
|
||||
uses: astral-sh/setup-uv@fac544c07dec837d0ccb6301d7b5580bf5edae39 # setup-uv v8.1.0; uv 0.11.14
|
||||
with:
|
||||
version: "0.11.14"
|
||||
enable-cache: true
|
||||
- name: Install dependencies
|
||||
run: make sync
|
||||
- name: Build package
|
||||
run: uv build
|
||||
- name: Publish to PyPI
|
||||
uses: pypa/gh-action-pypi-publish@ed0c53931b1dc9bd32cbe73a98c7f6766f8a527e
|
||||
uses: pypa/gh-action-pypi-publish@cef221092ed1bacb1cc03d23a2d87d1d172e277b
|
||||
|
||||
@@ -1,109 +0,0 @@
|
||||
name: Update release PR on main updates
|
||||
|
||||
on:
|
||||
push:
|
||||
branches:
|
||||
- main
|
||||
|
||||
concurrency:
|
||||
group: release-pr-update
|
||||
cancel-in-progress: true
|
||||
|
||||
permissions:
|
||||
contents: write
|
||||
pull-requests: write
|
||||
|
||||
jobs:
|
||||
update-release-pr:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
|
||||
with:
|
||||
fetch-depth: 0
|
||||
- name: Fetch tags
|
||||
run: git fetch origin --tags --prune
|
||||
- name: Configure git
|
||||
run: |
|
||||
git config user.name "github-actions[bot]"
|
||||
git config user.email "github-actions[bot]@users.noreply.github.com"
|
||||
- name: Find release PR
|
||||
id: find
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
base_branch="main"
|
||||
prs_json="$(gh pr list \
|
||||
--base "$base_branch" \
|
||||
--state open \
|
||||
--search "head:release/v" \
|
||||
--limit 200 \
|
||||
--json number,headRefName,isCrossRepository,headRepositoryOwner)"
|
||||
count="$(echo "$prs_json" | jq '[.[] | select(.isCrossRepository == false) | select(.headRefName|startswith("release/v"))] | length')"
|
||||
if [ "$count" -eq 0 ]; then
|
||||
echo "found=false" >> "$GITHUB_OUTPUT"
|
||||
exit 0
|
||||
fi
|
||||
if [ "$count" -gt 1 ]; then
|
||||
echo "Multiple release PRs found; expected a single release PR." >&2
|
||||
exit 1
|
||||
fi
|
||||
number="$(echo "$prs_json" | jq -r '.[] | select(.isCrossRepository == false) | select(.headRefName|startswith("release/v")) | .number')"
|
||||
branch="$(echo "$prs_json" | jq -r '.[] | select(.isCrossRepository == false) | select(.headRefName|startswith("release/v")) | .headRefName')"
|
||||
echo "found=true" >> "$GITHUB_OUTPUT"
|
||||
echo "number=$number" >> "$GITHUB_OUTPUT"
|
||||
echo "branch=$branch" >> "$GITHUB_OUTPUT"
|
||||
- name: Rebase release branch
|
||||
if: steps.find.outputs.found == 'true'
|
||||
env:
|
||||
RELEASE_BRANCH: ${{ steps.find.outputs.branch }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
git fetch origin main "$RELEASE_BRANCH"
|
||||
git checkout -B "$RELEASE_BRANCH" "origin/$RELEASE_BRANCH"
|
||||
git rebase origin/main
|
||||
- name: Prepare Codex output
|
||||
if: steps.find.outputs.found == 'true'
|
||||
id: codex-output
|
||||
run: |
|
||||
set -euo pipefail
|
||||
output_dir=".tmp/codex/outputs"
|
||||
output_file="${output_dir}/release-review.md"
|
||||
mkdir -p "$output_dir"
|
||||
echo "output_file=${output_file}" >> "$GITHUB_OUTPUT"
|
||||
- name: Run Codex release review
|
||||
if: steps.find.outputs.found == 'true'
|
||||
uses: openai/codex-action@c25d10f3f498316d4b2496cc4c6dd58057a7b031
|
||||
with:
|
||||
openai-api-key: ${{ secrets.PROD_OPENAI_API_KEY }}
|
||||
prompt-file: .github/codex/prompts/release-review.md
|
||||
output-file: ${{ steps.codex-output.outputs.output_file }}
|
||||
# Keep the legacy Linux sandbox path until the default bubblewrap path
|
||||
# works reliably on GitHub-hosted Ubuntu runners.
|
||||
codex-args: '["--enable","use_legacy_landlock"]'
|
||||
safety-strategy: drop-sudo
|
||||
sandbox: read-only
|
||||
- name: Update PR body and push
|
||||
if: steps.find.outputs.found == 'true'
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
PR_NUMBER: ${{ steps.find.outputs.number }}
|
||||
RELEASE_BRANCH: ${{ steps.find.outputs.branch }}
|
||||
RELEASE_REVIEW_PATH: ${{ steps.codex-output.outputs.output_file }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
git push --force-with-lease origin "$RELEASE_BRANCH"
|
||||
gh pr edit "$PR_NUMBER" --body-file "$RELEASE_REVIEW_PATH"
|
||||
version="${RELEASE_BRANCH#release/v}"
|
||||
milestone_name="$(python .github/scripts/select-release-milestone.py --version "$version")"
|
||||
if [ -n "$milestone_name" ]; then
|
||||
if ! gh pr edit "$PR_NUMBER" --add-label "project" --milestone "$milestone_name"; then
|
||||
echo "PR label/milestone update failed; continuing without changes." >&2
|
||||
fi
|
||||
else
|
||||
if ! gh pr edit "$PR_NUMBER" --add-label "project"; then
|
||||
echo "PR label update failed; continuing without changes." >&2
|
||||
fi
|
||||
fi
|
||||
@@ -16,13 +16,14 @@ jobs:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
|
||||
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0
|
||||
with:
|
||||
fetch-depth: 0
|
||||
ref: main
|
||||
- name: Setup uv
|
||||
uses: astral-sh/setup-uv@cec208311dfd045dd5311c1add060b2062131d57
|
||||
uses: astral-sh/setup-uv@fac544c07dec837d0ccb6301d7b5580bf5edae39 # setup-uv v8.1.0; uv 0.11.14
|
||||
with:
|
||||
version: "0.11.14"
|
||||
enable-cache: true
|
||||
- name: Fetch tags
|
||||
run: git fetch origin --tags --prune
|
||||
@@ -92,36 +93,11 @@ jobs:
|
||||
fi
|
||||
git commit -m "Bump version to ${RELEASE_VERSION}"
|
||||
git push --set-upstream origin "$branch"
|
||||
- name: Prepare Codex output
|
||||
id: codex-output
|
||||
run: |
|
||||
set -euo pipefail
|
||||
output_dir=".tmp/codex/outputs"
|
||||
output_file="${output_dir}/release-review.md"
|
||||
mkdir -p "$output_dir"
|
||||
echo "output_file=${output_file}" >> "$GITHUB_OUTPUT"
|
||||
- name: Run Codex release review
|
||||
uses: openai/codex-action@c25d10f3f498316d4b2496cc4c6dd58057a7b031
|
||||
with:
|
||||
openai-api-key: ${{ secrets.PROD_OPENAI_API_KEY }}
|
||||
prompt-file: .github/codex/prompts/release-review.md
|
||||
output-file: ${{ steps.codex-output.outputs.output_file }}
|
||||
# Keep the legacy Linux sandbox path until the default bubblewrap path
|
||||
# works reliably on GitHub-hosted Ubuntu runners.
|
||||
codex-args: '["--enable","use_legacy_landlock"]'
|
||||
safety-strategy: drop-sudo
|
||||
sandbox: read-only
|
||||
- name: Build PR body
|
||||
env:
|
||||
RELEASE_REVIEW_PATH: ${{ steps.codex-output.outputs.output_file }}
|
||||
RELEASE_VERSION: ${{ inputs.version }}
|
||||
run: |
|
||||
python - <<'PY'
|
||||
import os
|
||||
import pathlib
|
||||
|
||||
report = pathlib.Path(os.environ["RELEASE_REVIEW_PATH"]).read_text()
|
||||
pathlib.Path("pr-body.md").write_text(report)
|
||||
PY
|
||||
printf 'Release PR for %s.\n\nThe release readiness report will be prepared manually.\n' "$RELEASE_VERSION" > pr-body.md
|
||||
- name: Create or update PR
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
|
||||
@@ -14,6 +14,7 @@ jobs:
|
||||
tag-release:
|
||||
if: >-
|
||||
github.event.pull_request.merged == true &&
|
||||
github.event.pull_request.head.repo.full_name == github.repository &&
|
||||
startsWith(github.event.pull_request.head.ref, 'release/v')
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
@@ -26,12 +27,12 @@ jobs:
|
||||
exit 1
|
||||
fi
|
||||
- name: Checkout merge commit
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
|
||||
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0
|
||||
with:
|
||||
fetch-depth: 0
|
||||
ref: ${{ github.event.pull_request.merge_commit_sha }}
|
||||
- name: Setup Python
|
||||
uses: actions/setup-python@a309ff8b426b58ec0e2a45f0f869d46889d02405
|
||||
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1
|
||||
with:
|
||||
python-version: "3.11"
|
||||
- name: Configure git
|
||||
|
||||
+15
-10
@@ -18,14 +18,15 @@ jobs:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
|
||||
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0
|
||||
- name: Detect code changes
|
||||
id: changes
|
||||
run: ./.github/scripts/detect-changes.sh code "${{ github.event.pull_request.base.sha || github.event.before }}" "${{ github.sha }}"
|
||||
- name: Setup uv
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
uses: astral-sh/setup-uv@cec208311dfd045dd5311c1add060b2062131d57
|
||||
uses: astral-sh/setup-uv@fac544c07dec837d0ccb6301d7b5580bf5edae39 # setup-uv v8.1.0; uv 0.11.14
|
||||
with:
|
||||
version: "0.11.14"
|
||||
enable-cache: true
|
||||
- name: Install dependencies
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
@@ -44,14 +45,15 @@ jobs:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
|
||||
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0
|
||||
- name: Detect code changes
|
||||
id: changes
|
||||
run: ./.github/scripts/detect-changes.sh code "${{ github.event.pull_request.base.sha || github.event.before }}" "${{ github.sha }}"
|
||||
- name: Setup uv
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
uses: astral-sh/setup-uv@cec208311dfd045dd5311c1add060b2062131d57
|
||||
uses: astral-sh/setup-uv@fac544c07dec837d0ccb6301d7b5580bf5edae39 # setup-uv v8.1.0; uv 0.11.14
|
||||
with:
|
||||
version: "0.11.14"
|
||||
enable-cache: true
|
||||
- name: Install dependencies
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
@@ -78,14 +80,15 @@ jobs:
|
||||
OPENAI_API_KEY: fake-for-tests
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
|
||||
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0
|
||||
- name: Detect code changes
|
||||
id: changes
|
||||
run: ./.github/scripts/detect-changes.sh code "${{ github.event.pull_request.base.sha || github.event.before }}" "${{ github.sha }}"
|
||||
- name: Setup uv
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
uses: astral-sh/setup-uv@cec208311dfd045dd5311c1add060b2062131d57
|
||||
uses: astral-sh/setup-uv@fac544c07dec837d0ccb6301d7b5580bf5edae39 # setup-uv v8.1.0; uv 0.11.14
|
||||
with:
|
||||
version: "0.11.14"
|
||||
enable-cache: true
|
||||
python-version: ${{ matrix.python-version }}
|
||||
- name: Install dependencies
|
||||
@@ -110,15 +113,16 @@ jobs:
|
||||
OPENAI_API_KEY: fake-for-tests
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
|
||||
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0
|
||||
- name: Detect code changes
|
||||
id: changes
|
||||
shell: bash
|
||||
run: ./.github/scripts/detect-changes.sh code "${{ github.event.pull_request.base.sha || github.event.before }}" "${{ github.sha }}"
|
||||
- name: Setup uv
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
uses: astral-sh/setup-uv@cec208311dfd045dd5311c1add060b2062131d57
|
||||
uses: astral-sh/setup-uv@fac544c07dec837d0ccb6301d7b5580bf5edae39 # setup-uv v8.1.0; uv 0.11.14
|
||||
with:
|
||||
version: "0.11.14"
|
||||
enable-cache: true
|
||||
python-version: "3.13"
|
||||
- name: Install dependencies
|
||||
@@ -137,14 +141,15 @@ jobs:
|
||||
OPENAI_API_KEY: fake-for-tests
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
|
||||
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0
|
||||
- name: Detect docs changes
|
||||
id: changes
|
||||
run: ./.github/scripts/detect-changes.sh docs "${{ github.event.pull_request.base.sha || github.event.before }}" "${{ github.sha }}"
|
||||
- name: Setup uv
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
uses: astral-sh/setup-uv@cec208311dfd045dd5311c1add060b2062131d57
|
||||
uses: astral-sh/setup-uv@fac544c07dec837d0ccb6301d7b5580bf5edae39 # setup-uv v8.1.0; uv 0.11.14
|
||||
with:
|
||||
version: "0.11.14"
|
||||
enable-cache: true
|
||||
- name: Install dependencies
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
|
||||
@@ -1,88 +0,0 @@
|
||||
name: "Update Translated Docs"
|
||||
|
||||
# This GitHub Actions job automates the process of updating all translated document pages. Please note the following:
|
||||
# 1. The translation results may vary each time; some differences in detail are expected.
|
||||
# 2. When you add a new page to the left-hand menu, **make sure to manually update mkdocs.yml** to include the new item.
|
||||
# 3. If you switch to a different LLM (for example, from o3 to a newer model), be sure to conduct thorough testing before making the switch.
|
||||
|
||||
# To add more languages, you will update the following:
|
||||
# 1. Add '!docs/{lang}/**' to `on.push.paths` in this file
|
||||
# 2. Update mkdocs.yml to have the new language
|
||||
# 3. Update docs/scripts/translate_docs.py to have the new language
|
||||
|
||||
on:
|
||||
push:
|
||||
branches:
|
||||
- main
|
||||
paths:
|
||||
- 'docs/**'
|
||||
- mkdocs.yml
|
||||
- '!docs/ja/**'
|
||||
- '!docs/ko/**'
|
||||
- '!docs/zh/**'
|
||||
workflow_dispatch:
|
||||
inputs:
|
||||
translate_mode:
|
||||
description: "Translation mode"
|
||||
type: choice
|
||||
options:
|
||||
- only-changes
|
||||
- full
|
||||
default: only-changes
|
||||
|
||||
permissions:
|
||||
contents: write
|
||||
pull-requests: write
|
||||
|
||||
jobs:
|
||||
update-docs:
|
||||
if: "!contains(github.event.head_commit.message, 'Update all translated document pages')"
|
||||
name: Build and Push Translated Docs
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 30
|
||||
env:
|
||||
PROD_OPENAI_API_KEY: ${{ secrets.PROD_OPENAI_API_KEY }}
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
|
||||
with:
|
||||
fetch-depth: 0
|
||||
- name: Setup uv
|
||||
uses: astral-sh/setup-uv@cec208311dfd045dd5311c1add060b2062131d57
|
||||
with:
|
||||
enable-cache: true
|
||||
- name: Install dependencies
|
||||
run: make sync
|
||||
- name: Build translated docs
|
||||
run: |
|
||||
mode="${{ inputs.translate_mode || 'only-changes' }}"
|
||||
uv run docs/scripts/translate_docs.py --mode "$mode"
|
||||
uv run mkdocs build
|
||||
|
||||
- name: Commit changes
|
||||
id: commit
|
||||
run: |
|
||||
git config user.name "github-actions[bot]"
|
||||
git config user.email "github-actions[bot]@users.noreply.github.com"
|
||||
git add docs/
|
||||
if git diff --cached --quiet; then
|
||||
echo "No changes to commit"
|
||||
echo "committed=false" >> "$GITHUB_OUTPUT"
|
||||
else
|
||||
git commit -m "Update all translated document pages"
|
||||
echo "committed=true" >> "$GITHUB_OUTPUT"
|
||||
fi
|
||||
|
||||
- name: Create Pull Request
|
||||
if: steps.commit.outputs.committed == 'true'
|
||||
uses: peter-evans/create-pull-request@c0f553fe549906ede9cf27b5156039d195d2ece0
|
||||
with:
|
||||
commit-message: "Update translated document pages"
|
||||
title: "docs: update translated document pages"
|
||||
body: |
|
||||
Automated update of translated documentation.
|
||||
|
||||
Triggered by commit: [${{ github.event.head_commit.id }}](${{ github.server_url }}/${{ github.repository }}/commit/${{ github.event.head_commit.id }}).
|
||||
Message: `${{ github.event.head_commit.message }}`
|
||||
branch: update-translated-docs-${{ github.run_id }}
|
||||
delete-branch: true
|
||||
@@ -143,6 +143,10 @@ cython_debug/
|
||||
# Ruff stuff:
|
||||
.ruff_cache/
|
||||
|
||||
# Example runtime state
|
||||
examples/sandbox/extensions/daytona/usaspending_text2sql/.audit_log.jsonl
|
||||
examples/sandbox/extensions/daytona/usaspending_text2sql/.session_state.json
|
||||
|
||||
# PyPI configuration file
|
||||
.pypirc
|
||||
.aider*
|
||||
@@ -154,3 +158,4 @@ tmp/
|
||||
|
||||
# execplans
|
||||
plans/
|
||||
.vercel
|
||||
|
||||
@@ -36,17 +36,19 @@ Before changing runtime code, exported APIs, external configuration, persisted s
|
||||
|
||||
#### `$pr-draft-summary`
|
||||
|
||||
When a task in this repo finishes with moderate-or-larger code changes, invoke `$pr-draft-summary` in the final handoff to generate the required PR summary block, branch suggestion, title, and draft description. Treat this as the default close-out step after runtime code, tests, examples, build/test configuration, or docs with behavior impact are changed.
|
||||
Before every final response for a task that changed runtime code, tests, examples, build/test configuration, or docs with behavior impact, invoke `$pr-draft-summary` to generate the required PR summary block, branch suggestion, title, and draft description. Determine whether to invoke it from the changed files, not from a subjective assessment of change size.
|
||||
|
||||
Skip `$pr-draft-summary` only for trivial or conversation-only tasks, repo-meta/doc-only tasks without behavior impact, or when the user explicitly says not to include the PR draft block.
|
||||
|
||||
Producing the PR draft block is part of the local final handoff. It is required for eligible local-only or uncommitted changes and does not authorize creating a branch, committing, pushing, or opening a pull request.
|
||||
|
||||
### ExecPlans
|
||||
|
||||
Call out compatibility risk early in your plan only when the change affects behavior shipped in the latest release tag or a released or explicitly supported durable external state boundary, and confirm the approach before implementing changes that could impact users.
|
||||
|
||||
Use an ExecPlan when work is multi-step, spans several files, involves new features or refactors, or is likely to take more than about an hour. Start with the template and rules in `PLANS.md`, keep milestones and living sections (Progress, Surprises & Discoveries, Decision Log, Outcomes & Retrospective) up to date as you execute, and rewrite the plan if scope shifts. Call out compatibility risk only when the plan changes behavior shipped in the latest release tag or a released or explicitly supported durable external state boundary. Do not treat branch-local interface churn or unreleased post-tag changes on `main` as breaking by default; prefer direct replacement over compatibility layers in those cases, and renumber or squash unreleased persisted schemas before release when the intermediate snapshots are intentionally unsupported. If you intentionally skip an ExecPlan for a complex task, note why in your response so reviewers understand the choice.
|
||||
|
||||
### Public API Positional Compatibility
|
||||
### Public API Compatibility
|
||||
|
||||
Treat the parameter and dataclass field order of exported runtime APIs as a compatibility contract.
|
||||
|
||||
@@ -54,6 +56,17 @@ Treat the parameter and dataclass field order of exported runtime APIs as a comp
|
||||
- When adding a new optional public field/parameter, append it to the end whenever possible and keep old fields in the same order.
|
||||
- If reordering is unavoidable, add an explicit compatibility layer and regression tests that exercise the old positional call pattern.
|
||||
- Prefer keyword arguments at call sites to reduce accidental breakage, but do not rely on this to justify breaking positional compatibility for public APIs.
|
||||
- Treat intended import paths and `__all__` membership as compatibility contracts. When adding or moving a public symbol, update the owning module, intended top-level or subpackage re-exports, and an import regression test. Keep top-level imports free of optional-dependency failures and runtime side effects; use lazy exports when needed.
|
||||
|
||||
### Platform, Docs, and Security Review
|
||||
|
||||
- Documentation is published to the live site, so coordinate SDK behavior changes and docs carefully. If docs describe behavior that is not released yet, either delay the docs change until the SDK release is available or split it into a follow-up PR.
|
||||
- Treat runnable docs snippets as API compatibility checks. Before adding OpenAI API, provider, Responses, Realtime, WebSocket, or SDK constructor examples, verify the shown arguments and call shape against the actual implementation.
|
||||
- Do not let untrusted sandbox manifests opt themselves out of host filesystem or base-directory boundaries. Escape hatches for local source materialization must be controlled by trusted application code at the call site, not by serialized manifest data.
|
||||
- When documenting sandbox or security grants, verify the actual implementation path enforces the grant or boundary. Do not claim a grant applies to `LocalDir`, `LocalFile`, archive extraction, or other materialization paths unless those paths actually consult it.
|
||||
- When redacting OpenAI tool, MCP, model, or provider payloads, consider traceback display, exception chaining, `__context__`, logs, and telemetry. Suppressing display with `raise ... from None` is not enough if the original exception object still carries sensitive input data.
|
||||
- For OpenAI platform or SDK-specific docs changes, prefer `$openai-knowledge` for authoritative platform behavior and inspect the local code path for SDK behavior. Do not rely on generic API assumptions when documenting Responses, Chat Completions, Realtime, tools, MCP, or provider adapters.
|
||||
- For Realtime tracing changes, read [Realtime tracing architecture](.agents/references/realtime-tracing.md) before proposing SDK spans. Realtime API server traces and Agents SDK client traces are separate; `group_id` can correlate them but does not create a shared trace hierarchy.
|
||||
|
||||
## Project Structure Guide
|
||||
|
||||
@@ -72,26 +85,28 @@ The OpenAI Agents Python repository provides the Python Agents SDK, examples, an
|
||||
- `Makefile`: Common developer commands.
|
||||
- `pyproject.toml`, `uv.lock`: Python dependencies and tool configuration.
|
||||
- `.github/PULL_REQUEST_TEMPLATE/pull_request_template.md`: Pull request template to use when opening PRs.
|
||||
- `.agents/references/`: Durable SDK maintainer architecture references. Start with [the reference map](.agents/references/README.md) and open only the files relevant to the affected runtime boundary.
|
||||
- `site/`: Built documentation output.
|
||||
|
||||
### Agents Core Runtime Guidelines
|
||||
|
||||
- For `Agent` fields, cloning, dynamic instructions, enabled tools or handoffs, output schemas, run context wrappers, usage aggregation, or public-versus-internal agent identity, read [Agent definition and run context](.agents/references/agent-definition-and-run-context.md).
|
||||
- `src/agents/run.py` is the runtime entrypoint (`Runner`, `AgentRunner`). Keep it focused on orchestration and public flow control. Put new runtime logic under `src/agents/run_internal/` and import it into `run.py`.
|
||||
- When `run.py` grows, refactor helpers into `run_internal/` modules (for example `run_loop.py`, `turn_resolution.py`, `tool_execution.py`, `session_persistence.py`) and leave only wiring and composition in `run.py`.
|
||||
- Keep streaming and non-streaming paths behaviorally aligned. Changes to `run_internal/run_loop.py` (`run_single_turn`, `run_single_turn_streamed`, `get_new_response`, `start_streaming`) should be mirrored, and any new streaming item types must be reflected in `src/agents/stream_events.py`.
|
||||
- Input guardrails run only on the first turn and only for the starting agent. Resuming an interruption from `RunState` must not increment the turn counter; only actual model calls advance turns.
|
||||
- Server-managed conversation (`conversation_id`, `previous_response_id`, `auto_previous_response_id`) uses `OpenAIServerConversationTracker` in `run_internal/oai_conversation.py`. Only deltas should be sent. If `call_model_input_filter` is used, it must return `ModelInputData` with a list input and the tracker must be updated with the filtered input (`mark_input_as_sent`). Session persistence is disabled when server-managed conversation is active.
|
||||
- Adding new tool/output/approval item types requires coordinated updates across:
|
||||
- `src/agents/items.py` (RunItem types and conversions)
|
||||
- `src/agents/run_internal/run_steps.py` (ProcessedResponse and tool run structs)
|
||||
- `src/agents/run_internal/turn_resolution.py` (model output processing, run item extraction)
|
||||
- `src/agents/run_internal/tool_execution.py` and `src/agents/run_internal/tool_planning.py`
|
||||
- `src/agents/run_internal/items.py` (normalization, dedupe, approval filtering)
|
||||
- `src/agents/stream_events.py` (stream event names)
|
||||
- `src/agents/run_state.py` (RunState serialization/deserialization)
|
||||
- `src/agents/run_internal/session_persistence.py` (session save/rewind)
|
||||
- If the serialized RunState shape changes, update `CURRENT_SCHEMA_VERSION` in `src/agents/run_state.py` and the related serialization/deserialization logic. Keep released schema versions readable, and feel free to renumber or squash unreleased schema versions before release when those intermediate snapshots are intentionally unsupported.
|
||||
- When bumping `CURRENT_SCHEMA_VERSION`, also add or update the matching entry in `SCHEMA_VERSION_SUMMARIES` in `src/agents/run_state.py` so every supported version keeps a short historical note describing what changed in that schema.
|
||||
- For turn accounting, guardrail ordering, handoffs, interruptions, cancellation, hooks, or streaming behavior, read [Runner lifecycle](.agents/references/runner-lifecycle.md). Keep streaming and non-streaming paths behaviorally aligned.
|
||||
- For new model output, tool call, approval, or run item variants, read [Run item lifecycle](.agents/references/run-item-lifecycle.md) and update every applicable processing, event, replay, persistence, tracing, and serialization surface.
|
||||
- For function-tool parameter schemas, `Annotated` or `Field` metadata, strict JSON schema conversion, or structured output schemas, read [Function and output schema](.agents/references/function-and-output-schema.md).
|
||||
- For function-tool naming, namespacing, lookup, approvals, tracing, or call-ID changes, read [Tool identity and routing](.agents/references/tool-identity.md) and use the canonical helpers in `src/agents/_tool_identity.py` instead of adding local normalization rules.
|
||||
- For function-tool planning, approval ordering, tool guardrails, concurrency, cancellation, timeouts, hooks, or failure conversion, read [Tool execution lifecycle](.agents/references/tool-execution-lifecycle.md).
|
||||
- For local MCP connection ownership, `MCPServerManager`, request serialization, tool caching or filtering, transport retries, cancellation, or cleanup, read [Local MCP server lifecycle](.agents/references/local-mcp-server-lifecycle.md).
|
||||
- For trace or span context, processors, export, flush, shutdown, sensitive data, or resumed trace state, read [Tracing lifecycle](.agents/references/tracing-lifecycle.md).
|
||||
- For `RealtimeSession` lifecycle, background-task, handoff, listener, connection, or cleanup changes, read [Realtime session lifecycle](.agents/references/realtime-session-lifecycle.md) and verify both normal and failure-path resource ownership.
|
||||
- For `VoicePipeline`, streamed audio input, STT session ownership, TTS task ordering, voice lifecycle events, PCM framing, or voice tracing changes, read [Voice pipeline lifecycle](.agents/references/voice-pipeline-lifecycle.md).
|
||||
- For server-managed conversation (`conversation_id`, `previous_response_id`, `auto_previous_response_id`), read [Conversation state ownership](.agents/references/conversation-state-ownership.md) before changing continuation, filtering, retry, compaction, handoffs, or resume behavior.
|
||||
- For client-managed session input, per-turn saves, retry rewind, backend atomicity, or compaction replacement, read [Session persistence](.agents/references/session-persistence.md).
|
||||
- For model resolution, `ModelSettings`, provider adapters, Responses versus Chat Completions capabilities, request conversion, terminal events, transport reuse, or model retries, read [Model and provider boundaries](.agents/references/model-provider-boundaries.md).
|
||||
- If the serialized `RunState` shape changes, read [RunState schema and resume boundary](.agents/references/runstate-schema.md) and follow its release-boundary, schema-version, backward-read, and regression-test rules.
|
||||
- For sandbox session ownership, agent preparation, manifests, host-path materialization, snapshots, resume state, or cleanup, read [Sandbox runtime boundary](.agents/references/sandbox-runtime-boundary.md).
|
||||
|
||||
## Operation Guide
|
||||
|
||||
@@ -116,7 +131,7 @@ The OpenAI Agents Python repository provides the Python Agents SDK, examples, an
|
||||
```
|
||||
6. When `$code-change-verification` applies, run it to execute the full verification stack before marking work complete.
|
||||
7. Commit with concise, imperative messages; keep commits small and focused, then open a pull request.
|
||||
8. When reporting code changes as complete (after substantial code work), invoke `$pr-draft-summary` as the final handoff step unless the task falls under the documented skip cases.
|
||||
8. Before reporting eligible code changes as complete, invoke `$pr-draft-summary` as the final handoff step unless the task falls under the documented skip cases. Do not omit it based on perceived change size or because the work remains local or uncommitted.
|
||||
|
||||
### Testing & Automated Checks
|
||||
|
||||
|
||||
@@ -17,7 +17,7 @@ The OpenAI Agents SDK is a lightweight yet powerful framework for building multi
|
||||
1. [**Human in the loop**](https://openai.github.io/openai-agents-python/human_in_the_loop/): Built-in mechanisms for involving humans across agent runs
|
||||
1. [**Sessions**](https://openai.github.io/openai-agents-python/sessions/): Automatic conversation history management across agent runs
|
||||
1. [**Tracing**](https://openai.github.io/openai-agents-python/tracing/): Built-in tracking of agent runs, allowing you to view, debug and optimize your workflows
|
||||
1. [**Realtime Agents**](https://openai.github.io/openai-agents-python/realtime/quickstart/): Build powerful voice agents with `gpt-realtime-1.5` and full agent features
|
||||
1. [**Realtime Agents**](https://openai.github.io/openai-agents-python/realtime/quickstart/): Build powerful voice agents with `gpt-realtime-2` and full agent features
|
||||
|
||||
Explore the [examples](https://github.com/openai/openai-agents-python/tree/main/examples) directory to see the SDK in action, and read our [documentation](https://openai.github.io/openai-agents-python/) for more details.
|
||||
|
||||
@@ -60,11 +60,7 @@ from agents.sandbox.sandboxes import UnixLocalSandboxClient
|
||||
agent = SandboxAgent(
|
||||
name="Workspace Assistant",
|
||||
instructions="Inspect the sandbox workspace before answering.",
|
||||
default_manifest=Manifest(
|
||||
entries={
|
||||
"repo": GitRepo(repo="openai/openai-agents-python", ref="main"),
|
||||
}
|
||||
),
|
||||
default_manifest=Manifest(entries={"repo": GitRepo(repo="openai/openai-agents-python", ref="main")}),
|
||||
)
|
||||
|
||||
result = Runner.run_sync(
|
||||
@@ -75,11 +71,28 @@ result = Runner.run_sync(
|
||||
)
|
||||
print(result.final_output)
|
||||
|
||||
# This project provides a Python SDK for building multi-agent workflows.
|
||||
# Output: "This project provides a Python SDK for building multi-agent workflows."
|
||||
```
|
||||
|
||||
(_If running this, ensure you set the `OPENAI_API_KEY` environment variable_)
|
||||
|
||||
## Run an agent without a sandbox
|
||||
|
||||
You can still use a regular `Agent` when your workflow does not need a filesystem workspace or sandbox lifecycle.
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
|
||||
agent = Agent(name="Assistant", instructions="You are a helpful assistant")
|
||||
|
||||
result = Runner.run_sync(agent, "Write a haiku about recursion in programming.")
|
||||
print(result.final_output)
|
||||
|
||||
# Code within the code,
|
||||
# Functions calling themselves,
|
||||
# Infinite loop's dance.
|
||||
```
|
||||
|
||||
(_For Jupyter notebook users, see [hello_world_jupyter.ipynb](https://github.com/openai/openai-agents-python/blob/main/examples/basic/hello_world_jupyter.ipynb)_)
|
||||
|
||||
Explore the [examples](https://github.com/openai/openai-agents-python/tree/main/examples) directory to see the SDK in action, and read our [documentation](https://openai.github.io/openai-agents-python/) for more details.
|
||||
|
||||
@@ -0,0 +1,5 @@
|
||||
# Security Policy
|
||||
|
||||
For a more in-depth look at our security policy, please check out our [Coordinated Vulnerability Disclosure Policy](https://openai.com/security/disclosure/#:~:text=Disclosure%20Policy,-Security%20is%20essential&text=OpenAI%27s%20coordinated%20vulnerability%20disclosure%20policy,expect%20from%20us%20in%20return.).
|
||||
|
||||
Our PGP key can located [at this address.](https://cdn.openai.com/security.txt)
|
||||
+3
-4
@@ -28,7 +28,7 @@ The most common properties of an agent are:
|
||||
| Property | Required | Description |
|
||||
| --- | --- | --- |
|
||||
| `name` | yes | Human-readable agent name. |
|
||||
| `instructions` | yes | System prompt or dynamic instructions callback. See [Dynamic instructions](#dynamic-instructions). |
|
||||
| `instructions` | no | System prompt or dynamic instructions callback. Strongly recommended. See [Dynamic instructions](#dynamic-instructions). |
|
||||
| `prompt` | no | OpenAI Responses API prompt configuration. Accepts a static prompt object or a function. See [Prompt templates](#prompt-templates). |
|
||||
| `handoff_description` | no | Short description exposed when this agent is offered as a handoff target. |
|
||||
| `handoffs` | no | Delegate the conversation to specialist agents. See [handoffs](handoffs.md). |
|
||||
@@ -261,8 +261,7 @@ Typical hook timing:
|
||||
|
||||
- `on_agent_start` / `on_agent_end`: when a specific agent begins or finishes producing a final output.
|
||||
- `on_llm_start` / `on_llm_end`: immediately around each model call.
|
||||
- `on_tool_start` / `on_tool_end`: around each local tool invocation.
|
||||
For function tools, the hook `context` is typically a `ToolContext`, so you can inspect tool-call metadata such as `tool_call_id`.
|
||||
- `on_tool_start` / `on_tool_end`: around each local tool invocation. For function tools, the hook `context` is typically a `ToolContext`, so you can inspect tool-call metadata such as `tool_call_id`.
|
||||
- `on_handoff`: when control moves from one agent to another.
|
||||
|
||||
Use `RunHooks` when you want a single observer for the whole workflow, and `AgentHooks` when one agent needs custom side effects.
|
||||
@@ -301,7 +300,7 @@ By using the `clone()` method on an agent, you can duplicate an Agent, and optio
|
||||
pirate_agent = Agent(
|
||||
name="Pirate",
|
||||
instructions="Write like a pirate",
|
||||
model="gpt-5.4",
|
||||
model="gpt-5.5",
|
||||
)
|
||||
|
||||
robot_agent = pirate_agent.clone(
|
||||
|
||||
@@ -47,6 +47,38 @@ from agents import set_default_openai_api
|
||||
set_default_openai_api("chat_completions")
|
||||
```
|
||||
|
||||
## OpenAI provider defaults
|
||||
|
||||
OpenAI-backed providers also read SDK-wide defaults when they resolve model names. Use [`set_default_openai_responses_transport()`][agents.set_default_openai_responses_transport] to make OpenAI Responses models use websocket transport by default:
|
||||
|
||||
```python
|
||||
from agents import set_default_openai_responses_transport
|
||||
|
||||
set_default_openai_responses_transport("websocket")
|
||||
```
|
||||
|
||||
This affects OpenAI Responses models resolved by the default OpenAI provider. For provider-level setup, connection reuse, keepalive options, and custom websocket endpoints, see [Responses WebSocket transport](models/index.md#responses-websocket-transport).
|
||||
|
||||
If your OpenAI setup expects provider-level agent registration metadata, configure a default harness ID once at startup:
|
||||
|
||||
```python
|
||||
from agents import set_default_openai_harness
|
||||
|
||||
set_default_openai_harness("your-harness-id")
|
||||
```
|
||||
|
||||
You can also pass the full registration object:
|
||||
|
||||
```python
|
||||
from agents import OpenAIAgentRegistrationConfig, set_default_openai_agent_registration
|
||||
|
||||
set_default_openai_agent_registration(
|
||||
OpenAIAgentRegistrationConfig(harness_id="your-harness-id")
|
||||
)
|
||||
```
|
||||
|
||||
If no SDK default is set, OpenAI-backed providers fall back to the `OPENAI_AGENT_HARNESS_ID` environment variable. When a harness ID is configured, the SDK adds it to trace metadata as `agent_harness_id` unless that key is already present in `RunConfig.trace_metadata`.
|
||||
|
||||
## Tracing
|
||||
|
||||
Tracing is enabled by default. By default it uses the same OpenAI API key as your model requests from the section above (that is, the environment variable or the default key you set). You can specifically set the API key used for tracing by using the [`set_tracing_export_api_key`][agents.set_tracing_export_api_key] function.
|
||||
|
||||
+22
-28
@@ -4,8 +4,7 @@ Check out a variety of sample implementations of the SDK in the examples section
|
||||
|
||||
## Categories
|
||||
|
||||
- **[agent_patterns](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns):**
|
||||
Examples in this category illustrate common agent design patterns, such as
|
||||
- **[agent_patterns](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns):** Examples in this category illustrate common agent design patterns, such as
|
||||
|
||||
- Deterministic workflows
|
||||
- Agents as tools
|
||||
@@ -22,8 +21,7 @@ Check out a variety of sample implementations of the SDK in the examples section
|
||||
- Human-in-the-loop with streaming (`examples/agent_patterns/human_in_the_loop_stream.py`)
|
||||
- Custom rejection messages for approval flows (`examples/agent_patterns/human_in_the_loop_custom_rejection.py`)
|
||||
|
||||
- **[basic](https://github.com/openai/openai-agents-python/tree/main/examples/basic):**
|
||||
These examples showcase foundational capabilities of the SDK, such as
|
||||
- **[basic](https://github.com/openai/openai-agents-python/tree/main/examples/basic):** These examples showcase foundational capabilities of the SDK, such as
|
||||
|
||||
- Hello world examples (Default model, GPT-5, open-weight model)
|
||||
- Agent lifecycle management
|
||||
@@ -42,28 +40,23 @@ Check out a variety of sample implementations of the SDK in the examples section
|
||||
- Non-strict output types
|
||||
- Previous response ID usage
|
||||
|
||||
- **[customer_service](https://github.com/openai/openai-agents-python/tree/main/examples/customer_service):**
|
||||
Example customer service system for an airline.
|
||||
- **[customer_service](https://github.com/openai/openai-agents-python/tree/main/examples/customer_service):** Example customer service system for an airline.
|
||||
|
||||
- **[financial_research_agent](https://github.com/openai/openai-agents-python/tree/main/examples/financial_research_agent):**
|
||||
A financial research agent that demonstrates structured research workflows with agents and tools for financial data analysis.
|
||||
- **[financial_research_agent](https://github.com/openai/openai-agents-python/tree/main/examples/financial_research_agent):** A financial research agent that demonstrates structured research workflows with agents and tools for financial data analysis.
|
||||
|
||||
- **[handoffs](https://github.com/openai/openai-agents-python/tree/main/examples/handoffs):**
|
||||
Practical examples of agent handoffs with message filtering, including:
|
||||
- **[handoffs](https://github.com/openai/openai-agents-python/tree/main/examples/handoffs):** Practical examples of agent handoffs with message filtering, including:
|
||||
|
||||
- Message filter example (`examples/handoffs/message_filter.py`)
|
||||
- Message filter with streaming (`examples/handoffs/message_filter_streaming.py`)
|
||||
|
||||
- **[hosted_mcp](https://github.com/openai/openai-agents-python/tree/main/examples/hosted_mcp):**
|
||||
Examples demonstrating how to use hosted MCP (Model Context Protocol) with the OpenAI Responses API, including:
|
||||
- **[hosted_mcp](https://github.com/openai/openai-agents-python/tree/main/examples/hosted_mcp):** Examples demonstrating how to use hosted MCP (Model Context Protocol) with the OpenAI Responses API, including:
|
||||
|
||||
- Simple hosted MCP without approval (`examples/hosted_mcp/simple.py`)
|
||||
- MCP connectors such as Google Calendar (`examples/hosted_mcp/connectors.py`)
|
||||
- Human-in-the-loop with interruption-based approvals (`examples/hosted_mcp/human_in_the_loop.py`)
|
||||
- On-approval callback for MCP tool calls (`examples/hosted_mcp/on_approval.py`)
|
||||
|
||||
- **[mcp](https://github.com/openai/openai-agents-python/tree/main/examples/mcp):**
|
||||
Learn how to build agents with MCP (Model Context Protocol), including:
|
||||
- **[mcp](https://github.com/openai/openai-agents-python/tree/main/examples/mcp):** Learn how to build agents with MCP (Model Context Protocol), including:
|
||||
|
||||
- Filesystem examples
|
||||
- Git examples
|
||||
@@ -77,8 +70,7 @@ Check out a variety of sample implementations of the SDK in the examples section
|
||||
- MCPServerManager with FastAPI (`examples/mcp/manager_example`)
|
||||
- MCP tool filtering (`examples/mcp/tool_filter_example`)
|
||||
|
||||
- **[memory](https://github.com/openai/openai-agents-python/tree/main/examples/memory):**
|
||||
Examples of different memory implementations for agents, including:
|
||||
- **[memory](https://github.com/openai/openai-agents-python/tree/main/examples/memory):** Examples of different memory implementations for agents, including:
|
||||
|
||||
- SQLite session storage
|
||||
- Advanced SQLite session storage
|
||||
@@ -95,29 +87,32 @@ Check out a variety of sample implementations of the SDK in the examples section
|
||||
- OpenAI Conversations session with human-in-the-loop (`examples/memory/openai_session_hitl_example.py`)
|
||||
- HITL approval/rejection scenario across sessions (`examples/memory/hitl_session_scenario.py`)
|
||||
|
||||
- **[model_providers](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers):**
|
||||
Explore how to use non-OpenAI models with the SDK, including custom providers and third-party adapters.
|
||||
- **[model_providers](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers):** Explore how to use non-OpenAI models with the SDK, including custom providers and third-party adapters.
|
||||
|
||||
- **[realtime](https://github.com/openai/openai-agents-python/tree/main/examples/realtime):**
|
||||
Examples showing how to build real-time experiences using the SDK, including:
|
||||
- **[realtime](https://github.com/openai/openai-agents-python/tree/main/examples/realtime):** Examples showing how to build real-time experiences using the SDK, including:
|
||||
|
||||
- Web application patterns with structured text and image messages
|
||||
- Command-line audio loops and playback handling
|
||||
- Twilio Media Streams integration over WebSocket
|
||||
- Twilio SIP integration using Realtime Calls API attach flows
|
||||
|
||||
- **[reasoning_content](https://github.com/openai/openai-agents-python/tree/main/examples/reasoning_content):**
|
||||
Examples demonstrating how to work with reasoning content, including:
|
||||
- **[reasoning_content](https://github.com/openai/openai-agents-python/tree/main/examples/reasoning_content):** Examples demonstrating how to work with reasoning content, including:
|
||||
|
||||
- Reasoning content with the Runner API, streaming and non-streaming (`examples/reasoning_content/runner_example.py`)
|
||||
- Reasoning content with OSS models via OpenRouter (`examples/reasoning_content/gpt_oss_stream.py`)
|
||||
- Basic reasoning content example (`examples/reasoning_content/main.py`)
|
||||
|
||||
- **[research_bot](https://github.com/openai/openai-agents-python/tree/main/examples/research_bot):**
|
||||
Simple deep research clone that demonstrates complex multi-agent research workflows.
|
||||
- **[research_bot](https://github.com/openai/openai-agents-python/tree/main/examples/research_bot):** Simple deep research clone that demonstrates complex multi-agent research workflows.
|
||||
|
||||
- **[tools](https://github.com/openai/openai-agents-python/tree/main/examples/tools):**
|
||||
Learn how to implement OAI hosted tools and experimental Codex tooling such as:
|
||||
- **[sandbox](https://github.com/openai/openai-agents-python/tree/main/examples/sandbox):** Examples for running agents in isolated workspaces, including:
|
||||
|
||||
- Basic sandbox agent setup (`examples/sandbox/basic.py`)
|
||||
- Unix-local and Docker sandbox lifecycle examples
|
||||
- Sandbox-backed handoffs (`examples/sandbox/handoffs.py`)
|
||||
- Sandbox memory and snapshot resume (`examples/sandbox/memory.py`)
|
||||
- Sandbox agents exposed as tools (`examples/sandbox/sandbox_agents_as_tools.py`)
|
||||
|
||||
- **[tools](https://github.com/openai/openai-agents-python/tree/main/examples/tools):** Learn how to implement OAI hosted tools and experimental Codex tooling such as:
|
||||
|
||||
- Web search and web search with filters
|
||||
- File search
|
||||
@@ -134,5 +129,4 @@ Check out a variety of sample implementations of the SDK in the examples section
|
||||
- Experimental Codex tool workflows (`examples/tools/codex.py`)
|
||||
- Experimental Codex same-thread workflows (`examples/tools/codex_same_thread.py`)
|
||||
|
||||
- **[voice](https://github.com/openai/openai-agents-python/tree/main/examples/voice):**
|
||||
See examples of voice agents, using our TTS and STT models, including streamed voice examples.
|
||||
- **[voice](https://github.com/openai/openai-agents-python/tree/main/examples/voice):** See examples of voice agents, using our TTS and STT models, including streamed voice examples.
|
||||
|
||||
@@ -57,6 +57,7 @@ Tool guardrails wrap **function tools** and let you validate or block tool calls
|
||||
|
||||
- Input tool guardrails run before the tool executes and can skip the call, replace the output with a message, or raise a tripwire.
|
||||
- Output tool guardrails run after the tool executes and can replace the output or raise a tripwire.
|
||||
- If a function tool requires approval, input tool guardrails normally run after approval and immediately before execution. Set [`RunConfig.tool_execution`][agents.run.RunConfig.tool_execution] to [`ToolExecutionConfig(pre_approval_tool_input_guardrails=True)`][agents.run.ToolExecutionConfig] when you want those input checks to run before the pending approval interruption is emitted. Calls that pass this pre-approval check are still checked again after approval before the tool executes.
|
||||
- Tool guardrails apply only to function tools created with [`function_tool`][agents.tool.function_tool]. Handoffs run through the SDK's handoff pipeline rather than the normal function-tool pipeline, so tool guardrails do not apply to the handoff call itself. Hosted tools (`WebSearchTool`, `FileSearchTool`, `HostedMCPTool`, `CodeInterpreterTool`, `ImageGenerationTool`) and built-in execution tools (`ComputerTool`, `ShellTool`, `ApplyPatchTool`, `LocalShellTool`) also do not use this guardrail pipeline, and [`Agent.as_tool()`][agents.agent.Agent.as_tool] does not currently expose tool-guardrail options directly.
|
||||
|
||||
See the code snippet below for details.
|
||||
|
||||
+1
-1
@@ -61,7 +61,7 @@ handoff_obj = handoff(
|
||||
|
||||
## Handoff inputs
|
||||
|
||||
In certain situations, you want the LLM to provide some data when it calls a handoff. For example, imagine a handoff to an "Escalation agent". You might want a reason to be provided, so you can log it.
|
||||
In certain situations, you want the LLM to provide some data when it calls a handoff. For example, imagine a handoff to an "Escalation agent". You might want the model to provide a reason so you can log it.
|
||||
|
||||
```python
|
||||
from pydantic import BaseModel
|
||||
|
||||
@@ -186,19 +186,11 @@ Useful serialization options:
|
||||
|
||||
- `context_serializer`: Customize how non-mapping context objects are serialized.
|
||||
- `context_deserializer`: Rebuild non-mapping context objects when loading state with `RunState.from_json(...)` or `RunState.from_string(...)`.
|
||||
- `strict_context=True`: Fail serialization or deserialization unless the context is already a
|
||||
mapping or you provide the appropriate serializer/deserializer.
|
||||
- `context_override`: Replace the serialized context when loading state. This is useful when you
|
||||
do not want to restore the original context object, but it does not remove that context from an
|
||||
already serialized payload.
|
||||
- `include_tracing_api_key=True`: Include the tracing API key in the serialized trace payload
|
||||
when you need resumed work to keep exporting traces with the same credentials.
|
||||
- `strict_context=True`: Fail serialization or deserialization unless the context is already a mapping or you provide the appropriate serializer/deserializer.
|
||||
- `context_override`: Replace the serialized context when loading state. This is useful when you do not want to restore the original context object, but it does not remove that context from an already serialized payload.
|
||||
- `include_tracing_api_key=True`: Include the tracing API key in the serialized trace payload when you need resumed work to keep exporting traces with the same credentials.
|
||||
|
||||
Serialized run state includes your app context plus SDK-managed runtime metadata such as approvals,
|
||||
usage, serialized `tool_input`, nested agent-as-tool resumptions, trace metadata, and server-managed
|
||||
conversation settings. If you plan to store or transmit serialized state, treat
|
||||
`RunContextWrapper.context` as persisted data and avoid placing secrets there unless you
|
||||
intentionally want them to travel with the state.
|
||||
Serialized run state includes your app context plus SDK-managed runtime metadata such as approvals, usage, serialized `tool_input`, nested agent-as-tool resumptions, trace metadata, and server-managed conversation settings. If you plan to store or transmit serialized state, treat `RunContextWrapper.context` as persisted data and avoid placing secrets there unless you intentionally want them to travel with the state.
|
||||
|
||||
## Versioning pending tasks
|
||||
|
||||
|
||||
+2
-2
@@ -27,7 +27,7 @@ Here are the main features of the SDK:
|
||||
- **Sessions**: A persistent memory layer for maintaining working context within an agent loop.
|
||||
- **Human in the loop**: Built-in mechanisms for involving humans across agent runs.
|
||||
- **Tracing**: Built-in tracing for visualizing, debugging, and monitoring workflows, with support for the OpenAI suite of evaluation, fine-tuning, and distillation tools.
|
||||
- **Realtime Agents**: Build powerful voice agents with `gpt-realtime-1.5`, automatic interruption detection, context management, guardrails, and more.
|
||||
- **Realtime Agents**: Build powerful voice agents with `gpt-realtime-2`, automatic interruption detection, context management, guardrails, and more.
|
||||
|
||||
## Agents SDK or Responses API?
|
||||
|
||||
@@ -93,5 +93,5 @@ Use this table when you know the job you want to do, but not which page explains
|
||||
| Keep memory across turns | [Running agents](running_agents.md#choose-a-memory-strategy) and [Sessions](sessions/index.md) |
|
||||
| Use OpenAI models, websocket transport, or non-OpenAI providers | [Models](models/index.md) |
|
||||
| Review outputs, run items, interruptions, and resume state | [Results](results.md) |
|
||||
| Build a low-latency voice agent with `gpt-realtime-1.5` | [Realtime agents quickstart](realtime/quickstart.md) and [Realtime transport](realtime/transport.md) |
|
||||
| Build a low-latency voice agent with `gpt-realtime-2` | [Realtime agents quickstart](realtime/quickstart.md) and [Realtime transport](realtime/transport.md) |
|
||||
| Build a speech-to-text / agent / text-to-speech pipeline | [Voice pipeline quickstart](voice/quickstart.md) |
|
||||
|
||||
+77
-78
@@ -4,26 +4,26 @@ search:
|
||||
---
|
||||
# エージェント
|
||||
|
||||
エージェントは、アプリ内の中核的な基本コンポーネントです。エージェントは、instructions、tools、およびハンドオフ、ガードレール、structured outputs などの任意の実行時動作で構成された大規模言語モデル (LLM) です。
|
||||
エージェントは、アプリ内の中核的な構成要素です。エージェントとは、設定として instructions と tools に加え、ハンドオフ、ガードレール、structured outputs など任意のランタイム動作を持つ大規模言語モデル (LLM) です。
|
||||
|
||||
このページは、単一のプレーンな `Agent` を定義またはカスタマイズしたい場合に使用します。複数のエージェントがどのように連携すべきかを決める場合は、[Agent orchestration](multi_agent.md) をお読みください。エージェントを、manifest で定義されたファイルと sandbox ネイティブ機能を備えた分離ワークスペース内で実行する必要がある場合は、[Sandbox agent concepts](sandbox/guide.md) をお読みください。
|
||||
このページは、単一のプレーンな `Agent` を定義またはカスタマイズしたい場合に使用します。複数のエージェントをどのように連携させるかを決める場合は、[エージェントオーケストレーション](multi_agent.md) を参照してください。エージェントを、マニフェストで定義されたファイルやサンドボックスネイティブの機能を持つ隔離されたワークスペース内で実行する必要がある場合は、[サンドボックスエージェントの概念](sandbox/guide.md) を参照してください。
|
||||
|
||||
SDK は、OpenAI モデルではデフォルトで Responses API を使用しますが、ここでの違いはオーケストレーションです。`Agent` と `Runner` により、SDK がターン、tools、ガードレール、ハンドオフ、セッションを管理します。このループを自分で管理したい場合は、代わりに Responses API を直接使用してください。
|
||||
この SDK は、OpenAI モデルではデフォルトで Responses API を使用しますが、ここでの違いはオーケストレーションです。`Agent` と `Runner` により、SDK がターン、ツール、ガードレール、ハンドオフ、セッションを管理できます。そのループを自分で管理したい場合は、代わりに Responses API を直接使用してください。
|
||||
|
||||
## 次のガイドの選択
|
||||
|
||||
このページをエージェント定義のハブとして使用してください。次に必要な判断に合う隣接ガイドへ移動できます。
|
||||
このページは、エージェント定義のハブとして使用してください。次に判断する内容に合った隣接ガイドへ進んでください。
|
||||
|
||||
| 次のことをしたい場合 | 次に読むもの |
|
||||
| やりたいこと | 次に読むもの |
|
||||
| --- | --- |
|
||||
| モデルまたはプロバイダー設定を選ぶ | [Models](models/index.md) |
|
||||
| エージェントに機能を追加する | [Tools](tools.md) |
|
||||
| 実際のリポジトリ、ドキュメントバンドル、または分離ワークスペースに対してエージェントを実行する | [Sandbox agents quickstart](sandbox_agents.md) |
|
||||
| manager 型オーケストレーションとハンドオフのどちらにするか決める | [Agent orchestration](multi_agent.md) |
|
||||
| ハンドオフの動作を設定する | [Handoffs](handoffs.md) |
|
||||
| ターンを実行する、イベントをストリーミングする、または会話状態を管理する | [Running agents](running_agents.md) |
|
||||
| 最終出力、実行項目、または再開可能な状態を確認する | [Results](results.md) |
|
||||
| ローカル依存関係と実行時状態を共有する | [Context management](context.md) |
|
||||
| モデルまたはプロバイダー設定を選択する | [モデル](models/index.md) |
|
||||
| エージェントに機能を追加する | [ツール](tools.md) |
|
||||
| 実際のリポジトリ、ドキュメントバンドル、または隔離されたワークスペースに対してエージェントを実行する | [サンドボックスエージェントクイックスタート](sandbox_agents.md) |
|
||||
| マネージャースタイルのオーケストレーションとハンドオフのどちらにするか決める | [エージェントオーケストレーション](multi_agent.md) |
|
||||
| ハンドオフ動作を設定する | [ハンドオフ](handoffs.md) |
|
||||
| ターンを実行する、イベントをストリーミングする、または会話状態を管理する | [エージェントの実行](running_agents.md) |
|
||||
| 最終出力、実行項目、または再開可能な状態を確認する | [実行結果](results.md) |
|
||||
| ローカル依存関係とランタイム状態を共有する | [コンテキスト管理](context.md) |
|
||||
|
||||
## 基本設定
|
||||
|
||||
@@ -32,21 +32,21 @@ SDK は、OpenAI モデルではデフォルトで Responses API を使用しま
|
||||
| プロパティ | 必須 | 説明 |
|
||||
| --- | --- | --- |
|
||||
| `name` | はい | 人間が読めるエージェント名です。 |
|
||||
| `instructions` | はい | システムプロンプト、または動的 instructions コールバックです。[Dynamic instructions](#dynamic-instructions) を参照してください。 |
|
||||
| `prompt` | いいえ | OpenAI Responses API の prompt 設定です。静的な prompt オブジェクトまたは関数を受け付けます。[Prompt templates](#prompt-templates) を参照してください。 |
|
||||
| `handoff_description` | いいえ | このエージェントがハンドオフ先として提示される際に公開される短い説明です。 |
|
||||
| `handoffs` | いいえ | 会話を専門エージェントへ委譲します。[handoffs](handoffs.md) を参照してください。 |
|
||||
| `model` | いいえ | 使用する LLM です。[Models](models/index.md) を参照してください。 |
|
||||
| `instructions` | いいえ | システムプロンプトまたは動的 instructions コールバックです。強く推奨されます。[動的 instructions](#dynamic-instructions) を参照してください。 |
|
||||
| `prompt` | いいえ | OpenAI Responses API のプロンプト設定です。静的なプロンプトオブジェクトまたは関数を受け付けます。[プロンプトテンプレート](#prompt-templates) を参照してください。 |
|
||||
| `handoff_description` | いいえ | このエージェントがハンドオフ先として提示されるときに公開される短い説明です。 |
|
||||
| `handoffs` | いいえ | 会話を専門エージェントに委任します。[ハンドオフ](handoffs.md) を参照してください。 |
|
||||
| `model` | いいえ | 使用する LLM です。[モデル](models/index.md) を参照してください。 |
|
||||
| `model_settings` | いいえ | `temperature`、`top_p`、`tool_choice` などのモデル調整パラメーターです。 |
|
||||
| `tools` | いいえ | エージェントが呼び出せる tools です。[Tools](tools.md) を参照してください。 |
|
||||
| `mcp_servers` | いいえ | エージェント用の MCP ベース tools です。[MCP guide](mcp.md) を参照してください。 |
|
||||
| `mcp_config` | いいえ | strict な schema 変換や MCP 失敗フォーマットなど、MCP tools の準備方法を微調整します。[MCP guide](mcp.md#agent-level-mcp-configuration) を参照してください。 |
|
||||
| `input_guardrails` | いいえ | このエージェントチェーンの最初のユーザー入力で実行されるガードレールです。[Guardrails](guardrails.md) を参照してください。 |
|
||||
| `output_guardrails` | いいえ | このエージェントの最終出力で実行されるガードレールです。[Guardrails](guardrails.md) を参照してください。 |
|
||||
| `output_type` | いいえ | プレーンテキストの代わりに使用する structured output 型です。[Output types](#output-types) を参照してください。 |
|
||||
| `hooks` | いいえ | エージェントスコープのライフサイクルコールバックです。[Lifecycle events (hooks)](#lifecycle-events-hooks) を参照してください。 |
|
||||
| `tool_use_behavior` | いいえ | ツール結果をモデルに戻すか、実行を終了するかを制御します。[Tool use behavior](#tool-use-behavior) を参照してください。 |
|
||||
| `reset_tool_choice` | いいえ | ツール使用ループを避けるため、ツール呼び出し後に `tool_choice` をリセットします (デフォルト: `True`)。[Forcing tool use](#forcing-tool-use) を参照してください。 |
|
||||
| `tools` | いいえ | エージェントが呼び出せるツールです。[ツール](tools.md) を参照してください。 |
|
||||
| `mcp_servers` | いいえ | エージェント向けの MCP に基づくツールです。[MCP ガイド](mcp.md) を参照してください。 |
|
||||
| `mcp_config` | いいえ | 厳密なスキーマ変換や MCP の失敗時のフォーマットなど、MCP ツールの準備方法を微調整します。[MCP ガイド](mcp.md#agent-level-mcp-configuration) を参照してください。 |
|
||||
| `input_guardrails` | いいえ | このエージェントチェーンの最初のユーザー入力に対して実行されるガードレールです。[ガードレール](guardrails.md) を参照してください。 |
|
||||
| `output_guardrails` | いいえ | このエージェントの最終出力に対して実行されるガードレールです。[ガードレール](guardrails.md) を参照してください。 |
|
||||
| `output_type` | いいえ | プレーンテキストの代わりに使用する構造化出力型です。[出力タイプ](#output-types) を参照してください。 |
|
||||
| `hooks` | いいえ | エージェントスコープのライフサイクルコールバックです。[ライフサイクルイベント (フック)](#lifecycle-events-hooks) を参照してください。 |
|
||||
| `tool_use_behavior` | いいえ | ツール結果をモデルに戻すか、実行を終了するかを制御します。[ツール使用動作](#tool-use-behavior) を参照してください。 |
|
||||
| `reset_tool_choice` | いいえ | ツール使用ループを避けるため、ツール呼び出し後に `tool_choice` をリセットします (デフォルト: `True`)。[ツール使用の強制](#forcing-tool-use) を参照してください。 |
|
||||
|
||||
```python
|
||||
from agents import Agent, ModelSettings, function_tool
|
||||
@@ -64,23 +64,23 @@ agent = Agent(
|
||||
)
|
||||
```
|
||||
|
||||
このセクションの内容はすべて `Agent` に適用されます。`SandboxAgent` は同じ考え方に基づき、さらにワークスペーススコープ実行向けに `default_manifest`、`base_instructions`、`capabilities`、`run_as` を追加します。[Sandbox agent concepts](sandbox/guide.md) を参照してください。
|
||||
このセクションの内容はすべて `Agent` に適用されます。SandboxAgent は同じ考え方を基に、ワークスペーススコープの実行向けに `default_manifest`、`base_instructions`、`capabilities`、`run_as` を追加します。[サンドボックスエージェントの概念](sandbox/guide.md) を参照してください。
|
||||
|
||||
## プロンプトテンプレート
|
||||
|
||||
`prompt` を設定することで、OpenAI プラットフォームで作成したプロンプトテンプレートを参照できます。これは Responses API を使用する OpenAI モデルで動作します。
|
||||
`prompt` を設定することで、OpenAI プラットフォームで作成したプロンプトテンプレートを参照できます。これは Responses API を使用する OpenAI モデルで機能します。
|
||||
|
||||
使用するには、次を行ってください。
|
||||
使用するには、次のようにします。
|
||||
|
||||
1. https://platform.openai.com/playground/prompts に移動します
|
||||
2. 新しい prompt 変数 `poem_style` を作成します。
|
||||
1. https://platform.openai.com/playground/prompts に移動します。
|
||||
2. 新しいプロンプト変数 `poem_style` を作成します。
|
||||
3. 次の内容でシステムプロンプトを作成します。
|
||||
|
||||
```
|
||||
Write a poem in {{poem_style}}
|
||||
```
|
||||
|
||||
4. `--prompt-id` フラグを付けて例を実行します。
|
||||
4. このコード例を `--prompt-id` フラグ付きで実行します。
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
@@ -127,9 +127,9 @@ result = await Runner.run(
|
||||
|
||||
## コンテキスト
|
||||
|
||||
エージェントは `context` 型に対してジェネリックです。コンテキストは依存性注入ツールです。これは、作成して `Runner.run()` に渡すオブジェクトで、すべてのエージェント、ツール、ハンドオフなどに渡され、エージェント実行の依存関係と状態をまとめる入れ物として機能します。コンテキストには任意の Python オブジェクトを提供できます。
|
||||
エージェントは `context` 型に対してジェネリックです。コンテキストは依存性注入のためのツールです。ユーザーが作成して `Runner.run()` に渡すオブジェクトであり、すべてのエージェント、ツール、ハンドオフなどに渡され、エージェント実行の依存関係や状態をまとめて保持するものとして機能します。任意の Python オブジェクトをコンテキストとして提供できます。
|
||||
|
||||
`RunContextWrapper` の完全な機能、共有使用量トラッキング、ネストされた `tool_input`、シリアライズ時の注意点については、[context guide](context.md) をお読みください。
|
||||
`RunContextWrapper` の完全な API サーフェス、共有された使用量追跡、ネストされた `tool_input`、シリアライズ時の注意点については、[コンテキストガイド](context.md) を参照してください。
|
||||
|
||||
```python
|
||||
@dataclass
|
||||
@@ -146,9 +146,9 @@ agent = Agent[UserContext](
|
||||
)
|
||||
```
|
||||
|
||||
## 出力型
|
||||
## 出力タイプ
|
||||
|
||||
デフォルトでは、エージェントはプレーンテキスト (つまり `str`) 出力を生成します。エージェントに特定の型の出力を生成させたい場合は、`output_type` パラメーターを使用できます。一般的な選択肢は [Pydantic](https://docs.pydantic.dev/) オブジェクトですが、Pydantic の [TypeAdapter](https://docs.pydantic.dev/latest/api/type_adapter/) でラップできる任意の型 (dataclasses、lists、TypedDict など) をサポートしています。
|
||||
デフォルトでは、エージェントはプレーンテキスト (つまり `str`) の出力を生成します。エージェントに特定の型の出力を生成させたい場合は、`output_type` パラメーターを使用できます。一般的な選択肢は [Pydantic](https://docs.pydantic.dev/) オブジェクトを使用することですが、Pydantic の [TypeAdapter](https://docs.pydantic.dev/latest/api/type_adapter/) でラップできる任意の型をサポートしています。dataclass、リスト、TypedDict などです。
|
||||
|
||||
```python
|
||||
from pydantic import BaseModel
|
||||
@@ -169,20 +169,20 @@ agent = Agent(
|
||||
|
||||
!!! note
|
||||
|
||||
`output_type` を渡すと、通常のプレーンテキスト応答ではなく [structured outputs](https://platform.openai.com/docs/guides/structured-outputs) を使用するようモデルに指示します。
|
||||
`output_type` を渡すと、通常のプレーンテキストレスポンスではなく [structured outputs](https://platform.openai.com/docs/guides/structured-outputs) を使用するようモデルに指示します。
|
||||
|
||||
## マルチエージェントシステム設計パターン
|
||||
## マルチエージェントシステムの設計パターン
|
||||
|
||||
マルチエージェントシステムの設計方法は多数ありますが、広く適用可能なパターンとしては主に次の 2 つがよく見られます。
|
||||
マルチエージェントシステムを設計する方法は多数ありますが、広く適用できるパターンとして主に次の 2 つがよく見られます。
|
||||
|
||||
1. Manager (Agents as tools): 中央の manager / orchestrator が、専門化されたサブエージェントを tools として呼び出し、会話の制御を保持します。
|
||||
2. ハンドオフ: ピアエージェントが、会話を引き継ぐ専門エージェントへ制御をハンドオフします。これは分散型です。
|
||||
1. マネージャー (agents as tools): 中央のマネージャー/オーケストレーターが、専門化されたサブエージェントをツールとして呼び出し、会話の制御を保持します。
|
||||
2. ハンドオフ: 同等の立場のエージェントが、会話を引き継ぐ専門エージェントに制御をハンドオフします。これは分散型です。
|
||||
|
||||
詳細は [our practical guide to building agents](https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf) を参照してください。
|
||||
詳細については、[エージェント構築の実践ガイド](https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf) を参照してください。
|
||||
|
||||
### Manager (Agents as tools)
|
||||
### マネージャー (agents as tools)
|
||||
|
||||
`customer_facing_agent` はすべてのユーザー対話を処理し、tools として公開された専門サブエージェントを呼び出します。詳細は [tools](tools.md#agents-as-tools) のドキュメントを参照してください。
|
||||
`customer_facing_agent` はすべてのユーザー操作を処理し、ツールとして公開された専門化されたサブエージェントを呼び出します。詳しくは [ツール](tools.md#agents-as-tools) ドキュメントを参照してください。
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
@@ -211,7 +211,7 @@ customer_facing_agent = Agent(
|
||||
|
||||
### ハンドオフ
|
||||
|
||||
ハンドオフは、エージェントが委譲できるサブエージェントです。ハンドオフが発生すると、委譲先エージェントが会話履歴を受け取り、会話を引き継ぎます。このパターンにより、単一タスクに特化して優れたモジュール型の専門エージェントを実現できます。詳細は [handoffs](handoffs.md) のドキュメントを参照してください。
|
||||
ハンドオフは、エージェントが委任できるサブエージェントです。ハンドオフが発生すると、委任先のエージェントは会話履歴を受け取り、会話を引き継ぎます。このパターンにより、単一のタスクに優れたモジュール型の専門エージェントを実現できます。詳しくは [ハンドオフ](handoffs.md) ドキュメントを参照してください。
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
@@ -232,7 +232,7 @@ triage_agent = Agent(
|
||||
|
||||
## 動的 instructions
|
||||
|
||||
ほとんどの場合、エージェント作成時に instructions を提供できます。ただし、関数を介して動的 instructions を提供することもできます。この関数はエージェントとコンテキストを受け取り、プロンプトを返す必要があります。通常の関数と `async` 関数の両方を受け付けます。
|
||||
多くの場合、エージェントの作成時に instructions を指定できます。ただし、関数を通じて動的な instructions を指定することもできます。この関数はエージェントとコンテキストを受け取り、プロンプトを返す必要があります。通常の関数と `async` 関数の両方を受け付けます。
|
||||
|
||||
```python
|
||||
def dynamic_instructions(
|
||||
@@ -247,29 +247,28 @@ agent = Agent[UserContext](
|
||||
)
|
||||
```
|
||||
|
||||
## ライフサイクルイベント (hooks)
|
||||
## ライフサイクルイベント (フック)
|
||||
|
||||
場合によっては、エージェントのライフサイクルを監視したいことがあります。たとえば、イベントをログに記録したり、データを事前取得したり、特定イベント発生時の使用状況を記録したりしたい場合です。
|
||||
エージェントのライフサイクルを監視したい場合があります。たとえば、特定のイベントが発生したときにイベントをログに記録したり、データを事前取得したり、使用量を記録したりできます。
|
||||
|
||||
hook のスコープは 2 つあります。
|
||||
フックのスコープは 2 つあります。
|
||||
|
||||
- [`RunHooks`][agents.lifecycle.RunHooks] は、他エージェントへのハンドオフを含む `Runner.run(...)` 呼び出し全体を監視します。
|
||||
- [`AgentHooks`][agents.lifecycle.AgentHooks] は `agent.hooks` を介して特定のエージェントインスタンスにアタッチされます。
|
||||
- [`RunHooks`][agents.lifecycle.RunHooks] は、他のエージェントへのハンドオフを含む、`Runner.run(...)` 呼び出し全体を監視します。
|
||||
- [`AgentHooks`][agents.lifecycle.AgentHooks] は、`agent.hooks` を通じて特定のエージェントインスタンスにアタッチされます。
|
||||
|
||||
コールバックコンテキストもイベントによって変わります。
|
||||
コールバックのコンテキストも、イベントによって変わります。
|
||||
|
||||
- エージェント開始 / 終了 hook は、元のコンテキストをラップし、共有実行使用量状態を保持する [`AgentHookContext`][agents.run_context.AgentHookContext] を受け取ります。
|
||||
- LLM、ツール、ハンドオフ hook は [`RunContextWrapper`][agents.run_context.RunContextWrapper] を受け取ります。
|
||||
- エージェントの開始/終了フックは [`AgentHookContext`][agents.run_context.AgentHookContext] を受け取ります。これは元のコンテキストをラップし、共有された実行使用量状態を保持します。
|
||||
- LLM、ツール、ハンドオフのフックは [`RunContextWrapper`][agents.run_context.RunContextWrapper] を受け取ります。
|
||||
|
||||
典型的な hook のタイミング:
|
||||
典型的なフックのタイミング:
|
||||
|
||||
- `on_agent_start` / `on_agent_end`: 特定エージェントが最終出力の生成を開始または終了したとき。
|
||||
- `on_llm_start` / `on_llm_end`: 各モデル呼び出しの直前 / 直後。
|
||||
- `on_tool_start` / `on_tool_end`: 各ローカルツール呼び出しの前後。
|
||||
関数ツールでは、hook の `context` は通常 `ToolContext` なので、`tool_call_id` などのツール呼び出しメタデータを確認できます。
|
||||
- `on_agent_start` / `on_agent_end`: 特定のエージェントが最終出力の生成を開始または完了したとき。
|
||||
- `on_llm_start` / `on_llm_end`: 各モデル呼び出しの直前と直後。
|
||||
- `on_tool_start` / `on_tool_end`: 各ローカルツール呼び出しの前後。関数ツールでは、フックの `context` は通常 `ToolContext` なので、`tool_call_id` などのツール呼び出しメタデータを確認できます。
|
||||
- `on_handoff`: 制御があるエージェントから別のエージェントに移るとき。
|
||||
|
||||
ワークフロー全体に対して単一の監視者が必要な場合は `RunHooks` を使用し、1 つのエージェントでカスタムな副作用が必要な場合は `AgentHooks` を使用してください。
|
||||
ワークフロー全体に対する単一のオブザーバーが必要な場合は `RunHooks` を使用し、1 つのエージェントにカスタムの副作用が必要な場合は `AgentHooks` を使用してください。
|
||||
|
||||
```python
|
||||
from agents import Agent, RunHooks, Runner
|
||||
@@ -291,21 +290,21 @@ result = await Runner.run(agent, "Explain quines", hooks=LoggingHooks())
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
コールバックの完全な仕様は、[Lifecycle API reference](ref/lifecycle.md) を参照してください。
|
||||
コールバック API 全体については、[ライフサイクル API リファレンス](ref/lifecycle.md) を参照してください。
|
||||
|
||||
## ガードレール
|
||||
|
||||
ガードレールを使うと、エージェント実行と並行してユーザー入力に対するチェック / 検証を実行し、さらに出力生成後にエージェントの出力に対するチェック / 検証も実行できます。たとえば、ユーザー入力とエージェント出力の関連性をスクリーニングできます。詳細は [guardrails](guardrails.md) のドキュメントを参照してください。
|
||||
ガードレールを使用すると、エージェントの実行と並行してユーザー入力に対するチェック/検証を実行し、またエージェントの出力が生成された後にその出力に対するチェック/検証を実行できます。たとえば、ユーザー入力とエージェント出力の関連性をスクリーニングできます。詳しくは [ガードレール](guardrails.md) ドキュメントを参照してください。
|
||||
|
||||
## エージェントの複製 / コピー
|
||||
## エージェントのクローン/コピー
|
||||
|
||||
エージェントで `clone()` メソッドを使用すると、Agent を複製し、必要に応じて任意のプロパティを変更できます。
|
||||
エージェントの `clone()` メソッドを使用すると、エージェントを複製し、必要に応じて任意のプロパティを変更できます。
|
||||
|
||||
```python
|
||||
pirate_agent = Agent(
|
||||
name="Pirate",
|
||||
instructions="Write like a pirate",
|
||||
model="gpt-5.4",
|
||||
model="gpt-5.5",
|
||||
)
|
||||
|
||||
robot_agent = pirate_agent.clone(
|
||||
@@ -316,14 +315,14 @@ robot_agent = pirate_agent.clone(
|
||||
|
||||
## ツール使用の強制
|
||||
|
||||
ツールのリストを提供しても、LLM が必ずツールを使用するとは限りません。[`ModelSettings.tool_choice`][agents.model_settings.ModelSettings.tool_choice] を設定することでツール使用を強制できます。有効な値は次のとおりです。
|
||||
ツールのリストを渡しても、LLM が必ずツールを使用するとは限りません。[`ModelSettings.tool_choice`][agents.model_settings.ModelSettings.tool_choice] を設定することで、ツール使用を強制できます。有効な値は次のとおりです。
|
||||
|
||||
1. `auto`: LLM がツールを使用するかどうかを判断します。
|
||||
2. `required`: LLM にツール使用を必須化します (ただし、どのツールを使うかは賢く判断できます)。
|
||||
3. `none`: LLM にツールを _使用しない_ ことを必須化します。
|
||||
4. 具体的な文字列 (例: `my_tool`) を設定: LLM にその特定ツールの使用を必須化します。
|
||||
1. `auto`: LLM がツールを使用するかどうかを判断できるようにします。
|
||||
2. `required`: LLM にツールの使用を必須にします (ただし、どのツールを使うかは賢く判断できます)。
|
||||
3. `none`: LLM にツールを使用 _しない_ ことを必須にします。
|
||||
4. `my_tool` などの特定の文字列を設定すると、LLM にその特定のツールの使用を必須にします。
|
||||
|
||||
OpenAI Responses の tool search を使用している場合、名前付き tool choice にはより多くの制限があります。`tool_choice` では素の namespace 名や deferred 専用ツールを指定できず、`tool_choice="tool_search"` は [`ToolSearchTool`][agents.tool.ToolSearchTool] を対象にしません。これらの場合は `auto` または `required` を推奨します。Responses 固有の制約は [Hosted tool search](tools.md#hosted-tool-search) を参照してください。
|
||||
OpenAI Responses のツール検索を使用する場合、名前付きツール選択にはより多くの制限があります。`tool_choice` では素の名前空間名や遅延専用ツールを対象にできず、`tool_choice="tool_search"` は [`ToolSearchTool`][agents.tool.ToolSearchTool] を対象にしません。このような場合は、`auto` または `required` を優先してください。Responses 固有の制約については [ホスト型ツール検索](tools.md#hosted-tool-search) を参照してください。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, function_tool, ModelSettings
|
||||
@@ -343,10 +342,10 @@ agent = Agent(
|
||||
|
||||
## ツール使用動作
|
||||
|
||||
`Agent` 設定内の `tool_use_behavior` パラメーターは、ツール出力の処理方法を制御します。
|
||||
`Agent` 設定の `tool_use_behavior` パラメーターは、ツール出力の処理方法を制御します。
|
||||
|
||||
- `"run_llm_again"`: デフォルトです。ツールを実行し、その結果を LLM が処理して最終応答を生成します。
|
||||
- `"stop_on_first_tool"`: 最初のツール呼び出しの出力を、追加の LLM 処理なしで最終応答として使用します。
|
||||
- `"run_llm_again"`: デフォルトです。ツールが実行され、LLM がその結果を処理して最終レスポンスを生成します。
|
||||
- `"stop_on_first_tool"`: 最初のツール呼び出しの出力が、それ以上の LLM 処理なしに最終レスポンスとして使用されます。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, function_tool, ModelSettings
|
||||
@@ -364,7 +363,7 @@ agent = Agent(
|
||||
)
|
||||
```
|
||||
|
||||
- `StopAtTools(stop_at_tool_names=[...])`: 指定したツールのいずれかが呼び出された場合に停止し、その出力を最終応答として使用します。
|
||||
- `StopAtTools(stop_at_tool_names=[...])`: 指定されたツールのいずれかが呼び出された場合に停止し、その出力を最終レスポンスとして使用します。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, function_tool
|
||||
@@ -388,7 +387,7 @@ agent = Agent(
|
||||
)
|
||||
```
|
||||
|
||||
- `ToolsToFinalOutputFunction`: ツール結果を処理し、停止するか LLM で継続するかを決定するカスタム関数です。
|
||||
- `ToolsToFinalOutputFunction`: ツール結果を処理し、停止するか LLM による処理を続行するかを判断するカスタム関数です。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, function_tool, FunctionToolResult, RunContextWrapper
|
||||
@@ -426,4 +425,4 @@ agent = Agent(
|
||||
|
||||
!!! note
|
||||
|
||||
無限ループを防ぐため、フレームワークはツール呼び出し後に `tool_choice` を自動的に "auto" にリセットします。この動作は [`agent.reset_tool_choice`][agents.agent.Agent.reset_tool_choice] で設定可能です。無限ループが起きる理由は、ツール結果が LLM に送信され、その後 `tool_choice` のために LLM が再びツール呼び出しを生成し、これが際限なく続くためです。
|
||||
無限ループを防ぐため、フレームワークはツール呼び出し後に `tool_choice` を自動的に "auto" にリセットします。この動作は [`agent.reset_tool_choice`][agents.agent.Agent.reset_tool_choice] で設定できます。無限ループが発生するのは、ツール結果が LLM に送信され、`tool_choice` のために LLM がさらに別のツール呼び出しを生成し、それが際限なく続くためです。
|
||||
+56
-24
@@ -4,21 +4,21 @@ search:
|
||||
---
|
||||
# 設定
|
||||
|
||||
このページでは、通常はアプリケーション起動時に 1 度だけ設定する SDK 全体のデフォルト(デフォルトの OpenAI キーまたはクライアント、デフォルトの OpenAI API 形式、トレーシングエクスポートのデフォルト、ログ動作など)を扱います。
|
||||
このページでは、デフォルトの OpenAI キーまたはクライアント、デフォルトの OpenAI API の形式、トレーシングエクスポートのデフォルト、ログ動作など、通常はアプリケーションの起動時に一度だけ設定する SDK 全体のデフォルトについて説明します。
|
||||
|
||||
これらのデフォルトは sandbox ベースのワークフローにも適用されますが、sandbox ワークスペース、sandbox クライアント、セッション再利用は別途設定します。
|
||||
これらのデフォルトはサンドボックスベースのワークフローにも引き続き適用されますが、サンドボックスワークスペース、サンドボックスクライアント、セッションの再利用は別途設定します。
|
||||
|
||||
代わりに特定のエージェントや実行を設定する必要がある場合は、次から始めてください:
|
||||
代わりに特定のエージェントまたは実行を設定する必要がある場合は、まず次を参照してください:
|
||||
|
||||
- 通常の `Agent` における instructions、ツール、出力タイプ、ハンドオフ、ガードレールについては [Agents](agents.md)。
|
||||
- `RunConfig`、セッション、会話状態オプションについては [エージェントの実行](running_agents.md)。
|
||||
- `SandboxRunConfig`、マニフェスト、機能、sandbox クライアント固有のワークスペース設定については [Sandbox エージェント](sandbox/guide.md)。
|
||||
- モデル選択とプロバイダー設定については [Models](models/index.md)。
|
||||
- 実行ごとのトレーシングメタデータとカスタムトレースプロセッサーについては [トレーシング](tracing.md)。
|
||||
- [エージェント](agents.md): 通常の `Agent` の instructions、tools、出力タイプ、ハンドオフ、ガードレールについて。
|
||||
- [エージェントの実行](running_agents.md): `RunConfig`、セッション、会話状態オプションについて。
|
||||
- [サンドボックスエージェント](sandbox/guide.md): `SandboxRunConfig`、マニフェスト、ケイパビリティ、サンドボックスクライアント固有のワークスペース設定について。
|
||||
- [モデル](models/index.md): モデル選択とプロバイダー設定について。
|
||||
- [トレーシング](tracing.md): 実行ごとのトレーシングメタデータとカスタムトレースプロセッサーについて。
|
||||
|
||||
## API キーとクライアント
|
||||
|
||||
デフォルトでは、SDK は LLM リクエストとトレーシングに `OPENAI_API_KEY` 環境変数を使用します。キーは SDK が最初に OpenAI クライアントを作成する際(遅延初期化)に解決されるため、最初のモデル呼び出し前に環境変数を設定してください。アプリ起動前にその環境変数を設定できない場合は、キーを設定するために [set_default_openai_key()][agents.set_default_openai_key] 関数を使用できます。
|
||||
デフォルトでは、SDK は LLM リクエストとトレーシングに `OPENAI_API_KEY` 環境変数を使用します。このキーは、SDK が初めて OpenAI クライアントを作成するときに解決されます(遅延初期化)。そのため、最初のモデル呼び出しの前に環境変数を設定してください。アプリの起動前にその環境変数を設定できない場合は、[set_default_openai_key()][agents.set_default_openai_key] 関数を使用してキーを設定できます。
|
||||
|
||||
```python
|
||||
from agents import set_default_openai_key
|
||||
@@ -26,7 +26,7 @@ from agents import set_default_openai_key
|
||||
set_default_openai_key("sk-...")
|
||||
```
|
||||
|
||||
または、使用する OpenAI クライアントを設定することもできます。デフォルトでは、SDK は環境変数の API キーまたは上記で設定したデフォルトキーを使用して `AsyncOpenAI` インスタンスを作成します。これは [set_default_openai_client()][agents.set_default_openai_client] 関数で変更できます。
|
||||
また、使用する OpenAI クライアントを設定することもできます。デフォルトでは、SDK は環境変数の API キー、または上で設定したデフォルトキーを使用して、`AsyncOpenAI` インスタンスを作成します。これは [set_default_openai_client()][agents.set_default_openai_client] 関数を使用して変更できます。
|
||||
|
||||
```python
|
||||
from openai import AsyncOpenAI
|
||||
@@ -36,14 +36,14 @@ custom_client = AsyncOpenAI(base_url="...", api_key="...")
|
||||
set_default_openai_client(custom_client)
|
||||
```
|
||||
|
||||
環境変数ベースのエンドポイント設定を使いたい場合、デフォルトの OpenAI プロバイダーは `OPENAI_BASE_URL` も読み取ります。Responses websocket トランスポートを有効にすると、websocket `/responses` エンドポイント用に `OPENAI_WEBSOCKET_BASE_URL` も読み取ります。
|
||||
環境変数ベースのエンドポイント設定を使用したい場合、デフォルトの OpenAI プロバイダーは `OPENAI_BASE_URL` も読み取ります。Responses WebSocket トランスポートを有効にすると、WebSocket の `/responses` エンドポイント用に `OPENAI_WEBSOCKET_BASE_URL` も読み取ります。
|
||||
|
||||
```bash
|
||||
export OPENAI_BASE_URL="https://your-openai-compatible-endpoint.example/v1"
|
||||
export OPENAI_WEBSOCKET_BASE_URL="wss://your-openai-compatible-endpoint.example/v1"
|
||||
```
|
||||
|
||||
最後に、使用する OpenAI API をカスタマイズすることもできます。デフォルトでは OpenAI Responses API を使用します。これは [set_default_openai_api()][agents.set_default_openai_api] 関数を使って Chat Completions API を使うように上書きできます。
|
||||
最後に、使用される OpenAI API もカスタマイズできます。デフォルトでは OpenAI Responses API を使用します。[set_default_openai_api()][agents.set_default_openai_api] 関数を使用すると、これを上書きして Chat Completions API を使用できます。
|
||||
|
||||
```python
|
||||
from agents import set_default_openai_api
|
||||
@@ -51,9 +51,41 @@ from agents import set_default_openai_api
|
||||
set_default_openai_api("chat_completions")
|
||||
```
|
||||
|
||||
## OpenAI プロバイダーのデフォルト
|
||||
|
||||
OpenAI を基盤とするプロバイダーは、モデル名を解決するときにも SDK 全体のデフォルトを読み取ります。OpenAI Responses モデルがデフォルトで WebSocket トランスポートを使用するようにするには、[`set_default_openai_responses_transport()`][agents.set_default_openai_responses_transport] を使用します:
|
||||
|
||||
```python
|
||||
from agents import set_default_openai_responses_transport
|
||||
|
||||
set_default_openai_responses_transport("websocket")
|
||||
```
|
||||
|
||||
これは、デフォルトの OpenAI プロバイダーによって解決される OpenAI Responses モデルに影響します。プロバイダーレベルのセットアップ、接続の再利用、キープアライブオプション、カスタム WebSocket エンドポイントについては、[Responses WebSocket トランスポート](models/index.md#responses-websocket-transport) を参照してください。
|
||||
|
||||
OpenAI セットアップでプロバイダーレベルのエージェント登録メタデータが必要な場合は、起動時にデフォルトのハーネス ID を一度設定してください:
|
||||
|
||||
```python
|
||||
from agents import set_default_openai_harness
|
||||
|
||||
set_default_openai_harness("your-harness-id")
|
||||
```
|
||||
|
||||
完全な登録オブジェクトを渡すこともできます:
|
||||
|
||||
```python
|
||||
from agents import OpenAIAgentRegistrationConfig, set_default_openai_agent_registration
|
||||
|
||||
set_default_openai_agent_registration(
|
||||
OpenAIAgentRegistrationConfig(harness_id="your-harness-id")
|
||||
)
|
||||
```
|
||||
|
||||
SDK のデフォルトが設定されていない場合、OpenAI を基盤とするプロバイダーは `OPENAI_AGENT_HARNESS_ID` 環境変数にフォールバックします。ハーネス ID が設定されている場合、`agent_harness_id` が `RunConfig.trace_metadata` にすでに存在する場合を除き、SDK はそれを `agent_harness_id` としてトレースメタデータに追加します。
|
||||
|
||||
## トレーシング
|
||||
|
||||
トレーシングはデフォルトで有効です。デフォルトでは、上のセクションのモデルリクエストと同じ OpenAI API キー(つまり環境変数または設定したデフォルトキー)を使用します。トレーシングに使用する API キーは [`set_tracing_export_api_key`][agents.set_tracing_export_api_key] 関数で明示的に設定できます。
|
||||
トレーシングはデフォルトで有効です。デフォルトでは、上のセクションで説明したモデルリクエストと同じ OpenAI API キー(つまり、環境変数または設定したデフォルトキー)を使用します。トレーシングに使用する API キーは、[`set_tracing_export_api_key`][agents.set_tracing_export_api_key] 関数を使用して個別に設定できます。
|
||||
|
||||
```python
|
||||
from agents import set_tracing_export_api_key
|
||||
@@ -61,7 +93,7 @@ from agents import set_tracing_export_api_key
|
||||
set_tracing_export_api_key("sk-...")
|
||||
```
|
||||
|
||||
モデル通信があるキーまたはクライアントを使い、トレーシングは別の OpenAI キーを使う必要がある場合、デフォルトキーまたはクライアント設定時に `use_for_tracing=False` を渡してから、トレーシングを個別に設定してください。カスタムクライアントを使わない場合は [`set_default_openai_key()`][agents.set_default_openai_key] でも同じパターンが使えます。
|
||||
モデルのトラフィックではあるキーまたはクライアントを使用し、トレーシングでは別の OpenAI キーを使用したい場合は、デフォルトのキーまたはクライアントを設定するときに `use_for_tracing=False` を渡し、そのうえでトレーシングを個別に設定してください。カスタムクライアントを使用していない場合は、同じパターンを [`set_default_openai_key()`][agents.set_default_openai_key] にも適用できます。
|
||||
|
||||
```python
|
||||
from openai import AsyncOpenAI
|
||||
@@ -76,7 +108,7 @@ set_default_openai_client(custom_client, use_for_tracing=False)
|
||||
set_tracing_export_api_key("sk-tracing")
|
||||
```
|
||||
|
||||
デフォルトのエクスポーター使用時に、トレースを特定の組織またはプロジェクトに紐付ける必要がある場合は、アプリ起動前に以下の環境変数を設定してください:
|
||||
デフォルトエクスポーターを使用するときに、トレースを特定の組織またはプロジェクトに関連付ける必要がある場合は、アプリの起動前に次の環境変数を設定してください:
|
||||
|
||||
```bash
|
||||
export OPENAI_ORG_ID="org_..."
|
||||
@@ -103,7 +135,7 @@ from agents import set_tracing_disabled
|
||||
set_tracing_disabled(True)
|
||||
```
|
||||
|
||||
トレーシングを有効のまま、トレースペイロードから機密性の高い可能性がある入出力を除外したい場合は、[`RunConfig.trace_include_sensitive_data`][agents.run.RunConfig.trace_include_sensitive_data] を `False` に設定してください:
|
||||
トレーシングを有効のままにしつつ、機微情報を含む可能性のある入力/出力をトレースペイロードから除外したい場合は、[`RunConfig.trace_include_sensitive_data`][agents.run.RunConfig.trace_include_sensitive_data] を `False` に設定してください:
|
||||
|
||||
```python
|
||||
from agents import Runner, RunConfig
|
||||
@@ -115,17 +147,17 @@ await Runner.run(
|
||||
)
|
||||
```
|
||||
|
||||
アプリ起動前にこの環境変数を設定すれば、コードなしでデフォルトを変更することもできます:
|
||||
コードを使わずにデフォルトを変更することもできます。アプリの起動前にこの環境変数を設定してください:
|
||||
|
||||
```bash
|
||||
export OPENAI_AGENTS_TRACE_INCLUDE_SENSITIVE_DATA=0
|
||||
```
|
||||
|
||||
トレーシング制御の全体については、[トレーシングガイド](tracing.md) を参照してください。
|
||||
トレーシング制御の詳細については、[トレーシングガイド](tracing.md) を参照してください。
|
||||
|
||||
## デバッグログ
|
||||
|
||||
SDK は 2 つの Python ロガー(`openai.agents` と `openai.agents.tracing`)を定義しており、デフォルトではハンドラーをアタッチしません。ログはアプリケーションの Python ログ設定に従います。
|
||||
SDK は 2 つの Python ロガー(`openai.agents` と `openai.agents.tracing`)を定義しており、デフォルトではハンドラーをアタッチしません。ログは、アプリケーションの Python ロギング設定に従います。
|
||||
|
||||
詳細ログを有効にするには、[`enable_verbose_stdout_logging()`][agents.enable_verbose_stdout_logging] 関数を使用します。
|
||||
|
||||
@@ -135,7 +167,7 @@ from agents import enable_verbose_stdout_logging
|
||||
enable_verbose_stdout_logging()
|
||||
```
|
||||
|
||||
または、ハンドラー、フィルター、フォーマッターなどを追加してログをカスタマイズできます。詳細は [Python logging guide](https://docs.python.org/3/howto/logging.html) を参照してください。
|
||||
また、ハンドラー、フィルター、フォーマッターなどを追加してログをカスタマイズすることもできます。詳しくは [Python ロギングガイド](https://docs.python.org/3/howto/logging.html) を参照してください。
|
||||
|
||||
```python
|
||||
import logging
|
||||
@@ -154,18 +186,18 @@ logger.setLevel(logging.WARNING)
|
||||
logger.addHandler(logging.StreamHandler())
|
||||
```
|
||||
|
||||
### ログ内の機密データ
|
||||
### ログ内の機微データ
|
||||
|
||||
特定のログには機密データ(たとえばユーザーデータ)が含まれる場合があります。
|
||||
一部のログには、機微データ(たとえば、ユーザーデータ)が含まれる場合があります。
|
||||
|
||||
デフォルトでは、SDK は LLM の入出力やツールの入出力を **ログに記録しません**。これらの保護は次によって制御されます:
|
||||
デフォルトでは、SDK は LLM の入力/出力やツールの入力/出力を **ログに記録しません** 。これらの保護は次の項目で制御されます:
|
||||
|
||||
```bash
|
||||
OPENAI_AGENTS_DONT_LOG_MODEL_DATA=1
|
||||
OPENAI_AGENTS_DONT_LOG_TOOL_DATA=1
|
||||
```
|
||||
|
||||
デバッグのために一時的にこれらのデータを含める必要がある場合は、アプリ起動前にいずれかの変数を `0`(または `false`)に設定してください:
|
||||
デバッグのためにこのデータを一時的に含める必要がある場合は、アプリの起動前にいずれかの変数を `0`(または `false`)に設定してください:
|
||||
|
||||
```bash
|
||||
export OPENAI_AGENTS_DONT_LOG_MODEL_DATA=0
|
||||
|
||||
+41
-41
@@ -4,49 +4,49 @@ search:
|
||||
---
|
||||
# コンテキスト管理
|
||||
|
||||
コンテキストは多義的な用語です。主に、重要になるコンテキストには 2 つの分類があります。
|
||||
コンテキストは多義的な用語です。考慮すべきコンテキストには、主に 2 つの種類があります:
|
||||
|
||||
1. コード内でローカルに利用可能なコンテキスト: これは、関数ツールの実行時、`on_handoff` のようなコールバック時、ライフサイクルフック時などに必要になる可能性があるデータや依存関係です。
|
||||
2. LLM が利用可能なコンテキスト: これは、LLM がレスポンスを生成するときに参照するデータです。
|
||||
1. コードからローカルに利用できるコンテキスト: これは、ツール関数の実行時、`on_handoff` のようなコールバック内、ライフサイクルフック内などで必要になる可能性のあるデータや依存関係です。
|
||||
2. LLM が利用できるコンテキスト: これは、LLM が応答を生成するときに参照するデータです。
|
||||
|
||||
## ローカルコンテキスト
|
||||
|
||||
これは [`RunContextWrapper`][agents.run_context.RunContextWrapper] クラスと、その内部の [`context`][agents.run_context.RunContextWrapper.context] プロパティで表現されます。動作は次のとおりです。
|
||||
これは [`RunContextWrapper`][agents.run_context.RunContextWrapper] クラス、およびその中の [`context`][agents.run_context.RunContextWrapper.context] プロパティで表現されます。仕組みは次のとおりです:
|
||||
|
||||
1. 任意の Python オブジェクトを作成します。一般的なパターンは、dataclass または Pydantic オブジェクトを使うことです。
|
||||
2. そのオブジェクトを各種 run メソッドに渡します(例: `Runner.run(..., context=whatever)`)。
|
||||
3. すべてのツール呼び出し、ライフサイクルフックなどには `RunContextWrapper[T]` のラッパーオブジェクトが渡されます。ここで `T` はコンテキストオブジェクトの型を表し、`wrapper.context` でアクセスできます。
|
||||
1. 任意の Python オブジェクトを作成します。一般的なパターンは dataclass や Pydantic オブジェクトを使用することです。
|
||||
2. そのオブジェクトを各種実行メソッドに渡します (例: `Runner.run(..., context=whatever)`)。
|
||||
3. すべてのツール呼び出し、ライフサイクルフックなどには、ラッパーオブジェクト `RunContextWrapper[T]` が渡されます。ここで `T` はコンテキストオブジェクトの型を表し、`wrapper.context` を通じてアクセスできます。
|
||||
|
||||
ランタイム固有の一部コールバックでは、SDK が `RunContextWrapper[T]` のより特化したサブクラスを渡す場合があります。たとえば、関数ツールのライフサイクルフックは通常 `ToolContext` を受け取り、`tool_call_id`、`tool_name`、`tool_arguments` などのツール呼び出しメタデータにもアクセスできます。
|
||||
一部のランタイム固有のコールバックでは、SDK はより特殊化された `RunContextWrapper[T]` のサブクラスを渡す場合があります。たとえば、関数ツールのライフサイクルフックは通常 `ToolContext` を受け取り、これは `tool_call_id`、`tool_name`、`tool_arguments` などのツール呼び出しメタデータも公開します。
|
||||
|
||||
認識しておくべき **最も重要** な点: 特定のエージェント実行におけるすべてのエージェント、関数ツール、ライフサイクルなどは、同じコンテキストの _型_ を使用する必要があります。
|
||||
認識すべき **最も重要な** 点は、あるエージェント実行におけるすべてのエージェント、ツール関数、ライフサイクルなどが、同じ _型_ のコンテキストを使用しなければならないということです。
|
||||
|
||||
コンテキストは次のような用途で使用できます。
|
||||
コンテキストは、たとえば次の用途に使用できます:
|
||||
|
||||
- 実行のためのコンテキストデータ(例: ユーザー名 / uid や、ユーザーに関するその他の情報)
|
||||
- 依存関係(例: logger オブジェクト、データ取得処理など)
|
||||
- 実行時のコンテキストデータ (例: ユーザー名 / uid や、ユーザーに関するその他の情報)
|
||||
- 依存関係 (例: ロガーオブジェクト、データ取得器など)
|
||||
- ヘルパー関数
|
||||
|
||||
!!! danger "注意"
|
||||
!!! danger "注記"
|
||||
|
||||
コンテキストオブジェクトは LLM に **送信されません**。これは純粋にローカルオブジェクトであり、読み取り、書き込み、メソッド呼び出しが可能です。
|
||||
コンテキストオブジェクトは LLM に **送信されません**。これは完全にローカルなオブジェクトであり、読み取り、書き込み、メソッドの呼び出しができます。
|
||||
|
||||
1 回の実行内では、派生ラッパーは同じ基盤のアプリコンテキスト、承認状態、使用量トラッキングを共有します。ネストした [`Agent.as_tool()`][agents.agent.Agent.as_tool] 実行では別の `tool_input` が付与される場合がありますが、デフォルトではアプリ状態の分離コピーは取得しません。
|
||||
1 回の実行内では、派生したラッパーは同じ基盤となるアプリコンテキスト、承認状態、使用状況の追跡を共有します。ネストされた [`Agent.as_tool()`][agents.agent.Agent.as_tool] の実行では異なる `tool_input` を付加する場合がありますが、デフォルトではアプリ状態の分離コピーは取得しません。
|
||||
|
||||
### `RunContextWrapper` の公開内容
|
||||
|
||||
[`RunContextWrapper`][agents.run_context.RunContextWrapper] は、アプリで定義したコンテキストオブジェクトのラッパーです。実際には、主に次を使用します。
|
||||
[`RunContextWrapper`][agents.run_context.RunContextWrapper] は、アプリで定義したコンテキストオブジェクトを包むラッパーです。実際には、ほとんどの場合、次のものを使用します:
|
||||
|
||||
- 独自の可変アプリ状態および依存関係には [`wrapper.context`][agents.run_context.RunContextWrapper.context]。
|
||||
- 現在の実行全体の集計されたリクエストおよびトークン使用量には [`wrapper.usage`][agents.run_context.RunContextWrapper.usage]。
|
||||
- 現在の実行が [`Agent.as_tool()`][agents.agent.Agent.as_tool] 内で実行されているときの構造化入力には [`wrapper.tool_input`][agents.run_context.RunContextWrapper.tool_input]。
|
||||
- 承認状態をプログラムで更新する必要がある場合は [`wrapper.approve_tool(...)`][agents.run_context.RunContextWrapper.approve_tool] / [`wrapper.reject_tool(...)`][agents.run_context.RunContextWrapper.reject_tool]。
|
||||
- [`wrapper.context`][agents.run_context.RunContextWrapper.context]: 独自の可変なアプリ状態と依存関係に使用します。
|
||||
- [`wrapper.usage`][agents.run_context.RunContextWrapper.usage]: 現在の実行全体で集計されたリクエストおよびトークン使用量に使用します。
|
||||
- [`wrapper.tool_input`][agents.run_context.RunContextWrapper.tool_input]: 現在の実行が [`Agent.as_tool()`][agents.agent.Agent.as_tool] の内部で実行されている場合の構造化入力に使用します。
|
||||
- [`wrapper.approve_tool(...)`][agents.run_context.RunContextWrapper.approve_tool] / [`wrapper.reject_tool(...)`][agents.run_context.RunContextWrapper.reject_tool]: 承認状態をプログラムで更新する必要がある場合に使用します。
|
||||
|
||||
アプリで定義したオブジェクトは `wrapper.context` のみです。その他のフィールドは SDK が管理するランタイムメタデータです。
|
||||
|
||||
後で human-in-the-loop や永続ジョブワークフロー向けに [`RunState`][agents.run_state.RunState] をシリアライズする場合、そのランタイムメタデータは状態とともに保存されます。シリアライズした状態を永続化または送信する予定がある場合は、[`RunContextWrapper.context`][agents.run_context.RunContextWrapper.context] にシークレットを入れないでください。
|
||||
後で human-in-the-loop や耐久ジョブワークフローのために [`RunState`][agents.run_state.RunState] をシリアライズする場合、そのランタイムメタデータは状態とともに保存されます。シリアライズされた状態を永続化または送信する予定がある場合、[`RunContextWrapper.context`][agents.run_context.RunContextWrapper.context] にシークレットを入れないでください。
|
||||
|
||||
会話状態は別の関心事項です。ターンをどのように引き継ぐかに応じて、`result.to_input_list()`、`session`、`conversation_id`、または `previous_response_id` を使用してください。この判断については [results](results.md)、[running agents](running_agents.md)、[sessions](sessions/index.md) を参照してください。
|
||||
会話状態は別の関心事です。ターンをどのように引き継ぐかに応じて、`result.to_input_list()`、`session`、`conversation_id`、または `previous_response_id` を使用してください。その判断については、[実行結果](results.md)、[エージェントの実行](running_agents.md)、[セッション](sessions/index.md) を参照してください。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -85,18 +85,18 @@ if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
1. これはコンテキストオブジェクトです。ここでは dataclass を使用していますが、任意の型を使用できます。
|
||||
2. これはツールです。`RunContextWrapper[UserInfo]` を受け取ることがわかります。ツール実装はコンテキストから読み取ります。
|
||||
3. 型チェッカーがエラーを検出できるように、エージェントをジェネリック `UserInfo` で指定します(たとえば、異なるコンテキスト型を受け取るツールを渡そうとした場合)。
|
||||
1. これがコンテキストオブジェクトです。ここでは dataclass を使用していますが、任意の型を使用できます。
|
||||
2. これはツールです。`RunContextWrapper[UserInfo]` を受け取っていることがわかります。ツール実装はコンテキストから読み取ります。
|
||||
3. 型チェッカーがエラーを検出できるように、エージェントにジェネリック `UserInfo` を指定します (たとえば、異なるコンテキスト型を受け取るツールを渡そうとした場合)。
|
||||
4. コンテキストは `run` 関数に渡されます。
|
||||
5. エージェントは正しくツールを呼び出し、年齢を取得します。
|
||||
|
||||
---
|
||||
|
||||
### 高度な使用法: `ToolContext`
|
||||
### 高度な内容: `ToolContext`
|
||||
|
||||
場合によっては、実行中のツールに関する追加メタデータ(名前、呼び出し ID、生の引数文字列など)にアクセスしたいことがあります。
|
||||
このために、`RunContextWrapper` を拡張する [`ToolContext`][agents.tool_context.ToolContext] クラスを使用できます。
|
||||
場合によっては、実行中のツールに関する追加メタデータ (名前、呼び出し ID、生の引数文字列など) にアクセスしたいことがあります。
|
||||
この場合、`RunContextWrapper` を拡張する [`ToolContext`][agents.tool_context.ToolContext] クラスを使用できます。
|
||||
|
||||
```python
|
||||
from typing import Annotated
|
||||
@@ -125,24 +125,24 @@ agent = Agent(
|
||||
```
|
||||
|
||||
`ToolContext` は `RunContextWrapper` と同じ `.context` プロパティを提供し、
|
||||
さらに現在のツール呼び出しに固有の追加フィールドも提供します。
|
||||
現在のツール呼び出しに固有の追加フィールドも提供します:
|
||||
|
||||
- `tool_name` – 呼び出されるツールの名前
|
||||
- `tool_call_id` – このツール呼び出しの一意識別子
|
||||
- `tool_arguments` – ツールに渡される生の引数文字列
|
||||
- `tool_namespace` – ツールが `tool_namespace()` または他の名前空間付きサーフェスを通じて読み込まれた場合の、ツール呼び出しの Responses 名前空間
|
||||
- `qualified_tool_name` – 名前空間が利用可能な場合に、その名前空間で修飾されたツール名
|
||||
- `tool_name` – 呼び出されているツールの名前
|
||||
- `tool_call_id` – このツール呼び出しの一意の識別子
|
||||
- `tool_arguments` – ツールに渡された生の引数文字列
|
||||
- `tool_namespace` – ツールが `tool_namespace()` または別の名前空間付きサーフェスを通じて読み込まれた場合の、ツール呼び出しに対する Responses 名前空間
|
||||
- `qualified_tool_name` – 名前空間が利用できる場合に、その名前空間で修飾されたツール名
|
||||
|
||||
実行中にツールレベルのメタデータが必要な場合は `ToolContext` を使用してください。
|
||||
エージェントとツール間の一般的なコンテキスト共有には、`RunContextWrapper` で十分です。`ToolContext` は `RunContextWrapper` を拡張しているため、ネストした `Agent.as_tool()` 実行が構造化入力を提供した場合は `.tool_input` も公開できます。
|
||||
実行中にツールレベルのメタデータが必要な場合は、`ToolContext` を使用してください。
|
||||
エージェントとツール間で一般的なコンテキスト共有を行うには、`RunContextWrapper` のままで十分です。`ToolContext` は `RunContextWrapper` を拡張しているため、ネストされた `Agent.as_tool()` 実行が構造化入力を提供した場合には `.tool_input` も公開できます。
|
||||
|
||||
---
|
||||
|
||||
## エージェント / LLM コンテキスト
|
||||
|
||||
LLM が呼び出されると、参照できるデータは会話履歴にあるもの **のみ** です。つまり、新しいデータを LLM で利用可能にしたい場合は、その履歴で利用できる形にする必要があります。方法はいくつかあります。
|
||||
LLM が呼び出されるとき、その LLM が参照できる **唯一の** データは会話履歴に含まれるものです。つまり、新しいデータを LLM に利用可能にしたい場合は、その履歴内で利用可能になるような方法で行う必要があります。これにはいくつかの方法があります:
|
||||
|
||||
1. エージェントの `instructions` に追加します。これは「システムプロンプト」または「開発者メッセージ」とも呼ばれます。システムプロンプトは静的文字列にもできますし、コンテキストを受け取って文字列を返す動的関数にもできます。これは、常に有用な情報(たとえばユーザー名や現在日付)に対する一般的な手法です。
|
||||
2. `Runner.run` 関数を呼び出す際の `input` に追加します。これは `instructions` の手法に似ていますが、[chain of command](https://cdn.openai.com/spec/model-spec-2024-05-08.html#follow-the-chain-of-command) でより下位のメッセージを持てます。
|
||||
3. 関数ツールを介して公開します。これは _オンデマンド_ のコンテキストに有用です。LLM がデータを必要とするタイミングを判断し、そのデータを取得するためにツールを呼び出せます。
|
||||
4. retrieval または Web 検索を使用します。これらは、ファイルやデータベース(retrieval)、または Web(Web 検索)から関連データを取得できる特別なツールです。これは、レスポンスを関連するコンテキストデータに「グラウンディング」するのに有用です。
|
||||
1. エージェントの `instructions` に追加できます。これは「システムプロンプト」または「開発者メッセージ」とも呼ばれます。システムプロンプトは静的文字列にも、コンテキストを受け取って文字列を出力する動的関数にもできます。これは、常に役立つ情報 (たとえば、ユーザーの名前や現在の日付) に対する一般的な手法です。
|
||||
2. `Runner.run` 関数を呼び出すときに `input` に追加します。これは `instructions` の手法に似ていますが、[指揮系統](https://cdn.openai.com/spec/model-spec-2024-05-08.html#follow-the-chain-of-command) においてより下位のメッセージにできます。
|
||||
3. 関数ツールを介して公開します。これは _オンデマンド_ のコンテキストに便利です。LLM がデータを必要とするタイミングを判断し、そのデータを取得するためにツールを呼び出せます。
|
||||
4. リトリーバルまたは Web 検索を使用します。これらは、ファイルやデータベースから関連データを取得する (リトリーバル)、または Web から取得する (Web 検索) ことができる特殊なツールです。これは、関連するコンテキストデータに基づいて応答を「グラウンディング」するのに便利です。
|
||||
+76
-90
@@ -4,139 +4,125 @@ search:
|
||||
---
|
||||
# コード例
|
||||
|
||||
[repo](https://github.com/openai/openai-agents-python/tree/main/examples) の examples セクションで、 SDK のさまざまなサンプル実装を確認できます。これらのコード例は、異なるパターンと機能を示す複数のカテゴリーに整理されています。
|
||||
SDK のさまざまなサンプル実装は、[リポジトリ](https://github.com/openai/openai-agents-python/tree/main/examples) のコード例セクションで確認できます。これらのコード例は、さまざまなパターンと機能を示す複数のカテゴリーに整理されています。
|
||||
|
||||
## カテゴリー
|
||||
|
||||
- **[agent_patterns](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns):**
|
||||
このカテゴリーのコード例では、次のような一般的なエージェント設計パターンを示します。
|
||||
- **[agent_patterns](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns):** このカテゴリーのコード例では、次のような一般的なエージェント設計パターンを示します。
|
||||
|
||||
- 決定論的ワークフロー
|
||||
- Agents as tools
|
||||
- ストリーミングイベントを伴う Agents as tools (`examples/agent_patterns/agents_as_tools_streaming.py`)
|
||||
- 構造化入力パラメーターを伴う Agents as tools (`examples/agent_patterns/agents_as_tools_structured.py`)
|
||||
- 並列エージェント実行
|
||||
- ストリーミングイベントを伴う Agents as tools(`examples/agent_patterns/agents_as_tools_streaming.py`)
|
||||
- 構造化された入力パラメーターを伴う Agents as tools(`examples/agent_patterns/agents_as_tools_structured.py`)
|
||||
- エージェントの並列実行
|
||||
- 条件付きツール使用
|
||||
- 異なる挙動でツール使用を強制する (`examples/agent_patterns/forcing_tool_use.py`)
|
||||
- 入力 / 出力ガードレール
|
||||
- 審査者としての LLM
|
||||
- 異なる動作でツール使用を強制(`examples/agent_patterns/forcing_tool_use.py`)
|
||||
- 入出力ガードレール
|
||||
- LLM を評価者として使用
|
||||
- ルーティング
|
||||
- ストリーミングガードレール
|
||||
- ツール承認と状態シリアライズを伴う Human-in-the-loop (`examples/agent_patterns/human_in_the_loop.py`)
|
||||
- ストリーミングを伴う Human-in-the-loop (`examples/agent_patterns/human_in_the_loop_stream.py`)
|
||||
- 承認フロー向けのカスタム拒否メッセージ (`examples/agent_patterns/human_in_the_loop_custom_rejection.py`)
|
||||
- ツール承認と状態のシリアライズを伴う人間参加型フロー(`examples/agent_patterns/human_in_the_loop.py`)
|
||||
- ストリーミングを伴う人間参加型フロー(`examples/agent_patterns/human_in_the_loop_stream.py`)
|
||||
- 承認フロー向けのカスタム拒否メッセージ(`examples/agent_patterns/human_in_the_loop_custom_rejection.py`)
|
||||
|
||||
- **[basic](https://github.com/openai/openai-agents-python/tree/main/examples/basic):**
|
||||
これらのコード例では、次のような SDK の基本機能を紹介します。
|
||||
- **[basic](https://github.com/openai/openai-agents-python/tree/main/examples/basic):** これらのコード例では、次のような SDK の基礎的な機能を紹介します。
|
||||
|
||||
- Hello world のコード例 (デフォルトモデル、 GPT-5、 open-weight モデル)
|
||||
- エージェントライフサイクル管理
|
||||
- Run hooks と agent hooks のライフサイクル例 (`examples/basic/lifecycle_example.py`)
|
||||
- Hello World のコード例(デフォルトモデル、GPT-5、オープンウェイトモデル)
|
||||
- エージェントのライフサイクル管理
|
||||
- 実行フックとエージェントフックのライフサイクルコード例(`examples/basic/lifecycle_example.py`)
|
||||
- 動的システムプロンプト
|
||||
- 基本的なツール使用 (`examples/basic/tools.py`)
|
||||
- ツール入力 / 出力ガードレール (`examples/basic/tool_guardrails.py`)
|
||||
- 画像ツール出力 (`examples/basic/image_tool_output.py`)
|
||||
- ストリーミング出力 (テキスト、項目、関数呼び出し引数)
|
||||
- 複数ターンで共有セッションヘルパーを使用する Responses websocket transport (`examples/basic/stream_ws.py`)
|
||||
- 基本的なツール使用(`examples/basic/tools.py`)
|
||||
- ツールの入出力ガードレール(`examples/basic/tool_guardrails.py`)
|
||||
- 画像ツール出力(`examples/basic/image_tool_output.py`)
|
||||
- ストリーミング出力(テキスト、項目、関数呼び出し引数)
|
||||
- ターンをまたいだ共有セッションヘルパーを使用する Responses WebSocket トランスポート(`examples/basic/stream_ws.py`)
|
||||
- プロンプトテンプレート
|
||||
- ファイル処理 (ローカルとリモート、画像と PDF)
|
||||
- 使用状況追跡
|
||||
- Runner 管理の再試行設定 (`examples/basic/retry.py`)
|
||||
- サードパーティアダプター経由の Runner 管理再試行 (`examples/basic/retry_litellm.py`)
|
||||
- 非 strict な出力型
|
||||
- 以前の response ID の使用
|
||||
- ファイル処理(ローカルとリモート、画像と PDF)
|
||||
- 使用状況トラッキング
|
||||
- Runner 管理の再試行設定(`examples/basic/retry.py`)
|
||||
- サードパーティアダプター経由の Runner 管理の再試行(`examples/basic/retry_litellm.py`)
|
||||
- 非厳密な出力型
|
||||
- 以前のレスポンス ID の使用
|
||||
|
||||
- **[customer_service](https://github.com/openai/openai-agents-python/tree/main/examples/customer_service):**
|
||||
航空会社向けのカスタマーサービスシステムのコード例です。
|
||||
- **[customer_service](https://github.com/openai/openai-agents-python/tree/main/examples/customer_service):** 航空会社向けのカスタマーサービスシステムのコード例です。
|
||||
|
||||
- **[financial_research_agent](https://github.com/openai/openai-agents-python/tree/main/examples/financial_research_agent):**
|
||||
金融データ分析のためのエージェントとツールを用いた、構造化された調査ワークフローを示す金融リサーチエージェントです。
|
||||
- **[financial_research_agent](https://github.com/openai/openai-agents-python/tree/main/examples/financial_research_agent):** 金融データ分析向けのエージェントとツールを使った、構造化されたリサーチワークフローを示す金融調査エージェントです。
|
||||
|
||||
- **[handoffs](https://github.com/openai/openai-agents-python/tree/main/examples/handoffs):**
|
||||
メッセージフィルタリングを含む、エージェントのハンドオフの実践的なコード例です。
|
||||
- **[handoffs](https://github.com/openai/openai-agents-python/tree/main/examples/handoffs):** メッセージフィルタリングを含む、エージェントのハンドオフの実践的なコード例です。内容は次のとおりです。
|
||||
|
||||
- メッセージフィルター例 (`examples/handoffs/message_filter.py`)
|
||||
- ストリーミングを伴うメッセージフィルター (`examples/handoffs/message_filter_streaming.py`)
|
||||
- メッセージフィルターのコード例(`examples/handoffs/message_filter.py`)
|
||||
- ストリーミングを伴うメッセージフィルター(`examples/handoffs/message_filter_streaming.py`)
|
||||
|
||||
- **[hosted_mcp](https://github.com/openai/openai-agents-python/tree/main/examples/hosted_mcp):**
|
||||
OpenAI Responses API で hosted MCP (Model Context Protocol) を使用する方法を示すコード例です。以下を含みます。
|
||||
- **[hosted_mcp](https://github.com/openai/openai-agents-python/tree/main/examples/hosted_mcp):** OpenAI Responses API でホスト型 MCP (Model Context Protocol) を使用する方法を示すコード例です。内容は次のとおりです。
|
||||
|
||||
- 承認なしのシンプルな hosted MCP (`examples/hosted_mcp/simple.py`)
|
||||
- Google Calendar などの MCP コネクター (`examples/hosted_mcp/connectors.py`)
|
||||
- 割り込みベース承認を伴う Human-in-the-loop (`examples/hosted_mcp/human_in_the_loop.py`)
|
||||
- MCP ツール呼び出しの on-approval コールバック (`examples/hosted_mcp/on_approval.py`)
|
||||
- 承認なしのシンプルなホスト型 MCP(`examples/hosted_mcp/simple.py`)
|
||||
- Google Calendar などの MCP コネクター(`examples/hosted_mcp/connectors.py`)
|
||||
- 割り込みベースの承認を伴う人間参加型フロー(`examples/hosted_mcp/human_in_the_loop.py`)
|
||||
- MCP ツール呼び出しに対する承認時コールバック(`examples/hosted_mcp/on_approval.py`)
|
||||
|
||||
- **[mcp](https://github.com/openai/openai-agents-python/tree/main/examples/mcp):**
|
||||
以下を含め、 MCP (Model Context Protocol) でエージェントを構築する方法を学べます。
|
||||
- **[mcp](https://github.com/openai/openai-agents-python/tree/main/examples/mcp):** MCP (Model Context Protocol) でエージェントを構築する方法を学べます。内容は次のとおりです。
|
||||
|
||||
- Filesystem のコード例
|
||||
- ファイルシステムのコード例
|
||||
- Git のコード例
|
||||
- MCP prompt server のコード例
|
||||
- MCP プロンプトサーバーのコード例
|
||||
- SSE (Server-Sent Events) のコード例
|
||||
- SSE リモートサーバー接続 (`examples/mcp/sse_remote_example`)
|
||||
- SSE リモートサーバー接続(`examples/mcp/sse_remote_example`)
|
||||
- Streamable HTTP のコード例
|
||||
- Streamable HTTP リモート接続 (`examples/mcp/streamable_http_remote_example`)
|
||||
- Streamable HTTP 向けカスタム HTTP client factory (`examples/mcp/streamablehttp_custom_client_example`)
|
||||
- `MCPUtil.get_all_function_tools` による全 MCP ツールの事前取得 (`examples/mcp/get_all_mcp_tools_example`)
|
||||
- FastAPI を使用した MCPServerManager (`examples/mcp/manager_example`)
|
||||
- MCP ツールフィルタリング (`examples/mcp/tool_filter_example`)
|
||||
- Streamable HTTP リモート接続(`examples/mcp/streamable_http_remote_example`)
|
||||
- Streamable HTTP 用のカスタム HTTP クライアントファクトリ(`examples/mcp/streamablehttp_custom_client_example`)
|
||||
- `MCPUtil.get_all_function_tools` によるすべての MCP ツールのプリフェッチ(`examples/mcp/get_all_mcp_tools_example`)
|
||||
- FastAPI を使用した MCPServerManager(`examples/mcp/manager_example`)
|
||||
- MCP ツールのフィルタリング(`examples/mcp/tool_filter_example`)
|
||||
|
||||
- **[memory](https://github.com/openai/openai-agents-python/tree/main/examples/memory):**
|
||||
エージェント向けのさまざまなメモリ実装のコード例です。以下を含みます。
|
||||
- **[memory](https://github.com/openai/openai-agents-python/tree/main/examples/memory):** エージェント向けのさまざまなメモリ実装のコード例です。内容は次のとおりです。
|
||||
|
||||
- SQLite セッションストレージ
|
||||
- 高度な SQLite セッションストレージ
|
||||
- Redis セッションストレージ
|
||||
- SQLAlchemy セッションストレージ
|
||||
- Dapr state store セッションストレージ
|
||||
- Dapr ステートストアセッションストレージ
|
||||
- 暗号化セッションストレージ
|
||||
- OpenAI Conversations セッションストレージ
|
||||
- Responses compaction セッションストレージ
|
||||
- `ModelSettings(store=False)` を使用したステートレスな Responses compaction (`examples/memory/compaction_session_stateless_example.py`)
|
||||
- ファイルベースのセッションストレージ (`examples/memory/file_session.py`)
|
||||
- Human-in-the-loop を伴うファイルベースセッション (`examples/memory/file_hitl_example.py`)
|
||||
- Human-in-the-loop を伴う SQLite インメモリセッション (`examples/memory/memory_session_hitl_example.py`)
|
||||
- Human-in-the-loop を伴う OpenAI Conversations セッション (`examples/memory/openai_session_hitl_example.py`)
|
||||
- セッションをまたぐ HITL 承認 / 拒否シナリオ (`examples/memory/hitl_session_scenario.py`)
|
||||
- Responses コンパクションセッションストレージ
|
||||
- `ModelSettings(store=False)` を使用したステートレスな Responses コンパクション(`examples/memory/compaction_session_stateless_example.py`)
|
||||
- ファイルバックエンドのセッションストレージ(`examples/memory/file_session.py`)
|
||||
- 人間参加型フローを伴うファイルバックエンドのセッション(`examples/memory/file_hitl_example.py`)
|
||||
- 人間参加型フローを伴う SQLite インメモリセッション(`examples/memory/memory_session_hitl_example.py`)
|
||||
- 人間参加型フローを伴う OpenAI Conversations セッション(`examples/memory/openai_session_hitl_example.py`)
|
||||
- セッションをまたぐ HITL 承認/拒否シナリオ(`examples/memory/hitl_session_scenario.py`)
|
||||
|
||||
- **[model_providers](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers):**
|
||||
カスタムプロバイダーやサードパーティアダプターを含め、 SDK で非 OpenAI モデルを使用する方法を確認できます。
|
||||
- **[model_providers](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers):** カスタムプロバイダーやサードパーティアダプターを含め、SDK で OpenAI 以外のモデルを使用する方法を確認できます。
|
||||
|
||||
- **[realtime](https://github.com/openai/openai-agents-python/tree/main/examples/realtime):**
|
||||
SDK を使用してリアルタイム体験を構築する方法を示すコード例です。以下を含みます。
|
||||
- **[realtime](https://github.com/openai/openai-agents-python/tree/main/examples/realtime):** SDK を使用してリアルタイム体験を構築する方法を示すコード例です。内容は次のとおりです。
|
||||
|
||||
- 構造化されたテキストおよび画像メッセージによる Web アプリケーションパターン
|
||||
- コマンドライン音声ループと再生処理
|
||||
- 構造化テキストメッセージと画像メッセージを使用する Web アプリケーションパターン
|
||||
- コマンドラインの音声ループと再生処理
|
||||
- WebSocket 経由の Twilio Media Streams 統合
|
||||
- Realtime Calls API attach フローを使用した Twilio SIP 統合
|
||||
- Realtime Calls API のアタッチフローを使用した Twilio SIP 統合
|
||||
|
||||
- **[reasoning_content](https://github.com/openai/openai-agents-python/tree/main/examples/reasoning_content):**
|
||||
reasoning content の扱い方を示すコード例です。以下を含みます。
|
||||
- **[reasoning_content](https://github.com/openai/openai-agents-python/tree/main/examples/reasoning_content):** 推論内容を扱う方法を示すコード例です。内容は次のとおりです。
|
||||
|
||||
- Runner API、ストリーミング、非ストリーミングでの reasoning content (`examples/reasoning_content/runner_example.py`)
|
||||
- OpenRouter 経由で OSS モデルを使用した reasoning content (`examples/reasoning_content/gpt_oss_stream.py`)
|
||||
- 基本的な reasoning content のコード例 (`examples/reasoning_content/main.py`)
|
||||
- Runner API による推論内容、ストリーミングおよび非ストリーミング(`examples/reasoning_content/runner_example.py`)
|
||||
- OpenRouter 経由の OSS モデルによる推論内容(`examples/reasoning_content/gpt_oss_stream.py`)
|
||||
- 基本的な推論内容のコード例(`examples/reasoning_content/main.py`)
|
||||
|
||||
- **[research_bot](https://github.com/openai/openai-agents-python/tree/main/examples/research_bot):**
|
||||
複雑なマルチエージェント調査ワークフローを示す、シンプルなディープリサーチクローンです。
|
||||
- **[research_bot](https://github.com/openai/openai-agents-python/tree/main/examples/research_bot):** 複雑なマルチエージェントリサーチワークフローを示す、シンプルなディープリサーチのクローンです。
|
||||
|
||||
- **[tools](https://github.com/openai/openai-agents-python/tree/main/examples/tools):**
|
||||
以下のような OpenAI がホストするツールと実験的な Codex ツール機能の実装方法を学べます。
|
||||
- **[tools](https://github.com/openai/openai-agents-python/tree/main/examples/tools):** OpenAI がホストするツールと、次のような実験的な Codex ツール機能の実装方法を学べます。
|
||||
|
||||
- Web 検索 とフィルター付き Web 検索
|
||||
- Web 検索とフィルター付き Web 検索
|
||||
- ファイル検索
|
||||
- Code interpreter
|
||||
- ファイル編集と承認を伴う apply patch ツール (`examples/tools/apply_patch.py`)
|
||||
- 承認コールバックを伴う shell ツール実行 (`examples/tools/shell.py`)
|
||||
- Human-in-the-loop 割り込みベース承認を伴う shell ツール (`examples/tools/shell_human_in_the_loop.py`)
|
||||
- インラインスキルを伴う hosted container shell (`examples/tools/container_shell_inline_skill.py`)
|
||||
- スキル参照を伴う hosted container shell (`examples/tools/container_shell_skill_reference.py`)
|
||||
- ローカルスキルを伴う local shell (`examples/tools/local_shell_skill.py`)
|
||||
- 名前空間と遅延ツールを伴うツール検索 (`examples/tools/tool_search.py`)
|
||||
- ファイル編集と承認を伴う Apply patch ツール(`examples/tools/apply_patch.py`)
|
||||
- 承認コールバックを伴うシェルツール実行(`examples/tools/shell.py`)
|
||||
- 割り込みベースの承認を伴う人間参加型シェルツール(`examples/tools/shell_human_in_the_loop.py`)
|
||||
- インラインスキルを伴うホスト型コンテナーシェル(`examples/tools/container_shell_inline_skill.py`)
|
||||
- スキル参照を伴うホスト型コンテナーシェル(`examples/tools/container_shell_skill_reference.py`)
|
||||
- ローカルスキルを伴うローカルシェル(`examples/tools/local_shell_skill.py`)
|
||||
- 名前空間と遅延ツールを伴うツール検索(`examples/tools/tool_search.py`)
|
||||
- コンピュータ操作
|
||||
- 画像生成
|
||||
- 実験的な Codex ツールワークフロー (`examples/tools/codex.py`)
|
||||
- 実験的な Codex 同一スレッドワークフロー (`examples/tools/codex_same_thread.py`)
|
||||
- 実験的な Codex ツールワークフロー(`examples/tools/codex.py`)
|
||||
- 実験的な Codex 同一スレッドワークフロー(`examples/tools/codex_same_thread.py`)
|
||||
|
||||
- **[voice](https://github.com/openai/openai-agents-python/tree/main/examples/voice):**
|
||||
ストリーミング音声のコード例を含む、 TTS および STT モデルを使用した音声エージェントのコード例を確認できます。
|
||||
- **[voice](https://github.com/openai/openai-agents-python/tree/main/examples/voice):** ストリーミング音声のコード例を含む、当社の TTS および STT モデルを使用した音声エージェントのコード例を確認できます。
|
||||
+43
-42
@@ -4,74 +4,75 @@ search:
|
||||
---
|
||||
# ガードレール
|
||||
|
||||
ガードレールを使うと、ユーザー入力とエージェント出力のチェックや検証を行えます。たとえば、顧客リクエスト対応のために非常に高性能(したがって低速 / 高コスト)なモデルを使うエージェントがあるとします。悪意のあるユーザーに、そのモデルで数学の宿題を手伝わせたくはありません。そのため、高速 / 低コストなモデルでガードレールを実行できます。ガードレールが悪意のある利用を検知した場合、すぐにエラーを発生させて高コストなモデルの実行を防げます。これにより時間とコストを節約できます( **blocking guardrails** を使う場合。並列ガードレールでは、ガードレール完了前に高コストなモデルがすでに実行を開始している可能性があります。詳細は下記の「実行モード」を参照してください)。
|
||||
ガードレールを使用すると、ユーザー入力とエージェント出力のチェックおよび検証を行えます。たとえば、顧客からのリクエストを支援するために非常に賢い(そのため低速で高コストな)モデルを使用するエージェントがあるとします。悪意のあるユーザーに、そのモデルへ数学の宿題を手伝わせたくはないはずです。そこで、高速 / 低コストなモデルでガードレールを実行できます。ガードレールが悪意のある使用を検出した場合、即座にエラーを送出して高価なモデルの実行を防ぎ、時間と費用を節約できます( **ブロッキングガードレールを使用する場合に限ります。並列ガードレールでは、ガードレールが完了する前に、高価なモデルがすでに実行を開始している可能性があります。詳細は下記の「実行モード」を参照してください** )。
|
||||
|
||||
ガードレールには 2 種類あります。
|
||||
|
||||
1. Input ガードレールは最初のユーザー入力で実行されます
|
||||
2. Output ガードレールは最終的なエージェント出力で実行されます
|
||||
1. 入力ガードレールは、最初のユーザー入力に対して実行されます。
|
||||
2. 出力ガードレールは、最終的なエージェント出力に対して実行されます。
|
||||
|
||||
## ワークフロー境界
|
||||
## ワークフローの境界
|
||||
|
||||
ガードレールはエージェントとツールにアタッチされますが、ワークフロー内の同じタイミングで実行されるわけではありません。
|
||||
ガードレールはエージェントとツールにアタッチされますが、すべてがワークフロー内の同じ時点で実行されるわけではありません。
|
||||
|
||||
- **Input ガードレール** はチェーン内の最初のエージェントに対してのみ実行されます。
|
||||
- **Output ガードレール** は最終出力を生成するエージェントに対してのみ実行されます。
|
||||
- **ツールガードレール** はカスタム関数ツールの呼び出しごとに実行され、Input ガードレールは実行前、Output ガードレールは実行後に実行されます。
|
||||
- **入力ガードレール** は、チェーン内の最初のエージェントに対してのみ実行されます。
|
||||
- **出力ガードレール** は、最終出力を生成するエージェントに対してのみ実行されます。
|
||||
- **ツールガードレール** は、すべてのカスタム関数ツール呼び出しで実行され、実行前に入力ガードレール、実行後に出力ガードレールが実行されます。
|
||||
|
||||
manager、ハンドオフ、または委譲された specialist を含むワークフローで、カスタム関数ツール呼び出しごとにチェックが必要な場合は、エージェントレベルの Input / Output ガードレールのみに頼るのではなく、ツールガードレールを使用してください。
|
||||
マネージャー、ハンドオフ、または委任先の専門家を含むワークフローで、各カスタム関数ツール呼び出しの前後にチェックが必要な場合は、エージェントレベルの入力 / 出力ガードレールだけに頼るのではなく、ツールガードレールを使用してください。
|
||||
|
||||
## Input ガードレール
|
||||
## 入力ガードレール
|
||||
|
||||
Input ガードレールは 3 ステップで実行されます。
|
||||
入力ガードレールは 3 ステップで実行されます。
|
||||
|
||||
1. まず、ガードレールはエージェントに渡されたものと同じ入力を受け取ります。
|
||||
2. 次に、ガードレール関数が実行されて [`GuardrailFunctionOutput`][agents.guardrail.GuardrailFunctionOutput] を生成し、それが [`InputGuardrailResult`][agents.guardrail.InputGuardrailResult] にラップされます
|
||||
3. 最後に、[`.tripwire_triggered`][agents.guardrail.GuardrailFunctionOutput.tripwire_triggered] が true かどうかを確認します。true の場合は [`InputGuardrailTripwireTriggered`][agents.exceptions.InputGuardrailTripwireTriggered] 例外が発生するため、ユーザーへの適切な応答や例外処理を行えます。
|
||||
2. 次に、ガードレール関数が実行され、 [`GuardrailFunctionOutput`][agents.guardrail.GuardrailFunctionOutput] が生成されます。これはその後 [`InputGuardrailResult`][agents.guardrail.InputGuardrailResult] にラップされます。
|
||||
3. 最後に、 [`.tripwire_triggered`][agents.guardrail.GuardrailFunctionOutput.tripwire_triggered] が true かどうかを確認します。true の場合、 [`InputGuardrailTripwireTriggered`][agents.exceptions.InputGuardrailTripwireTriggered] 例外が送出されるため、ユーザーに適切に応答したり、例外を処理したりできます。
|
||||
|
||||
!!! Note
|
||||
|
||||
Input ガードレールはユーザー入力に対して実行することを想定しているため、エージェントのガードレールはそのエージェントが *最初* のエージェントである場合にのみ実行されます。`guardrails` プロパティが `Runner.run` に渡されるのではなくエージェント側にある理由は何か、と疑問に思うかもしれません。これは、ガードレールが実際の Agent に関連することが多く、エージェントごとに異なるガードレールを実行するため、コードを同じ場所に置くことで可読性が向上するためです。
|
||||
入力ガードレールはユーザー入力に対して実行されることを想定しているため、エージェントのガードレールは、そのエージェントが *最初の* エージェントである場合にのみ実行されます。なぜ `guardrails` プロパティが `Runner.run` に渡されるのではなく、エージェント上にあるのか疑問に思うかもしれません。これは、ガードレールが実際のエージェントに関連していることが多いためです。エージェントごとに異なるガードレールを実行するため、コードを同じ場所に配置しておくと可読性の面で役立ちます。
|
||||
|
||||
### 実行モード
|
||||
|
||||
Input ガードレールは 2 つの実行モードをサポートしています。
|
||||
入力ガードレールは 2 つの実行モードをサポートします。
|
||||
|
||||
- **並列実行**(デフォルト、`run_in_parallel=True`): ガードレールはエージェント実行と同時に並行して実行されます。両方が同時に開始されるため、レイテンシの面で最も有利です。ただし、ガードレールが失敗した場合、キャンセルされる前にエージェントがすでにトークンを消費し、ツールを実行している可能性があります。
|
||||
- **並列実行** (デフォルト、 `run_in_parallel=True` ): ガードレールはエージェントの実行と並行して実行されます。両方が同時に開始されるため、レイテンシが最も良くなります。ただし、ガードレールが失敗した場合、キャンセルされる前にエージェントがすでにトークンを消費し、ツールを実行している可能性があります。
|
||||
|
||||
- **ブロッキング実行**(`run_in_parallel=False`): ガードレールはエージェント開始 *前* に実行され、完了します。ガードレールの tripwire がトリガーされた場合、エージェントは実行されないため、トークン消費とツール実行を防げます。これはコスト最適化に理想的で、ツール呼び出しによる潜在的な副作用を避けたい場合にも適しています。
|
||||
- **ブロッキング実行** ( `run_in_parallel=False` ): ガードレールはエージェントが開始する *前に* 実行され、完了します。ガードレールのトリップワイヤーが発火した場合、エージェントは一切実行されないため、トークン消費とツール実行を防げます。これは、コスト最適化や、ツール呼び出しによる潜在的な副作用を避けたい場合に最適です。
|
||||
|
||||
## Output ガードレール
|
||||
## 出力ガードレール
|
||||
|
||||
Output ガードレールは 3 ステップで実行されます。
|
||||
出力ガードレールは 3 ステップで実行されます。
|
||||
|
||||
1. まず、ガードレールはエージェントが生成した出力を受け取ります。
|
||||
2. 次に、ガードレール関数が実行されて [`GuardrailFunctionOutput`][agents.guardrail.GuardrailFunctionOutput] を生成し、それが [`OutputGuardrailResult`][agents.guardrail.OutputGuardrailResult] にラップされます
|
||||
3. 最後に、[`.tripwire_triggered`][agents.guardrail.GuardrailFunctionOutput.tripwire_triggered] が true かどうかを確認します。true の場合は [`OutputGuardrailTripwireTriggered`][agents.exceptions.OutputGuardrailTripwireTriggered] 例外が発生するため、ユーザーへの適切な応答や例外処理を行えます。
|
||||
1. まず、ガードレールはエージェントによって生成された出力を受け取ります。
|
||||
2. 次に、ガードレール関数が実行され、 [`GuardrailFunctionOutput`][agents.guardrail.GuardrailFunctionOutput] が生成されます。これはその後 [`OutputGuardrailResult`][agents.guardrail.OutputGuardrailResult] にラップされます。
|
||||
3. 最後に、 [`.tripwire_triggered`][agents.guardrail.GuardrailFunctionOutput.tripwire_triggered] が true かどうかを確認します。true の場合、 [`OutputGuardrailTripwireTriggered`][agents.exceptions.OutputGuardrailTripwireTriggered] 例外が送出されるため、ユーザーに適切に応答したり、例外を処理したりできます。
|
||||
|
||||
!!! Note
|
||||
|
||||
Output ガードレールは最終的なエージェント出力に対して実行することを想定しているため、エージェントのガードレールはそのエージェントが *最後* のエージェントである場合にのみ実行されます。Input ガードレールと同様に、これはガードレールが実際の Agent に関連することが多く、エージェントごとに異なるガードレールを実行するため、コードを同じ場所に置くことで可読性が向上するためです。
|
||||
出力ガードレールは最終的なエージェント出力に対して実行されることを想定しているため、エージェントのガードレールは、そのエージェントが *最後の* エージェントである場合にのみ実行されます。入力ガードレールと同様に、これはガードレールが実際のエージェントに関連していることが多いためです。エージェントごとに異なるガードレールを実行するため、コードを同じ場所に配置しておくと可読性の面で役立ちます。
|
||||
|
||||
Output ガードレールは常にエージェント完了後に実行されるため、`run_in_parallel` パラメーターはサポートしていません。
|
||||
出力ガードレールは必ずエージェントの完了後に実行されるため、 `run_in_parallel` パラメーターはサポートしません。
|
||||
|
||||
## ツールガードレール
|
||||
|
||||
ツールガードレールは **function tools** をラップし、実行の前後でツール呼び出しを検証またはブロックできます。設定はツール自体に対して行い、そのツールが呼び出されるたびに実行されます。
|
||||
ツールガードレールは **関数ツール** をラップし、実行の前後でツール呼び出しを検証またはブロックできるようにします。これはツール自体に設定され、そのツールが呼び出されるたびに実行されます。
|
||||
|
||||
- Input ツールガードレールはツール実行前に実行され、呼び出しをスキップする、メッセージで出力を置き換える、または tripwire を発生させることができます。
|
||||
- Output ツールガードレールはツール実行後に実行され、出力を置き換えるか、tripwire を発生させることができます。
|
||||
- ツールガードレールは [`function_tool`][agents.tool.function_tool] で作成された関数ツールにのみ適用されます。ハンドオフは通常の関数ツールパイプラインではなく SDK のハンドオフパイプラインを通るため、ツールガードレールはハンドオフ呼び出し自体には適用されません。Hosted ツール(`WebSearchTool`、`FileSearchTool`、`HostedMCPTool`、`CodeInterpreterTool`、`ImageGenerationTool`)および組み込み実行ツール(`ComputerTool`、`ShellTool`、`ApplyPatchTool`、`LocalShellTool`)もこのガードレールパイプラインを使用せず、[`Agent.as_tool()`][agents.agent.Agent.as_tool] でも現在はツールガードレールオプションを直接公開していません。
|
||||
- 入力ツールガードレールはツール実行前に実行され、呼び出しをスキップしたり、出力をメッセージに置き換えたり、トリップワイヤーを送出したりできます。
|
||||
- 出力ツールガードレールはツール実行後に実行され、出力を置き換えたり、トリップワイヤーを送出したりできます。
|
||||
- 関数ツールに承認が必要な場合、入力ツールガードレールは通常、承認後かつ実行直前に実行されます。保留中の承認割り込みが発行される前にこれらの入力チェックを実行したい場合は、 [`RunConfig.tool_execution`][agents.run.RunConfig.tool_execution] を [`ToolExecutionConfig(pre_approval_tool_input_guardrails=True)`][agents.run.ToolExecutionConfig] に設定してください。この承認前チェックに合格した呼び出しも、ツールが実行される前に、承認後に再度チェックされます。
|
||||
- ツールガードレールは、 [`function_tool`][agents.tool.function_tool] で作成された関数ツールにのみ適用されます。ハンドオフは通常の関数ツールパイプラインではなく、 SDK のハンドオフパイプラインを通じて実行されるため、ツールガードレールはハンドオフ呼び出し自体には適用されません。ホスト型ツール( `WebSearchTool` 、 `FileSearchTool` 、 `HostedMCPTool` 、 `CodeInterpreterTool` 、 `ImageGenerationTool` )と組み込み実行ツール( `ComputerTool` 、 `ShellTool` 、 `ApplyPatchTool` 、 `LocalShellTool` )もこのガードレールパイプラインを使用しません。また、 [`Agent.as_tool()`][agents.agent.Agent.as_tool] は現在、ツールガードレールのオプションを直接公開していません。
|
||||
|
||||
詳細は以下のコードスニペットを参照してください。
|
||||
詳細は、下記のコードスニペットを参照してください。
|
||||
|
||||
## トリップワイヤー
|
||||
|
||||
入力または出力がガードレールに失敗した場合、Guardrail は tripwire でこれを通知できます。tripwire がトリガーされたガードレールを検知すると、直ちに `{Input,Output}GuardrailTripwireTriggered` 例外を発生させ、Agent の実行を停止します。
|
||||
入力または出力がガードレールに合格しなかった場合、ガードレールはトリップワイヤーでこれを知らせることができます。トリップワイヤーが発火したガードレールを検出した時点で、即座に `{Input,Output}GuardrailTripwireTriggered` 例外を送出し、エージェントの実行を停止します。
|
||||
|
||||
## ガードレール実装
|
||||
## ガードレールの実装
|
||||
|
||||
入力を受け取り、[`GuardrailFunctionOutput`][agents.guardrail.GuardrailFunctionOutput] を返す関数を提供する必要があります。この例では、内部で Agent を実行してこれを実現します。
|
||||
入力を受け取り、 [`GuardrailFunctionOutput`][agents.guardrail.GuardrailFunctionOutput] を返す関数を用意する必要があります。この例では、内部でエージェントを実行することでこれを行います。
|
||||
|
||||
```python
|
||||
from pydantic import BaseModel
|
||||
@@ -124,12 +125,12 @@ async def main():
|
||||
print("Math homework guardrail tripped")
|
||||
```
|
||||
|
||||
1. このエージェントをガードレール関数内で使用します。
|
||||
2. これはエージェントの入力 / コンテキストを受け取り、結果を返すガードレール関数です。
|
||||
3. ガードレール結果には追加情報を含められます。
|
||||
4. これはワークフローを定義する実際のエージェントです。
|
||||
1. このエージェントをガードレール関数で使用します。
|
||||
2. これは、エージェントの入力 / コンテキストを受け取り、実行結果を返すガードレール関数です。
|
||||
3. ガードレールの実行結果に追加情報を含めることができます。
|
||||
4. これは、ワークフローを定義する実際のエージェントです。
|
||||
|
||||
Output ガードレールも同様です。
|
||||
出力ガードレールも同様です。
|
||||
|
||||
```python
|
||||
from pydantic import BaseModel
|
||||
@@ -182,12 +183,12 @@ async def main():
|
||||
print("Math output guardrail tripped")
|
||||
```
|
||||
|
||||
1. これは実際のエージェントの出力型です。
|
||||
2. これはガードレールの出力型です。
|
||||
3. これはエージェントの出力を受け取り、結果を返すガードレール関数です。
|
||||
4. これはワークフローを定義する実際のエージェントです。
|
||||
1. これは、実際のエージェントの出力型です。
|
||||
2. これは、ガードレールの出力型です。
|
||||
3. これは、エージェントの出力を受け取り、実行結果を返すガードレール関数です。
|
||||
4. これは、ワークフローを定義する実際のエージェントです。
|
||||
|
||||
最後に、ツールガードレールの例を示します。
|
||||
最後に、ツールガードレールのコード例を示します。
|
||||
|
||||
```python
|
||||
import json
|
||||
|
||||
+41
-41
@@ -4,21 +4,21 @@ search:
|
||||
---
|
||||
# ハンドオフ
|
||||
|
||||
ハンドオフを使うと、あるエージェントが別のエージェントにタスクを委譲できます。これは、異なるエージェントがそれぞれ異なる領域を専門にしているシナリオで特に有用です。たとえば、カスタマーサポートアプリでは、注文状況、返金、 FAQ などのタスクをそれぞれ専任で処理するエージェントを用意できます。
|
||||
ハンドオフにより、エージェントはタスクを別のエージェントに委任できます。これは、異なるエージェントがそれぞれ別の領域を専門とするシナリオで特に役立ちます。たとえば、カスタマーサポートアプリには、注文ステータス、返金、 FAQ などのタスクをそれぞれ専門に扱うエージェントがあるかもしれません。
|
||||
|
||||
ハンドオフは LLM に対してツールとして表現されます。したがって、`Refund Agent` という名前のエージェントへのハンドオフがある場合、そのツール名は `transfer_to_refund_agent` になります。
|
||||
ハンドオフは LLM に対してツールとして表現されます。そのため、 `Refund Agent` という名前のエージェントへのハンドオフがある場合、そのツールは `transfer_to_refund_agent` と呼ばれます。
|
||||
|
||||
## ハンドオフの作成
|
||||
|
||||
すべてのエージェントには [`handoffs`][agents.agent.Agent.handoffs] パラメーターがあり、`Agent` を直接渡すことも、ハンドオフをカスタマイズする `Handoff` オブジェクトを渡すこともできます。
|
||||
すべてのエージェントには [`handoffs`][agents.agent.Agent.handoffs] パラメーターがあり、 `Agent` を直接受け取ることも、ハンドオフをカスタマイズする `Handoff` オブジェクトを受け取ることもできます。
|
||||
|
||||
プレーンな `Agent` インスタンスを渡す場合、[`handoff_description`][agents.agent.Agent.handoff_description](設定されている場合)がデフォルトのツール説明に追記されます。これを使うと、完全な `handoff()` オブジェクトを書かなくても、どのときにそのハンドオフをモデルが選ぶべきかを示せます。
|
||||
通常の `Agent` インスタンスを渡す場合、その [`handoff_description`][agents.agent.Agent.handoff_description] (設定されている場合)がデフォルトのツール説明に追加されます。完全な `handoff()` オブジェクトを書かずに、そのハンドオフをモデルが選ぶべきタイミングを示唆するために使用してください。
|
||||
|
||||
Agents SDK が提供する [`handoff()`][agents.handoffs.handoff] 関数を使ってハンドオフを作成できます。この関数では、ハンドオフ先のエージェントに加えて、任意のオーバーライドや input filter を指定できます。
|
||||
Agents SDK が提供する [`handoff()`][agents.handoffs.handoff] 関数を使用してハンドオフを作成できます。この関数では、必要に応じた上書きや入力フィルターとともに、引き渡し先のエージェントを指定できます。
|
||||
|
||||
### 基本的な使い方
|
||||
### 基本的な使用法
|
||||
|
||||
シンプルなハンドオフは次のように作成できます。
|
||||
シンプルなハンドオフを作成する方法は次のとおりです。
|
||||
|
||||
```python
|
||||
from agents import Agent, handoff
|
||||
@@ -30,22 +30,22 @@ refund_agent = Agent(name="Refund agent")
|
||||
triage_agent = Agent(name="Triage agent", handoffs=[billing_agent, handoff(refund_agent)])
|
||||
```
|
||||
|
||||
1. エージェントを直接(`billing_agent` のように)使うことも、`handoff()` 関数を使うこともできます。
|
||||
1. エージェントを直接使用することも( `billing_agent` のように)、 `handoff()` 関数を使用することもできます。
|
||||
|
||||
### `handoff()` 関数によるハンドオフのカスタマイズ
|
||||
|
||||
[`handoff()`][agents.handoffs.handoff] 関数を使うと、さまざまなカスタマイズができます。
|
||||
[`handoff()`][agents.handoffs.handoff] 関数を使用すると、さまざまな項目をカスタマイズできます。
|
||||
|
||||
- `agent`: ハンドオフ先のエージェントです。
|
||||
- `tool_name_override`: デフォルトでは `Handoff.default_tool_name()` 関数が使われ、`transfer_to_<agent_name>` に解決されます。これをオーバーライドできます。
|
||||
- `tool_description_override`: `Handoff.default_tool_description()` のデフォルトツール説明をオーバーライドします。
|
||||
- `on_handoff`: ハンドオフが呼び出されたときに実行されるコールバック関数です。ハンドオフ呼び出しが分かった時点でデータ取得を開始する、といった用途に有用です。この関数はエージェントコンテキストを受け取り、任意で LLM が生成した入力も受け取れます。入力データは `input_type` パラメーターで制御されます。
|
||||
- `input_type`: ハンドオフのツール呼び出し引数のスキーマです。設定すると、パース済みペイロードが `on_handoff` に渡されます。
|
||||
- `input_filter`: 次のエージェントが受け取る入力をフィルタリングできます。詳細は下記を参照してください。
|
||||
- `is_enabled`: ハンドオフを有効にするかどうかです。boolean または boolean を返す関数を指定でき、実行時に動的に有効 / 無効を切り替えられます。
|
||||
- `nest_handoff_history`: RunConfig レベルの `nest_handoff_history` 設定を呼び出し単位で上書きする任意設定です。`None` の場合、アクティブな実行設定で定義された値が代わりに使われます。
|
||||
- `agent`: 処理を引き渡す先のエージェントです。
|
||||
- `tool_name_override`: デフォルトでは `Handoff.default_tool_name()` 関数が使用され、 `transfer_to_<agent_name>` に解決されます。これは上書きできます。
|
||||
- `tool_description_override`: `Handoff.default_tool_description()` から得られるデフォルトのツール説明を上書きします。
|
||||
- `on_handoff`: ハンドオフが呼び出されたときに実行されるコールバック関数です。ハンドオフが呼び出されることが分かった時点ですぐにデータ取得を開始する、といった用途に便利です。この関数はエージェントコンテキストを受け取り、任意で LLM が生成した入力も受け取れます。入力データは `input_type` パラメーターによって制御されます。
|
||||
- `input_type`: ハンドオフツール呼び出し引数のスキーマです。設定されている場合、解析されたペイロードが `on_handoff` に渡されます。
|
||||
- `input_filter`: これにより、次のエージェントが受け取る入力をフィルタリングできます。詳細は以下を参照してください。
|
||||
- `is_enabled`: ハンドオフが有効かどうかです。これはブール値、またはブール値を返す関数にでき、実行時にハンドオフを動的に有効化または無効化できます。
|
||||
- `nest_handoff_history`: RunConfig レベルの `nest_handoff_history` 設定に対する、呼び出しごとの任意の上書きです。 `None` の場合は、アクティブな実行設定で定義された値が代わりに使用されます。
|
||||
|
||||
[`handoff()`][agents.handoffs.handoff] ヘルパーは、常に渡された特定の `agent` に制御を移します。遷移先候補が複数ある場合は、遷移先ごとにハンドオフを 1 つずつ登録し、モデルにその中から選ばせてください。独自のハンドオフコードが呼び出し時に返すエージェントを決定する必要がある場合にのみ、カスタム [`Handoff`][agents.handoffs.Handoff] を使用してください。
|
||||
[`handoff()`][agents.handoffs.handoff] ヘルパーは、渡された特定の `agent` に常に制御を移します。複数の宛先候補がある場合は、宛先ごとに 1 つのハンドオフを登録し、モデルにその中から選ばせてください。独自のハンドオフコードが呼び出し時にどのエージェントを返すかを決定する必要がある場合にのみ、カスタム [`Handoff`][agents.handoffs.Handoff] を使用してください。
|
||||
|
||||
```python
|
||||
from agents import Agent, handoff, RunContextWrapper
|
||||
@@ -65,7 +65,7 @@ handoff_obj = handoff(
|
||||
|
||||
## ハンドオフ入力
|
||||
|
||||
状況によっては、ハンドオフを呼び出すときに LLM にデータを渡してほしいことがあります。たとえば「Escalation agent」へのハンドオフを考えてみてください。ログに記録できるよう、理由を渡してほしい場合があります。
|
||||
状況によっては、 LLM がハンドオフを呼び出すときに何らかのデータを提供してほしい場合があります。たとえば、「エスカレーションエージェント」へのハンドオフを想像してみてください。ログに記録できるように、モデルに理由を提供してほしい場合があります。
|
||||
|
||||
```python
|
||||
from pydantic import BaseModel
|
||||
@@ -87,44 +87,44 @@ handoff_obj = handoff(
|
||||
)
|
||||
```
|
||||
|
||||
`input_type` は、ハンドオフツール呼び出し自体の引数を記述します。SDK はそのスキーマをハンドオフツールの `parameters` としてモデルに公開し、返された JSON をローカルで検証して、パース済みの値を `on_handoff` に渡します。
|
||||
`input_type` は、ハンドオフツール呼び出し自体の引数を表します。 SDK はそのスキーマをハンドオフツールの `parameters` としてモデルに公開し、返された JSON をローカルで検証して、解析済みの値を `on_handoff` に渡します。
|
||||
|
||||
これは次のエージェントのメイン入力を置き換えるものではなく、遷移先を変更するものでもありません。[`handoff()`][agents.handoffs.handoff] ヘルパーは、引き続きラップした特定のエージェントへハンドオフします。また、受信側エージェントは、[`input_filter`][agents.handoffs.Handoff.input_filter] やネストされたハンドオフ履歴設定で変更しない限り、会話履歴を引き続き参照します。
|
||||
これは次のエージェントのメイン入力を置き換えるものではなく、別の宛先を選択するものでもありません。 [`handoff()`][agents.handoffs.handoff] ヘルパーは引き続き、ラップした特定のエージェントへ転送し、受け取り側のエージェントは [`input_filter`][agents.handoffs.Handoff.input_filter] またはネストされたハンドオフ履歴設定で変更しない限り、引き続き会話履歴を参照します。
|
||||
|
||||
`input_type` は [`RunContextWrapper.context`][agents.run_context.RunContextWrapper.context] とも別物です。`input_type` は、ハンドオフ時にモデルが決定するメタデータに使い、ローカルですでに持っているアプリケーション状態や依存関係には使わないでください。
|
||||
`input_type` は [`RunContextWrapper.context`][agents.run_context.RunContextWrapper.context] とも別のものです。ローカルにすでにあるアプリケーション状態や依存関係ではなく、ハンドオフ時にモデルが決定するメタデータには `input_type` を使用してください。
|
||||
|
||||
### `input_type` を使うタイミング
|
||||
### `input_type` の使用タイミング
|
||||
|
||||
ハンドオフに `reason`、`language`、`priority`、`summary` のような、モデル生成の小さなメタデータが必要な場合に `input_type` を使ってください。たとえば、トリアージエージェントは `{ "reason": "duplicate_charge", "priority": "high" }` を付けて返金エージェントへハンドオフでき、`on_handoff` は返金エージェントに制御が移る前にそのメタデータをログ化または永続化できます。
|
||||
ハンドオフに `reason` 、 `language` 、 `priority` 、 `summary` など、モデルが生成する小さなメタデータが必要な場合に `input_type` を使用してください。たとえば、トリアージエージェントは `{ "reason": "duplicate_charge", "priority": "high" }` とともに返金エージェントへハンドオフでき、返金エージェントが引き継ぐ前に `on_handoff` でそのメタデータをログに記録したり永続化したりできます。
|
||||
|
||||
目的が異なる場合は、別の仕組みを選んでください。
|
||||
|
||||
- 既存のアプリケーション状態と依存関係は [`RunContextWrapper.context`][agents.run_context.RunContextWrapper.context] に入れてください。[context ガイド](context.md)を参照してください。
|
||||
- 受信側エージェントが見る履歴を変更したい場合は、[`input_filter`][agents.handoffs.Handoff.input_filter]、[`RunConfig.nest_handoff_history`][agents.run.RunConfig.nest_handoff_history]、または [`RunConfig.handoff_history_mapper`][agents.run.RunConfig.handoff_history_mapper] を使ってください。
|
||||
- 複数の専門エージェントが候補にある場合は、遷移先ごとにハンドオフを 1 つずつ登録してください。`input_type` は選ばれたハンドオフにメタデータを追加できますが、遷移先の振り分けはしません。
|
||||
- 会話を転送せずにネストされた専門エージェント向けの構造化入力が欲しい場合は、[`Agent.as_tool(parameters=...)`][agents.agent.Agent.as_tool] を優先してください。[tools](tools.md#structured-input-for-tool-agents)を参照してください。
|
||||
- 既存のアプリケーション状態と依存関係は [`RunContextWrapper.context`][agents.run_context.RunContextWrapper.context] に置いてください。[コンテキストガイド](context.md)を参照してください。
|
||||
- 受け取り側のエージェントが参照する履歴を変更したい場合は、 [`input_filter`][agents.handoffs.Handoff.input_filter] 、 [`RunConfig.nest_handoff_history`][agents.run.RunConfig.nest_handoff_history] 、または [`RunConfig.handoff_history_mapper`][agents.run.RunConfig.handoff_history_mapper] を使用してください。
|
||||
- 複数の専門エージェント候補がある場合は、宛先ごとに 1 つのハンドオフを登録してください。 `input_type` は選択されたハンドオフにメタデータを追加できますが、宛先間の振り分けは行いません。
|
||||
- 会話を引き渡さずに、ネストされた専門エージェントに構造化入力を渡したい場合は、 [`Agent.as_tool(parameters=...)`][agents.agent.Agent.as_tool] を優先してください。[ツール](tools.md#structured-input-for-tool-agents)を参照してください。
|
||||
|
||||
## input filter
|
||||
## 入力フィルター
|
||||
|
||||
ハンドオフが発生すると、新しいエージェントが会話を引き継ぎ、以前の会話履歴全体を参照できる状態になります。これを変更したい場合は、[`input_filter`][agents.handoffs.Handoff.input_filter] を設定できます。input filter は、既存入力を [`HandoffInputData`][agents.handoffs.HandoffInputData] 経由で受け取り、新しい `HandoffInputData` を返す関数です。
|
||||
ハンドオフが発生すると、新しいエージェントが会話を引き継ぎ、以前の会話履歴全体を参照できるようになります。これを変更したい場合は、 [`input_filter`][agents.handoffs.Handoff.input_filter] を設定できます。入力フィルターは、 [`HandoffInputData`][agents.handoffs.HandoffInputData] を通じて既存の入力を受け取り、新しい `HandoffInputData` を返す必要がある関数です。
|
||||
|
||||
[`HandoffInputData`][agents.handoffs.HandoffInputData] には次が含まれます。
|
||||
|
||||
- `input_history`: `Runner.run(...)` 開始前の入力履歴。
|
||||
- `pre_handoff_items`: ハンドオフが呼び出されたエージェントターンより前に生成されたアイテム。
|
||||
- `new_items`: 現在のターン中に生成されたアイテム(ハンドオフ呼び出しとハンドオフ出力アイテムを含む)。
|
||||
- `input_items`: `new_items` の代わりに次のエージェントへ渡す任意のアイテム。これにより、セッション履歴用に `new_items` を保ったまま、モデル入力をフィルタリングできます。
|
||||
- `run_context`: ハンドオフ呼び出し時点でアクティブな [`RunContextWrapper`][agents.run_context.RunContextWrapper]。
|
||||
- `input_history`: `Runner.run(...)` が開始する前の入力履歴です。
|
||||
- `pre_handoff_items`: ハンドオフが呼び出されたエージェントターンより前に生成されたアイテムです。
|
||||
- `new_items`: ハンドオフ呼び出しとハンドオフ出力アイテムを含む、現在のターン中に生成されたアイテムです。
|
||||
- `input_items`: セッション履歴用に `new_items` をそのまま保ちながらモデル入力をフィルタリングできるよう、 `new_items` の代わりに次のエージェントへ転送する任意のアイテムです。
|
||||
- `run_context`: ハンドオフが呼び出された時点でアクティブな [`RunContextWrapper`][agents.run_context.RunContextWrapper] です。
|
||||
|
||||
ネストされたハンドオフは opt-in のベータとして提供されており、安定化のためデフォルトでは無効です。[`RunConfig.nest_handoff_history`][agents.run.RunConfig.nest_handoff_history] を有効にすると、runner はそれまでの transcript を 1 つの assistant 要約メッセージに折りたたみ、同一 run 中に複数のハンドオフが起きると新しいターンが追記され続ける `<CONVERSATION HISTORY>` ブロックに包みます。完全な `input_filter` を書かずに生成メッセージを置き換えたい場合は、[`RunConfig.handoff_history_mapper`][agents.run.RunConfig.handoff_history_mapper] で独自のマッピング関数を渡せます。この opt-in は、ハンドオフ側と run 側のいずれも明示的な `input_filter` を指定していない場合にのみ適用されるため、すでにペイロードをカスタマイズしている既存コード(このリポジトリのコード例を含む)は変更なしで現在の挙動を維持します。[`handoff(...)`][agents.handoffs.handoff] に `nest_handoff_history=True` または `False` を渡すことで、単一ハンドオフのネスト挙動を上書きできます(これは [`Handoff.nest_handoff_history`][agents.handoffs.Handoff.nest_handoff_history] を設定します)。生成要約のラッパーテキストだけを変更したい場合は、エージェント実行前に [`set_conversation_history_wrappers`][agents.handoffs.set_conversation_history_wrappers](必要に応じて [`reset_conversation_history_wrappers`][agents.handoffs.reset_conversation_history_wrappers])を呼び出してください。
|
||||
ネストされたハンドオフはオプトインのベータとして利用でき、安定化が進むまではデフォルトで無効です。 [`RunConfig.nest_handoff_history`][agents.run.RunConfig.nest_handoff_history] を有効にすると、ランナーは以前の会話記録を 1 つの assistant 要約メッセージにまとめ、それを `<CONVERSATION HISTORY>` ブロックで包みます。このブロックには、同じ実行中に複数のハンドオフが発生した場合に新しいターンが追加され続けます。 [`RunConfig.handoff_history_mapper`][agents.run.RunConfig.handoff_history_mapper] を通じて独自のマッピング関数を提供し、完全な `input_filter` を書くことなく、生成されたメッセージを置き換えることができます。このオプトインは、ハンドオフと実行のどちらも明示的な `input_filter` を指定していない場合にのみ適用されます。そのため、ペイロードをすでにカスタマイズしている既存のコード(このリポジトリ内のコード例を含む)は、変更なしで現在の動作を維持します。単一のハンドオフに対してネスト動作を上書きするには、 [`handoff(...)`][agents.handoffs.handoff] に `nest_handoff_history=True` または `False` を渡します。これにより [`Handoff.nest_handoff_history`][agents.handoffs.Handoff.nest_handoff_history] が設定されます。生成された要約のラッパーテキストだけを変更したい場合は、エージェントを実行する前に [`set_conversation_history_wrappers`][agents.handoffs.set_conversation_history_wrappers] を呼び出してください(必要に応じて [`reset_conversation_history_wrappers`][agents.handoffs.reset_conversation_history_wrappers] も呼び出せます)。
|
||||
|
||||
ハンドオフ側とアクティブな [`RunConfig.handoff_input_filter`][agents.run.RunConfig.handoff_input_filter] の両方でフィルターが定義されている場合、その特定ハンドオフではハンドオフ単位の [`input_filter`][agents.handoffs.Handoff.input_filter] が優先されます。
|
||||
ハンドオフとアクティブな [`RunConfig.handoff_input_filter`][agents.run.RunConfig.handoff_input_filter] の両方がフィルターを定義している場合、その特定のハンドオフではハンドオフごとの [`input_filter`][agents.handoffs.Handoff.input_filter] が優先されます。
|
||||
|
||||
!!! note
|
||||
|
||||
ハンドオフは単一の run 内に留まります。入力ガードレールは依然としてチェーン内の最初のエージェントにのみ適用され、出力ガードレールは最終出力を生成するエージェントにのみ適用されます。ワークフロー内の各カスタム function-tool 呼び出しごとにチェックが必要な場合は、ツールガードレールを使用してください。
|
||||
ハンドオフは単一の実行内にとどまります。入力ガードレールは引き続きチェーン内の最初のエージェントにのみ適用され、出力ガードレールは最終出力を生成するエージェントにのみ適用されます。ワークフロー内の各カスタム関数ツール呼び出しの周囲でチェックが必要な場合は、ツールガードレールを使用してください。
|
||||
|
||||
一般的なパターン(たとえば履歴からすべてのツール呼び出しを削除するなど)は、[`agents.extensions.handoff_filters`][] に実装されています。
|
||||
一般的なパターン(たとえば、履歴からすべてのツール呼び出しを削除するなど)がいくつかあり、 [`agents.extensions.handoff_filters`][] に実装されています。
|
||||
|
||||
```python
|
||||
from agents import Agent, handoff
|
||||
@@ -138,11 +138,11 @@ handoff_obj = handoff(
|
||||
)
|
||||
```
|
||||
|
||||
1. これにより、`FAQ agent` が呼び出されたときに履歴からすべてのツールが自動的に削除されます。
|
||||
1. これにより、 `FAQ agent` が呼び出されたときに、履歴からすべてのツールが自動的に削除されます。
|
||||
|
||||
## 推奨プロンプト
|
||||
|
||||
LLM がハンドオフを適切に理解できるように、エージェントにハンドオフ情報を含めることを推奨します。[`agents.extensions.handoff_prompt.RECOMMENDED_PROMPT_PREFIX`][] に推奨プレフィックスがあり、または [`agents.extensions.handoff_prompt.prompt_with_handoff_instructions`][] を呼び出して、推奨データをプロンプトに自動追加できます。
|
||||
LLM がハンドオフを適切に理解できるように、エージェントにハンドオフに関する情報を含めることを推奨します。推奨されるプレフィックスを [`agents.extensions.handoff_prompt.RECOMMENDED_PROMPT_PREFIX`][] に用意しています。または、 [`agents.extensions.handoff_prompt.prompt_with_handoff_instructions`][] を呼び出して、推奨データをプロンプトに自動的に追加できます。
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
|
||||
@@ -2,19 +2,19 @@
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# Human-in-the-loop
|
||||
# ヒューマンインザループ
|
||||
|
||||
human-in-the-loop ( HITL ) フローを使用すると、機密性の高いツール呼び出しを人が承認または拒否するまで、エージェント実行を一時停止できます。ツールは承認が必要なタイミングを宣言し、実行結果は保留中の承認を中断として表示し、`RunState` によって判断後に実行をシリアライズおよび再開できます。
|
||||
ヒューマンインザループ (HITL) フローを使用すると、人が慎重な扱いが必要なツール呼び出しを承認または拒否するまで、エージェントの実行を一時停止できます。ツールは承認が必要なタイミングを宣言し、実行結果は保留中の承認を中断として提示し、`RunState` によって判定後に実行をシリアライズして再開できます。
|
||||
|
||||
この承認サーフェスは実行全体に適用され、現在のトップレベルエージェントに限定されません。同じパターンは、ツールが現在のエージェントに属する場合、ハンドオフで到達したエージェントに属する場合、またはネストされた [`Agent.as_tool()`][agents.agent.Agent.as_tool] 実行に属する場合にも適用されます。ネストされた `Agent.as_tool()` の場合でも、中断は外側の実行に表示されるため、外側の `RunState` で承認または拒否し、元のトップレベル実行を再開します。
|
||||
その承認の提示先は実行全体であり、現在のトップレベルのエージェントに限定されません。同じパターンは、ツールが現在のエージェントに属する場合、ハンドオフを通じて到達したエージェントに属する場合、またはネストされた [`Agent.as_tool()`][agents.agent.Agent.as_tool] 実行に属する場合にも適用されます。ネストされた `Agent.as_tool()` の場合でも、中断は外側の実行に提示されるため、外側の `RunState` で承認または拒否し、元のトップレベルの実行を再開します。
|
||||
|
||||
`Agent.as_tool()` では、承認は 2 つの異なるレイヤーで発生する可能性があります。エージェントツール自体が `Agent.as_tool(..., needs_approval=...)` によって承認を要求でき、さらにネストされたエージェント内のツールがネスト実行開始後に独自の承認を発生させることもできます。どちらも同じ外側実行の中断フローで処理されます。
|
||||
`Agent.as_tool()` では、承認が 2 つの異なるレイヤーで発生する可能性があります。エージェントツール自体が `Agent.as_tool(..., needs_approval=...)` によって承認を要求でき、ネストされたエージェント内のツールも、ネストされた実行が開始した後に独自の承認を要求できます。どちらも同じ外側の実行の中断フローを通じて処理されます。
|
||||
|
||||
このページでは、`interruptions` を介した手動承認フローに焦点を当てます。アプリがコードで判断できる場合、一部のツールタイプはプログラムによる承認コールバックもサポートしており、実行を一時停止せずに継続できます。
|
||||
このページでは、`interruptions` を介した手動承認フローに焦点を当てます。アプリがコード内で判定できる場合、一部のツールタイプはプログラムによる承認コールバックにも対応しているため、実行を一時停止せずに続行できます。
|
||||
|
||||
## 承認が必要なツールのマーキング
|
||||
## 承認が必要なツールの指定
|
||||
|
||||
`needs_approval` を `True` に設定すると常に承認が必要になり、呼び出しごとに判断する非同期関数を渡すこともできます。呼び出し可能オブジェクトは、実行コンテキスト、解析済みツールパラメーター、ツール呼び出し ID を受け取ります。
|
||||
常に承認を要求するには `needs_approval` を `True` に設定するか、呼び出しごとに判定する async 関数を指定します。この呼び出し可能オブジェクトは、実行コンテキスト、解析済みのツールパラメーター、ツール呼び出し ID を受け取ります。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, function_tool
|
||||
@@ -41,28 +41,28 @@ agent = Agent(
|
||||
)
|
||||
```
|
||||
|
||||
`needs_approval` は [`function_tool`][agents.tool.function_tool]、[`Agent.as_tool`][agents.agent.Agent.as_tool]、[`ShellTool`][agents.tool.ShellTool]、[`ApplyPatchTool`][agents.tool.ApplyPatchTool] で利用できます。ローカル MCP サーバーも、[`MCPServerStdio`][agents.mcp.server.MCPServerStdio]、[`MCPServerSse`][agents.mcp.server.MCPServerSse]、[`MCPServerStreamableHttp`][agents.mcp.server.MCPServerStreamableHttp] の `require_approval` を通じて承認をサポートします。ホスト型 MCP サーバーは、[`HostedMCPTool`][agents.tool.HostedMCPTool] の `tool_config={"require_approval": "always"}` と、任意の `on_approval_request` コールバックを介して承認をサポートします。 shell および apply_patch ツールは、割り込みを表示せずに自動承認または自動拒否したい場合に `on_approval` コールバックを受け付けます。
|
||||
`needs_approval` は、[`function_tool`][agents.tool.function_tool]、[`Agent.as_tool`][agents.agent.Agent.as_tool]、[`ShellTool`][agents.tool.ShellTool]、[`ApplyPatchTool`][agents.tool.ApplyPatchTool] で利用できます。ローカル MCP サーバーも、[`MCPServerStdio`][agents.mcp.server.MCPServerStdio]、[`MCPServerSse`][agents.mcp.server.MCPServerSse]、[`MCPServerStreamableHttp`][agents.mcp.server.MCPServerStreamableHttp] の `require_approval` を通じて承認に対応しています。ホスト型 MCP サーバーは、`tool_config={"require_approval": "always"}` と任意の `on_approval_request` コールバックを設定した [`HostedMCPTool`][agents.tool.HostedMCPTool] によって承認に対応します。Shell と apply_patch ツールは、中断を提示せずに自動承認または自動拒否したい場合に `on_approval` コールバックを受け付けます。
|
||||
|
||||
## 承認フローの仕組み
|
||||
|
||||
1. モデルがツール呼び出しを出力すると、ランナーはその承認ルール (`needs_approval`、`require_approval`、またはホスト型 MCP の同等機能) を評価します。
|
||||
2. そのツール呼び出しに対する承認判断がすでに [`RunContextWrapper`][agents.run_context.RunContextWrapper] に保存されている場合、ランナーは確認なしで続行します。呼び出し単位の承認は特定の呼び出し ID にスコープされます。実行の残り期間における同ツールへの今後の呼び出しにも同じ判断を保持するには、`always_approve=True` または `always_reject=True` を渡します。
|
||||
3. それ以外の場合、実行は一時停止し、`RunResult.interruptions` (または `RunResultStreaming.interruptions`) に `agent.name`、`tool_name`、`arguments` などの詳細を含む [`ToolApprovalItem`][agents.items.ToolApprovalItem] エントリーが入ります。これには、ハンドオフ後またはネストされた `Agent.as_tool()` 実行内で発生した承認も含まれます。
|
||||
4. `result.to_state()` で結果を `RunState` に変換し、`state.approve(...)` または `state.reject(...)` を呼び出した後、`Runner.run(agent, state)` または `Runner.run_streamed(agent, state)` で再開します。ここで `agent` は、その実行の元のトップレベルエージェントです。
|
||||
5. 再開された実行は中断地点から継続し、新たな承認が必要であればこのフローに再度入ります。
|
||||
2. そのツール呼び出しの承認判定がすでに [`RunContextWrapper`][agents.run_context.RunContextWrapper] に保存されている場合、ランナーは確認を求めずに処理を続行します。呼び出しごとの承認は特定の呼び出し ID にスコープされます。そのツールに対する今後の呼び出しに、実行の残りの間同じ判定を保持するには、`always_approve=True` または `always_reject=True` を渡します。
|
||||
3. それ以外の場合、実行は一時停止し、`RunResult.interruptions` (または `RunResultStreaming.interruptions`) に、`agent.name`、`tool_name`、`arguments` などの詳細を含む [`ToolApprovalItem`][agents.items.ToolApprovalItem] エントリが入ります。これには、ハンドオフ後やネストされた `Agent.as_tool()` 実行内で発生した承認も含まれます。
|
||||
4. 実行結果を `result.to_state()` で `RunState` に変換し、`state.approve(...)` または `state.reject(...)` を呼び出してから、`Runner.run(agent, state)` または `Runner.run_streamed(agent, state)` で再開します。ここで `agent` は、その実行における元のトップレベルのエージェントです。
|
||||
5. 再開された実行は中断した場所から続行し、新しい承認が必要になった場合はこのフローに再び入ります。
|
||||
|
||||
`always_approve=True` または `always_reject=True` で作成された固定判断は実行状態に保存されるため、同じ一時停止済み実行を後で再開する際に `state.to_string()` / `RunState.from_string(...)` および `state.to_json()` / `RunState.from_json(...)` をまたいで保持されます。
|
||||
`always_approve=True` または `always_reject=True` で作成された固定判定は実行状態に保存されるため、後で同じ一時停止中の実行を再開するときに `state.to_string()` / `RunState.from_string(...)` および `state.to_json()` / `RunState.from_json(...)` を使っても保持されます。
|
||||
|
||||
同じパスで保留中の承認をすべて解決する必要はありません。`interruptions` には、通常の関数ツール、ホスト型 MCP 承認、ネストされた `Agent.as_tool()` 承認が混在する可能性があります。一部の項目のみ承認または拒否して再実行した場合、解決済みの呼び出しは継続し、未解決のものは `interruptions` に残って実行を再び一時停止します。
|
||||
すべての保留中承認を同じ 1 回の処理で解決する必要はありません。`interruptions` には、通常の関数ツール、ホスト型 MCP の承認、ネストされた `Agent.as_tool()` の承認が混在する場合があります。一部の項目だけを承認または拒否した後に再実行すると、解決済みの呼び出しは続行でき、未解決のものは `interruptions` に残って実行を再び一時停止します。
|
||||
|
||||
## 拒否メッセージのカスタマイズ
|
||||
## カスタム拒否メッセージ
|
||||
|
||||
デフォルトでは、拒否されたツール呼び出しは SDK の標準拒否テキストを実行に返します。このメッセージは 2 つのレイヤーでカスタマイズできます。
|
||||
既定では、拒否されたツール呼び出しは SDK 標準の拒否テキストを実行内に返します。このメッセージは 2 つのレイヤーでカスタマイズできます。
|
||||
|
||||
- 実行全体のフォールバック: [`RunConfig.tool_error_formatter`][agents.run.RunConfig.tool_error_formatter] を設定し、実行全体の承認拒否に対するモデル可視のデフォルトメッセージを制御します。
|
||||
- 呼び出し単位の上書き: 特定の拒否ツール呼び出しだけ別メッセージを表示したい場合、`state.reject(...)` に `rejection_message=...` を渡します。
|
||||
- 実行全体のフォールバック: 実行全体で承認拒否に対するモデルに見える既定メッセージを制御するには、[`RunConfig.tool_error_formatter`][agents.run.RunConfig.tool_error_formatter] を設定します。
|
||||
- 呼び出しごとのオーバーライド: 特定の拒否されたツール呼び出しだけに異なるメッセージを提示したい場合は、`state.reject(...)` に `rejection_message=...` を渡します。
|
||||
|
||||
両方が指定された場合、呼び出し単位の `rejection_message` が実行全体フォーマッターより優先されます。
|
||||
両方が指定されている場合、呼び出しごとの `rejection_message` が実行全体のフォーマッターより優先されます。
|
||||
|
||||
```python
|
||||
from agents import RunConfig, ToolErrorFormatterArgs
|
||||
@@ -83,27 +83,27 @@ state.reject(
|
||||
)
|
||||
```
|
||||
|
||||
両レイヤーを組み合わせて示す完全な例は [`examples/agent_patterns/human_in_the_loop_custom_rejection.py`](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns/human_in_the_loop_custom_rejection.py) を参照してください。
|
||||
両方のレイヤーをまとめて示す完全なコード例については、[`examples/agent_patterns/human_in_the_loop_custom_rejection.py`](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns/human_in_the_loop_custom_rejection.py) を参照してください。
|
||||
|
||||
## 自動承認判断
|
||||
## 自動承認判定
|
||||
|
||||
手動 `interruptions` は最も汎用的なパターンですが、唯一ではありません。
|
||||
手動の `interruptions` は最も汎用的なパターンですが、唯一の方法ではありません。
|
||||
|
||||
- ローカル [`ShellTool`][agents.tool.ShellTool] と [`ApplyPatchTool`][agents.tool.ApplyPatchTool] は `on_approval` を使用してコード内で即時に承認または拒否できます。
|
||||
- [`HostedMCPTool`][agents.tool.HostedMCPTool] は、同種のプログラムによる判断のために `tool_config={"require_approval": "always"}` と `on_approval_request` を併用できます。
|
||||
- 通常の [`function_tool`][agents.tool.function_tool] ツールと [`Agent.as_tool()`][agents.agent.Agent.as_tool] は、このページの手動中断フローを使用します。
|
||||
- ローカルの [`ShellTool`][agents.tool.ShellTool] と [`ApplyPatchTool`][agents.tool.ApplyPatchTool] は、`on_approval` を使用してコード内で即座に承認または拒否できます。
|
||||
- [`HostedMCPTool`][agents.tool.HostedMCPTool] は、`tool_config={"require_approval": "always"}` と `on_approval_request` を組み合わせて、同じ種類のプログラムによる判定を行えます。
|
||||
- 通常の [`function_tool`][agents.tool.function_tool] ツールと [`Agent.as_tool()`][agents.agent.Agent.as_tool] は、このページの手動中断フローを使用します。
|
||||
|
||||
これらのコールバックが判断を返すと、実行は人の応答を待って一時停止せずに継続します。 Realtime および音声セッション API については、[Realtime ガイド](realtime/guide.md) の承認フローを参照してください。
|
||||
これらのコールバックが判定を返すと、人間の応答を待って一時停止することなく実行が続行されます。Realtime および音声セッション API については、[Realtime ガイド](realtime/guide.md) の承認フローを参照してください。
|
||||
|
||||
## ストリーミングとセッション
|
||||
|
||||
同じ中断フローはストリーミング実行でも機能します。ストリーミング実行が一時停止したら、イテレーターが終了するまで [`RunResultStreaming.stream_events()`][agents.result.RunResultStreaming.stream_events] を消費し、[`RunResultStreaming.interruptions`][agents.result.RunResultStreaming.interruptions] を確認して解決し、再開後の出力もストリーミングを継続したい場合は [`Runner.run_streamed(...)`][agents.run.Runner.run_streamed] で再開します。このパターンのストリーミング版は [ストリーミング](streaming.md) を参照してください。
|
||||
同じ中断フローはストリーミング実行でも機能します。ストリーミング実行が一時停止した後は、イテレーターが終了するまで [`RunResultStreaming.stream_events()`][agents.result.RunResultStreaming.stream_events] を消費し続け、[`RunResultStreaming.interruptions`][agents.result.RunResultStreaming.interruptions] を確認して解決し、再開後の出力もストリーミングし続けたい場合は [`Runner.run_streamed(...)`][agents.run.Runner.run_streamed] で再開します。このパターンのストリーミング版については、[ストリーミング](streaming.md) を参照してください。
|
||||
|
||||
セッションも使用している場合は、`RunState` から再開する際に同じセッションインスタンスを渡し続けるか、同じバックエンドストアを指す別のセッションオブジェクトを渡してください。再開されたターンは同じ保存済み会話履歴に追加されます。セッションライフサイクルの詳細は [セッション](sessions/index.md) を参照してください。
|
||||
セッションも使用している場合は、`RunState` から再開するときに同じセッションインスタンスを渡し続けるか、同じバッキングストアを指す別のセッションオブジェクトを渡します。これにより、再開されたターンは同じ保存済み会話履歴に追加されます。セッションのライフサイクル詳細については、[セッション](sessions/index.md) を参照してください。
|
||||
|
||||
## 例: 一時停止、承認、再開
|
||||
## 例: 一時停止・承認・再開
|
||||
|
||||
以下のスニペットは JavaScript の HITL ガイドを踏襲しています。ツールに承認が必要なときに一時停止し、状態をディスクに保存し、再読み込みして、判断を収集した後に再開します。
|
||||
以下のスニペットは JavaScript の HITL ガイドと同じ流れです。ツールに承認が必要な場合に一時停止し、状態をディスクに永続化して再読み込みし、判定を収集した後に再開します。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -167,43 +167,35 @@ if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
この例では、`prompt_approval` は `input()` を使用し `run_in_executor(...)` で実行されるため同期的です。承認ソースがすでに非同期 ( 例: HTTP リクエストや非同期データベースクエリ) の場合は、`async def` 関数を使用して直接 `await` できます。
|
||||
この例では、`prompt_approval` は `input()` を使用し、`run_in_executor(...)` で実行されるため同期的です。承認の取得元がすでに非同期である場合 (たとえば、HTTP リクエストや非同期データベースクエリ)、代わりに `async def` 関数を使用して直接 `await` できます。
|
||||
|
||||
承認待ち中にも出力をストリーミングしたい場合は、`Runner.run_streamed` を呼び出し、完了まで `result.stream_events()` を消費し、その後は上記と同じ `result.to_state()` と再開手順に従ってください。
|
||||
承認を待つ間に出力をストリーミングするには、`Runner.run_streamed` を呼び出し、完了するまで `result.stream_events()` を消費してから、上記と同じ `result.to_state()` と再開手順に従います。
|
||||
|
||||
## リポジトリのパターンと例
|
||||
## リポジトリのパターンとコード例
|
||||
|
||||
- **ストリーミング承認**: `examples/agent_patterns/human_in_the_loop_stream.py` は、`stream_events()` を最後まで処理し、保留中ツール呼び出しを承認してから `Runner.run_streamed(agent, state)` で再開する方法を示します。
|
||||
- **カスタム拒否テキスト**: `examples/agent_patterns/human_in_the_loop_custom_rejection.py` は、承認が拒否されたときに実行レベルの `tool_error_formatter` と呼び出し単位の `rejection_message` 上書きを組み合わせる方法を示します。
|
||||
- **Agent as tool 承認**: `Agent.as_tool(..., needs_approval=...)` は、委譲されたエージェントタスクにレビューが必要な場合にも同じ中断フローを適用します。ネストされた中断も外側の実行に表示されるため、ネスト側ではなく元のトップレベルエージェントを再開してください。
|
||||
- **ローカル shell / apply_patch ツール**: `ShellTool` と `ApplyPatchTool` も `needs_approval` をサポートします。将来の呼び出しのために判断をキャッシュするには `state.approve(interruption, always_approve=True)` または `state.reject(..., always_reject=True)` を使用します。自動判断には `on_approval` を指定します ( `examples/tools/shell.py` を参照)。手動判断には中断を処理します ( `examples/tools/shell_human_in_the_loop.py` を参照)。ホスト型 shell 環境は `needs_approval` または `on_approval` をサポートしません。[ツールガイド](tools.md) を参照してください。
|
||||
- **ローカル MCP サーバー**: `MCPServerStdio` / `MCPServerSse` / `MCPServerStreamableHttp` で `require_approval` を使用し、MCP ツール呼び出しを制御します ( `examples/mcp/get_all_mcp_tools_example/main.py` および `examples/mcp/tool_filter_example/main.py` を参照)。
|
||||
- **ホスト型 MCP サーバー**: HITL を強制するには `HostedMCPTool` で `require_approval` を `"always"` に設定し、必要に応じて `on_approval_request` を指定して自動承認または拒否します ( `examples/hosted_mcp/human_in_the_loop.py` および `examples/hosted_mcp/on_approval.py` を参照)。信頼済みサーバーには `"never"` を使用します (`examples/hosted_mcp/simple.py`)。
|
||||
- **セッションとメモリ**: 複数ターンにわたり承認と会話履歴を保持するには `Runner.run` にセッションを渡します。 SQLite および OpenAI Conversations セッションのバリアントは `examples/memory/memory_session_hitl_example.py` と `examples/memory/openai_session_hitl_example.py` にあります。
|
||||
- **Realtime エージェント**: realtime デモは `RealtimeSession` の `approve_tool_call` / `reject_tool_call` を介してツール呼び出しを承認または拒否する WebSocket メッセージを公開します ( サーバー側ハンドラーは `examples/realtime/app/server.py`、API サーフェスは [Realtime ガイド](realtime/guide.md#tool-approvals) を参照)。
|
||||
- **ストリーミング承認**: `examples/agent_patterns/human_in_the_loop_stream.py` は、`stream_events()` を最後まで読み出し、その後 `Runner.run_streamed(agent, state)` で再開する前に保留中のツール呼び出しを承認する方法を示します。
|
||||
- **カスタム拒否テキスト**: `examples/agent_patterns/human_in_the_loop_custom_rejection.py` は、承認が拒否された場合に、実行レベルの `tool_error_formatter` と呼び出しごとの `rejection_message` オーバーライドを組み合わせる方法を示します。
|
||||
- **ツールとしてのエージェントの承認**: `Agent.as_tool(..., needs_approval=...)` は、委譲されたエージェントタスクにレビューが必要な場合に同じ中断フローを適用します。ネストされた中断も外側の実行に提示されるため、ネストされたエージェントではなく元のトップレベルのエージェントを再開してください。
|
||||
- **ローカル shell と apply_patch ツール**: `ShellTool` と `ApplyPatchTool` も `needs_approval` に対応しています。将来の呼び出しに備えて判定をキャッシュするには、`state.approve(interruption, always_approve=True)` または `state.reject(..., always_reject=True)` を使用します。自動判定には `on_approval` を指定します (`examples/tools/shell.py` を参照)。手動判定には中断を処理します (`examples/tools/shell_human_in_the_loop.py` を参照)。ホスト型 shell 環境は `needs_approval` または `on_approval` に対応していません。[ツールガイド](tools.md) を参照してください。
|
||||
- **ローカル MCP サーバー**: MCP ツール呼び出しを制御するには、`MCPServerStdio` / `MCPServerSse` / `MCPServerStreamableHttp` の `require_approval` を使用します (`examples/mcp/get_all_mcp_tools_example/main.py` と `examples/mcp/tool_filter_example/main.py` を参照)。
|
||||
- **ホスト型 MCP サーバー**: HITL を強制するには、`HostedMCPTool` で `require_approval` を `"always"` に設定し、必要に応じて自動承認または拒否のために `on_approval_request` を指定します (`examples/hosted_mcp/human_in_the_loop.py` と `examples/hosted_mcp/on_approval.py` を参照)。信頼済みサーバーには `"never"` を使用します (`examples/hosted_mcp/simple.py`)。
|
||||
- **セッションとメモリ**: セッションを `Runner.run` に渡すと、承認と会話履歴が複数ターンにわたって保持されます。SQLite と OpenAI Conversations のセッション版は、`examples/memory/memory_session_hitl_example.py` と `examples/memory/openai_session_hitl_example.py` にあります。
|
||||
- **Realtime エージェント**: Realtime デモでは、`RealtimeSession` の `approve_tool_call` / `reject_tool_call` を介してツール呼び出しを承認または拒否する WebSocket メッセージを公開しています (サーバー側ハンドラーについては `examples/realtime/app/server.py`、API サーフェスについては [Realtime ガイド](realtime/guide.md#tool-approvals) を参照)。
|
||||
|
||||
## 長時間実行承認
|
||||
## 長時間にわたる承認
|
||||
|
||||
`RunState` は永続性を考慮して設計されています。保留中作業をデータベースやキューに保存するには `state.to_json()` または `state.to_string()` を使用し、後で `RunState.from_json(...)` または `RunState.from_string(...)` で再作成します。
|
||||
`RunState` は耐久性を持つように設計されています。`state.to_json()` または `state.to_string()` を使用して保留中の作業をデータベースまたはキューに保存し、後で `RunState.from_json(...)` または `RunState.from_string(...)` で再作成します。
|
||||
|
||||
有用なシリアライズオプション:
|
||||
便利なシリアライズオプション:
|
||||
|
||||
- `context_serializer`: マッピング以外のコンテキストオブジェクトをどのようにシリアライズするかをカスタマイズします。
|
||||
- `context_deserializer`: `RunState.from_json(...)` または `RunState.from_string(...)` で状態をロードするときに、マッピング以外のコンテキストオブジェクトを再構築します。
|
||||
- `strict_context=True`: コンテキストがすでに
|
||||
マッピングであるか、適切な serializer / deserializer を提供しない限り、シリアライズまたはデシリアライズを失敗させます。
|
||||
- `context_override`: 状態ロード時にシリアライズ済みコンテキストを置き換えます。これは
|
||||
元のコンテキストオブジェクトを復元したくない場合に有用ですが、すでに
|
||||
シリアライズ済みペイロードからそのコンテキストを削除するものではありません。
|
||||
- `include_tracing_api_key=True`: 再開作業でも同じ認証情報でトレースをエクスポートし続ける必要がある場合に、
|
||||
シリアライズされたトレースペイロードに tracing API キーを含めます。
|
||||
- `context_serializer`: 非マッピングのコンテキストオブジェクトのシリアライズ方法をカスタマイズします。
|
||||
- `context_deserializer`: `RunState.from_json(...)` または `RunState.from_string(...)` で状態を読み込むときに、非マッピングのコンテキストオブジェクトを再構築します。
|
||||
- `strict_context=True`: コンテキストがすでにマッピングであるか、適切なシリアライザー / デシリアライザーを指定している場合を除き、シリアライズまたはデシリアライズを失敗させます。
|
||||
- `context_override`: 状態を読み込むときに、シリアライズされたコンテキストを置き換えます。これは、元のコンテキストオブジェクトを復元したくない場合に便利ですが、すでにシリアライズ済みのペイロードからそのコンテキストを削除するわけではありません。
|
||||
- `include_tracing_api_key=True`: 再開された作業で同じ認証情報を使ってトレースのエクスポートを継続する必要がある場合、シリアライズされたトレースペイロードにトレーシング API キーを含めます。
|
||||
|
||||
シリアライズされた実行状態には、アプリコンテキストに加えて、承認、
|
||||
使用量、シリアライズされた `tool_input`、ネストされた agent-as-tool 再開、トレースメタデータ、サーバー管理の
|
||||
会話設定など、SDK 管理の実行時メタデータが含まれます。シリアライズ状態を保存または転送する予定がある場合は、
|
||||
`RunContextWrapper.context` を永続化データとして扱い、意図的に
|
||||
状態と一緒に移動させたい場合を除き、そこに秘密情報を置かないでください。
|
||||
シリアライズされた実行状態には、アプリのコンテキストに加えて、承認、使用量、シリアライズ済みの `tool_input`、ネストされた agent-as-tool の再開情報、トレースメタデータ、サーバー管理の会話設定など、SDK 管理のランタイムメタデータが含まれます。シリアライズされた状態を保存または送信する予定がある場合は、`RunContextWrapper.context` を永続化データとして扱い、状態と一緒に移動させる意図がある場合を除き、そこにシークレットを置かないでください。
|
||||
|
||||
## 保留タスクのバージョニング
|
||||
## 保留中タスクのバージョニング
|
||||
|
||||
承認がしばらく保留される可能性がある場合は、シリアライズ状態と一緒にエージェント定義または SDK のバージョンマーカーを保存してください。これにより、デシリアライズを対応するコードパスに振り分け、モデル、プロンプト、またはツール定義が変更された際の非互換性を回避できます。
|
||||
承認がしばらく保留される可能性がある場合は、エージェント定義または SDK のバージョンマーカーを、シリアライズされた状態と一緒に保存してください。これにより、モデル、プロンプト、ツール定義が変更された場合の非互換性を避けるために、対応するコードパスへデシリアライズ処理を振り分けられます。
|
||||
+47
-47
@@ -4,51 +4,51 @@ search:
|
||||
---
|
||||
# OpenAI Agents SDK
|
||||
|
||||
[OpenAI Agents SDK](https://github.com/openai/openai-agents-python) を使うと、ごく少数の抽象化だけを備えた軽量で使いやすいパッケージで、エージェント型 AI アプリを構築できます。これは、以前のエージェント向け実験プロジェクトである [Swarm](https://github.com/openai/swarm/tree/main) を本番対応に進化させたものです。Agents SDK には、ごく少数の基本コンポーネントがあります。
|
||||
[OpenAI Agents SDK](https://github.com/openai/openai-agents-python) は、抽象化をほとんど持たない軽量で使いやすいパッケージで、エージェント型 AI アプリを構築できるようにします。これは、以前のエージェント向け実験プロジェクトである [Swarm](https://github.com/openai/swarm/tree/main) を本番環境対応に発展させたものです。Agents SDK は、非常に少数の基本コンポーネントで構成されています:
|
||||
|
||||
- **エージェント**。instructions と tools を備えた LLM です
|
||||
- **Agents as tools / ハンドオフ**。特定のタスクについて、エージェントがほかのエージェントに委任できるようにします
|
||||
- **ガードレール**。エージェントの入力と出力の検証を可能にします
|
||||
- **エージェント**: 指示とツールを備えた LLM です
|
||||
- **Agents as tools / ハンドオフ**: エージェントが特定のタスクを他のエージェントに委任できるようにします
|
||||
- **ガードレール**: エージェントの入力と出力の検証を可能にします
|
||||
|
||||
これらの基本コンポーネントは Python と組み合わせることで、ツールとエージェントの複雑な関係を表現するのに十分な力を発揮し、学習コストを大きくかけることなく実運用のアプリケーションを構築できます。さらに、この SDK には組み込みの **トレーシング** があり、エージェントフローの可視化やデバッグに加えて、評価や、アプリケーション向けのモデルのファインチューニングまで行えます。
|
||||
Python と組み合わせることで、これらの基本コンポーネントは、ツールとエージェント間の複雑な関係を表現するのに十分強力であり、習得のハードルを高くすることなく実世界のアプリケーションを構築できます。さらに、SDK には組み込みの **トレーシング** が含まれており、エージェント型フローの可視化とデバッグ、評価、さらにはアプリケーション向けのモデルのファインチューニングも可能です。
|
||||
|
||||
## Agents SDK を使う理由
|
||||
## Agents SDK の利用理由
|
||||
|
||||
この SDK には、設計上の主要な原則が 2 つあります。
|
||||
SDK の設計を支える原則は 2 つあります:
|
||||
|
||||
1. 使う価値があるだけの十分な機能を備えつつ、素早く学べるよう基本コンポーネントは少数にとどめること。
|
||||
2. そのままですぐに使えて、しかも何が起きるかを正確にカスタマイズできること。
|
||||
1. 使用する価値がある十分な機能を備えつつ、すばやく学べるだけの少数の基本コンポーネントに抑えること。
|
||||
2. そのままでも優れた動作をしつつ、何が起こるかを正確にカスタマイズできること。
|
||||
|
||||
以下は、この SDK の主な機能です。
|
||||
SDK の主な機能は次のとおりです:
|
||||
|
||||
- **エージェントループ**: ツール呼び出しを処理し、結果を LLM に返し、タスクが完了するまで継続する組み込みのエージェントループです。
|
||||
- **Python ファースト**: 新しい抽象化を学ぶ必要はなく、組み込みの言語機能を使ってエージェントオーケストレーションや連携を行えます。
|
||||
- **Agents as tools / ハンドオフ**: 複数のエージェント間で作業を調整および委任するための強力な仕組みです。
|
||||
- **Sandbox エージェント**: manifest で定義されたファイル、sandbox client の選択、再開可能な sandbox session を備えた、実際に分離されたワークスペース内で専門エージェントを実行します。
|
||||
- **ガードレール**: エージェントの実行と並行して入力検証と安全性チェックを実行し、チェックに通らなかった場合は即座に失敗させます。
|
||||
- **関数ツール**: 自動スキーマ生成と Pydantic ベースの検証により、任意の Python 関数をツールに変換します。
|
||||
- **MCP サーバーツール呼び出し**: 関数ツールと同じ方法で動作する、組み込みの MCP サーバーツール統合です。
|
||||
- **エージェントループ**: ツール呼び出しを処理し、結果を LLM に送り返し、タスクが完了するまで継続する組み込みのエージェントループです。
|
||||
- **Python ファースト**: 新しい抽象化を学ぶ必要なく、組み込みの言語機能を使ってエージェントをオーケストレーションし、連鎖させます。
|
||||
- **Agents as tools / ハンドオフ**: 複数のエージェント間で作業を調整し、委任するための強力な仕組みです。
|
||||
- **Sandbox エージェント**: マニフェストで定義されたファイル、Sandbox クライアントの選択、再開可能なサンドボックスセッションを備えた、実際の隔離ワークスペース内で専門エージェントを実行します。
|
||||
- **ガードレール**: エージェント実行と並行して入力検証と安全性チェックを実行し、チェックに通らない場合は即座に失敗として終了します。
|
||||
- **関数ツール**: スキーマの自動生成と Pydantic によるバリデーションにより、任意の Python 関数をツールに変換します。
|
||||
- **MCP サーバーのツール呼び出し**: 関数ツールと同じように動作する、組み込みの MCP サーバーツール統合です。
|
||||
- **セッション**: エージェントループ内で作業コンテキストを維持するための永続的なメモリレイヤーです。
|
||||
- **Human in the loop**: エージェント実行全体で人間を関与させるための組み込みの仕組みです。
|
||||
- **トレーシング**: ワークフローの可視化、デバッグ、監視のための組み込みトレーシングで、OpenAI の評価、ファインチューニング、蒸留ツール群をサポートします。
|
||||
- **Realtime Agents**: `gpt-realtime-1.5`、自動割り込み検出、コンテキスト管理、ガードレールなどを使用して、強力な音声エージェントを構築できます。
|
||||
- **ヒューマンインザループ**: エージェント実行の各所に人間を関与させるための組み込みの仕組みです。
|
||||
- **トレーシング**: ワークフローを可視化、デバッグ、監視するための組み込みのトレーシングで、OpenAI の評価、ファインチューニング、蒸留ツール群をサポートします。
|
||||
- **Realtime エージェント**: `gpt-realtime-2` を使い、自動割り込み検出、コンテキスト管理、ガードレールなどを備えた強力な音声エージェントを構築します。
|
||||
|
||||
## Agents SDK と Responses API の比較
|
||||
## Agents SDK と Responses API の選択
|
||||
|
||||
この SDK は、OpenAI モデルに対してはデフォルトで Responses API を使用しますが、モデル呼び出しの上により高水準のランタイムを追加します。
|
||||
SDK は OpenAI モデルに対してデフォルトで Responses API を使用しますが、モデル呼び出しの周りに高レベルのランタイムを追加します。
|
||||
|
||||
次のような場合は、Responses API を直接使用してください。
|
||||
次の場合は Responses API を直接使用します:
|
||||
|
||||
- ループ、ツールのディスパッチ、状態管理を自分で扱いたい
|
||||
- ワークフローが短命で、主にモデルの応答を返すことが目的である
|
||||
- ループ、ツールのディスパッチ、状態処理を自分で管理したい場合
|
||||
- ワークフローが短期間で、主にモデルの応答を返すことが目的の場合
|
||||
|
||||
次のような場合は、Agents SDK を使用してください。
|
||||
次の場合は Agents SDK を使用します:
|
||||
|
||||
- ランタイムにターン管理、ツール実行、ガードレール、ハンドオフ、またはセッションを管理させたい
|
||||
- エージェントに成果物を生成させたい、または複数の協調したステップにまたがって動作させたい
|
||||
- [Sandbox エージェント](sandbox_agents.md) を通じて、実際のワークスペースや再開可能な実行が必要である
|
||||
- ランタイムにターン、ツール実行、ガードレール、ハンドオフ、またはセッションを管理させたい場合
|
||||
- エージェントが成果物を生成する、または複数の協調したステップにわたって動作する必要がある場合
|
||||
- 実際のワークスペース、または [Sandbox エージェント](sandbox_agents.md) による再開可能な実行が必要な場合
|
||||
|
||||
どちらか一方を全体で選ぶ必要はありません。多くのアプリケーションでは、管理されたワークフローには SDK を使い、より低水準の経路には Responses API を直接呼び出しています。
|
||||
アプリケーション全体でどちらか一方を選ぶ必要はありません。多くのアプリケーションでは、管理されたワークフローには SDK を使用し、低レベルの処理経路では Responses API を直接呼び出します。
|
||||
|
||||
## インストール
|
||||
|
||||
@@ -56,7 +56,7 @@ search:
|
||||
pip install openai-agents
|
||||
```
|
||||
|
||||
## Hello World の例
|
||||
## Hello world の例
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
@@ -71,7 +71,7 @@ print(result.final_output)
|
||||
# Infinite loop's dance.
|
||||
```
|
||||
|
||||
(_これを実行する場合は、`OPENAI_API_KEY` 環境変数を設定していることを確認してください_)
|
||||
(_これを実行する場合は、 `OPENAI_API_KEY` 環境変数を設定していることを確認してください_)
|
||||
|
||||
```bash
|
||||
export OPENAI_API_KEY=sk-...
|
||||
@@ -79,23 +79,23 @@ export OPENAI_API_KEY=sk-...
|
||||
|
||||
## 開始ポイント
|
||||
|
||||
- [Quickstart](quickstart.md) で最初のテキストベースのエージェントを構築します。
|
||||
- 次に、[Running agents](running_agents.md#choose-a-memory-strategy) でターン間の状態の持ち方を決めます。
|
||||
- タスクが実際のファイル、リポジトリ、またはエージェントごとに分離されたワークスペース状態に依存する場合は、[Sandbox agents quickstart](sandbox_agents.md) を参照してください。
|
||||
- ハンドオフと manager 型のオーケストレーションのどちらにするかを決める場合は、[Agent orchestration](multi_agent.md) を参照してください。
|
||||
- [クイックスタート](quickstart.md) で、最初のテキストベースのエージェントを構築します。
|
||||
- 次に、[エージェントの実行](running_agents.md#choose-a-memory-strategy) で、ターン間で状態をどのように引き継ぐかを決定します。
|
||||
- タスクが実際のファイル、リポジトリ、またはエージェントごとに隔離されたワークスペース状態に依存する場合は、[Sandbox エージェントのクイックスタート](sandbox_agents.md) を参照してください。
|
||||
- ハンドオフとマネージャースタイルのオーケストレーションのどちらにするかを決める場合は、[エージェントオーケストレーション](multi_agent.md) を参照してください。
|
||||
|
||||
## パスの選択
|
||||
|
||||
やりたいことは分かっているが、それを説明しているページが分からない場合は、この表を使ってください。
|
||||
実行したい作業は分かっているものの、どのページで説明されているか分からない場合は、この表を使用してください。
|
||||
|
||||
| 目標 | 開始ポイント |
|
||||
| 目的 | 参照先 |
|
||||
| --- | --- |
|
||||
| 最初のテキストエージェントを構築し、完全な 1 回の実行を見る | [Quickstart](quickstart.md) |
|
||||
| 関数ツール、ホストされたツール、または Agents as tools を追加する | [Tools](tools.md) |
|
||||
| 実際に分離されたワークスペース内で、コーディング、レビュー、またはドキュメント用エージェントを実行する | [Sandbox agents quickstart](sandbox_agents.md) と [Sandbox clients](sandbox/clients.md) |
|
||||
| ハンドオフと manager 型のエージェントオーケストレーションのどちらにするかを決める | [Agent orchestration](multi_agent.md) |
|
||||
| ターンをまたいでメモリを維持する | [Running agents](running_agents.md#choose-a-memory-strategy) と [Sessions](sessions/index.md) |
|
||||
| OpenAI モデル、websocket トランスポート、または OpenAI 以外のプロバイダーを使う | [Models](models/index.md) |
|
||||
| 出力、実行項目、割り込み、再開状態を確認する | [Results](results.md) |
|
||||
| `gpt-realtime-1.5` を使った低レイテンシの音声エージェントを構築する | [Realtime agents quickstart](realtime/quickstart.md) と [Realtime transport](realtime/transport.md) |
|
||||
| speech-to-text / agent / text-to-speech パイプラインを構築する | [Voice pipeline quickstart](voice/quickstart.md) |
|
||||
| 最初のテキストエージェントを構築し、完全な 1 回の実行を確認する | [クイックスタート](quickstart.md) |
|
||||
| 関数ツール、OpenAI がホストするツール、または agents as tools を追加する | [ツール](tools.md) |
|
||||
| 実際の隔離ワークスペース内で、コーディング、レビュー、またはドキュメント処理のエージェントを実行する | [Sandbox エージェントのクイックスタート](sandbox_agents.md) and [Sandbox クライアント](sandbox/clients.md) |
|
||||
| ハンドオフとマネージャースタイルのオーケストレーションのどちらを使うか決める | [エージェントオーケストレーション](multi_agent.md) |
|
||||
| ターン間でメモリを保持する | [エージェントの実行](running_agents.md#choose-a-memory-strategy) and [セッション](sessions/index.md) |
|
||||
| OpenAI モデル、WebSocket トランスポート、または OpenAI 以外のプロバイダーを使用する | [モデル](models/index.md) |
|
||||
| 出力、実行アイテム、割り込み、再開状態を確認する | [実行結果](results.md) |
|
||||
| `gpt-realtime-2` を使って低レイテンシの音声エージェントを構築する | [Realtime エージェントのクイックスタート](realtime/quickstart.md) and [Realtime トランスポート](realtime/transport.md) |
|
||||
| 音声認識 / エージェント / 音声合成のパイプラインを構築する | [音声パイプラインのクイックスタート](voice/quickstart.md) |
|
||||
+105
-94
@@ -4,28 +4,31 @@ search:
|
||||
---
|
||||
# Model context protocol (MCP)
|
||||
|
||||
[Model context protocol](https://modelcontextprotocol.io/introduction) (MCP) は、アプリケーションが言語モデルにツールやコンテキストを公開する方法を標準化します。公式ドキュメントより:
|
||||
[Model context protocol](https://modelcontextprotocol.io/introduction) (MCP) は、アプリケーションがツールや
|
||||
コンテキストを言語モデルに公開する方法を標準化します。公式ドキュメントより:
|
||||
|
||||
> MCP は、アプリケーションが LLM にコンテキストを提供する方法を標準化するオープンプロトコルです。MCP は AI アプリケーション向けの USB-C ポートのようなものだと考えてください。USB-C がデバイスをさまざまな周辺機器やアクセサリーに接続するための標準化された方法を提供するのと同様に、MCP は AI モデルを異なるデータソースやツールに接続するための標準化された方法を提供します。
|
||||
> MCP は、アプリケーションが LLM にコンテキストを提供する方法を標準化するオープンプロトコルです。MCP は、AI
|
||||
> アプリケーションにおける USB-C ポートのようなものだと考えてください。USB-C がデバイスをさまざまな周辺機器やアクセサリに接続する標準化された方法を提供するのと同様に、MCP
|
||||
> は AI モデルをさまざまなデータソースやツールに接続する標準化された方法を提供します。
|
||||
|
||||
Agents Python SDK は複数の MCP トランスポートを理解します。これにより、既存の MCP サーバーを再利用したり、独自に構築してファイルシステム、 HTTP 、またはコネクタをバックエンドとするツールをエージェントに公開したりできます。
|
||||
Agents Python SDK は複数の MCP トランスポートに対応しています。これにより、既存の MCP サーバーを再利用したり、独自に構築して、ファイルシステム、HTTP、またはコネクターをバックエンドとするツールをエージェントに公開できます。
|
||||
|
||||
## MCP 統合の選択
|
||||
|
||||
MCP サーバーをエージェントに接続する前に、ツール呼び出しをどこで実行するか、到達可能なトランスポートはどれかを決めてください。以下のマトリクスは、 Python SDK がサポートする選択肢を要約したものです。
|
||||
MCP サーバーをエージェントに組み込む前に、ツール呼び出しをどこで実行すべきか、どのトランスポートに到達できるかを決めてください。次の表は、Python SDK がサポートする選択肢の概要です。
|
||||
|
||||
| 必要なもの | 推奨オプション |
|
||||
| 必要なこと | 推奨オプション |
|
||||
| ------------------------------------------------------------------------------------ | ----------------------------------------------------- |
|
||||
| モデルの代わりに OpenAI の Responses API から公開到達可能な MCP サーバーを呼び出す | [`HostedMCPTool`][agents.tool.HostedMCPTool] による **Hosted MCP server tools** |
|
||||
| ローカルまたはリモートで実行している Streamable HTTP サーバーに接続する | [`MCPServerStreamableHttp`][agents.mcp.server.MCPServerStreamableHttp] による **Streamable HTTP MCP servers** |
|
||||
| Server-Sent Events を使う HTTP を実装したサーバーと通信する | [`MCPServerSse`][agents.mcp.server.MCPServerSse] による **HTTP with SSE MCP servers** |
|
||||
| ローカルプロセスを起動し stdin/stdout 経由で通信する | [`MCPServerStdio`][agents.mcp.server.MCPServerStdio] による **stdio MCP servers** |
|
||||
| OpenAI の Responses API に、モデルに代わって公開到達可能な MCP サーバーを呼び出させる| **ホスト型 MCP サーバーツール** ([`HostedMCPTool`][agents.tool.HostedMCPTool] 経由) |
|
||||
| ローカルまたはリモートで実行している Streamable HTTP サーバーに接続する | **Streamable HTTP MCP サーバー** ([`MCPServerStreamableHttp`][agents.mcp.server.MCPServerStreamableHttp] 経由) |
|
||||
| Server-Sent Events を用いた HTTP を実装しているサーバーと通信する | **SSE を用いた HTTP MCP サーバー** ([`MCPServerSse`][agents.mcp.server.MCPServerSse] 経由) |
|
||||
| ローカルプロセスを起動し、stdin/stdout 経由で通信する | **stdio MCP サーバー** ([`MCPServerStdio`][agents.mcp.server.MCPServerStdio] 経由) |
|
||||
|
||||
以下のセクションでは、各オプション、設定方法、どのトランスポートを優先すべきかを説明します。
|
||||
以降のセクションでは、各オプション、その設定方法、あるトランスポートを別のトランスポートより優先すべきタイミングについて説明します。
|
||||
|
||||
## エージェントレベルの MCP 設定
|
||||
|
||||
トランスポートの選択に加えて、 `Agent.mcp_config` を設定して MCP ツールの準備方法を調整できます。
|
||||
トランスポートの選択に加えて、`Agent.mcp_config` を設定することで MCP ツールの準備方法を調整できます。
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
@@ -39,35 +42,38 @@ agent = Agent(
|
||||
# If None, MCP tool failures are raised as exceptions instead of
|
||||
# returning model-visible error text.
|
||||
"failure_error_function": None,
|
||||
# Prefix local MCP tool names with their server name.
|
||||
"include_server_in_tool_names": True,
|
||||
},
|
||||
)
|
||||
```
|
||||
|
||||
注記:
|
||||
注:
|
||||
|
||||
- `convert_schemas_to_strict` はベストエフォートです。スキーマを変換できない場合は元のスキーマが使われます。
|
||||
- `failure_error_function` は MCP ツール呼び出し失敗をモデルへどのように提示するかを制御します。
|
||||
- `failure_error_function` が未設定の場合、 SDK はデフォルトのツールエラーフォーマッターを使います。
|
||||
- サーバーレベルの `failure_error_function` は、そのサーバーに対して `Agent.mcp_config["failure_error_function"]` を上書きします。
|
||||
- `convert_schemas_to_strict` はベストエフォートです。スキーマを変換できない場合は、元のスキーマが使用されます。
|
||||
- `failure_error_function` は、MCP ツール呼び出しの失敗をモデルにどのように提示するかを制御します。
|
||||
- `failure_error_function` が未設定の場合、SDK はデフォルトのツールエラーフォーマッターを使用します。
|
||||
- サーバーレベルの `failure_error_function` は、そのサーバーについて `Agent.mcp_config["failure_error_function"]` を上書きします。
|
||||
- `include_server_in_tool_names` はオプトインです。有効にすると、各ローカル MCP ツールは、決定論的なサーバー接頭辞付きの名前でモデルに公開されます。これにより、複数の MCP サーバーが同じ名前のツールを公開する場合の衝突を避けやすくなります。生成される名前は ASCII セーフで、関数ツール名の長さ制限内に収まり、同じエージェント上の既存のローカル関数ツール名および有効化されたハンドオフ名を避けます。それでも SDK は元のサーバー上で元の MCP ツール名を呼び出します。
|
||||
|
||||
## トランスポート間の共通パターン
|
||||
## トランスポート共通のパターン
|
||||
|
||||
トランスポートを選んだ後、ほとんどの統合で同じ追加判断が必要です:
|
||||
トランスポートを選択した後、多くの統合では同じ追加判断が必要になります。
|
||||
|
||||
- ツールの一部だけを公開する方法 ([Tool filtering](#tool-filtering))。
|
||||
- サーバーが再利用可能なプロンプトも提供するかどうか ([Prompts](#prompts))。
|
||||
- `list_tools()` をキャッシュすべきかどうか ([Caching](#caching))。
|
||||
- MCP アクティビティがトレースにどう表示されるか ([Tracing](#tracing))。
|
||||
- ツールのサブセットのみを公開する方法([ツールフィルタリング](#tool-filtering))。
|
||||
- サーバーが再利用可能なプロンプトも提供するかどうか([プロンプト](#prompts))。
|
||||
- `list_tools()` をキャッシュすべきかどうか([キャッシュ](#caching))。
|
||||
- MCP アクティビティがトレースにどのように表示されるか([トレーシング](#tracing))。
|
||||
|
||||
ローカル MCP サーバー (`MCPServerStdio` 、 `MCPServerSse` 、 `MCPServerStreamableHttp`) では、承認ポリシーと呼び出しごとの `_meta` ペイロードも共通概念です。 Streamable HTTP セクションが最も完全なコード例を示しており、同じパターンが他のローカルトランスポートにも適用されます。
|
||||
ローカル MCP サーバー(`MCPServerStdio`、`MCPServerSse`、`MCPServerStreamableHttp`)では、承認ポリシーと呼び出しごとの `_meta` ペイロードも共通の概念です。Streamable HTTP セクションでは最も完全な例を示しており、同じパターンは他のローカルトランスポートにも適用されます。
|
||||
|
||||
## 1. Hosted MCP server tools
|
||||
## 1. ホスト型 MCP サーバーツール
|
||||
|
||||
Hosted ツールは、ツールの往復全体を OpenAI のインフラに委ねます。コード側でツールを列挙・呼び出す代わりに、[`HostedMCPTool`][agents.tool.HostedMCPTool] がサーバーラベル(および任意のコネクタメタデータ)を Responses API に転送します。モデルはリモートサーバーのツールを列挙し、 Python プロセスへの追加コールバックなしで実行します。 Hosted ツールは現在、 Responses API の hosted MCP 統合をサポートする OpenAI モデルで動作します。
|
||||
ホスト型ツールでは、ツールのラウンドトリップ全体を OpenAI のインフラに委ねます。ツールの一覧取得と呼び出しをコード側で行う代わりに、[`HostedMCPTool`][agents.tool.HostedMCPTool] がサーバーラベル(および任意のコネクターメタデータ)を Responses API に転送します。モデルはリモートサーバーのツールを一覧表示し、Python プロセスへの追加のコールバックなしにそれらを呼び出します。現在、ホスト型ツールは Responses API のホスト型 MCP 統合をサポートする OpenAI モデルで動作します。
|
||||
|
||||
### 基本の Hosted MCP ツール
|
||||
### 基本的なホスト型 MCP ツール
|
||||
|
||||
エージェントの `tools` リストに [`HostedMCPTool`][agents.tool.HostedMCPTool] を追加して Hosted ツールを作成します。 `tool_config` 辞書は REST API に送る JSON を反映します:
|
||||
エージェントの `tools` リストに [`HostedMCPTool`][agents.tool.HostedMCPTool] を追加してホスト型ツールを作成します。`tool_config` 辞書は REST API に送信する JSON と同じ構造です:
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -77,31 +83,36 @@ from agents import Agent, HostedMCPTool, Runner
|
||||
async def main() -> None:
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
instructions="Use the DeepWiki hosted MCP server to inspect openai/openai-agents-python.",
|
||||
tools=[
|
||||
HostedMCPTool(
|
||||
tool_config={
|
||||
"type": "mcp",
|
||||
"server_label": "gitmcp",
|
||||
"server_url": "https://gitmcp.io/openai/codex",
|
||||
"server_label": "deepwiki",
|
||||
"server_url": "https://mcp.deepwiki.com/mcp",
|
||||
"require_approval": "never",
|
||||
}
|
||||
)
|
||||
],
|
||||
)
|
||||
|
||||
result = await Runner.run(agent, "Which language is this repository written in?")
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Which language is the repository openai/openai-agents-python written in?",
|
||||
)
|
||||
print(result.final_output)
|
||||
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
Hosted サーバーはツールを自動公開するため、 `mcp_servers` に追加する必要はありません。
|
||||
ホスト型サーバーはツールを自動的に公開します。`mcp_servers` に追加する必要はありません。
|
||||
|
||||
Hosted ツール検索で hosted MCP サーバーを遅延読み込みしたい場合は、 `tool_config["defer_loading"] = True` を設定し、エージェントに [`ToolSearchTool`][agents.tool.ToolSearchTool] を追加してください。これは OpenAI Responses モデルでのみサポートされます。完全なツール検索の設定と制約は [Tools](tools.md#hosted-tool-search) を参照してください。
|
||||
ホスト型ツール検索にホスト型 MCP サーバーを遅延読み込みさせたい場合は、`tool_config["defer_loading"] = True` を設定し、エージェントに [`ToolSearchTool`][agents.tool.ToolSearchTool] を追加します。これは OpenAI Responses モデルでのみサポートされます。ツール検索の完全な設定と制約については、[ツール](tools.md#hosted-tool-search) を参照してください。
|
||||
|
||||
### Hosted MCP 結果のストリーミング
|
||||
### ホスト型 MCP の結果のストリーミング
|
||||
|
||||
Hosted ツールは、関数ツールとまったく同じ方法で結果のストリーミングをサポートします。 `Runner.run_streamed` を使うと、モデルがまだ処理中でも増分 MCP 出力を消費できます:
|
||||
ホスト型ツールは、関数ツールとまったく同じ方法で結果のストリーミングをサポートします。モデルがまだ処理中でも、`Runner.run_streamed` を使用して
|
||||
増分的な MCP 出力を受け取れます:
|
||||
|
||||
```python
|
||||
result = Runner.run_streamed(agent, "Summarise this repository's top languages")
|
||||
@@ -113,12 +124,12 @@ print(result.final_output)
|
||||
|
||||
### 任意の承認フロー
|
||||
|
||||
サーバーが機密操作を実行可能な場合、各ツール実行前に人手またはプログラムによる承認を要求できます。 `tool_config` の `require_approval` に、単一ポリシー (`"always"` 、 `"never"`) またはツール名からポリシーへの辞書を設定します。 Python 側で判断するには `on_approval_request` コールバックを提供します。
|
||||
サーバーが機密性の高い操作を実行できる場合、各ツール実行の前に人間またはプログラムによる承認を必須にできます。`tool_config` で `require_approval` を、単一のポリシー(`"always"`、`"never"`)またはツール名をポリシーにマッピングする辞書として設定します。Python 内で判断するには、`on_approval_request` コールバックを指定します。
|
||||
|
||||
```python
|
||||
from agents import MCPToolApprovalFunctionResult, MCPToolApprovalRequest
|
||||
|
||||
SAFE_TOOLS = {"read_project_metadata"}
|
||||
SAFE_TOOLS = {"read_wiki_structure", "read_wiki_contents", "ask_question"}
|
||||
|
||||
def approve_tool(request: MCPToolApprovalRequest) -> MCPToolApprovalFunctionResult:
|
||||
if request.data.name in SAFE_TOOLS:
|
||||
@@ -131,8 +142,8 @@ agent = Agent(
|
||||
HostedMCPTool(
|
||||
tool_config={
|
||||
"type": "mcp",
|
||||
"server_label": "gitmcp",
|
||||
"server_url": "https://gitmcp.io/openai/codex",
|
||||
"server_label": "deepwiki",
|
||||
"server_url": "https://mcp.deepwiki.com/mcp",
|
||||
"require_approval": "always",
|
||||
},
|
||||
on_approval_request=approve_tool,
|
||||
@@ -141,11 +152,11 @@ agent = Agent(
|
||||
)
|
||||
```
|
||||
|
||||
このコールバックは同期・非同期のどちらでもよく、モデルが実行継続のために承認データを必要とするたびに呼び出されます。
|
||||
このコールバックは同期または非同期にでき、モデルが実行を継続するための承認データを必要とするたびに呼び出されます。
|
||||
|
||||
### コネクタをバックエンドとする Hosted サーバー
|
||||
### コネクター対応のホスト型サーバー
|
||||
|
||||
Hosted MCP は OpenAI コネクタもサポートします。 `server_url` を指定する代わりに、 `connector_id` とアクセストークンを渡します。 Responses API が認証を処理し、 hosted サーバーがコネクタのツールを公開します。
|
||||
ホスト型 MCP は OpenAI コネクターにも対応しています。`server_url` を指定する代わりに、`connector_id` とアクセストークンを指定します。Responses API が認証を処理し、ホスト型サーバーがコネクターのツールを公開します。
|
||||
|
||||
```python
|
||||
import os
|
||||
@@ -161,11 +172,11 @@ HostedMCPTool(
|
||||
)
|
||||
```
|
||||
|
||||
ストリーミング、承認、コネクタを含む完全動作する Hosted ツールのサンプルは、[`examples/hosted_mcp`](https://github.com/openai/openai-agents-python/tree/main/examples/hosted_mcp) にあります。
|
||||
完全に動作するホスト型ツールのサンプル(ストリーミング、承認、コネクターを含む)は [`examples/hosted_mcp`](https://github.com/openai/openai-agents-python/tree/main/examples/hosted_mcp) にあります。
|
||||
|
||||
## 2. Streamable HTTP MCP servers
|
||||
## 2. Streamable HTTP MCP サーバー
|
||||
|
||||
ネットワーク接続を自分で管理したい場合は、[`MCPServerStreamableHttp`][agents.mcp.server.MCPServerStreamableHttp] を使用します。 Streamable HTTP サーバーは、トランスポートを制御したい場合や、低遅延を保ちながら独自インフラ内でサーバーを実行したい場合に最適です。
|
||||
ネットワーク接続を自分で管理したい場合は、[`MCPServerStreamableHttp`][agents.mcp.server.MCPServerStreamableHttp] を使用します。Streamable HTTP サーバーは、トランスポートを制御したい場合や、レイテンシを低く保ちながら自分のインフラ内でサーバーを実行したい場合に最適です。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -200,27 +211,26 @@ async def main() -> None:
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
コンストラクターは追加オプションを受け取ります:
|
||||
コンストラクターは追加オプションを受け取ります。
|
||||
|
||||
- `client_session_timeout_seconds` は HTTP の読み取りタイムアウトを制御します。
|
||||
- `use_structured_content` はテキスト出力より `tool_result.structured_content` を優先するかを切り替えます。
|
||||
- `max_retry_attempts` と `retry_backoff_seconds_base` は `list_tools()` と `call_tool()` の自動リトライを追加します。
|
||||
- `tool_filter` はツールの一部だけを公開できます([Tool filtering](#tool-filtering) 参照)。
|
||||
- `require_approval` はローカル MCP ツールで human-in-the-loop 承認ポリシーを有効化します。
|
||||
- `failure_error_function` はモデルに見える MCP ツール失敗メッセージをカスタマイズします。代わりにエラーを送出したい場合は `None` を設定します。
|
||||
- `tool_meta_resolver` は `call_tool()` 前に呼び出しごとの MCP `_meta` ペイロードを注入します。
|
||||
- `client_session_timeout_seconds` は HTTP 読み取りタイムアウトを制御します。
|
||||
- `use_structured_content` は、`tool_result.structured_content` をテキスト出力より優先するかどうかを切り替えます。
|
||||
- `max_retry_attempts` と `retry_backoff_seconds_base` は、`list_tools()` と `call_tool()` に自動リトライを追加します。
|
||||
- `tool_filter` は、ツールのサブセットのみを公開できるようにします([ツールフィルタリング](#tool-filtering) を参照)。
|
||||
- `require_approval` は、ローカル MCP ツールでヒューマンインザループの承認ポリシーを有効にします。
|
||||
- `failure_error_function` は、モデルに表示される MCP ツール失敗メッセージをカスタマイズします。`None` に設定すると、代わりにエラーを送出します。
|
||||
- `tool_meta_resolver` は、呼び出しごとの MCP `_meta` ペイロードを `call_tool()` の前に注入します。
|
||||
|
||||
### ローカル MCP サーバーの承認ポリシー
|
||||
|
||||
`MCPServerStdio` 、 `MCPServerSse` 、 `MCPServerStreamableHttp` はすべて `require_approval` を受け付けます。
|
||||
`MCPServerStdio`、`MCPServerSse`、`MCPServerStreamableHttp` はいずれも `require_approval` を受け取ります。
|
||||
|
||||
サポートされる形式:
|
||||
|
||||
- すべてのツールに対する `"always"` または `"never"` 。
|
||||
- `True` / `False` ( always/never と同等)。
|
||||
- ツールごとのマップ。例: `{"delete_file": "always", "read_file": "never"}` 。
|
||||
- グループ化オブジェクト:
|
||||
`{"always": {"tool_names": [...]}, "never": {"tool_names": [...]}}` 。
|
||||
- すべてのツールに対する `"always"` または `"never"`。
|
||||
- `True` / `False`(always/never と同等)。
|
||||
- ツールごとのマップ。例: `{"delete_file": "always", "read_file": "never"}`。
|
||||
- グループ化されたオブジェクト: `{"always": {"tool_names": [...]}, "never": {"tool_names": [...]}}`。
|
||||
|
||||
```python
|
||||
async with MCPServerStreamableHttp(
|
||||
@@ -231,11 +241,11 @@ async with MCPServerStreamableHttp(
|
||||
...
|
||||
```
|
||||
|
||||
完全な一時停止/再開フローは、 [Human-in-the-loop](human_in_the_loop.md) と `examples/mcp/get_all_mcp_tools_example/main.py` を参照してください。
|
||||
完全な一時停止 / 再開フローについては、[ヒューマンインザループ](human_in_the_loop.md) と `examples/mcp/get_all_mcp_tools_example/main.py` を参照してください。
|
||||
|
||||
### `tool_meta_resolver` による呼び出しごとのメタデータ
|
||||
|
||||
MCP サーバーが `_meta` のリクエストメタデータ(例: テナント ID やトレースコンテキスト)を必要とする場合は `tool_meta_resolver` を使います。以下の例は、 `Runner.run(...)` に `context` として `dict` を渡すことを前提にしています。
|
||||
MCP サーバーが `_meta` にリクエストメタデータ(たとえばテナント ID やトレースコンテキスト)を期待する場合は、`tool_meta_resolver` を使用します。下の例では、`Runner.run(...)` に `context` として `dict` を渡すことを前提としています。
|
||||
|
||||
```python
|
||||
from agents.mcp import MCPServerStreamableHttp, MCPToolMetaContext
|
||||
@@ -256,19 +266,19 @@ server = MCPServerStreamableHttp(
|
||||
)
|
||||
```
|
||||
|
||||
実行コンテキストが Pydantic モデル、 dataclass 、またはカスタムクラスの場合は、代わりに属性アクセスでテナント ID を読み取ってください。
|
||||
実行コンテキストが Pydantic モデル、データクラス、またはカスタムクラスの場合は、属性アクセスでテナント ID を読み取ってください。
|
||||
|
||||
### MCP ツール出力: テキストと画像
|
||||
|
||||
MCP ツールが画像コンテンツを返す場合、 SDK はそれを自動的に画像ツール出力エントリにマップします。テキスト/画像混在レスポンスは出力項目のリストとして転送されるため、エージェントは通常の関数ツールからの画像出力と同じ方法で MCP 画像結果を処理できます。
|
||||
MCP ツールが画像コンテンツを返すと、SDK はそれを画像ツール出力エントリーに自動的にマッピングします。テキスト / 画像の混在レスポンスは出力項目のリストとして転送されるため、エージェントは通常の関数ツールからの画像出力を扱うのと同じ方法で MCP の画像結果を扱えます。
|
||||
|
||||
## 3. HTTP with SSE MCP servers
|
||||
## 3. SSE を用いた HTTP MCP サーバー
|
||||
|
||||
!!! warning
|
||||
|
||||
MCP プロジェクトは Server-Sent Events トランスポートを非推奨にしています。新規統合では Streamable HTTP または stdio を優先し、 SSE はレガシーサーバー用のみにしてください。
|
||||
MCP プロジェクトでは Server-Sent Events トランスポートが非推奨になりました。新しい統合では Streamable HTTP または stdio を優先し、SSE はレガシーサーバーにのみ使用してください。
|
||||
|
||||
MCP サーバーが HTTP with SSE トランスポートを実装している場合は、[`MCPServerSse`][agents.mcp.server.MCPServerSse] をインスタンス化します。トランスポート以外の API は Streamable HTTP サーバーと同一です。
|
||||
MCP サーバーが SSE を用いた HTTP トランスポートを実装している場合は、[`MCPServerSse`][agents.mcp.server.MCPServerSse] をインスタンス化します。トランスポート以外は、API は Streamable HTTP サーバーと同一です。
|
||||
|
||||
```python
|
||||
|
||||
@@ -295,9 +305,9 @@ async with MCPServerSse(
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
## 4. stdio MCP servers
|
||||
## 4. stdio MCP サーバー
|
||||
|
||||
ローカルサブプロセスとして実行される MCP サーバーには、 [`MCPServerStdio`][agents.mcp.server.MCPServerStdio] を使います。 SDK はプロセスを起動し、パイプを開いたまま維持し、コンテキストマネージャー終了時に自動で閉じます。このオプションは、素早い概念実証や、サーバーがコマンドラインエントリポイントしか公開していない場合に有用です。
|
||||
ローカルサブプロセスとして実行される MCP サーバーには、[`MCPServerStdio`][agents.mcp.server.MCPServerStdio] を使用します。SDK はプロセスを起動し、パイプを開いたままにし、コンテキストマネージャーを抜けると自動的に閉じます。このオプションは、簡単な概念実証や、サーバーがコマンドラインエントリーポイントのみを公開する場合に役立ちます。
|
||||
|
||||
```python
|
||||
from pathlib import Path
|
||||
@@ -325,7 +335,7 @@ async with MCPServerStdio(
|
||||
|
||||
## 5. MCP サーバーマネージャー
|
||||
|
||||
複数の MCP サーバーがある場合は、 `MCPServerManager` を使って事前に接続し、接続済みサブセットをエージェントに公開します。コンストラクターオプションと再接続動作は [MCPServerManager API reference](ref/mcp/manager.md) を参照してください。
|
||||
複数の MCP サーバーがある場合は、`MCPServerManager` を使用して事前に接続し、接続済みのサブセットをエージェントに公開します。コンストラクターオプションと再接続の動作については、[MCPServerManager API リファレンス](ref/mcp/manager.md) を参照してください。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
@@ -346,25 +356,25 @@ async with MCPServerManager(servers) as manager:
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
主な挙動:
|
||||
主な動作:
|
||||
|
||||
- `active_servers` は `drop_failed_servers=True` (デフォルト)時に接続成功したサーバーのみを含みます。
|
||||
- `active_servers` には、`drop_failed_servers=True`(デフォルト)の場合、正常に接続されたサーバーのみが含まれます。
|
||||
- 失敗は `failed_servers` と `errors` で追跡されます。
|
||||
- 最初の接続失敗で例外を発生させるには `strict=True` を設定します。
|
||||
- 失敗サーバーのみ再試行するには `reconnect(failed_only=True)` 、全サーバーを再起動するには `reconnect(failed_only=False)` を呼びます。
|
||||
- ライフサイクル動作を調整するには `connect_timeout_seconds` 、 `cleanup_timeout_seconds` 、 `connect_in_parallel` を使います。
|
||||
- `strict=True` を設定すると、最初の接続失敗時にエラーを送出します。
|
||||
- `reconnect(failed_only=True)` を呼び出すと失敗したサーバーを再試行し、`reconnect(failed_only=False)` を呼び出すとすべてのサーバーを再起動します。
|
||||
- `connect_timeout_seconds`、`cleanup_timeout_seconds`、`connect_in_parallel` を使用してライフサイクル動作を調整します。
|
||||
|
||||
## 共通サーバー機能
|
||||
## 共通のサーバー機能
|
||||
|
||||
以下のセクションは MCP サーバートランスポート全体に適用されます(正確な API 表面はサーバークラスに依存します)。
|
||||
以下のセクションは MCP サーバートランスポート全体に適用されます(正確な API の範囲はサーバークラスによって異なります)。
|
||||
|
||||
## Tool filtering
|
||||
## ツールフィルタリング
|
||||
|
||||
各 MCP サーバーはツールフィルターをサポートしており、エージェントに必要な関数だけを公開できます。フィルタリングは構築時または実行ごとに動的に行えます。
|
||||
各 MCP サーバーはツールフィルターに対応しているため、エージェントに必要な関数だけを公開できます。フィルタリングは構築時に行うことも、実行ごとに動的に行うこともできます。
|
||||
|
||||
### 静的ツールフィルタリング
|
||||
### 静的なツールフィルタリング
|
||||
|
||||
シンプルな許可/ブロックリストを設定するには [`create_static_tool_filter`][agents.mcp.create_static_tool_filter] を使います:
|
||||
[`create_static_tool_filter`][agents.mcp.create_static_tool_filter] を使用して、シンプルな許可 / ブロックリストを設定します:
|
||||
|
||||
```python
|
||||
from pathlib import Path
|
||||
@@ -382,11 +392,11 @@ filesystem_server = MCPServerStdio(
|
||||
)
|
||||
```
|
||||
|
||||
`allowed_tool_names` と `blocked_tool_names` の両方が与えられた場合、 SDK はまず許可リストを適用し、その残り集合からブロック対象ツールを除外します。
|
||||
`allowed_tool_names` と `blocked_tool_names` の両方が指定された場合、SDK はまず許可リストを適用し、その後、残ったセットからブロックされたツールを削除します。
|
||||
|
||||
### 動的ツールフィルタリング
|
||||
### 動的なツールフィルタリング
|
||||
|
||||
より高度なロジックには [`ToolFilterContext`][agents.mcp.ToolFilterContext] を受け取る callable を渡します。 callable は同期・非同期のいずれでもよく、ツールを公開すべき場合に `True` を返します。
|
||||
より高度なロジックには、[`ToolFilterContext`][agents.mcp.ToolFilterContext] を受け取るコール可能オブジェクトを渡します。このコール可能オブジェクトは同期または非同期にでき、ツールを公開すべき場合に `True` を返します。
|
||||
|
||||
```python
|
||||
from pathlib import Path
|
||||
@@ -410,14 +420,15 @@ async with MCPServerStdio(
|
||||
...
|
||||
```
|
||||
|
||||
フィルターコンテキストは、アクティブな `run_context` 、ツールを要求する `agent` 、および `server_name` を公開します。
|
||||
フィルターコンテキストは、アクティブな `run_context`、ツールを要求している `agent`、および `server_name` を公開します。
|
||||
|
||||
## Prompts
|
||||
## プロンプト
|
||||
|
||||
MCP サーバーは、エージェント指示を動的生成するプロンプトも提供できます。プロンプト対応サーバーは次の 2 つのメソッドを公開します:
|
||||
MCP サーバーは、エージェントの指示を動的に生成するプロンプトも提供できます。プロンプトに対応するサーバーは 2 つの
|
||||
メソッドを公開します:
|
||||
|
||||
- `list_prompts()` は利用可能なプロンプトテンプレートを列挙します。
|
||||
- `get_prompt(name, arguments)` は具体的なプロンプトを取得します(必要に応じてパラメーター付き)。
|
||||
- `get_prompt(name, arguments)` は具体的なプロンプトを取得します。任意でパラメーターを指定できます。
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
@@ -435,21 +446,21 @@ agent = Agent(
|
||||
)
|
||||
```
|
||||
|
||||
## Caching
|
||||
## キャッシュ
|
||||
|
||||
各エージェント実行は各 MCP サーバーで `list_tools()` を呼びます。リモートサーバーは目立つレイテンシを生む可能性があるため、すべての MCP サーバークラスは `cache_tools_list` オプションを公開しています。ツール定義が頻繁に変わらないと確信できる場合にのみ `True` に設定してください。後で最新リストを強制したい場合は、サーバーインスタンスで `invalidate_tools_cache()` を呼びます。
|
||||
エージェントを実行するたびに、各 MCP サーバーで `list_tools()` が呼び出されます。リモートサーバーでは目に見えるレイテンシが発生する可能性があるため、すべての MCP サーバークラスは `cache_tools_list` オプションを公開しています。ツール定義が頻繁に変更されないと確信できる場合にのみ、`True` に設定してください。後で最新のリストを強制的に取得するには、サーバーインスタンスで `invalidate_tools_cache()` を呼び出します。
|
||||
|
||||
## Tracing
|
||||
## トレーシング
|
||||
|
||||
[Tracing](./tracing.md) は、以下を含む MCP アクティビティを自動で記録します:
|
||||
[トレーシング](./tracing.md) は、次を含む MCP アクティビティを自動的にキャプチャします。
|
||||
|
||||
1. ツール一覧取得のための MCP サーバー呼び出し。
|
||||
1. ツールを一覧表示するための MCP サーバーへの呼び出し。
|
||||
2. ツール呼び出し上の MCP 関連情報。
|
||||
|
||||

|
||||

|
||||
|
||||
## 参考情報
|
||||
|
||||
- [Model Context Protocol](https://modelcontextprotocol.io/) – 仕様と設計ガイド。
|
||||
- [examples/mcp](https://github.com/openai/openai-agents-python/tree/main/examples/mcp) – 実行可能な stdio 、 SSE 、 Streamable HTTP サンプル。
|
||||
- [examples/hosted_mcp](https://github.com/openai/openai-agents-python/tree/main/examples/hosted_mcp) – 承認とコネクタを含む完全な hosted MCP デモ。
|
||||
- [examples/mcp](https://github.com/openai/openai-agents-python/tree/main/examples/mcp) – 実行可能な stdio、SSE、Streamable HTTP のサンプル。
|
||||
- [examples/hosted_mcp](https://github.com/openai/openai-agents-python/tree/main/examples/hosted_mcp) – 承認とコネクターを含む、完全なホスト型 MCP デモ。
|
||||
+194
-151
@@ -4,42 +4,42 @@ search:
|
||||
---
|
||||
# モデル
|
||||
|
||||
Agents SDK には、OpenAI モデル向けの即時利用可能なサポートが 2 つの形式で含まれています:
|
||||
Agents SDK には、すぐに使える OpenAI モデルのサポートが 2 種類用意されています:
|
||||
|
||||
- **推奨**: 新しい [Responses API](https://platform.openai.com/docs/api-reference/responses) を使用して OpenAI API を呼び出す [`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel]。
|
||||
- [Chat Completions API](https://platform.openai.com/docs/api-reference/chat) を使用して OpenAI API を呼び出す [`OpenAIChatCompletionsModel`][agents.models.openai_chatcompletions.OpenAIChatCompletionsModel]。
|
||||
- **推奨**: 新しい [Responses API](https://platform.openai.com/docs/api-reference/responses) を使用して OpenAI API を呼び出す [`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel]。
|
||||
- [Chat Completions API](https://platform.openai.com/docs/api-reference/chat) を使用して OpenAI API を呼び出す [`OpenAIChatCompletionsModel`][agents.models.openai_chatcompletions.OpenAIChatCompletionsModel]。
|
||||
|
||||
## モデル設定の選択
|
||||
|
||||
ご利用環境に合う最もシンプルな経路から開始してください:
|
||||
セットアップに合う最もシンプルな方法から始めてください:
|
||||
|
||||
| If you are trying to... | Recommended path | Read more |
|
||||
| 目的 | 推奨パス | 詳細 |
|
||||
| --- | --- | --- |
|
||||
| OpenAI モデルのみを使用する | 既定の OpenAI provider と Responses model path を使用する | [OpenAI モデル](#openai-models) |
|
||||
| websocket transport で OpenAI Responses API を使用する | Responses model path を維持し、websocket transport を有効化する | [Responses WebSocket transport](#responses-websocket-transport) |
|
||||
| 1 つの non-OpenAI provider を使用する | 組み込み provider 統合ポイントから開始する | [Non-OpenAI モデル](#non-openai-models) |
|
||||
| エージェント間でモデルまたは provider を混在させる | 実行ごとまたはエージェントごとに provider を選択し、機能差を確認する | [1 つのワークフローでのモデル混在](#mixing-models-in-one-workflow) と [provider 間でのモデル混在](#mixing-models-across-providers) |
|
||||
| 高度な OpenAI Responses リクエスト設定を調整する | OpenAI Responses path で `ModelSettings` を使用する | [高度な OpenAI Responses 設定](#advanced-openai-responses-settings) |
|
||||
| non-OpenAI または mixed-provider ルーティング用にサードパーティ adapter を使用する | サポートされる beta adapter を比較し、提供予定の provider path を検証する | [サードパーティ adapter](#third-party-adapters) |
|
||||
| OpenAI モデルのみを使用する | デフォルトの OpenAI プロバイダーを Responses モデルパスで使用する | [OpenAI モデル](#openai-models) |
|
||||
| websocket トランスポート経由で OpenAI Responses API を使用する | Responses モデルパスを維持し、websocket トランスポートを有効にする | [Responses WebSocket トランスポート](#responses-websocket-transport) |
|
||||
| 1 つの非 OpenAI プロバイダーを使用する | 組み込みのプロバイダー統合ポイントから始める | [非 OpenAI モデル](#non-openai-models) |
|
||||
| エージェント間でモデルまたはプロバイダーを混在させる | 実行ごとまたはエージェントごとにプロバイダーを選択し、機能差異を確認する | [1 つのワークフロー内でのモデルの混在](#mixing-models-in-one-workflow) と [プロバイダー間でのモデルの混在](#mixing-models-across-providers) |
|
||||
| 高度な OpenAI Responses リクエスト設定を調整する | OpenAI Responses パスで `ModelSettings` を使用する | [高度な OpenAI Responses 設定](#advanced-openai-responses-settings) |
|
||||
| 非 OpenAI または混在プロバイダーのルーティングにサードパーティアダプターを使用する | サポートされているベータ版アダプターを比較し、出荷予定のプロバイダーパスを検証する | [サードパーティアダプター](#third-party-adapters) |
|
||||
|
||||
## OpenAI モデル
|
||||
|
||||
ほとんどの OpenAI 専用アプリでは、推奨経路は既定の OpenAI provider で文字列のモデル名を使い、Responses model path を維持することです。
|
||||
ほとんどの OpenAI のみのアプリでは、デフォルトの OpenAI プロバイダーで文字列のモデル名を使用し、Responses モデルパスを維持する方法をお勧めします。
|
||||
|
||||
`Agent` の初期化時にモデルを指定しない場合、既定モデルが使用されます。現在の既定は互換性と低遅延のため [`gpt-4.1`](https://developers.openai.com/api/docs/models/gpt-4.1) です。利用可能であれば、明示的な `model_settings` を維持したまま、より高品質な [`gpt-5.4`](https://developers.openai.com/api/docs/models/gpt-5.4) をエージェントに設定することを推奨します。
|
||||
`Agent` の初期化時にモデルを指定しない場合、デフォルトモデルが使用されます。デフォルトは現在、低レイテンシのエージェントワークフロー向けに `reasoning.effort="none"` と `verbosity="low"` を指定した [`gpt-5.4-mini`](https://developers.openai.com/api/docs/models/gpt-5.4-mini) です。アクセス権がある場合は、明示的な `model_settings` を維持しながら、より高い品質のためにエージェントを [`gpt-5.5`](https://developers.openai.com/api/docs/models/gpt-5.5) に設定することをお勧めします。
|
||||
|
||||
[`gpt-5.4`](https://developers.openai.com/api/docs/models/gpt-5.4) のような他モデルへ切り替える場合、エージェントを設定する方法は 2 つあります。
|
||||
[`gpt-5.5`](https://developers.openai.com/api/docs/models/gpt-5.5) のような他のモデルに切り替えたい場合、エージェントを設定する方法は 2 つあります。
|
||||
|
||||
### 既定モデル
|
||||
### デフォルトモデル
|
||||
|
||||
まず、カスタムモデルを設定していないすべてのエージェントで特定モデルを一貫して使いたい場合は、エージェント実行前に `OPENAI_DEFAULT_MODEL` 環境変数を設定します。
|
||||
まず、カスタムモデルを設定していないすべてのエージェントで特定のモデルを一貫して使用したい場合は、エージェントを実行する前に `OPENAI_DEFAULT_MODEL` 環境変数を設定します。
|
||||
|
||||
```bash
|
||||
export OPENAI_DEFAULT_MODEL=gpt-5.4
|
||||
export OPENAI_DEFAULT_MODEL=gpt-5.5
|
||||
python3 my_awesome_agent.py
|
||||
```
|
||||
|
||||
次に、`RunConfig` を通じて実行単位の既定モデルを設定できます。エージェントにモデルを設定しない場合は、この実行のモデルが使われます。
|
||||
次に、`RunConfig` を通じて実行のデフォルトモデルを設定できます。エージェントにモデルを設定しない場合、この実行のモデルが使用されます。
|
||||
|
||||
```python
|
||||
from agents import Agent, RunConfig, Runner
|
||||
@@ -52,13 +52,13 @@ agent = Agent(
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Hello",
|
||||
run_config=RunConfig(model="gpt-5.4"),
|
||||
run_config=RunConfig(model="gpt-5.5"),
|
||||
)
|
||||
```
|
||||
|
||||
#### GPT-5 モデル
|
||||
|
||||
この方法で [`gpt-5.4`](https://developers.openai.com/api/docs/models/gpt-5.4) など任意の GPT-5 モデルを使うと、SDK は既定の `ModelSettings` を適用します。ほとんどのユースケースで最適に動作する設定です。既定モデルの推論 effort を調整するには、独自の `ModelSettings` を渡します:
|
||||
この方法で [`gpt-5.5`](https://developers.openai.com/api/docs/models/gpt-5.5) など任意の GPT-5 モデルを使用すると、SDK はデフォルトの `ModelSettings` を適用します。ほとんどのユースケースで最も適切に機能する設定が適用されます。デフォルトモデルの推論 effort を調整するには、独自の `ModelSettings` を渡します:
|
||||
|
||||
```python
|
||||
from openai.types.shared import Reasoning
|
||||
@@ -67,42 +67,42 @@ from agents import Agent, ModelSettings
|
||||
my_agent = Agent(
|
||||
name="My Agent",
|
||||
instructions="You're a helpful agent.",
|
||||
# If OPENAI_DEFAULT_MODEL=gpt-5.4 is set, passing only model_settings works.
|
||||
# If OPENAI_DEFAULT_MODEL=gpt-5.5 is set, passing only model_settings works.
|
||||
# It's also fine to pass a GPT-5 model name explicitly:
|
||||
model="gpt-5.4",
|
||||
model="gpt-5.5",
|
||||
model_settings=ModelSettings(reasoning=Reasoning(effort="high"), verbosity="low")
|
||||
)
|
||||
```
|
||||
|
||||
より低遅延にするには、`gpt-5.4` で `reasoning.effort="none"` の使用が推奨されます。gpt-4.1 ファミリー( mini / nano バリアントを含む)も、対話型エージェントアプリ構築における堅実な選択肢です。
|
||||
低レイテンシにするには、GPT-5 モデルで `reasoning.effort="none"` を使用することをお勧めします。
|
||||
|
||||
#### ComputerTool モデル選択
|
||||
#### ComputerTool モデルの選択
|
||||
|
||||
エージェントに [`ComputerTool`][agents.tool.ComputerTool] が含まれる場合、実際の Responses リクエストで有効なモデルにより、SDK が送信するコンピュータツール payload が決まります。明示的な `gpt-5.4` リクエストでは GA の組み込み `computer` ツールを使用し、明示的な `computer-use-preview` リクエストでは旧 `computer_use_preview` payload を維持します。
|
||||
エージェントに [`ComputerTool`][agents.tool.ComputerTool] が含まれる場合、実際の Responses リクエスト上の有効なモデルによって、SDK が送信する computer-tool ペイロードが決まります。明示的な `gpt-5.5` リクエストでは GA 組み込みの `computer` ツールが使用され、明示的な `computer-use-preview` リクエストでは古い `computer_use_preview` ペイロードが維持されます。
|
||||
|
||||
主な例外は prompt 管理呼び出しです。prompt template がモデルを管理し、SDK がリクエストから `model` を省略する場合、SDK は prompt 固定モデルを推測しないよう preview 互換のコンピュータ payload を既定で使います。このフローで GA path を維持するには、リクエストで `model="gpt-5.4"` を明示するか、`ModelSettings(tool_choice="computer")` または `ModelSettings(tool_choice="computer_use")` で GA セレクターを強制します。
|
||||
主な例外は、プロンプト管理の呼び出しです。プロンプトテンプレートがモデルを所有し、SDK がリクエストから `model` を省略する場合、SDK はプロンプトがどのモデルを固定しているかを推測しないよう、プレビュー互換のコンピューターペイロードをデフォルトにします。そのフローで GA パスを維持するには、リクエストで `model="gpt-5.5"` を明示するか、`ModelSettings(tool_choice="computer")` または `ModelSettings(tool_choice="computer_use")` で GA セレクターを強制します。
|
||||
|
||||
[`ComputerTool`][agents.tool.ComputerTool] が登録されている場合、`tool_choice="computer"`、`"computer_use"`、`"computer_use_preview"` は有効リクエストモデルに一致する組み込みセレクターに正規化されます。`ComputerTool` が登録されていない場合、これらの文字列は通常の関数名として動作し続けます。
|
||||
登録済みの [`ComputerTool`][agents.tool.ComputerTool] がある場合、`tool_choice="computer"`、`"computer_use"`、`"computer_use_preview"` は、有効なリクエストモデルに一致する組み込みセレクターに正規化されます。`ComputerTool` が登録されていない場合、これらの文字列は通常の関数名として引き続き動作します。
|
||||
|
||||
preview 互換リクエストでは `environment` と表示寸法を事前に serialize する必要があるため、[`ComputerProvider`][agents.tool.ComputerProvider] ファクトリーを使う prompt 管理フローでは、具体的な `Computer` または `AsyncComputer` インスタンスを渡すか、リクエスト送信前に GA セレクターを強制する必要があります。移行の詳細は [Tools](../tools.md#computertool-and-the-responses-computer-tool) を参照してください。
|
||||
プレビュー互換リクエストでは、`environment` と表示サイズを事前にシリアライズする必要があるため、[`ComputerProvider`][agents.tool.ComputerProvider] ファクトリを使用するプロンプト管理フローでは、具象の `Computer` または `AsyncComputer` インスタンスを渡すか、リクエスト送信前に GA セレクターを強制する必要があります。移行の詳細については、[ツール](../tools.md#computertool-and-the-responses-computer-tool)を参照してください。
|
||||
|
||||
#### 非 GPT-5 モデル
|
||||
|
||||
カスタム `model_settings` なしで非 GPT-5 モデル名を渡した場合、SDK は任意モデル互換の汎用 `ModelSettings` に戻ります。
|
||||
カスタムの `model_settings` なしで非 GPT-5 モデル名を渡すと、SDK は任意のモデルと互換性のある汎用の `ModelSettings` に戻ります。
|
||||
|
||||
### Responses 専用ツール検索機能
|
||||
### Responses 専用のツール検索機能
|
||||
|
||||
以下のツール機能は OpenAI Responses モデルでのみサポートされます:
|
||||
次のツール機能は、OpenAI Responses モデルでのみサポートされています:
|
||||
|
||||
- [`ToolSearchTool`][agents.tool.ToolSearchTool]
|
||||
- [`tool_namespace()`][agents.tool.tool_namespace]
|
||||
- `@function_tool(defer_loading=True)` およびその他の deferred-loading Responses ツール surface
|
||||
- [`ToolSearchTool`][agents.tool.ToolSearchTool]
|
||||
- [`tool_namespace()`][agents.tool.tool_namespace]
|
||||
- `@function_tool(defer_loading=True)` と、その他の遅延読み込みの Responses ツールサーフェス
|
||||
|
||||
これらの機能は Chat Completions モデルと non-Responses backend では拒否されます。deferred-loading ツールを使う場合は、エージェントに `ToolSearchTool()` を追加し、素の namespace 名や deferred 専用関数名を強制せず、`auto` または `required` の tool choice でモデルにツールをロードさせてください。設定詳細と現在の制約は [Tools](../tools.md#hosted-tool-search) を参照してください。
|
||||
これらの機能は、Chat Completions モデルおよび非 Responses バックエンドでは拒否されます。遅延読み込みツールを使用する場合は、`ToolSearchTool()` をエージェントに追加し、素の名前空間名や遅延専用の関数名を強制する代わりに、`auto` または `required` のツール選択を通じてモデルにツールを読み込ませます。設定の詳細と現在の制約については、[ツール](../tools.md#hosted-tool-search)を参照してください。
|
||||
|
||||
### Responses WebSocket transport
|
||||
### Responses WebSocket トランスポート
|
||||
|
||||
既定では、OpenAI Responses API リクエストは HTTP transport を使います。OpenAI バックエンドモデル使用時に websocket transport を有効化できます。
|
||||
デフォルトでは、OpenAI Responses API リクエストは HTTP トランスポートを使用します。OpenAI を基盤とするモデルを使用する場合、websocket トランスポートをオプトインできます。
|
||||
|
||||
#### 基本設定
|
||||
|
||||
@@ -112,13 +112,13 @@ from agents import set_default_openai_responses_transport
|
||||
set_default_openai_responses_transport("websocket")
|
||||
```
|
||||
|
||||
これは既定の OpenAI provider により解決される OpenAI Responses モデル(`"gpt-5.4"` などの文字列モデル名を含む)に影響します。
|
||||
これは、デフォルトの OpenAI プロバイダーによって解決される OpenAI Responses モデル(`"gpt-5.5"` などの文字列モデル名を含む)に影響します。
|
||||
|
||||
transport の選択は、SDK がモデル名をモデルインスタンスへ解決する時点で行われます。具体的な [`Model`][agents.models.interface.Model] オブジェクトを渡す場合、その transport はすでに固定です: [`OpenAIResponsesWSModel`][agents.models.openai_responses.OpenAIResponsesWSModel] は websocket、[`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel] は HTTP、[`OpenAIChatCompletionsModel`][agents.models.openai_chatcompletions.OpenAIChatCompletionsModel] は Chat Completions のままです。`RunConfig(model_provider=...)` を渡した場合、global default ではなくその provider が transport 選択を制御します。
|
||||
トランスポートの選択は、SDK がモデル名をモデルインスタンスへ解決するときに行われます。具象の [`Model`][agents.models.interface.Model] オブジェクトを渡す場合、そのトランスポートはすでに固定されています。[`OpenAIResponsesWSModel`][agents.models.openai_responses.OpenAIResponsesWSModel] は websocket を使用し、[`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel] は HTTP を使用し、[`OpenAIChatCompletionsModel`][agents.models.openai_chatcompletions.OpenAIChatCompletionsModel] は Chat Completions のままです。`RunConfig(model_provider=...)` を渡す場合、グローバルデフォルトではなく、そのプロバイダーがトランスポート選択を制御します。
|
||||
|
||||
#### provider / 実行レベル設定
|
||||
#### プロバイダーまたは実行レベルの設定
|
||||
|
||||
websocket transport は provider 単位または実行単位でも設定できます:
|
||||
websocket トランスポートは、プロバイダーごと、または実行ごとにも設定できます:
|
||||
|
||||
```python
|
||||
from agents import Agent, OpenAIProvider, RunConfig, Runner
|
||||
@@ -127,6 +127,8 @@ provider = OpenAIProvider(
|
||||
use_responses_websocket=True,
|
||||
# Optional; if omitted, OPENAI_WEBSOCKET_BASE_URL is used when set.
|
||||
websocket_base_url="wss://your-proxy.example/v1",
|
||||
# Optional low-level websocket keepalive settings.
|
||||
responses_websocket_options={"ping_interval": 20.0, "ping_timeout": 60.0},
|
||||
)
|
||||
|
||||
agent = Agent(name="Assistant")
|
||||
@@ -137,7 +139,7 @@ result = await Runner.run(
|
||||
)
|
||||
```
|
||||
|
||||
OpenAI バックエンド provider は任意のエージェント登録設定も受け付けます。これは OpenAI 設定が harness ID などの provider レベル登録メタデータを期待するケース向けの高度なオプションです。
|
||||
OpenAI を基盤とするプロバイダーは、任意のエージェント登録設定も受け付けます。これは、OpenAI のセットアップで harness ID などのプロバイダーレベルの登録メタデータが想定される場合向けの高度なオプションです。
|
||||
|
||||
```python
|
||||
from agents import (
|
||||
@@ -163,14 +165,14 @@ result = await Runner.run(
|
||||
|
||||
#### `MultiProvider` による高度なルーティング
|
||||
|
||||
prefix ベースのモデルルーティング(例: 1 回の実行で `openai/...` と `any-llm/...` モデル名を混在)を必要とする場合は、[`MultiProvider`][agents.MultiProvider] を使用し、そこで `openai_use_responses_websocket=True` を設定してください。
|
||||
プレフィックスベースのモデルルーティングが必要な場合(たとえば、1 つの実行内で `openai/...` と `any-llm/...` のモデル名を混在させる場合)は、[`MultiProvider`][agents.MultiProvider] を使用し、そこで `openai_use_responses_websocket=True` を設定します。
|
||||
|
||||
`MultiProvider` は 2 つの履歴的既定値を維持します:
|
||||
`MultiProvider` は、歴史的なデフォルトを 2 つ維持しています:
|
||||
|
||||
- `openai/...` は OpenAI provider の alias として扱われるため、`openai/gpt-4.1` はモデル `gpt-4.1` としてルーティングされます。
|
||||
- 不明な prefix は pass-through されず `UserError` を発生させます。
|
||||
- `openai/...` は OpenAI プロバイダーのエイリアスとして扱われるため、`openai/gpt-4.1` はモデル `gpt-4.1` としてルーティングされます。
|
||||
- 不明なプレフィックスは、パススルーされる代わりに `UserError` を発生させます。
|
||||
|
||||
OpenAI provider を、文字通り namespaced モデル ID を期待する OpenAI 互換 endpoint に向ける場合は、明示的に pass-through 動作を有効化してください。websocket 有効構成では、`MultiProvider` 側でも `openai_use_responses_websocket=True` を維持します:
|
||||
リテラルな名前空間付きモデル ID を想定する OpenAI 互換エンドポイントに OpenAI プロバイダーを向ける場合は、パススルー動作を明示的にオプトインしてください。websocket が有効なセットアップでは、`MultiProvider` でも `openai_use_responses_websocket=True` を維持します:
|
||||
|
||||
```python
|
||||
from agents import Agent, MultiProvider, RunConfig, Runner
|
||||
@@ -196,38 +198,40 @@ result = await Runner.run(
|
||||
)
|
||||
```
|
||||
|
||||
backend が文字列 `openai/...` をそのまま期待する場合は `openai_prefix_mode="model_id"` を使います。`openrouter/openai/gpt-4.1-mini` のような他の namespaced モデル ID を backend が期待する場合は `unknown_prefix_mode="model_id"` を使います。これらのオプションは websocket transport 外の `MultiProvider` でも動作します。この例で websocket を有効にしているのは、この節で説明している transport 設定の一部だからです。同じオプションは [`responses_websocket_session()`][agents.responses_websocket_session] でも利用可能です。
|
||||
バックエンドがリテラルな `openai/...` 文字列を想定する場合は、`openai_prefix_mode="model_id"` を使用します。バックエンドが `openrouter/openai/gpt-4.1-mini` のような他の名前空間付きモデル ID を想定する場合は、`unknown_prefix_mode="model_id"` を使用します。これらのオプションは、websocket トランスポート以外の `MultiProvider` でも機能します。この例で websocket を有効にしているのは、このセクションで説明しているトランスポート設定の一部であるためです。同じオプションは [`responses_websocket_session()`][agents.responses_websocket_session] でも利用できます。
|
||||
|
||||
`MultiProvider` 経由ルーティング時にも同じ provider レベル登録メタデータが必要な場合は、`openai_agent_registration=OpenAIAgentRegistrationConfig(...)` を渡すと、基盤の OpenAI provider へ転送されます。
|
||||
`MultiProvider` 経由でルーティングしながら同じプロバイダーレベルの登録メタデータが必要な場合は、`openai_agent_registration=OpenAIAgentRegistrationConfig(...)` を渡すと、基盤となる OpenAI プロバイダーに転送されます。
|
||||
|
||||
カスタム OpenAI 互換 endpoint または proxy を使う場合、websocket transport には互換 websocket `/responses` endpoint も必要です。これらの構成では `websocket_base_url` を明示設定する必要がある場合があります。
|
||||
カスタムの OpenAI 互換エンドポイントまたはプロキシを使用する場合、websocket トランスポートには互換性のある websocket `/responses` エンドポイントも必要です。そのようなセットアップでは、`websocket_base_url` を明示的に設定する必要がある場合があります。
|
||||
|
||||
#### 注記
|
||||
|
||||
- これは websocket transport 上の Responses API であり、[Realtime API](../realtime/guide.md) ではありません。Chat Completions や、Responses websocket `/responses` endpoint をサポートしない non-OpenAI provider には適用されません。
|
||||
- 環境に未導入であれば `websockets` パッケージをインストールしてください。
|
||||
- websocket transport 有効化後は [`Runner.run_streamed()`][agents.run.Runner.run_streamed] を直接使用できます。複数ターンのワークフローで同一 websocket 接続をターン間(およびネストした agent-as-tool 呼び出し間)で再利用したい場合は、[`responses_websocket_session()`][agents.responses_websocket_session] ヘルパーを推奨します。[Running agents](../running_agents.md) ガイドと [`examples/basic/stream_ws.py`](https://github.com/openai/openai-agents-python/tree/main/examples/basic/stream_ws.py) を参照してください。
|
||||
- これは [Realtime API](../realtime/guide.md) ではなく、websocket トランスポート経由の Responses API です。Responses websocket `/responses` エンドポイントをサポートしていない限り、Chat Completions や非 OpenAI プロバイダーには適用されません。
|
||||
- 環境でまだ利用できない場合は、`websockets` パッケージをインストールしてください。
|
||||
- websocket トランスポートを有効にした後、[`Runner.run_streamed()`][agents.run.Runner.run_streamed] を直接使用できます。複数ターンのワークフローで、ターン間(およびネストされた agent-as-tool 呼び出し)で同じ websocket 接続を再利用したい場合は、[`responses_websocket_session()`][agents.responses_websocket_session] ヘルパーをお勧めします。[エージェントの実行](../running_agents.md)ガイドと [`examples/basic/stream_ws.py`](https://github.com/openai/openai-agents-python/tree/main/examples/basic/stream_ws.py) を参照してください。
|
||||
- 長い推論ターンやレイテンシスパイクのあるネットワークでは、`responses_websocket_options` で websocket の keepalive 動作をカスタマイズします。遅延した pong フレームを許容するには `ping_timeout` を増やすか、ping を有効にしたままハートビートタイムアウトを無効にするには `ping_timeout=None` を設定します。信頼性が websocket のレイテンシより重要な場合は、HTTP/SSE トランスポートを優先してください。
|
||||
- デフォルトでは、SDK は受信メッセージサイズの上限を無効にします(`max_size=None`)。プロキシ背後の長寿命エージェントプロセスや、メモリ制約のあるコンテナーでは、メッセージごとのメモリ使用量に上限を設けるために `responses_websocket_options={"max_size": 8 * 1024 * 1024}` を設定します。
|
||||
|
||||
## Non-OpenAI モデル
|
||||
## 非 OpenAI モデル
|
||||
|
||||
non-OpenAI provider が必要な場合、まず SDK 組み込みの provider 統合ポイントから始めてください。多くの構成ではサードパーティ adapter を追加せずに十分です。各パターンの例は [examples/model_providers](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/) にあります。
|
||||
非 OpenAI プロバイダーが必要な場合は、SDK に組み込まれたプロバイダー統合ポイントから始めてください。多くのセットアップでは、サードパーティアダプターを追加しなくてもこれで十分です。各パターンのコード例は [examples/model_providers](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/) にあります。
|
||||
|
||||
### non-OpenAI provider 統合方法
|
||||
### 非 OpenAI プロバイダーの統合方法
|
||||
|
||||
| Approach | Use it when | Scope |
|
||||
| アプローチ | 使用する場合 | 範囲 |
|
||||
| --- | --- | --- |
|
||||
| [`set_default_openai_client`][agents.set_default_openai_client] | 1 つの OpenAI 互換 endpoint を大半または全エージェントの既定にしたい | グローバル既定 |
|
||||
| [`ModelProvider`][agents.models.interface.ModelProvider] | 1 つのカスタム provider を単一実行に適用したい | 実行単位 |
|
||||
| [`Agent.model`][agents.agent.Agent.model] | エージェントごとに異なる provider または具体モデルオブジェクトが必要 | エージェント単位 |
|
||||
| サードパーティ adapter | 組み込み経路で提供されない adapter 管理の provider カバレッジまたはルーティングが必要 | [サードパーティ adapters](#third-party-adapters) を参照 |
|
||||
| [`set_default_openai_client`][agents.set_default_openai_client] | 1 つの OpenAI 互換エンドポイントを、ほとんどまたはすべてのエージェントのデフォルトにする必要がある場合 | グローバルデフォルト |
|
||||
| [`ModelProvider`][agents.models.interface.ModelProvider] | 1 つのカスタムプロバイダーを単一の実行に適用する必要がある場合 | 実行ごと |
|
||||
| [`Agent.model`][agents.agent.Agent.model] | 異なるエージェントが異なるプロバイダーまたは具象モデルオブジェクトを必要とする場合 | エージェントごと |
|
||||
| サードパーティアダプター | 組み込みパスでは提供されない、アダプター管理のプロバイダー対応範囲またはルーティングが必要な場合 | [サードパーティアダプター](#third-party-adapters) を参照 |
|
||||
|
||||
これらの組み込み経路で他の LLM provider を統合できます:
|
||||
これらの組み込みパスを使って、他の LLM プロバイダーを統合できます:
|
||||
|
||||
1. [`set_default_openai_client`][agents.set_default_openai_client] は、`AsyncOpenAI` インスタンスを LLM クライアントとしてグローバル利用したい場合に有用です。これは LLM provider が OpenAI 互換 API endpoint を持ち、`base_url` と `api_key` を設定できるケース向けです。設定可能な例は [examples/model_providers/custom_example_global.py](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/custom_example_global.py) を参照してください。
|
||||
2. [`ModelProvider`][agents.models.interface.ModelProvider] は `Runner.run` レベルです。これにより「この実行の全エージェントでカスタムモデル provider を使う」と指定できます。設定可能な例は [examples/model_providers/custom_example_provider.py](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/custom_example_provider.py) を参照してください。
|
||||
3. [`Agent.model`][agents.agent.Agent.model] は特定 Agent インスタンスでモデルを指定できます。これによりエージェントごとに異なる provider を混在できます。設定可能な例は [examples/model_providers/custom_example_agent.py](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/custom_example_agent.py) を参照してください。
|
||||
1. [`set_default_openai_client`][agents.set_default_openai_client] は、`AsyncOpenAI` のインスタンスを LLM クライアントとしてグローバルに使用したい場合に便利です。これは、LLM プロバイダーが OpenAI 互換 API エンドポイントを持ち、`base_url` と `api_key` を設定できる場合向けです。設定可能なコード例は [examples/model_providers/custom_example_global.py](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/custom_example_global.py) を参照してください。
|
||||
2. [`ModelProvider`][agents.models.interface.ModelProvider] は `Runner.run` レベルのものです。これにより、「この実行内のすべてのエージェントでカスタムモデルプロバイダーを使用する」と指定できます。設定可能なコード例は [examples/model_providers/custom_example_provider.py](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/custom_example_provider.py) を参照してください。
|
||||
3. [`Agent.model`][agents.agent.Agent.model] では、特定の Agent インスタンスでモデルを指定できます。これにより、異なるエージェントに対して異なるプロバイダーを組み合わせて使用できます。設定可能なコード例は [examples/model_providers/custom_example_agent.py](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/custom_example_agent.py) を参照してください。
|
||||
|
||||
`platform.openai.com` の API key がない場合は、`set_tracing_disabled()` でトレーシングを無効化するか、[別のトレーシングプロセッサー](../tracing.md) を設定することを推奨します。
|
||||
`platform.openai.com` の API キーがない場合は、`set_tracing_disabled()` でトレーシングを無効にするか、[別のトレーシングプロセッサー](../tracing.md)を設定することをお勧めします。
|
||||
|
||||
``` python
|
||||
from agents import Agent, AsyncOpenAI, OpenAIChatCompletionsModel, set_tracing_disabled
|
||||
@@ -242,19 +246,19 @@ agent= Agent(name="Helping Agent", instructions="You are a Helping Agent", model
|
||||
|
||||
!!! note
|
||||
|
||||
これらの例では、多くの LLM provider がまだ Responses API をサポートしていないため、Chat Completions API / model を使用しています。LLM provider が対応している場合は Responses の使用を推奨します。
|
||||
これらのコード例では Chat Completions API/モデルを使用しています。多くの LLM プロバイダーは、まだ Responses API をサポートしていないためです。LLM プロバイダーが Responses API をサポートしている場合は、Responses の使用をお勧めします。
|
||||
|
||||
## 1 つのワークフローでのモデル混在
|
||||
## 1 つのワークフロー内でのモデルの混在
|
||||
|
||||
単一ワークフロー内で、エージェントごとに異なるモデルを使いたい場合があります。たとえば、トリアージにはより小型で高速なモデル、複雑タスクにはより大型で高性能なモデルを使えます。[`Agent`][agents.Agent] 設定時は、次のいずれかで特定モデルを選択できます:
|
||||
単一のワークフロー内で、エージェントごとに異なるモデルを使用したい場合があります。たとえば、トリアージには小型で高速なモデルを使用し、複雑なタスクには大型で高性能なモデルを使用できます。[`Agent`][agents.Agent] を設定する際、次のいずれかの方法で特定のモデルを選択できます:
|
||||
|
||||
1. モデル名を渡す。
|
||||
2. 任意のモデル名 + その名前を Model インスタンスへマッピングできる [`ModelProvider`][agents.models.interface.ModelProvider] を渡す。
|
||||
3. [`Model`][agents.models.interface.Model] 実装を直接渡す。
|
||||
2. 任意のモデル名と、その名前を Model インスタンスへマッピングできる [`ModelProvider`][agents.models.interface.ModelProvider] を渡す。
|
||||
3. [`Model`][agents.models.interface.Model] 実装を直接提供する。
|
||||
|
||||
!!! note
|
||||
|
||||
SDK は [`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel] と [`OpenAIChatCompletionsModel`][agents.models.openai_chatcompletions.OpenAIChatCompletionsModel] の両 shape をサポートしますが、2 つの shape は対応機能とツール集合が異なるため、各ワークフローでは単一 shape の使用を推奨します。shape を混在させる必要がある場合は、使用する全機能が両方で利用可能であることを確認してください。
|
||||
SDK は [`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel] と [`OpenAIChatCompletionsModel`][agents.models.openai_chatcompletions.OpenAIChatCompletionsModel] の両方の形式をサポートしていますが、2 つの形式はサポートする機能やツールのセットが異なるため、各ワークフローでは 1 つのモデル形式を使用することをお勧めします。ワークフローでモデル形式を混在させる必要がある場合は、使用しているすべての機能が両方で利用可能であることを確認してください。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, AsyncOpenAI, OpenAIChatCompletionsModel
|
||||
@@ -279,7 +283,7 @@ triage_agent = Agent(
|
||||
name="Triage agent",
|
||||
instructions="Handoff to the appropriate agent based on the language of the request.",
|
||||
handoffs=[spanish_agent, english_agent],
|
||||
model="gpt-5.4",
|
||||
model="gpt-5.5",
|
||||
)
|
||||
|
||||
async def main():
|
||||
@@ -287,10 +291,10 @@ async def main():
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
1. OpenAI モデル名を直接設定します。
|
||||
2. [`Model`][agents.models.interface.Model] 実装を提供します。
|
||||
1. OpenAI モデルの名前を直接設定します。
|
||||
2. [`Model`][agents.models.interface.Model] 実装を提供します。
|
||||
|
||||
エージェントで使用するモデルをさらに設定したい場合は、temperature などの任意モデル設定パラメーターを提供する [`ModelSettings`][agents.models.interface.ModelSettings] を渡せます。
|
||||
エージェントで使用するモデルをさらに設定したい場合は、temperature などの任意のモデル設定パラメーターを提供する [`ModelSettings`][agents.models.interface.ModelSettings] を渡すことができます。
|
||||
|
||||
```python
|
||||
from agents import Agent, ModelSettings
|
||||
@@ -305,30 +309,32 @@ english_agent = Agent(
|
||||
|
||||
## 高度な OpenAI Responses 設定
|
||||
|
||||
OpenAI Responses path でより細かい制御が必要な場合は、まず `ModelSettings` から始めてください。
|
||||
OpenAI Responses パスを使用していて、より細かく制御する必要がある場合は、`ModelSettings` から始めてください。
|
||||
|
||||
### 一般的な高度 `ModelSettings` オプション
|
||||
### 一般的な高度な `ModelSettings` オプション
|
||||
|
||||
OpenAI Responses API 使用時、いくつかのリクエストフィールドには対応する `ModelSettings` フィールドがすでにあるため、それらには `extra_args` は不要です。
|
||||
OpenAI Responses API を使用している場合、いくつかのリクエストフィールドはすでに直接の `ModelSettings` フィールドを持つため、それらに `extra_args` は必要ありません。
|
||||
|
||||
- `parallel_tool_calls`: 同一ターンでの複数ツール呼び出しを許可または禁止します。
|
||||
- `truncation`: context あふれ時に失敗させる代わりに、Responses API が最も古い会話項目を削除するよう `"auto"` を設定します。
|
||||
- `store`: 生成応答を後で取得できるようサーバー側に保存するかを制御します。これは response ID に依存するフォローアップワークフローや、`store=False` 時にローカル入力へフォールバックが必要になり得るセッション圧縮フローで重要です。
|
||||
- `prompt_cache_retention`: たとえば `"24h"` でキャッシュ済み prompt prefix をより長く保持します。
|
||||
- `response_include`: `web_search_call.action.sources`、`file_search_call.results`、`reasoning.encrypted_content` など、よりリッチな応答 payload を要求します。
|
||||
- `top_logprobs`: 出力テキストの top-token logprobs を要求します。SDK は `message.output_text.logprobs` も自動追加します。
|
||||
- `retry`: モデル呼び出しに runner 管理リトライ設定を opt in します。[Runner 管理リトライ](#runner-managed-retries) を参照してください。
|
||||
- `parallel_tool_calls`: 同じターンで複数のツール呼び出しを許可または禁止します。
|
||||
- `truncation`: コンテキストがあふれる場合に失敗する代わりに、Responses API が最も古い会話項目を削除できるようにするには、`"auto"` を設定します。
|
||||
- `store`: 生成されたレスポンスを後で取得できるようサーバー側に保存するかどうかを制御します。これは、レスポンス ID に依存するフォローアップワークフローや、`store=False` の場合にローカル入力へフォールバックする必要があるセッション圧縮フローで重要です。
|
||||
- `context_management`: `compact_threshold` を使った Responses 圧縮など、サーバー側のコンテキスト処理を設定します。
|
||||
- `prompt_cache_retention`: たとえば `"24h"` を使って、キャッシュされたプロンプトプレフィックスをより長く保持します。
|
||||
- `response_include`: `web_search_call.action.sources`、`file_search_call.results`、`reasoning.encrypted_content` など、よりリッチなレスポンスペイロードをリクエストします。
|
||||
- `top_logprobs`: 出力テキストの上位トークン logprobs をリクエストします。SDK は `message.output_text.logprobs` も自動的に追加します。
|
||||
- `retry`: モデル呼び出しに対する Runner 管理のリトライ設定をオプトインします。[Runner 管理リトライ](#runner-managed-retries)を参照してください。
|
||||
|
||||
```python
|
||||
from agents import Agent, ModelSettings
|
||||
|
||||
research_agent = Agent(
|
||||
name="Research agent",
|
||||
model="gpt-5.4",
|
||||
model="gpt-5.5",
|
||||
model_settings=ModelSettings(
|
||||
parallel_tool_calls=False,
|
||||
truncation="auto",
|
||||
store=True,
|
||||
context_management=[{"type": "compaction", "compact_threshold": 200000}],
|
||||
prompt_cache_retention="24h",
|
||||
response_include=["web_search_call.action.sources"],
|
||||
top_logprobs=5,
|
||||
@@ -336,13 +342,15 @@ research_agent = Agent(
|
||||
)
|
||||
```
|
||||
|
||||
`store=False` を設定すると、Responses API はその応答を後でサーバー側取得できる状態で保持しません。これは stateless またはゼロデータ保持スタイルのフローに有用ですが、通常 response ID を再利用する機能が、代わりにローカル管理状態へ依存することも意味します。たとえば [`OpenAIResponsesCompactionSession`][agents.memory.openai_responses_compaction_session.OpenAIResponsesCompactionSession] は、最後の応答が保存されていない場合、既定 `"auto"` 圧縮経路を input ベース圧縮へ切り替えます。[Sessions ガイド](../sessions/index.md#openai-responses-compaction-sessions) を参照してください。
|
||||
`store=False` を設定すると、Responses API はそのレスポンスを後でサーバー側で取得できる状態に保持しません。これは、ステートレスまたはゼロデータ保持スタイルのフローに有用ですが、レスポンス ID を再利用する機能が、代わりにローカル管理の状態に依存する必要があることも意味します。たとえば、[`OpenAIResponsesCompactionSession`][agents.memory.openai_responses_compaction_session.OpenAIResponsesCompactionSession] は、最後のレスポンスが保存されていない場合、デフォルトの `"auto"` 圧縮パスを入力ベースの圧縮に切り替えます。[セッションガイド](../sessions/index.md#openai-responses-compaction-sessions)を参照してください。
|
||||
|
||||
### `extra_args` の受け渡し
|
||||
サーバー側圧縮は [`OpenAIResponsesCompactionSession`][agents.memory.openai_responses_compaction_session.OpenAIResponsesCompactionSession] とは異なります。`context_management=[{"type": "compaction", "compact_threshold": ...}]` は各 Responses API リクエストとともに送信され、レンダリングされたコンテキストがしきい値を超えると、API はレスポンスの一部として圧縮項目を出力できます。`OpenAIResponsesCompactionSession` はターン間でスタンドアロンの `responses.compact` エンドポイントを呼び出し、ローカルのセッション履歴を書き換えます。
|
||||
|
||||
SDK がまだトップレベルで直接公開していない provider 固有または新しいリクエストフィールドが必要な場合は `extra_args` を使います。
|
||||
### `extra_args` の渡し方
|
||||
|
||||
また OpenAI の Responses API 使用時は、[他にもいくつか任意パラメーター](https://platform.openai.com/docs/api-reference/responses/create)(例: `user`、`service_tier` など)があります。トップレベルで利用できない場合は、`extra_args` で渡せます。
|
||||
SDK がまだトップレベルで直接公開していない、プロバイダー固有または新しいリクエストフィールドが必要な場合は、`extra_args` を使用します。
|
||||
|
||||
また、OpenAI の Responses API を使用する場合、[他にもいくつかの任意パラメーターがあります](https://platform.openai.com/docs/api-reference/responses/create)(例: `user`、`service_tier` など)。トップレベルで利用できない場合は、それらも `extra_args` で渡せます。同じリクエストフィールドを直接の `ModelSettings` フィールドでも設定しないでください。
|
||||
|
||||
```python
|
||||
from agents import Agent, ModelSettings
|
||||
@@ -360,14 +368,14 @@ english_agent = Agent(
|
||||
|
||||
## Runner 管理リトライ
|
||||
|
||||
リトライは実行時専用で opt in です。`ModelSettings(retry=...)` を設定し、かつリトライポリシーがリトライを選択しない限り、SDK は一般的なモデルリクエストをリトライしません。
|
||||
リトライは実行時専用で、オプトインです。`ModelSettings(retry=...)` を設定し、リトライポリシーがリトライを選択しない限り、SDK は一般的なモデルリクエストをリトライしません。
|
||||
|
||||
```python
|
||||
from agents import Agent, ModelRetrySettings, ModelSettings, retry_policies
|
||||
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
model="gpt-5.4",
|
||||
model="gpt-5.5",
|
||||
model_settings=ModelSettings(
|
||||
retry=ModelRetrySettings(
|
||||
max_retries=4,
|
||||
@@ -392,81 +400,116 @@ agent = Agent(
|
||||
|
||||
<div class="field-table" markdown="1">
|
||||
|
||||
| Field | Type | Notes |
|
||||
| フィールド | 型 | 注記 |
|
||||
| --- | --- | --- |
|
||||
| `max_retries` | `int | None` | 初回リクエスト後に許可されるリトライ試行回数。 |
|
||||
| `backoff` | `ModelRetryBackoffSettings | dict | None` | ポリシーが明示遅延を返さずリトライする場合の既定遅延戦略。 |
|
||||
| `policy` | `RetryPolicy | None` | リトライするか決定するコールバック。このフィールドは実行時専用で serialize されません。 |
|
||||
| `backoff` | `ModelRetryBackoffSettings | dict | None` | ポリシーが明示的な遅延を返さずにリトライする場合のデフォルト遅延戦略。`backoff.max_delay` は、この計算されたバックオフ遅延だけに上限を設けます。ポリシーが返す明示的な遅延や retry-after ヒントには上限を設けません。 |
|
||||
| `policy` | `RetryPolicy | None` | リトライするかどうかを決定するコールバック。このフィールドは実行時専用で、シリアライズされません。 |
|
||||
|
||||
</div>
|
||||
|
||||
リトライポリシーは [`RetryPolicyContext`][agents.retry.RetryPolicyContext] を受け取ります。内容:
|
||||
リトライポリシーは、次を持つ [`RetryPolicyContext`][agents.retry.RetryPolicyContext] を受け取ります:
|
||||
|
||||
- 試行回数依存の判断に使える `attempt` と `max_retries`。
|
||||
- ストリーミング / 非ストリーミング動作を分岐できる `stream`。
|
||||
- raw 検査用の `error`。
|
||||
- `status_code`、`retry_after`、`error_code`、`is_network_error`、`is_timeout`、`is_abort` など正規化情報の `normalized`。
|
||||
- 基盤モデル adapter がリトライ指針を提供できる場合の `provider_advice`。
|
||||
- `attempt` と `max_retries`: 試行回数を考慮した判断に使用できます。
|
||||
- `stream`: ストリーミングと非ストリーミングの挙動を分岐できます。
|
||||
- `error`: raw な確認に使用します。
|
||||
- `normalized`: `status_code`、`retry_after`、`error_code`、`is_network_error`、`is_timeout`、`is_abort` などの正規化済み情報です。
|
||||
- `provider_advice`: 下位のモデルアダプターがリトライのガイダンスを提供できる場合に使用します。
|
||||
|
||||
ポリシーは次のいずれかを返せます:
|
||||
|
||||
- 単純なリトライ判定の `True` / `False`。
|
||||
- 遅延上書きや診断理由付与が必要な場合の [`RetryDecision`][agents.retry.RetryDecision]。
|
||||
- シンプルなリトライ判断として `True` / `False`。
|
||||
- 遅延を上書きしたり診断理由を付加したりしたい場合の [`RetryDecision`][agents.retry.RetryDecision]。
|
||||
|
||||
SDK は `retry_policies` に既製ヘルパーを公開しています:
|
||||
SDK は、`retry_policies` でそのまま使えるヘルパーをエクスポートしています:
|
||||
|
||||
| Helper | Behavior |
|
||||
| ヘルパー | 動作 |
|
||||
| --- | --- |
|
||||
| `retry_policies.never()` | 常に opt out します。 |
|
||||
| `retry_policies.provider_suggested()` | 利用可能な場合 provider のリトライ助言に従います。 |
|
||||
| `retry_policies.network_error()` | 一時的な transport / timeout 失敗に一致します。 |
|
||||
| `retry_policies.http_status([...])` | 選択した HTTP status code に一致します。 |
|
||||
| `retry_policies.retry_after()` | retry-after ヒントがある場合のみ、その遅延でリトライします。 |
|
||||
| `retry_policies.any(...)` | ネストした任意ポリシーが opt in したときにリトライします。 |
|
||||
| `retry_policies.all(...)` | ネストしたすべてのポリシーが opt in したときのみリトライします。 |
|
||||
| `retry_policies.never()` | 常にオプトアウトします。 |
|
||||
| `retry_policies.provider_suggested()` | 利用可能な場合、プロバイダーのリトライ助言に従います。 |
|
||||
| `retry_policies.network_error()` | 一時的なトランスポート障害とタイムアウト障害に一致します。 |
|
||||
| `retry_policies.http_status([...])` | 選択した HTTP ステータスコードに一致します。 |
|
||||
| `retry_policies.retry_after()` | retry-after ヒントが利用可能な場合にのみ、その遅延を使ってリトライします。このヘルパーは retry-after 値を明示的なポリシー遅延として扱うため、`backoff.max_delay` は上限を設けません。 |
|
||||
| `retry_policies.any(...)` | ネストされたポリシーのいずれかがオプトインした場合にリトライします。 |
|
||||
| `retry_policies.all(...)` | ネストされたすべてのポリシーがオプトインした場合にのみリトライします。 |
|
||||
|
||||
ポリシーを合成する場合、`provider_suggested()` は provider veto と replay-safe 承認を維持できるため、最も安全な最初の構成要素です。
|
||||
ポリシーを合成する場合、`provider_suggested()` は最初の構成要素として最も安全です。プロバイダーが区別できる場合に、プロバイダーの拒否と再実行安全性の承認を保持するためです。
|
||||
|
||||
##### 安全境界
|
||||
##### 安全性の境界
|
||||
|
||||
一部失敗は自動リトライされません:
|
||||
一部の失敗は自動では決してリトライされません:
|
||||
|
||||
- Abort エラー。
|
||||
- provider 助言が replay を unsafe と判定したリクエスト。
|
||||
- 出力開始後で replay が unsafe になるストリーミング実行。
|
||||
- 中止エラー。
|
||||
- プロバイダーの助言によって再実行が安全でないと示されたリクエスト。
|
||||
- 出力がすでに開始されており、再実行が安全でなくなる形のストリーミング実行。
|
||||
|
||||
`previous_response_id` または `conversation_id` を使う stateful なフォローアップリクエストも、より保守的に扱われます。これらのリクエストでは、`network_error()` や `http_status([500])` のような non-provider 条件だけでは不十分です。リトライポリシーには通常 `retry_policies.provider_suggested()` を通じた provider の replay-safe 承認を含めるべきです。
|
||||
`previous_response_id` または `conversation_id` を使用するステートフルなフォローアップリクエストも、より保守的に扱われます。これらのリクエストでは、`network_error()` や `http_status([500])` などのプロバイダー由来ではない条件だけでは不十分です。リトライポリシーには、通常 `retry_policies.provider_suggested()` を通じた、プロバイダーからの再実行安全性の承認を含める必要があります。
|
||||
|
||||
##### Runner とエージェントのマージ動作
|
||||
|
||||
`retry` は runner レベルとエージェントレベルの `ModelSettings` 間で deep-merge されます:
|
||||
`retry` は、Runner レベルとエージェントレベルの `ModelSettings` の間でディープマージされます:
|
||||
|
||||
- エージェントは `retry.max_retries` のみ上書きし、runner の `policy` を継承できます。
|
||||
- エージェントは `retry.backoff` の一部のみ上書きし、兄弟 backoff フィールドを runner から維持できます。
|
||||
- `policy` は実行時専用のため、serialize された `ModelSettings` は `max_retries` と `backoff` を保持し、コールバック自体は省略します。
|
||||
- エージェントは `retry.max_retries` だけを上書きし、Runner の `policy` を引き続き継承できます。
|
||||
- エージェントは `retry.backoff` の一部だけを上書きし、同階層のバックオフフィールドを Runner から維持できます。
|
||||
- `policy` は実行時専用であるため、シリアライズされた `ModelSettings` では `max_retries` と `backoff` は保持されますが、コールバック自体は省略されます。
|
||||
|
||||
より完全な例は [`examples/basic/retry.py`](https://github.com/openai/openai-agents-python/tree/main/examples/basic/retry.py) と [adapter-backed retry 例](https://github.com/openai/openai-agents-python/tree/main/examples/basic/retry_litellm.py) を参照してください。
|
||||
より詳しいコード例については、[`examples/basic/retry.py`](https://github.com/openai/openai-agents-python/tree/main/examples/basic/retry.py) と [アダプターを利用したリトライ例](https://github.com/openai/openai-agents-python/tree/main/examples/basic/retry_litellm.py)を参照してください。
|
||||
|
||||
## non-OpenAI provider のトラブルシューティング
|
||||
## 非 OpenAI プロバイダーのトラブルシューティング
|
||||
|
||||
### トレーシングクライアントエラー 401
|
||||
|
||||
トレーシング関連エラーが出る場合、trace は OpenAI サーバーへアップロードされるため、OpenAI API key がないことが原因です。解決方法は 3 つあります:
|
||||
トレーシングに関連するエラーが発生する場合、これはトレースが OpenAI サーバーにアップロードされる一方で、OpenAI API キーがないことが原因です。これを解決するには、3 つの選択肢があります:
|
||||
|
||||
1. トレーシングを完全に無効化: [`set_tracing_disabled(True)`][agents.set_tracing_disabled]。
|
||||
2. トレーシング用 OpenAI key を設定: [`set_tracing_export_api_key(...)`][agents.set_tracing_export_api_key]。この API key は trace アップロード専用で、[platform.openai.com](https://platform.openai.com/) 由来である必要があります。
|
||||
3. non-OpenAI の trace プロセッサーを使用。詳細は [tracing docs](../tracing.md#custom-tracing-processors) を参照してください。
|
||||
1. トレーシングを完全に無効化する: [`set_tracing_disabled(True)`][agents.set_tracing_disabled]。
|
||||
2. トレーシング用の OpenAI キーを設定する: [`set_tracing_export_api_key(...)`][agents.set_tracing_export_api_key]。この API キーはトレースのアップロードにのみ使用され、[platform.openai.com](https://platform.openai.com/) のものである必要があります。
|
||||
3. 非 OpenAI のトレースプロセッサーを使用する。[トレーシングドキュメント](../tracing.md#custom-tracing-processors)を参照してください。
|
||||
|
||||
### Responses API サポート
|
||||
### Responses API のサポート
|
||||
|
||||
SDK は既定で Responses API を使いますが、他の多くの LLM provider はまだサポートしていません。その結果 404 などの問題が発生することがあります。解決方法は 2 つあります:
|
||||
SDK はデフォルトで Responses API を使用しますが、他の多くの LLM プロバイダーはまだサポートしていません。その結果として 404 エラーや類似の問題が発生する場合があります。解決するには、2 つの選択肢があります:
|
||||
|
||||
1. [`set_default_openai_api("chat_completions")`][agents.set_default_openai_api] を呼び出す。これは環境変数で `OPENAI_API_KEY` と `OPENAI_BASE_URL` を設定している場合に機能します。
|
||||
2. [`OpenAIChatCompletionsModel`][agents.models.openai_chatcompletions.OpenAIChatCompletionsModel] を使う。例は [こちら](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/) にあります。
|
||||
1. [`set_default_openai_api("chat_completions")`][agents.set_default_openai_api] を呼び出す。これは、環境変数で `OPENAI_API_KEY` と `OPENAI_BASE_URL` を設定している場合に機能します。
|
||||
2. [`OpenAIChatCompletionsModel`][agents.models.openai_chatcompletions.OpenAIChatCompletionsModel] を使用する。コード例は [こちら](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/)にあります。
|
||||
|
||||
### structured outputs サポート
|
||||
### Chat Completions 互換性オプション
|
||||
|
||||
一部モデル provider は [structured outputs](https://platform.openai.com/docs/guides/structured-outputs) をサポートしていません。これにより、次のようなエラーが発生することがあります:
|
||||
Chat Completions 経由でルーティングする場合、SDK は Chat Completions が送信できない Responses 専用フィールド(`previous_response_id`、`conversation_id`、プロンプト、テキストのみではないツール出力など)を黙って削除することで互換性を維持します。開発中にこれらの不一致を早期に失敗させたい場合は、OpenAI プロバイダーで厳格な機能検証を有効にしてください:
|
||||
|
||||
```python
|
||||
from agents import Agent, OpenAIProvider, RunConfig, Runner
|
||||
|
||||
provider = OpenAIProvider(
|
||||
use_responses=False,
|
||||
strict_feature_validation=True,
|
||||
)
|
||||
|
||||
agent = Agent(name="Assistant")
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Hello",
|
||||
run_config=RunConfig(model_provider=provider),
|
||||
)
|
||||
```
|
||||
|
||||
[`MultiProvider`][agents.MultiProvider] を使用する場合は、代わりに `openai_strict_feature_validation=True` を渡します。
|
||||
|
||||
一部の OpenAI 互換 Chat Completions プロバイダーは、ツール呼び出しの差分をチャンクでストリーミングしますが、SDK のインクリメンタル処理に十分な信頼性がない場合があります。その場合は、ストリーミングされたツール呼び出しのバッファリングを有効にし、プロバイダーのストリームが終了した後にのみ SDK がツール呼び出しを出力するようにします:
|
||||
|
||||
```python
|
||||
from agents import OpenAIProvider
|
||||
|
||||
provider = OpenAIProvider(
|
||||
use_responses=False,
|
||||
buffer_streamed_tool_calls=True,
|
||||
)
|
||||
```
|
||||
|
||||
[`MultiProvider`][agents.MultiProvider] では、`openai_buffer_streamed_tool_calls=True` を使用します。
|
||||
|
||||
### structured outputs のサポート
|
||||
|
||||
一部のモデルプロバイダーは [structured outputs](https://platform.openai.com/docs/guides/structured-outputs) に対応していません。これにより、次のようなエラーが発生することがあります:
|
||||
|
||||
```
|
||||
|
||||
@@ -474,34 +517,34 @@ BadRequestError: Error code: 400 - {'error': {'message': "'response_format.type'
|
||||
|
||||
```
|
||||
|
||||
これは一部モデル provider の制限です。JSON 出力はサポートしていても、出力に使う `json_schema` 指定を許可しません。この問題の修正を進めていますが、JSON schema 出力をサポートする provider への依存を推奨します。そうでない場合、不正 JSON によりアプリが頻繁に壊れる可能性があります。
|
||||
これは一部のモデルプロバイダーの制約です。JSON 出力はサポートしていても、出力に使用する `json_schema` を指定できないためです。現在この修正に取り組んでいますが、JSON スキーマ出力をサポートしているプロバイダーに依存することをお勧めします。そうしないと、不正な JSON によってアプリが頻繁に失敗する可能性があります。
|
||||
|
||||
## provider 間でのモデル混在
|
||||
## プロバイダー間でのモデルの混在
|
||||
|
||||
モデル provider 間の機能差を理解していないと、エラーに遭遇する可能性があります。たとえば OpenAI は structured outputs、マルチモーダル入力、ホスト型ファイル検索と Web 検索をサポートしますが、多くの他 provider はこれらをサポートしません。次の制約に注意してください:
|
||||
モデルプロバイダー間の機能差異を把握しておく必要があります。そうしないと、エラーが発生する可能性があります。たとえば、OpenAI は structured outputs、マルチモーダル入力、ホスト型のファイル検索と Web 検索をサポートしていますが、他の多くのプロバイダーはこれらの機能をサポートしていません。次の制限に注意してください:
|
||||
|
||||
- サポートしない provider に未対応の `tools` を送らない
|
||||
- テキスト専用モデル呼び出し前にマルチモーダル入力を除外する
|
||||
- structured JSON 出力非対応 provider は無効 JSON を時折生成する点を認識する
|
||||
- サポートされていない `tools` を、それらを理解しないプロバイダーに送信しないでください
|
||||
- テキストのみのモデルを呼び出す前に、マルチモーダル入力を除外してください
|
||||
- 構造化 JSON 出力をサポートしていないプロバイダーは、ときどき無効な JSON を生成することに注意してください。
|
||||
|
||||
## サードパーティ adapters
|
||||
## サードパーティアダプター
|
||||
|
||||
SDK の組み込み provider 統合ポイントで不十分な場合にのみ、サードパーティ adapter を使用してください。この SDK で OpenAI モデルのみを使う場合、Any-LLM や LiteLLM ではなく、組み込み [`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel] 経路を優先してください。サードパーティ adapter は、OpenAI モデルと non-OpenAI provider の組み合わせ、または組み込み経路で提供されない adapter 管理の provider カバレッジ / ルーティングが必要なケース向けです。adapter は SDK と上流モデル provider の間に別の互換レイヤーを追加するため、機能サポートとリクエスト意味論は provider により変動します。SDK は現在、Any-LLM と LiteLLM を best-effort の beta adapter 統合として含みます。
|
||||
サードパーティアダプターは、SDK に組み込まれたプロバイダー統合ポイントだけでは不十分な場合にのみ使用してください。この SDK で OpenAI モデルのみを使用している場合は、Any-LLM や LiteLLM ではなく、組み込みの [`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel] パスを優先してください。サードパーティアダプターは、OpenAI モデルを非 OpenAI プロバイダーと組み合わせる必要がある場合、または組み込みパスでは提供されないアダプター管理のプロバイダー対応範囲やルーティングが必要な場合向けです。アダプターは SDK と上流のモデルプロバイダーの間に別の互換性レイヤーを追加するため、機能サポートとリクエストセマンティクスはプロバイダーによって異なる場合があります。SDK には現在、ベストエフォートのベータ版アダプター統合として Any-LLM と LiteLLM が含まれています。
|
||||
|
||||
### Any-LLM
|
||||
|
||||
Any-LLM サポートは、Any-LLM 管理の provider カバレッジまたはルーティングが必要なケース向けに、best-effort な beta として含まれます。
|
||||
Any-LLM のサポートは、Any-LLM が管理するプロバイダー対応範囲またはルーティングが必要な場合向けに、ベストエフォートのベータ版として含まれています。
|
||||
|
||||
上流 provider 経路により、Any-LLM は Responses API、Chat Completions 互換 API、または provider 固有の互換レイヤーを使う場合があります。
|
||||
上流のプロバイダーパスに応じて、Any-LLM は Responses API、Chat Completions 互換 API、またはプロバイダー固有の互換性レイヤーを使用する場合があります。
|
||||
|
||||
Any-LLM が必要な場合は `openai-agents[any-llm]` をインストールし、[`examples/model_providers/any_llm_auto.py`](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/any_llm_auto.py) または [`examples/model_providers/any_llm_provider.py`](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/any_llm_provider.py) から開始してください。[`MultiProvider`][agents.MultiProvider] で `any-llm/...` モデル名を使う、`AnyLLMModel` を直接インスタンス化する、または実行スコープで `AnyLLMProvider` を使うことができます。モデル surface を明示固定したい場合は、`AnyLLMModel` 構築時に `api="responses"` または `api="chat_completions"` を渡します。
|
||||
Any-LLM が必要な場合は、`openai-agents[any-llm]` をインストールし、[`examples/model_providers/any_llm_auto.py`](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/any_llm_auto.py) または [`examples/model_providers/any_llm_provider.py`](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/any_llm_provider.py) から始めます。[`MultiProvider`][agents.MultiProvider] で `any-llm/...` モデル名を使用したり、`AnyLLMModel` を直接インスタンス化したり、実行スコープで `AnyLLMProvider` を使用したりできます。モデルサーフェスを明示的に固定する必要がある場合は、`AnyLLMModel` を構築する際に `api="responses"` または `api="chat_completions"` を渡します。
|
||||
|
||||
Any-LLM はサードパーティ adapter レイヤーであり、provider 依存関係と機能ギャップは SDK ではなく Any-LLM 側で定義されます。使用量メトリクスは上流 provider が返す場合に自動伝搬されますが、ストリーミング Chat Completions backend では usage chunk 出力前に `ModelSettings(include_usage=True)` が必要な場合があります。structured outputs、ツール呼び出し、使用量レポート、Responses 固有動作に依存する場合は、デプロイ予定の正確な provider backend を検証してください。
|
||||
Any-LLM はサードパーティアダプターレイヤーであるため、プロバイダーの依存関係と機能ギャップは SDK ではなく Any-LLM によって上流で定義されます。使用量メトリクスは、上流プロバイダーが返す場合に自動的に伝播されますが、ストリーミングされた Chat Completions バックエンドでは、使用量チャンクを出力する前に `ModelSettings(include_usage=True)` が必要になる場合があります。structured outputs、ツール呼び出し、使用量レポート、または Responses 固有の動作に依存する場合は、デプロイ予定の正確なプロバイダーバックエンドを検証してください。
|
||||
|
||||
### LiteLLM
|
||||
|
||||
LiteLLM サポートは、LiteLLM 固有の provider カバレッジまたはルーティングが必要なケース向けに、best-effort な beta として含まれます。
|
||||
LiteLLM のサポートは、LiteLLM 固有のプロバイダー対応範囲またはルーティングが必要な場合向けに、ベストエフォートのベータ版として含まれています。
|
||||
|
||||
LiteLLM が必要な場合は `openai-agents[litellm]` をインストールし、[`examples/model_providers/litellm_auto.py`](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/litellm_auto.py) または [`examples/model_providers/litellm_provider.py`](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/litellm_provider.py) から開始してください。`litellm/...` モデル名を使用するか、[`LitellmModel`][agents.extensions.models.litellm_model.LitellmModel] を直接インスタンス化できます。
|
||||
LiteLLM が必要な場合は、`openai-agents[litellm]` をインストールし、[`examples/model_providers/litellm_auto.py`](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/litellm_auto.py) または [`examples/model_providers/litellm_provider.py`](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/litellm_provider.py) から始めます。`litellm/...` モデル名を使用したり、[`LitellmModel`][agents.extensions.models.litellm_model.LitellmModel] を直接インスタンス化したりできます。
|
||||
|
||||
一部 LiteLLM バックエンド provider は、既定では SDK 使用量メトリクスを設定しません。使用量レポートが必要な場合は `ModelSettings(include_usage=True)` を渡し、structured outputs、ツール呼び出し、使用量レポート、adapter 固有ルーティング動作に依存する場合は、デプロイ予定の正確な provider backend を検証してください。
|
||||
LiteLLM を利用する一部のプロバイダーは、デフォルトでは SDK の使用量メトリクスを設定しません。使用量レポートが必要な場合は、`ModelSettings(include_usage=True)` を渡し、structured outputs、ツール呼び出し、使用量レポート、またはアダプター固有のルーティング動作に依存する場合は、デプロイ予定の正確なプロバイダーバックエンドを検証してください。
|
||||
@@ -8,6 +8,6 @@ search:
|
||||
window.location.replace("../#third-party-adapters");
|
||||
</script>
|
||||
|
||||
このページは [Models の Third-party adapters セクション](index.md#third-party-adapters)に移動しました。
|
||||
このページは [Models のサードパーティアダプターセクション](index.md#third-party-adapters) に移動しました。
|
||||
|
||||
自動的にリダイレクトされない場合は、上記のリンクを使用してください。
|
||||
+33
-33
@@ -4,61 +4,61 @@ search:
|
||||
---
|
||||
# エージェントオーケストレーション
|
||||
|
||||
オーケストレーションとは、アプリ内でのエージェントの流れを指します。どのエージェントが、どの順序で実行され、次に何が起こるかをどのように決定するか、ということです。エージェントをオーケストレーションする主な方法は 2 つあります。
|
||||
オーケストレーションとは、アプリ内でのエージェントの流れを指します。どのエージェントを、どの順序で実行し、次に何が起こるかをどのように決定するか、ということです。エージェントをオーケストレーションする主な方法は 2 つあります:
|
||||
|
||||
1. LLM に意思決定させる: LLM の知性を使って計画・推論を行い、それに基づいてどのステップを取るかを決定します。
|
||||
2. コードでオーケストレーションする: コードによってエージェントの流れを決定します。
|
||||
1. LLM に判断を任せる:LLM の知能を使って計画と推論を行い、それに基づいて取る手順を決定します。
|
||||
2. コードによるオーケストレーション:コードでエージェントの流れを決定します。
|
||||
|
||||
これらのパターンは組み合わせて使えます。それぞれにトレードオフがあり、以下で説明します。
|
||||
これらのパターンは組み合わせることができます。それぞれにトレードオフがあり、以下で説明します。
|
||||
|
||||
## LLM によるオーケストレーション
|
||||
|
||||
エージェントは、instructions、tools、ハンドオフを備えた LLM です。つまり、オープンエンドなタスクが与えられた場合、LLM はそのタスクへの取り組み方を自律的に計画でき、tools を使ってアクションを実行しデータを取得し、ハンドオフを使ってサブエージェントにタスクを委譲できます。たとえば、リサーチエージェントには次のようなツールを備えられます。
|
||||
エージェントとは、instructions、tools、ハンドオフを備えた LLM です。つまり、自由度の高いタスクが与えられた場合、LLM は、tools を使ってアクションを実行しデータを取得し、ハンドオフを使ってサブエージェントにタスクを委任しながら、そのタスクへの取り組み方を自律的に計画できます。たとえば、リサーチエージェントには次のようなツールを備えられます:
|
||||
|
||||
- オンライン情報を見つけるための Web 検索
|
||||
- オンラインで情報を見つけるための Web 検索
|
||||
- 独自データや接続先を検索するためのファイル検索と取得
|
||||
- コンピュータ上でアクションを実行するためのコンピュータ操作
|
||||
- コンピューター上でアクションを実行するためのコンピュータ操作
|
||||
- データ分析を行うためのコード実行
|
||||
- 計画、レポート作成などに優れた専門エージェントへのハンドオフ
|
||||
- 計画、レポート作成などに優れた専門エージェントへのハンドオフ。
|
||||
|
||||
### SDK の中核パターン
|
||||
### コア SDK パターン
|
||||
|
||||
Python SDK では、次の 2 つのオーケストレーションパターンが最もよく使われます。
|
||||
Python SDK では、次の 2 つのオーケストレーションパターンが最もよく登場します:
|
||||
|
||||
| パターン | 仕組み | 最適な場面 |
|
||||
| パターン | 仕組み | 最適な場合 |
|
||||
| --- | --- | --- |
|
||||
| Agents as tools | マネージャーエージェントが会話の制御を維持し、`Agent.as_tool()` を通じて専門エージェントを呼び出します。 | 1 つのエージェントに最終回答を担わせたい、複数の専門家の出力を統合したい、または共通のガードレールを 1 か所で適用したい場合。 |
|
||||
| ハンドオフ | トリアージエージェントが会話を専門エージェントへ振り分け、その専門エージェントがそのターンの残りでアクティブなエージェントになります。 | 専門エージェントに直接応答させたい、プロンプトを集中させたい、またはマネージャーが結果を説明せずに instructions を切り替えたい場合。 |
|
||||
| Agents as tools | マネージャーエージェントが会話の制御を維持し、`Agent.as_tool()` を通じて専門エージェントを呼び出します。 | 1 つのエージェントに最終回答を担わせたい場合、複数の専門エージェントからの出力を統合したい場合、または共通のガードレールを 1 か所で適用したい場合。 |
|
||||
| ハンドオフ | トリアージエージェントが会話を専門エージェントにルーティングし、その専門エージェントがそのターンの残りでアクティブなエージェントになります。 | 専門エージェントに直接応答させたい場合、プロンプトを焦点の絞られた状態に保ちたい場合、またはマネージャーが結果を説明することなく instructions を切り替えたい場合。 |
|
||||
|
||||
専門エージェントが限定的なサブタスクを支援すべきで、ユーザー向け会話を引き継ぐべきではない場合は **agents as tools** を使います。ルーティング自体がワークフローの一部であり、選ばれた専門エージェントに次のやり取りを担わせたい場合は **handoffs** を使います。
|
||||
専門エージェントが範囲の限定されたサブタスクを支援するべきだが、ユーザー向けの会話を引き継ぐべきではない場合は、 **agents as tools** を使用します。ルーティング自体がワークフローの一部であり、選ばれた専門エージェントにインタラクションの次の部分を担わせたい場合は、 **ハンドオフ** を使用します。
|
||||
|
||||
2 つを組み合わせることもできます。トリアージエージェントが専門エージェントにハンドオフし、その専門エージェントがさらに限定的なサブタスクのために他のエージェントをツールとして呼び出すことも可能です。
|
||||
この 2 つを組み合わせることもできます。トリアージエージェントが専門エージェントにハンドオフし、その専門エージェントがさらに狭いサブタスクのために他のエージェントをツールとして呼び出すこともできます。
|
||||
|
||||
このパターンは、タスクがオープンエンドで、LLM の知性に依存したい場合に非常に有効です。ここで最も重要な戦術は次のとおりです。
|
||||
このパターンは、タスクの自由度が高く、LLM の知能に頼りたい場合に最適です。ここで最も重要な戦術は次のとおりです:
|
||||
|
||||
1. 良いプロンプトに投資する。どのツールが利用可能か、どう使うか、どのパラメーター範囲内で動作すべきかを明確にします。
|
||||
2. アプリを監視し、反復改善する。どこで問題が起こるかを確認し、プロンプトを改善します。
|
||||
3. エージェントに内省と改善を許可する。たとえば、ループで実行して自己批評させる、またはエラーメッセージを与えて改善させます。
|
||||
4. どんなタスクにも対応する汎用エージェントを期待するより、1 つのタスクに優れた専門エージェントを用意します。
|
||||
5. [evals](https://platform.openai.com/docs/guides/evals) に投資する。これによりエージェントを改善するための訓練ができ、タスク性能を向上させられます。
|
||||
1. 優れたプロンプトに投資します。利用できるツール、その使い方、そしてエージェントが従うべきパラメーターを明確にします。
|
||||
2. アプリを監視し、反復改善します。どこで問題が起こるかを確認し、プロンプトを改善します。
|
||||
3. エージェントが内省して改善できるようにします。たとえば、ループ内で実行して自己批評させる、またはエラーメッセージを提供して改善させます。
|
||||
4. 何でも得意であることを期待される汎用エージェントではなく、1 つのタスクに秀でた専門エージェントを用意します。
|
||||
5. [evals](https://platform.openai.com/docs/guides/evals) に投資します。これにより、エージェントをトレーニングして改善し、タスクの遂行能力を高めることができます。
|
||||
|
||||
このスタイルのオーケストレーションを支える SDK の基本コンポーネントを確認したい場合は、[tools](tools.md)、[handoffs](handoffs.md)、[running agents](running_agents.md) から始めてください。
|
||||
このスタイルのオーケストレーションを支えるコア SDK の基本コンポーネントを知りたい場合は、[ツール](tools.md)、[ハンドオフ](handoffs.md)、[エージェントの実行](running_agents.md) から始めてください。
|
||||
|
||||
## コードによるオーケストレーション
|
||||
|
||||
LLM によるオーケストレーションは強力ですが、コードによるオーケストレーションは、速度・コスト・性能の面でタスクをより決定的で予測可能にします。ここで一般的なパターンは次のとおりです。
|
||||
LLM によるオーケストレーションは強力ですが、コードによるオーケストレーションは、速度、コスト、パフォーマンスの観点でタスクをより決定論的で予測可能にします。ここでの一般的なパターンは次のとおりです:
|
||||
|
||||
- [structured outputs](https://platform.openai.com/docs/guides/structured-outputs) を使い、コードで検査可能な適切な形式のデータを生成する。たとえば、タスクをいくつかのカテゴリーに分類するようエージェントに求め、そのカテゴリーに基づいて次のエージェントを選択できます。
|
||||
- 1 つの出力を次の入力に変換して複数エージェントを連結する。ブログ記事執筆のようなタスクを、リサーチ、アウトライン作成、記事執筆、批評、改善という一連のステップに分解できます。
|
||||
- 評価とフィードバックを行うエージェントと組み合わせて、タスク実行エージェントを `while` ループで実行し、評価側が出力が特定の基準を満たしたと言うまで続ける。
|
||||
- 複数エージェントを並列実行する。たとえば `asyncio.gather` のような Python の基本機能を使います。これは、相互依存しない複数タスクがある場合の高速化に有用です。
|
||||
- [structured outputs](https://platform.openai.com/docs/guides/structured-outputs) を使って、コードで検査できる適切な形式のデータを生成します。たとえば、エージェントにタスクをいくつかのカテゴリーに分類させ、そのカテゴリーに基づいて次のエージェントを選択できます。
|
||||
- 複数のエージェントをチェーンし、あるエージェントの出力を次のエージェントの入力に変換します。ブログ記事を書くようなタスクを一連のステップに分解できます - リサーチする、アウトラインを書く、ブログ記事を書く、批評し、それから改善します。
|
||||
- タスクを実行するエージェントを `while` ループ内で、評価してフィードバックを提供するエージェントと一緒に実行し、評価者が出力が特定の基準を満たしたと言うまで続けます。
|
||||
- 複数のエージェントを並列に実行します。たとえば、`asyncio.gather` のような Python の基本コンポーネントを使います。これは、互いに依存しない複数のタスクがある場合に高速化に役立ちます。
|
||||
|
||||
[`examples/agent_patterns`](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns) に多数のコード例があります。
|
||||
[`examples/agent_patterns`](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns) には多数のコード例があります。
|
||||
|
||||
## 関連ガイド
|
||||
|
||||
- 構成パターンとエージェント設定については [Agents](agents.md)。
|
||||
- `Agent.as_tool()` とマネージャースタイルのオーケストレーションについては [Tools](tools.md#agents-as-tools)。
|
||||
- 専門エージェント間の委譲については [Handoffs](handoffs.md)。
|
||||
- 実行ごとのオーケストレーション制御と会話状態については [Running agents](running_agents.md)。
|
||||
- 最小のエンドツーエンドなハンドオフ例については [Quickstart](quickstart.md)。
|
||||
- 構成パターンとエージェント設定については、[エージェント](agents.md) を参照してください。
|
||||
- `Agent.as_tool()` とマネージャースタイルのオーケストレーションについては、[ツール](tools.md#agents-as-tools) を参照してください。
|
||||
- 専門エージェント間の委任については、[ハンドオフ](handoffs.md) を参照してください。
|
||||
- 実行ごとのオーケストレーション制御と会話状態については、[エージェントの実行](running_agents.md) を参照してください。
|
||||
- 最小限のエンドツーエンドのハンドオフ例については、[クイックスタート](quickstart.md) を参照してください。
|
||||
+55
-31
@@ -6,7 +6,7 @@ search:
|
||||
|
||||
## プロジェクトと仮想環境の作成
|
||||
|
||||
これは一度だけ実行すれば十分です。
|
||||
これは一度だけ行えば十分です。
|
||||
|
||||
```bash
|
||||
mkdir my_project
|
||||
@@ -16,12 +16,20 @@ python -m venv .venv
|
||||
|
||||
### 仮想環境の有効化
|
||||
|
||||
新しいターミナルセッションを開始するたびに実行してください。
|
||||
新しいターミナルセッションを開始するたびに行ってください。
|
||||
|
||||
macOS または Linux の場合:
|
||||
|
||||
```bash
|
||||
source .venv/bin/activate
|
||||
```
|
||||
|
||||
Windows の場合:
|
||||
|
||||
```cmd
|
||||
.venv\Scripts\activate
|
||||
```
|
||||
|
||||
### Agents SDK のインストール
|
||||
|
||||
```bash
|
||||
@@ -30,15 +38,31 @@ pip install openai-agents # or `uv add openai-agents`, etc
|
||||
|
||||
### OpenAI API キーの設定
|
||||
|
||||
まだお持ちでない場合は、OpenAI API キーを作成するために [こちらの手順](https://platform.openai.com/docs/quickstart#create-and-export-an-api-key) に従ってください。
|
||||
まだ持っていない場合は、[こちらの手順](https://platform.openai.com/docs/quickstart#create-and-export-an-api-key)に従って OpenAI API キーを作成してください。
|
||||
|
||||
以下のコマンドは、現在のターミナルセッションにキーを設定します。
|
||||
|
||||
macOS または Linux の場合:
|
||||
|
||||
```bash
|
||||
export OPENAI_API_KEY=sk-...
|
||||
```
|
||||
|
||||
Windows PowerShell の場合:
|
||||
|
||||
```powershell
|
||||
$env:OPENAI_API_KEY = "sk-..."
|
||||
```
|
||||
|
||||
Windows コマンドプロンプトの場合:
|
||||
|
||||
```cmd
|
||||
set "OPENAI_API_KEY=sk-..."
|
||||
```
|
||||
|
||||
## 最初のエージェントの作成
|
||||
|
||||
エージェントは instructions、名前、および特定のモデルなどの任意の設定で定義します。
|
||||
エージェントは、instructions、名前、および特定のモデルなどの任意の設定で定義します。
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
@@ -70,23 +94,23 @@ if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
2 回目のターンでは、`result.to_input_list()` を `Runner.run(...)` に戻して渡すか、[session](sessions/index.md) をアタッチするか、`conversation_id` / `previous_response_id` で OpenAI のサーバー管理状態を再利用できます。[running agents](running_agents.md) ガイドでは、これらのアプローチを比較しています。
|
||||
2 回目のターンでは、`result.to_input_list()` を `Runner.run(...)` に戻して渡すか、[セッション](sessions/index.md)をアタッチするか、`conversation_id` / `previous_response_id` で OpenAI によりサーバー側で管理される状態を再利用できます。[エージェントの実行](running_agents.md)ガイドでは、これらのアプローチを比較しています。
|
||||
|
||||
次の目安を使ってください。
|
||||
次の目安を使ってください:
|
||||
|
||||
| 望んでいること | まず使うもの |
|
||||
| 実現したいこと... | まず使うもの... |
|
||||
| --- | --- |
|
||||
| 完全な手動制御とプロバイダー非依存の履歴 | `result.to_input_list()` |
|
||||
| SDK に履歴の読み込みと保存を任せる | [`session=...`](sessions/index.md) |
|
||||
| OpenAI 管理のサーバー側継続 | `previous_response_id` または `conversation_id` |
|
||||
|
||||
トレードオフと正確な動作については、[Running agents](running_agents.md#choose-a-memory-strategy) を参照してください。
|
||||
トレードオフと正確な挙動については、[エージェントの実行](running_agents.md#choose-a-memory-strategy)を参照してください。
|
||||
|
||||
タスクが主にプロンプト、ツール、会話状態で完結する場合は、プレーンな `Agent` と `Runner` を使用してください。エージェントが分離されたワークスペース内の実ファイルを検査または変更する必要がある場合は、[Sandbox agents quickstart](sandbox_agents.md) に進んでください。
|
||||
タスクが主にプロンプト、ツール、会話状態で完結する場合は、シンプルな `Agent` と `Runner` を使います。エージェントが分離されたワークスペース内の実ファイルを検査または変更する必要がある場合は、[Sandbox エージェントのクイックスタート](sandbox_agents.md)に進んでください。
|
||||
|
||||
## エージェントへのツール付与
|
||||
## エージェントへのツールの付与
|
||||
|
||||
エージェントに、情報を調べたりアクションを実行したりするためのツールを与えることができます。
|
||||
エージェントにツールを与えることで、情報を調べたりアクションを実行したりできます。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -118,16 +142,16 @@ if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
## 追加エージェント
|
||||
## さらにいくつかのエージェントの追加
|
||||
|
||||
マルチエージェントパターンを選ぶ前に、最終回答を誰が担当するかを決めてください。
|
||||
マルチエージェントパターンを選ぶ前に、最終回答の主導権を誰が持つべきかを決めてください。
|
||||
|
||||
- **ハンドオフ**: そのターンの該当部分では、専門エージェントが会話を引き継ぎます。
|
||||
- **Agents as tools**: オーケストレーターが制御を維持し、専門エージェントをツールとして呼び出します。
|
||||
- **ハンドオフ**: スペシャリストが、そのターンの該当部分について会話を引き継ぎます。
|
||||
- **Agents as tools**: オーケストレーターが制御を維持し、スペシャリストをツールとして呼び出します。
|
||||
|
||||
このクイックスタートでは、最初の例として最短であるため **ハンドオフ** を続けて扱います。マネージャースタイルのパターンについては、[Agent orchestration](multi_agent.md) と [Tools: agents as tools](tools.md#agents-as-tools) を参照してください。
|
||||
このクイックスタートでは、最初の例として最も短いため、 **ハンドオフ** で続けます。マネージャースタイルのパターンについては、[エージェントオーケストレーション](multi_agent.md)と[ツール: agents as tools](tools.md#agents-as-tools)を参照してください。
|
||||
|
||||
追加のエージェントも同じ方法で定義できます。`handoff_description` は、いつ委譲するかについてルーティングエージェントに追加コンテキストを与えます。
|
||||
追加のエージェントも同じ方法で定義できます。`handoff_description` は、いつ委譲すべきかについて、ルーティングエージェントに追加のコンテキストを提供します。
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
@@ -147,7 +171,7 @@ math_tutor_agent = Agent(
|
||||
|
||||
## ハンドオフの定義
|
||||
|
||||
エージェントでは、タスク解決中に選択可能な送信先ハンドオフオプションの一覧を定義できます。
|
||||
エージェントには、タスクを解決する際に選択できるハンドオフ先の選択肢の一覧を定義できます。
|
||||
|
||||
```python
|
||||
triage_agent = Agent(
|
||||
@@ -159,7 +183,7 @@ triage_agent = Agent(
|
||||
|
||||
## エージェントオーケストレーションの実行
|
||||
|
||||
ランナーは、個々のエージェント実行、ハンドオフ、ツール呼び出しを処理します。
|
||||
ランナーは、個々のエージェントの実行、すべてのハンドオフ、すべてのツール呼び出しを処理します。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -179,23 +203,23 @@ if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
## 参照コード例
|
||||
## 参考コード例
|
||||
|
||||
リポジトリには、同じ主要パターンの完全なスクリプトが含まれています。
|
||||
このリポジトリには、同じ主要パターンに対応する完全なスクリプトが含まれています:
|
||||
|
||||
- 最初の実行向け: [`examples/basic/hello_world.py`](https://github.com/openai/openai-agents-python/tree/main/examples/basic/hello_world.py)
|
||||
- 関数ツール向け: [`examples/basic/tools.py`](https://github.com/openai/openai-agents-python/tree/main/examples/basic/tools.py)
|
||||
- マルチエージェントルーティング向け: [`examples/agent_patterns/routing.py`](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns/routing.py)
|
||||
- [`examples/basic/hello_world.py`](https://github.com/openai/openai-agents-python/tree/main/examples/basic/hello_world.py) は最初の実行の例です。
|
||||
- [`examples/basic/tools.py`](https://github.com/openai/openai-agents-python/tree/main/examples/basic/tools.py) は関数ツールの例です。
|
||||
- [`examples/agent_patterns/routing.py`](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns/routing.py) はマルチエージェントルーティングの例です。
|
||||
|
||||
## トレースの確認
|
||||
## トレースの表示
|
||||
|
||||
エージェント実行中に何が起きたかを確認するには、[OpenAI ダッシュボードの Trace viewer](https://platform.openai.com/traces) に移動して、エージェント実行のトレースを表示してください。
|
||||
エージェントの実行中に何が起きたかを確認するには、[OpenAI ダッシュボードのトレースビューアー](https://platform.openai.com/traces)に移動して、エージェント実行のトレースを表示してください。
|
||||
|
||||
## 次のステップ
|
||||
|
||||
より複雑なエージェントフローの構築方法を学びます。
|
||||
より複雑なエージェント型フローの構築方法を学びましょう:
|
||||
|
||||
- [Agents](agents.md) の設定方法を学ぶ。
|
||||
- [running agents](running_agents.md) と [sessions](sessions/index.md) を学ぶ。
|
||||
- 作業を実際のワークスペース内で行うべき場合は [Sandbox agents](sandbox_agents.md) を学ぶ。
|
||||
- [tools](tools.md)、[guardrails](guardrails.md)、[models](models/index.md) を学ぶ。
|
||||
- [エージェント](agents.md)の設定方法について学びます。
|
||||
- [エージェントの実行](running_agents.md)と[セッション](sessions/index.md)について学びます。
|
||||
- 作業を実際のワークスペース内で行う必要がある場合は、[Sandbox エージェント](sandbox_agents.md)について学びます。
|
||||
- [ツール](tools.md)、[ガードレール](guardrails.md)、[モデル](models/index.md)について学びます。
|
||||
+76
-71
@@ -4,59 +4,55 @@ search:
|
||||
---
|
||||
# Realtime エージェントガイド
|
||||
|
||||
このガイドでは、 OpenAI Agents SDK の realtime レイヤーが OpenAI Realtime API にどのように対応しているか、そして Python SDK がその上にどのような追加動作を加えるかを説明します。
|
||||
このガイドでは、 OpenAI Agents SDK の realtime レイヤーが OpenAI Realtime API にどのように対応するか、および Python SDK がその上に追加する挙動について説明します。
|
||||
|
||||
!!! warning "Beta 機能"
|
||||
!!! note "はじめに"
|
||||
|
||||
Realtime エージェントは beta 段階です。実装の改善に伴い、破壊的変更が入る可能性があります。
|
||||
|
||||
!!! note "開始ポイント"
|
||||
|
||||
デフォルトの Python パスを使いたい場合は、まず [quickstart](quickstart.md) を読んでください。アプリでサーバーサイド WebSocket と SIP のどちらを使うべきか判断したい場合は、[Realtime transport](transport.md) を読んでください。ブラウザの WebRTC transport は Python SDK の対象外です。
|
||||
デフォルトの Python パスを使いたい場合は、まず [クイックスタート](quickstart.md) を読んでください。アプリでサーバー側 WebSocket または SIP のどちらを使うべきかを判断している場合は、 [Realtime トランスポート](transport.md) を読んでください。ブラウザー WebRTC トランスポートは Python SDK には含まれていません。
|
||||
|
||||
## 概要
|
||||
|
||||
Realtime エージェントは Realtime API への長時間接続を維持するため、モデルはテキストと音声を段階的に処理し、音声出力をストリーミングし、ツールを呼び出し、毎ターン新しいリクエストを再開せずに割り込みを処理できます。
|
||||
Realtime エージェントは Realtime API への長寿命の接続を開いたままにするため、モデルはテキストと音声を段階的に処理し、音声出力をストリーミングし、ツールを呼び出し、中断を処理できます。ターンごとに新しいリクエストを再開始する必要はありません。
|
||||
|
||||
主な SDK コンポーネントは次のとおりです。
|
||||
|
||||
- **RealtimeAgent**: 1 つの realtime 専門エージェント向けの instructions、ツール、出力ガードレール、ハンドオフ
|
||||
- **RealtimeRunner**: 開始エージェントを realtime transport に接続するセッションファクトリー
|
||||
- **RealtimeSession**: 入力送信、イベント受信、履歴追跡、ツール実行を行うライブセッション
|
||||
- **RealtimeModel**: transport 抽象化。デフォルトは OpenAI のサーバーサイド WebSocket 実装です。
|
||||
- **RealtimeAgent**: 1 つの realtime 専門エージェント向けの instructions、tools、出力ガードレール、ハンドオフ
|
||||
- **RealtimeRunner**: 開始エージェントを realtime トランスポートに接続するセッションファクトリ
|
||||
- **RealtimeSession**: 入力を送信し、イベントを受信し、履歴を追跡し、ツールを実行するライブセッション
|
||||
- **RealtimeModel**: トランスポート抽象化。デフォルトは OpenAI のサーバー側 WebSocket 実装です。
|
||||
|
||||
## セッションライフサイクル
|
||||
|
||||
典型的な realtime セッションは次のようになります。
|
||||
一般的な realtime セッションは次のようになります。
|
||||
|
||||
1. 1 つ以上の `RealtimeAgent` を作成します。
|
||||
2. 開始エージェントで `RealtimeRunner` を作成します。
|
||||
2. 開始エージェントを指定して `RealtimeRunner` を作成します。
|
||||
3. `await runner.run()` を呼び出して `RealtimeSession` を取得します。
|
||||
4. `async with session:` または `await session.enter()` でセッションに入ります。
|
||||
5. `send_message()` または `send_audio()` でユーザー入力を送信します。
|
||||
6. 会話が終了するまでセッションイベントを反復処理します。
|
||||
|
||||
テキスト専用 run とは異なり、`runner.run()` は最終 result を即時には生成しません。transport レイヤーと同期を保ちながら、ローカル履歴、バックグラウンドツール実行、ガードレール状態、アクティブなエージェント設定を保持するライブセッションオブジェクトを返します。
|
||||
テキストのみの実行とは異なり、 `runner.run()` は最終的な実行結果をすぐには生成しません。ローカル履歴、バックグラウンドでのツール実行、ガードレール状態、アクティブなエージェント設定をトランスポート層と同期し続けるライブセッションオブジェクトを返します。
|
||||
|
||||
デフォルトでは、`RealtimeRunner` は `OpenAIRealtimeWebSocketModel` を使用します。そのため、デフォルトの Python パスは Realtime API へのサーバーサイド WebSocket 接続です。別の `RealtimeModel` を渡した場合でも、同じセッションライフサイクルとエージェント機能が適用され、接続メカニズムのみ変更できます。
|
||||
デフォルトでは、 `RealtimeRunner` は `OpenAIRealtimeWebSocketModel` を使用するため、デフォルトの Python パスは Realtime API へのサーバー側 WebSocket 接続です。別の `RealtimeModel` を渡した場合でも、接続の仕組みは変わる可能性がありますが、同じセッションライフサイクルとエージェント機能が引き続き適用されます。
|
||||
|
||||
## エージェントとセッション設定
|
||||
|
||||
`RealtimeAgent` は通常の `Agent` 型より意図的に範囲が狭くなっています。
|
||||
`RealtimeAgent` は通常の `Agent` 型より意図的に機能範囲が絞られています。
|
||||
|
||||
- モデル選択はエージェントごとではなくセッションレベルで設定します。
|
||||
- structured outputs はサポートされていません。
|
||||
- Voice は設定できますが、セッションがすでに音声を生成した後は変更できません。
|
||||
- Instructions、関数ツール、ハンドオフ、フック、出力ガードレールはすべて引き続き利用できます。
|
||||
- モデル選択はエージェントごとではなく、セッションレベルで設定されます。
|
||||
- Structured outputs はサポートされていません。
|
||||
- voice は設定できますが、セッションがすでに発話音声を生成した後は変更できません。
|
||||
- instructions、関数ツール、ハンドオフ、フック、出力ガードレールはすべて引き続き機能します。
|
||||
|
||||
`RealtimeSessionModelSettings` は、新しいネストされた `audio` 設定と古いフラットなエイリアスの両方をサポートします。新規コードではネスト形式を推奨し、新しい realtime エージェントには `gpt-realtime-1.5` から始めてください。
|
||||
`RealtimeSessionModelSettings` は、新しいネストされた `audio` 設定と、古いフラットなエイリアスの両方をサポートします。新しいコードではネストされた形式を推奨し、新しい realtime エージェントでは `gpt-realtime-2` から始めてください。
|
||||
|
||||
```python
|
||||
runner = RealtimeRunner(
|
||||
starting_agent=agent,
|
||||
config={
|
||||
"model_settings": {
|
||||
"model_name": "gpt-realtime-1.5",
|
||||
"model_name": "gpt-realtime-2",
|
||||
"audio": {
|
||||
"input": {
|
||||
"format": "pcm16",
|
||||
@@ -71,7 +67,7 @@ runner = RealtimeRunner(
|
||||
)
|
||||
```
|
||||
|
||||
有用なセッションレベル設定には次が含まれます。
|
||||
有用なセッションレベル設定には次のものがあります。
|
||||
|
||||
- `audio.input.format`, `audio.output.format`
|
||||
- `audio.input.transcription`
|
||||
@@ -83,7 +79,7 @@ runner = RealtimeRunner(
|
||||
- `prompt`
|
||||
- `tracing`
|
||||
|
||||
`RealtimeRunner(config=...)` での有用な run レベル設定には次が含まれます。
|
||||
`RealtimeRunner(config=...)` の有用な run レベル設定には次のものがあります。
|
||||
|
||||
- `async_tool_calls`
|
||||
- `output_guardrails`
|
||||
@@ -91,13 +87,13 @@ runner = RealtimeRunner(
|
||||
- `tool_error_formatter`
|
||||
- `tracing_disabled`
|
||||
|
||||
型付きの完全な仕様は [`RealtimeRunConfig`][agents.realtime.config.RealtimeRunConfig] と [`RealtimeSessionModelSettings`][agents.realtime.config.RealtimeSessionModelSettings] を参照してください。
|
||||
完全な型付きインターフェースについては、 [`RealtimeRunConfig`][agents.realtime.config.RealtimeRunConfig] と [`RealtimeSessionModelSettings`][agents.realtime.config.RealtimeSessionModelSettings] を参照してください。
|
||||
|
||||
## 入力と出力
|
||||
|
||||
### テキストと構造化ユーザーメッセージ
|
||||
|
||||
プレーンテキストまたは構造化 realtime メッセージには [`session.send_message()`][agents.realtime.session.RealtimeSession.send_message] を使用します。
|
||||
プレーンテキストまたは構造化された realtime メッセージには、 [`session.send_message()`][agents.realtime.session.RealtimeSession.send_message] を使用します。
|
||||
|
||||
```python
|
||||
from agents.realtime import RealtimeUserInputMessage
|
||||
@@ -115,31 +111,31 @@ message: RealtimeUserInputMessage = {
|
||||
await session.send_message(message)
|
||||
```
|
||||
|
||||
構造化メッセージは、realtime 会話に画像入力を含める主要な方法です。[`examples/realtime/app/server.py`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/app/server.py) の Web デモ例では、この方法で `input_image` メッセージを転送しています。
|
||||
構造化メッセージは、 realtime 会話に画像入力を含める主な方法です。 [`examples/realtime/app/server.py`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/app/server.py) のサンプル Web デモでは、この方法で `input_image` メッセージを転送しています。
|
||||
|
||||
### 音声入力
|
||||
|
||||
raw 音声バイトをストリーミングするには [`session.send_audio()`][agents.realtime.session.RealtimeSession.send_audio] を使用します。
|
||||
生の音声バイトをストリーミングするには、 [`session.send_audio()`][agents.realtime.session.RealtimeSession.send_audio] を使用します。
|
||||
|
||||
```python
|
||||
await session.send_audio(audio_bytes)
|
||||
```
|
||||
|
||||
サーバーサイドの turn detection が無効な場合、ターン境界の指定はユーザー側の責任です。高レベルの簡易手段は次のとおりです。
|
||||
サーバー側のターン検出が無効な場合、ターン境界をマークする責任は開発者側にあります。高レベルの便利な方法は次のとおりです。
|
||||
|
||||
```python
|
||||
await session.send_audio(audio_bytes, commit=True)
|
||||
```
|
||||
|
||||
より低レベルな制御が必要な場合は、基盤となる model transport を通じて `input_audio_buffer.commit` などの raw client event も送信できます。
|
||||
より低レベルの制御が必要な場合は、基盤となるモデルトランスポートを通じて `input_audio_buffer.commit` などの raw クライアントイベントを送信することもできます。
|
||||
|
||||
### 手動レスポンス制御
|
||||
### 手動応答制御
|
||||
|
||||
`session.send_message()` は高レベルパスでユーザー入力を送信し、レスポンス開始も自動で行います。raw 音声バッファリングでは、すべての設定で同様に自動実行される **わけではありません** 。
|
||||
`session.send_message()` は高レベルパスを使ってユーザー入力を送信し、応答を開始します。生の音声バッファリングは、すべての設定で同じことを自動的に行うわけでは **ありません** 。
|
||||
|
||||
Realtime API レベルでは、手動ターン制御は raw `session.update` で `turn_detection` をクリアし、その後 `input_audio_buffer.commit` と `response.create` を自分で送信することを意味します。
|
||||
Realtime API レベルでは、手動のターン制御とは、 raw `session.update` で `turn_detection` をクリアし、その後に `input_audio_buffer.commit` と `response.create` を自分で送信することを意味します。
|
||||
|
||||
ターンを手動管理する場合は、model transport 経由で raw client event を送信できます。
|
||||
ターンを手動で管理している場合は、モデルのトランスポートを通じて raw クライアントイベントを送信できます。
|
||||
|
||||
```python
|
||||
from agents.realtime.model_inputs import RealtimeModelSendRawMessage
|
||||
@@ -155,17 +151,17 @@ await session.model.send_event(
|
||||
|
||||
このパターンは次の場合に有用です。
|
||||
|
||||
- `turn_detection` が無効で、モデルがいつ応答するかを自分で決めたい場合
|
||||
- レスポンスをトリガーする前にユーザー入力を検査またはゲートしたい場合
|
||||
- out-of-band レスポンス向けにカスタムプロンプトが必要な場合
|
||||
- `turn_detection` が無効で、モデルがいつ応答すべきかを自分で決めたい場合
|
||||
- 応答をトリガーする前にユーザー入力を検査またはゲートしたい場合
|
||||
- アウトオブバンド応答にカスタムプロンプトが必要な場合
|
||||
|
||||
[`examples/realtime/twilio_sip/server.py`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/twilio_sip/server.py) の SIP 例では、raw `response.create` を使って開始時の挨拶を強制しています。
|
||||
[`examples/realtime/twilio_sip/server.py`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/twilio_sip/server.py) の SIP の例では、開始時の挨拶を強制するために raw `response.create` を使用しています。
|
||||
|
||||
## イベント、履歴、割り込み
|
||||
## イベント、履歴、中断
|
||||
|
||||
`RealtimeSession` は高レベル SDK イベントを発行しつつ、必要時には raw model event も転送します。
|
||||
`RealtimeSession` は、必要なときには raw モデルイベントも転送しつつ、より高レベルの SDK イベントを発行します。
|
||||
|
||||
価値の高いセッションイベントには次が含まれます。
|
||||
有用性の高いセッションイベントには次のものがあります。
|
||||
|
||||
- `audio`, `audio_end`, `audio_interrupted`
|
||||
- `agent_start`, `agent_end`
|
||||
@@ -177,21 +173,21 @@ await session.model.send_event(
|
||||
- `error`
|
||||
- `raw_model_event`
|
||||
|
||||
UI 状態管理で特に有用なのは通常 `history_added` と `history_updated` です。これらは、ユーザーメッセージ、assistant メッセージ、ツール呼び出しを含むセッションのローカル履歴を `RealtimeItem` オブジェクトとして公開します。
|
||||
UI 状態に最も有用なイベントは通常 `history_added` と `history_updated` です。これらは、ユーザーメッセージ、アシスタントメッセージ、ツール呼び出しを含む、セッションのローカル履歴を `RealtimeItem` オブジェクトとして公開します。
|
||||
|
||||
### 割り込みと再生追跡
|
||||
### 中断と再生トラッキング
|
||||
|
||||
ユーザーが assistant を割り込んだ場合、セッションは `audio_interrupted` を発行し、サーバーサイド会話がユーザーの実際の聴取内容と一致するよう履歴を更新します。
|
||||
ユーザーがアシスタントを中断すると、セッションは `audio_interrupted` を発行し、履歴を更新して、サーバー側の会話がユーザーが実際に聞いた内容と一致し続けるようにします。
|
||||
|
||||
低遅延のローカル再生では、デフォルトの再生トラッカーで十分なことが多いです。リモート再生や遅延再生のシナリオ、特に電話では、すべての生成音声がすでに聴取済みと仮定するのではなく、実際の再生進捗に基づいて割り込み切り詰めを行うために [`RealtimePlaybackTracker`][agents.realtime.model.RealtimePlaybackTracker] を使用してください。
|
||||
低レイテンシーのローカル再生では、デフォルトの再生トラッカーで十分なことが多いです。リモート再生や遅延再生のシナリオ、特にテレフォニーでは、 [`RealtimePlaybackTracker`][agents.realtime.model.RealtimePlaybackTracker] を使用してください。これにより、生成された音声がすべてすでに聞かれたと仮定するのではなく、実際の再生進行に基づいて中断時の切り詰めが行われます。
|
||||
|
||||
[`examples/realtime/twilio/twilio_handler.py`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/twilio/twilio_handler.py) の Twilio 例はこのパターンを示しています。
|
||||
[`examples/realtime/twilio/twilio_handler.py`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/twilio/twilio_handler.py) の Twilio の例はこのパターンを示しています。
|
||||
|
||||
## ツール、承認、ハンドオフ、ガードレール
|
||||
|
||||
### 関数ツール
|
||||
|
||||
Realtime エージェントはライブ会話中の関数ツールをサポートします。
|
||||
Realtime エージェントはライブ会話中に関数ツールをサポートします。
|
||||
|
||||
```python
|
||||
from agents import function_tool
|
||||
@@ -212,7 +208,9 @@ agent = RealtimeAgent(
|
||||
|
||||
### ツール承認
|
||||
|
||||
関数ツールは、実行前に人間の承認を必要とするようにできます。その場合、セッションは `tool_approval_required` を発行し、`approve_tool_call()` または `reject_tool_call()` を呼び出すまでツール実行を一時停止します。
|
||||
関数ツールは実行前に人間の承認を要求できます。その場合、セッションは `tool_approval_required` を発行し、 `approve_tool_call()` または `reject_tool_call()` を呼び出すまでツール実行を一時停止します。
|
||||
|
||||
ツールに入力ガードレールもある場合、それらのガードレールは承認後、実行の直前に実行されます。承認イベントが発行される前にそれらを実行するには、 `RealtimeRunner(..., config={"tool_execution": {"pre_approval_tool_input_guardrails": True}})` で runner を作成します。この事前承認チェックに合格した呼び出しも、承認後、実行前にもう一度チェックされます。
|
||||
|
||||
```python
|
||||
async for event in session:
|
||||
@@ -220,11 +218,11 @@ async for event in session:
|
||||
await session.approve_tool_call(event.call_id)
|
||||
```
|
||||
|
||||
具体的なサーバーサイド承認ループは [`examples/realtime/app/server.py`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/app/server.py) を参照してください。human-in-the-loop ドキュメントでも [Human in the loop](../human_in_the_loop.md) でこのフローを参照しています。
|
||||
具体的なサーバー側承認ループについては、 [`examples/realtime/app/server.py`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/app/server.py) を参照してください。人間参加型のドキュメントも、 [人間を介した処理](../human_in_the_loop.md) でこのフローを参照しています。
|
||||
|
||||
### ハンドオフ
|
||||
|
||||
Realtime ハンドオフでは、あるエージェントがライブ会話を別の専門エージェントへ転送できます。
|
||||
Realtime ハンドオフにより、 1 つのエージェントがライブ会話を別の専門エージェントへ転送できます。
|
||||
|
||||
```python
|
||||
from agents.realtime import RealtimeAgent, realtime_handoff
|
||||
@@ -237,15 +235,20 @@ billing_agent = RealtimeAgent(
|
||||
main_agent = RealtimeAgent(
|
||||
name="Customer Service",
|
||||
instructions="Triage the request and hand off when needed.",
|
||||
handoffs=[realtime_handoff(billing_agent, tool_description="Transfer to billing support")],
|
||||
handoffs=[
|
||||
realtime_handoff(
|
||||
billing_agent,
|
||||
tool_description_override="Transfer to billing support",
|
||||
)
|
||||
],
|
||||
)
|
||||
```
|
||||
|
||||
素の `RealtimeAgent` ハンドオフは自動ラップされ、`realtime_handoff(...)` では名前、説明、検証、コールバック、可用性をカスタマイズできます。Realtime ハンドオフは通常の handoff `input_filter` をサポートしません。
|
||||
そのままの `RealtimeAgent` ハンドオフは自動的にラップされ、 `realtime_handoff(...)` では名前、説明、検証、コールバック、利用可否をカスタマイズできます。Realtime ハンドオフは通常のハンドオフの `input_filter` を **サポートしていません** 。
|
||||
|
||||
### ガードレール
|
||||
|
||||
Realtime エージェントでサポートされるのは出力ガードレールのみです。これらは各部分 token ごとではなく、デバウンスされた transcript 蓄積に対して実行され、例外を送出する代わりに `guardrail_tripped` を発行します。
|
||||
Realtime エージェントは、エージェント応答に対する出力ガードレールと、関数ツール呼び出しに対する入力ガードレールをサポートします。出力ガードレールは、部分的なトークンごとではなく、デバウンスされたトランスクリプトの蓄積に対して実行され、例外を発生させる代わりに `guardrail_tripped` を発行します。
|
||||
|
||||
```python
|
||||
from agents.guardrail import GuardrailFunctionOutput, OutputGuardrail
|
||||
@@ -265,11 +268,13 @@ agent = RealtimeAgent(
|
||||
)
|
||||
```
|
||||
|
||||
realtime 出力ガードレールが作動すると、セッションはアクティブな応答を中断し、 `response.cancel` を強制し、 `guardrail_tripped` を発行し、作動したガードレールの名前を示す後続のユーザーメッセージを送信して、モデルが代替応答を生成できるようにします。音声プレーヤーはそれでも `audio_interrupted` をリッスンし、ローカル再生をただちに停止する必要があります。これは、ガードレールがデバウンスされたトランスクリプトテキストに対して実行され、トリップワイヤーが作動した時点で一部の音声がすでにバッファリングされている可能性があるためです。
|
||||
|
||||
## SIP とテレフォニー
|
||||
|
||||
Python SDK には [`OpenAIRealtimeSIPModel`][agents.realtime.openai_realtime.OpenAIRealtimeSIPModel] による第一級の SIP 接続フローが含まれています。
|
||||
Python SDK には、 [`OpenAIRealtimeSIPModel`][agents.realtime.openai_realtime.OpenAIRealtimeSIPModel] を介したファーストクラスの SIP アタッチフローが含まれています。
|
||||
|
||||
Realtime Calls API 経由で着信し、結果として得られる `call_id` にエージェントセッションを接続したい場合に使用します。
|
||||
Realtime Calls API 経由で通話が到着し、得られた `call_id` にエージェントセッションをアタッチしたい場合に使用します。
|
||||
|
||||
```python
|
||||
from agents.realtime import RealtimeRunner
|
||||
@@ -286,18 +291,18 @@ async with await runner.run(
|
||||
...
|
||||
```
|
||||
|
||||
まず通話を受け付ける必要があり、受け付けペイロードをエージェント由来のセッション設定に一致させたい場合は、`OpenAIRealtimeSIPModel.build_initial_session_payload(...)` を使用してください。完全なフローは [`examples/realtime/twilio_sip/server.py`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/twilio_sip/server.py) にあります。
|
||||
先に通話を受け付ける必要があり、 accept ペイロードをエージェント由来のセッション設定と一致させたい場合は、 `OpenAIRealtimeSIPModel.build_initial_session_payload(...)` を使用します。完全なフローは [`examples/realtime/twilio_sip/server.py`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/twilio_sip/server.py) に示されています。
|
||||
|
||||
## 低レベルアクセスとカスタムエンドポイント
|
||||
|
||||
`session.model` から基盤 transport オブジェクトにアクセスできます。
|
||||
`session.model` を通じて基盤となるトランスポートオブジェクトにアクセスできます。
|
||||
|
||||
必要な場合に使用します。
|
||||
次が必要な場合に使用します。
|
||||
|
||||
- `session.model.add_listener(...)` によるカスタムリスナー
|
||||
- `response.create` や `session.update` などの raw client event
|
||||
- `model_config` 経由のカスタム `url`、`headers`、`api_key` 処理
|
||||
- 既存 realtime 通話への `call_id` 接続
|
||||
- `response.create` や `session.update` などの raw クライアントイベント
|
||||
- `model_config` によるカスタムの `url`、 `headers`、 `api_key` 処理
|
||||
- 既存の realtime 通話への `call_id` アタッチ
|
||||
|
||||
`RealtimeModelConfig` は次をサポートします。
|
||||
|
||||
@@ -308,9 +313,9 @@ async with await runner.run(
|
||||
- `playback_tracker`
|
||||
- `call_id`
|
||||
|
||||
このリポジトリに含まれる `call_id` の例は SIP です。より広い Realtime API では一部のサーバーサイド制御フローにも `call_id` を使いますが、ここでは Python 例としては提供されていません。
|
||||
このリポジトリに同梱されている `call_id` のコード例は SIP です。より広範な Realtime API でも、一部のサーバー側制御フローに `call_id` を使用しますが、それらはここでは Python のコード例としてパッケージ化されていません。
|
||||
|
||||
Azure OpenAI に接続する場合は、 GA Realtime endpoint URL と明示的な headers を渡してください。例:
|
||||
Azure OpenAI に接続する場合は、 GA Realtime エンドポイント URL と明示的なヘッダーを渡します。例:
|
||||
|
||||
```python
|
||||
session = await runner.run(
|
||||
@@ -321,7 +326,7 @@ session = await runner.run(
|
||||
)
|
||||
```
|
||||
|
||||
トークンベース認証では、`headers` に bearer token を使用します。
|
||||
トークンベースの認証では、 `headers` に Bearer トークンを使用します。
|
||||
|
||||
```python
|
||||
session = await runner.run(
|
||||
@@ -332,12 +337,12 @@ session = await runner.run(
|
||||
)
|
||||
```
|
||||
|
||||
`headers` を渡した場合、SDK は `Authorization` を自動追加しません。realtime エージェントではレガシー beta パス(`/openai/realtime?api-version=...`)を避けてください。
|
||||
`headers` を渡すと、 SDK は `Authorization` を自動的には追加しません。realtime エージェントでは、従来の beta パス (`/openai/realtime?api-version=...`) は避けてください。
|
||||
|
||||
## 参考資料
|
||||
## 関連資料
|
||||
|
||||
- [Realtime transport](transport.md)
|
||||
- [Quickstart](quickstart.md)
|
||||
- [OpenAI Realtime conversations](https://developers.openai.com/api/docs/guides/realtime-conversations/)
|
||||
- [OpenAI Realtime server-side controls](https://developers.openai.com/api/docs/guides/realtime-server-controls/)
|
||||
- [Realtime トランスポート](transport.md)
|
||||
- [クイックスタート](quickstart.md)
|
||||
- [OpenAI Realtime 会話](https://developers.openai.com/api/docs/guides/realtime-conversations/)
|
||||
- [OpenAI Realtime サーバー側コントロール](https://developers.openai.com/api/docs/guides/realtime-server-controls/)
|
||||
- [`examples/realtime`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime)
|
||||
@@ -4,33 +4,29 @@ search:
|
||||
---
|
||||
# クイックスタート
|
||||
|
||||
Python SDK の Realtime エージェントは、WebSocket トランスポート経由の OpenAI Realtime API 上に構築された、サーバーサイドの低レイテンシなエージェントです。
|
||||
Python SDK のリアルタイムエージェントは、WebSocket トランスポート経由の OpenAI Realtime API 上に構築された、サーバー側の低レイテンシーエージェントです。
|
||||
|
||||
!!! warning "Beta 機能"
|
||||
!!! note "Python SDK の境界"
|
||||
|
||||
Realtime エージェントは beta です。実装の改善に伴い、破壊的変更が発生する可能性があります。
|
||||
|
||||
!!! note "Python SDK の範囲"
|
||||
|
||||
Python SDK はブラウザー向けの WebRTC トランスポートを **提供しません** 。このページでは、サーバーサイド WebSocket 経由で Python が管理する realtime session のみを扱います。サーバーサイドのオーケストレーション、ツール、承認、テレフォニー統合にはこの SDK を使用してください。あわせて [Realtime transport](transport.md) も参照してください。
|
||||
Python SDK は、ブラウザーの WebRTC トランスポートを **提供していません** 。このページでは、サーバー側 WebSocket を介して Python で管理されるリアルタイムセッションのみを扱います。この SDK は、サーバー側のオーケストレーション、ツール、承認、テレフォニー連携に使用してください。併せて [リアルタイムトランスポート](transport.md) も参照してください。
|
||||
|
||||
## 前提条件
|
||||
|
||||
- Python 3.10 以上
|
||||
- OpenAI API キー
|
||||
- OpenAI Agents SDK の基本的な理解
|
||||
- OpenAI Agents SDK の基本的な知識
|
||||
|
||||
## インストール
|
||||
|
||||
まだの場合は、OpenAI Agents SDK をインストールします。
|
||||
まだインストールしていない場合は、OpenAI Agents SDK をインストールしてください:
|
||||
|
||||
```bash
|
||||
pip install openai-agents
|
||||
```
|
||||
|
||||
## サーバーサイド realtime session の作成
|
||||
## サーバー側リアルタイムセッションの作成
|
||||
|
||||
### 1. Realtime コンポーネントのインポート
|
||||
### 1. リアルタイムコンポーネントのインポート
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -47,16 +43,16 @@ agent = RealtimeAgent(
|
||||
)
|
||||
```
|
||||
|
||||
### 3. runner の設定
|
||||
### 3. ランナーの設定
|
||||
|
||||
新しいコードでは、ネストされた `audio.input` / `audio.output` session 設定の形式を推奨します。新しい Realtime エージェントでは、`gpt-realtime-1.5` から始めてください。
|
||||
新しいコードでは、ネストされた `audio.input` / `audio.output` のセッション設定形式を推奨します。新しいリアルタイムエージェントでは、`gpt-realtime-2` から始めてください。
|
||||
|
||||
```python
|
||||
runner = RealtimeRunner(
|
||||
starting_agent=agent,
|
||||
config={
|
||||
"model_settings": {
|
||||
"model_name": "gpt-realtime-1.5",
|
||||
"model_name": "gpt-realtime-2",
|
||||
"audio": {
|
||||
"input": {
|
||||
"format": "pcm16",
|
||||
@@ -76,9 +72,9 @@ runner = RealtimeRunner(
|
||||
)
|
||||
```
|
||||
|
||||
### 4. session の開始と入力の送信
|
||||
### 4. セッションの開始と入力の送信
|
||||
|
||||
`runner.run()` は `RealtimeSession` を返します。session context に入ると接続が開かれます。
|
||||
`runner.run()` は `RealtimeSession` を返します。セッションコンテキストに入ると接続が開かれます。
|
||||
|
||||
```python
|
||||
async def main() -> None:
|
||||
@@ -104,59 +100,59 @@ if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
`session.send_message()` はプレーンな文字列または構造化された realtime message のいずれかを受け取ります。raw audio chunk には [`session.send_audio()`][agents.realtime.session.RealtimeSession.send_audio] を使用してください。
|
||||
`session.send_message()` は、プレーン文字列または構造化されたリアルタイムメッセージのいずれかを受け取ります。生の音声チャンクには、[`session.send_audio()`][agents.realtime.session.RealtimeSession.send_audio] を使用してください。
|
||||
|
||||
## このクイックスタートに含まれない内容
|
||||
|
||||
- マイク入力とスピーカー再生のコード。[`examples/realtime`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime) の realtime コード例を参照してください。
|
||||
- SIP / テレフォニー接続フロー。[Realtime transport](transport.md) と [SIP セクション](guide.md#sip-and-telephony) を参照してください。
|
||||
- マイクのキャプチャおよびスピーカー再生のコード。[`examples/realtime`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime) のリアルタイムのコード例を参照してください。
|
||||
- SIP / テレフォニーのアタッチフロー。[リアルタイムトランスポート](transport.md) と [SIP セクション](guide.md#sip-and-telephony) を参照してください。
|
||||
|
||||
## 主要設定
|
||||
|
||||
基本的な session が動作したら、次によく使われる設定は以下です。
|
||||
基本的なセッションが動作するようになったら、多くの方が次に利用する設定は次のとおりです:
|
||||
|
||||
- `model_name`
|
||||
- `audio.input.format`, `audio.output.format`
|
||||
- `audio.input.transcription`
|
||||
- `audio.input.noise_reduction`
|
||||
- 自動ターン検出のための `audio.input.turn_detection`
|
||||
- `audio.input.turn_detection`(自動ターン検出用)
|
||||
- `audio.output.voice`
|
||||
- `tool_choice`, `prompt`, `tracing`
|
||||
- `async_tool_calls`, `guardrails_settings.debounce_text_length`, `tool_error_formatter`
|
||||
- `async_tool_calls`, `tool_execution.pre_approval_tool_input_guardrails`, `guardrails_settings.debounce_text_length`, `tool_error_formatter`
|
||||
|
||||
`input_audio_format`、`output_audio_format`、`input_audio_transcription`、`turn_detection` などの古いフラットな別名も引き続き動作しますが、新しいコードではネストされた `audio` 設定を推奨します。
|
||||
`input_audio_format`、`output_audio_format`、`input_audio_transcription`、`turn_detection` などの古いフラットなエイリアスも引き続き機能しますが、新しいコードではネストされた `audio` 設定が推奨されます。
|
||||
|
||||
手動でターン制御を行う場合は、[Realtime agents guide](guide.md#manual-response-control) にある説明のとおり、raw の `session.update` / `input_audio_buffer.commit` / `response.create` フローを使用してください。
|
||||
手動のターン制御には、[リアルタイムエージェントガイド](guide.md#manual-response-control) で説明されている raw な `session.update` / `input_audio_buffer.commit` / `response.create` フローを使用してください。
|
||||
|
||||
完全なスキーマについては、[`RealtimeRunConfig`][agents.realtime.config.RealtimeRunConfig] と [`RealtimeSessionModelSettings`][agents.realtime.config.RealtimeSessionModelSettings] を参照してください。
|
||||
完全なスキーマについては、[`RealtimeRunConfig`][agents.realtime.config.RealtimeRunConfig] および [`RealtimeSessionModelSettings`][agents.realtime.config.RealtimeSessionModelSettings] を参照してください。
|
||||
|
||||
## 接続オプション
|
||||
|
||||
環境変数に API キーを設定します。
|
||||
API キーを環境変数に設定してください:
|
||||
|
||||
```bash
|
||||
export OPENAI_API_KEY="your-api-key-here"
|
||||
```
|
||||
|
||||
または、session 開始時に直接渡します。
|
||||
または、セッションの開始時に直接渡します:
|
||||
|
||||
```python
|
||||
session = await runner.run(model_config={"api_key": "your-api-key"})
|
||||
```
|
||||
|
||||
`model_config` は次もサポートします。
|
||||
`model_config` は以下にも対応しています:
|
||||
|
||||
- `url`: カスタム WebSocket endpoint
|
||||
- `headers`: カスタム request header
|
||||
- `call_id`: 既存の realtime call に接続します。このリポジトリで文書化されている接続フローは SIP です。
|
||||
- `playback_tracker`: ユーザーが実際に聞いた audio の量を報告します
|
||||
- `url`: カスタム WebSocket エンドポイント
|
||||
- `headers`: カスタムリクエストヘッダー
|
||||
- `call_id`: 既存のリアルタイムコールにアタッチします。このリポジトリでドキュメント化されているアタッチフローは SIP です。
|
||||
- `playback_tracker`: ユーザーが実際に聞いた音声量を報告します
|
||||
|
||||
`headers` を明示的に渡した場合、SDK は `Authorization` header を **自動挿入しません** 。
|
||||
`headers` を明示的に渡す場合、SDK は `Authorization` ヘッダーを **挿入しません** 。
|
||||
|
||||
Azure OpenAI に接続する場合は、`model_config["url"]` に GA Realtime endpoint URL と明示的な headers を渡してください。realtime エージェントでは、legacy beta path (`/openai/realtime?api-version=...`) を避けてください。詳細は [Realtime agents guide](guide.md#low-level-access-and-custom-endpoints) を参照してください。
|
||||
Azure OpenAI に接続する場合は、`model_config["url"]` に GA Realtime エンドポイント URL を指定し、明示的なヘッダーも渡してください。リアルタイムエージェントでは、レガシーな beta パス(`/openai/realtime?api-version=...`)の使用は避けてください。詳細については、[リアルタイムエージェントガイド](guide.md#low-level-access-and-custom-endpoints) を参照してください。
|
||||
|
||||
## 次のステップ
|
||||
|
||||
- サーバーサイド WebSocket と SIP のどちらを選ぶか判断するために [Realtime transport](transport.md) を読んでください。
|
||||
- ライフサイクル、構造化入力、承認、ハンドオフ、ガードレール、低レベル制御について [Realtime agents guide](guide.md) を読んでください。
|
||||
- [`examples/realtime`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime) のコード例を確認してください。
|
||||
- サーバー側 WebSocket と SIP のどちらを選ぶかを判断するには、[リアルタイムトランスポート](transport.md) をお読みください。
|
||||
- ライフサイクル、構造化入力、承認、ハンドオフ、ガードレール、低レベル制御については、[リアルタイムエージェントガイド](guide.md) をお読みください。
|
||||
- [`examples/realtime`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime) のコード例をご覧ください。
|
||||
@@ -4,32 +4,32 @@ search:
|
||||
---
|
||||
# Realtime トランスポート
|
||||
|
||||
このページは、realtime エージェントを Python アプリケーションにどのように組み込むかを判断するために使用します。
|
||||
このページは、realtime エージェントを Python アプリケーションにどのように組み込むかを判断するために使用してください。
|
||||
|
||||
!!! note "Python SDK の境界"
|
||||
|
||||
Python SDK にはブラウザー WebRTC トランスポートは **含まれていません** 。このページは Python SDK のトランスポート選択、つまりサーバーサイド WebSocket と SIP アタッチフローのみを対象としています。ブラウザー WebRTC は別のプラットフォームトピックであり、公式の [Realtime API with WebRTC](https://developers.openai.com/api/docs/guides/realtime-webrtc/) ガイドに記載されています。
|
||||
Python SDK には、ブラウザー WebRTC トランスポートは含まれて **いません**。このページは、Python SDK のトランスポート選択肢であるサーバー側 WebSocket と SIP アタッチフローのみを扱います。ブラウザー WebRTC は別のプラットフォームトピックであり、公式の [WebRTC による Realtime API](https://developers.openai.com/api/docs/guides/realtime-webrtc/) ガイドに記載されています。
|
||||
|
||||
## 判断ガイド
|
||||
|
||||
| Goal | Start with | Why |
|
||||
| 目的 | はじめに | 理由 |
|
||||
| --- | --- | --- |
|
||||
| サーバー管理の realtime アプリを構築する | [Quickstart](quickstart.md) | デフォルトの Python パスは、`RealtimeRunner` で管理されるサーバーサイド WebSocket セッションです。 |
|
||||
| どのトランスポートとデプロイ形状を選ぶべきか理解する | このページ | トランスポートやデプロイ形状を確定する前に、このページを使用してください。 |
|
||||
| エージェントを電話または SIP 通話にアタッチする | [Realtime guide](guide.md) と [`examples/realtime/twilio_sip`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/twilio_sip) | このリポジトリには、`call_id` で駆動する SIP アタッチフローが含まれています。 |
|
||||
| サーバー管理の realtime アプリを構築する | [クイックスタート](quickstart.md) | デフォルトの Python パスは、`RealtimeRunner` によって管理されるサーバー側 WebSocket セッションです。 |
|
||||
| 選択すべきトランスポートとデプロイ形態を理解する | このページ | トランスポートやデプロイ形態を決定する前に使用してください。 |
|
||||
| エージェントを電話または SIP 通話にアタッチする | [Realtime ガイド](guide.md) と [`examples/realtime/twilio_sip`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/twilio_sip) | このリポジトリには、`call_id` によって駆動される SIP アタッチフローが含まれています。 |
|
||||
|
||||
## サーバーサイド WebSocket というデフォルトの Python パス
|
||||
## デフォルトの Python パスであるサーバー側 WebSocket
|
||||
|
||||
`RealtimeRunner` は、カスタム `RealtimeModel` を渡さない限り `OpenAIRealtimeWebSocketModel` を使用します。
|
||||
カスタム `RealtimeModel` を渡さない限り、`RealtimeRunner` は `OpenAIRealtimeWebSocketModel` を使用します。
|
||||
|
||||
つまり、標準的な Python トポロジーは次のようになります。
|
||||
|
||||
1. Python サービスが `RealtimeRunner` を作成します。
|
||||
2. `await runner.run()` は `RealtimeSession` を返します。
|
||||
2. `await runner.run()` が `RealtimeSession` を返します。
|
||||
3. セッションに入り、テキスト、構造化メッセージ、または音声を送信します。
|
||||
4. `RealtimeSessionEvent` 項目を消費し、音声またはトランスクリプトをアプリケーションに転送します。
|
||||
|
||||
このトポロジーは、コアデモアプリ、CLI 例、Twilio Media Streams 例で使用されています。
|
||||
これは、コアデモアプリ、CLI の例、Twilio Media Streams の例で使用されているトポロジーです。
|
||||
|
||||
- [`examples/realtime/app`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/app)
|
||||
- [`examples/realtime/cli`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/cli)
|
||||
@@ -37,40 +37,40 @@ search:
|
||||
|
||||
サーバーが音声パイプライン、ツール実行、承認フロー、履歴処理を管理する場合は、このパスを使用してください。
|
||||
|
||||
## SIP アタッチというテレフォニーパス
|
||||
## テレフォニー向けパスとしての SIP アタッチ
|
||||
|
||||
このリポジトリで文書化されているテレフォニーフローでは、Python SDK は `call_id` を介して既存の realtime 通話にアタッチします。
|
||||
このリポジトリで説明されているテレフォニーフローでは、Python SDK は `call_id` を介して既存の realtime 通話にアタッチします。
|
||||
|
||||
このトポロジーは次のようになります。
|
||||
|
||||
1. OpenAI が `realtime.call.incoming` などの webhook をサービスに送信します。
|
||||
2. サービスが Realtime Calls API を通じて通話を受け付けます。
|
||||
2. サービスが Realtime Calls API を通じて通話を受け入れます。
|
||||
3. Python サービスが `RealtimeRunner(..., model=OpenAIRealtimeSIPModel())` を開始します。
|
||||
4. セッションは `model_config={"call_id": ...}` で接続し、その後は他の realtime セッションと同様にイベントを処理します。
|
||||
|
||||
これは [`examples/realtime/twilio_sip`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/twilio_sip) で示されているトポロジーです。
|
||||
これは [`examples/realtime/twilio_sip`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/twilio_sip) に示されているトポロジーです。
|
||||
|
||||
より広い Realtime API でも一部のサーバーサイド制御パターンで `call_id` を使用しますが、このリポジトリで提供されているアタッチ例は SIP です。
|
||||
より広範な Realtime API でも、一部のサーバー側制御パターンで `call_id` を使用しますが、このリポジトリに含まれるアタッチ例は SIP です。
|
||||
|
||||
## この SDK の対象外であるブラウザー WebRTC
|
||||
## この SDK の範囲外であるブラウザー WebRTC
|
||||
|
||||
アプリの主要クライアントが Realtime WebRTC を使用するブラウザーである場合:
|
||||
アプリの主なクライアントが Realtime WebRTC を使用するブラウザーである場合:
|
||||
|
||||
- このリポジトリの Python SDK ドキュメントの対象外として扱ってください。
|
||||
- クライアントサイドフローとイベントモデルについては、公式の [Realtime API with WebRTC](https://developers.openai.com/api/docs/guides/realtime-webrtc/) と [Realtime conversations](https://developers.openai.com/api/docs/guides/realtime-conversations/) のドキュメントを使用してください。
|
||||
- ブラウザー WebRTC クライアントに加えてサイドバンドのサーバー接続が必要な場合は、公式の [Realtime server-side controls](https://developers.openai.com/api/docs/guides/realtime-server-controls/) ガイドを使用してください。
|
||||
- このリポジトリがブラウザーサイド `RTCPeerConnection` 抽象化や、すぐに使えるブラウザー WebRTC サンプルを提供することは期待しないでください。
|
||||
- このリポジトリの Python SDK ドキュメントの範囲外として扱ってください。
|
||||
- クライアント側のフローとイベントモデルについては、公式の [WebRTC による Realtime API](https://developers.openai.com/api/docs/guides/realtime-webrtc/) および [Realtime conversations](https://developers.openai.com/api/docs/guides/realtime-conversations/) ドキュメントを使用してください。
|
||||
- ブラウザー WebRTC クライアントの上にサイドバンドのサーバー接続が必要な場合は、公式の [Realtime server-side controls](https://developers.openai.com/api/docs/guides/realtime-server-controls/) ガイドを使用してください。
|
||||
- このリポジトリが、ブラウザー側の `RTCPeerConnection` 抽象化や、すぐに使えるブラウザー WebRTC サンプルを提供することは期待しないでください。
|
||||
|
||||
このリポジトリには現在、ブラウザー WebRTC と Python サイドバンドを組み合わせた例も含まれていません。
|
||||
このリポジトリには、現在、ブラウザー WebRTC と Python サイドバンドを組み合わせた例も含まれていません。
|
||||
|
||||
## カスタムエンドポイントとアタッチポイント
|
||||
|
||||
[`RealtimeModelConfig`][agents.realtime.model.RealtimeModelConfig] のトランスポート設定インターフェースにより、デフォルトパスを調整できます。
|
||||
[`RealtimeModelConfig`][agents.realtime.model.RealtimeModelConfig] のトランスポート設定サーフェスを使用すると、デフォルトのパスを調整できます。
|
||||
|
||||
- `url`: WebSocket エンドポイントを上書きします
|
||||
- `headers`: Azure 認証ヘッダーなどの明示的なヘッダーを提供します
|
||||
- `headers`: Azure 認証ヘッダーなどの明示的なヘッダーを指定します
|
||||
- `api_key`: API キーを直接、またはコールバック経由で渡します
|
||||
- `call_id`: 既存の realtime 通話にアタッチします。このリポジトリで文書化されている例は SIP です。
|
||||
- `playback_tracker`: 割り込み処理のために実際の再生進行を報告します
|
||||
- `call_id`: 既存の realtime 通話にアタッチします。このリポジトリで記載されている例は SIP です。
|
||||
- `playback_tracker`: 割り込み処理のために実際の再生進捗を報告します
|
||||
|
||||
トポロジーを選択した後の詳細なライフサイクルと機能インターフェースについては、[Realtime agents guide](guide.md) を参照してください。
|
||||
トポロジーを選択した後の詳細なライフサイクルと機能サーフェスについては、[Realtime エージェントガイド](guide.md) を参照してください。
|
||||
+110
-46
@@ -4,111 +4,175 @@ search:
|
||||
---
|
||||
# リリースプロセス / 変更履歴
|
||||
|
||||
このプロジェクトでは、`0.Y.Z` 形式を使用する、semantic versioning をやや修正したバージョニングを採用しています。先頭の `0` は、この SDK がまだ急速に進化していることを示します。各コンポーネントは次のように増分されます。
|
||||
このプロジェクトは、形式 `0.Y.Z` を使用するセマンティックバージョニングを少し修正したものに従います。先頭の `0` は、SDK がまだ急速に進化中であることを示します。各構成要素は次のように増加させます:
|
||||
|
||||
## マイナー (`Y`) バージョン
|
||||
|
||||
ベータとしてマークされていない公開インターフェースに **破壊的変更** がある場合、マイナーバージョン `Y` を上げます。たとえば、`0.0.x` から `0.1.x` への移行には破壊的変更が含まれる可能性があります。
|
||||
ベータとしてマークされていない公開インターフェイスに対する **破壊的変更** の場合、マイナーバージョン `Y` を増やします。たとえば、`0.0.x` から `0.1.x` への移行には破壊的変更が含まれる可能性があります。
|
||||
|
||||
破壊的変更を望まない場合は、プロジェクト内で `0.0.x` バージョンに固定することを推奨します。
|
||||
破壊的変更を避けたい場合は、プロジェクトで `0.0.x` バージョンに固定することをおすすめします。
|
||||
|
||||
## パッチ (`Z`) バージョン
|
||||
|
||||
破壊的ではない変更については `Z` を増やします。
|
||||
非破壊的変更の場合は `Z` を増やします:
|
||||
|
||||
- バグ修正
|
||||
- 新機能
|
||||
- 非公開インターフェースの変更
|
||||
- ベータ機能の更新
|
||||
- バグ修正
|
||||
- 新機能
|
||||
- 非公開インターフェイスの変更
|
||||
- ベータ機能の更新
|
||||
|
||||
## 破壊的変更の変更履歴
|
||||
|
||||
### 0.17.0
|
||||
|
||||
このバージョンでは、サンドボックスのローカルソースの実体化において、ソースパスが `Manifest.extra_path_grants` によってカバーされていない限り、`LocalFile.src` と `LocalDir.src` は実体化時の `base_dir` 内に収められます。`base_dir` は、マニフェストが適用される時点での SDK プロセスの現在の作業ディレクトリです。相対ローカルソースはそのディレクトリから解決され、絶対ローカルソースはすでにそのディレクトリ内にあるか、明示的な許可の配下にある必要があります。これによりローカルアーティファクトの境界に関する問題は解消されますが、そのベースディレクトリ外から信頼済みホストのファイルまたはディレクトリをサンドボックスワークスペースへ意図的にコピーするアプリケーションに影響する可能性があります。
|
||||
|
||||
移行するには、信頼済みホストルートをマニフェストレベルで `SandboxPathGrant` により許可してください。サンドボックスがそれらのファイルを読み取るだけでよい場合は、読み取り専用にすることをおすすめします:
|
||||
|
||||
```python
|
||||
from pathlib import Path
|
||||
|
||||
from agents.sandbox import Manifest, SandboxPathGrant
|
||||
from agents.sandbox.entries import Dir, LocalDir
|
||||
|
||||
# This is an absolute host path outside the SDK process base_dir.
|
||||
TRUSTED_DOCS_ROOT = Path("/opt/my-app/docs")
|
||||
|
||||
manifest = Manifest(
|
||||
extra_path_grants=(
|
||||
# This host root is outside the SDK process base_dir, so the manifest must grant it.
|
||||
SandboxPathGrant(path=str(TRUSTED_DOCS_ROOT), read_only=True),
|
||||
),
|
||||
entries={
|
||||
# No grant is needed for local sources that stay under the SDK process base_dir.
|
||||
"fixtures": LocalDir(src=Path("fixtures"), description="Local test fixtures."),
|
||||
# This entry reads from the granted host root and copies it into the sandbox workspace.
|
||||
"docs": LocalDir(src=TRUSTED_DOCS_ROOT, description="Trusted local documents."),
|
||||
# Dir creates a sandbox workspace directory; it does not read from the host filesystem.
|
||||
"output": Dir(description="Generated artifacts."),
|
||||
},
|
||||
)
|
||||
```
|
||||
|
||||
`extra_path_grants` は信頼済みアプリケーション設定として扱ってください。アプリケーションがそれらのホストパスをすでに承認していない限り、モデル出力やその他の信頼できないマニフェスト入力から許可を設定しないでください。
|
||||
|
||||
### 0.16.0
|
||||
|
||||
このバージョンでは、SDK のデフォルトモデルが `gpt-4.1` ではなく `gpt-5.4-mini` になりました。これは、モデルを明示的に設定していないエージェントと実行に影響します。新しいデフォルトは GPT-5 モデルであるため、暗黙的なデフォルトモデル設定には `reasoning.effort="none"` や `verbosity="low"` などの GPT-5 デフォルトが含まれるようになりました。
|
||||
|
||||
以前のデフォルトモデルの挙動を維持する必要がある場合は、エージェントまたは実行設定でモデルを明示的に設定するか、`OPENAI_DEFAULT_MODEL` 環境変数を設定してください:
|
||||
|
||||
```python
|
||||
agent = Agent(name="Assistant", model="gpt-4.1")
|
||||
```
|
||||
|
||||
主な変更点:
|
||||
|
||||
- `Runner.run`、`Runner.run_sync`、`Runner.run_streamed` は、ターン制限を無効にするために `max_turns=None` を受け取れるようになりました。
|
||||
- サンドボックスワークスペースのハイドレーションは、ローカル、Docker、およびプロバイダーがバックするサンドボックス実装全体で、絶対シンボリックリンクターゲットを含む、アーカイブルートの外部を指すシンボリックリンクを含む tar アーカイブを拒否するようになりました。
|
||||
|
||||
### 0.15.0
|
||||
|
||||
このバージョンでは、モデルの拒否応答は、空のテキスト出力として扱われたり、structured outputs の場合に実行ループが `MaxTurnsExceeded` まで再試行したりするのではなく、`ModelRefusalError` として明示的に表面化されるようになりました。
|
||||
|
||||
これは、以前に拒否のみのモデル応答が `final_output == ""` で完了することを期待していたコードに影響します。例外を発生させずに拒否を処理するには、`model_refusal` 実行エラーハンドラーを提供してください:
|
||||
|
||||
```python
|
||||
result = Runner.run_sync(
|
||||
agent,
|
||||
input,
|
||||
error_handlers={"model_refusal": lambda data: data.error.refusal},
|
||||
)
|
||||
```
|
||||
|
||||
structured-output エージェントの場合、ハンドラーはエージェントの出力スキーマに一致する値を返すことができ、SDK は他の実行エラーハンドラーの最終出力と同様に検証します。
|
||||
|
||||
### 0.14.0
|
||||
|
||||
このマイナーリリースでは **破壊的変更** は導入されませんが、新しい主要なベータ機能領域として Sandbox Agents が追加されています。また、ローカル環境、コンテナ化環境、ホスト環境でそれらを使用するために必要なランタイム、バックエンド、ドキュメントのサポートも含まれています。
|
||||
このマイナーリリースでは、破壊的変更は **導入しません** が、主要な新しいベータ機能領域である Sandbox エージェントに加え、ローカル、コンテナ化、ホスト環境全体でそれらを使用するために必要なランタイム、バックエンド、ドキュメントのサポートが追加されています。
|
||||
|
||||
主なポイント:
|
||||
主な変更点:
|
||||
|
||||
- `SandboxAgent`、`Manifest`、`SandboxRunConfig` を中心とした新しいベータ sandbox runtime surface を追加し、ファイル、ディレクトリ、Git リポジトリ、マウント、スナップショット、再開サポートを備えた永続的で隔離されたワークスペース内でエージェントが動作できるようにしました。
|
||||
- `UnixLocalSandboxClient` と `DockerSandboxClient` により、ローカルおよびコンテナ化された開発向けの sandbox 実行バックエンドを追加しました。さらに、オプションの extra を通じて Blaxel、Cloudflare、Daytona、E2B、Modal、Runloop、Vercel 向けのホスト型プロバイダー統合も追加しました。
|
||||
- 将来の実行で過去の実行から得た学びを再利用できるように sandbox memory support を追加しました。これには progressive disclosure、複数ターンのグルーピング、設定可能な分離境界、S3 ベースのワークフローを含む永続化メモリーの例が含まれます。
|
||||
- より広範なワークスペースおよび再開モデルを追加しました。これには、ローカルおよび合成ワークスペースエントリー、S3 / R2 / GCS / Azure Blob Storage / S3 Files 向けのリモートストレージマウント、ポータブルなスナップショット、`RunState`、`SandboxSessionState`、または保存済みスナップショットを介した再開フローが含まれます。
|
||||
- `examples/sandbox/` 以下に充実した sandbox のコード例とチュートリアルを追加しました。skills、ハンドオフ、メモリー、プロバイダー固有のセットアップ、コードレビュー、dataroom QA、Web サイトのクローン作成などのエンドツーエンドワークフローを用いたコーディングタスクを扱っています。
|
||||
- sandbox 対応のセッション準備、capability binding、状態のシリアライズ、統合トレーシング、prompt cache key のデフォルト、および機微な MCP 出力のより安全な秘匿化を含めて、コアランタイムとトレーシングスタックを拡張しました。
|
||||
- `SandboxAgent`、`Manifest`、`SandboxRunConfig` を中心とした新しいベータのサンドボックスランタイムサーフェスを追加しました。これにより、エージェントはファイル、ディレクトリ、Git リポジトリ、マウント、スナップショット、再開サポートを備えた永続的な隔離ワークスペース内で動作できます。
|
||||
- `UnixLocalSandboxClient` と `DockerSandboxClient` によるローカルおよびコンテナ化開発向けのサンドボックス実行バックエンドに加え、任意の extras を通じた Blaxel、Cloudflare、Daytona、E2B、Modal、Runloop、Vercel 向けのホスト型プロバイダー連携を追加しました。
|
||||
- 将来の実行が以前の実行から得た知見を再利用できるように、サンドボックスメモリサポートを追加しました。段階的開示、複数ターンのグループ化、設定可能な隔離境界、および S3 バックのワークフローを含む永続化メモリのコード例が含まれます。
|
||||
- ローカルおよび合成ワークスペースエントリー、S3/R2/GCS/Azure Blob Storage/S3 Files 向けのリモートストレージマウント、移植可能なスナップショット、`RunState`、`SandboxSessionState`、または保存済みスナップショットによる再開フローを含む、より広範なワークスペースと再開モデルを追加しました。
|
||||
- `examples/sandbox/` 配下に、充実したサンドボックスのコード例とチュートリアルを追加しました。スキル、ハンドオフ、メモリを用いたコーディングタスク、プロバイダー固有のセットアップ、コードレビュー、データルーム QA、Web サイトのクローン作成などのエンドツーエンドのワークフローを扱います。
|
||||
- サンドボックス対応のセッション準備、ケイパビリティバインディング、状態のシリアライズ、統合トレーシング、プロンプトキャッシュキーのデフォルト、より安全な機密 MCP 出力のマスキングにより、コアランタイムとトレーシングスタックを拡張しました。
|
||||
|
||||
### 0.13.0
|
||||
|
||||
このマイナーリリースでは **破壊的変更** は導入されませんが、注目すべき Realtime のデフォルト更新に加えて、新しい MCP 機能とランタイム安定性の修正が含まれています。
|
||||
このマイナーリリースでは、破壊的変更は **導入しません** が、注目すべき Realtime のデフォルト更新に加え、新しい MCP 機能とランタイム安定性の修正が含まれます。
|
||||
|
||||
主なポイント:
|
||||
主な変更点:
|
||||
|
||||
- デフォルトの websocket Realtime モデルが `gpt-realtime-1.5` になり、新しい Realtime エージェント構成では追加設定なしで新しいモデルが使用されるようになりました。
|
||||
- `MCPServer` は `list_resources()`、`list_resource_templates()`、`read_resource()` を公開するようになり、`MCPServerStreamableHttp` は `session_id` を公開するようになったため、streamable HTTP セッションを再接続時やステートレスなワーカー間で再開できるようになりました。
|
||||
- Chat Completions 統合で `should_replay_reasoning_content` による reasoning-content の再生を選択できるようになり、LiteLLM / DeepSeek などのアダプターにおいて、プロバイダー固有の reasoning / tool-call の継続性が向上しました。
|
||||
- `SQLAlchemySession` における同時の最初の書き込み、reasoning の除去後に assistant message ID が孤立した compaction リクエスト、`remove_all_tools()` で MCP / reasoning 項目が残る問題、関数ツールのバッチエグゼキューターにおける競合など、複数のランタイムおよびセッションのエッジケースを修正しました。
|
||||
- デフォルトの WebSocket Realtime モデルは `gpt-realtime-1.5` になりました。そのため、新しい Realtime エージェントのセットアップでは、追加設定なしで新しいモデルが使用されます。
|
||||
- `MCPServer` は `list_resources()`、`list_resource_templates()`、`read_resource()` を公開するようになりました。また、`MCPServerStreamableHttp` は `session_id` を公開するようになったため、ストリーム可能な HTTP セッションを再接続やステートレスワーカーをまたいで再開できます。
|
||||
- Chat Completions 連携は、`should_replay_reasoning_content` によって推論コンテンツの再生をオプトインできるようになりました。これにより、LiteLLM/DeepSeek などのアダプターで、プロバイダー固有の推論 / ツール呼び出しの連続性が向上します。
|
||||
- `SQLAlchemySession` における初回書き込みの同時実行、推論の除去後に孤立した assistant メッセージ ID を持つ圧縮リクエスト、`remove_all_tools()` が MCP/reasoning 項目を残す問題、関数ツールバッチ実行器の競合など、複数のランタイムおよびセッションのエッジケースを修正しました。
|
||||
|
||||
### 0.12.0
|
||||
|
||||
このマイナーリリースでは **破壊的変更** は導入されません。主要な機能追加については [リリースノート](https://github.com/openai/openai-agents-python/releases/tag/v0.12.0) を確認してください。
|
||||
このマイナーリリースでは、破壊的変更は **導入しません**。主要な機能追加については、[リリースノート](https://github.com/openai/openai-agents-python/releases/tag/v0.12.0)を確認してください。
|
||||
|
||||
### 0.11.0
|
||||
|
||||
このマイナーリリースでは **破壊的変更** は導入されません。主要な機能追加については [リリースノート](https://github.com/openai/openai-agents-python/releases/tag/v0.11.0) を確認してください。
|
||||
このマイナーリリースでは、破壊的変更は **導入しません**。主要な機能追加については、[リリースノート](https://github.com/openai/openai-agents-python/releases/tag/v0.11.0)を確認してください。
|
||||
|
||||
### 0.10.0
|
||||
|
||||
このマイナーリリースでは **破壊的変更** は導入されませんが、OpenAI Responses ユーザー向けの重要な新機能領域として Responses API の websocket transport support が含まれています。
|
||||
このマイナーリリースでは、破壊的変更は **導入しません** が、OpenAI Responses ユーザー向けの重要な新機能領域である Responses API の WebSocket トランスポートサポートが含まれます。
|
||||
|
||||
主なポイント:
|
||||
主な変更点:
|
||||
|
||||
- OpenAI Responses モデル向けに websocket transport support を追加しました(オプトイン方式で、既定の transport は引き続き HTTP です)。
|
||||
- 複数ターンの実行にまたがって websocket 対応の共有プロバイダーと `RunConfig` を再利用するための `responses_websocket_session()` ヘルパー / `ResponsesWebSocketSession` を追加しました。
|
||||
- ストリーミング、tools、承認、フォローアップターンを扱う新しい websocket ストリーミングのコード例 (`examples/basic/stream_ws.py`) を追加しました。
|
||||
- OpenAI Responses モデル向けの WebSocket トランスポートサポートを追加しました(オプトインです。HTTP は引き続きデフォルトのトランスポートです)。
|
||||
- 複数ターンの実行全体で共有の WebSocket 対応プロバイダーと `RunConfig` を再利用するための `responses_websocket_session()` ヘルパー / `ResponsesWebSocketSession` を追加しました。
|
||||
- ストリーミング、ツール、承認、フォローアップターンを扱う新しい WebSocket ストリーミングコード例(`examples/basic/stream_ws.py`)を追加しました。
|
||||
|
||||
### 0.9.0
|
||||
|
||||
このバージョンでは、Python 3.9 は 3 か月前に EOL に達したため、サポート対象外となりました。より新しいランタイムバージョンにアップグレードしてください。
|
||||
このバージョンでは、このメジャーバージョンが 3 か月前に EOL に達したため、Python 3.9 はサポートされなくなりました。より新しいランタイムバージョンへアップグレードしてください。
|
||||
|
||||
さらに、`Agent#as_tool()` メソッドから返される値の型ヒントは、`Tool` から `FunctionTool` に絞り込まれました。この変更は通常、破壊的な問題を引き起こすことはありませんが、コードがより広い union type に依存している場合は、利用側でいくつか調整が必要になる可能性があります。
|
||||
さらに、`Agent#as_tool()` メソッドから返される値の型ヒントが、`Tool` から `FunctionTool` へ狭められました。この変更は通常、破壊的な問題を引き起こすことはありませんが、コードがより広い Union 型に依存している場合は、側でいくつか調整が必要になる可能性があります。
|
||||
|
||||
### 0.8.0
|
||||
|
||||
このバージョンでは、ランタイム動作の 2 つの変更により、移行作業が必要になる場合があります。
|
||||
このバージョンでは、2 つのランタイム挙動の変更により、移行作業が必要になる可能性があります:
|
||||
|
||||
- **同期的な** Python callable をラップする関数ツールは、イベントループスレッド上で実行されるのではなく、`asyncio.to_thread(...)` を介してワーカースレッド上で実行されるようになりました。ツールロジックがスレッドローカルな状態やスレッドに紐づくリソースに依存している場合は、非同期ツール実装へ移行するか、ツールコード内でスレッド親和性を明示してください。
|
||||
- ローカル MCP ツールの失敗処理が設定可能になり、デフォルト動作では実行全体を失敗させる代わりに、モデルから見えるエラー出力を返す場合があります。fail-fast の意味論に依存している場合は、`mcp_config={"failure_error_function": None}` を設定してください。サーバーレベルの `failure_error_function` の値はエージェントレベルの設定を上書きするため、明示的なハンドラーを持つ各ローカル MCP サーバーで `failure_error_function=None` を設定してください。
|
||||
- **同期** Python 呼び出し可能オブジェクトをラップする関数ツールは、イベントループスレッドで実行されるのではなく、`asyncio.to_thread(...)` を介してワーカースレッドで実行されるようになりました。ツールのロジックがスレッドローカル状態やスレッドアフィンなリソースに依存している場合は、async ツール実装へ移行するか、ツールコード内でスレッドアフィニティを明示してください。
|
||||
- ローカル MCP ツールの失敗処理は設定可能になり、デフォルトの挙動では実行全体を失敗させる代わりに、モデルから見えるエラー出力を返す場合があります。フェイルファストのセマンティクスに依存している場合は、`mcp_config={"failure_error_function": None}` を設定してください。サーバーレベルの `failure_error_function` 値はエージェントレベルの設定を上書きするため、明示的なハンドラーを持つ各ローカル MCP サーバーで `failure_error_function=None` を設定してください。
|
||||
|
||||
### 0.7.0
|
||||
|
||||
このバージョンでは、既存のアプリケーションに影響する可能性のある動作変更がいくつかあります。
|
||||
このバージョンでは、既存のアプリケーションに影響する可能性がある挙動の変更がいくつかありました:
|
||||
|
||||
- ネストされたハンドオフ履歴は現在 **オプトイン** です(デフォルトでは無効)。v0.6.x のデフォルトのネスト動作に依存していた場合は、明示的に `RunConfig(nest_handoff_history=True)` を設定してください。
|
||||
- `gpt-5.1` / `gpt-5.2` に対するデフォルトの `reasoning.effort` は `"none"` に変更されました(SDK デフォルトで設定されていた従来の `"low"` から変更)。プロンプトや品質 / コストプロファイルが `"low"` に依存していた場合は、`model_settings` で明示的に設定してください。
|
||||
- ネストされたハンドオフ履歴は **オプトイン** になりました(デフォルトでは無効)。v0.6.x のデフォルトのネスト挙動に依存していた場合は、`RunConfig(nest_handoff_history=True)` を明示的に設定してください。
|
||||
- `gpt-5.1` / `gpt-5.2` のデフォルトの `reasoning.effort` は、`"none"` に変更されました(SDK デフォルトで設定されていた以前のデフォルト `"low"` からの変更です)。プロンプトや品質 / コストのプロファイルが `"low"` に依存していた場合は、`model_settings` で明示的に設定してください。
|
||||
|
||||
### 0.6.0
|
||||
|
||||
このバージョンでは、デフォルトのハンドオフ履歴は、生の user / assistant ターンを公開する代わりに、単一の assistant メッセージにまとめられるようになり、下流エージェントに簡潔で予測可能な要約を提供します。
|
||||
- 既存の単一メッセージのハンドオフトランスクリプトは、デフォルトで `<CONVERSATION HISTORY>` ブロックの前に "For context, here is the conversation so far between the user and the previous agent:" で始まるようになり、下流エージェントが明確にラベル付けされた要約を受け取れるようになりました。
|
||||
このバージョンでは、デフォルトのハンドオフ履歴は、生のユーザー / アシスタントターンを公開するのではなく、単一の assistant メッセージにまとめられるようになり、後続のエージェントに簡潔で予測可能な要約を提供します
|
||||
- 既存の単一メッセージのハンドオフトランスクリプトは、デフォルトで `<CONVERSATION HISTORY>` ブロックの前に "For context, here is the conversation so far between the user and the previous agent:" で始まるようになったため、後続のエージェントは明確なラベル付きの要約を受け取れます
|
||||
|
||||
### 0.5.0
|
||||
|
||||
このバージョンでは、目に見える破壊的変更は導入されませんが、新機能と内部的な重要更新がいくつか含まれています。
|
||||
このバージョンでは、目に見える破壊的変更は導入されませんが、新機能と内部のいくつかの重要な更新が含まれます:
|
||||
|
||||
- `RealtimeRunner` が [SIP protocol connections](https://platform.openai.com/docs/guides/realtime-sip) を扱えるようサポートを追加しました
|
||||
- Python 3.14 互換性のために `Runner#run_sync` の内部ロジックを大幅に改訂しました
|
||||
- `RealtimeRunner` が [SIP プロトコル接続](https://platform.openai.com/docs/guides/realtime-sip)を処理するためのサポートを追加しました
|
||||
- Python 3.14 互換性のために、`Runner#run_sync` の内部ロジックを大幅に改訂しました
|
||||
|
||||
### 0.4.0
|
||||
|
||||
このバージョンでは、[openai](https://pypi.org/project/openai/) パッケージの v1.x 系はサポート対象外となりました。この SDK と合わせて openai v2.x を使用してください。
|
||||
このバージョンでは、[openai](https://pypi.org/project/openai/) パッケージの v1.x バージョンはサポートされなくなりました。この SDK とともに openai v2.x を使用してください。
|
||||
|
||||
### 0.3.0
|
||||
|
||||
このバージョンでは、Realtime API のサポートが gpt-realtime モデルおよびその API インターフェース( GA 版)に移行します。
|
||||
このバージョンでは、Realtime API サポートは gpt-realtime モデルとその API インターフェイス(GA バージョン)へ移行します。
|
||||
|
||||
### 0.2.0
|
||||
|
||||
このバージョンでは、これまで引数として `Agent` を受け取っていたいくつかの箇所が、代わりに `AgentBase` を受け取るようになりました。たとえば、MCP サーバー内の `list_tools()` 呼び出しです。これは純粋に型に関する変更であり、引き続き `Agent` オブジェクトを受け取ります。更新するには、`Agent` を `AgentBase` に置き換えて型エラーを修正してください。
|
||||
このバージョンでは、以前は `Agent` を引数として受け取っていたいくつかの箇所が、代わりに `AgentBase` を引数として受け取るようになりました。たとえば、MCP サーバーの `list_tools()` 呼び出しです。これは純粋に型付け上の変更であり、引き続き `Agent` オブジェクトを受け取ります。更新するには、`Agent` を `AgentBase` に置き換えて型エラーを修正するだけです。
|
||||
|
||||
### 0.1.0
|
||||
|
||||
このバージョンでは、[`MCPServer.list_tools()`][agents.mcp.server.MCPServer] に 2 つの新しい params が追加されています: `run_context` と `agent` です。`MCPServer` をサブクラス化しているすべてのクラスに、これらの params を追加する必要があります。
|
||||
このバージョンでは、[`MCPServer.list_tools()`][agents.mcp.server.MCPServer] に `run_context` と `agent` という 2 つの新しいパラメーターが追加されました。`MCPServer` をサブクラス化しているすべてのクラスに、これらのパラメーターを追加する必要があります。
|
||||
+3
-3
@@ -4,7 +4,7 @@ search:
|
||||
---
|
||||
# REPL ユーティリティ
|
||||
|
||||
この SDK は、ターミナル上でエージェントの挙動を素早く対話的にテストできる `run_demo_loop` を提供します。
|
||||
SDK は、ターミナルでエージェントの動作を直接すばやく対話的にテストするための `run_demo_loop` を提供します。
|
||||
|
||||
|
||||
```python
|
||||
@@ -19,6 +19,6 @@ if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
`run_demo_loop` はループでユーザー入力を促し、ターン間で会話履歴を保持します。デフォルトでは、生成されたモデル出力をストリーミングします。上記の例を実行すると、`run_demo_loop` は対話型のチャットセッションを開始します。入力を継続的に求め、これまでの会話履歴全体を保持することで(エージェントが何について話したかを把握できます)、生成と同時にエージェントの応答をリアルタイムで自動的にストリーミングします。
|
||||
`run_demo_loop` はループ内でユーザー入力を求め、ターン間の会話履歴を保持します。デフォルトでは、生成されるモデル出力をストリーミングします。上記の例を実行すると、run_demo_loop は対話型チャットセッションを開始します。入力を継続的に求め、ターン間の会話履歴全体を記憶し(そのためエージェントはこれまでに話し合われた内容を把握できます)、エージェントの応答が生成されるとリアルタイムで自動的にストリーミングします。
|
||||
|
||||
このチャットセッションを終了するには、`quit` または `exit` と入力して Enter を押すか、`Ctrl-D` のキーボードショートカットを使用します。
|
||||
このチャットセッションを終了するには、単に `quit` または `exit` と入力して Enter キーを押すか、`Ctrl-D` キーボードショートカットを使用します。
|
||||
+67
-67
@@ -4,95 +4,95 @@ search:
|
||||
---
|
||||
# 実行結果
|
||||
|
||||
`Runner.run` メソッドを呼び出すと、次の 2 種類の結果タイプのいずれかを受け取ります。
|
||||
`Runner.run` メソッドを呼び出すと、次の 2 つの実行結果型のいずれかを受け取ります。
|
||||
|
||||
- `Runner.run(...)` または `Runner.run_sync(...)` からの [`RunResult`][agents.result.RunResult]
|
||||
- `Runner.run_streamed(...)` からの [`RunResultStreaming`][agents.result.RunResultStreaming]
|
||||
|
||||
どちらも [`RunResultBase`][agents.result.RunResultBase] を継承しており、`final_output`、`new_items`、`last_agent`、`raw_responses`、`to_state()` などの共通の結果サーフェスを公開します。
|
||||
どちらも [`RunResultBase`][agents.result.RunResultBase] を継承しており、`final_output`、`new_items`、`last_agent`、`raw_responses`、`to_state()` などの共通の実行結果サーフェスを公開します。
|
||||
|
||||
`RunResultStreaming` には、[`stream_events()`][agents.result.RunResultStreaming.stream_events]、[`current_agent`][agents.result.RunResultStreaming.current_agent]、[`is_complete`][agents.result.RunResultStreaming.is_complete]、[`cancel(...)`][agents.result.RunResultStreaming.cancel] などのストリーミング固有の制御が追加されています。
|
||||
`RunResultStreaming` は、[`stream_events()`][agents.result.RunResultStreaming.stream_events]、[`current_agent`][agents.result.RunResultStreaming.current_agent]、[`is_complete`][agents.result.RunResultStreaming.is_complete]、[`cancel(...)`][agents.result.RunResultStreaming.cancel] など、ストリーミング固有の制御機能を追加します。
|
||||
|
||||
## 適切な結果サーフェスの選択
|
||||
## 適切な実行結果サーフェスの選択
|
||||
|
||||
ほとんどのアプリケーションで必要なのは、いくつかの結果プロパティまたはヘルパーだけです。
|
||||
ほとんどのアプリケーションでは、いくつかの実行結果プロパティまたはヘルパーだけで十分です。
|
||||
|
||||
| 必要なもの | 使用先 |
|
||||
| 必要なもの... | 使用するもの |
|
||||
| --- | --- |
|
||||
| ユーザーに表示する最終回答 | `final_output` |
|
||||
| ローカルの完全なトランスクリプトを含む、再生可能な次ターン入力リスト | `to_input_list()` |
|
||||
| エージェント、ツール、ハンドオフ、承認メタデータを含むリッチな実行アイテム | `new_items` |
|
||||
| 完全なローカルトランスクリプトを含む、リプレイ可能な次ターン入力リスト | `to_input_list()` |
|
||||
| エージェント、ツール、ハンドオフ、承認メタデータを含む詳細な実行項目 | `new_items` |
|
||||
| 通常、次のユーザーターンを処理すべきエージェント | `last_agent` |
|
||||
| `previous_response_id` を用いた OpenAI Responses API チェーン | `last_response_id` |
|
||||
| 保留中の承認と再開可能なスナップショット | `interruptions` と `to_state()` |
|
||||
| `previous_response_id` による OpenAI Responses API チェーン | `last_response_id` |
|
||||
| 保留中の承認と再開可能なスナップショット | `interruptions` and `to_state()` |
|
||||
| 現在のネストされた `Agent.as_tool()` 呼び出しに関するメタデータ | `agent_tool_invocation` |
|
||||
| 生のモデル呼び出しまたはガードレール診断 | `raw_responses` とガードレール結果配列 |
|
||||
| raw モデル呼び出しまたはガードレール診断 | `raw_responses` and the guardrail result arrays |
|
||||
|
||||
## 最終出力
|
||||
|
||||
[`final_output`][agents.result.RunResultBase.final_output] プロパティには、最後に実行されたエージェントの最終出力が含まれます。これは次のいずれかです。
|
||||
|
||||
- 最後のエージェントに `output_type` が定義されていない場合は `str`
|
||||
- 最後のエージェントに出力型が定義されている場合は `last_agent.output_type` 型のオブジェクト
|
||||
- 承認による割り込みで一時停止した場合など、最終出力が生成される前に実行が停止した場合は `None`
|
||||
- 最後のエージェントに `output_type` が定義されていなかった場合は `str`
|
||||
- 最後のエージェントに出力型が定義されていた場合は `last_agent.output_type` 型のオブジェクト
|
||||
- 承認中断で一時停止した場合など、最終出力が生成される前に実行が停止した場合は `None`
|
||||
|
||||
!!! note
|
||||
|
||||
`final_output` は `Any` 型です。ハンドオフにより実行を完了するエージェントが変わる可能性があるため、SDK は取り得る出力型の完全な集合を静的に把握できません。
|
||||
`final_output` は `Any` として型付けされています。ハンドオフによって、どのエージェントが実行を終了するかが変わる可能性があるため、SDK は考えられる出力型の全体集合を静的に把握できません。
|
||||
|
||||
ストリーミングモードでは、ストリームの処理が完了するまで `final_output` は `None` のままです。イベントごとの流れは [Streaming](streaming.md) を参照してください。
|
||||
ストリーミングモードでは、ストリームの処理が完了するまで `final_output` は `None` のままです。イベントごとのフローについては、[ストリーミング](streaming.md)を参照してください。
|
||||
|
||||
## 入力、次ターン履歴、new items
|
||||
## 入力、次ターン履歴、新規項目
|
||||
|
||||
これらのサーフェスは、それぞれ異なる問いに答えます。
|
||||
これらのサーフェスは、それぞれ異なる問いに対応します。
|
||||
|
||||
| プロパティまたはヘルパー | 含まれる内容 | 最適な用途 |
|
||||
| --- | --- | --- |
|
||||
| [`input`][agents.result.RunResultBase.input] | この実行セグメントのベース入力。ハンドオフ入力フィルターが履歴を書き換えた場合、実行が継続したフィルター後の入力が反映されます。 | この実行が実際に入力として何を使ったかの監査 |
|
||||
| [`to_input_list()`][agents.result.RunResultBase.to_input_list] | 実行の入力アイテムビュー。既定の `mode="preserve_all"` は `new_items` から変換された完全な履歴を保持し、`mode="normalized"` はハンドオフフィルタリングでモデル履歴が書き換えられた際に正規の継続入力を優先します。 | 手動チャットループ、クライアント管理の会話状態、プレーンアイテム履歴の確認 |
|
||||
| [`new_items`][agents.result.RunResultBase.new_items] | エージェント、ツール、ハンドオフ、承認メタデータを持つリッチな [`RunItem`][agents.items.RunItem] ラッパー。 | ログ、UI、監査、デバッグ |
|
||||
| [`raw_responses`][agents.result.RunResultBase.raw_responses] | 実行内の各モデル呼び出しから得られる生の [`ModelResponse`][agents.items.ModelResponse] オブジェクト。 | プロバイダーレベルの診断や生レスポンスの確認 |
|
||||
| [`input`][agents.result.RunResultBase.input] | この実行セグメントの基本入力です。ハンドオフ入力フィルターが履歴を書き換えた場合、実行が継続されたフィルター済み入力がここに反映されます。 | この実行が実際に入力として使用した内容の監査 |
|
||||
| [`to_input_list()`][agents.result.RunResultBase.to_input_list] | 実行を入力項目として見たビューです。デフォルトの `mode="preserve_all"` は、`new_items` から変換された完全な履歴を保持します。`mode="normalized"` は、ハンドオフフィルタリングによってモデル履歴が書き換えられた場合に、正規の継続入力を優先します。 | 手動のチャットループ、クライアント管理の会話状態、プレーンな項目履歴の確認 |
|
||||
| [`new_items`][agents.result.RunResultBase.new_items] | エージェント、ツール、ハンドオフ、承認メタデータを含む詳細な [`RunItem`][agents.items.RunItem] ラッパーです。 | ログ、UI、監査、デバッグ |
|
||||
| [`raw_responses`][agents.result.RunResultBase.raw_responses] | 実行内の各モデル呼び出しからの raw [`ModelResponse`][agents.items.ModelResponse] オブジェクトです。 | プロバイダーレベルの診断または raw レスポンスの確認 |
|
||||
|
||||
実運用では次のとおりです。
|
||||
実際には、次のように使い分けます。
|
||||
|
||||
- 実行のプレーンな入力アイテムビューが必要な場合は `to_input_list()` を使います。
|
||||
- ハンドオフフィルタリングやネストされたハンドオフ履歴書き換え後、次の `Runner.run(..., input=...)` 呼び出し向けの正規ローカル入力が必要な場合は `to_input_list(mode="normalized")` を使います。
|
||||
- SDK に履歴の読み書きを任せたい場合は [`session=...`](sessions/index.md) を使います。
|
||||
- `conversation_id` や `previous_response_id` による OpenAI のサーバー管理状態を使っている場合、通常は `to_input_list()` を再送せず、新しいユーザー入力のみを渡して保存済み ID を再利用します。
|
||||
- ログ、UI、監査のために完全な変換済み履歴が必要な場合は、既定の `to_input_list()` モードまたは `new_items` を使います。
|
||||
- 実行のプレーンな入力項目ビューが必要な場合は、`to_input_list()` を使用します。
|
||||
- ハンドオフフィルタリングまたはネストされたハンドオフ履歴の書き換え後に、次の `Runner.run(..., input=...)` 呼び出しに渡す正規のローカル入力が必要な場合は、`to_input_list(mode="normalized")` を使用します。
|
||||
- SDK に履歴の読み込みと保存を任せたい場合は、[`session=...`](sessions/index.md) を使用します。
|
||||
- `conversation_id` または `previous_response_id` を使って OpenAI のサーバー管理状態を使用している場合、通常は `to_input_list()` を再送信する代わりに、新しいユーザー入力のみを渡して保存済み ID を再利用します。
|
||||
- ログ、UI、監査向けに完全な変換済み履歴が必要な場合は、デフォルトの `to_input_list()` モードまたは `new_items` を使用します。
|
||||
|
||||
JavaScript SDK と異なり、Python はモデル形状の差分のみを表す独立した `output` プロパティを公開しません。SDK メタデータが必要なら `new_items` を使い、生のモデルペイロードが必要なら `raw_responses` を確認してください。
|
||||
JavaScript SDK とは異なり、Python ではモデル形式の差分のみを表す個別の `output` プロパティは公開されません。SDK メタデータが必要な場合は `new_items` を使用し、raw モデルペイロードが必要な場合は `raw_responses` を確認してください。
|
||||
|
||||
コンピュータツールのリプレイは、生の Responses ペイロード形状に従います。プレビュー版モデルの `computer_call` アイテムは単一の `action` を保持し、`gpt-5.4` のコンピュータ呼び出しはバッチ化された `actions[]` を保持できます。[`to_input_list()`][agents.result.RunResultBase.to_input_list] と [`RunState`][agents.run_state.RunState] は、モデルが生成した形状をそのまま保持するため、手動リプレイ、一時停止/再開フロー、保存済みトランスクリプトはプレビュー版と GA の両方のコンピュータツール呼び出しで継続して機能します。ローカルの実行結果は引き続き `new_items` 内で `computer_call_output` アイテムとして現れます。
|
||||
コンピュータツールのリプレイは、raw Responses ペイロードの形状に従います。プレビューモデルの `computer_call` 項目は単一の `action` を保持しますが、`gpt-5.5` のコンピュータ呼び出しではバッチ化された `actions[]` を保持できます。[`to_input_list()`][agents.result.RunResultBase.to_input_list] と [`RunState`][agents.run_state.RunState] は、モデルが生成した形状をそのまま保持するため、手動リプレイ、一時停止/再開フロー、保存済みトランスクリプトは、プレビュー版と GA 版の両方のコンピュータツール呼び出しで引き続き機能します。ローカル実行結果は引き続き `new_items` 内の `computer_call_output` 項目として表示されます。
|
||||
|
||||
### New items
|
||||
### 新規項目
|
||||
|
||||
[`new_items`][agents.result.RunResultBase.new_items] は、実行中に何が起きたかを最もリッチに把握できるビューです。一般的なアイテムタイプは次のとおりです。
|
||||
[`new_items`][agents.result.RunResultBase.new_items] は、実行中に何が起きたかを最も詳細に確認できるビューです。一般的な項目型は次のとおりです。
|
||||
|
||||
- アシスタントメッセージ用の [`MessageOutputItem`][agents.items.MessageOutputItem]
|
||||
- 推論アイテム用の [`ReasoningItem`][agents.items.ReasoningItem]
|
||||
- Responses ツール検索リクエストおよび読み込まれたツール検索結果用の [`ToolSearchCallItem`][agents.items.ToolSearchCallItem] と [`ToolSearchOutputItem`][agents.items.ToolSearchOutputItem]
|
||||
- ツール呼び出しとその結果用の [`ToolCallItem`][agents.items.ToolCallItem] と [`ToolCallOutputItem`][agents.items.ToolCallOutputItem]
|
||||
- 推論項目用の [`ReasoningItem`][agents.items.ReasoningItem]
|
||||
- Responses のツール検索リクエストと、ロードされたツール検索の実行結果用の [`ToolSearchCallItem`][agents.items.ToolSearchCallItem] および [`ToolSearchOutputItem`][agents.items.ToolSearchOutputItem]
|
||||
- ツール呼び出しとその実行結果用の [`ToolCallItem`][agents.items.ToolCallItem] および [`ToolCallOutputItem`][agents.items.ToolCallOutputItem]
|
||||
- 承認待ちで一時停止したツール呼び出し用の [`ToolApprovalItem`][agents.items.ToolApprovalItem]
|
||||
- ハンドオフ要求と完了した転送用の [`HandoffCallItem`][agents.items.HandoffCallItem] と [`HandoffOutputItem`][agents.items.HandoffOutputItem]
|
||||
- ハンドオフリクエストと完了済みの引き継ぎ用の [`HandoffCallItem`][agents.items.HandoffCallItem] および [`HandoffOutputItem`][agents.items.HandoffOutputItem]
|
||||
|
||||
エージェントとの関連付け、ツール出力、ハンドオフ境界、承認境界が必要な場合は、`to_input_list()` より `new_items` を選んでください。
|
||||
エージェントの関連付け、ツール出力、ハンドオフの境界、承認の境界が必要な場合は、常に `to_input_list()` よりも `new_items` を選択してください。
|
||||
|
||||
ホストされたツール検索を使う場合、モデルが出力した検索リクエストは `ToolSearchCallItem.raw_item` を、当該ターンでどの名前空間・関数・ホストされた MCP サーバーが読み込まれたかは `ToolSearchOutputItem.raw_item` を確認してください。
|
||||
ホスト型ツール検索を使用する場合、モデルが生成した検索リクエストを確認するには `ToolSearchCallItem.raw_item` を確認し、そのターンでどの名前空間、関数、またはホスト型 MCP サーバーがロードされたかを確認するには `ToolSearchOutputItem.raw_item` を確認してください。
|
||||
|
||||
## 会話の継続または再開
|
||||
|
||||
### 次ターンのエージェント
|
||||
|
||||
[`last_agent`][agents.result.RunResultBase.last_agent] には、最後に実行されたエージェントが含まれます。これはハンドオフ後の次のユーザーターンで再利用するエージェントとして最適なことがよくあります。
|
||||
[`last_agent`][agents.result.RunResultBase.last_agent] には、最後に実行されたエージェントが含まれます。これは多くの場合、ハンドオフ後の次のユーザーターンで再利用するのに最適なエージェントです。
|
||||
|
||||
ストリーミングモードでは、[`RunResultStreaming.current_agent`][agents.result.RunResultStreaming.current_agent] は実行進行に応じて更新されるため、ストリーム完了前にハンドオフを観察できます。
|
||||
ストリーミングモードでは、[`RunResultStreaming.current_agent`][agents.result.RunResultStreaming.current_agent] が実行の進行に合わせて更新されるため、ストリームが終了する前にハンドオフを観察できます。
|
||||
|
||||
### 割り込みと実行状態
|
||||
### 中断と実行状態
|
||||
|
||||
ツールに承認が必要な場合、保留中の承認は [`RunResult.interruptions`][agents.result.RunResult.interruptions] または [`RunResultStreaming.interruptions`][agents.result.RunResultStreaming.interruptions] で公開されます。これには、直接ツールで発生した承認、ハンドオフ後に到達したツールで発生した承認、ネストされた [`Agent.as_tool()`][agents.agent.Agent.as_tool] 実行で発生した承認が含まれる場合があります。
|
||||
ツールに承認が必要な場合、保留中の承認は [`RunResult.interruptions`][agents.result.RunResult.interruptions] または [`RunResultStreaming.interruptions`][agents.result.RunResultStreaming.interruptions] で公開されます。これには、直接呼び出されたツール、ハンドオフ後に到達したツール、またはネストされた [`Agent.as_tool()`][agents.agent.Agent.as_tool] 実行によって発生した承認が含まれることがあります。
|
||||
|
||||
[`to_state()`][agents.result.RunResult.to_state] を呼び出して再開可能な [`RunState`][agents.run_state.RunState] を取得し、保留中アイテムを承認または拒否してから、`Runner.run(...)` または `Runner.run_streamed(...)` で再開します。
|
||||
[`to_state()`][agents.result.RunResult.to_state] を呼び出して、再開可能な [`RunState`][agents.run_state.RunState] を取得し、保留中の項目を承認または拒否してから、`Runner.run(...)` または `Runner.run_streamed(...)` で再開します。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
@@ -107,17 +107,17 @@ if result.interruptions:
|
||||
result = await Runner.run(agent, state)
|
||||
```
|
||||
|
||||
ストリーミング実行では、まず [`stream_events()`][agents.result.RunResultStreaming.stream_events] の消費を完了し、その後 `result.interruptions` を確認して `result.to_state()` から再開してください。承認フロー全体は [Human-in-the-loop](human_in_the_loop.md) を参照してください。
|
||||
ストリーミング実行では、まず [`stream_events()`][agents.result.RunResultStreaming.stream_events] の消費を完了してから `result.interruptions` を確認し、`result.to_state()` から再開してください。承認フロー全体については、[ヒューマンインザループ](human_in_the_loop.md)を参照してください。
|
||||
|
||||
### サーバー管理の継続
|
||||
|
||||
[`last_response_id`][agents.result.RunResultBase.last_response_id] は、この実行における最新のモデルレスポンス ID です。OpenAI Responses API チェーンを継続したい場合は、次ターンでこれを `previous_response_id` として渡します。
|
||||
[`last_response_id`][agents.result.RunResultBase.last_response_id] は、実行から得られた最新のモデルレスポンス ID です。OpenAI Responses API チェーンを継続したい場合は、次のターンで `previous_response_id` として渡してください。
|
||||
|
||||
すでに `to_input_list()`、`session`、または `conversation_id` で会話を継続している場合、通常は `last_response_id` は不要です。マルチステップ実行のすべてのモデルレスポンスが必要な場合は、代わりに `raw_responses` を確認してください。
|
||||
すでに `to_input_list()`、`session`、または `conversation_id` で会話を継続している場合、通常は `last_response_id` は不要です。複数ステップの実行におけるすべてのモデルレスポンスが必要な場合は、代わりに `raw_responses` を確認してください。
|
||||
|
||||
## Agent-as-tool メタデータ
|
||||
## ツールとしてのエージェントのメタデータ
|
||||
|
||||
結果がネストされた [`Agent.as_tool()`][agents.agent.Agent.as_tool] 実行から来ている場合、[`agent_tool_invocation`][agents.result.RunResultBase.agent_tool_invocation] は外側ツール呼び出しの不変メタデータを公開します。
|
||||
実行結果がネストされた [`Agent.as_tool()`][agents.agent.Agent.as_tool] 実行に由来する場合、[`agent_tool_invocation`][agents.result.RunResultBase.agent_tool_invocation] は外側のツール呼び出しに関する変更不可のメタデータを公開します。
|
||||
|
||||
- `tool_name`
|
||||
- `tool_call_id`
|
||||
@@ -125,41 +125,41 @@ if result.interruptions:
|
||||
|
||||
通常のトップレベル実行では、`agent_tool_invocation` は `None` です。
|
||||
|
||||
これは特に `custom_output_extractor` 内で有用で、ネスト結果を後処理する際に外側のツール名、呼び出し ID、または生の引数が必要になることがあります。周辺の `Agent.as_tool()` パターンは [Tools](tools.md) を参照してください。
|
||||
これは、ネストされた実行結果を後処理する際に、外側のツール名、呼び出し ID、または生の引数が必要になることがある `custom_output_extractor` 内で特に便利です。関連する `Agent.as_tool()` パターンについては、[ツール](tools.md)を参照してください。
|
||||
|
||||
そのネスト実行のパース済み structured outputs 入力も必要な場合は、`context_wrapper.tool_input` を読んでください。これは [`RunState`][agents.run_state.RunState] がネストツール入力向けに汎用的にシリアライズするフィールドであり、`agent_tool_invocation` は現在のネスト呼び出し向けのライブ結果アクセサです。
|
||||
そのネストされた実行のパース済み構造化入力も必要な場合は、`context_wrapper.tool_input` を読み取ってください。これは、ネストされたツール入力に対して [`RunState`][agents.run_state.RunState] が汎用的にシリアライズするフィールドです。一方、`agent_tool_invocation` は、現在のネストされた呼び出しに対するライブの実行結果アクセサーです。
|
||||
|
||||
## ストリーミングライフサイクルと診断
|
||||
## ストリーミングのライフサイクルと診断
|
||||
|
||||
[`RunResultStreaming`][agents.result.RunResultStreaming] は上記と同じ結果サーフェスを継承しますが、ストリーミング固有の制御を追加します。
|
||||
[`RunResultStreaming`][agents.result.RunResultStreaming] は、上記と同じ実行結果サーフェスを継承しますが、ストリーミング固有の制御機能を追加します。
|
||||
|
||||
- セマンティックなストリームイベントを消費する [`stream_events()`][agents.result.RunResultStreaming.stream_events]
|
||||
- 実行途中のアクティブエージェントを追跡する [`current_agent`][agents.result.RunResultStreaming.current_agent]
|
||||
- ストリーミング実行が完全に終了したかを確認する [`is_complete`][agents.result.RunResultStreaming.is_complete]
|
||||
- 実行を即時または現在ターン後に停止する [`cancel(...)`][agents.result.RunResultStreaming.cancel]
|
||||
- セマンティックなストリームイベントを消費するための [`stream_events()`][agents.result.RunResultStreaming.stream_events]
|
||||
- 実行途中でアクティブなエージェントを追跡するための [`current_agent`][agents.result.RunResultStreaming.current_agent]
|
||||
- ストリーミング実行が完全に終了したかどうかを確認するための [`is_complete`][agents.result.RunResultStreaming.is_complete]
|
||||
- 実行を即時または現在のターン後に停止するための [`cancel(...)`][agents.result.RunResultStreaming.cancel]
|
||||
|
||||
非同期イテレーターが終了するまで `stream_events()` を消費し続けてください。ストリーミング実行はそのイテレーターが終わるまで完了しません。また、`final_output`、`interruptions`、`raw_responses`、セッション永続化の副作用などの要約プロパティは、最後に見えるトークン到着後も確定中である可能性があります。
|
||||
非同期イテレーターが終了するまで `stream_events()` を消費し続けてください。そのイテレーターが終了するまで、ストリーミング実行は完了していません。また、`final_output`、`interruptions`、`raw_responses` などの要約プロパティや、セッション永続化の副作用は、目に見える最後のトークンが到着した後もまだ確定中の場合があります。
|
||||
|
||||
`cancel()` を呼び出した場合も、キャンセルとクリーンアップを正しく完了させるために `stream_events()` の消費を続けてください。
|
||||
`cancel()` を呼び出した場合は、キャンセルとクリーンアップが正しく完了できるように、`stream_events()` を消費し続けてください。
|
||||
|
||||
Python は、ストリーミング専用の `completed` promise や `error` プロパティを別途公開しません。終端のストリーミング失敗は `stream_events()` からの例外送出として表面化し、`is_complete` は実行が終端状態に達したかどうかを反映します。
|
||||
Python では、ストリーミング用の個別の `completed` プロミスや `error` プロパティは公開されません。終端的なストリーミング失敗は `stream_events()` から例外が送出されることで表面化し、`is_complete` は実行が終端状態に到達したかどうかを反映します。
|
||||
|
||||
### Raw responses
|
||||
### raw レスポンス
|
||||
|
||||
[`raw_responses`][agents.result.RunResultBase.raw_responses] には、実行中に収集された生のモデルレスポンスが含まれます。マルチステップ実行では、たとえばハンドオフやモデル/ツール/モデルの反復サイクルをまたいで、複数のレスポンスが生成されることがあります。
|
||||
[`raw_responses`][agents.result.RunResultBase.raw_responses] には、実行中に収集された raw モデルレスポンスが含まれます。複数ステップの実行では、ハンドオフをまたいだり、モデル/ツール/モデルのサイクルが繰り返されたりする場合など、複数のレスポンスが生成されることがあります。
|
||||
|
||||
[`last_response_id`][agents.result.RunResultBase.last_response_id] は、`raw_responses` の最後のエントリの ID にすぎません。
|
||||
|
||||
### ガードレール結果
|
||||
### ガードレールの実行結果
|
||||
|
||||
エージェントレベルのガードレールは [`input_guardrail_results`][agents.result.RunResultBase.input_guardrail_results] と [`output_guardrail_results`][agents.result.RunResultBase.output_guardrail_results] として公開されます。
|
||||
エージェントレベルのガードレールは、[`input_guardrail_results`][agents.result.RunResultBase.input_guardrail_results] および [`output_guardrail_results`][agents.result.RunResultBase.output_guardrail_results] として公開されます。
|
||||
|
||||
ツールのガードレールは、[`tool_input_guardrail_results`][agents.result.RunResultBase.tool_input_guardrail_results] と [`tool_output_guardrail_results`][agents.result.RunResultBase.tool_output_guardrail_results] として別途公開されます。
|
||||
ツールガードレールは、[`tool_input_guardrail_results`][agents.result.RunResultBase.tool_input_guardrail_results] および [`tool_output_guardrail_results`][agents.result.RunResultBase.tool_output_guardrail_results] として個別に公開されます。
|
||||
|
||||
これらの配列は実行全体で蓄積されるため、判定のログ化、追加ガードレールメタデータの保存、実行がブロックされた理由のデバッグに有用です。
|
||||
これらの配列は実行全体で蓄積されるため、判定のログ記録、追加のガードレールメタデータの保存、または実行がブロックされた理由のデバッグに役立ちます。
|
||||
|
||||
### コンテキストと使用量
|
||||
|
||||
[`context_wrapper`][agents.result.RunResultBase.context_wrapper] は、承認、使用量、ネストされた `tool_input` などの SDK 管理ランタイムメタデータとともに、アプリコンテキストを公開します。
|
||||
[`context_wrapper`][agents.result.RunResultBase.context_wrapper] は、アプリのコンテキストと、承認、使用量、ネストされた `tool_input` など SDK が管理するランタイムメタデータを公開します。
|
||||
|
||||
使用量は `context_wrapper.usage` で追跡されます。ストリーミング実行では、ストリーム最終チャンクの処理が終わるまで使用量合計が遅延する場合があります。ラッパーの完全な形状と永続化時の注意点は [Context management](context.md) を参照してください。
|
||||
使用量は `context_wrapper.usage` で追跡されます。ストリーミング実行では、ストリームの最終チャンクが処理されるまで、使用量の合計値が遅れて反映される場合があります。ラッパーの完全な形状と永続化に関する注意事項については、[コンテキスト管理](context.md)を参照してください。
|
||||
+227
-141
@@ -4,11 +4,11 @@ search:
|
||||
---
|
||||
# エージェントの実行
|
||||
|
||||
エージェントは [`Runner`][agents.run.Runner] クラス経由で実行できます。選択肢は 3 つあります。
|
||||
エージェントは [`Runner`][agents.run.Runner] クラスを通じて実行できます。選択肢は 3 つあります。
|
||||
|
||||
1. [`Runner.run()`][agents.run.Runner.run]。非同期で実行され、[`RunResult`][agents.result.RunResult] を返します。
|
||||
2. [`Runner.run_sync()`][agents.run.Runner.run_sync]。同期メソッドで、内部では `.run()` を実行するだけです。
|
||||
3. [`Runner.run_streamed()`][agents.run.Runner.run_streamed]。非同期で実行され、[`RunResultStreaming`][agents.result.RunResultStreaming] を返します。ストリーミングモードで LLM を呼び出し、受信したイベントをそのままストリーミングします。
|
||||
1. [`Runner.run()`][agents.run.Runner.run] は非同期で実行され、[`RunResult`][agents.result.RunResult] を返します。
|
||||
2. [`Runner.run_sync()`][agents.run.Runner.run_sync] は同期メソッドで、内部では単に `.run()` を実行します。
|
||||
3. [`Runner.run_streamed()`][agents.run.Runner.run_streamed] は非同期で実行され、[`RunResultStreaming`][agents.result.RunResultStreaming] を返します。これはストリーミングモードで LLM を呼び出し、受信したイベントをそのままストリーミングします。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
@@ -23,46 +23,46 @@ async def main():
|
||||
# Infinite loop's dance
|
||||
```
|
||||
|
||||
詳細は [results ガイド](results.md) を参照してください。
|
||||
詳しくは [実行結果ガイド](results.md) を参照してください。
|
||||
|
||||
## Runner ライフサイクルと設定
|
||||
## Runner のライフサイクルと設定
|
||||
|
||||
### エージェントループ
|
||||
|
||||
`Runner` の run メソッドを使うときは、開始エージェントと入力を渡します。入力には以下を指定できます。
|
||||
`Runner` の run メソッドを使用するときは、開始エージェントと入力を渡します。入力には次を指定できます。
|
||||
|
||||
- 文字列(ユーザーメッセージとして扱われます)
|
||||
- OpenAI Responses API 形式の入力アイテムのリスト
|
||||
- 中断した実行を再開する際の [`RunState`][agents.run_state.RunState]
|
||||
- 文字列(ユーザーメッセージとして扱われます)、
|
||||
- OpenAI Responses API 形式の入力項目のリスト、または
|
||||
- 中断された実行を再開する場合の [`RunState`][agents.run_state.RunState]。
|
||||
|
||||
その後、Runner は次のループを実行します。
|
||||
runner は次にループを実行します。
|
||||
|
||||
1. 現在の入力を使って、現在のエージェントに対して LLM を呼び出します。
|
||||
1. 現在のエージェントに対して、現在の入力で LLM を呼び出します。
|
||||
2. LLM が出力を生成します。
|
||||
1. LLM が `final_output` を返した場合、ループを終了して結果を返します。
|
||||
2. LLM がハンドオフを行った場合、現在のエージェントと入力を更新してループを再実行します。
|
||||
3. LLM がツール呼び出しを生成した場合、それらを実行して結果を追加し、ループを再実行します。
|
||||
3. 渡された `max_turns` を超えた場合、[`MaxTurnsExceeded`][agents.exceptions.MaxTurnsExceeded] 例外を送出します。
|
||||
1. LLM が `final_output` を返した場合、ループは終了し、実行結果を返します。
|
||||
2. LLM がハンドオフを行った場合、現在のエージェントと入力を更新し、ループを再実行します。
|
||||
3. LLM がツール呼び出しを生成した場合、それらのツール呼び出しを実行し、実行結果を追加して、ループを再実行します。
|
||||
3. 渡された `max_turns` を超えた場合、[`MaxTurnsExceeded`][agents.exceptions.MaxTurnsExceeded] 例外を発生させます。このターン制限を無効にするには `max_turns=None` を渡します。
|
||||
|
||||
!!! note
|
||||
|
||||
LLM 出力を「最終出力」と見なすルールは、期待する型のテキスト出力が生成され、かつツール呼び出しがないことです。
|
||||
LLM の出力が「最終出力」とみなされる条件は、望ましい型のテキスト出力を生成し、かつツール呼び出しがないことです。
|
||||
|
||||
### ストリーミング
|
||||
|
||||
ストリーミングを使うと、LLM 実行中のストリーミングイベントも受け取れます。ストリーム完了後、[`RunResultStreaming`][agents.result.RunResultStreaming] には、生成されたすべての新しい出力を含む実行情報全体が格納されます。ストリーミングイベントは `.stream_events()` で取得できます。詳細は [ストリーミングガイド](streaming.md) を参照してください。
|
||||
ストリーミングを使用すると、LLM の実行中にストリーミングイベントも受け取れます。ストリームが完了すると、[`RunResultStreaming`][agents.result.RunResultStreaming] には、生成されたすべての新しい出力を含む、実行に関する完全な情報が含まれます。ストリーミングイベントには `.stream_events()` を呼び出せます。詳しくは [ストリーミングガイド](streaming.md) を参照してください。
|
||||
|
||||
#### Responses WebSocket トランスポート(任意ヘルパー)
|
||||
#### Responses WebSocket トランスポート(任意のヘルパー)
|
||||
|
||||
OpenAI Responses websocket トランスポートを有効化しても、通常の `Runner` API をそのまま使えます。接続再利用には websocket session helper の利用を推奨しますが、必須ではありません。
|
||||
OpenAI Responses websocket トランスポートを有効にしても、通常の `Runner` API を引き続き使用できます。接続の再利用には websocket セッションヘルパーの使用を推奨しますが、必須ではありません。
|
||||
|
||||
これは websocket トランスポート上の Responses API であり、[Realtime API](realtime/guide.md) ではありません。
|
||||
|
||||
トランスポート選択ルールや、具体的なモデルオブジェクト/カスタムプロバイダーに関する注意点は、[Models](models/index.md#responses-websocket-transport) を参照してください。
|
||||
トランスポート選択ルール、および具象モデルオブジェクトやカスタムプロバイダーに関する注意点については、[モデル](models/index.md#responses-websocket-transport) を参照してください。
|
||||
|
||||
##### パターン 1: session helper なし(動作します)
|
||||
##### パターン 1: セッションヘルパーなし(動作可)
|
||||
|
||||
websocket トランスポートだけを使いたく、SDK に共有 provider / session 管理を任せる必要がない場合に使います。
|
||||
websocket トランスポートだけが必要で、共有プロバイダーやセッションを SDK に管理させる必要がない場合に使用します。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -85,11 +85,11 @@ async def main():
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
このパターンは単発実行には問題ありません。`Runner.run()` / `Runner.run_streamed()` を繰り返し呼ぶ場合、同じ `RunConfig` / provider インスタンスを手動で再利用しない限り、実行ごとに再接続が発生する可能性があります。
|
||||
このパターンは単発の実行には問題ありません。`Runner.run()` / `Runner.run_streamed()` を繰り返し呼び出す場合、同じ `RunConfig` / プロバイダーインスタンスを手動で再利用しない限り、各実行で再接続される可能性があります。
|
||||
|
||||
##### パターン 2: `responses_websocket_session()` を使用(複数ターン再利用に推奨)
|
||||
##### パターン 2: `responses_websocket_session()` の使用(複数ターンでの再利用に推奨)
|
||||
|
||||
複数回の実行で websocket 対応 provider と `RunConfig` を共有したい場合(同じ `run_config` を継承するネストした agent-as-tool 呼び出しを含む)は、[`responses_websocket_session()`][agents.responses_websocket_session] を使います。
|
||||
複数の実行にわたって共有の websocket 対応プロバイダーと `RunConfig` を使用したい場合(同じ `run_config` を継承するネストされた agent-as-tool 呼び出しを含む)は、[`responses_websocket_session()`][agents.responses_websocket_session] を使用します。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -100,7 +100,9 @@ from agents import Agent, responses_websocket_session
|
||||
async def main():
|
||||
agent = Agent(name="Assistant", instructions="Be concise.")
|
||||
|
||||
async with responses_websocket_session() as ws:
|
||||
async with responses_websocket_session(
|
||||
responses_websocket_options={"ping_interval": 20.0, "ping_timeout": 60.0},
|
||||
) as ws:
|
||||
first = ws.run_streamed(agent, "Say hello in one short sentence.")
|
||||
async for _event in first.stream_events():
|
||||
pass
|
||||
@@ -117,63 +119,114 @@ async def main():
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
コンテキストを抜ける前に、ストリーミング結果の消費を完了してください。websocket リクエストが進行中のままコンテキストを終了すると、共有接続が強制クローズされる場合があります。
|
||||
コンテキストを抜ける前に、ストリーミングされた実行結果の消費を完了してください。websocket リクエストがまだ進行中の状態でコンテキストを抜けると、共有接続が強制的に閉じられる可能性があります。
|
||||
|
||||
### RunConfig
|
||||
長い推論ターンで websocket の keepalive タイムアウトに達する場合は、`ping_timeout` を増やすか、`ping_timeout=None` を設定してハートビートタイムアウトを無効にしてください。websocket のレイテンシより信頼性が重要な実行では、HTTP/SSE トランスポートを使用してください。
|
||||
|
||||
`run_config` パラメーターを使うと、エージェント実行のグローバル設定をいくつか構成できます。
|
||||
### 実行設定
|
||||
|
||||
#### 共通 RunConfig カテゴリー
|
||||
`run_config` パラメーターを使用すると、エージェント実行の一部のグローバル設定を構成できます。
|
||||
|
||||
`RunConfig` を使うと、各エージェント定義を変更せずに単一の実行に対して動作を上書きできます。
|
||||
#### 一般的な実行設定カテゴリー
|
||||
|
||||
##### モデル、プロバイダー、セッションの既定値
|
||||
各エージェント定義を変更せずに単一の実行の挙動を上書きするには、`RunConfig` を使用します。
|
||||
|
||||
- [`model`][agents.run.RunConfig.model]: 各 Agent の `model` 設定に関係なく、グローバルに使用する LLM モデルを設定できます。
|
||||
- [`model_provider`][agents.run.RunConfig.model_provider]: モデル名を解決するモデルプロバイダーです。既定値は OpenAI です。
|
||||
- [`model_settings`][agents.run.RunConfig.model_settings]: エージェント固有設定を上書きします。たとえば、グローバルな `temperature` や `top_p` を設定できます。
|
||||
- [`session_settings`][agents.run.RunConfig.session_settings]: 実行中に履歴を取得する際のセッションレベル既定値(例: `SessionSettings(limit=...)`)を上書きします。
|
||||
- [`session_input_callback`][agents.run.RunConfig.session_input_callback]: Sessions 使用時に、各ターン前に新しいユーザー入力をセッション履歴へどうマージするかをカスタマイズします。コールバックは同期/非同期どちらでも可能です。
|
||||
##### モデル、プロバイダー、セッションのデフォルト
|
||||
|
||||
##### ガードレール、ハンドオフ、モデル入力整形
|
||||
- [`model`][agents.run.RunConfig.model]: 各 Agent が持つ `model` に関係なく、使用するグローバルな LLM モデルを設定できます。
|
||||
- [`model_provider`][agents.run.RunConfig.model_provider]: モデル名の検索に使用するモデルプロバイダーです。デフォルトは OpenAI です。
|
||||
- [`model_settings`][agents.run.RunConfig.model_settings]: エージェント固有の設定を上書きします。たとえば、グローバルな `temperature` や `top_p` を設定できます。
|
||||
- [`session_settings`][agents.run.RunConfig.session_settings]: 実行中に履歴を取得する際のセッションレベルのデフォルト(例: `SessionSettings(limit=...)`)を上書きします。
|
||||
- [`session_input_callback`][agents.run.RunConfig.session_input_callback]: セッションを使用する場合に、各ターンの前に新しいユーザー入力をセッション履歴とどのようにマージするかをカスタマイズします。コールバックは同期または非同期にできます。
|
||||
|
||||
- [`input_guardrails`][agents.run.RunConfig.input_guardrails], [`output_guardrails`][agents.run.RunConfig.output_guardrails]: すべての実行に含める入力/出力ガードレールのリストです。
|
||||
- [`handoff_input_filter`][agents.run.RunConfig.handoff_input_filter]: ハンドオフ側に未設定の場合、すべてのハンドオフに適用するグローバル入力フィルターです。新しいエージェントへ送る入力を編集できます。詳細は [`Handoff.input_filter`][agents.handoffs.Handoff.input_filter] のドキュメントを参照してください。
|
||||
- [`nest_handoff_history`][agents.run.RunConfig.nest_handoff_history]: 次エージェント呼び出し前に、直前までの transcript を単一の assistant メッセージへ折りたたむ opt-in beta 機能です。ネストしたハンドオフの安定化中のため既定で無効です。有効化は `True`、raw transcript をそのまま通すには `False` を使います。[Runner メソッド][agents.run.Runner] は `RunConfig` 未指定時に自動作成されるため、quickstart や examples では既定の無効状態が維持され、明示的な [`Handoff.input_filter`][agents.handoffs.Handoff.input_filter] コールバックは引き続き優先されます。個々のハンドオフは [`Handoff.nest_handoff_history`][agents.handoffs.Handoff.nest_handoff_history] で上書きできます。
|
||||
- [`handoff_history_mapper`][agents.run.RunConfig.handoff_history_mapper]: `nest_handoff_history` を有効化した際に、正規化された transcript(履歴 + ハンドオフアイテム)を受け取る任意 callable です。次エージェントへ渡す入力アイテムの**正確なリスト**を返す必要があり、完全なハンドオフフィルターを書かずに組み込み要約を置き換えられます。
|
||||
- [`call_model_input_filter`][agents.run.RunConfig.call_model_input_filter]: モデル呼び出し直前に、完全に準備済みのモデル入力(instructions と入力アイテム)を編集するフックです。例: 履歴のトリミングやシステムプロンプトの注入。
|
||||
- [`reasoning_item_id_policy`][agents.run.RunConfig.reasoning_item_id_policy]: Runner が過去出力を次ターンのモデル入力へ変換する際に、reasoning item ID を保持するか省略するかを制御します。
|
||||
##### ガードレール、ハンドオフ、モデル入力の整形
|
||||
|
||||
- [`input_guardrails`][agents.run.RunConfig.input_guardrails], [`output_guardrails`][agents.run.RunConfig.output_guardrails]: すべての実行に含める入力または出力ガードレールのリストです。
|
||||
- [`handoff_input_filter`][agents.run.RunConfig.handoff_input_filter]: ハンドオフに入力フィルターがまだない場合に、すべてのハンドオフへ適用するグローバル入力フィルターです。入力フィルターを使用すると、新しいエージェントへ送信される入力を編集できます。詳細は [`Handoff.input_filter`][agents.handoffs.Handoff.input_filter] のドキュメントを参照してください。
|
||||
- [`nest_handoff_history`][agents.run.RunConfig.nest_handoff_history]: 次のエージェントを呼び出す前に、以前のトランスクリプトを単一の assistant メッセージに折りたたむオプトインのベータ機能です。ネストされたハンドオフを安定化している間、これはデフォルトで無効です。有効にするには `True` に設定し、raw トランスクリプトをそのまま渡すには `False` のままにします。すべての [Runner メソッド][agents.run.Runner] は、渡されていない場合に自動的に `RunConfig` を作成するため、クイックスタートやコード例ではデフォルトがオフのままになります。また、明示的な [`Handoff.input_filter`][agents.handoffs.Handoff.input_filter] コールバックは引き続きこれを上書きします。個別のハンドオフでは [`Handoff.nest_handoff_history`][agents.handoffs.Handoff.nest_handoff_history] を通じてこの設定を上書きできます。
|
||||
- [`handoff_history_mapper`][agents.run.RunConfig.handoff_history_mapper]: `nest_handoff_history` をオプトインしたときに、正規化されたトランスクリプト(履歴 + ハンドオフ項目)を受け取る任意の callable です。次のエージェントへ転送する入力項目の正確なリストを返す必要があり、完全なハンドオフフィルターを書かずに組み込みの要約を置き換えられます。
|
||||
- [`call_model_input_filter`][agents.run.RunConfig.call_model_input_filter]: モデル呼び出しの直前に、完全に準備されたモデル入力(instructions と入力項目)を編集するためのフックです。たとえば、履歴をトリミングしたり、システムプロンプトを注入したりできます。
|
||||
- [`reasoning_item_id_policy`][agents.run.RunConfig.reasoning_item_id_policy]: runner が以前の出力を次ターンのモデル入力に変換するときに、reasoning 項目の ID を保持するか省略するかを制御します。
|
||||
|
||||
##### トレーシングと可観測性
|
||||
|
||||
- [`tracing_disabled`][agents.run.RunConfig.tracing_disabled]: 実行全体の [トレーシング](tracing.md) を無効化できます。
|
||||
- [`tracing`][agents.run.RunConfig.tracing]: [`TracingConfig`][agents.tracing.TracingConfig] を渡し、実行単位のトレーシング API key などの trace export 設定を上書きします。
|
||||
- [`trace_include_sensitive_data`][agents.run.RunConfig.trace_include_sensitive_data]: trace に LLM やツール呼び出しの入力/出力などの機微データを含めるかを設定します。
|
||||
- [`workflow_name`][agents.run.RunConfig.workflow_name], [`trace_id`][agents.run.RunConfig.trace_id], [`group_id`][agents.run.RunConfig.group_id]: 実行のトレーシング workflow 名、trace ID、trace group ID を設定します。少なくとも `workflow_name` の設定を推奨します。group ID は任意で、複数実行間の trace を関連付けられます。
|
||||
- [`trace_metadata`][agents.run.RunConfig.trace_metadata]: すべての trace に含めるメタデータです。
|
||||
- [`tracing_disabled`][agents.run.RunConfig.tracing_disabled]: 実行全体の [トレーシング](tracing.md) を無効にできます。
|
||||
- [`tracing`][agents.run.RunConfig.tracing]: 実行ごとのトレーシング API キーなど、トレースエクスポート設定を上書きするには [`TracingConfig`][agents.tracing.TracingConfig] を渡します。
|
||||
- [`trace_include_sensitive_data`][agents.run.RunConfig.trace_include_sensitive_data]: トレースに LLM やツール呼び出しの入力/出力など、機微である可能性のあるデータを含めるかどうかを設定します。
|
||||
- [`workflow_name`][agents.run.RunConfig.workflow_name], [`trace_id`][agents.run.RunConfig.trace_id], [`group_id`][agents.run.RunConfig.group_id]: 実行のトレーシングワークフロー名、トレース ID、トレースグループ ID を設定します。少なくとも `workflow_name` を設定することを推奨します。グループ ID は、複数の実行にわたってトレースを関連付けられる任意フィールドです。
|
||||
- [`trace_metadata`][agents.run.RunConfig.trace_metadata]: すべてのトレースに含めるメタデータです。
|
||||
|
||||
##### ツール承認とツールエラー動作
|
||||
##### ツール実行、承認、ツールエラーの挙動
|
||||
|
||||
- [`tool_error_formatter`][agents.run.RunConfig.tool_error_formatter]: 承認フロー中にツール呼び出しが拒否された場合、モデルに見えるメッセージをカスタマイズします。
|
||||
- [`tool_execution`][agents.run.RunConfig.tool_execution]: 一度に実行される関数ツールの数を制限するなど、ローカルツール呼び出しに対する SDK 側の実行挙動を設定します。
|
||||
- [`tool_not_found_behavior`][agents.run.RunConfig.tool_not_found_behavior]: モデルによって出力された未解決の関数ツール呼び出しを runner がどのように扱うかを設定します。デフォルトでは `ModelBehaviorError` が発生します。代わりにモデルから見えるエラー出力を返すようにオプトインできます。
|
||||
- [`tool_error_formatter`][agents.run.RunConfig.tool_error_formatter]: 承認拒否やオプトインのツール未検出出力など、モデルから見えるツールエラーメッセージをカスタマイズします。
|
||||
|
||||
ネストしたハンドオフは opt-in beta として利用できます。折りたたみ transcript 動作を有効にするには `RunConfig(nest_handoff_history=True)` を渡すか、特定ハンドオフで `handoff(..., nest_handoff_history=True)` を設定してください。raw transcript(既定)を維持したい場合は、フラグを未設定のままにするか、必要な形で会話を正確に転送する `handoff_input_filter`(または `handoff_history_mapper`)を指定してください。カスタム mapper を書かずに生成要約で使うラッパーテキストを変更するには、[`set_conversation_history_wrappers`][agents.handoffs.set_conversation_history_wrappers] を呼び出してください(既定へ戻すには [`reset_conversation_history_wrappers`][agents.handoffs.reset_conversation_history_wrappers])。
|
||||
ネストされたハンドオフはオプトインのベータ機能として利用できます。折りたたみトランスクリプトの挙動を有効にするには `RunConfig(nest_handoff_history=True)` を渡すか、特定のハンドオフで有効にするには `handoff(..., nest_handoff_history=True)` を設定します。raw トランスクリプトを維持したい場合(デフォルト)は、フラグを未設定のままにするか、必要なとおりに会話を転送する `handoff_input_filter`(または `handoff_history_mapper`)を指定します。カスタムマッパーを書かずに、生成される要約で使用されるラッパーテキストを変更するには、[`set_conversation_history_wrappers`][agents.handoffs.set_conversation_history_wrappers] を呼び出します(デフォルトを復元するには [`reset_conversation_history_wrappers`][agents.handoffs.reset_conversation_history_wrappers] を呼び出します)。
|
||||
|
||||
#### RunConfig 詳細
|
||||
#### 実行設定の詳細
|
||||
|
||||
##### `tool_execution`
|
||||
|
||||
実行におけるローカル関数ツールの同時実行数の制限など、ローカル関数ツールに対する SDK 側の挙動を設定したい場合は、`tool_execution` を使用します。
|
||||
|
||||
```python
|
||||
from agents import Agent, RunConfig, Runner, ToolExecutionConfig
|
||||
|
||||
agent = Agent(name="Assistant", tools=[...])
|
||||
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Run the required tool calls.",
|
||||
run_config=RunConfig(
|
||||
tool_execution=ToolExecutionConfig(
|
||||
max_function_tool_concurrency=2,
|
||||
pre_approval_tool_input_guardrails=True,
|
||||
),
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
`max_function_tool_concurrency=None` はデフォルトの挙動を維持します。モデルが 1 ターンで複数の関数ツール呼び出しを出力した場合、SDK は出力されたすべてのローカル関数ツール呼び出しを開始します。整数値を設定すると、それらのローカル関数ツールのうち同時に実行される数に上限を設けられます。
|
||||
|
||||
これはプロバイダー側の [`ModelSettings.parallel_tool_calls`][agents.model_settings.ModelSettings.parallel_tool_calls] とは別です。`parallel_tool_calls` は、モデルが 1 つのレスポンスで複数のツール呼び出しを出力できるかどうかを制御します。`tool_execution.max_function_tool_concurrency` は、モデルがそれらを出力した後に、SDK がローカル関数ツール呼び出しをどのように実行するかを制御します。
|
||||
|
||||
`pre_approval_tool_input_guardrails=False` はデフォルトの承認フローを維持します。関数ツールに承認が必要な場合、実行はまず一時停止し、ツール入力ガードレールは承認後、実行直前にのみ実行されます。保留中の承認割り込みが出力される前に関数ツール入力ガードレールを実行したい場合は、`True` に設定します。この事前承認チェックに合格した呼び出しでも、承認後に同じ入力ガードレールが再度実行されるため、時間依存のチェックは実行前に再検証されます。
|
||||
|
||||
##### `tool_not_found_behavior`
|
||||
|
||||
デフォルトでは、モデルが現在のエージェントで利用可能な関数ツールのいずれにも一致しない関数ツール呼び出しを出力した場合、runner は `ModelBehaviorError` を発生させます。
|
||||
|
||||
実行を回復可能なままにしたい場合は、`tool_not_found_behavior="return_error_to_model"` を設定します。このモードでは、SDK は未解決のツール呼び出しに対する `function_call_output` を追加し、モデルを再度実行します。これにより、モデルは利用可能なツールを選択するか、そのツールを使用せずに回答できます。
|
||||
|
||||
```python
|
||||
from agents import Agent, RunConfig, Runner
|
||||
|
||||
agent = Agent(name="Assistant", tools=[...])
|
||||
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Handle this request with the available tools.",
|
||||
run_config=RunConfig(tool_not_found_behavior="return_error_to_model"),
|
||||
)
|
||||
```
|
||||
|
||||
このオプションは現在、未解決の関数ツール呼び出しにのみ適用されます。その他の無効なツールペイロードでは、既存のエラー挙動が引き続き使用されます。
|
||||
|
||||
##### `tool_error_formatter`
|
||||
|
||||
`tool_error_formatter` を使うと、承認フローでツール呼び出しが拒否された際にモデルへ返すメッセージをカスタマイズできます。
|
||||
SDK がモデルから見えるツールエラー出力を作成するときにモデルへ返されるメッセージをカスタマイズするには、`tool_error_formatter` を使用します。
|
||||
|
||||
formatter には以下を含む [`ToolErrorFormatterArgs`][agents.run_config.ToolErrorFormatterArgs] が渡されます。
|
||||
formatter は、次を含む [`ToolErrorFormatterArgs`][agents.run_config.ToolErrorFormatterArgs] を受け取ります。
|
||||
|
||||
- `kind`: エラーカテゴリー。現時点では `"approval_rejected"` です。
|
||||
- `tool_type`: ツールランタイム(`"function"`、`"computer"`、`"shell"`、`"apply_patch"`、`"custom"`)。
|
||||
- `tool_name`: ツール名。
|
||||
- `call_id`: ツール呼び出し ID。
|
||||
- `default_message`: SDK 既定のモデル可視メッセージ。
|
||||
- `run_context`: 現在の run context wrapper。
|
||||
- `kind`: `"approval_rejected"` や `"tool_not_found"` などのエラーカテゴリーです。
|
||||
- `tool_type`: ツールランタイム(`"function"`, `"computer"`, `"shell"`, `"apply_patch"`, または `"custom"`)です。
|
||||
- `tool_name`: ツール名です。
|
||||
- `call_id`: ツール呼び出し ID です。
|
||||
- `default_message`: SDK のデフォルトのモデルから見えるメッセージです。
|
||||
- `run_context`: アクティブな実行コンテキストラッパーです。
|
||||
|
||||
メッセージを置き換える文字列を返すか、SDK 既定を使う場合は `None` を返します。
|
||||
メッセージを置き換えるには文字列を返し、SDK のデフォルトを使用するには `None` を返します。
|
||||
|
||||
```python
|
||||
from agents import Agent, RunConfig, Runner, ToolErrorFormatterArgs
|
||||
@@ -185,6 +238,8 @@ def format_rejection(args: ToolErrorFormatterArgs[None]) -> str | None:
|
||||
f"Tool call '{args.tool_name}' was rejected by a human reviewer. "
|
||||
"Ask for confirmation or propose a safer alternative."
|
||||
)
|
||||
if args.kind == "tool_not_found":
|
||||
return f"Tool '{args.tool_name}' is not available. Choose one of the listed tools."
|
||||
return None
|
||||
|
||||
|
||||
@@ -198,57 +253,56 @@ result = Runner.run_sync(
|
||||
|
||||
##### `reasoning_item_id_policy`
|
||||
|
||||
`reasoning_item_id_policy` は、Runner が履歴を引き継ぐ際(例: `RunResult.to_input_list()` やセッションバック実行)に reasoning items を次ターンのモデル入力へどう変換するかを制御します。
|
||||
`reasoning_item_id_policy` は、runner が履歴を次へ持ち越すとき(たとえば、`RunResult.to_input_list()` やセッションに基づく実行を使用するとき)に、reasoning 項目を次ターンのモデル入力へどのように変換するかを制御します。
|
||||
|
||||
- `None` または `"preserve"`(既定): reasoning item ID を保持します。
|
||||
- `"omit"`: 生成される次ターン入力から reasoning item ID を除去します。
|
||||
- `None` または `"preserve"`(デフォルト): reasoning 項目 ID を保持します。
|
||||
- `"omit"`: 生成される次ターン入力から reasoning 項目 ID を削除します。
|
||||
|
||||
`"omit"` は主に、reasoning item に `id` があるが必須の後続 item がない場合に発生する Responses API 400 エラー群への opt-in 緩和策として使います(例: `Item 'rs_...' of type 'reasoning' was provided without its required following item.`)。
|
||||
`"omit"` は主に、reasoning 項目が `id` 付きで送信されたものの、必須の後続項目がない場合に発生する Responses API 400 エラーの一種に対するオプトインの緩和策として使用します(例: `Item 'rs_...' of type 'reasoning' was provided without its required following item.`)。
|
||||
|
||||
これは、SDK が過去出力から後続入力を構築する複数ターンエージェント実行(セッション永続化、サーバー管理会話 delta、ストリーミング/非ストリーミング後続ターン、再開経路を含む)で、reasoning item ID が保持される一方、プロバイダー側でその ID を対応する後続 item とペアで維持することを要求する場合に発生し得ます。
|
||||
これは、SDK が以前の出力から後続入力を構築する複数ターンのエージェント実行(セッション永続化、サーバー管理の会話差分、ストリーミング/非ストリーミングの後続ターン、再開パスを含む)において、reasoning 項目 ID が保持されている一方で、プロバイダーがその ID を対応する後続項目とペアのままにすることを要求する場合に発生する可能性があります。
|
||||
|
||||
`reasoning_item_id_policy="omit"` を設定すると、reasoning 内容は保持しつつ reasoning item の `id` を除去するため、SDK 生成の後続入力でその API 不変条件の違反を回避できます。
|
||||
`reasoning_item_id_policy="omit"` を設定すると、reasoning の内容は保持しつつ reasoning 項目の `id` を削除します。これにより、SDK が生成する後続入力でその API 不変条件に抵触することを回避できます。
|
||||
|
||||
スコープに関する注意:
|
||||
|
||||
- 変更対象は、SDK が後続入力を構築する際に生成/転送する reasoning items のみです。
|
||||
- ユーザー提供の初期入力 items は書き換えません。
|
||||
- `call_model_input_filter` により、このポリシー適用後に意図的に reasoning ID を再導入することは可能です。
|
||||
- これは、SDK が後続入力を構築するときに生成または転送する reasoning 項目のみを変更します。
|
||||
- ユーザーが指定した初期入力項目は書き換えません。
|
||||
- このポリシーが適用された後でも、`call_model_input_filter` は意図的に reasoning ID を再導入できます。
|
||||
|
||||
## 状態と会話管理
|
||||
|
||||
### メモリ戦略の選択
|
||||
|
||||
状態を次ターンへ渡す一般的な方法は 4 つあります。
|
||||
次のターンへ状態を持ち越す一般的な方法は 4 つあります。
|
||||
|
||||
| Strategy | Where state lives | Best for | What you pass on the next turn |
|
||||
| 戦略 | 状態の所在 | 最適な用途 | 次のターンで渡すもの |
|
||||
| --- | --- | --- | --- |
|
||||
| `result.to_input_list()` | アプリのメモリ | 小規模チャットループ、完全な手動制御、任意のプロバイダー | `result.to_input_list()` のリスト + 次のユーザーメッセージ |
|
||||
| `session` | ユーザーのストレージ + SDK | 永続チャット状態、再開可能実行、カスタムストア | 同じ `session` インスタンス、または同じストアを指す別インスタンス |
|
||||
| `conversation_id` | OpenAI Conversations API | 複数ワーカー/サービス間で共有したい名前付きサーバー側会話 | 同じ `conversation_id` + 新しいユーザーターンのみ |
|
||||
| `previous_response_id` | OpenAI Responses API | 会話リソースを作らない軽量サーバー管理継続 | `result.last_response_id` + 新しいユーザーターンのみ |
|
||||
| `result.to_input_list()` | アプリのメモリ | 小規模なチャットループ、完全な手動制御、任意のプロバイダー | `result.to_input_list()` からのリストに次のユーザーメッセージを加えたもの |
|
||||
| `session` | アプリのストレージと SDK | 永続的なチャット状態、再開可能な実行、カスタムストア | 同じ `session` インスタンス、または同じストアを指す別のインスタンス |
|
||||
| `conversation_id` | OpenAI Conversations API | ワーカーやサービス間で共有したい名前付きのサーバー側会話 | 同じ `conversation_id` と、新しいユーザーターンのみ |
|
||||
| `previous_response_id` | OpenAI Responses API | 会話リソースを作成しない、軽量なサーバー管理の継続 | `result.last_response_id` と、新しいユーザーターンのみ |
|
||||
|
||||
`result.to_input_list()` と `session` はクライアント管理です。`conversation_id` と `previous_response_id` は OpenAI 管理で、OpenAI Responses API 使用時のみ適用されます。多くのアプリでは、会話ごとに永続化戦略を 1 つ選んでください。クライアント管理履歴と OpenAI 管理状態を混在させると、意図的に両レイヤーを調整していない限りコンテキストが重複する場合があります。
|
||||
`result.to_input_list()` と `session` はクライアント管理です。`conversation_id` と `previous_response_id` は OpenAI 管理であり、OpenAI Responses API を使用している場合にのみ適用されます。ほとんどのアプリケーションでは、会話ごとに 1 つの永続化戦略を選択してください。クライアント管理の履歴と OpenAI 管理の状態を混在させると、両方のレイヤーを意図的に調整している場合を除き、コンテキストが重複する可能性があります。
|
||||
|
||||
!!! note
|
||||
|
||||
セッション永続化はサーバー管理会話設定
|
||||
(`conversation_id`、`previous_response_id`、`auto_previous_response_id`)と
|
||||
同一実行で併用できません。
|
||||
呼び出しごとにどちらか 1 つの方式を選んでください。
|
||||
セッション永続化は、サーバー管理の会話設定
|
||||
(`conversation_id`, `previous_response_id`, または `auto_previous_response_id`)と同じ実行内で
|
||||
組み合わせることはできません。呼び出しごとに 1 つの方法を選択してください。
|
||||
|
||||
### Conversations/chat threads
|
||||
### 会話/チャットスレッド
|
||||
|
||||
どの run メソッドを呼び出しても、結果として 1 つ以上のエージェント実行(つまり 1 回以上の LLM 呼び出し)が発生する可能性がありますが、チャット会話上は 1 つの論理ターンを表します。例:
|
||||
いずれかの run メソッドを呼び出すと、1 つ以上のエージェントが実行される(したがって 1 回以上の LLM 呼び出しが行われる)可能性がありますが、チャット会話における単一の論理ターンを表します。例:
|
||||
|
||||
1. ユーザーターン: ユーザーがテキスト入力
|
||||
2. Runner 実行: 最初のエージェントが LLM を呼び出し、ツールを実行し、2 つ目のエージェントへハンドオフし、2 つ目のエージェントがさらにツールを実行して出力を生成
|
||||
1. ユーザーターン: ユーザーがテキストを入力します
|
||||
2. Runner 実行: 最初のエージェントが LLM を呼び出し、ツールを実行し、2 番目のエージェントへハンドオフし、2 番目のエージェントがさらにツールを実行してから出力を生成します。
|
||||
|
||||
エージェント実行の最後に、ユーザーへ何を表示するかを選べます。たとえば、エージェントが生成した新規アイテムをすべて表示することも、最終出力のみ表示することもできます。いずれの場合も、その後ユーザーがフォローアップ質問をしたら、run メソッドを再度呼び出せます。
|
||||
エージェントの実行が終了した時点で、ユーザーに何を表示するかを選択できます。たとえば、エージェントによって生成されたすべての新しい項目をユーザーに表示することも、最終出力だけを表示することもできます。いずれの場合も、その後ユーザーがフォローアップ質問をする可能性があり、その場合は run メソッドを再度呼び出せます。
|
||||
|
||||
#### 手動の会話管理
|
||||
|
||||
[`RunResultBase.to_input_list()`][agents.result.RunResultBase.to_input_list] メソッドを使うと、次ターン用入力を取得して会話履歴を手動管理できます。
|
||||
[`RunResultBase.to_input_list()`][agents.result.RunResultBase.to_input_list] メソッドを使用して次ターンの入力を取得することで、会話履歴を手動で管理できます。
|
||||
|
||||
```python
|
||||
async def main():
|
||||
@@ -268,9 +322,9 @@ async def main():
|
||||
# California
|
||||
```
|
||||
|
||||
#### Sessions による自動会話管理
|
||||
#### セッションによる自動会話管理
|
||||
|
||||
より簡単な方法として、[Sessions](sessions/index.md) を使うと `.to_input_list()` を手動で呼ばずに会話履歴を自動処理できます。
|
||||
より簡単な方法として、`.to_input_list()` を手動で呼び出さずに会話履歴を自動処理するために [セッション](sessions/index.md) を使用できます。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, SQLiteSession
|
||||
@@ -294,24 +348,24 @@ async def main():
|
||||
# California
|
||||
```
|
||||
|
||||
Sessions は自動で次を行います。
|
||||
セッションは次を自動的に行います。
|
||||
|
||||
- 各実行前に会話履歴を取得
|
||||
- 各実行後に新規メッセージを保存
|
||||
- 異なるセッション ID ごとに別会話を維持
|
||||
- 各実行の前に会話履歴を取得します
|
||||
- 各実行の後に新しいメッセージを保存します
|
||||
- 異なるセッション ID ごとに別々の会話を維持します
|
||||
|
||||
詳細は [Sessions ドキュメント](sessions/index.md) を参照してください。
|
||||
詳細については、[セッションのドキュメント](sessions/index.md) を参照してください。
|
||||
|
||||
|
||||
#### サーバー管理会話
|
||||
#### サーバー管理の会話
|
||||
|
||||
`to_input_list()` や `Sessions` でローカル管理する代わりに、OpenAI の会話状態機能でサーバー側管理することもできます。これにより、過去メッセージを毎回手動で再送せずに会話履歴を保持できます。以下いずれのサーバー管理方式でも、各リクエストでは新規ターン入力のみを渡し、保存済み ID を再利用してください。詳細は [OpenAI Conversation state ガイド](https://platform.openai.com/docs/guides/conversation-state?api-mode=responses) を参照してください。
|
||||
`to_input_list()` や `Sessions` でローカルに処理する代わりに、OpenAI の会話状態機能にサーバー側で会話状態を管理させることもできます。これにより、過去のすべてのメッセージを手動で再送信せずに会話履歴を保持できます。以下のいずれのサーバー管理アプローチでも、各リクエストでは新しいターンの入力のみを渡し、保存した ID を再利用してください。詳細については、[OpenAI Conversation state guide](https://platform.openai.com/docs/guides/conversation-state?api-mode=responses) を参照してください。
|
||||
|
||||
OpenAI ではターン間状態追跡に 2 つの方法があります。
|
||||
OpenAI は、ターン間で状態を追跡する 2 つの方法を提供しています。
|
||||
|
||||
##### 1. `conversation_id` を使用
|
||||
##### 1. `conversation_id` の使用
|
||||
|
||||
最初に OpenAI Conversations API で会話を作成し、以降の呼び出しごとにその ID を再利用します。
|
||||
まず OpenAI Conversations API を使用して会話を作成し、その後のすべての呼び出しでその ID を再利用します。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
@@ -332,9 +386,9 @@ async def main():
|
||||
print(f"Assistant: {result.final_output}")
|
||||
```
|
||||
|
||||
##### 2. `previous_response_id` を使用
|
||||
##### 2. `previous_response_id` の使用
|
||||
|
||||
もう 1 つは **response chaining** で、各ターンが前ターンの response ID に明示的にリンクします。
|
||||
もう 1 つの選択肢は **レスポンスチェーン** で、各ターンが前のターンのレスポンス ID に明示的にリンクします。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
@@ -359,33 +413,31 @@ async def main():
|
||||
print(f"Assistant: {result.final_output}")
|
||||
```
|
||||
|
||||
実行が承認待ちで一時停止し、[`RunState`][agents.run_state.RunState] から再開した場合、
|
||||
SDK は保存済みの `conversation_id` / `previous_response_id` / `auto_previous_response_id`
|
||||
設定を維持するため、再開ターンも同じサーバー管理会話で継続されます。
|
||||
実行が承認のために一時停止し、[`RunState`][agents.run_state.RunState] から再開する場合、SDK は保存された `conversation_id` / `previous_response_id` / `auto_previous_response_id` 設定を維持するため、再開されたターンは同じサーバー管理の会話内で継続します。
|
||||
|
||||
`conversation_id` と `previous_response_id` は排他的です。システム間で共有可能な名前付き会話リソースが必要なら `conversation_id` を使ってください。ターン間継続の最も軽量な Responses API プリミティブが必要なら `previous_response_id` を使ってください。
|
||||
`conversation_id` と `previous_response_id` は相互に排他的です。システム間で共有できる名前付きの会話リソースが必要な場合は `conversation_id` を使用します。あるターンから次のターンへ最も軽量な Responses API 継続の基本コンポーネントが必要な場合は `previous_response_id` を使用します。
|
||||
|
||||
!!! note
|
||||
|
||||
SDK は `conversation_locked` エラーをバックオフ付きで自動再試行します。サーバー管理
|
||||
会話実行では、再試行前に内部の conversation-tracker 入力を巻き戻し、同じ
|
||||
準備済みアイテムをクリーンに再送できるようにします。
|
||||
SDK は `conversation_locked` エラーをバックオフ付きで自動的に再試行します。サーバー管理の
|
||||
会話実行では、再試行前に内部の会話トラッカー入力を巻き戻し、
|
||||
同じ準備済み項目をクリーンに再送信できるようにします。
|
||||
|
||||
ローカルのセッションベース実行(`conversation_id`、
|
||||
`previous_response_id`、`auto_previous_response_id` と併用不可)でも、
|
||||
SDK は再試行後の履歴重複を減らすため、直近で永続化した入力アイテムの
|
||||
ベストエフォートなロールバックを行います。
|
||||
ローカルのセッションベースの実行(`conversation_id`,
|
||||
`previous_response_id`, または `auto_previous_response_id` と組み合わせることはできません)では、SDK は
|
||||
再試行後に履歴エントリが重複するのを減らすため、最近永続化された入力項目についてもベストエフォートの
|
||||
ロールバックを行います。
|
||||
|
||||
この互換性再試行は、`ModelSettings.retry` を設定していなくても実行されます。より
|
||||
広範な opt-in モデルリクエスト再試行については、[Runner 管理再試行](models/index.md#runner-managed-retries) を参照してください。
|
||||
この互換性のための再試行は、`ModelSettings.retry` を設定していない場合でも行われます。モデルリクエストに対する
|
||||
より広範なオプトインの再試行挙動については、[Runner 管理の再試行](models/index.md#runner-managed-retries) を参照してください。
|
||||
|
||||
## フックとカスタマイズ
|
||||
|
||||
### call model input filter
|
||||
### モデル呼び出し入力フィルター
|
||||
|
||||
`call_model_input_filter` を使うと、モデル呼び出し直前にモデル入力を編集できます。このフックは現在のエージェント、コンテキスト、結合済み入力アイテム(存在する場合はセッション履歴を含む)を受け取り、新しい `ModelInputData` を返します。
|
||||
モデル呼び出しの直前にモデル入力を編集するには、`call_model_input_filter` を使用します。このフックは現在のエージェント、コンテキスト、および結合済みの入力項目(存在する場合はセッション履歴を含む)を受け取り、新しい `ModelInputData` を返します。
|
||||
|
||||
戻り値は [`ModelInputData`][agents.run.ModelInputData] オブジェクトである必要があります。`input` フィールドは必須で、入力アイテムのリストでなければなりません。これ以外の形を返すと `UserError` が発生します。
|
||||
戻り値は [`ModelInputData`][agents.run.ModelInputData] オブジェクトである必要があります。その `input` フィールドは必須で、入力項目のリストでなければなりません。それ以外の形式を返すと `UserError` が発生します。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, RunConfig
|
||||
@@ -404,19 +456,19 @@ result = Runner.run_sync(
|
||||
)
|
||||
```
|
||||
|
||||
Runner は準備済み入力リストのコピーをこのフックに渡すため、呼び出し元の元リストを直接変更せずに、トリミング、置換、並べ替えができます。
|
||||
runner は準備済み入力リストのコピーをフックに渡すため、呼び出し元の元のリストをその場で変更せずに、トリミング、置換、並べ替えができます。
|
||||
|
||||
session 使用時、`call_model_input_filter` はセッション履歴の読み込みと現在ターンへのマージが完了した後に実行されます。この前段のマージ処理自体をカスタマイズしたい場合は [`session_input_callback`][agents.run.RunConfig.session_input_callback] を使ってください。
|
||||
セッションを使用している場合、`call_model_input_filter` はセッション履歴がすでに読み込まれ、現在のターンとマージされた後に実行されます。その前段のマージ手順自体をカスタマイズしたい場合は、[`session_input_callback`][agents.run.RunConfig.session_input_callback] を使用してください。
|
||||
|
||||
`conversation_id`、`previous_response_id`、`auto_previous_response_id` による OpenAI サーバー管理会話状態を使う場合、このフックは次の Responses API 呼び出し用に準備されたペイロードに対して実行されます。そのペイロードは、過去履歴の完全再送ではなく新規ターン差分のみを表すことがあります。サーバー管理継続で送信済みとしてマークされるのは、あなたが返したアイテムのみです。
|
||||
`conversation_id`、`previous_response_id`、または `auto_previous_response_id` を使って OpenAI のサーバー管理の会話状態を使用している場合、このフックは次の Responses API 呼び出し向けに準備されたペイロード上で実行されます。そのペイロードは、以前の履歴全体の再生ではなく、すでに新しいターンの差分のみを表している場合があります。返した項目のみが、そのサーバー管理の継続に対して送信済みとしてマークされます。
|
||||
|
||||
このフックは `run_config` 経由で実行ごとに設定でき、機微データのマスキング、長い履歴のトリミング、追加のシステムガイダンス注入に使えます。
|
||||
機微データのマスク、長い履歴のトリミング、追加のシステムガイダンスの注入を行うには、`run_config` を通じて実行ごとにフックを設定します。
|
||||
|
||||
## エラーと復旧
|
||||
|
||||
### エラーハンドラー
|
||||
|
||||
すべての `Runner` エントリーポイントは、エラー種別をキーにした dict `error_handlers` を受け取れます。現時点でサポートされるキーは `"max_turns"` です。`MaxTurnsExceeded` を送出せず、制御された最終出力を返したい場合に使用します。
|
||||
すべての `Runner` エントリーポイントは、エラー種別をキーとする dict である `error_handlers` を受け取ります。サポートされているキーは `"max_turns"` と `"model_refusal"` です。`MaxTurnsExceeded` や `ModelRefusalError` を発生させる代わりに、制御された最終出力を返したい場合に使用します。
|
||||
|
||||
```python
|
||||
from agents import (
|
||||
@@ -445,35 +497,69 @@ result = Runner.run_sync(
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
フォールバック出力を会話履歴に追加したくない場合は、`include_in_history=False` を設定してください。
|
||||
フォールバック出力を会話履歴に追加したくない場合は、`include_in_history=False` を設定します。
|
||||
|
||||
## 耐久実行連携と human-in-the-loop
|
||||
モデルの拒否によって `ModelRefusalError` で実行を終了するのではなく、アプリケーション固有のフォールバックを生成したい場合は、`"model_refusal"` を使用します。
|
||||
|
||||
ツール承認の pause / resume パターンについては、専用の [Human-in-the-loop ガイド](human_in_the_loop.md) から始めてください。
|
||||
以下の連携は、実行が長時間待機、再試行、プロセス再起動をまたぐ場合の耐久オーケストレーション向けです。
|
||||
```python
|
||||
from pydantic import BaseModel
|
||||
|
||||
from agents import Agent, ModelRefusalError, RunErrorHandlerInput, Runner
|
||||
|
||||
|
||||
class Recipe(BaseModel):
|
||||
ingredients: list[str]
|
||||
refusal_reason: str | None = None
|
||||
|
||||
|
||||
def on_model_refusal(data: RunErrorHandlerInput[None]) -> Recipe:
|
||||
assert isinstance(data.error, ModelRefusalError)
|
||||
return Recipe(ingredients=[], refusal_reason=data.error.refusal)
|
||||
|
||||
|
||||
agent = Agent(
|
||||
name="Recipe assistant",
|
||||
instructions="Return a structured recipe.",
|
||||
output_type=Recipe,
|
||||
)
|
||||
|
||||
result = Runner.run_sync(
|
||||
agent,
|
||||
"Make me something unsafe.",
|
||||
error_handlers={"model_refusal": on_model_refusal},
|
||||
)
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
## 耐久実行の統合とヒューマンインザループ
|
||||
|
||||
ツール承認の一時停止/再開パターンについては、専用の [Human-in-the-loop ガイド](human_in_the_loop.md) から始めてください。以下の統合は、実行が長い待機、再試行、またはプロセス再起動にまたがる可能性がある場合の耐久オーケストレーション向けです。
|
||||
|
||||
### Dapr
|
||||
|
||||
Agents SDK の [Dapr](https://dapr.io) Diagrid 統合を使用すると、ヒューマンインザループをサポートしつつ、障害から自動的に復旧する、耐久性のある長時間実行エージェントを実行できます。Dapr はベンダー中立の [CNCF](https://cncf.io) ワークフローオーケストレーターです。Dapr と OpenAI エージェントの利用開始は [こちら](https://docs.diagrid.io/getting-started/quickstarts/ai-agents/?agentframework=openai) から行えます。
|
||||
|
||||
### Temporal
|
||||
|
||||
Agents SDK の [Temporal](https://temporal.io/) 連携を使うと、human-in-the-loop タスクを含む耐久的な長時間ワークフローを実行できます。Temporal と Agents SDK が連携して長時間タスクを完了するデモは [この動画](https://www.youtube.com/watch?v=fFBZqzT4DD8) を参照し、[ドキュメントはこちら](https://github.com/temporalio/sdk-python/tree/main/temporalio/contrib/openai_agents) です。
|
||||
Agents SDK の [Temporal](https://temporal.io/) 統合を使用すると、ヒューマンインザループタスクを含む、耐久性のある長時間実行ワークフローを実行できます。Temporal と Agents SDK が連携して長時間実行タスクを完了するデモは [この動画](https://www.youtube.com/watch?v=fFBZqzT4DD8) で確認でき、[ドキュメントはこちら](https://github.com/temporalio/sdk-python/tree/main/temporalio/contrib/openai_agents) で参照できます。
|
||||
|
||||
### Restate
|
||||
|
||||
Agents SDK の [Restate](https://restate.dev/) 連携を使うと、human approval、ハンドオフ、セッション管理を含む軽量で耐久性のあるエージェントを利用できます。この連携は依存関係として Restate の single-binary runtime を必要とし、プロセス/コンテナまたはサーバーレス関数としてエージェント実行をサポートします。
|
||||
詳細は [概要](https://www.restate.dev/blog/durable-orchestration-for-ai-agents-with-restate-and-openai-sdk) または [ドキュメント](https://docs.restate.dev/ai) を参照してください。
|
||||
Agents SDK の [Restate](https://restate.dev/) 統合は、人による承認、ハンドオフ、セッション管理を含む、軽量で耐久性のあるエージェントに使用できます。この統合では Restate の単一バイナリランタイムが依存関係として必要であり、エージェントをプロセス/コンテナーまたはサーバーレス関数として実行することをサポートしています。詳細については、[概要](https://www.restate.dev/blog/durable-orchestration-for-ai-agents-with-restate-and-openai-sdk) を読むか、[ドキュメント](https://docs.restate.dev/ai) を参照してください。
|
||||
|
||||
### DBOS
|
||||
|
||||
Agents SDK の [DBOS](https://dbos.dev/) 連携を使うと、障害や再起動をまたいで進捗を保持する信頼性の高いエージェントを実行できます。長時間実行エージェント、human-in-the-loop ワークフロー、ハンドオフをサポートします。同期/非同期メソッドの両方に対応しています。この連携に必要なのは SQLite または Postgres データベースのみです。詳細は連携 [repo](https://github.com/dbos-inc/dbos-openai-agents) と [ドキュメント](https://docs.dbos.dev/integrations/openai-agents) を参照してください。
|
||||
Agents SDK の [DBOS](https://dbos.dev/) 統合を使用すると、障害や再起動をまたいで進行状況を保持する信頼性の高いエージェントを実行できます。長時間実行エージェント、ヒューマンインザループワークフロー、ハンドオフをサポートしています。同期メソッドと非同期メソッドの両方をサポートしています。この統合に必要なのは SQLite または Postgres データベースのみです。詳細については、統合の [repo](https://github.com/dbos-inc/dbos-openai-agents) と [ドキュメント](https://docs.dbos.dev/integrations/openai-agents) を参照してください。
|
||||
|
||||
## 例外
|
||||
|
||||
SDK は特定のケースで例外を送出します。完全な一覧は [`agents.exceptions`][] にあります。概要は次のとおりです。
|
||||
SDK は特定の場合に例外を発生させます。完全な一覧は [`agents.exceptions`][] にあります。概要は次のとおりです。
|
||||
|
||||
- [`AgentsException`][agents.exceptions.AgentsException]: SDK 内で発生するすべての例外の基底クラスです。他のすべての具体的な例外はこの汎用型から派生します。
|
||||
- [`MaxTurnsExceeded`][agents.exceptions.MaxTurnsExceeded]: エージェント実行が `Runner.run`、`Runner.run_sync`、`Runner.run_streamed` に渡した `max_turns` 制限を超えたときに送出されます。指定された対話ターン数内でタスクを完了できなかったことを示します。
|
||||
- [`ModelBehaviorError`][agents.exceptions.ModelBehaviorError]: 基盤モデル(LLM)が予期しない、または無効な出力を生成したときに発生します。例:
|
||||
- 不正な JSON: モデルがツール呼び出し用、または直接出力で不正な JSON 構造を返した場合(特に特定の `output_type` が定義されている場合)。
|
||||
- 予期しないツール関連の失敗: モデルが期待される方法でツールを使用しない場合
|
||||
- [`ToolTimeoutError`][agents.exceptions.ToolTimeoutError]: 関数ツール呼び出しが設定したタイムアウトを超え、かつツールが `timeout_behavior="raise_exception"` を使用している場合に送出されます。
|
||||
- [`UserError`][agents.exceptions.UserError]: SDK 使用時に(SDK を使ったコードを書く人が)誤りをした場合に送出されます。通常は不正なコード実装、無効な設定、または SDK API の誤用が原因です。
|
||||
- [`InputGuardrailTripwireTriggered`][agents.exceptions.InputGuardrailTripwireTriggered], [`OutputGuardrailTripwireTriggered`][agents.exceptions.OutputGuardrailTripwireTriggered]: それぞれ入力ガードレールまたは出力ガードレールの条件が満たされたときに送出されます。入力ガードレールは処理前の受信メッセージを検査し、出力ガードレールは配信前のエージェント最終応答を検査します。
|
||||
- [`AgentsException`][agents.exceptions.AgentsException]: これは SDK 内で発生するすべての例外の基底クラスです。他のすべての具体的な例外が派生する汎用型として機能します。
|
||||
- [`MaxTurnsExceeded`][agents.exceptions.MaxTurnsExceeded]: この例外は、エージェントの実行が `Runner.run`、`Runner.run_sync`、または `Runner.run_streamed` メソッドに渡された `max_turns` 制限を超えた場合に発生します。これは、指定された対話ターン数の範囲内でエージェントがタスクを完了できなかったことを示します。制限を無効にするには `max_turns=None` を設定します。
|
||||
- [`ModelBehaviorError`][agents.exceptions.ModelBehaviorError]: この例外は、基盤モデル(LLM)が予期しない出力または無効な出力を生成した場合に発生します。これには次が含まれます。
|
||||
- 不正な形式の JSON: モデルがツール呼び出しまたは直接出力で不正な形式の JSON 構造を提供した場合です。特に特定の `output_type` が定義されている場合に該当します。
|
||||
- 予期しないツール関連の失敗: モデルが期待された方法でツールを使用できなかった場合です
|
||||
- [`ToolTimeoutError`][agents.exceptions.ToolTimeoutError]: この例外は、関数ツール呼び出しが設定されたタイムアウトを超え、そのツールが `timeout_behavior="raise_exception"` を使用している場合に発生します。
|
||||
- [`UserError`][agents.exceptions.UserError]: この例外は、あなた(SDK を使用してコードを書く人)が SDK の使用中に誤りを犯した場合に発生します。通常は、不正なコード実装、無効な設定、または SDK の API の誤用が原因です。
|
||||
- [`InputGuardrailTripwireTriggered`][agents.exceptions.InputGuardrailTripwireTriggered], [`OutputGuardrailTripwireTriggered`][agents.exceptions.OutputGuardrailTripwireTriggered]: この例外は、入力ガードレールまたは出力ガードレールの条件がそれぞれ満たされた場合に発生します。入力ガードレールは処理前に受信メッセージをチェックし、出力ガードレールは配信前にエージェントの最終応答をチェックします。
|
||||
+52
-52
@@ -2,42 +2,42 @@
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# Sandbox クライアント
|
||||
# サンドボックスクライアント
|
||||
|
||||
このページでは、 sandbox の作業をどこで実行するかを選択します。ほとんどの場合、 `SandboxAgent` の定義は同じままで、 sandbox クライアントとクライアント固有のオプションのみが [`SandboxRunConfig`][agents.run_config.SandboxRunConfig] で変わります。
|
||||
このページでは、サンドボックスでの作業をどこで実行するかを選択します。ほとんどの場合、`SandboxAgent` の定義は同じままにし、サンドボックスクライアントとクライアント固有のオプションを [`SandboxRunConfig`][agents.run_config.SandboxRunConfig] で変更します。
|
||||
|
||||
!!! warning "ベータ機能"
|
||||
|
||||
sandbox エージェントはベータ版です。一般提供までに API の詳細、デフォルト値、対応機能は変更される可能性があり、今後さらに高度な機能が追加される予定です。
|
||||
サンドボックスエージェントはベータ版です。一般提供までに API の詳細、デフォルト値、サポートされる機能が変更される可能性があります。また、時間とともにより高度な機能が追加される見込みです。
|
||||
|
||||
## 判断ガイド
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| 目的 | 開始時の選択肢 | 理由 |
|
||||
| 目的 | まず使うもの | 理由 |
|
||||
| --- | --- | --- |
|
||||
| macOS または Linux で最速のローカル反復 | `UnixLocalSandboxClient` | 追加インストール不要で、シンプルなローカルファイルシステム開発ができます。 |
|
||||
| macOS または Linux での最速のローカル反復 | `UnixLocalSandboxClient` | 追加インストール不要で、シンプルなローカルファイルシステム開発ができます。 |
|
||||
| 基本的なコンテナ分離 | `DockerSandboxClient` | 特定のイメージを使って Docker 内で作業を実行します。 |
|
||||
| ホスト型実行または本番環境に近い分離 | ホスト型 sandbox クライアント | ワークスペースの境界を、プロバイダー管理の環境に移します。 |
|
||||
| ホスト型実行または本番環境スタイルの分離 | ホスト型サンドボックスクライアント | ワークスペース境界をプロバイダー管理環境へ移します。 |
|
||||
|
||||
</div>
|
||||
|
||||
## ローカルクライアント
|
||||
|
||||
ほとんどのユーザーは、まず次の 2 つの sandbox クライアントのいずれかから始めてください。
|
||||
ほとんどのユーザーは、これら 2 つのサンドボックスクライアントのいずれかから始めることをおすすめします。
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| クライアント | インストール | 選ぶタイミング | 例 |
|
||||
| クライアント | インストール | 選ぶ場面 | 例 |
|
||||
| --- | --- | --- | --- |
|
||||
| `UnixLocalSandboxClient` | なし | macOS または Linux で最速のローカル反復が必要な場合。ローカル開発のよいデフォルトです。 | [Unix-local スターター](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/unix_local_runner.py) |
|
||||
| `DockerSandboxClient` | `openai-agents[docker]` | コンテナ分離や、ローカル環境の整合性のために特定のイメージが必要な場合。 | [Docker スターター](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/docker/docker_runner.py) |
|
||||
| `UnixLocalSandboxClient` | なし | macOS または Linux で最速のローカル反復が必要な場合。ローカル開発の既定として適しています。 | [Unix-local スターター](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/unix_local_runner.py) |
|
||||
| `DockerSandboxClient` | `openai-agents[docker]` | コンテナ分離、またはローカルで同等性を保つための特定のイメージが必要な場合。 | [Docker スターター](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/docker/docker_runner.py) |
|
||||
|
||||
</div>
|
||||
|
||||
Unix-local は、ローカルファイルシステムを対象に開発を始める最も簡単な方法です。より強い環境分離や本番環境に近い整合性が必要になったら、 Docker またはホスト型プロバイダーに移行してください。
|
||||
Unix-local は、ローカルファイルシステムに対して開発を始める最も簡単な方法です。より強い環境分離や本番環境スタイルの同等性が必要になったら、Docker またはホスト型プロバイダーへ移行してください。
|
||||
|
||||
Unix-local から Docker に切り替えるには、エージェント定義はそのままにして、 run config のみを変更します。
|
||||
Unix-local から Docker に切り替えるには、エージェント定義は同じままにして、実行設定だけを変更します。
|
||||
|
||||
```python
|
||||
from docker import from_env as docker_from_env
|
||||
@@ -54,41 +54,41 @@ run_config = RunConfig(
|
||||
)
|
||||
```
|
||||
|
||||
これは、コンテナ分離またはイメージの整合性が必要な場合に使用します。[examples/sandbox/docker/docker_runner.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/docker/docker_runner.py) を参照してください。
|
||||
コンテナ分離またはイメージの同等性が必要な場合に使用してください。[examples/sandbox/docker/docker_runner.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/docker/docker_runner.py) を参照してください。
|
||||
|
||||
## マウントとリモートストレージ
|
||||
|
||||
mount エントリーは公開するストレージを記述し、 mount strategy は sandbox バックエンドがそのストレージをどのように接続するかを記述します。組み込みの mount エントリーと汎用 strategy は `agents.sandbox.entries` から import します。ホスト型プロバイダーの strategy は `agents.extensions.sandbox` またはプロバイダー固有の拡張パッケージから利用できます。
|
||||
マウントエントリーはどのストレージを公開するかを表し、マウント戦略はサンドボックスバックエンドがそのストレージをどのようにアタッチするかを表します。組み込みのマウントエントリーと汎用戦略は `agents.sandbox.entries` からインポートします。ホスト型プロバイダーの戦略は `agents.extensions.sandbox` またはプロバイダー固有の拡張パッケージから利用できます。
|
||||
|
||||
一般的な mount オプション:
|
||||
一般的なマウントオプション:
|
||||
|
||||
- `mount_path`: sandbox 内でストレージが現れる場所です。相対パスは manifest ルート配下で解決され、絶対パスはそのまま使用されます。
|
||||
- `read_only`: デフォルトは `True` です。 sandbox がマウントされたストレージに書き戻す必要がある場合にのみ `False` を設定してください。
|
||||
- `mount_strategy`: 必須です。 mount エントリーと sandbox バックエンドの両方に一致する strategy を使用してください。
|
||||
- `mount_path`: ストレージがサンドボックス内で表示される場所です。相対パスはマニフェストルート配下で解決され、絶対パスはそのまま使用されます。
|
||||
- `read_only`: 既定は `True` です。サンドボックスがマウントされたストレージへ書き戻す必要がある場合にのみ `False` に設定してください。
|
||||
- `mount_strategy`: 必須です。マウントエントリーとサンドボックスバックエンドの両方に合う戦略を使用してください。
|
||||
|
||||
mount は一時的なワークスペースエントリーとして扱われます。スナップショットと永続化のフローでは、マウントされたリモートストレージを保存済みワークスペースにコピーするのではなく、マウントされたパスを切り離すかスキップします。
|
||||
マウントは一時的なワークスペースエントリーとして扱われます。スナップショットと永続化のフローでは、マウントされたリモートストレージを保存済みワークスペースへコピーするのではなく、マウントされたパスをデタッチするかスキップします。
|
||||
|
||||
汎用のローカル / コンテナ strategy:
|
||||
汎用ローカル / コンテナ戦略:
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| Strategy またはパターン | 使用するタイミング | 注記 |
|
||||
| 戦略またはパターン | 使用する場面 | 備考 |
|
||||
| --- | --- | --- |
|
||||
| `InContainerMountStrategy(pattern=RcloneMountPattern(...))` | sandbox イメージで `rclone` を実行できる場合。 | S3 、 GCS 、 R2 、 Azure Blob をサポートします。`RcloneMountPattern` は `fuse` モードまたは `nfs` モードで実行できます。 |
|
||||
| `InContainerMountStrategy(pattern=MountpointMountPattern(...))` | イメージに `mount-s3` があり、 Mountpoint スタイルの S3 または S3 互換アクセスが必要な場合。 | `S3Mount` と `GCSMount` をサポートします。 |
|
||||
| `InContainerMountStrategy(pattern=FuseMountPattern(...))` | イメージに `blobfuse2` と FUSE サポートがある場合。 | `AzureBlobMount` をサポートします。 |
|
||||
| `InContainerMountStrategy(pattern=S3FilesMountPattern(...))` | イメージに `mount.s3files` があり、既存の S3 Files マウント先に到達できる場合。 | `S3FilesMount` をサポートします。 |
|
||||
| `DockerVolumeMountStrategy(driver=...)` | コンテナ起動前に Docker が volume-driver ベースのマウントを接続すべき場合。 | Docker 専用です。 S3 、 GCS 、 R2 、 Azure Blob は `rclone` をサポートし、 S3 と GCS は `mountpoint` もサポートします。 |
|
||||
| `InContainerMountStrategy(pattern=RcloneMountPattern(...))` | サンドボックスイメージで `rclone` を実行できる場合。 | S3、GCS、R2、Azure Blob、Box をサポートします。`RcloneMountPattern` は `fuse` モードまたは `nfs` モードで実行できます。 |
|
||||
| `InContainerMountStrategy(pattern=MountpointMountPattern(...))` | イメージに `mount-s3` があり、Mountpoint スタイルの S3 または S3 互換アクセスが必要な場合。 | `S3Mount` と `GCSMount` をサポートします。 |
|
||||
| `InContainerMountStrategy(pattern=FuseMountPattern(...))` | イメージに `blobfuse2` があり、FUSE サポートがある場合。 | `AzureBlobMount` をサポートします。 |
|
||||
| `InContainerMountStrategy(pattern=S3FilesMountPattern(...))` | イメージに `mount.s3files` があり、既存の S3 Files マウントターゲットに到達できる場合。 | `S3FilesMount` をサポートします。 |
|
||||
| `DockerVolumeMountStrategy(driver=...)` | Docker がコンテナ起動前にボリュームドライバー対応のマウントをアタッチする必要がある場合。 | Docker のみです。`rclone` は S3、GCS、R2、Azure Blob、Box をサポートし、`mountpoint` は S3 と GCS もサポートします。 |
|
||||
|
||||
</div>
|
||||
|
||||
## 対応するホスト型プラットフォーム
|
||||
## サポートされるホスト型プラットフォーム
|
||||
|
||||
ホスト型環境が必要な場合、通常は同じ `SandboxAgent` の定義をそのまま使え、 [`SandboxRunConfig`][agents.run_config.SandboxRunConfig] で sandbox クライアントだけを変更します。
|
||||
ホスト型環境が必要な場合、通常は同じ `SandboxAgent` 定義をそのまま引き継ぎ、[`SandboxRunConfig`][agents.run_config.SandboxRunConfig] でサンドボックスクライアントだけを変更します。
|
||||
|
||||
このリポジトリのチェックアウト版ではなく公開済み SDK を使用している場合は、対応する package extra を通じて sandbox-client の依存関係をインストールしてください。
|
||||
このリポジトリのチェックアウトではなく公開されている SDK を使用している場合は、対応するパッケージ extra を通じてサンドボックスクライアントの依存関係をインストールしてください。
|
||||
|
||||
プロバイダー固有のセットアップに関する注意事項と、リポジトリに含まれている拡張コード例へのリンクについては、[examples/sandbox/extensions/README.md](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/extensions/README.md) を参照してください。
|
||||
プロバイダー固有のセットアップメモと、チェックイン済みの拡張コード例へのリンクについては、[examples/sandbox/extensions/README.md](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/extensions/README.md) を参照してください。
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
@@ -104,38 +104,38 @@ mount は一時的なワークスペースエントリーとして扱われま
|
||||
|
||||
</div>
|
||||
|
||||
ホスト型 sandbox クライアントは、プロバイダー固有の mount strategy を公開しています。ストレージプロバイダーに最も適したバックエンドと mount strategy を選択してください。
|
||||
ホスト型サンドボックスクライアントは、プロバイダー固有のマウント戦略を公開します。ストレージプロバイダーに最も合うバックエンドとマウント戦略を選択してください。
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| バックエンド | mount に関する注記 |
|
||||
| バックエンド | マウントに関する注記 |
|
||||
| --- | --- |
|
||||
| Docker | `InContainerMountStrategy` や `DockerVolumeMountStrategy` などのローカル strategy を使って、 `S3Mount` 、 `GCSMount` 、 `R2Mount` 、 `AzureBlobMount` 、 `S3FilesMount` をサポートします。 |
|
||||
| `ModalSandboxClient` | `S3Mount` 、 `R2Mount` 、 HMAC 認証付き `GCSMount` に対して、 `ModalCloudBucketMountStrategy` による Modal cloud bucket mount をサポートします。インライン認証情報または名前付きの Modal Secret を使用できます。 |
|
||||
| `CloudflareSandboxClient` | `S3Mount` 、 `R2Mount` 、 HMAC 認証付き `GCSMount` に対して、 `CloudflareBucketMountStrategy` による Cloudflare bucket mount をサポートします。 |
|
||||
| `BlaxelSandboxClient` | `S3Mount` 、 `R2Mount` 、 `GCSMount` に対して、 `BlaxelCloudBucketMountStrategy` による cloud bucket mount をサポートします。また、 `agents.extensions.sandbox.blaxel` の `BlaxelDriveMount` と `BlaxelDriveMountStrategy` による永続的な Blaxel Drive もサポートします。 |
|
||||
| `DaytonaSandboxClient` | `DaytonaCloudBucketMountStrategy` による cloud bucket mount をサポートします。`S3Mount` 、 `GCSMount` 、 `R2Mount` 、 `AzureBlobMount` と組み合わせて使用します。 |
|
||||
| `E2BSandboxClient` | `E2BCloudBucketMountStrategy` による cloud bucket mount をサポートします。`S3Mount` 、 `GCSMount` 、 `R2Mount` 、 `AzureBlobMount` と組み合わせて使用します。 |
|
||||
| `RunloopSandboxClient` | `RunloopCloudBucketMountStrategy` による cloud bucket mount をサポートします。`S3Mount` 、 `GCSMount` 、 `R2Mount` 、 `AzureBlobMount` と組み合わせて使用します。 |
|
||||
| `VercelSandboxClient` | 現時点ではホスト型固有の mount strategy は公開されていません。代わりに manifest ファイル、リポジトリ、またはその他のワークスペース入力を使用してください。 |
|
||||
| Docker | `InContainerMountStrategy` や `DockerVolumeMountStrategy` などのローカル戦略で、`S3Mount`、`GCSMount`、`R2Mount`、`AzureBlobMount`、`BoxMount`、`S3FilesMount` をサポートします。 |
|
||||
| `ModalSandboxClient` | `S3Mount`、`R2Mount`、HMAC 認証済みの `GCSMount` で、`ModalCloudBucketMountStrategy` による Modal のクラウドバケットマウントをサポートします。インライン認証情報、または名前付きの Modal Secret を使用できます。 |
|
||||
| `CloudflareSandboxClient` | `S3Mount`、`R2Mount`、HMAC 認証済みの `GCSMount` で、`CloudflareBucketMountStrategy` による Cloudflare バケットマウントをサポートします。 |
|
||||
| `BlaxelSandboxClient` | `S3Mount`、`R2Mount`、`GCSMount` で、`BlaxelCloudBucketMountStrategy` によるクラウドバケットマウントをサポートします。`agents.extensions.sandbox.blaxel` の `BlaxelDriveMount` と `BlaxelDriveMountStrategy` による永続的な Blaxel Drives もサポートします。 |
|
||||
| `DaytonaSandboxClient` | `DaytonaCloudBucketMountStrategy` による `rclone` ベースのクラウドストレージマウントをサポートします。`S3Mount`、`GCSMount`、`R2Mount`、`AzureBlobMount`、`BoxMount` と組み合わせて使用してください。 |
|
||||
| `E2BSandboxClient` | `E2BCloudBucketMountStrategy` による `rclone` ベースのクラウドストレージマウントをサポートします。`S3Mount`、`GCSMount`、`R2Mount`、`AzureBlobMount`、`BoxMount` と組み合わせて使用してください。 |
|
||||
| `RunloopSandboxClient` | `RunloopCloudBucketMountStrategy` による `rclone` ベースのクラウドストレージマウントをサポートします。`S3Mount`、`GCSMount`、`R2Mount`、`AzureBlobMount`、`BoxMount` と組み合わせて使用してください。 |
|
||||
| `VercelSandboxClient` | 現時点ではホスト型固有のマウント戦略は公開されていません。代わりにマニフェストファイル、リポジトリ、またはその他のワークスペース入力を使用してください。 |
|
||||
|
||||
</div>
|
||||
|
||||
以下の表は、各バックエンドがどのリモートストレージエントリーを直接マウントできるかをまとめたものです。
|
||||
以下の表は、各バックエンドが直接マウントできるリモートストレージエントリーをまとめたものです。
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| バックエンド | AWS S3 | Cloudflare R2 | GCS | Azure Blob Storage | S3 Files |
|
||||
| --- | --- | --- | --- | --- | --- |
|
||||
| Docker | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| `ModalSandboxClient` | ✓ | ✓ | ✓ | - | - |
|
||||
| `CloudflareSandboxClient` | ✓ | ✓ | ✓ | - | - |
|
||||
| `BlaxelSandboxClient` | ✓ | ✓ | ✓ | - | - |
|
||||
| `DaytonaSandboxClient` | ✓ | ✓ | ✓ | ✓ | - |
|
||||
| `E2BSandboxClient` | ✓ | ✓ | ✓ | ✓ | - |
|
||||
| `RunloopSandboxClient` | ✓ | ✓ | ✓ | ✓ | - |
|
||||
| `VercelSandboxClient` | - | - | - | - | - |
|
||||
| バックエンド | AWS S3 | Cloudflare R2 | GCS | Azure Blob Storage | Box | S3 Files |
|
||||
| --- | --- | --- | --- | --- | --- | --- |
|
||||
| Docker | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| `ModalSandboxClient` | ✓ | ✓ | ✓ | - | - | - |
|
||||
| `CloudflareSandboxClient` | ✓ | ✓ | ✓ | - | - | - |
|
||||
| `BlaxelSandboxClient` | ✓ | ✓ | ✓ | - | - | - |
|
||||
| `DaytonaSandboxClient` | ✓ | ✓ | ✓ | ✓ | ✓ | - |
|
||||
| `E2BSandboxClient` | ✓ | ✓ | ✓ | ✓ | ✓ | - |
|
||||
| `RunloopSandboxClient` | ✓ | ✓ | ✓ | ✓ | ✓ | - |
|
||||
| `VercelSandboxClient` | - | - | - | - | - | - |
|
||||
|
||||
</div>
|
||||
|
||||
さらに実行可能な例については、ローカル、コーディング、メモリ、ハンドオフ、エージェント合成のパターンは [examples/sandbox/](https://github.com/openai/openai-agents-python/tree/main/examples/sandbox) を参照し、ホスト型 sandbox クライアントについては [examples/sandbox/extensions/](https://github.com/openai/openai-agents-python/tree/main/examples/sandbox/extensions) を参照してください。
|
||||
実行可能なコード例をさらに見るには、ローカル、コーディング、メモリ、ハンドオフ、エージェント合成パターンについては [examples/sandbox/](https://github.com/openai/openai-agents-python/tree/main/examples/sandbox) を、ホスト型サンドボックスクライアントについては [examples/sandbox/extensions/](https://github.com/openai/openai-agents-python/tree/main/examples/sandbox/extensions) を参照してください。
|
||||
+226
-216
@@ -4,37 +4,37 @@ search:
|
||||
---
|
||||
# 概念
|
||||
|
||||
!!! warning "Beta 機能"
|
||||
!!! warning "ベータ機能"
|
||||
|
||||
Sandbox Agents は beta です。API の詳細、デフォルト値、対応機能は一般提供前に変更される可能性があり、時間の経過とともにより高度な機能が追加される予定です。
|
||||
サンドボックスエージェントはベータ版です。一般提供までに API の詳細、デフォルト、サポートされる機能が変更される可能性があり、時間とともにより高度な機能が追加される予定です。
|
||||
|
||||
現代的なエージェントは、ファイルシステム上の実際のファイルを操作できると最も効果的に動作します。**Sandbox Agents** は、特化したツールやシェルコマンドを利用して、大規模なドキュメント集合の検索や操作、ファイル編集、成果物の生成、コマンド実行を行えます。sandbox は、モデルに永続的なワークスペースを提供し、エージェントがユーザーに代わって作業できるようにします。Agents SDK の Sandbox Agents は、sandbox 環境と組み合わせたエージェントの実行を容易にし、適切なファイルをファイルシステム上に配置し、sandbox をオーケストレーションして、大規模にタスクを開始、停止、再開しやすくします。
|
||||
現代的なエージェントは、ファイルシステム上の実ファイルを操作できる場合に最もよく機能します。 **サンドボックスエージェント** は、専用ツールとシェルコマンドを利用して、大規模なドキュメントセットの検索や操作、ファイル編集、成果物の生成、コマンド実行を行えます。サンドボックスは、エージェントがユーザーに代わって作業するために使用できる永続的なワークスペースをモデルに提供します。Agents SDK のサンドボックスエージェントは、サンドボックス環境とペアになったエージェントを簡単に実行できるようにし、ファイルシステム上に適切なファイルを配置し、サンドボックスをオーケストレーションして、大規模にタスクを開始、停止、再開しやすくします。
|
||||
|
||||
エージェントに必要なデータを中心にワークスペースを定義します。GitHub リポジトリ、ローカルファイルやディレクトリ、合成されたタスクファイル、S3 や Azure Blob Storage などのリモートファイルシステム、その他ユーザーが提供する sandbox 入力から開始できます。
|
||||
エージェントが必要とするデータを中心にワークスペースを定義します。GitHub リポジトリ、ローカルファイルとディレクトリ、合成タスクファイル、S3 や Azure Blob Storage などのリモートファイルシステム、およびユーザーが提供するその他のサンドボックス入力から開始できます。
|
||||
|
||||
<div class="sandbox-harness-image" markdown="1">
|
||||
|
||||

|
||||

|
||||
|
||||
</div>
|
||||
|
||||
`SandboxAgent` は依然として `Agent` です。`instructions`、`prompt`、`tools`、`handoffs`、`mcp_servers`、`model_settings`、`output_type`、ガードレール、hooks など、通常のエージェントのインターフェースを維持し、通常の `Runner` API を通じて実行されます。変わるのは実行境界です。
|
||||
`SandboxAgent` は引き続き `Agent` です。`instructions`、`prompt`、`tools`、`handoffs`、`mcp_servers`、`model_settings`、`output_type`、ガードレール、フックといった通常のエージェントサーフェスを保持し、通常の `Runner` API を通じて実行されます。変わるのは実行境界です。
|
||||
|
||||
- `SandboxAgent` はエージェント自体を定義します。通常のエージェント設定に加えて、`default_manifest`、`base_instructions`、`run_as` といった sandbox 固有のデフォルトや、ファイルシステムツール、シェルアクセス、skills、memory、compaction などの機能を含みます。
|
||||
- `Manifest` は、新しい sandbox ワークスペースの望ましい初期内容とレイアウトを宣言します。これには、ファイル、リポジトリ、mount、環境が含まれます。
|
||||
- sandbox session は、コマンドが実行され、ファイルが変更されるライブな分離環境です。
|
||||
- [`SandboxRunConfig`][agents.run_config.SandboxRunConfig] は、その実行がどのように sandbox session を取得するかを決定します。たとえば、直接注入する、シリアライズ済みの sandbox session state から再接続する、sandbox client を通じて新しい sandbox session を作成する、などです。
|
||||
- 保存された sandbox state と snapshots により、後続の実行で以前の作業に再接続したり、保存済みの内容から新しい sandbox session を初期化したりできます。
|
||||
- `SandboxAgent` はエージェント自体を定義します。通常のエージェント設定に加え、`default_manifest`、`base_instructions`、`run_as` などのサンドボックス固有のデフォルト、およびファイルシステムツール、シェルアクセス、スキル、メモリ、コンパクションなどの機能を含みます。
|
||||
- `Manifest` は、新しいサンドボックスワークスペースの望ましい初期コンテンツとレイアウトを宣言します。これにはファイル、リポジトリ、マウント、環境が含まれます。
|
||||
- サンドボックスセッションは、コマンドが実行されファイルが変更される、ライブの隔離環境です。
|
||||
- [`SandboxRunConfig`][agents.run_config.SandboxRunConfig] は、その実行がサンドボックスセッションをどのように取得するかを決定します。たとえば、直接注入する、シリアライズ済みのサンドボックスセッション状態から再接続する、サンドボックスクライアントを通じて新しいサンドボックスセッションを作成する、などです。
|
||||
- 保存されたサンドボックス状態とスナップショットにより、後続の実行で以前の作業へ再接続したり、保存済みコンテンツから新しいサンドボックスセッションを初期化したりできます。
|
||||
|
||||
`Manifest` は新規 session 用ワークスペースの契約であり、すべてのライブ sandbox の完全な正本ではありません。実行時の実効ワークスペースは、再利用される sandbox session、シリアライズ済み sandbox session state、または実行時に選択された snapshot から決まることがあります。
|
||||
`Manifest` は新規セッションのワークスペース契約であり、すべてのライブサンドボックスに対する完全な信頼できる情報源ではありません。ある実行における実効ワークスペースは、再利用されたサンドボックスセッション、シリアライズ済みのサンドボックスセッション状態、または実行時に選択されたスナップショットから得られる場合があります。
|
||||
|
||||
このページ全体でいう "sandbox session" とは、sandbox client が管理するライブ実行環境を指します。これは [Sessions](../sessions/index.md) で説明されている SDK の会話用 [`Session`][agents.memory.session.Session] インターフェースとは異なります。
|
||||
このページ全体で、「サンドボックスセッション」とは、サンドボックスクライアントによって管理されるライブ実行環境を意味します。これは、[Sessions](../sessions/index.md) で説明されている SDK の会話用 [`Session`][agents.memory.session.Session] インターフェイスとは異なります。
|
||||
|
||||
外側のランタイムは引き続き approvals、トレーシング、ハンドオフ、再開 bookkeeping を管理します。sandbox session はコマンド、ファイル変更、環境分離を管理します。この分離はモデルの中核的な部分です。
|
||||
外側のランタイムは引き続き、承認、トレーシング、ハンドオフ、再開のブックキーピングを所有します。サンドボックスセッションは、コマンド、ファイル変更、環境の隔離を所有します。この分離はモデルの中核的な部分です。
|
||||
|
||||
### 構成要素の適合
|
||||
### 各要素の関係
|
||||
|
||||
sandbox 実行は、エージェント定義と実行ごとの sandbox 設定を組み合わせたものです。runner はエージェントを準備し、ライブな sandbox session にバインドし、後続の実行のために state を保存できます。
|
||||
サンドボックス実行は、エージェント定義と実行ごとのサンドボックス設定を組み合わせます。Runner はエージェントを準備し、それをライブサンドボックスセッションにバインドし、後続の実行のために状態を保存できます。
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
@@ -50,198 +50,198 @@ flowchart LR
|
||||
sandbox --> saved
|
||||
```
|
||||
|
||||
sandbox 固有のデフォルトは `SandboxAgent` に置かれます。実行ごとの sandbox-session の選択は `SandboxRunConfig` に置かれます。
|
||||
サンドボックス固有のデフォルトは `SandboxAgent` に保持します。実行ごとのサンドボックスセッション選択は `SandboxRunConfig` に保持します。
|
||||
|
||||
ライフサイクルは 3 つのフェーズで考えるとよいです。
|
||||
ライフサイクルは 3 つのフェーズで考えてください。
|
||||
|
||||
1. `SandboxAgent`、`Manifest`、および capabilities で、エージェントと新規ワークスペース契約を定義します。
|
||||
2. `Runner` に `SandboxRunConfig` を渡して sandbox session を注入、再開、または作成し、実行します。
|
||||
3. runner 管理の `RunState`、明示的な sandbox `session_state`、または保存済みワークスペース snapshot から後で継続します。
|
||||
1. `SandboxAgent`、`Manifest`、機能を使って、エージェントと新規ワークスペース契約を定義します。
|
||||
2. `Runner` に、サンドボックスセッションを注入、再開、または作成する `SandboxRunConfig` を渡して実行します。
|
||||
3. Runner 管理の `RunState`、明示的なサンドボックス `session_state`、または保存済みワークスペーススナップショットから後で継続します。
|
||||
|
||||
シェルアクセスがたまに使うツールの 1 つにすぎない場合は、まず [tools ガイド](../tools.md) の hosted shell を使ってください。ワークスペース分離、sandbox client の選択、sandbox-session の再開動作が設計の一部である場合に sandbox agents を使ってください。
|
||||
シェルアクセスが一時的に使うツールの 1 つにすぎない場合は、[tools guide](../tools.md) のホスト型シェルから始めてください。ワークスペースの隔離、サンドボックスクライアントの選択、またはサンドボックスセッションの再開動作が設計の一部である場合は、サンドボックスエージェントを使用してください。
|
||||
|
||||
## 利用場面
|
||||
## 利用する場面
|
||||
|
||||
sandbox agents は、たとえば次のようなワークスペース中心のワークフローに適しています。
|
||||
サンドボックスエージェントは、ワークスペース中心のワークフローに適しています。たとえば次のような場合です。
|
||||
|
||||
- コーディングやデバッグ。たとえば、GitHub リポジトリの issue レポートに対する自動修正をエージェントオーケストレーションし、対象を絞ったテストを実行する
|
||||
- ドキュメント処理や編集。たとえば、ユーザーの財務書類から情報を抽出し、記入済みの納税フォーム草案を作成する
|
||||
- ファイルに基づくレビューや分析。たとえば、オンボーディング資料、生成されたレポート、成果物バンドルを確認してから回答する
|
||||
- 分離されたマルチエージェントパターン。たとえば、各レビュー担当またはコーディング用サブエージェントに専用ワークスペースを与える
|
||||
- 複数段階のワークスペースタスク。たとえば、ある実行でバグを修正し、後でリグレッションテストを追加する、または snapshot や sandbox session state から再開する
|
||||
- コーディングとデバッグ。たとえば、GitHub リポジトリ内の Issue レポートに対する自動修正をオーケストレーションし、対象テストを実行する場合
|
||||
- ドキュメント処理と編集。たとえば、ユーザーの財務書類から情報を抽出し、完成した税務フォームのドラフトを作成する場合
|
||||
- ファイルに基づくレビューや分析。たとえば、オンボーディングパケット、生成されたレポート、成果物バンドルを確認してから回答する場合
|
||||
- 隔離されたマルチエージェントパターン。たとえば、各レビュアーやコーディングサブエージェントに独自のワークスペースを与える場合
|
||||
- 複数ステップのワークスペースタスク。たとえば、ある実行でバグを修正し、後で回帰テストを追加する場合や、スナップショットまたはサンドボックスセッション状態から再開する場合
|
||||
|
||||
ファイルや動的なファイルシステムへのアクセスが不要であれば、引き続き `Agent` を使ってください。シェルアクセスがたまに必要な機能であれば hosted shell を追加してください。ワークスペース境界自体が機能の一部であるなら sandbox agents を使ってください。
|
||||
ファイルや動的なファイルシステムへのアクセスが不要な場合は、引き続き `Agent` を使用してください。シェルアクセスが一時的な機能にすぎない場合は、ホスト型シェルを追加します。ワークスペース境界自体が機能の一部である場合は、サンドボックスエージェントを使用します。
|
||||
|
||||
## sandbox client の選択
|
||||
## サンドボックスクライアントの選択
|
||||
|
||||
ローカル開発では `UnixLocalSandboxClient` から始めてください。コンテナ分離やイメージの一致が必要になったら `DockerSandboxClient` に移行してください。プロバイダー管理の実行が必要なら hosted provider に移行してください。
|
||||
ローカル開発では `UnixLocalSandboxClient` から始めてください。コンテナ隔離やイメージの同等性が必要になったら `DockerSandboxClient` に移行します。プロバイダー管理の実行が必要な場合は、ホスト型プロバイダーに移行します。
|
||||
|
||||
ほとんどの場合、`SandboxAgent` の定義は同じままで、変更されるのは sandbox client とそのオプションだけであり、それらは [`SandboxRunConfig`][agents.run_config.SandboxRunConfig] で指定します。ローカル、Docker、hosted、remote-mount の各オプションについては [Sandbox clients](clients.md) を参照してください。
|
||||
ほとんどの場合、[`SandboxRunConfig`][agents.run_config.SandboxRunConfig] 内のサンドボックスクライアントとそのオプションが変わっても、`SandboxAgent` 定義は同じままです。ローカル、Docker、ホスト型、リモートマウントの各オプションについては、[Sandbox clients](clients.md) を参照してください。
|
||||
|
||||
## 中核要素
|
||||
## 主要な構成要素
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| Layer | Main SDK pieces | What it answers |
|
||||
| レイヤー | 主な SDK 構成要素 | 回答する内容 |
|
||||
| --- | --- | --- |
|
||||
| エージェント定義 | `SandboxAgent`、`Manifest`、capabilities | どのエージェントが実行され、新規 session のワークスペース契約は何から開始されるべきですか。 |
|
||||
| sandbox 実行 | `SandboxRunConfig`、sandbox client、ライブな sandbox session | この実行はどのようにライブな sandbox session を取得し、どこで作業が実行されますか。 |
|
||||
| 保存された sandbox state | `RunState` の sandbox payload、`session_state`、snapshots | このワークフローはどのように以前の sandbox 作業に再接続するか、または保存済み内容から新しい sandbox session を初期化しますか。 |
|
||||
| エージェント定義 | `SandboxAgent`、`Manifest`、機能 | どのエージェントを実行し、どの新規セッションのワークスペース契約から開始すべきか。 |
|
||||
| サンドボックス実行 | `SandboxRunConfig`、サンドボックスクライアント、ライブサンドボックスセッション | この実行はどのようにライブサンドボックスセッションを取得し、作業はどこで実行されるか。 |
|
||||
| 保存済みサンドボックス状態 | `RunState` サンドボックスペイロード、`session_state`、スナップショット | このワークフローはどのように以前のサンドボックス作業へ再接続するか、または保存済みコンテンツから新しいサンドボックスセッションを初期化するか。 |
|
||||
|
||||
</div>
|
||||
|
||||
主な SDK 要素は、これらのレイヤーに次のように対応します。
|
||||
主な SDK 構成要素は、次のようにこれらのレイヤーに対応します。
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| Piece | What it owns | Ask this question |
|
||||
| 構成要素 | 所有するもの | この質問をします |
|
||||
| --- | --- | --- |
|
||||
| [`SandboxAgent`][agents.sandbox.sandbox_agent.SandboxAgent] | エージェント定義 | このエージェントは何を行うべきで、どのデフォルトを一緒に持ち運ぶべきですか。 |
|
||||
| [`Manifest`][agents.sandbox.manifest.Manifest] | 新規 session のワークスペースファイルとフォルダー | 実行開始時に、どのファイルとフォルダーがファイルシステム上に存在すべきですか。 |
|
||||
| [`Capability`][agents.sandbox.capabilities.capability.Capability] | sandbox ネイティブな動作 | どのツール、instructions の断片、またはランタイム動作をこのエージェントに付与すべきですか。 |
|
||||
| [`SandboxRunConfig`][agents.run_config.SandboxRunConfig] | 実行ごとの sandbox client と sandbox-session の取得元 | この実行では sandbox session を注入、再開、または作成すべきですか。 |
|
||||
| [`RunState`][agents.run_state.RunState] | runner 管理の保存済み sandbox state | 以前の runner 管理ワークフローを再開し、その sandbox state を自動的に引き継いでいますか。 |
|
||||
| [`SandboxRunConfig.session_state`][agents.run_config.SandboxRunConfig.session_state] | 明示的にシリアライズされた sandbox session state | `RunState` の外で既にシリアライズした sandbox state から再開したいですか。 |
|
||||
| [`SandboxRunConfig.snapshot`][agents.run_config.SandboxRunConfig.snapshot] | 新しい sandbox session 用の保存済みワークスペース内容 | 新しい sandbox session を保存済みのファイルや成果物から開始すべきですか。 |
|
||||
| [`SandboxAgent`][agents.sandbox.sandbox_agent.SandboxAgent] | エージェント定義 | このエージェントは何を行うべきか、どのデフォルトを一緒に持たせるべきか。 |
|
||||
| [`Manifest`][agents.sandbox.manifest.Manifest] | 新規セッションのワークスペースファイルとフォルダー | 実行開始時にファイルシステム上にどのファイルとフォルダーが存在すべきか。 |
|
||||
| [`Capability`][agents.sandbox.capabilities.capability.Capability] | サンドボックスネイティブの動作 | どのツール、指示フラグメント、またはランタイム動作をこのエージェントに付与すべきか。 |
|
||||
| [`SandboxRunConfig`][agents.run_config.SandboxRunConfig] | 実行ごとのサンドボックスクライアントとサンドボックスセッションソース | この実行はサンドボックスセッションを注入、再開、または作成すべきか。 |
|
||||
| [`RunState`][agents.run_state.RunState] | Runner 管理の保存済みサンドボックス状態 | 以前の Runner 管理ワークフローを再開し、そのサンドボックス状態を自動的に引き継ぐか。 |
|
||||
| [`SandboxRunConfig.session_state`][agents.run_config.SandboxRunConfig.session_state] | 明示的にシリアライズされたサンドボックスセッション状態 | `RunState` の外部で既にシリアライズしたサンドボックス状態から再開したいか。 |
|
||||
| [`SandboxRunConfig.snapshot`][agents.run_config.SandboxRunConfig.snapshot] | 新規サンドボックスセッション用の保存済みワークスペースコンテンツ | 新しいサンドボックスセッションを保存済みファイルと成果物から開始すべきか。 |
|
||||
|
||||
</div>
|
||||
|
||||
実践的な設計順序は次のとおりです。
|
||||
実用的な設計順序は次のとおりです。
|
||||
|
||||
1. `Manifest` で新規 session のワークスペース契約を定義します。
|
||||
1. `Manifest` で新規セッションのワークスペース契約を定義します。
|
||||
2. `SandboxAgent` でエージェントを定義します。
|
||||
3. 組み込みまたはカスタム capabilities を追加します。
|
||||
4. 各実行が `RunConfig(sandbox=SandboxRunConfig(...))` でどのように sandbox session を取得するか決めます。
|
||||
3. 組み込みまたはカスタムの機能を追加します。
|
||||
4. `RunConfig(sandbox=SandboxRunConfig(...))` で、各実行がサンドボックスセッションをどのように取得するかを決定します。
|
||||
|
||||
## sandbox 実行の準備
|
||||
## サンドボックス実行の準備
|
||||
|
||||
実行時に、runner はその定義を具体的な sandbox-backed 実行へ変換します。
|
||||
実行時、Runner はその定義を具体的なサンドボックス対応の実行に変換します。
|
||||
|
||||
1. `SandboxRunConfig` から sandbox session を解決します。
|
||||
`session=...` を渡した場合、そのライブな sandbox session を再利用します。
|
||||
それ以外の場合は `client=...` を使って作成または再開します。
|
||||
2. 実行用の実効ワークスペース入力を決定します。
|
||||
実行が sandbox session を注入または再開する場合、その既存の sandbox state が優先されます。
|
||||
それ以外の場合、runner は一時的な manifest override または `agent.default_manifest` から開始します。
|
||||
これが、`Manifest` だけではすべての実行の最終的なライブワークスペースを定義しない理由です。
|
||||
3. capabilities に、生成された manifest を処理させます。
|
||||
これにより、最終的なエージェント準備の前に、capabilities がファイル、mount、またはその他のワークスペーススコープの動作を追加できます。
|
||||
4. 最終的な instructions を固定順で構築します。
|
||||
SDK のデフォルト sandbox prompt、または明示的に上書きした場合は `base_instructions`、その後に `instructions`、次に capability の instructions 断片、次に remote-mount policy のテキスト、最後にレンダリングされたファイルシステムツリーを追加します。
|
||||
5. capability のツールをライブな sandbox session にバインドし、準備されたエージェントを通常の `Runner` API で実行します。
|
||||
1. `SandboxRunConfig` からサンドボックスセッションを解決します。`session=...` を渡した場合、そのライブサンドボックスセッションを再利用します。それ以外の場合は、`client=...` を使って作成または再開します。
|
||||
2. 実行の実効ワークスペース入力を決定します。実行がサンドボックスセッションを注入または再開する場合、その既存のサンドボックス状態が優先されます。それ以外の場合、Runner は 1 回限りの manifest オーバーライドまたは `agent.default_manifest` から開始します。このため、`Manifest` だけでは、すべての実行における最終的なライブワークスペースは定義されません。
|
||||
3. 機能が結果の manifest を処理できるようにします。これにより、最終的なエージェントが準備される前に、機能がファイル、マウント、その他のワークスペーススコープの動作を追加できます。
|
||||
4. 固定された順序で最終的な指示を構築します。SDK のデフォルトサンドボックスプロンプト、または明示的に上書きした場合は `base_instructions`、次に `instructions`、次に機能の指示フラグメント、次にリモートマウントポリシーテキスト、最後にレンダリングされたファイルシステムツリーです。
|
||||
5. 機能ツールをライブサンドボックスセッションにバインドし、通常の `Runner` API を通じて準備済みエージェントを実行します。
|
||||
|
||||
sandbox 化しても、turn の意味は変わりません。turn は依然としてモデルの 1 ステップであり、単一のシェルコマンドや sandbox 操作ではありません。sandbox 側の操作と turn の間に固定の 1:1 対応はありません。作業の一部は sandbox 実行レイヤー内に留まり、別のアクションはツール結果、approval、または別のモデルステップを必要とするその他の state を返すことがあります。実用上のルールとしては、sandbox 作業の後にエージェントランタイムが別のモデル応答を必要とする場合にのみ、追加の turn が消費されます。
|
||||
サンドボックス化は、ターンの意味を変えません。ターンは引き続きモデルステップであり、単一のシェルコマンドやサンドボックスアクションではありません。サンドボックス側の操作とターンの間に固定された 1:1 の対応はありません。一部の作業はサンドボックス実行レイヤー内にとどまり、他のアクションはツール結果、承認、または別のモデルステップを必要とするその他の状態を返す場合があります。実用上のルールとして、サンドボックス作業の後にエージェントランタイムが別のモデル応答を必要とする場合にのみ、別のターンが消費されます。
|
||||
|
||||
これらの準備手順があるため、`default_manifest`、`instructions`、`base_instructions`、`capabilities`、`run_as` が `SandboxAgent` を設計する際に考えるべき主な sandbox 固有オプションです。
|
||||
これらの準備ステップがあるため、`SandboxAgent` を設計する際には、`default_manifest`、`instructions`、`base_instructions`、`capabilities`、`run_as` が主なサンドボックス固有オプションになります。
|
||||
|
||||
## `SandboxAgent` のオプション
|
||||
## `SandboxAgent` オプション
|
||||
|
||||
通常の `Agent` フィールドに加えて、sandbox 固有のオプションは次のとおりです。
|
||||
通常の `Agent` フィールドに加わるサンドボックス固有のオプションは次のとおりです。
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| Option | Best use |
|
||||
| オプション | 最適な用途 |
|
||||
| --- | --- |
|
||||
| `default_manifest` | runner が作成する新しい sandbox session のデフォルトワークスペース。 |
|
||||
| `instructions` | SDK の sandbox prompt の後に追加される、追加の役割、ワークフロー、成功基準。 |
|
||||
| `base_instructions` | SDK の sandbox prompt を置き換えるための高度な escape hatch。 |
|
||||
| `capabilities` | このエージェントとともに持ち運ばれるべき sandbox ネイティブなツールと動作。 |
|
||||
| `run_as` | シェルコマンド、ファイル読み取り、patch など、モデル向け sandbox ツールで使うユーザー ID。 |
|
||||
| `default_manifest` | Runner によって作成される新規サンドボックスセッションのデフォルトワークスペース。 |
|
||||
| `instructions` | SDK サンドボックスプロンプトの後に追加される追加の役割、ワークフロー、成功基準。 |
|
||||
| `base_instructions` | SDK サンドボックスプロンプトを置き換える高度なエスケープハッチ。 |
|
||||
| `capabilities` | このエージェントと一緒に移動すべきサンドボックスネイティブのツールと動作。 |
|
||||
| `run_as` | シェルコマンド、ファイル読み取り、パッチなど、モデル向けサンドボックスツールのユーザー ID。 |
|
||||
|
||||
</div>
|
||||
|
||||
sandbox client の選択、sandbox-session の再利用、manifest override、snapshot の選択は、エージェントではなく [`SandboxRunConfig`][agents.run_config.SandboxRunConfig] に属します。
|
||||
サンドボックスクライアントの選択、サンドボックスセッションの再利用、manifest オーバーライド、スナップショット選択は、エージェントではなく [`SandboxRunConfig`][agents.run_config.SandboxRunConfig] に属します。
|
||||
|
||||
### `default_manifest`
|
||||
|
||||
`default_manifest` は、このエージェント用に runner が新しい sandbox session を作成する際に使われるデフォルトの [`Manifest`][agents.sandbox.manifest.Manifest] です。エージェントが通常開始時に持つべきファイル、リポジトリ、補助資料、出力ディレクトリ、mount のために使ってください。
|
||||
`default_manifest` は、Runner がこのエージェント用に新規サンドボックスセッションを作成するときに使用されるデフォルトの [`Manifest`][agents.sandbox.manifest.Manifest] です。エージェントが通常開始すべきファイル、リポジトリ、補助資料、出力ディレクトリ、マウントに使用します。
|
||||
|
||||
これはあくまでデフォルトです。実行ごとに `SandboxRunConfig(manifest=...)` で上書きできますし、再利用または再開された sandbox session は既存のワークスペース state を保持します。
|
||||
これはデフォルトにすぎません。実行は `SandboxRunConfig(manifest=...)` で上書きできます。また、再利用または再開されたサンドボックスセッションは既存のワークスペース状態を保持します。
|
||||
|
||||
### `instructions` と `base_instructions`
|
||||
|
||||
`instructions` は、異なる prompt でも維持したい短いルールに使ってください。`SandboxAgent` では、これらの instructions は SDK の sandbox base prompt の後に追加されるため、組み込みの sandbox ガイダンスを保持しつつ、独自の役割、ワークフロー、成功基準を追加できます。
|
||||
異なるプロンプトでも維持すべき短いルールには `instructions` を使用します。`SandboxAgent` では、これらの instructions は SDK のサンドボックスベースプロンプトの後に追加されるため、組み込みのサンドボックスガイダンスを保ちながら、独自の役割、ワークフロー、成功基準を追加できます。
|
||||
|
||||
`base_instructions` は、SDK の sandbox base prompt を置き換えたい場合にのみ使ってください。ほとんどのエージェントでは設定しないほうがよいです。
|
||||
SDK サンドボックスベースプロンプトを置き換えたい場合にのみ、`base_instructions` を使用します。ほとんどのエージェントでは設定すべきではありません。
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| Put it in... | Use it for | Examples |
|
||||
| 配置先 | 用途 | 例 |
|
||||
| --- | --- | --- |
|
||||
| `instructions` | エージェントの安定した役割、ワークフロールール、成功基準。 | 「オンボーディング文書を確認してからハンドオフしてください。」「最終ファイルは `output/` に書き込んでください。」 |
|
||||
| `base_instructions` | SDK の sandbox base prompt の完全な置き換え。 | カスタムの低レベル sandbox wrapper prompt。 |
|
||||
| ユーザープロンプト | この実行の一度限りのリクエスト。 | 「このワークスペースを要約してください。」 |
|
||||
| manifest 内のワークスペースファイル | 長めのタスク仕様、リポジトリローカルの instructions、または範囲が限定された参考資料。 | `repo/task.md`、document bundles、sample packets。 |
|
||||
| `instructions` | エージェントの安定した役割、ワークフロールール、成功基準。 | 「オンボーディング文書を検査してからハンドオフする。」、「最終ファイルを `output/` に書き込む。」 |
|
||||
| `base_instructions` | SDK サンドボックスベースプロンプトの完全な置き換え。 | カスタムの低レベルサンドボックスラッパープロンプト。 |
|
||||
| ユーザープロンプト | この実行に対する 1 回限りのリクエスト。 | 「このワークスペースを要約してください。」 |
|
||||
| manifest 内のワークスペースファイル | より長いタスク仕様、リポジトリローカルの指示、または範囲が限定された参考資料。 | `repo/task.md`、ドキュメントバンドル、サンプルパケット。 |
|
||||
|
||||
</div>
|
||||
|
||||
`instructions` の適切な使い方の例は次のとおりです。
|
||||
`instructions` の適切な使用例は次のとおりです。
|
||||
|
||||
- [examples/sandbox/unix_local_pty.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/unix_local_pty.py) は、PTY state が重要な場合にエージェントを 1 つの対話的プロセス内に保ちます。
|
||||
- [examples/sandbox/handoffs.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/handoffs.py) は、sandbox reviewer が確認後にユーザーへ直接回答することを禁止します。
|
||||
- [examples/sandbox/tax_prep.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/tax_prep.py) は、最終的な記入済みファイルが実際に `output/` に配置されることを要求します。
|
||||
- [examples/sandbox/docs/coding_task.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/docs/coding_task.py) は、正確な検証コマンドを固定し、ワークスペースルート相対の patch path を明確にします。
|
||||
- [examples/sandbox/unix_local_pty.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/unix_local_pty.py) は、PTY 状態が重要な場合にエージェントを 1 つの対話型プロセス内に保ちます。
|
||||
- [examples/sandbox/handoffs.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/handoffs.py) は、サンドボックスレビュアーが検査後にユーザーへ直接回答することを禁止します。
|
||||
- [examples/sandbox/tax_prep.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/tax_prep.py) は、最終的に入力済みのファイルが実際に `output/` に配置されることを要求します。
|
||||
- [examples/sandbox/docs/coding_task.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/docs/coding_task.py) は、正確な検証コマンドを固定し、ワークスペースルート相対のパッチパスを明確にします。
|
||||
|
||||
ユーザーの一度限りのタスクを `instructions` にコピーしたり、manifest に置くべき長い参考資料を埋め込んだり、組み込み capabilities がすでに注入するツールドキュメントを繰り返したり、実行時にモデルが必要としないローカルインストール手順を混在させたりするのは避けてください。
|
||||
ユーザーの 1 回限りのタスクを `instructions` にコピーすること、manifest に属する長い参考資料を埋め込むこと、組み込み機能が既に注入するツールドキュメントを言い換えること、実行時にモデルが必要としないローカルインストールメモを混ぜることは避けてください。
|
||||
|
||||
`instructions` を省略しても、SDK はデフォルトの sandbox prompt を含めます。これは低レベル wrapper には十分ですが、ほとんどのユーザー向けエージェントでは明示的な `instructions` を提供するべきです。
|
||||
`instructions` を省略しても、SDK はデフォルトのサンドボックスプロンプトを含めます。これは低レベルラッパーには十分ですが、ほとんどのユーザー向けエージェントでは明示的な `instructions` も提供すべきです。
|
||||
|
||||
### `capabilities`
|
||||
|
||||
capabilities は `SandboxAgent` に sandbox ネイティブな動作を付与します。実行開始前のワークスペース形成、sandbox 固有 instructions の追加、ライブな sandbox session にバインドされるツールの公開、そのエージェント向けのモデル動作や入力処理の調整を行えます。
|
||||
機能は、サンドボックスネイティブの動作を `SandboxAgent` に付与します。実行開始前にワークスペースを形成し、サンドボックス固有の指示を追加し、ライブサンドボックスセッションにバインドされるツールを公開し、そのエージェントのモデル動作や入力処理を調整できます。
|
||||
|
||||
組み込み capabilities には次が含まれます。
|
||||
組み込み機能には次のものがあります。
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| Capability | Add it when | Notes |
|
||||
| 機能 | 追加する場合 | 注記 |
|
||||
| --- | --- | --- |
|
||||
| `Shell` | エージェントにシェルアクセスが必要。 | `exec_command` を追加し、sandbox client が PTY 対話をサポートする場合は `write_stdin` も追加します。 |
|
||||
| `Filesystem` | エージェントがファイルを編集したりローカル画像を調べたりする必要がある。 | `apply_patch` と `view_image` を追加します。patch path はワークスペースルート相対です。 |
|
||||
| `Skills` | sandbox 内で skill の発見と materialization を行いたい。 | sandbox ローカルの `SKILL.md` skills には、`.agents` や `.agents/skills` を手動で mount するよりこれを推奨します。 |
|
||||
| `Memory` | 後続の実行で memory artifacts を読み取ったり生成したりすべき。 | `Shell` が必要です。ライブ更新には `Filesystem` も必要です。 |
|
||||
| `Compaction` | 長時間実行フローで compaction items の後にコンテキスト圧縮が必要。 | モデル sampling と入力処理を調整します。 |
|
||||
| `Shell` | エージェントにシェルアクセスが必要な場合。 | `exec_command` を追加します。サンドボックスクライアントが PTY 対話をサポートする場合は `write_stdin` も追加します。 |
|
||||
| `Filesystem` | エージェントにファイル編集やローカル画像の検査が必要な場合。 | `apply_patch` と `view_image` を追加します。パッチパスはワークスペースルート相対です。 |
|
||||
| `Skills` | サンドボックス内でスキルの検出とマテリアライズを行いたい場合。 | `.agents` または `.agents/skills` を手動でマウントするよりも、これを優先してください。`Skills` はスキルをインデックス化し、サンドボックスにマテリアライズします。 |
|
||||
| `Memory` | 後続の実行がメモリ成果物を読み取る、または生成する必要がある場合。 | `Shell` が必要です。ライブ更新には `Filesystem` も必要です。 |
|
||||
| `Compaction` | 長時間実行されるフローで、コンパクション項目の後にコンテキストのトリミングが必要な場合。 | モデルサンプリングと入力処理を調整します。 |
|
||||
|
||||
</div>
|
||||
|
||||
デフォルトでは、`SandboxAgent.capabilities` は `Capabilities.default()` を使い、これには `Filesystem()`、`Shell()`、`Compaction()` が含まれます。`capabilities=[...]` を渡すと、そのリストがデフォルトを置き換えるため、引き続き必要なデフォルト capability は含めてください。
|
||||
デフォルトでは、`SandboxAgent.capabilities` は `Capabilities.default()` を使用します。これには `Filesystem()`、`Shell()`、`Compaction()` が含まれます。`capabilities=[...]` を渡すと、そのリストがデフォルトを置き換えるため、引き続き必要なデフォルト機能を含めてください。
|
||||
|
||||
skills については、どのように materialize したいかに応じて source を選んでください。
|
||||
スキルについては、どのようにマテリアライズしたいかに基づいてソースを選択します。
|
||||
|
||||
- `Skills(lazy_from=LocalDirLazySkillSource(...))` は、より大きなローカル skill ディレクトリに対するよいデフォルトです。モデルがまず index を発見し、必要なものだけを読み込めるためです。
|
||||
- `Skills(from_=LocalDir(src=...))` は、事前に一括配置したい小さなローカル bundle に適しています。
|
||||
- `Skills(from_=GitRepo(repo=..., ref=...))` は、skills 自体をリポジトリから取得したい場合に適しています。
|
||||
- `Skills(lazy_from=LocalDirLazySkillSource(...))` は、モデルがまずインデックスを検出し、必要なものだけをロードできるため、大きなローカルスキルディレクトリの適切なデフォルトです。
|
||||
- `LocalDirLazySkillSource(source=LocalDir(src=...))` は、SDK プロセスが実行されているファイルシステムから読み取ります。サンドボックスイメージまたはワークスペース内にしか存在しないパスではなく、元のホスト側スキルディレクトリを渡してください。
|
||||
- `Skills(from_=LocalDir(src=...))` は、事前にステージングしたい小さなローカルバンドルに適しています。
|
||||
- `Skills(from_=GitRepo(repo=..., ref=...))` は、スキル自体をリポジトリから取得すべき場合に適しています。
|
||||
|
||||
skills がすでに `.agents/skills/<name>/SKILL.md` のような場所に存在する場合は、`LocalDir(...)` をその source root に向け、公開には引き続き `Skills(...)` を使ってください。sandbox 内レイアウトを別にする既存のワークスペース契約に依存していない限り、デフォルトの `skills_path=".agents"` を維持してください。
|
||||
`LocalDir.src` は SDK ホスト上のソースパスです。`skills_path` は、`load_skill` が呼び出されたときにスキルがステージングされる、サンドボックスワークスペース内の相対的な宛先パスです。
|
||||
|
||||
適合する場合は、組み込み capabilities を優先してください。組み込みで対応できない sandbox 固有のツールや instructions の表面が必要な場合にのみ、カスタム capability を書いてください。
|
||||
スキルが既に `.agents/skills/<name>/SKILL.md` のような場所にディスク上で存在する場合は、そのソースルートを `LocalDir(...)` に指定し、それでも `Skills(...)` を使って公開してください。サンドボックス内の別のレイアウトに依存する既存のワークスペース契約がない限り、デフォルトの `skills_path=".agents"` を維持してください。
|
||||
|
||||
適合する場合は組み込み機能を優先してください。組み込みでは対応できないサンドボックス固有のツールまたは指示サーフェスが必要な場合にのみ、カスタム機能を作成します。
|
||||
|
||||
## 概念
|
||||
|
||||
### Manifest
|
||||
|
||||
[`Manifest`][agents.sandbox.manifest.Manifest] は、新しい sandbox session のワークスペースを記述します。ワークスペースの `root` を設定し、ファイルやディレクトリを宣言し、ローカルファイルをコピーし、Git リポジトリを clone し、リモートストレージ mount を接続し、環境変数を設定し、ユーザーやグループを定義し、ワークスペース外の特定の絶対 path へのアクセスを許可できます。
|
||||
[`Manifest`][agents.sandbox.manifest.Manifest] は、新規サンドボックスセッションのワークスペースを記述します。ワークスペース `root` の設定、ファイルとディレクトリの宣言、ローカルファイルのコピー、Git リポジトリのクローン、リモートストレージマウントの接続、環境変数の設定、ユーザーまたはグループの定義、ワークスペース外の特定の絶対パスへのアクセス許可を行えます。
|
||||
|
||||
Manifest エントリの path はワークスペース相対です。絶対 path にしたり、`..` でワークスペース外へ出たりはできないため、ローカル、Docker、hosted client 間でワークスペース契約の移植性が保たれます。
|
||||
Manifest エントリのパスはワークスペース相対です。絶対パスにしたり、`..` でワークスペースから抜け出したりすることはできません。これにより、ワークスペース契約はローカル、Docker、ホスト型クライアント間でポータブルに保たれます。
|
||||
|
||||
作業開始前にエージェントが必要とする資料には manifest エントリを使ってください。
|
||||
作業開始前にエージェントが必要とする資料には、manifest エントリを使用します。
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| Manifest entry | Use it for |
|
||||
| Manifest エントリ | 用途 |
|
||||
| --- | --- |
|
||||
| `File`、`Dir` | 小さな合成入力、補助ファイル、出力ディレクトリ。 |
|
||||
| `LocalFile`、`LocalDir` | sandbox 内に materialize すべきホストファイルまたはディレクトリ。 |
|
||||
| `File`、`Dir` | 小さな合成入力、補助ファイル、または出力ディレクトリ。 |
|
||||
| `LocalFile`、`LocalDir` | サンドボックスにマテリアライズすべきホストファイルまたはディレクトリ。 |
|
||||
| `GitRepo` | ワークスペースに取得すべきリポジトリ。 |
|
||||
| `S3Mount`、`GCSMount`、`R2Mount`、`AzureBlobMount`、`S3FilesMount` などの mounts | sandbox 内に表示すべき外部ストレージ。 |
|
||||
| `S3Mount`、`GCSMount`、`R2Mount`、`AzureBlobMount`、`BoxMount`、`S3FilesMount` などのマウント | サンドボックス内に表示すべき外部ストレージ。 |
|
||||
|
||||
</div>
|
||||
|
||||
mount エントリは、どのストレージを公開するかを記述します。mount strategy は、sandbox backend がそのストレージをどのように接続するかを記述します。mount オプションとプロバイダー対応については [Sandbox clients](clients.md#mounts-and-remote-storage) を参照してください。
|
||||
`Dir` は、合成子要素から、または出力場所として、サンドボックスワークスペース内にディレクトリを作成します。ホストファイルシステムから読み取ることはありません。既存のホストディレクトリをサンドボックスワークスペースにコピーする必要がある場合は、`LocalDir` を使用します。
|
||||
|
||||
よい manifest 設計とは通常、ワークスペース契約を狭く保ち、長いタスク手順を `repo/task.md` のようなワークスペースファイルに置き、instructions では `repo/task.md` や `output/report.md` のようにワークスペース相対 path を使うことです。エージェントが `Filesystem` capability の `apply_patch` ツールでファイルを編集する場合、patch path はシェルの `workdir` ではなく sandbox ワークスペースルート相対であることを忘れないでください。
|
||||
`LocalFile.src` と `LocalDir.src` は、デフォルトで SDK プロセスの作業ディレクトリに対して解決されます。ソースは、`extra_path_grants` でカバーされていない限り、そのベースディレクトリ配下に留まる必要があります。これにより、ローカルソースのマテリアライズは、サンドボックス manifest の他の部分と同じホストパス信頼境界内に保たれます。
|
||||
|
||||
`extra_path_grants` は、エージェントがワークスペース外の具体的な絶対 path を必要とする場合にのみ使ってください。たとえば、一時的なツール出力のための `/tmp` や、読み取り専用ランタイムのための `/opt/toolchain` などです。grant は、backend がファイルシステムポリシーを適用できる場所では、SDK ファイル API とシェル実行の両方に適用されます。
|
||||
マウントエントリは公開するストレージを記述し、マウント戦略はサンドボックスバックエンドがそのストレージを接続する方法を記述します。マウントオプションとプロバイダーサポートについては、[Sandbox clients](clients.md#mounts-and-remote-storage) を参照してください。
|
||||
|
||||
適切な manifest 設計では通常、ワークスペース契約を狭く保ち、長いタスク手順を `repo/task.md` などのワークスペースファイルに置き、instructions 内では `repo/task.md` や `output/report.md` などの相対ワークスペースパスを使用します。エージェントが `Filesystem` 機能の `apply_patch` ツールでファイルを編集する場合、パッチパスはシェルの `workdir` ではなく、サンドボックスワークスペースルートからの相対パスであることに注意してください。
|
||||
|
||||
`extra_path_grants` は、エージェントがワークスペース外の具体的な絶対パスを必要とする場合、または manifest が SDK プロセス作業ディレクトリ外の信頼済みローカルソースをコピーする必要がある場合にのみ使用します。例として、一時的なツール出力用の `/tmp`、読み取り専用ランタイム用の `/opt/toolchain`、サンドボックスにマテリアライズすべき生成済みスキルディレクトリなどがあります。グラントは、ローカルソースのマテリアライズ、SDK ファイル API、およびバックエンドがファイルシステムポリシーを強制できる場合のシェル実行に適用されます。
|
||||
|
||||
```python
|
||||
from agents.sandbox import Manifest, SandboxPathGrant
|
||||
@@ -254,20 +254,22 @@ manifest = Manifest(
|
||||
)
|
||||
```
|
||||
|
||||
snapshots と `persist_workspace()` には、引き続きワークスペースルートのみが含まれます。追加で許可された path は実行時アクセスであり、永続的なワークスペース state ではありません。
|
||||
`extra_path_grants` を含む manifest は、信頼済み設定として扱ってください。アプリケーションがそれらのホストパスを既に承認していない限り、モデル出力やその他の信頼できないペイロードからグラントを読み込まないでください。
|
||||
|
||||
### 権限
|
||||
スナップショットと `persist_workspace()` には、引き続きワークスペースルートのみが含まれます。追加で許可されたパスはランタイムアクセスであり、永続的なワークスペース状態ではありません。
|
||||
|
||||
`Permissions` は manifest エントリのファイルシステム権限を制御します。これは sandbox が materialize するファイルに関するものであり、モデル権限、approval policy、API 資格情報に関するものではありません。
|
||||
### Permissions
|
||||
|
||||
デフォルトでは、manifest エントリは owner に対して読み取り、書き込み、実行が可能であり、group と others に対しては読み取りと実行が可能です。ステージングされたファイルを非公開、読み取り専用、または実行可能にしたい場合は、これを上書きしてください。
|
||||
`Permissions` は、manifest エントリのファイルシステム権限を制御します。これはサンドボックスがマテリアライズするファイルに関するものであり、モデル権限、承認ポリシー、API 認証情報に関するものではありません。
|
||||
|
||||
デフォルトでは、manifest エントリは所有者が読み取り/書き込み/実行可能で、グループとその他ユーザーが読み取り/実行可能です。ステージングされたファイルをプライベート、読み取り専用、または実行可能にすべき場合は、これを上書きします。
|
||||
|
||||
```python
|
||||
from agents.sandbox import FileMode, Permissions
|
||||
from agents.sandbox.entries import File
|
||||
|
||||
private_notes = File(
|
||||
text="internal notes",
|
||||
content=b"internal notes",
|
||||
permissions=Permissions(
|
||||
owner=FileMode.READ | FileMode.WRITE,
|
||||
group=FileMode.NONE,
|
||||
@@ -276,9 +278,9 @@ private_notes = File(
|
||||
)
|
||||
```
|
||||
|
||||
`Permissions` は、owner、group、other ごとの個別の bit と、そのエントリがディレクトリかどうかを保持します。直接構築することも、`Permissions.from_str(...)` で mode 文字列から解析することも、`Permissions.from_mode(...)` で OS mode から導出することもできます。
|
||||
`Permissions` は、所有者、グループ、その他のビットと、そのエントリがディレクトリかどうかを別々に保存します。直接構築することも、`Permissions.from_str(...)` でモード文字列から解析することも、`Permissions.from_mode(...)` で OS モードから導出することもできます。
|
||||
|
||||
Users は、作業を実行できる sandbox ID です。その ID を sandbox 内に存在させたい場合は manifest に `User` を追加し、シェルコマンド、ファイル読み取り、patch などのモデル向け sandbox ツールをそのユーザーで実行したい場合は `SandboxAgent.run_as` を設定してください。`run_as` が manifest にまだないユーザーを指している場合、runner はそのユーザーを実効 manifest に自動で追加します。
|
||||
ユーザーは、作業を実行できるサンドボックス ID です。その ID をサンドボックス内に存在させたい場合は、manifest に `User` を追加し、シェルコマンド、ファイル読み取り、パッチなどのモデル向けサンドボックスツールをそのユーザーとして実行すべき場合は `SandboxAgent.run_as` を設定します。`run_as` が manifest にまだ存在しないユーザーを指している場合、Runner は有効な manifest にそのユーザーを追加します。
|
||||
|
||||
```python
|
||||
from agents import Runner
|
||||
@@ -330,13 +332,13 @@ result = await Runner.run(
|
||||
)
|
||||
```
|
||||
|
||||
ファイルレベルの共有ルールも必要な場合は、users と manifest groups、およびエントリの `group` metadata を組み合わせてください。`run_as` ユーザーは誰が sandbox ネイティブアクションを実行するかを制御し、`Permissions` は sandbox がワークスペースを materialize した後に、そのユーザーがどのファイルを読み取り、書き込み、実行できるかを制御します。
|
||||
ファイルレベルの共有ルールも必要な場合は、ユーザーを manifest グループおよびエントリの `group` メタデータと組み合わせてください。`run_as` ユーザーは、サンドボックスネイティブアクションを誰が実行するかを制御します。`Permissions` は、サンドボックスがワークスペースをマテリアライズした後、そのユーザーがどのファイルを読み取り、書き込み、実行できるかを制御します。
|
||||
|
||||
### SnapshotSpec
|
||||
|
||||
`SnapshotSpec` は、新しい sandbox session に対して、保存済みワークスペース内容をどこから復元し、どこへ永続化するかを指定します。これは sandbox ワークスペースの snapshot policy であり、`session_state` は特定の sandbox backend を再開するためのシリアライズ済み接続 state です。
|
||||
`SnapshotSpec` は、新規サンドボックスセッションが保存済みワークスペースコンテンツをどこから復元し、どこへ永続化すべきかを示します。これはサンドボックスワークスペースのスナップショットポリシーであり、`session_state` は特定のサンドボックスバックエンドを再開するためのシリアライズ済み接続状態です。
|
||||
|
||||
ローカルの永続 snapshots には `LocalSnapshotSpec` を使い、アプリが remote snapshot client を提供する場合は `RemoteSnapshotSpec` を使ってください。ローカル snapshot のセットアップが利用できない場合は no-op snapshot が fallback として使われ、ワークスペース snapshot の永続化が不要な高度な呼び出し側はそれを明示的に使うこともできます。
|
||||
ローカルの永続スナップショットには `LocalSnapshotSpec` を使用し、アプリがリモートスナップショットクライアントを提供する場合は `RemoteSnapshotSpec` を使用します。ローカルスナップショット設定が利用できない場合はフォールバックとして no-op スナップショットが使用され、高度な呼び出し元はワークスペーススナップショットの永続化を望まない場合に明示的に使用できます。
|
||||
|
||||
```python
|
||||
from pathlib import Path
|
||||
@@ -353,13 +355,13 @@ run_config = RunConfig(
|
||||
)
|
||||
```
|
||||
|
||||
runner が新しい sandbox session を作成すると、その session 用の snapshot instance が sandbox client によって構築されます。開始時に snapshot が復元可能であれば、sandbox は保存済みワークスペース内容を復元してから実行を続行します。cleanup 時には、runner が所有する sandbox session がワークスペースをアーカイブし、snapshot を通じて再び永続化します。
|
||||
Runner が新規サンドボックスセッションを作成すると、サンドボックスクライアントはそのセッション用のスナップショットインスタンスを構築します。開始時に、スナップショットが復元可能であれば、サンドボックスは実行が続行される前に保存済みワークスペースコンテンツを復元します。クリーンアップ時には、Runner 所有のサンドボックスセッションがワークスペースをアーカイブし、スナップショットを通じて永続化します。
|
||||
|
||||
`snapshot` を省略すると、ランタイムは可能であればデフォルトのローカル snapshot 保存先を使おうとします。設定できない場合は no-op snapshot に fallback します。mount された path や一時的な path は、永続的なワークスペース内容として snapshots にコピーされません。
|
||||
`snapshot` を省略した場合、ランタイムは可能であればデフォルトのローカルスナップショット場所を使用しようとします。それを設定できない場合は、no-op スナップショットにフォールバックします。マウントされたパスと一時パスは、永続的なワークスペースコンテンツとしてスナップショットにコピーされません。
|
||||
|
||||
### sandbox ライフサイクル
|
||||
### サンドボックスライフサイクル
|
||||
|
||||
ライフサイクルモードは **SDK 所有** と **開発者所有** の 2 つです。
|
||||
ライフサイクルモードには、 **SDK 所有** と **開発者所有** の 2 つがあります。
|
||||
|
||||
<div class="sandbox-lifecycle-diagram" markdown="1">
|
||||
|
||||
@@ -387,7 +389,7 @@ sequenceDiagram
|
||||
|
||||
</div>
|
||||
|
||||
sandbox を 1 回の実行だけ存続させればよい場合は、SDK 所有ライフサイクルを使ってください。`client`、任意の `manifest`、任意の `snapshot`、client の `options` を渡すと、runner が sandbox を作成または再開し、開始し、エージェントを実行し、snapshot-backed なワークスペース state を永続化し、sandbox を停止し、runner 所有リソースを client に cleanup させます。
|
||||
サンドボックスを 1 回の実行だけ存続させればよい場合は、SDK 所有ライフサイクルを使用します。`client`、任意の `manifest`、任意の `snapshot`、およびクライアントの `options` を渡します。Runner はサンドボックスを作成または再開し、開始し、エージェントを実行し、スナップショット対応のワークスペース状態を永続化し、サンドボックスをシャットダウンし、Runner 所有リソースをクライアントにクリーンアップさせます。
|
||||
|
||||
```python
|
||||
result = await Runner.run(
|
||||
@@ -399,7 +401,7 @@ result = await Runner.run(
|
||||
)
|
||||
```
|
||||
|
||||
sandbox を事前に作成したい、1 つのライブ sandbox を複数実行で再利用したい、実行後にファイルを確認したい、自分で作成した sandbox 上で stream したい、または cleanup のタイミングを厳密に制御したい場合は、開発者所有ライフサイクルを使ってください。`session=...` を渡すと、runner はそのライブ sandbox を使いますが、代わりに閉じることはしません。
|
||||
サンドボックスを事前に作成したい場合、1 つのライブサンドボックスを複数の実行で再利用したい場合、実行後にファイルを検査したい場合、自分で作成したサンドボックス越しにストリーミングしたい場合、またはクリーンアップのタイミングを正確に決めたい場合は、開発者所有ライフサイクルを使用します。`session=...` を渡すと、Runner はそのライブサンドボックスを使用しますが、代わりに閉じることはありません。
|
||||
|
||||
```python
|
||||
sandbox = await client.create(manifest=agent.default_manifest)
|
||||
@@ -410,7 +412,7 @@ async with sandbox:
|
||||
await Runner.run(agent, "Write the final report.", run_config=run_config)
|
||||
```
|
||||
|
||||
通常の形は context manager です。entry 時に sandbox を開始し、exit 時に session cleanup ライフサイクルを実行します。アプリで context manager を使えない場合は、ライフサイクルメソッドを直接呼び出してください。
|
||||
通常はコンテキストマネージャー形式です。入力時にサンドボックスを開始し、終了時にセッションクリーンアップライフサイクルを実行します。アプリがコンテキストマネージャーを使用できない場合は、ライフサイクルメソッドを直接呼び出してください。
|
||||
|
||||
```python
|
||||
sandbox = await client.create(
|
||||
@@ -431,62 +433,64 @@ finally:
|
||||
await sandbox.aclose()
|
||||
```
|
||||
|
||||
`stop()` は snapshot-backed なワークスペース内容を永続化するだけで、sandbox 自体は破棄しません。`aclose()` は完全な session cleanup 経路です。pre-stop hooks を実行し、`stop()` を呼び出し、sandbox リソースを停止し、session スコープの依存関係を閉じます。
|
||||
`stop()` はスナップショット対応のワークスペースコンテンツのみを永続化します。サンドボックスを破棄するわけではありません。`aclose()` は完全なセッションクリーンアップパスです。pre-stop フックを実行し、`stop()` を呼び出し、サンドボックスリソースをシャットダウンし、セッションスコープの依存関係を閉じます。
|
||||
|
||||
## `SandboxRunConfig` のオプション
|
||||
## `SandboxRunConfig` オプション
|
||||
|
||||
[`SandboxRunConfig`][agents.run_config.SandboxRunConfig] は、sandbox session の取得元と、新しい session をどのように初期化するかを決める実行ごとのオプションを保持します。
|
||||
[`SandboxRunConfig`][agents.run_config.SandboxRunConfig] は、サンドボックスセッションの取得元と、新規セッションをどのように初期化するかを決定する実行ごとのオプションを保持します。
|
||||
|
||||
### sandbox の取得元
|
||||
### サンドボックスソース
|
||||
|
||||
これらのオプションは、runner が sandbox session を再利用、再開、または作成するかどうかを決定します。
|
||||
これらのオプションは、Runner がサンドボックスセッションを再利用、再開、または作成すべきかを決定します。
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| Option | Use it when | Notes |
|
||||
| オプション | 使用する場合 | 注記 |
|
||||
| --- | --- | --- |
|
||||
| `client` | runner に sandbox session の作成、再開、cleanup を任せたい。 | ライブな sandbox `session` を渡さない限り必須です。 |
|
||||
| `session` | すでにライブな sandbox session を自分で作成している。 | ライフサイクルは呼び出し側が所有し、runner はそのライブ sandbox session を再利用します。 |
|
||||
| `session_state` | シリアライズ済み sandbox session state はあるが、ライブな sandbox session object はない。 | `client` が必要で、runner はその明示的 state から所有 session として再開します。 |
|
||||
| `client` | Runner にサンドボックスセッションの作成、再開、クリーンアップを任せたい場合。 | ライブサンドボックス `session` を提供しない限り必須です。 |
|
||||
| `session` | 既に自分でライブサンドボックスセッションを作成している場合。 | 呼び出し元がライフサイクルを所有します。Runner はそのライブサンドボックスセッションを再利用します。 |
|
||||
| `session_state` | シリアライズ済みのサンドボックスセッション状態はあるが、ライブサンドボックスセッションオブジェクトはない場合。 | `client` が必要です。Runner はその明示的な状態から所有セッションとして再開します。 |
|
||||
|
||||
</div>
|
||||
|
||||
実際には、runner は次の順序で sandbox session を解決します。
|
||||
実際には、Runner は次の順序でサンドボックスセッションを解決します。
|
||||
|
||||
1. `run_config.sandbox.session` を注入した場合、そのライブな sandbox session を直接再利用します。
|
||||
2. それ以外で、実行が `RunState` から再開される場合は、保存された sandbox session state を再開します。
|
||||
3. それ以外で、`run_config.sandbox.session_state` を渡した場合は、その明示的にシリアライズされた sandbox session state から runner が再開します。
|
||||
4. それ以外の場合、runner は新しい sandbox session を作成します。その新しい session には、指定があれば `run_config.sandbox.manifest`、なければ `agent.default_manifest` を使います。
|
||||
1. `run_config.sandbox.session` を注入した場合、そのライブサンドボックスセッションが直接再利用されます。
|
||||
2. それ以外で、実行が `RunState` から再開している場合、保存されたサンドボックスセッション状態が再開されます。
|
||||
3. それ以外で、`run_config.sandbox.session_state` を渡した場合、Runner はその明示的にシリアライズされたサンドボックスセッション状態から再開します。
|
||||
4. それ以外の場合、Runner は新規サンドボックスセッションを作成します。その新規セッションでは、提供されていれば `run_config.sandbox.manifest` を使用し、そうでなければ `agent.default_manifest` を使用します。
|
||||
|
||||
### 新規 session の入力
|
||||
### 新規セッション入力
|
||||
|
||||
これらのオプションは、runner が新しい sandbox session を作成するときにのみ重要です。
|
||||
これらのオプションは、Runner が新規サンドボックスセッションを作成する場合にのみ意味があります。
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| Option | Use it when | Notes |
|
||||
| オプション | 使用する場合 | 注記 |
|
||||
| --- | --- | --- |
|
||||
| `manifest` | 一度限りの新規 session ワークスペース override が必要。 | 省略時は `agent.default_manifest` に fallback します。 |
|
||||
| `snapshot` | 新しい sandbox session を snapshot から初期化すべき。 | 再開に近いフローや remote snapshot client に有用です。 |
|
||||
| `options` | sandbox client に作成時オプションが必要。 | Docker image、Modal app 名、E2B template、timeouts など、client 固有設定でよく使われます。 |
|
||||
| `manifest` | 1 回限りの新規セッションワークスペース上書きを行いたい場合。 | 省略された場合は `agent.default_manifest` にフォールバックします。 |
|
||||
| `snapshot` | 新規サンドボックスセッションをスナップショットから初期化すべき場合。 | 再開に似たフローやリモートスナップショットクライアントに便利です。 |
|
||||
| `options` | サンドボックスクライアントが作成時オプションを必要とする場合。 | Docker イメージ、Modal アプリ名、E2B テンプレート、タイムアウト、および同様のクライアント固有設定で一般的です。 |
|
||||
|
||||
</div>
|
||||
|
||||
### materialization 制御
|
||||
### マテリアライズ制御
|
||||
|
||||
`concurrency_limits` は、どれだけの sandbox materialization 作業を並列実行できるかを制御します。大きな manifests やローカルディレクトリのコピーで、より厳密なリソース制御が必要な場合は `SandboxConcurrencyLimits(manifest_entries=..., local_dir_files=...)` を使ってください。どちらかの値を `None` にすると、その特定の制限を無効化できます。
|
||||
`concurrency_limits` は、並列に実行できるサンドボックスマテリアライズ作業の量を制御します。大きな manifest やローカルディレクトリコピーに対してより厳密なリソース制御が必要な場合は、`SandboxConcurrencyLimits(manifest_entries=..., local_dir_files=...)` を使用します。いずれかの値を `None` に設定すると、その特定の制限を無効化できます。
|
||||
|
||||
いくつかの含意を覚えておくとよいです。
|
||||
`archive_limits` は、アーカイブ展開に対する SDK 側のリソースチェックを制御します。SDK デフォルトしきい値を有効にするには `archive_limits=SandboxArchiveLimits()` を設定します。アーカイブにより厳密なリソース制御が必要な場合は、`SandboxArchiveLimits(max_input_bytes=..., max_extracted_bytes=..., max_members=...)` のような明示的な値を渡します。SDK のアーカイブリソース制限なしというデフォルト動作を維持するには `archive_limits=None` のままにします。または、個別フィールドを `None` に設定して、その制限だけを無効化します。
|
||||
|
||||
- 新しい session: `manifest=` と `snapshot=` は、runner が新しい sandbox session を作成する場合にのみ適用されます。
|
||||
- 再開と snapshot: `session_state=` は以前にシリアライズした sandbox state に再接続するものであり、`snapshot=` は保存済みワークスペース内容から新しい sandbox session を初期化するものです。
|
||||
- client 固有オプション: `options=` は sandbox client に依存します。Docker や多くの hosted client では必須です。
|
||||
- 注入されたライブ session: 実行中の sandbox `session` を渡した場合、capability による manifest 更新で、互換性のある非 mount エントリを追加できます。ただし `manifest.root`、`manifest.environment`、`manifest.users`、`manifest.groups` を変更したり、既存エントリを削除したり、エントリ型を置き換えたり、mount エントリを追加または変更したりはできません。
|
||||
- Runner API: `SandboxAgent` の実行は、引き続き通常の `Runner.run()`、`Runner.run_sync()`、`Runner.run_streamed()` API を使います。
|
||||
覚えておく価値のある含意がいくつかあります。
|
||||
|
||||
- 新規セッション: `manifest=` と `snapshot=` は、Runner が新規サンドボックスセッションを作成する場合にのみ適用されます。
|
||||
- 再開とスナップショット: `session_state=` は以前にシリアライズされたサンドボックス状態へ再接続します。一方、`snapshot=` は保存済みワークスペースコンテンツから新しいサンドボックスセッションを初期化します。
|
||||
- クライアント固有オプション: `options=` はサンドボックスクライアントに依存します。Docker と多くのホスト型クライアントでは必須です。
|
||||
- 注入されたライブセッション: 実行中のサンドボックス `session` を渡した場合、機能による manifest 更新は互換性のある非マウントエントリを追加できます。`manifest.root`、`manifest.environment`、`manifest.users`、`manifest.groups` を変更すること、既存エントリを削除すること、エントリ型を置き換えること、マウントエントリを追加または変更することはできません。
|
||||
- Runner API: `SandboxAgent` の実行は引き続き通常の `Runner.run()`、`Runner.run_sync()`、`Runner.run_streamed()` API を使用します。
|
||||
|
||||
## 完全な例: コーディングタスク
|
||||
|
||||
このコーディングスタイルの例は、よいデフォルトの出発点です。
|
||||
このコーディング形式の例は、適切なデフォルトの出発点です。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -529,9 +533,10 @@ def build_agent(model: str) -> SandboxAgent[None]:
|
||||
}
|
||||
),
|
||||
capabilities=Capabilities.default() + [
|
||||
# Let Skills(...) stage and index sandbox-local skills for you.
|
||||
Skills(
|
||||
lazy_from=LocalDirLazySkillSource(
|
||||
# This is a host path read by the SDK process.
|
||||
# Requested skills are copied into `skills_path` in the sandbox.
|
||||
source=LocalDir(src=HOST_SKILLS_DIR),
|
||||
)
|
||||
),
|
||||
@@ -555,7 +560,7 @@ async def main(model: str, prompt: str) -> None:
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(
|
||||
main(
|
||||
model="gpt-5.4",
|
||||
model="gpt-5.5",
|
||||
prompt=(
|
||||
"Open `repo/task.md`, use the `$credit-note-fixer` skill, fix the bug, "
|
||||
f"run `{TARGET_TEST_CMD}`, and summarize the change."
|
||||
@@ -564,19 +569,19 @@ if __name__ == "__main__":
|
||||
)
|
||||
```
|
||||
|
||||
[examples/sandbox/docs/coding_task.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/docs/coding_task.py) を参照してください。この例では、小さなシェルベースのリポジトリを使っているため、Unix ローカル実行全体で決定的に検証できます。もちろん実際のタスクリポジトリは Python、JavaScript、その他何でも構いません。
|
||||
[examples/sandbox/docs/coding_task.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/docs/coding_task.py) を参照してください。この例では、Unix ローカル実行間で決定論的に検証できるように、小さなシェルベースのリポジトリを使用します。実際のタスクリポジトリは、もちろん Python、JavaScript、その他何でも構いません。
|
||||
|
||||
## 一般的なパターン
|
||||
|
||||
上記の完全な例から始めてください。多くの場合、同じ `SandboxAgent` はそのままで、sandbox client、sandbox-session の取得元、またはワークスペースの取得元だけを変更できます。
|
||||
上記の完全な例から始めてください。多くの場合、同じ `SandboxAgent` をそのまま維持し、サンドボックスクライアント、サンドボックスセッションソース、またはワークスペースソースだけを変更できます。
|
||||
|
||||
### sandbox client の切り替え
|
||||
### サンドボックスクライアントの切り替え
|
||||
|
||||
エージェント定義はそのままにして、run config だけを変更します。コンテナ分離やイメージの一致が必要なら Docker を使い、プロバイダー管理の実行が必要なら hosted provider を使ってください。例とプロバイダーオプションについては [Sandbox clients](clients.md) を参照してください。
|
||||
エージェント定義は同じままにし、実行設定だけを変更します。コンテナ隔離やイメージの同等性が必要な場合は Docker を使用し、プロバイダー管理の実行が必要な場合はホスト型プロバイダーを使用します。例とプロバイダーオプションについては、[Sandbox clients](clients.md) を参照してください。
|
||||
|
||||
### ワークスペースの上書き
|
||||
|
||||
エージェント定義はそのままにして、新規 session の manifest だけを差し替えます。
|
||||
エージェント定義は同じままにし、新規セッションの manifest だけを入れ替えます。
|
||||
|
||||
```python
|
||||
from agents.run import RunConfig
|
||||
@@ -596,11 +601,11 @@ run_config = RunConfig(
|
||||
)
|
||||
```
|
||||
|
||||
同じエージェントの役割を、異なるリポジトリ、資料パケット、タスクバンドルに対して、エージェントを作り直さずに実行したい場合に使います。上の検証可能なコーディング例では、一度限りの override の代わりに `default_manifest` を使って同じパターンを示しています。
|
||||
エージェントを再構築せずに、同じエージェントの役割を異なるリポジトリ、パケット、またはタスクバンドルに対して実行すべき場合に使用します。上記の検証済みコーディング例は、1 回限りの上書きではなく `default_manifest` で同じパターンを示しています。
|
||||
|
||||
### sandbox session の注入
|
||||
### サンドボックスセッションの注入
|
||||
|
||||
明示的なライフサイクル制御、実行後の確認、または出力コピーが必要な場合は、ライブな sandbox session を注入します。
|
||||
明示的なライフサイクル制御、実行後の検査、または出力コピーが必要な場合は、ライブサンドボックスセッションを注入します。
|
||||
|
||||
```python
|
||||
from agents import Runner
|
||||
@@ -621,11 +626,11 @@ async with sandbox:
|
||||
)
|
||||
```
|
||||
|
||||
実行後にワークスペースを確認したい場合や、すでに開始済みの sandbox session 上で stream したい場合に使います。[examples/sandbox/docs/coding_task.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/docs/coding_task.py) と [examples/sandbox/docker/docker_runner.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/docker/docker_runner.py) を参照してください。
|
||||
実行後にワークスペースを検査したい場合や、既に開始済みのサンドボックスセッション越しにストリーミングしたい場合に使用します。[examples/sandbox/docs/coding_task.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/docs/coding_task.py) と [examples/sandbox/docker/docker_runner.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/docker/docker_runner.py) を参照してください。
|
||||
|
||||
### session state からの再開
|
||||
### セッション状態からの再開
|
||||
|
||||
すでに `RunState` の外で sandbox state をシリアライズしている場合は、その state から runner に再接続させてください。
|
||||
`RunState` の外部で既にサンドボックス状態をシリアライズしている場合は、Runner にその状態から再接続させます。
|
||||
|
||||
```python
|
||||
from agents.run import RunConfig
|
||||
@@ -642,11 +647,11 @@ run_config = RunConfig(
|
||||
)
|
||||
```
|
||||
|
||||
sandbox state が独自のストレージや job system にあり、`Runner` にそこから直接再開させたい場合に使います。serialize / deserialize の流れについては [examples/sandbox/extensions/blaxel_runner.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/extensions/blaxel_runner.py) を参照してください。
|
||||
サンドボックス状態が独自のストレージまたはジョブシステムに存在し、`Runner` にそれを直接再開させたい場合に使用します。シリアライズ/デシリアライズフローについては、[examples/sandbox/extensions/blaxel_runner.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/extensions/blaxel_runner.py) を参照してください。
|
||||
|
||||
### snapshot からの開始
|
||||
### スナップショットからの開始
|
||||
|
||||
保存済みファイルや成果物から新しい sandbox を初期化します。
|
||||
保存済みファイルと成果物から新しいサンドボックスを初期化します。
|
||||
|
||||
```python
|
||||
from pathlib import Path
|
||||
@@ -663,11 +668,11 @@ run_config = RunConfig(
|
||||
)
|
||||
```
|
||||
|
||||
新しい実行を、`agent.default_manifest` だけでなく保存済みワークスペース内容から開始したい場合に使います。ローカル snapshot フローについては [examples/sandbox/memory.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/memory.py)、remote snapshot client については [examples/sandbox/sandbox_agent_with_remote_snapshot.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/sandbox_agent_with_remote_snapshot.py) を参照してください。
|
||||
新規実行を `agent.default_manifest` だけでなく、保存済みワークスペースコンテンツから開始すべき場合に使用します。ローカルスナップショットフローについては [examples/sandbox/memory.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/memory.py) を、リモートスナップショットクライアントについては [examples/sandbox/sandbox_agent_with_remote_snapshot.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/sandbox_agent_with_remote_snapshot.py) を参照してください。
|
||||
|
||||
### Git から skills を読み込む
|
||||
### Git からのスキル読み込み
|
||||
|
||||
ローカル skill source を、リポジトリベースのものに差し替えます。
|
||||
ローカルスキルソースを、リポジトリをバックエンドとするものに入れ替えます。
|
||||
|
||||
```python
|
||||
from agents.sandbox.capabilities import Capabilities, Skills
|
||||
@@ -678,11 +683,11 @@ capabilities = Capabilities.default() + [
|
||||
]
|
||||
```
|
||||
|
||||
skills bundle に独自のリリースサイクルがある場合や、複数の sandbox 間で共有したい場合に使います。[examples/sandbox/tax_prep.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/tax_prep.py) を参照してください。
|
||||
スキルバンドルに独自のリリースサイクルがある場合や、サンドボックス間で共有すべき場合に使用します。[examples/sandbox/tax_prep.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/tax_prep.py) を参照してください。
|
||||
|
||||
### ツールとして公開
|
||||
### ツールとしての公開
|
||||
|
||||
tool-agent は、独自の sandbox 境界を持つことも、親実行からライブな sandbox を再利用することもできます。再利用は、高速な読み取り専用 explorer agent に有用です。別の sandbox を作成、hydrate、snapshot するコストを払わずに、親が使っている正確なワークスペースを確認できます。
|
||||
ツールエージェントは、独自のサンドボックス境界を持つことも、親実行のライブサンドボックスを再利用することもできます。再利用は、高速な読み取り専用エクスプローラーエージェントに便利です。別のサンドボックスを作成、ハイドレート、スナップショットするコストを払わずに、親が使用している正確なワークスペースを検査できます。
|
||||
|
||||
```python
|
||||
from agents import Runner
|
||||
@@ -764,17 +769,22 @@ async with sandbox:
|
||||
)
|
||||
```
|
||||
|
||||
ここでは、親エージェントは `coordinator` として実行され、explorer tool-agent は同じライブな sandbox session 内で `explorer` として実行されます。`pricing_packet/` のエントリは `other` ユーザーに対して読み取り可能なので、explorer はそれらを素早く確認できますが、書き込み bit はありません。`work/` ディレクトリは coordinator の user / group にのみ利用可能なので、親は最終成果物を書き込めますが、explorer は読み取り専用のままです。
|
||||
ここでは、親エージェントが `coordinator` として実行され、エクスプローラーツールエージェントが同じライブサンドボックスセッション内で `explorer` として実行されます。`pricing_packet/` エントリは `other` ユーザーが読み取り可能なので、explorer はそれらをすばやく検査できますが、書き込みビットはありません。`work/` ディレクトリは coordinator のユーザー/グループだけが利用できるため、親は最終成果物を書き込めますが、explorer は読み取り専用のままです。
|
||||
|
||||
tool-agent に本当の分離が必要な場合は、独自の sandbox `RunConfig` を与えてください。
|
||||
ツールエージェントに実際の隔離が必要な場合は、独自のサンドボックス `RunConfig` を与えます。
|
||||
|
||||
```python
|
||||
from docker import from_env as docker_from_env
|
||||
|
||||
from agents.run import RunConfig
|
||||
from agents.sandbox import SandboxRunConfig
|
||||
from agents.sandbox import SandboxAgent, SandboxRunConfig
|
||||
from agents.sandbox.sandboxes.docker import DockerSandboxClient, DockerSandboxClientOptions
|
||||
|
||||
rollout_agent = SandboxAgent(
|
||||
name="Rollout Reviewer",
|
||||
instructions="Inspect the rollout packet and summarize implementation risk.",
|
||||
)
|
||||
|
||||
rollout_agent.as_tool(
|
||||
tool_name="review_rollout_risk",
|
||||
tool_description="Inspect the rollout packet and summarize implementation risk.",
|
||||
@@ -787,11 +797,11 @@ rollout_agent.as_tool(
|
||||
)
|
||||
```
|
||||
|
||||
tool-agent が自由に変更を加えるべき場合、信頼できないコマンドを実行すべき場合、または別の backend / image を使うべき場合は、別の sandbox を使ってください。[examples/sandbox/sandbox_agents_as_tools.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/sandbox_agents_as_tools.py) を参照してください。
|
||||
ツールエージェントが自由に変更する、信頼できないコマンドを実行する、または別のバックエンド/イメージを使用する必要がある場合は、別のサンドボックスを使用します。[examples/sandbox/sandbox_agents_as_tools.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/sandbox_agents_as_tools.py) を参照してください。
|
||||
|
||||
### ローカルツールおよび MCP との組み合わせ
|
||||
### ローカルツールと MCP との組み合わせ
|
||||
|
||||
同じエージェント上で通常のツールを使いつつ、sandbox ワークスペースも維持します。
|
||||
同じエージェントで通常のツールも使用しながら、サンドボックスワークスペースを維持します。
|
||||
|
||||
```python
|
||||
from agents.sandbox import SandboxAgent
|
||||
@@ -806,46 +816,46 @@ agent = SandboxAgent(
|
||||
)
|
||||
```
|
||||
|
||||
ワークスペース確認がエージェントの仕事の一部にすぎない場合に使います。[examples/sandbox/sandbox_agent_with_tools.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/sandbox_agent_with_tools.py) を参照してください。
|
||||
ワークスペース検査がエージェントの仕事の一部にすぎない場合に使用します。[examples/sandbox/sandbox_agent_with_tools.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/sandbox_agent_with_tools.py) を参照してください。
|
||||
|
||||
## Memory
|
||||
## メモリ
|
||||
|
||||
将来の sandbox-agent 実行が過去の実行から学習すべき場合は、`Memory` capability を使ってください。Memory は SDK の会話用 `Session` memory とは別物です。学びを sandbox ワークスペース内のファイルに要約し、後続の実行でそれらのファイルを読み取れるようにします。
|
||||
将来のサンドボックスエージェント実行が過去の実行から学習すべき場合は、`Memory` 機能を使用します。メモリは SDK の会話用 `Session` メモリとは別のものです。教訓をサンドボックスワークスペース内のファイルに抽出し、後続の実行がそれらのファイルを読み取れるようにします。
|
||||
|
||||
セットアップ、読み取り / 生成の動作、複数 turn の会話、レイアウト分離については [Agent memory](memory.md) を参照してください。
|
||||
セットアップ、読み取り/生成動作、マルチターン会話、レイアウト隔離については、[Agent memory](memory.md) を参照してください。
|
||||
|
||||
## 構成パターン
|
||||
|
||||
単一エージェントのパターンが明確になったら、次の設計上の問いは、より大きなシステムの中で sandbox 境界をどこに置くかです。
|
||||
単一エージェントパターンが明確になったら、次の設計上の問いは、より大きなシステムのどこにサンドボックス境界を置くかです。
|
||||
|
||||
sandbox agents は、SDK の他の部分とも引き続き組み合わせられます。
|
||||
サンドボックスエージェントは、引き続き SDK の他の部分と組み合わせられます。
|
||||
|
||||
- [Handoffs](../handoffs.md): sandbox でない intake agent から、ドキュメント量の多い作業を sandbox reviewer に handoff します。
|
||||
- [Agents as tools](../tools.md#agents-as-tools): 複数の sandbox agent をツールとして公開します。通常は各 `Agent.as_tool(...)` 呼び出しに `run_config=RunConfig(sandbox=SandboxRunConfig(...))` を渡し、各ツールが独自の sandbox 境界を持つようにします。
|
||||
- [MCP](../mcp.md) と通常の関数ツール: sandbox capabilities は `mcp_servers` や通常の Python ツールと共存できます。
|
||||
- [Running agents](../running_agents.md): sandbox 実行も引き続き通常の `Runner` API を使います。
|
||||
- [Handoffs](../handoffs.md): ドキュメント量の多い作業を、非サンドボックスの受付エージェントからサンドボックスレビュアーへハンドオフします。
|
||||
- [Agents as tools](../tools.md#agents-as-tools): 複数のサンドボックスエージェントをツールとして公開します。通常は各 `Agent.as_tool(...)` 呼び出しで `run_config=RunConfig(sandbox=SandboxRunConfig(...))` を渡し、各ツールが独自のサンドボックス境界を得るようにします。
|
||||
- [MCP](../mcp.md) と通常の関数ツール: サンドボックス機能は `mcp_servers` や通常の Python ツールと共存できます。
|
||||
- [エージェントの実行](../running_agents.md): サンドボックス実行は引き続き通常の `Runner` API を使用します。
|
||||
|
||||
特によくあるパターンは次の 2 つです。
|
||||
特に一般的なパターンは 2 つあります。
|
||||
|
||||
- sandbox でないエージェントが、ワークスペース分離が必要なワークフロー部分だけ sandbox agent に handoff する
|
||||
- オーケストレーターが複数の sandbox agent をツールとして公開し、通常は各 `Agent.as_tool(...)` 呼び出しごとに別個の sandbox `RunConfig` を使って、各ツールが独自の分離ワークスペースを持つようにする
|
||||
- ワークフローのうちワークスペース隔離が必要な部分だけ、非サンドボックスエージェントがサンドボックスエージェントへハンドオフする
|
||||
- オーケストレーターが複数のサンドボックスエージェントをツールとして公開する。通常は各 `Agent.as_tool(...)` 呼び出しに別々のサンドボックス `RunConfig` を使用し、各ツールが独自の隔離ワークスペースを得るようにする
|
||||
|
||||
### turn と sandbox 実行
|
||||
### ターンとサンドボックス実行
|
||||
|
||||
handoff と agent-as-tool の呼び出しは分けて説明すると理解しやすくなります。
|
||||
ハンドオフと agent-as-tool 呼び出しは分けて説明すると理解しやすくなります。
|
||||
|
||||
handoff の場合、トップレベルの実行は 1 つで、トップレベルの turn loop も 1 つのままです。アクティブなエージェントは変わりますが、実行はネストしません。sandbox でない intake agent が sandbox reviewer に handoff すると、その同じ実行内の次のモデル呼び出しは sandbox agent 向けに準備され、その sandbox agent が次の turn を担当するエージェントになります。つまり handoff は、同じ実行の次の turn をどのエージェントが担当するかを変えるものです。[examples/sandbox/handoffs.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/handoffs.py) を参照してください。
|
||||
ハンドオフの場合、トップレベルの実行とトップレベルのターンループは引き続き 1 つです。アクティブなエージェントは変わりますが、実行がネストされるわけではありません。非サンドボックスの受付エージェントがサンドボックスレビュアーへハンドオフすると、同じ実行内の次のモデル呼び出しはサンドボックスエージェント向けに準備され、そのサンドボックスエージェントが次のターンを担当するエージェントになります。言い換えると、ハンドオフは同じ実行の次のターンをどのエージェントが所有するかを変更します。[examples/sandbox/handoffs.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/handoffs.py) を参照してください。
|
||||
|
||||
`Agent.as_tool(...)` の場合は関係が異なります。外側のオーケストレーターは 1 回の外側 turn を使ってツール呼び出しを決定し、そのツール呼び出しによって sandbox agent のネストされた実行が開始されます。ネストされた実行は独自の turn loop、`max_turns`、approvals、そして通常は独自の sandbox `RunConfig` を持ちます。1 回のネスト turn で終わることもあれば、複数回かかることもあります。外側のオーケストレーターの観点では、そのすべての作業は依然として 1 回のツール呼び出しの背後にあるため、ネストされた turn は外側の実行の turn counter を増やしません。[examples/sandbox/sandbox_agents_as_tools.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/sandbox_agents_as_tools.py) を参照してください。
|
||||
`Agent.as_tool(...)` の場合、関係は異なります。外側のオーケストレーターは 1 つの外側ターンを使ってツール呼び出しを決定し、そのツール呼び出しがサンドボックスエージェントのネストされた実行を開始します。ネストされた実行には、独自のターンループ、`max_turns`、承認、通常は独自のサンドボックス `RunConfig` があります。1 つのネストターンで完了する場合もあれば、複数かかる場合もあります。外側のオーケストレーターから見ると、そのすべての作業は 1 つのツール呼び出しの背後にあるため、ネストされたターンは外側実行のターンカウンターを増やしません。[examples/sandbox/sandbox_agents_as_tools.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/sandbox_agents_as_tools.py) を参照してください。
|
||||
|
||||
approval の動作も同じ分割に従います。
|
||||
承認動作も同じ分離に従います。
|
||||
|
||||
- handoff では、sandbox agent がその実行のアクティブエージェントになるため、approvals は同じトップレベル実行に留まります
|
||||
- `Agent.as_tool(...)` では、sandbox tool-agent 内で発生した approvals も外側の実行に現れますが、それらは保存されたネスト実行 state から来るものであり、外側の実行が再開されるとネストされた sandbox 実行も再開されます
|
||||
- ハンドオフでは、サンドボックスエージェントがその実行のアクティブなエージェントになるため、承認は同じトップレベル実行に残ります
|
||||
- `Agent.as_tool(...)` では、サンドボックスツールエージェント内で発生した承認も外側の実行に表面化しますが、それらは保存されたネスト実行状態から来ており、外側の実行が再開するとネストされたサンドボックス実行を再開します
|
||||
|
||||
## 参考資料
|
||||
## 関連情報
|
||||
|
||||
- [Quickstart](quickstart.md): 1 つの sandbox agent を動かします。
|
||||
- [Sandbox clients](clients.md): ローカル、Docker、hosted、mount のオプションを選びます。
|
||||
- [Agent memory](memory.md): 過去の sandbox 実行から得た学びを保存して再利用します。
|
||||
- [examples/sandbox/](https://github.com/openai/openai-agents-python/tree/main/examples/sandbox): 実行可能なローカル、コーディング、memory、handoff、エージェント構成パターンです。
|
||||
- [Quickstart](../sandbox_agents.md): 1 つのサンドボックスエージェントを実行します。
|
||||
- [Sandbox clients](clients.md): ローカル、Docker、ホスト型、マウントの各オプションを選択します。
|
||||
- [Agent memory](memory.md): 以前のサンドボックス実行からの教訓を保持し、再利用します。
|
||||
- [examples/sandbox/](https://github.com/openai/openai-agents-python/tree/main/examples/sandbox): 実行可能なローカル、コーディング、メモリ、ハンドオフ、エージェント構成パターン。
|
||||
+30
-30
@@ -4,23 +4,23 @@ search:
|
||||
---
|
||||
# エージェントメモリ
|
||||
|
||||
メモリを使うと、今後の sandbox-agent の実行が過去の実行から学習できるようになります。これは、メッセージ履歴を保存する SDK の会話用 [`Session`](../sessions/index.md) メモリとは別のものです。メモリは、過去の実行から得られた学びを sandbox ワークスペース内のファイルに要約します。
|
||||
メモリにより、今後の sandbox エージェントの実行は過去の実行から学習できます。これは、メッセージ履歴を保存する SDK の会話用 [`Session`](../sessions/index.md) メモリとは別のものです。メモリは、過去の実行から得た学びを sandbox ワークスペース内のファイルに要約します。
|
||||
|
||||
!!! warning "ベータ機能"
|
||||
|
||||
Sandbox エージェントはベータ版です。一般提供までに API の詳細、デフォルト設定、サポートされる機能は変更される可能性があり、今後さらに高度な機能も追加される予定です。
|
||||
Sandbox エージェントはベータ版です。一般提供までに API の詳細、デフォルト値、サポートされる機能が変更される可能性があり、時間とともにさらに高度な機能が追加されることも想定してください。
|
||||
|
||||
メモリは、将来の実行における次の 3 種類のコストを削減できます。
|
||||
メモリは、今後の実行における 3 種類のコストを削減できます。
|
||||
|
||||
1. エージェントコスト: エージェントがワークフローの完了に長い時間を要した場合、次回の実行では探索が少なくて済むはずです。これにより、トークン使用量と完了までの時間を削減できます。
|
||||
2. ユーザーコスト: ユーザーがエージェントを修正したり、好みを示したりした場合、今後の実行ではそのフィードバックを記憶できます。これにより、人手による介入を減らせます。
|
||||
3. コンテキストコスト: エージェントが以前にタスクを完了していて、ユーザーがそのタスクを引き継いで進めたい場合、ユーザーは以前のスレッドを探したり、すべてのコンテキストを再入力したりする必要がありません。これにより、タスクの説明を短くできます。
|
||||
1. エージェントのコスト: エージェントがワークフローの完了に長い時間を要した場合、次回の実行では探索が少なくて済むはずです。これにより、トークン使用量と完了までの時間を削減できます。
|
||||
2. ユーザーのコスト: ユーザーがエージェントを修正したり好みを表明したりした場合、今後の実行でそのフィードバックを記憶できます。これにより、人による介入を削減できます。
|
||||
3. コンテキストのコスト: エージェントが以前にタスクを完了していて、ユーザーがそのタスクを発展させたい場合、ユーザーは以前のスレッドを探したり、すべてのコンテキストを再入力したりする必要がないはずです。これにより、タスク説明を短くできます。
|
||||
|
||||
バグを修正し、メモリを生成し、スナップショットを再開し、そのメモリを後続の verifier 実行で使用する 2 回実行の完全な例については、[examples/sandbox/memory.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/memory.py) を参照してください。別々のメモリレイアウトを使ったマルチターン・マルチエージェントの例については、[examples/sandbox/memory_multi_agent_multiturn.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/memory_multi_agent_multiturn.py) を参照してください。
|
||||
バグを修正し、メモリを生成し、スナップショットを再開し、そのメモリを後続の検証実行で使用する、2 回の実行からなる完全なコード例については、[examples/sandbox/memory.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/memory.py) を参照してください。独立したメモリレイアウトを持つマルチターン、マルチエージェントのコード例については、[examples/sandbox/memory_multi_agent_multiturn.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/memory_multi_agent_multiturn.py) を参照してください。
|
||||
|
||||
## メモリの有効化
|
||||
|
||||
sandbox エージェントの capability として `Memory()` を追加します。
|
||||
sandbox エージェントに機能として `Memory()` を追加します。
|
||||
|
||||
```python
|
||||
from pathlib import Path
|
||||
@@ -42,28 +42,28 @@ with tempfile.TemporaryDirectory(prefix="sandbox-memory-example-") as snapshot_d
|
||||
)
|
||||
```
|
||||
|
||||
読み取りが有効な場合、`Memory()` には `Shell()` が必要です。これにより、注入された要約だけでは不十分なときに、エージェントがメモリファイルを読み取り、検索できます。ライブメモリ更新が有効な場合(デフォルト)、`Filesystem()` も必要です。これにより、エージェントが古いメモリを見つけた場合や、ユーザーがメモリの更新を求めた場合に、`memories/MEMORY.md` を更新できます。
|
||||
読み取りが有効な場合、`Memory()` には `Shell()` が必要です。これにより、挿入されたサマリーだけでは不十分なときに、エージェントがメモリファイルを読み取り、検索できます。ライブメモリ更新が有効な場合(デフォルト)、`Filesystem()` も必要です。これにより、エージェントが古くなったメモリを発見した場合や、ユーザーがメモリの更新を依頼した場合に、`memories/MEMORY.md` を更新できます。
|
||||
|
||||
デフォルトでは、メモリアーティファクトは sandbox ワークスペースの `memories/` 以下に保存されます。後続の実行でそれらを再利用するには、同じライブ sandbox セッションを維持するか、永続化されたセッション状態またはスナップショットから再開することで、設定された memories ディレクトリー全体を保持して再利用してください。新しい空の sandbox は空のメモリで開始します。
|
||||
デフォルトでは、メモリアーティファクトは sandbox ワークスペースの `memories/` 配下に保存されます。後の実行で再利用するには、同じライブ sandbox セッションを維持するか、永続化されたセッション状態またはスナップショットから再開することで、設定済みのメモリディレクトリ全体を保持して再利用してください。新しい空の sandbox は空のメモリで開始されます。
|
||||
|
||||
`Memory()` は、メモリの読み取りと生成の両方を有効にします。メモリを読み取るが新しいメモリを生成すべきではないエージェントには `Memory(generate=None)` を使用します。たとえば、内部エージェント、subagent、checker、またはシグナルをあまり追加しない単発のツールエージェントです。実行で後のためにメモリを生成すべきだが、既存のメモリの影響は受けたくない場合は、`Memory(read=None)` を使用します。
|
||||
`Memory()` は、メモリの読み取りと生成の両方を有効にします。メモリを読み取るが新しいメモリは生成すべきでないエージェントには、`Memory(generate=None)` を使用します。たとえば、内部エージェント、サブエージェント、チェッカー、または実行から得られるシグナルが多くない 1 回限りのツールエージェントです。後で使うメモリを生成する必要はあるものの、ユーザーが既存メモリによる影響を望まない場合は、`Memory(read=None)` を使用します。
|
||||
|
||||
## メモリの読み取り
|
||||
|
||||
メモリの読み取りでは段階的開示を使用します。実行開始時に、SDK は一般的に有用なヒント、ユーザーの好み、利用可能なメモリの小さな要約(`memory_summary.md`)をエージェントの開発者プロンプトに注入します。これにより、過去の作業が関連しそうかどうかをエージェントが判断するための十分なコンテキストが与えられます。
|
||||
メモリ読み取りでは段階的開示を使用します。実行の開始時に、SDK は一般的に役立つヒント、ユーザーの好み、利用可能なメモリの小さなサマリー(`memory_summary.md`)を、エージェントの developer プロンプトに挿入します。これにより、エージェントは過去の作業が関連しそうかどうかを判断するのに十分なコンテキストを得られます。
|
||||
|
||||
過去の作業が関連していそうな場合、エージェントは現在のタスクのキーワードを使って、設定されたメモリインデックス(`memories_dir` 配下の `MEMORY.md`)を検索します。さらに詳しい情報が必要な場合にのみ、設定された `rollout_summaries/` ディレクトリー配下の対応する過去の rollout 要約を開きます。
|
||||
過去の作業が関連しそうな場合、エージェントは現在のタスクからキーワードを抽出して、設定されたメモリインデックス(`memories_dir` 配下の `MEMORY.md`)を検索します。より詳細が必要な場合にのみ、設定された `rollout_summaries/` ディレクトリ配下にある対応する過去のロールアウトサマリーを開きます。
|
||||
|
||||
メモリは古くなることがあります。エージェントには、メモリはあくまで参考情報として扱い、現在の環境を信頼するよう指示されています。デフォルトでは、メモリ読み取りでは `live_update` が有効になっているため、エージェントが古いメモリを見つけた場合、同じ実行内で設定された `MEMORY.md` を更新できます。たとえば、その実行がレイテンシーに敏感な場合など、エージェントがメモリを読み取るだけで実行中に変更すべきでない場合は、ライブ更新を無効にしてください。
|
||||
メモリは古くなることがあります。エージェントには、メモリをガイダンスとしてのみ扱い、現在の環境を信頼するよう指示されています。デフォルトでは、メモリ読み取りでは `live_update` が有効です。そのため、エージェントが古くなったメモリを発見した場合、同じ実行内で設定済みの `MEMORY.md` を更新できます。実行中にメモリを読み取るが変更してほしくない場合、たとえばレイテンシに敏感な実行では、ライブ更新を無効にしてください。
|
||||
|
||||
## メモリの生成
|
||||
|
||||
実行が終了すると、sandbox ランタイムはその実行セグメントを会話ファイルに追記します。蓄積された会話ファイルは、sandbox セッションが閉じられるときに処理されます。
|
||||
実行が完了すると、sandbox ランタイムはその実行セグメントを会話ファイルに追記します。蓄積された会話ファイルは、sandbox セッションが閉じられるときに処理されます。
|
||||
|
||||
メモリ生成には 2 つのフェーズがあります。
|
||||
|
||||
1. フェーズ 1: 会話抽出。メモリ生成モデルが蓄積された 1 つの会話ファイルを処理し、会話要約を生成します。system、developer、および reasoning の内容は省略されます。会話が長すぎる場合は、先頭と末尾を保持したまま、コンテキストウィンドウに収まるように切り詰められます。また、フェーズ 2 で統合できるよう、会話からの簡潔なメモである raw メモリ抽出も生成されます。
|
||||
2. フェーズ 2: レイアウト統合。統合エージェントが 1 つのメモリレイアウトの raw メモリを読み取り、さらに証拠が必要な場合は会話要約を開き、パターンを `MEMORY.md` と `memory_summary.md` に抽出します。
|
||||
1. フェーズ 1: 会話の抽出。メモリ生成モデルが、蓄積された 1 つの会話ファイルを処理し、会話サマリーを生成します。system、developer、reasoning のコンテンツは省略されます。会話が長すぎる場合は、先頭と末尾を保持したうえで、コンテキストウィンドウに収まるよう切り詰められます。また、未加工のメモリ抽出も生成します。これは、フェーズ 2 が統合できる会話からの簡潔なメモです。
|
||||
2. フェーズ 2: レイアウトの統合。統合エージェントは、1 つのメモリレイアウトに対応する未加工のメモリを読み取り、より多くの根拠が必要な場合は会話サマリーを開き、パターンを `MEMORY.md` と `memory_summary.md` に抽出します。
|
||||
|
||||
デフォルトのワークスペースレイアウトは次のとおりです。
|
||||
|
||||
@@ -83,7 +83,7 @@ workspace/
|
||||
└── skills/
|
||||
```
|
||||
|
||||
`MemoryGenerateConfig` を使ってメモリ生成を設定できます。
|
||||
`MemoryGenerateConfig` でメモリ生成を設定できます。
|
||||
|
||||
```python
|
||||
from agents.sandbox import MemoryGenerateConfig
|
||||
@@ -97,13 +97,13 @@ memory = Memory(
|
||||
)
|
||||
```
|
||||
|
||||
`extra_prompt` を使うと、GTM エージェント向けの顧客情報や企業情報のように、どのシグナルがユースケースで最も重要かをメモリ生成器に伝えられます。
|
||||
`extra_prompt` を使用して、GTM エージェント向けの顧客や会社の詳細など、ユースケースで最も重要なシグナルをメモリ生成器に伝えます。
|
||||
|
||||
最近の raw メモリが `max_raw_memories_for_consolidation`(デフォルトは 256)を超える場合、フェーズ 2 は最新の会話のメモリだけを保持し、古いものを削除します。新しさは、その会話が最後に更新された時刻に基づきます。この忘却メカニズムにより、メモリは最新の環境を反映しやすくなります。
|
||||
最近の未加工メモリが `max_raw_memories_for_consolidation`(デフォルトは 256)を超える場合、フェーズ 2 は最新の会話のメモリだけを保持し、古いものを削除します。新しさは、会話が最後に更新された時刻に基づきます。この忘却メカニズムにより、メモリが最新の環境を反映しやすくなります。
|
||||
|
||||
## マルチターン会話
|
||||
|
||||
マルチターンの sandbox チャットでは、通常の SDK `Session` を同じライブ sandbox セッションと組み合わせて使用します。
|
||||
マルチターンの sandbox チャットでは、同じライブ sandbox セッションとともに通常の SDK `Session` を使用します。
|
||||
|
||||
```python
|
||||
from agents import Runner, SQLiteSession
|
||||
@@ -132,20 +132,20 @@ async with sandbox:
|
||||
)
|
||||
```
|
||||
|
||||
両方の実行は同じメモリ会話ファイルに追記されます。これは、同じ SDK 会話セッション(`session=conversation_session`)を渡すことで、同じ `session.session_id` を共有するためです。これは、ライブワークスペースを識別する sandbox(`sandbox`)とは異なり、メモリ会話 ID としては使用されません。フェーズ 1 は sandbox セッションが閉じられたときに蓄積された会話を参照するため、分離された 2 つのターンではなく、やり取り全体からメモリを抽出できます。
|
||||
どちらの実行も、同じ SDK 会話セッション(`session=conversation_session`)を渡すため、1 つのメモリ会話ファイルに追記され、したがって同じ `session.session_id` を共有します。これはライブワークスペースを識別する sandbox(`sandbox`)とは異なります。`sandbox` はメモリ会話 ID としては使用されません。sandbox セッションが閉じられると、フェーズ 1 は蓄積された会話を参照するため、2 つの孤立したターンではなく、やり取り全体からメモリを抽出できます。
|
||||
|
||||
複数の `Runner.run(...)` 呼び出しを 1 つのメモリ会話にしたい場合は、それらの呼び出しにまたがって安定した識別子を渡してください。メモリが実行を会話に関連付けるときは、次の順序で解決されます。
|
||||
複数の `Runner.run(...)` 呼び出しを 1 つのメモリ会話にしたい場合は、それらの呼び出し全体で安定した識別子を渡してください。メモリが実行を会話に関連付けるときは、次の順序で解決します。
|
||||
|
||||
1. `Runner.run(...)` に渡した `conversation_id`
|
||||
2. `SQLiteSession` などの SDK `Session` を渡した場合の `session.session_id`
|
||||
3. 上記のいずれも存在しない場合の `RunConfig.group_id`
|
||||
4. 安定した識別子が存在しない場合の、実行ごとに生成される ID
|
||||
1. `conversation_id`(`Runner.run(...)` に渡した場合)
|
||||
2. `session.session_id`(`SQLiteSession` などの SDK `Session` を渡した場合)
|
||||
3. `RunConfig.group_id`(上記のどちらも存在しない場合)
|
||||
4. 生成された実行ごとの ID(安定した識別子が存在しない場合)
|
||||
|
||||
## 異なるエージェント向けのメモリ分離用レイアウト
|
||||
## エージェントごとのメモリ分離における異なるレイアウトの利用
|
||||
|
||||
メモリの分離は、エージェント名ではなく `MemoryLayoutConfig` に基づきます。同じレイアウトと同じメモリ会話 ID を持つエージェントは、1 つのメモリ会話と 1 つの統合メモリを共有します。異なるレイアウトを持つエージェントは、同じ sandbox ワークスペースを共有していても、別々の rollout ファイル、raw メモリ、`MEMORY.md`、および `memory_summary.md` を保持します。
|
||||
メモリの分離はエージェント名ではなく `MemoryLayoutConfig` に基づきます。同じレイアウトと同じメモリ会話 ID を持つエージェントは、1 つのメモリ会話と 1 つの統合済みメモリを共有します。異なるレイアウトを持つエージェントは、同じ sandbox ワークスペースを共有している場合でも、別々のロールアウトファイル、未加工メモリ、`MEMORY.md`、`memory_summary.md` を保持します。
|
||||
|
||||
複数のエージェントが 1 つの sandbox を共有しているが、メモリを共有すべきでない場合は、別々のレイアウトを使用します。
|
||||
複数のエージェントが 1 つの sandbox を共有するものの、メモリは共有すべきでない場合は、別々のレイアウトを使用します。
|
||||
|
||||
```python
|
||||
from agents import SQLiteSession
|
||||
|
||||
+26
-24
@@ -6,35 +6,35 @@ search:
|
||||
|
||||
!!! warning "ベータ機能"
|
||||
|
||||
Sandbox Agents はベータ版です。一般提供までの間に API の詳細、デフォルト値、対応機能は変更される可能性があり、また時間の経過とともにより高度な機能が追加される予定です。
|
||||
サンドボックスエージェントはベータ版です。一般提供前に API の詳細、デフォルト値、サポートされる機能が変更される可能性があり、今後より高度な機能が追加されることが想定されます。
|
||||
|
||||
現代的なエージェントは、ファイルシステム内の実際のファイルを操作できるときに最も効果的に動作します。Agents SDK の **Sandbox Agents** は、モデルに永続的なワークスペースを提供し、そこでは大規模なドキュメント群の検索、ファイル編集、コマンド実行、成果物の生成、保存された sandbox 状態からの作業再開が可能です。
|
||||
最新のエージェントは、ファイルシステム上の実際のファイルを操作できるときに最も効果を発揮します。Agents SDK の **サンドボックスエージェント** は、モデルに永続的なワークスペースを提供し、大規模なドキュメントセットの検索、ファイル編集、コマンド実行、成果物の生成、保存済みのサンドボックス状態からの作業再開を可能にします。
|
||||
|
||||
SDK は、ファイルのステージング、ファイルシステムツール、シェルアクセス、sandbox のライフサイクル、スナップショット、プロバイダー固有の接続処理を自分で組み合わせることなく、その実行ハーネスを提供します。通常の `Agent` と `Runner` のフローはそのまま維持しつつ、ワークスペース用の `Manifest` 、 sandbox ネイティブツール用の capabilities 、そして作業の実行場所を指定する `SandboxRunConfig` を追加できます。
|
||||
SDK は、ファイルのステージング、ファイルシステムツール、シェルアクセス、サンドボックスのライフサイクル、スナップショット、プロバイダー固有の連携を自分で組み合わせる必要なく、その実行ハーネスを提供します。通常の `Agent` と `Runner` のフローはそのままに、ワークスペース用の `Manifest`、サンドボックスネイティブツール用の機能、作業の実行場所を指定する `SandboxRunConfig` を追加します。
|
||||
|
||||
## 前提条件
|
||||
|
||||
- Python 3.10 以上
|
||||
- OpenAI Agents SDK の基本的な知識
|
||||
- sandbox クライアント。ローカル開発では、まず `UnixLocalSandboxClient` から始めてください。
|
||||
- OpenAI Agents SDK に関する基本的な理解
|
||||
- サンドボックスクライアント。ローカル開発では、`UnixLocalSandboxClient` から始めてください。
|
||||
|
||||
## インストール
|
||||
|
||||
まだ SDK をインストールしていない場合は、次を実行してください。
|
||||
SDK をまだインストールしていない場合:
|
||||
|
||||
```bash
|
||||
pip install openai-agents
|
||||
```
|
||||
|
||||
Docker ベースの sandbox の場合:
|
||||
Docker ベースのサンドボックスの場合:
|
||||
|
||||
```bash
|
||||
pip install "openai-agents[docker]"
|
||||
```
|
||||
|
||||
## ローカル sandbox エージェントの作成
|
||||
## ローカルサンドボックスエージェントの作成
|
||||
|
||||
この例では、ローカルのリポジトリを `repo/` 配下にステージングし、ローカル skills を遅延読み込みし、 runner が実行時に Unix ローカル sandbox セッションを作成できるようにします。
|
||||
この例では、ローカルリポジトリを `repo/` 配下にステージングし、ローカルスキルを遅延ロードし、ランナーが実行用の Unix ローカルサンドボックスセッションを作成できるようにします。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -69,6 +69,8 @@ def build_agent(model: str) -> SandboxAgent[None]:
|
||||
capabilities=Capabilities.default() + [
|
||||
Skills(
|
||||
lazy_from=LocalDirLazySkillSource(
|
||||
# This is a host path read by the SDK process.
|
||||
# Requested skills are copied into `skills_path` in the sandbox.
|
||||
source=LocalDir(src=HOST_SKILLS_DIR),
|
||||
)
|
||||
),
|
||||
@@ -78,7 +80,7 @@ def build_agent(model: str) -> SandboxAgent[None]:
|
||||
|
||||
async def main() -> None:
|
||||
result = await Runner.run(
|
||||
build_agent("gpt-5.4"),
|
||||
build_agent("gpt-5.5"),
|
||||
"Open `repo/task.md`, fix the issue, run the targeted test, and summarize the change.",
|
||||
run_config=RunConfig(
|
||||
sandbox=SandboxRunConfig(client=UnixLocalSandboxClient()),
|
||||
@@ -92,24 +94,24 @@ if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
[examples/sandbox/docs/coding_task.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/docs/coding_task.py) を参照してください。この例では小さなシェルベースのリポジトリを使用しているため、Unix ローカル実行全体で決定論的に検証できます。
|
||||
[examples/sandbox/docs/coding_task.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/docs/coding_task.py) を参照してください。この例は小さなシェルベースのリポジトリを使用しているため、Unix ローカル実行間で決定論的に検証できます。
|
||||
|
||||
## 主な選択肢
|
||||
|
||||
基本的な実行が動作したら、次に多くの人が選ぶ項目は以下です。
|
||||
基本的な実行が動作したら、次に多くの人が検討する選択肢は次のとおりです。
|
||||
|
||||
- `default_manifest`: 新しい sandbox セッション用のファイル、リポジトリ、ディレクトリ、マウント
|
||||
- `instructions`: プロンプト全体に適用される短いワークフロールール
|
||||
- `base_instructions`: SDK の sandbox プロンプトを置き換えるための高度なエスケープハッチ
|
||||
- `capabilities`: ファイルシステム編集 / 画像検査、シェル、 skills 、メモリ、 compaction などの sandbox ネイティブツール
|
||||
- `run_as`: モデル向けツールにおける sandbox ユーザー ID
|
||||
- `SandboxRunConfig.client`: sandbox バックエンド
|
||||
- `SandboxRunConfig.session` 、 `session_state` 、または `snapshot`: 後続の実行を以前の作業に再接続する方法
|
||||
- `default_manifest`: 新しいサンドボックスセッション用のファイル、リポジトリ、ディレクトリ、マウント
|
||||
- `instructions`: 複数のプロンプトにわたって適用すべき短いワークフロールール
|
||||
- `base_instructions`: SDK のサンドボックスプロンプトを置き換えるための高度なエスケープハッチ
|
||||
- `capabilities`: ファイルシステム編集 / 画像検査、シェル、スキル、メモリ、コンパクションなどのサンドボックスネイティブツール
|
||||
- `run_as`: モデル向けツールで使用するサンドボックスのユーザー ID
|
||||
- `SandboxRunConfig.client`: サンドボックスバックエンド
|
||||
- `SandboxRunConfig.session`、`session_state`、または `snapshot`: 後続の実行が以前の作業に再接続する方法
|
||||
|
||||
## 次の参照先
|
||||
## 次のステップ
|
||||
|
||||
- [概念](sandbox/guide.md): manifests 、 capabilities 、 permissions 、 snapshots 、 run config 、構成パターンを理解します。
|
||||
- [Sandbox クライアント](sandbox/clients.md): Unix ローカル、 Docker 、ホスト型プロバイダー、マウント戦略を選びます。
|
||||
- [エージェントメモリ](sandbox/memory.md): 以前の sandbox 実行から得た学びを保持し、再利用します。
|
||||
- [概念](sandbox/guide.md): マニフェスト、機能、権限、スナップショット、実行設定、構成パターンを理解します。
|
||||
- [サンドボックスクライアント](sandbox/clients.md): Unix ローカル、Docker、ホスト型プロバイダー、マウント戦略を選択します。
|
||||
- [エージェントメモリ](sandbox/memory.md): 以前のサンドボックス実行から得た知見を保存し、再利用します。
|
||||
|
||||
シェルアクセスが時々使うツールの 1 つにすぎない場合は、[ツールガイド](tools.md) のホスト型シェルから始めてください。ワークスペースの分離、sandbox クライアントの選択、または sandbox セッションの再開動作が設計の一部である場合は、sandbox エージェントを使用してください。
|
||||
シェルアクセスがたまに使うツールの 1 つにすぎない場合は、[ツールガイド](tools.md) のホスト型シェルから始めてください。ワークスペース分離、サンドボックスクライアントの選択、またはサンドボックスセッションの再開動作が設計に含まれる場合は、サンドボックスエージェントを選択してください。
|
||||
@@ -4,15 +4,15 @@ search:
|
||||
---
|
||||
# 高度な SQLite セッション
|
||||
|
||||
`AdvancedSQLiteSession` は、基本的な `SQLiteSession` の拡張版であり、会話の分岐、詳細な使用状況分析、構造化された会話クエリなどの高度な会話管理機能を提供します。
|
||||
`AdvancedSQLiteSession` は基本的な `SQLiteSession` の拡張版であり、会話の分岐、詳細な使用状況分析、構造化された会話クエリなど、高度な会話管理機能を提供します。
|
||||
|
||||
## 機能
|
||||
|
||||
- **会話の分岐**: 任意のユーザーメッセージから代替の会話パスを作成
|
||||
- **使用状況トラッキング**: 各ターンごとの詳細なトークン使用状況分析(完全な JSON 内訳付き)
|
||||
- **構造化クエリ**: ターン単位の会話、ツール使用統計などを取得
|
||||
- **ブランチ管理**: 独立したブランチ切り替えと管理
|
||||
- **メッセージ構造メタデータ**: メッセージタイプ、ツール使用状況、会話フローを追跡
|
||||
- **会話の分岐**: 任意のユーザーメッセージから別の会話パスを作成します
|
||||
- **使用状況の追跡**: ターンごとの詳細なトークン使用状況分析を、完全な JSON 内訳付きで提供します
|
||||
- **構造化クエリ**: ターン別の会話、ツール使用状況の統計などを取得します
|
||||
- **ブランチ管理**: 独立したブランチ切り替えと管理を行います
|
||||
- **メッセージ構造メタデータ**: メッセージタイプ、ツール使用、会話フローを追跡します
|
||||
|
||||
## クイックスタート
|
||||
|
||||
@@ -84,16 +84,16 @@ session = AdvancedSQLiteSession(
|
||||
|
||||
### パラメーター
|
||||
|
||||
- `session_id` (str): 会話セッションの一意な識別子
|
||||
- `db_path` (str | Path): SQLite データベースファイルへのパス。デフォルトはメモリ内ストレージ用の `:memory:`
|
||||
- `create_tables` (bool): 高度なテーブルを自動作成するかどうか。デフォルトは `False`
|
||||
- `logger` (logging.Logger | None): セッション用のカスタムロガー。デフォルトはモジュールロガー
|
||||
- `session_id` (str): 会話セッションの一意の識別子
|
||||
- `db_path` (str | Path): SQLite データベースファイルへのパス。デフォルトはインメモリストレージ用の `:memory:` です
|
||||
- `create_tables` (bool): 高度なテーブルを自動的に作成するかどうか。デフォルトは `False` です
|
||||
- `logger` (logging.Logger | None): セッション用のカスタムロガー。デフォルトはモジュールロガーです
|
||||
|
||||
## 使用状況トラッキング
|
||||
## 使用状況の追跡
|
||||
|
||||
AdvancedSQLiteSession は、会話ターンごとのトークン使用データを保存することで、詳細な使用状況分析を提供します。**これは各エージェント実行後に `store_run_usage` メソッドが呼び出されることに完全に依存します。**
|
||||
AdvancedSQLiteSession は、会話ターンごとにトークン使用状況データを保存することで、詳細な使用状況分析を提供します。 **これは、各エージェント実行後に `store_run_usage` メソッドが呼び出されることに完全に依存します。**
|
||||
|
||||
### 使用データの保存
|
||||
### 使用状況データの保存
|
||||
|
||||
```python
|
||||
# After each agent run, store the usage data
|
||||
@@ -107,7 +107,7 @@ await session.store_run_usage(result)
|
||||
# - Detailed JSON token information (if available)
|
||||
```
|
||||
|
||||
### 使用統計の取得
|
||||
### 使用状況統計の取得
|
||||
|
||||
```python
|
||||
# Get session-level usage (all branches)
|
||||
@@ -137,7 +137,7 @@ turn_2_usage = await session.get_turn_usage(user_turn_number=2)
|
||||
|
||||
## 会話の分岐
|
||||
|
||||
AdvancedSQLiteSession の主要機能の 1 つは、任意のユーザーメッセージから会話ブランチを作成できることです。これにより、代替の会話パスを探索できます。
|
||||
AdvancedSQLiteSession の主要機能の 1 つは、任意のユーザーメッセージから会話ブランチを作成し、別の会話パスを探索できることです。
|
||||
|
||||
### ブランチの作成
|
||||
|
||||
@@ -245,17 +245,17 @@ for turn in matching_turns:
|
||||
|
||||
### メッセージ構造
|
||||
|
||||
セッションは、以下を含むメッセージ構造を自動的に追跡します。
|
||||
セッションは、次を含むメッセージ構造を自動的に追跡します。
|
||||
|
||||
- メッセージタイプ (user, assistant, tool_call など)
|
||||
- ツール呼び出し用のツール名
|
||||
- メッセージタイプ(ユーザー、assistant、tool_call など)
|
||||
- ツール呼び出しのツール名
|
||||
- ターン番号とシーケンス番号
|
||||
- ブランチ関連付け
|
||||
- ブランチとの関連付け
|
||||
- タイムスタンプ
|
||||
|
||||
## データベーススキーマ
|
||||
|
||||
AdvancedSQLiteSession は、基本的な SQLite スキーマを次の 2 つの追加テーブルで拡張します。
|
||||
AdvancedSQLiteSession は、基本的な SQLite スキーマを 2 つの追加テーブルで拡張します。
|
||||
|
||||
### message_structure テーブル
|
||||
|
||||
@@ -298,7 +298,7 @@ CREATE TABLE turn_usage (
|
||||
|
||||
## 完全な例
|
||||
|
||||
すべての機能を包括的に示すデモについては、[完全な例](https://github.com/openai/openai-agents-python/tree/main/examples/memory/advanced_sqlite_session_example.py)をご確認ください。
|
||||
すべての機能を包括的に示す [完全な例](https://github.com/openai/openai-agents-python/tree/main/examples/memory/advanced_sqlite_session_example.py) を確認してください。
|
||||
|
||||
|
||||
## API リファレンス
|
||||
|
||||
@@ -4,18 +4,18 @@ search:
|
||||
---
|
||||
# 暗号化セッション
|
||||
|
||||
`EncryptedSession` は、あらゆるセッション実装に対して透過的な暗号化を提供し、古い項目の自動有効期限切れによって会話データを保護します。
|
||||
`EncryptedSession` は、任意のセッション実装に透過的な暗号化を提供し、古いアイテムの自動期限切れによって会話データを保護します。
|
||||
|
||||
## 機能
|
||||
|
||||
- **透過的な暗号化**: あらゆるセッションを Fernet 暗号化でラップします
|
||||
- **セッションごとのキー**: HKDF 鍵導出を使用して、セッションごとに一意の暗号化を行います
|
||||
- **自動有効期限切れ**: TTL が期限切れになると、古い項目は自動的にスキップされます
|
||||
- **そのまま置き換え可能**: 既存のあらゆるセッション実装で動作します
|
||||
- **透過的な暗号化**: 任意のセッションを Fernet 暗号化でラップします
|
||||
- **セッションごとのキー**: HKDF キー導出を使用して、セッションごとに一意の暗号化を行います
|
||||
- **自動期限切れ**: TTL が期限切れになると、古いアイテムは取得時に黙ってスキップされます
|
||||
- **ドロップイン置換**: 既存の任意のセッション実装で動作します
|
||||
|
||||
## インストール
|
||||
|
||||
暗号化セッションには `encrypt` 追加機能が必要です:
|
||||
暗号化セッションには `encrypt` extra が必要です。
|
||||
|
||||
```bash
|
||||
pip install openai-agents[encrypt]
|
||||
@@ -57,7 +57,7 @@ if __name__ == "__main__":
|
||||
|
||||
### 暗号化キー
|
||||
|
||||
暗号化キーには、Fernet キーまたは任意の文字列を使用できます:
|
||||
暗号化キーには、Fernet キーまたは任意の文字列を指定できます。
|
||||
|
||||
```python
|
||||
from agents.extensions.memory import EncryptedSession
|
||||
@@ -81,7 +81,7 @@ session = EncryptedSession(
|
||||
|
||||
### TTL (有効期間)
|
||||
|
||||
暗号化された項目を有効とする期間を設定します:
|
||||
暗号化されたアイテムが有効であり続ける期間を設定します。
|
||||
|
||||
```python
|
||||
# Items expire after 1 hour
|
||||
@@ -101,7 +101,7 @@ session = EncryptedSession(
|
||||
)
|
||||
```
|
||||
|
||||
## 異なるセッションタイプでの使用
|
||||
## さまざまなセッションタイプでの使用
|
||||
|
||||
### SQLite セッションでの使用
|
||||
|
||||
@@ -140,30 +140,30 @@ session = EncryptedSession(
|
||||
|
||||
!!! warning "高度なセッション機能"
|
||||
|
||||
`AdvancedSQLiteSession` のような高度なセッション実装で `EncryptedSession` を使用する場合は、次の点に注意してください:
|
||||
`EncryptedSession` を `AdvancedSQLiteSession` のような高度なセッション実装と使用する場合は、次の点に注意してください。
|
||||
|
||||
- メッセージ内容は暗号化されるため、`find_turns_by_content()` のようなメソッドは効果的に機能しません
|
||||
- コンテンツベースの検索は暗号化データに対して実行されるため、有効性が制限されます
|
||||
- メッセージ内容が暗号化されるため、`find_turns_by_content()` のようなメソッドは効果的に動作しません
|
||||
- コンテンツベースの検索は暗号化されたデータに対して実行されるため、有効性が制限されます
|
||||
|
||||
|
||||
|
||||
## 鍵導出
|
||||
## キー導出
|
||||
|
||||
EncryptedSession は HKDF (HMAC-based Key Derivation Function) を使用して、セッションごとに一意の暗号化キーを導出します:
|
||||
EncryptedSession は HKDF (HMAC-based Key Derivation Function) を使用して、セッションごとに一意の暗号化キーを導出します。
|
||||
|
||||
- **マスターキー**: 提供した暗号化キー
|
||||
- **マスターキー**: 提供された暗号化キー
|
||||
- **セッションソルト**: セッション ID
|
||||
- **Info 文字列**: `"agents.session-store.hkdf.v1"`
|
||||
- **情報文字列**: `"agents.session-store.hkdf.v1"`
|
||||
- **出力**: 32 バイトの Fernet キー
|
||||
|
||||
これにより、次が保証されます:
|
||||
- 各セッションが一意の暗号化キーを持つこと
|
||||
- マスターキーなしではキーを導出できないこと
|
||||
- セッションデータを異なるセッション間で復号できないこと
|
||||
これにより、次のことが保証されます。
|
||||
- 各セッションに一意の暗号化キーがあります
|
||||
- マスターキーがなければキーを導出できません
|
||||
- セッションデータを異なるセッション間で復号できません
|
||||
|
||||
## 自動有効期限切れ
|
||||
## 自動期限切れ
|
||||
|
||||
項目が TTL を超えると、取得時に自動的にスキップされます:
|
||||
アイテムが TTL を超えると、取得時に自動的にスキップされます。
|
||||
|
||||
```python
|
||||
# Items older than TTL are silently ignored
|
||||
|
||||
+128
-93
@@ -4,11 +4,11 @@ search:
|
||||
---
|
||||
# セッション
|
||||
|
||||
Agents SDK は、複数のエージェント実行にまたがって会話履歴を自動的に維持する組み込みのセッションメモリを提供しており、ターン間で `.to_input_list()` を手動で扱う必要をなくします。
|
||||
Agents SDK は、複数のエージェント実行にまたがって会話履歴を自動的に維持する組み込みのセッションメモリを提供し、ターン間で `.to_input_list()` を手動で扱う必要をなくします。
|
||||
|
||||
Sessions は特定のセッションの会話履歴を保存し、明示的な手動メモリ管理を必要とせずにエージェントがコンテキストを維持できるようにします。これは、エージェントに過去のやり取りを記憶させたいチャットアプリケーションや複数ターンの会話を構築する際に特に有用です。
|
||||
セッションは特定のセッションの会話履歴を保存し、明示的な手動メモリ管理を必要とせずに、エージェントがコンテキストを維持できるようにします。これは、エージェントに以前のやり取りを覚えておいてほしいチャットアプリケーションや複数ターンの会話を構築する場合に特に便利です。
|
||||
|
||||
SDK にクライアント側メモリ管理を任せたい場合は sessions を使用してください。Sessions は同一実行内で `conversation_id`、`previous_response_id`、`auto_previous_response_id` と組み合わせることはできません。代わりに OpenAI のサーバー管理による継続を使いたい場合は、session を重ねるのではなくそれらの仕組みのいずれかを選択してください。
|
||||
SDK にクライアント側メモリを管理させたい場合は、セッションを使用します。セッションは、同じ実行内で `conversation_id`、`previous_response_id`、または `auto_previous_response_id` と組み合わせることはできません。代わりに OpenAI のサーバー管理による継続を使用したい場合は、セッションを重ねて使うのではなく、それらの仕組みのいずれかを選択してください。
|
||||
|
||||
## クイックスタート
|
||||
|
||||
@@ -49,9 +49,9 @@ result = Runner.run_sync(
|
||||
print(result.final_output) # "Approximately 39 million"
|
||||
```
|
||||
|
||||
## 同一セッションで中断実行を再開
|
||||
## 同じセッションによる中断された実行の再開
|
||||
|
||||
実行が承認待ちで一時停止した場合は、同じ session インスタンス(または同じバックエンドストアを指す別の session インスタンス)で再開してください。そうすることで、再開したターンは同じ保存済み会話履歴を継続します。
|
||||
実行が承認待ちで一時停止した場合は、同じセッションインスタンス(または同じバッキングストアを指す別のセッションインスタンス)で再開し、再開されたターンが同じ保存済み会話履歴を継続するようにします。
|
||||
|
||||
```python
|
||||
result = await Runner.run(agent, "Delete temporary files that are no longer needed.", session=session)
|
||||
@@ -63,31 +63,31 @@ if result.interruptions:
|
||||
result = await Runner.run(agent, state, session=session)
|
||||
```
|
||||
|
||||
## セッションのコア動作
|
||||
## コアセッション動作
|
||||
|
||||
セッションメモリが有効な場合:
|
||||
|
||||
1. **各実行前**: runner はセッションの会話履歴を自動取得し、入力アイテムの先頭に追加します。
|
||||
2. **各実行後**: 実行中に生成されたすべての新規アイテム(ユーザー入力、assistant 応答、ツール呼び出しなど)が自動的にセッションへ保存されます。
|
||||
3. **コンテキスト保持**: 同じ session を使う後続の各実行には完全な会話履歴が含まれ、エージェントがコンテキストを維持できます。
|
||||
1. **各実行の前**: ランナーはセッションの会話履歴を自動的に取得し、入力アイテムの前に追加します。
|
||||
2. **各実行の後**: 実行中に生成されたすべての新しいアイテム(ユーザー入力、アシスタントの応答、ツール呼び出しなど)がセッションに自動的に保存されます。
|
||||
3. **コンテキストの保持**: 同じセッションでの後続の各実行には完全な会話履歴が含まれるため、エージェントはコンテキストを維持できます。
|
||||
|
||||
これにより、`.to_input_list()` を手動で呼び出して実行間の会話状態を管理する必要がなくなります。
|
||||
これにより、`.to_input_list()` を手動で呼び出したり、実行間の会話状態を管理したりする必要がなくなります。
|
||||
|
||||
## 履歴と新規入力のマージ方法の制御
|
||||
## 履歴と新しい入力のマージ方法の制御
|
||||
|
||||
session を渡すと、runner は通常次のようにモデル入力を準備します:
|
||||
セッションを渡すと、ランナーは通常、モデル入力を次のように準備します。
|
||||
|
||||
1. セッション履歴(`session.get_items(...)` から取得)
|
||||
2. 新しいターンの入力
|
||||
2. 新しいターン入力
|
||||
|
||||
モデル呼び出し前のこのマージ処理をカスタマイズするには [`RunConfig.session_input_callback`][agents.run.RunConfig.session_input_callback] を使用します。コールバックは 2 つのリストを受け取ります:
|
||||
モデル呼び出しの前にこのマージ手順をカスタマイズするには、[`RunConfig.session_input_callback`][agents.run.RunConfig.session_input_callback] を使用します。コールバックは次の 2 つのリストを受け取ります。
|
||||
|
||||
- `history`: 取得されたセッション履歴(すでに入力アイテム形式に正規化済み)
|
||||
- `new_input`: 現在ターンの新しい入力アイテム
|
||||
- `new_input`: 現在のターンの新しい入力アイテム
|
||||
|
||||
モデルに送信する最終的な入力アイテムのリストを返してください。
|
||||
モデルに送信する最終的な入力アイテムのリストを返します。
|
||||
|
||||
コールバックは両方のリストのコピーを受け取るため、安全に変更できます。返されたリストはそのターンのモデル入力を制御しますが、SDK が永続化するのは引き続き新しいターンに属するアイテムのみです。したがって、古い履歴を並べ替えたりフィルタしたりしても、古いセッションアイテムが新しい入力として再保存されることはありません。
|
||||
コールバックは両方のリストのコピーを受け取るため、安全に変更できます。返されたリストはそのターンのモデル入力を制御しますが、SDK は新しいターンに属するアイテムのみを永続化します。そのため、古い履歴を並べ替えたりフィルタリングしたりしても、古いセッションアイテムが新しい入力として再度保存されることはありません。
|
||||
|
||||
```python
|
||||
from agents import Agent, RunConfig, Runner, SQLiteSession
|
||||
@@ -109,16 +109,16 @@ result = await Runner.run(
|
||||
)
|
||||
```
|
||||
|
||||
これは、セッションの保存方法を変更せずに、履歴のカスタムな間引き、並べ替え、または選択的な取り込みが必要な場合に使用します。モデル呼び出し直前にさらに後段の最終処理が必要な場合は、[running agents guide](../running_agents.md) の [`call_model_input_filter`][agents.run.RunConfig.call_model_input_filter] を使用してください。
|
||||
セッションがアイテムを保存する方法を変更せずに、カスタムの枝刈り、並べ替え、または履歴の選択的な取り込みが必要な場合に使用します。モデル呼び出しの直前にさらに最終的な処理が必要な場合は、[エージェント実行ガイド](../running_agents.md)の [`call_model_input_filter`][agents.run.RunConfig.call_model_input_filter] を使用してください。
|
||||
|
||||
## 取得履歴の制限
|
||||
## 取得する履歴の制限
|
||||
|
||||
各実行前にどの程度の履歴を取得するかを制御するには [`SessionSettings`][agents.memory.SessionSettings] を使用します。
|
||||
各実行の前に取得する履歴の量を制御するには、[`SessionSettings`][agents.memory.SessionSettings] を使用します。
|
||||
|
||||
- `SessionSettings(limit=None)`(デフォルト): 利用可能なセッションアイテムをすべて取得
|
||||
- `SessionSettings(limit=N)`: 直近 `N` 件のアイテムのみ取得
|
||||
- `SessionSettings(limit=None)`(デフォルト): 利用可能なすべてのセッションアイテムを取得します
|
||||
- `SessionSettings(limit=N)`: 直近の `N` アイテムのみを取得します
|
||||
|
||||
これは [`RunConfig.session_settings`][agents.run.RunConfig.session_settings] で実行ごとに適用できます:
|
||||
これは、[`RunConfig.session_settings`][agents.run.RunConfig.session_settings] を介して実行ごとに適用できます。
|
||||
|
||||
```python
|
||||
from agents import Agent, RunConfig, Runner, SessionSettings, SQLiteSession
|
||||
@@ -134,13 +134,13 @@ result = await Runner.run(
|
||||
)
|
||||
```
|
||||
|
||||
セッション実装がデフォルトの session settings を公開している場合、`RunConfig.session_settings` はその実行において `None` 以外の値を上書きします。これは、セッションのデフォルト動作を変更せずに取得サイズの上限を設けたい長い会話で有用です。
|
||||
セッション実装がデフォルトのセッション設定を公開している場合、`RunConfig.session_settings` はその実行について `None` ではない値を上書きします。これは、セッションのデフォルト動作を変更せずに取得サイズを上限設定したい長い会話で便利です。
|
||||
|
||||
## メモリ操作
|
||||
|
||||
### 基本操作
|
||||
|
||||
Sessions は会話履歴を管理するための複数の操作をサポートしています:
|
||||
セッションは、会話履歴を管理するためのいくつかの操作をサポートしています。
|
||||
|
||||
```python
|
||||
from agents import SQLiteSession
|
||||
@@ -167,7 +167,7 @@ await session.clear_session()
|
||||
|
||||
### 修正のための pop_item の使用
|
||||
|
||||
`pop_item` メソッドは、会話の最後のアイテムを取り消したり変更したりしたい場合に特に有用です:
|
||||
`pop_item` メソッドは、会話内の最後のアイテムを取り消したり変更したりしたい場合に特に便利です。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, SQLiteSession
|
||||
@@ -202,27 +202,28 @@ SDK は、さまざまなユースケース向けに複数のセッション実
|
||||
|
||||
### 組み込みセッション実装の選択
|
||||
|
||||
以下の詳細な例を読む前に、この表を使って開始点を選んでください。
|
||||
以下の詳細な例を読む前に、開始点を選ぶためにこの表を使用してください。
|
||||
|
||||
| Session type | Best for | Notes |
|
||||
| セッションタイプ | 最適な用途 | 注記 |
|
||||
| --- | --- | --- |
|
||||
| `SQLiteSession` | ローカル開発とシンプルなアプリ | 組み込み、軽量、ファイル永続化またはインメモリ |
|
||||
| `AsyncSQLiteSession` | `aiosqlite` を使った非同期 SQLite | 非同期ドライバー対応の拡張バックエンド |
|
||||
| `RedisSession` | ワーカー / サービス間での共有メモリ | 低レイテンシな分散デプロイに適しています |
|
||||
| `SQLAlchemySession` | 既存データベースを持つ本番アプリ | SQLAlchemy 対応データベースで動作 |
|
||||
| `DaprSession` | Dapr sidecar を使うクラウドネイティブデプロイ | 複数の state store に加え TTL と整合性制御をサポート |
|
||||
| `OpenAIConversationsSession` | OpenAI でのサーバー管理ストレージ | OpenAI Conversations API ベースの履歴 |
|
||||
| `OpenAIResponsesCompactionSession` | 自動圧縮付きの長い会話 | 別のセッションバックエンドをラップ |
|
||||
| `AdvancedSQLiteSession` | 分岐 / 分析機能付き SQLite | 機能セットが大きめ。専用ページを参照 |
|
||||
| `EncryptedSession` | 別セッションの上に暗号化 + TTL | ラッパー。先に基盤バックエンドを選択 |
|
||||
| `SQLiteSession` | ローカル開発とシンプルなアプリ | 組み込み、軽量、ファイルバックまたはインメモリ |
|
||||
| `AsyncSQLiteSession` | `aiosqlite` を使用した非同期 SQLite | 非同期ドライバー対応の拡張バックエンド |
|
||||
| `RedisSession` | ワーカーやサービス間で共有するメモリ | 低レイテンシの分散デプロイに適しています |
|
||||
| `SQLAlchemySession` | 既存データベースを使用する本番アプリ | SQLAlchemy がサポートするデータベースで動作します |
|
||||
| `MongoDBSession` | すでに MongoDB を使用しているアプリ、またはマルチプロセスストレージが必要なアプリ | 非同期 pymongo;順序付け用のアトミックシーケンスカウンター |
|
||||
| `DaprSession` | Dapr サイドカーを使用するクラウドネイティブデプロイ | 複数のステートストアに加え、TTL と整合性制御をサポートします |
|
||||
| `OpenAIConversationsSession` | OpenAI でのサーバー管理ストレージ | OpenAI Conversations API をバックエンドとする履歴 |
|
||||
| `OpenAIResponsesCompactionSession` | 自動圧縮を伴う長い会話 | 別のセッションバックエンドをラップします |
|
||||
| `AdvancedSQLiteSession` | SQLite に加えて分岐や分析 | より多機能です。専用ページを参照してください |
|
||||
| `EncryptedSession` | 別のセッション上での暗号化と TTL | ラッパーです。まず基盤となるバックエンドを選択してください |
|
||||
|
||||
一部の実装には追加の詳細を説明した専用ページがあり、それらは各サブセクション内でリンクされています。
|
||||
一部の実装には、追加の詳細を含む専用ページがあります。それらは各サブセクション内でリンクされています。
|
||||
|
||||
ChatKit 用の Python サーバーを実装する場合は、ChatKit のスレッドとアイテム永続化に `chatkit.store.Store` 実装を使用してください。`SQLAlchemySession` などの Agents SDK セッションは SDK 側の会話履歴を管理しますが、ChatKit の store のそのままの置き換えにはなりません。[ChatKit データストアの実装に関する `chatkit-python` ガイド](https://github.com/openai/chatkit-python/blob/main/docs/guides/respond-to-user-message.md#implement-your-chatkit-data-store) を参照してください。
|
||||
ChatKit 用の Python サーバーを実装している場合は、ChatKit のスレッドとアイテムの永続化に `chatkit.store.Store` 実装を使用してください。`SQLAlchemySession` などの Agents SDK セッションは SDK 側の会話履歴を管理しますが、ChatKit のストアのドロップイン置き換えではありません。[ChatKit データストアの実装に関する `chatkit-python` ガイド](https://github.com/openai/chatkit-python/blob/main/docs/guides/respond-to-user-message.md#implement-your-chatkit-data-store)を参照してください。
|
||||
|
||||
### OpenAI Conversations API セッション
|
||||
|
||||
`OpenAIConversationsSession` を通じて [OpenAI's Conversations API](https://platform.openai.com/docs/api-reference/conversations) を使用します。
|
||||
`OpenAIConversationsSession` を通じて [OpenAI の Conversations API](https://platform.openai.com/docs/api-reference/conversations)を使用します。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, OpenAIConversationsSession
|
||||
@@ -258,7 +259,7 @@ print(result.final_output) # "California"
|
||||
|
||||
### OpenAI Responses 圧縮セッション
|
||||
|
||||
Responses API(`responses.compact`)で保存済み会話履歴を圧縮するには `OpenAIResponsesCompactionSession` を使用します。これは基盤となる session をラップし、`should_trigger_compaction` に基づいて各ターン後に自動圧縮できます。`OpenAIConversationsSession` をこれでラップしないでください。これら 2 つの機能は履歴を異なる方法で管理します。
|
||||
Responses API(`responses.compact`)で保存済みの会話履歴を圧縮するには、`OpenAIResponsesCompactionSession` を使用します。これは基盤となるセッションをラップし、`should_trigger_compaction` に基づいて各ターンの後に自動的に圧縮できます。`OpenAIConversationsSession` をこれでラップしないでください。この 2 つの機能は異なる方法で履歴を管理します。
|
||||
|
||||
#### 一般的な使用方法(自動圧縮)
|
||||
|
||||
@@ -277,17 +278,17 @@ result = await Runner.run(agent, "Hello", session=session)
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
デフォルトでは、候補しきい値に達すると各ターン後に圧縮が実行されます。
|
||||
デフォルトでは、候補しきい値に達すると各ターンの後に圧縮が実行されます。
|
||||
|
||||
`compaction_mode="previous_response_id"` は、すでに Responses API の response ID でターンを連結している場合に最適です。`compaction_mode="input"` は代わりに現在のセッションアイテムから圧縮リクエストを再構築します。これは response chain が利用できない場合や、セッション内容を信頼できる唯一の情報源にしたい場合に有用です。デフォルトの `"auto"` は、利用可能な中で最も安全な選択肢を選びます。
|
||||
`compaction_mode="previous_response_id"` は、Responses API の応答 ID でターンをすでに連鎖させている場合に最も適しています。`compaction_mode="input"` は、代わりに現在のセッションアイテムから圧縮リクエストを再構築します。これは、応答チェーンが利用できない場合や、セッション内容を信頼できる情報源にしたい場合に便利です。デフォルトの `"auto"` は、利用可能な中で最も安全な選択肢を選びます。
|
||||
|
||||
エージェント実行で `ModelSettings(store=False)` を使うと、Responses API は後で参照するための最新 response を保持しません。このステートレス構成では、デフォルトの `"auto"` モードは `previous_response_id` に依存せず、入力ベース圧縮にフォールバックします。完全な例は [`examples/memory/compaction_session_stateless_example.py`](https://github.com/openai/openai-agents-python/tree/main/examples/memory/compaction_session_stateless_example.py) を参照してください。
|
||||
エージェントが `ModelSettings(store=False)` で実行される場合、Responses API は後で検索するための最後の応答を保持しません。このステートレスな構成では、デフォルトの `"auto"` モードは `previous_response_id` に依存するのではなく、入力ベースの圧縮にフォールバックします。完全な例については、[`examples/memory/compaction_session_stateless_example.py`](https://github.com/openai/openai-agents-python/tree/main/examples/memory/compaction_session_stateless_example.py) を参照してください。
|
||||
|
||||
#### 自動圧縮はストリーミングをブロックする場合があります
|
||||
#### auto-compaction によるストリーミングのブロック
|
||||
|
||||
圧縮はセッション履歴をクリアして再書き込みするため、SDK は圧縮完了前に実行完了と見なしません。ストリーミングモードでは、圧縮が重い場合、最後の出力トークンの後も `run.stream_events()` が数秒開いたままになることがあります。
|
||||
圧縮はセッション履歴をクリアして書き換えるため、SDK は実行完了とみなす前に圧縮の完了を待ちます。ストリーミングモードでは、圧縮が重い場合、最後の出力トークンの後も `run.stream_events()` が数秒間開いたままになることがあります。
|
||||
|
||||
低レイテンシなストリーミングや高速なターン交代が必要な場合は、自動圧縮を無効化し、ターン間(またはアイドル時間)に `run_compaction()` を手動で呼び出してください。圧縮を強制するタイミングは独自の基準で決められます。
|
||||
低レイテンシのストリーミングや高速なターン処理が必要な場合は、自動圧縮を無効にし、ターン間(またはアイドル時間中)に自分で `run_compaction()` を呼び出してください。独自の基準に基づいて、いつ圧縮を強制するかを決めることができます。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, SQLiteSession
|
||||
@@ -310,7 +311,7 @@ await session.run_compaction({"force": True})
|
||||
|
||||
### SQLite セッション
|
||||
|
||||
SQLite を使用したデフォルトの軽量セッション実装です:
|
||||
SQLite を使用するデフォルトの軽量セッション実装です。
|
||||
|
||||
```python
|
||||
from agents import SQLiteSession
|
||||
@@ -331,7 +332,7 @@ result = await Runner.run(
|
||||
|
||||
### 非同期 SQLite セッション
|
||||
|
||||
`aiosqlite` をバックエンドにした SQLite 永続化が必要な場合は `AsyncSQLiteSession` を使用します。
|
||||
`aiosqlite` をバックエンドとする SQLite の永続化が必要な場合は、`AsyncSQLiteSession` を使用します。
|
||||
|
||||
```bash
|
||||
pip install aiosqlite
|
||||
@@ -348,7 +349,7 @@ result = await Runner.run(agent, "Hello", session=session)
|
||||
|
||||
### Redis セッション
|
||||
|
||||
複数のワーカーやサービス間でセッションメモリを共有するには `RedisSession` を使用します。
|
||||
複数のワーカーまたはサービス間で共有セッションメモリを使用するには、`RedisSession` を使用します。
|
||||
|
||||
```bash
|
||||
pip install openai-agents[redis]
|
||||
@@ -368,7 +369,7 @@ result = await Runner.run(agent, "Hello", session=session)
|
||||
|
||||
### SQLAlchemy セッション
|
||||
|
||||
SQLAlchemy 対応の任意のデータベースを使用した、本番対応の Agents SDK セッション永続化:
|
||||
SQLAlchemy がサポートする任意のデータベースを使用した、本番対応の Agents SDK セッション永続化です。
|
||||
|
||||
```python
|
||||
from agents.extensions.memory import SQLAlchemySession
|
||||
@@ -386,11 +387,11 @@ engine = create_async_engine("postgresql+asyncpg://user:pass@localhost/db")
|
||||
session = SQLAlchemySession("user_123", engine=engine, create_tables=True)
|
||||
```
|
||||
|
||||
詳細は [SQLAlchemy Sessions](sqlalchemy_session.md) を参照してください。
|
||||
詳細なドキュメントについては、[SQLAlchemy セッション](sqlalchemy_session.md)を参照してください。
|
||||
|
||||
### Dapr セッション
|
||||
|
||||
すでに Dapr sidecar を運用している場合、またはエージェントコードを変更せずに異なる state-store バックエンド間で移行可能なセッションストレージが必要な場合は `DaprSession` を使用します。
|
||||
すでに Dapr サイドカーを実行している場合、またはエージェントコードを変更せずに異なるステートストアバックエンドへ移行できるセッションストレージが必要な場合は、`DaprSession` を使用します。
|
||||
|
||||
```bash
|
||||
pip install openai-agents[dapr]
|
||||
@@ -411,18 +412,50 @@ async with DaprSession.from_address(
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
注意:
|
||||
注記:
|
||||
|
||||
- `from_address(...)` は Dapr クライアントを作成して所有します。アプリですでに管理している場合は、`dapr_client=...` を指定して直接 `DaprSession(...)` を構築してください。
|
||||
- 基盤 state store が TTL をサポートしている場合、`ttl=...` を渡すと古いセッションデータを自動期限切れにできます。
|
||||
- より強い read-after-write 保証が必要な場合は `consistency=DAPR_CONSISTENCY_STRONG` を渡してください。
|
||||
- Dapr Python SDK は HTTP sidecar endpoint も確認します。ローカル開発では、`dapr_address` で使用する gRPC ポートに加えて、`--dapr-http-port 3500` でも Dapr を起動してください。
|
||||
- ローカルコンポーネントやトラブルシューティングを含む完全なセットアップ手順は [`examples/memory/dapr_session_example.py`](https://github.com/openai/openai-agents-python/tree/main/examples/memory/dapr_session_example.py) を参照してください。
|
||||
- `from_address(...)` は Dapr クライアントを作成し、所有します。アプリがすでにクライアントを管理している場合は、`dapr_client=...` を指定して `DaprSession(...)` を直接構築してください。
|
||||
- 基盤となるステートストアが TTL をサポートしている場合に古いセッションデータを自動的に期限切れにするには、`ttl=...` を渡します。
|
||||
- より強い read-after-write 保証が必要な場合は、`consistency=DAPR_CONSISTENCY_STRONG` を渡します。
|
||||
- Dapr Python SDK は HTTP サイドカーエンドポイントもチェックします。ローカル開発では、`dapr_address` で使用する gRPC ポートに加えて、`--dapr-http-port 3500` でも Dapr を起動してください。
|
||||
- ローカルコンポーネントやトラブルシューティングを含む完全なセットアップ手順については、[`examples/memory/dapr_session_example.py`](https://github.com/openai/openai-agents-python/tree/main/examples/memory/dapr_session_example.py) を参照してください。
|
||||
|
||||
|
||||
### Advanced SQLite セッション
|
||||
### MongoDB セッション
|
||||
|
||||
会話分岐、使用状況分析、構造化クエリを備えた拡張 SQLite セッション:
|
||||
すでに MongoDB を使用しているアプリケーション、または水平スケーラブルでマルチプロセス対応のセッションストレージが必要なアプリケーションには、`MongoDBSession` を使用します。
|
||||
|
||||
```bash
|
||||
pip install openai-agents[mongodb]
|
||||
```
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
from agents.extensions.memory import MongoDBSession
|
||||
|
||||
agent = Agent(name="Assistant")
|
||||
|
||||
# Create from URI — owns the client and closes it when session.close() is called
|
||||
session = MongoDBSession.from_uri(
|
||||
"user-123",
|
||||
uri="mongodb://localhost:27017",
|
||||
database="agents",
|
||||
)
|
||||
result = await Runner.run(agent, "Hello", session=session)
|
||||
print(result.final_output)
|
||||
await session.close()
|
||||
```
|
||||
|
||||
注記:
|
||||
|
||||
- `from_uri(...)` は `AsyncMongoClient` を作成し、所有し、`session.close()` で閉じます。アプリケーションがすでにクライアントを管理している場合は、`client=...` を指定して `MongoDBSession(...)` を直接構築してください。その場合、`session.close()` は no-op となり、ライフサイクルは呼び出し元が保持します。
|
||||
- ほかの変更なしに、`from_uri(...)` に `mongodb+srv://user:password@cluster.example.mongodb.net` URI を渡すことで [MongoDB Atlas](https://www.mongodb.com/products/platform) に接続できます。
|
||||
- 2 つのコレクションが使用され、どちらの名前も `sessions_collection=`(デフォルトは `agent_sessions`)と `messages_collection=`(デフォルトは `agent_messages`)で設定できます。インデックスは初回使用時に自動的に作成されます。各メッセージドキュメントは、同時実行の書き込み元やプロセスをまたいで順序を保持する単調増加の `seq` カウンターを持ちます。
|
||||
- 最初の実行前に接続性を確認するには、`await session.ping()` を使用します。
|
||||
|
||||
### 高度な SQLite セッション
|
||||
|
||||
会話の分岐、使用状況分析、構造化クエリを備えた拡張 SQLite セッションです。
|
||||
|
||||
```python
|
||||
from agents.extensions.memory import AdvancedSQLiteSession
|
||||
@@ -442,11 +475,11 @@ await session.store_run_usage(result) # Track token usage
|
||||
await session.create_branch_from_turn(2) # Branch from turn 2
|
||||
```
|
||||
|
||||
詳細は [Advanced SQLite Sessions](advanced_sqlite_session.md) を参照してください。
|
||||
詳細なドキュメントについては、[高度な SQLite セッション](advanced_sqlite_session.md)を参照してください。
|
||||
|
||||
### Encrypted セッション
|
||||
### 暗号化セッション
|
||||
|
||||
任意のセッション実装向け透過的暗号化ラッパー:
|
||||
任意のセッション実装向けの透過的な暗号化ラッパーです。
|
||||
|
||||
```python
|
||||
from agents.extensions.memory import EncryptedSession, SQLAlchemySession
|
||||
@@ -469,33 +502,34 @@ session = EncryptedSession(
|
||||
result = await Runner.run(agent, "Hello", session=session)
|
||||
```
|
||||
|
||||
詳細は [Encrypted Sessions](encrypted_session.md) を参照してください。
|
||||
詳細なドキュメントについては、[暗号化セッション](encrypted_session.md)を参照してください。
|
||||
|
||||
### その他のセッションタイプ
|
||||
|
||||
このほかにもいくつかの組み込みオプションがあります。`examples/memory/` と `extensions/memory/` 配下のソースコードを参照してください。
|
||||
組み込みの選択肢はほかにもいくつかあります。`examples/memory/` と `extensions/memory/` 以下のソースコードを参照してください。
|
||||
|
||||
## 運用パターン
|
||||
|
||||
### セッション ID 命名
|
||||
### セッション ID の命名
|
||||
|
||||
会話の整理に役立つ、意味のあるセッション ID を使用してください:
|
||||
会話を整理しやすい、意味のあるセッション ID を使用してください。
|
||||
|
||||
- ユーザーベース: `"user_12345"`
|
||||
- スレッドベース: `"thread_abc123"`
|
||||
- コンテキストベース: `"support_ticket_456"`
|
||||
|
||||
### メモリ永続化
|
||||
### メモリの永続化
|
||||
|
||||
- 一時的な会話にはインメモリ SQLite(`SQLiteSession("session_id")`)を使用
|
||||
- 永続的な会話にはファイルベース SQLite(`SQLiteSession("session_id", "path/to/db.sqlite")`)を使用
|
||||
- `aiosqlite` ベース実装が必要な場合は非同期 SQLite(`AsyncSQLiteSession("session_id", db_path="...")`)を使用
|
||||
- 共有の低レイテンシなセッションメモリには Redis バックエンドセッション(`RedisSession.from_url("session_id", url="redis://...")`)を使用
|
||||
- SQLAlchemy が対応する既存データベースを持つ本番システムには SQLAlchemy ベースセッション(`SQLAlchemySession("session_id", engine=engine, create_tables=True)`)を使用
|
||||
- 組み込みテレメトリ、トレーシング、データ分離に加え 30 以上のデータベースバックエンドをサポートする本番クラウドネイティブデプロイには Dapr state store セッション(`DaprSession.from_address("session_id", state_store_name="statestore", dapr_address="localhost:50001")`)を使用
|
||||
- 履歴を OpenAI Conversations API に保存したい場合は OpenAI ホスト型ストレージ(`OpenAIConversationsSession()`)を使用
|
||||
- 任意のセッションを透過的暗号化と TTL ベース期限切れでラップするには暗号化セッション(`EncryptedSession(session_id, underlying_session, encryption_key)`)を使用
|
||||
- より高度なユースケース向けに、他の本番システム(例: Django)向けカスタムセッションバックエンドの実装も検討してください
|
||||
- 一時的な会話にはインメモリ SQLite(`SQLiteSession("session_id")`)を使用します
|
||||
- 永続的な会話にはファイルベース SQLite(`SQLiteSession("session_id", "path/to/db.sqlite")`)を使用します
|
||||
- `aiosqlite` ベースの実装が必要な場合は、非同期 SQLite(`AsyncSQLiteSession("session_id", db_path="...")`)を使用します
|
||||
- 共有された低レイテンシのセッションメモリには、Redis バックのセッション(`RedisSession.from_url("session_id", url="redis://...")`)を使用します
|
||||
- SQLAlchemy がサポートする既存データベースを持つ本番システムには、SQLAlchemy を利用したセッション(`SQLAlchemySession("session_id", engine=engine, create_tables=True)`) を使用します
|
||||
- すでに MongoDB を使用しているアプリケーション、またはマルチプロセスで水平スケーラブルなセッションストレージが必要なアプリケーションには、MongoDB セッション(`MongoDBSession.from_uri("session_id", uri="mongodb://localhost:27017")`)を使用します
|
||||
- 組み込みのテレメトリ、トレーシング、データ分離を備えた 30 以上のデータベースバックエンドをサポートする本番クラウドネイティブデプロイには、Dapr ステートストアセッション(`DaprSession.from_address("session_id", state_store_name="statestore", dapr_address="localhost:50001")`)を使用します
|
||||
- OpenAI Conversations API に履歴を保存したい場合は、OpenAI がホストするストレージ(`OpenAIConversationsSession()`)を使用します
|
||||
- 透過的な暗号化と TTL ベースの有効期限で任意のセッションをラップするには、暗号化セッション(`EncryptedSession(session_id, underlying_session, encryption_key)`)を使用します
|
||||
- より高度なユースケースでは、ほかの本番システム(たとえば Django)向けのカスタムセッションバックエンドの実装を検討してください
|
||||
|
||||
### 複数セッション
|
||||
|
||||
@@ -543,7 +577,7 @@ result2 = await Runner.run(
|
||||
|
||||
## 完全な例
|
||||
|
||||
セッションメモリの動作を示す完全な例です:
|
||||
セッションメモリの動作を示す完全な例を以下に示します。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -607,7 +641,7 @@ if __name__ == "__main__":
|
||||
|
||||
## カスタムセッション実装
|
||||
|
||||
[`Session`][agents.memory.session.Session] プロトコルに従うクラスを作成することで、独自のセッションメモリを実装できます:
|
||||
[`Session`][agents.memory.session.Session] プロトコルに従うクラスを作成することで、独自のセッションメモリを実装できます。
|
||||
|
||||
```python
|
||||
from agents.memory.session import SessionABC
|
||||
@@ -650,27 +684,28 @@ result = await Runner.run(
|
||||
)
|
||||
```
|
||||
|
||||
## コミュニティセッション実装
|
||||
## コミュニティによるセッション実装
|
||||
|
||||
コミュニティでは追加のセッション実装が開発されています:
|
||||
コミュニティは追加のセッション実装を開発しています。
|
||||
|
||||
| Package | Description |
|
||||
| パッケージ | 説明 |
|
||||
|---------|-------------|
|
||||
| [openai-django-sessions](https://pypi.org/project/openai-django-sessions/) | 任意の Django 対応データベース( PostgreSQL、 MySQL、 SQLite など)向けの Django ORM ベースセッション |
|
||||
| [openai-django-sessions](https://pypi.org/project/openai-django-sessions/) | 任意の Django 対応データベース(PostgreSQL、MySQL、SQLite など)向けの Django ORM ベースのセッション |
|
||||
|
||||
セッション実装を作成した場合は、ここに追加するためのドキュメント PR をぜひ送ってください。
|
||||
セッション実装を構築した場合は、ぜひドキュメント PR を送ってここに追加してください。
|
||||
|
||||
## API リファレンス
|
||||
|
||||
詳細な API ドキュメントは以下を参照してください:
|
||||
詳細な API ドキュメントについては、以下を参照してください。
|
||||
|
||||
- [`Session`][agents.memory.session.Session] - プロトコルインターフェース
|
||||
- [`Session`][agents.memory.session.Session] - プロトコルインターフェイス
|
||||
- [`OpenAIConversationsSession`][agents.memory.OpenAIConversationsSession] - OpenAI Conversations API 実装
|
||||
- [`OpenAIResponsesCompactionSession`][agents.memory.openai_responses_compaction_session.OpenAIResponsesCompactionSession] - Responses API 圧縮ラッパー
|
||||
- [`SQLiteSession`][agents.memory.sqlite_session.SQLiteSession] - 基本 SQLite 実装
|
||||
- [`AsyncSQLiteSession`][agents.extensions.memory.async_sqlite_session.AsyncSQLiteSession] - `aiosqlite` ベースの非同期 SQLite 実装
|
||||
- [`RedisSession`][agents.extensions.memory.redis_session.RedisSession] - Redis バックエンドセッション実装
|
||||
- [`SQLAlchemySession`][agents.extensions.memory.sqlalchemy_session.SQLAlchemySession] - SQLAlchemy ベース実装
|
||||
- [`DaprSession`][agents.extensions.memory.dapr_session.DaprSession] - Dapr state store 実装
|
||||
- [`SQLiteSession`][agents.memory.sqlite_session.SQLiteSession] - 基本的な SQLite 実装
|
||||
- [`AsyncSQLiteSession`][agents.extensions.memory.async_sqlite_session.AsyncSQLiteSession] - `aiosqlite` に基づく非同期 SQLite 実装
|
||||
- [`RedisSession`][agents.extensions.memory.redis_session.RedisSession] - Redis バックのセッション実装
|
||||
- [`SQLAlchemySession`][agents.extensions.memory.sqlalchemy_session.SQLAlchemySession] - SQLAlchemy を利用した実装
|
||||
- [`MongoDBSession`][agents.extensions.memory.mongodb_session.MongoDBSession] - MongoDB バックのセッション実装
|
||||
- [`DaprSession`][agents.extensions.memory.dapr_session.DaprSession] - Dapr ステートストア実装
|
||||
- [`AdvancedSQLiteSession`][agents.extensions.memory.advanced_sqlite_session.AdvancedSQLiteSession] - 分岐と分析を備えた拡張 SQLite
|
||||
- [`EncryptedSession`][agents.extensions.memory.encrypt_session.EncryptedSession] - 任意のセッション向け暗号化ラッパー
|
||||
- [`EncryptedSession`][agents.extensions.memory.encrypt_session.EncryptedSession] - 任意のセッション向けの暗号化ラッパー
|
||||
@@ -4,11 +4,11 @@ search:
|
||||
---
|
||||
# SQLAlchemy セッション
|
||||
|
||||
`SQLAlchemySession` は SQLAlchemy を使用して本番運用対応のセッション実装を提供し、セッションストレージに SQLAlchemy がサポートする任意のデータベース ( PostgreSQL 、 MySQL 、 SQLite など ) を使用できます。
|
||||
`SQLAlchemySession` は SQLAlchemy を使用して、本番環境に対応したセッション実装を提供します。これにより、セッションストレージとして SQLAlchemy がサポートする任意のデータベース (PostgreSQL、MySQL、SQLite など) を使用できます。
|
||||
|
||||
## インストール
|
||||
|
||||
SQLAlchemy セッションには `sqlalchemy` extra が必要です。
|
||||
SQLAlchemy セッションには `sqlalchemy` extra が必要です:
|
||||
|
||||
```bash
|
||||
pip install openai-agents[sqlalchemy]
|
||||
@@ -18,7 +18,7 @@ pip install openai-agents[sqlalchemy]
|
||||
|
||||
### データベース URL の使用
|
||||
|
||||
開始する最も簡単な方法です。
|
||||
始めるための最も簡単な方法です:
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -42,9 +42,9 @@ if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
### 既存 engine の使用
|
||||
### 既存エンジンの使用
|
||||
|
||||
既存の SQLAlchemy engine があるアプリケーション向けです。
|
||||
既存の SQLAlchemy エンジンを持つアプリケーション向けです:
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -77,4 +77,4 @@ if __name__ == "__main__":
|
||||
## API リファレンス
|
||||
|
||||
- [`SQLAlchemySession`][agents.extensions.memory.sqlalchemy_session.SQLAlchemySession] - メインクラス
|
||||
- [`Session`][agents.memory.session.Session] - ベースセッションプロトコル
|
||||
- [`Session`][agents.memory.session.Session] - 基本セッションプロトコル
|
||||
+22
-22
@@ -4,19 +4,19 @@ search:
|
||||
---
|
||||
# ストリーミング
|
||||
|
||||
ストリーミングを使うと、エージェントの実行が進行する間の更新を購読できます。これは、エンドユーザーに進捗更新や部分的な応答を表示するのに役立ちます。
|
||||
ストリーミングにより、エージェントの実行が進むにつれて更新を購読できます。これは、エンドユーザーに進捗状況の更新や部分的なレスポンスを表示する場合に役立ちます。
|
||||
|
||||
ストリーミングするには、[`Runner.run_streamed()`][agents.run.Runner.run_streamed] を呼び出します。これにより [`RunResultStreaming`][agents.result.RunResultStreaming] が得られます。`result.stream_events()` を呼び出すと、以下で説明する [`StreamEvent`][agents.stream_events.StreamEvent] オブジェクトの非同期ストリームが得られます。
|
||||
ストリーミングするには、[`Runner.run_streamed()`][agents.run.Runner.run_streamed] を呼び出します。これにより [`RunResultStreaming`][agents.result.RunResultStreaming] が返されます。`result.stream_events()` を呼び出すと、以下で説明する [`StreamEvent`][agents.stream_events.StreamEvent] オブジェクトの非同期ストリームが得られます。
|
||||
|
||||
非同期イテレーターが終了するまで `result.stream_events()` の消費を続けてください。ストリーミング実行は、イテレーターが終了するまで完了しません。また、セッション永続化、承認の記録管理、履歴の圧縮といった後処理は、最後の可視トークン到着後に完了する場合があります。ループを抜けた時点で、`result.is_complete` が最終的な実行状態を反映します。
|
||||
非同期イテレーターが終了するまで、`result.stream_events()` を消費し続けてください。ストリーミング実行は、イテレーターが終了するまで完了しません。また、セッションの永続化、承認の記録管理、履歴の圧縮などの後処理は、最後の可視トークンが到着した後に完了する場合があります。ループが終了すると、`result.is_complete` は最終的な実行状態を反映します。
|
||||
|
||||
## raw response イベント
|
||||
## raw レスポンスイベント
|
||||
|
||||
[`RawResponsesStreamEvent`][agents.stream_events.RawResponsesStreamEvent] は、LLM から直接渡される raw イベントです。これらは OpenAI Responses API 形式であり、各イベントはタイプ(`response.created`、`response.output_text.delta` など)とデータを持ちます。これらのイベントは、生成され次第すぐにレスポンスメッセージをユーザーへストリーミングしたい場合に有用です。
|
||||
[`RawResponsesStreamEvent`][agents.stream_events.RawResponsesStreamEvent] は、LLM から直接渡される raw イベントです。これらは OpenAI Responses API 形式であり、各イベントには型(`response.created`、`response.output_text.delta` など)とデータがあります。これらのイベントは、レスポンスメッセージが生成され次第、ユーザーにストリーミングしたい場合に役立ちます。
|
||||
|
||||
コンピュータツールの raw イベントは、保存済み結果と同じく preview と GA の区別を維持します。Preview フローでは 1 つの `action` を含む `computer_call` アイテムをストリーミングし、`gpt-5.4` ではバッチ化された `actions[]` を含む `computer_call` アイテムをストリーミングできます。より高レベルの [`RunItemStreamEvent`][agents.stream_events.RunItemStreamEvent] サーフェスでは、このためのコンピュータ専用イベント名は追加されません。どちらの形も引き続き `tool_called` として表出し、スクリーンショット結果は `computer_call_output` アイテムをラップした `tool_output` として返されます。
|
||||
コンピュータツールの raw イベントは、保存された実行結果と同じ preview と GA の区別を維持します。Preview フローでは、1 つの `action` を持つ `computer_call` アイテムをストリーミングします。一方、`gpt-5.5` では、バッチ化された `actions[]` を持つ `computer_call` アイテムをストリーミングできます。高レベルの [`RunItemStreamEvent`][agents.stream_events.RunItemStreamEvent] サーフェスは、このために特別なコンピュータ専用イベント名を追加しません。どちらの形式も引き続き `tool_called` として表面化し、スクリーンショットの実行結果は `computer_call_output` アイテムをラップする `tool_output` として返されます。
|
||||
|
||||
たとえば、これは LLM が生成するテキストをトークン単位で出力します。
|
||||
たとえば、これは LLM によって生成されたテキストをトークンごとに出力します。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -41,7 +41,7 @@ if __name__ == "__main__":
|
||||
|
||||
## ストリーミングと承認
|
||||
|
||||
ストリーミングは、ツール承認のために一時停止する実行とも互換性があります。ツールに承認が必要な場合、`result.stream_events()` は終了し、保留中の承認は [`RunResultStreaming.interruptions`][agents.result.RunResultStreaming.interruptions] に公開されます。`result.to_state()` で結果を [`RunState`][agents.run_state.RunState] に変換し、割り込みを承認または拒否してから、`Runner.run_streamed(...)` で再開します。
|
||||
ストリーミングは、ツール承認のために一時停止する実行と互換性があります。ツールに承認が必要な場合、`result.stream_events()` は終了し、保留中の承認は [`RunResultStreaming.interruptions`][agents.result.RunResultStreaming.interruptions] で公開されます。`result.to_state()` を使って実行結果を [`RunState`][agents.run_state.RunState] に変換し、中断を承認または拒否してから、`Runner.run_streamed(...)` で再開します。
|
||||
|
||||
```python
|
||||
result = Runner.run_streamed(agent, "Delete temporary files if they are no longer needed.")
|
||||
@@ -57,25 +57,25 @@ if result.interruptions:
|
||||
pass
|
||||
```
|
||||
|
||||
一時停止 / 再開の完全な手順は、[human-in-the-loop ガイド](human_in_the_loop.md) を参照してください。
|
||||
一時停止/再開の完全なウォークスルーについては、[human-in-the-loop ガイド](human_in_the_loop.md)を参照してください。
|
||||
|
||||
## 現在のターン後のストリーミングキャンセル
|
||||
## 現在のターン後のストリーミングのキャンセル
|
||||
|
||||
ストリーミング実行を途中で停止する必要がある場合は、[`result.cancel()`][agents.result.RunResultStreaming.cancel] を呼び出します。デフォルトでは、これにより実行は即時停止します。停止前に現在のターンをきれいに完了させるには、代わりに `result.cancel(mode="after_turn")` を呼び出してください。
|
||||
途中でストリーミング実行を停止する必要がある場合は、[`result.cancel()`][agents.result.RunResultStreaming.cancel] を呼び出します。デフォルトでは、これにより実行はすぐに停止します。停止する前に現在のターンを正常に完了させるには、代わりに `result.cancel(mode="after_turn")` を呼び出します。
|
||||
|
||||
ストリーミング実行は、`result.stream_events()` が終了するまで完了しません。SDK は、最後の可視トークンの後でも、セッション項目の永続化、承認状態の確定、履歴の圧縮を続ける場合があります。
|
||||
ストリーミング実行は、`result.stream_events()` が終了するまで完了しません。最後の可視トークンの後も、SDK がセッションアイテムを永続化したり、承認状態を確定したり、履歴を圧縮したりしている場合があります。
|
||||
|
||||
[`result.to_input_list(mode="normalized")`][agents.result.RunResultBase.to_input_list] から手動で継続していて、`cancel(mode="after_turn")` がツールターン後に停止した場合は、新しいユーザーターンをすぐ追加するのではなく、その正規化済み入力で `result.last_agent` を再実行して未完了ターンを継続してください。
|
||||
- ストリーミング実行がツール承認で停止した場合、それを新しいターンとして扱わないでください。ストリームの消費を最後まで完了し、`result.interruptions` を確認してから、`result.to_state()` から再開してください。
|
||||
- 次のモデル呼び出し前に、取得したセッション履歴と新しいユーザー入力をどのようにマージするかをカスタマイズするには [`RunConfig.session_input_callback`][agents.run.RunConfig.session_input_callback] を使用します。そこで新規ターン項目を書き換えた場合、そのターンで永続化されるのは書き換え後のバージョンです。
|
||||
[`result.to_input_list(mode="normalized")`][agents.result.RunResultBase.to_input_list] から手動で継続しており、`cancel(mode="after_turn")` がツールターンの後で停止した場合は、すぐに新しいユーザーターンを追加するのではなく、その正規化された入力で `result.last_agent` を再実行して、未完了のターンを継続してください。
|
||||
- ストリーミング実行がツール承認のために停止した場合、それを新しいターンとして扱わないでください。ストリームの読み出しを最後まで完了し、`result.interruptions` を確認して、代わりに `result.to_state()` から再開してください。
|
||||
- 次のモデル呼び出しの前に、取得したセッション履歴と新しいユーザー入力をどのようにマージするかをカスタマイズするには、[`RunConfig.session_input_callback`][agents.run.RunConfig.session_input_callback] を使用します。そこで新しいターンのアイテムを書き換えた場合、その書き換え後のバージョンがそのターンとして永続化されます。
|
||||
|
||||
## 実行項目イベントとエージェントイベント
|
||||
## 実行アイテムイベントとエージェントイベント
|
||||
|
||||
[`RunItemStreamEvent`][agents.stream_events.RunItemStreamEvent] はより高レベルのイベントです。項目が完全に生成されたときに通知します。これにより、各トークン単位ではなく、「メッセージ生成済み」「ツール実行済み」などのレベルで進捗更新を送れます。同様に、[`AgentUpdatedStreamEvent`][agents.stream_events.AgentUpdatedStreamEvent] は、現在のエージェントが変わったとき(例: ハンドオフの結果)に更新を提供します。
|
||||
[`RunItemStreamEvent`][agents.stream_events.RunItemStreamEvent] は、より高レベルのイベントです。アイテムが完全に生成されたタイミングを通知します。これにより、各トークン単位ではなく、「メッセージが生成された」「ツールが実行された」などのレベルで進捗更新を送信できます。同様に、[`AgentUpdatedStreamEvent`][agents.stream_events.AgentUpdatedStreamEvent] は、現在のエージェントが変更されたとき(例: ハンドオフの結果として)に更新を提供します。
|
||||
|
||||
### 実行項目イベント名
|
||||
### 実行アイテムイベント名
|
||||
|
||||
`RunItemStreamEvent.name` は、固定のセマンティックなイベント名セットを使用します。
|
||||
`RunItemStreamEvent.name` は、固定された一連のセマンティックなイベント名を使用します。
|
||||
|
||||
- `message_output_created`
|
||||
- `handoff_requested`
|
||||
@@ -89,11 +89,11 @@ if result.interruptions:
|
||||
- `mcp_approval_response`
|
||||
- `mcp_list_tools`
|
||||
|
||||
`handoff_occured` は、後方互換性のため意図的にスペルミスのままです。
|
||||
`handoff_occured` は、後方互換性のため意図的にスペルミスのままになっています。
|
||||
|
||||
ホスト型ツール検索を使用すると、モデルがツール検索リクエストを発行したときに `tool_search_called` が発行され、Responses API が読み込まれたサブセットを返したときに `tool_search_output_created` が発行されます。
|
||||
ホストされたツール検索を使用する場合、モデルがツール検索リクエストを発行すると `tool_search_called` が送出され、Responses API が読み込まれたサブセットを返すと `tool_search_output_created` が送出されます。
|
||||
|
||||
たとえば、これは raw イベントを無視して、ユーザーへの更新をストリーミングします。
|
||||
たとえば、これは raw イベントを無視し、更新をユーザーにストリーミングします。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
|
||||
+162
-165
@@ -4,42 +4,42 @@ search:
|
||||
---
|
||||
# ツール
|
||||
|
||||
ツールを使うと、エージェントはアクションを実行できます。たとえば、データ取得、コード実行、外部 API 呼び出し、さらにはコンピュータ操作などです。 SDK は 5 つのカテゴリーをサポートしています。
|
||||
ツールにより、エージェントは、データの取得、コードの実行、外部 API の呼び出し、さらにはコンピュータ操作といったアクションを実行できます。SDK は 5 つのカテゴリーをサポートしています:
|
||||
|
||||
- OpenAI がホストするツール: OpenAI サーバー上でモデルと並行して実行されます。
|
||||
- ローカル / ランタイム実行ツール: `ComputerTool` と `ApplyPatchTool` は常にあなたの環境で実行され、`ShellTool` はローカルまたはホストコンテナで実行できます。
|
||||
- Function Calling: 任意の Python 関数をツールとしてラップします。
|
||||
- ローカル / ランタイム実行ツール: `ComputerTool` と `ApplyPatchTool` は常にお使いの環境で実行され、`ShellTool` はローカルまたはホスト型コンテナーで実行できます。
|
||||
- Function calling: 任意の Python 関数をツールとしてラップします。
|
||||
- Agents as tools: 完全なハンドオフなしで、エージェントを呼び出し可能なツールとして公開します。
|
||||
- Experimental: Codex tool: ツール呼び出しから、ワークスペーススコープの Codex タスクを実行します。
|
||||
- 実験的: Codex ツール: ツール呼び出しからワークスペーススコープの Codex タスクを実行します。
|
||||
|
||||
## ツールタイプの選択
|
||||
|
||||
このページをカタログとして使い、次に自分が制御するランタイムに合うセクションへ進んでください。
|
||||
このページをカタログとして使用し、管理するランタイムに一致するセクションへ進んでください。
|
||||
|
||||
| 次をしたい場合... | ここから開始 |
|
||||
| やりたいこと | 開始場所 |
|
||||
| --- | --- |
|
||||
| OpenAI 管理ツールを使う ( Web 検索、ファイル検索、Code Interpreter、ホスト型 MCP、画像生成 ) | [Hosted tools](#hosted-tools) |
|
||||
| ツール検索で、実行時まで大規模なツール面を遅延させる | [Hosted tool search](#hosted-tool-search) |
|
||||
| 自分のプロセスまたは環境でツールを実行する | [Local runtime tools](#local-runtime-tools) |
|
||||
| Python 関数をツールとしてラップする | [Function tools](#function-tools) |
|
||||
| ハンドオフなしで、あるエージェントから別のエージェントを呼ぶ | [Agents as tools](#agents-as-tools) |
|
||||
| エージェントからワークスペーススコープの Codex タスクを実行する | [Experimental: Codex tool](#experimental-codex-tool) |
|
||||
| OpenAI 管理のツール (Web 検索、ファイル検索、code interpreter、ホスト型 MCP、画像生成) を使用する | [ホスト型ツール](#hosted-tools) |
|
||||
| ツール検索で大規模なツールサーフェスをランタイムまで遅延させる | [ホスト型ツール検索](#hosted-tool-search) |
|
||||
| 自身のプロセスまたは環境でツールを実行する | [ローカルランタイムツール](#local-runtime-tools) |
|
||||
| Python 関数をツールとしてラップする | [関数ツール](#function-tools) |
|
||||
| ハンドオフなしで 1 つのエージェントに別のエージェントを呼び出させる | [Agents as tools](#agents-as-tools) |
|
||||
| エージェントからワークスペーススコープの Codex タスクを実行する | [実験的: Codex ツール](#experimental-codex-tool) |
|
||||
|
||||
## Hosted tools
|
||||
## ホスト型ツール
|
||||
|
||||
[`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel] を使用する場合、 OpenAI はいくつかの組み込みツールを提供しています。
|
||||
OpenAI は、[`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel] を使用する際に、いくつかの組み込みツールを提供しています:
|
||||
|
||||
- [`WebSearchTool`][agents.tool.WebSearchTool] は、エージェントが Web 検索を行えるようにします。
|
||||
- [`FileSearchTool`][agents.tool.FileSearchTool] は、 OpenAI ベクトルストアから情報を取得できるようにします。
|
||||
- [`CodeInterpreterTool`][agents.tool.CodeInterpreterTool] は、 LLM がサンドボックス環境でコードを実行できるようにします。
|
||||
- [`WebSearchTool`][agents.tool.WebSearchTool] により、エージェントは Web 検索を実行できます。
|
||||
- [`FileSearchTool`][agents.tool.FileSearchTool] により、OpenAI ベクトルストアから情報を取得できます。
|
||||
- [`CodeInterpreterTool`][agents.tool.CodeInterpreterTool] により、LLM はサンドボックス環境でコードを実行できます。
|
||||
- [`HostedMCPTool`][agents.tool.HostedMCPTool] は、リモート MCP サーバーのツールをモデルに公開します。
|
||||
- [`ImageGenerationTool`][agents.tool.ImageGenerationTool] は、プロンプトから画像を生成します。
|
||||
- [`ToolSearchTool`][agents.tool.ToolSearchTool] は、モデルが必要に応じて遅延ツール、名前空間、またはホスト MCP サーバーを読み込めるようにします。
|
||||
- [`ToolSearchTool`][agents.tool.ToolSearchTool] により、モデルは遅延読み込み対象のツール、名前空間、またはホスト型 MCP サーバーをオンデマンドで読み込めます。
|
||||
|
||||
高度なホスト検索オプション:
|
||||
高度なホスト型検索オプション:
|
||||
|
||||
- `FileSearchTool` は、`vector_store_ids` と `max_num_results` に加えて、`filters`、`ranking_options`、`include_search_results` をサポートします。
|
||||
- `WebSearchTool` は、`filters`、`user_location`、`search_context_size` をサポートします。
|
||||
- `FileSearchTool` は、`vector_store_ids` と `max_num_results` に加えて、`filters`、`ranking_options`、`include_search_results` をサポートしています。
|
||||
- `WebSearchTool` は、`filters`、`user_location`、`search_context_size` をサポートしています。
|
||||
|
||||
```python
|
||||
from agents import Agent, FileSearchTool, Runner, WebSearchTool
|
||||
@@ -60,11 +60,11 @@ async def main():
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
### Hosted tool search
|
||||
### ホスト型ツール検索
|
||||
|
||||
ツール検索により、 OpenAI Responses モデルは大規模なツール面を実行時まで遅延できるため、モデルは現在のターンに必要なサブセットだけを読み込みます。これは、多数の関数ツール、名前空間グループ、またはホスト MCP サーバーがあり、すべてのツールを事前公開せずにツールスキーマのトークンを削減したい場合に有用です。
|
||||
ツール検索により、OpenAI Responses モデルは大規模なツールサーフェスをランタイムまで遅延できるため、モデルは現在のターンで必要なサブセットのみを読み込みます。これは、多数の関数ツール、名前空間グループ、またはホスト型 MCP サーバーがあり、すべてのツールを事前に公開せずにツールスキーマトークンを削減したい場合に便利です。
|
||||
|
||||
候補ツールがエージェント構築時に既知である場合は、 hosted tool search から開始してください。アプリケーションが動的に読み込む対象を判断する必要がある場合、 Responses API はクライアント実行のツール検索もサポートしますが、標準の `Runner` はそのモードを自動実行しません。
|
||||
エージェントを構築する時点で候補ツールが既知の場合は、ホスト型ツール検索から始めてください。アプリケーションが読み込む内容を動的に決定する必要がある場合、Responses API はクライアント実行型ツール検索もサポートしていますが、標準の `Runner` はそのモードを自動実行しません。
|
||||
|
||||
```python
|
||||
from typing import Annotated
|
||||
@@ -97,7 +97,7 @@ crm_tools = tool_namespace(
|
||||
|
||||
agent = Agent(
|
||||
name="Operations assistant",
|
||||
model="gpt-5.4",
|
||||
model="gpt-5.5",
|
||||
instructions="Load the crm namespace before using CRM tools.",
|
||||
tools=[*crm_tools, ToolSearchTool()],
|
||||
)
|
||||
@@ -106,26 +106,26 @@ result = await Runner.run(agent, "Look up customer_42 and list their open orders
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
知っておくべき点:
|
||||
知っておくべきこと:
|
||||
|
||||
- Hosted tool search は OpenAI Responses モデルでのみ利用可能です。現在の Python SDK サポートは `openai>=2.25.0` に依存します。
|
||||
- エージェントで遅延読み込み面を設定する場合は、`ToolSearchTool()` を正確に 1 つ追加してください。
|
||||
- 検索可能な面には、`@function_tool(defer_loading=True)`、`tool_namespace(name=..., description=..., tools=[...])`、`HostedMCPTool(tool_config={..., "defer_loading": True})` が含まれます。
|
||||
- 遅延読み込み関数ツールは `ToolSearchTool()` と組み合わせる必要があります。名前空間のみの構成でも、モデルが必要時に適切なグループを読み込めるよう `ToolSearchTool()` を使用できます。
|
||||
- `tool_namespace()` は、`FunctionTool` インスタンスを共有の名前空間名と説明の下にグループ化します。これは通常、`crm`、`billing`、`shipping` のように関連ツールが多い場合に最適です。
|
||||
- OpenAI の公式ベストプラクティスガイドは [Use namespaces where possible](https://developers.openai.com/api/docs/guides/tools-tool-search#use-namespaces-where-possible) です。
|
||||
- 可能な場合は、多数の個別遅延関数よりも名前空間またはホスト MCP サーバーを優先してください。通常、モデルにとってより良い高レベル検索面と、より高いトークン削減効果が得られます。
|
||||
- 名前空間には即時ツールと遅延ツールを混在できます。`defer_loading=True` がないツールは即時呼び出し可能なままで、同じ名前空間内の遅延ツールはツール検索経由で読み込まれます。
|
||||
- 目安として、各名前空間は比較的小さく保ち、理想的には 10 関数未満にしてください。
|
||||
- 名前付き `tool_choice` は、裸の名前空間名や遅延専用ツールを対象にできません。`auto`、`required`、または実在するトップレベル呼び出し可能ツール名を優先してください。
|
||||
- `ToolSearchTool(execution="client")` は手動 Responses オーケストレーション用です。モデルがクライアント実行の `tool_search_call` を出力した場合、標準 `Runner` はあなたの代わりに実行せずエラーにします。
|
||||
- ツール検索アクティビティは [`RunResult.new_items`](results.md#new-items) と、専用のアイテム / イベント型を持つ [`RunItemStreamEvent`](streaming.md#run-item-event-names) に表示されます。
|
||||
- 名前空間読み込みとトップレベル遅延ツールの両方を網羅した実行可能な完全例は `examples/tools/tool_search.py` を参照してください。
|
||||
- 公式プラットフォームガイド: [Tool search](https://developers.openai.com/api/docs/guides/tools-tool-search)。
|
||||
- ホスト型ツール検索は、OpenAI Responses モデルでのみ利用できます。現在の Python SDK サポートは `openai>=2.25.0` に依存します。
|
||||
- エージェントで遅延読み込みサーフェスを設定する場合は、`ToolSearchTool()` をちょうど 1 つ追加してください。
|
||||
- 検索可能なサーフェスには、`@function_tool(defer_loading=True)`、`tool_namespace(name=..., description=..., tools=[...])`、`HostedMCPTool(tool_config={..., "defer_loading": True})` が含まれます。
|
||||
- 遅延読み込みの関数ツールは、`ToolSearchTool()` と組み合わせる必要があります。名前空間のみの設定でも、モデルがオンデマンドで適切なグループを読み込めるように `ToolSearchTool()` を使用できます。
|
||||
- `tool_namespace()` は、`FunctionTool` インスタンスを共有の名前空間名と説明の下にグループ化します。これは通常、`crm`、`billing`、`shipping` など、多くの関連ツールがある場合に最適です。
|
||||
- OpenAI の公式ベストプラクティスガイダンスは [可能な場合は名前空間を使用する](https://developers.openai.com/api/docs/guides/tools-tool-search#use-namespaces-where-possible) です。
|
||||
- 可能な場合は、個別に遅延読み込みする多数の関数よりも、名前空間またはホスト型 MCP サーバーを優先してください。通常、モデルに対してより優れた高レベルの検索サーフェスと、より大きなトークン削減効果を提供します。
|
||||
- 名前空間には、即時ツールと遅延読み込みツールを混在させることができます。`defer_loading=True` のないツールはすぐに呼び出し可能なままであり、同じ名前空間内の遅延読み込みツールはツール検索を通じて読み込まれます。
|
||||
- 目安として、各名前空間はかなり小さく保ち、理想的には 10 個未満の関数にしてください。
|
||||
- 名前付き `tool_choice` は、単独の名前空間名や遅延読み込み専用ツールをターゲットにできません。`auto`、`required`、または実在するトップレベルの呼び出し可能なツール名を優先してください。
|
||||
- `ToolSearchTool(execution="client")` は手動の Responses オーケストレーション用です。モデルがクライアント実行型の `tool_search_call` を出力した場合、標準の `Runner` はそれを実行せずに例外を送出します。
|
||||
- ツール検索アクティビティは、[`RunResult.new_items`](results.md#new-items) と [`RunItemStreamEvent`](streaming.md#run-item-event-names) に、専用の item および event タイプとともに表示されます。
|
||||
- 名前空間付き読み込みとトップレベルの遅延読み込みツールの両方を扱う、完全に実行可能なコード例については、`examples/tools/tool_search.py` を参照してください。
|
||||
- 公式プラットフォームガイド: [ツール検索](https://developers.openai.com/api/docs/guides/tools-tool-search)。
|
||||
|
||||
### ホストコンテナ shell + skills
|
||||
### ホスト型コンテナーシェル + スキル
|
||||
|
||||
`ShellTool` は OpenAI ホストコンテナ実行もサポートします。モデルにローカルランタイムではなく管理コンテナで shell コマンドを実行させたい場合は、このモードを使用してください。
|
||||
`ShellTool` は、OpenAI がホストするコンテナー実行にも対応しています。ローカルランタイムではなく管理対象コンテナーでモデルにシェルコマンドを実行させたい場合は、このモードを使用してください。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, ShellTool, ShellToolSkillReference
|
||||
@@ -138,7 +138,7 @@ csv_skill: ShellToolSkillReference = {
|
||||
|
||||
agent = Agent(
|
||||
name="Container shell agent",
|
||||
model="gpt-5.4",
|
||||
model="gpt-5.5",
|
||||
instructions="Use the mounted skill when helpful.",
|
||||
tools=[
|
||||
ShellTool(
|
||||
@@ -158,52 +158,52 @@ result = await Runner.run(
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
後続の run で既存コンテナを再利用するには、`environment={"type": "container_reference", "container_id": "cntr_..."}` を設定します。
|
||||
後続の実行で既存のコンテナーを再利用するには、`environment={"type": "container_reference", "container_id": "cntr_..."}` を設定します。
|
||||
|
||||
知っておくべき点:
|
||||
知っておくべきこと:
|
||||
|
||||
- ホスト shell は Responses API の shell ツール経由で利用可能です。
|
||||
- `container_auto` はリクエスト用にコンテナをプロビジョニングし、`container_reference` は既存コンテナを再利用します。
|
||||
- ホスト型シェルは、Responses API のシェルツールを通じて利用できます。
|
||||
- `container_auto` はリクエスト用のコンテナーをプロビジョニングします。`container_reference` は既存のコンテナーを再利用します。
|
||||
- `container_auto` には `file_ids` と `memory_limit` も含められます。
|
||||
- `environment.skills` は skill 参照とインライン skill バンドルを受け付けます。
|
||||
- ホスト環境では、`ShellTool` に `executor`、`needs_approval`、`on_approval` を設定しないでください。
|
||||
- `network_policy` は `disabled` と `allowlist` モードをサポートします。
|
||||
- allowlist モードでは、`network_policy.domain_secrets` でドメインスコープのシークレットを名前で注入できます。
|
||||
- 完全な例は `examples/tools/container_shell_skill_reference.py` と `examples/tools/container_shell_inline_skill.py` を参照してください。
|
||||
- OpenAI プラットフォームガイド: [Shell](https://platform.openai.com/docs/guides/tools-shell) と [Skills](https://platform.openai.com/docs/guides/tools-skills)。
|
||||
- `environment.skills` は、スキル参照とインラインスキルバンドルを受け付けます。
|
||||
- ホスト型環境では、`ShellTool` に `executor`、`needs_approval`、`on_approval` を設定しないでください。
|
||||
- `network_policy` は `disabled` と `allowlist` モードをサポートしています。
|
||||
- allowlist モードでは、`network_policy.domain_secrets` により、名前でドメインスコープのシークレットを注入できます。
|
||||
- 完全なコード例については、`examples/tools/container_shell_skill_reference.py` と `examples/tools/container_shell_inline_skill.py` を参照してください。
|
||||
- OpenAI プラットフォームガイド: [シェル](https://platform.openai.com/docs/guides/tools-shell) と [スキル](https://platform.openai.com/docs/guides/tools-skills)。
|
||||
|
||||
## ローカルランタイムツール
|
||||
|
||||
ローカルランタイムツールは、モデル応答自体の外側で実行されます。モデルはいつ呼び出すかを決定しますが、実際の処理はアプリケーションまたは設定済み実行環境が行います。
|
||||
ローカルランタイムツールは、モデルレスポンス自体の外部で実行されます。モデルは引き続きいつ呼び出すかを決定しますが、アプリケーションまたは設定された実行環境が実際の作業を行います。
|
||||
|
||||
`ComputerTool` と `ApplyPatchTool` は常に、あなたが提供するローカル実装を必要とします。`ShellTool` は両モードにまたがります。管理実行が必要なら上記ホストコンテナ構成を使い、自分のプロセスでコマンドを実行したいなら以下のローカルランタイム構成を使ってください。
|
||||
`ComputerTool` と `ApplyPatchTool` には、常にユーザーが提供するローカル実装が必要です。`ShellTool` は両方のモードにまたがります。管理対象実行を行いたい場合は上記のホスト型コンテナー設定を使用し、自身のプロセスでコマンドを実行したい場合は以下のローカルランタイム設定を使用してください。
|
||||
|
||||
ローカルランタイムツールでは実装の提供が必要です:
|
||||
ローカルランタイムツールでは、実装を提供する必要があります:
|
||||
|
||||
- [`ComputerTool`][agents.tool.ComputerTool]: GUI / ブラウザ自動化を有効にするには [`Computer`][agents.computer.Computer] または [`AsyncComputer`][agents.computer.AsyncComputer] インターフェースを実装します。
|
||||
- [`ShellTool`][agents.tool.ShellTool]: ローカル実行とホストコンテナ実行の両方に対応する最新 shell ツールです。
|
||||
- [`LocalShellTool`][agents.tool.LocalShellTool]: レガシーのローカル shell 統合です。
|
||||
- [`ApplyPatchTool`][agents.tool.ApplyPatchTool]: 差分をローカル適用するには [`ApplyPatchEditor`][agents.editor.ApplyPatchEditor] を実装します。
|
||||
- ローカル shell skills は `ShellTool(environment={"type": "local", "skills": [...]})` で利用できます。
|
||||
- [`ComputerTool`][agents.tool.ComputerTool]: GUI / ブラウザー自動化を有効にするために、[`Computer`][agents.computer.Computer] または [`AsyncComputer`][agents.computer.AsyncComputer] インターフェイスを実装します。
|
||||
- [`ShellTool`][agents.tool.ShellTool]: ローカル実行とホスト型コンテナー実行の両方に対応する最新のシェルツールです。
|
||||
- [`LocalShellTool`][agents.tool.LocalShellTool]: レガシーなローカルシェル統合です。
|
||||
- [`ApplyPatchTool`][agents.tool.ApplyPatchTool]: diff をローカルに適用するために、[`ApplyPatchEditor`][agents.editor.ApplyPatchEditor] を実装します。
|
||||
- ローカルシェルスキルは、`ShellTool(environment={"type": "local", "skills": [...]})` で利用できます。
|
||||
|
||||
### ComputerTool と Responses computer tool
|
||||
### ComputerTool と Responses コンピュータツール
|
||||
|
||||
`ComputerTool` は依然としてローカルハーネスです。あなたが [`Computer`][agents.computer.Computer] または [`AsyncComputer`][agents.computer.AsyncComputer] 実装を提供し、 SDK がそのハーネスを OpenAI Responses API の computer 面にマッピングします。
|
||||
`ComputerTool` は引き続きローカルハーネスです。ユーザーが [`Computer`][agents.computer.Computer] または [`AsyncComputer`][agents.computer.AsyncComputer] 実装を提供し、SDK はそのハーネスを OpenAI Responses API のコンピュータサーフェスにマッピングします。
|
||||
|
||||
明示的な [`gpt-5.4`](https://developers.openai.com/api/docs/models/gpt-5.4) リクエストでは、 SDK は GA 組み込みツールペイロード `{"type": "computer"}` を送信します。古い `computer-use-preview` モデルでは、プレビュー用ペイロード `{"type": "computer_use_preview", "environment": ..., "display_width": ..., "display_height": ...}` を維持します。これは OpenAI の [Computer use guide](https://developers.openai.com/api/docs/guides/tools-computer-use/) で説明されているプラットフォーム移行を反映しています。
|
||||
明示的な [`gpt-5.5`](https://developers.openai.com/api/docs/models/gpt-5.5) リクエストでは、SDK は GA 組み込みツールペイロード `{"type": "computer"}` を送信します。古い `computer-use-preview` モデルは、プレビューペイロード `{"type": "computer_use_preview", "environment": ..., "display_width": ..., "display_height": ...}` を維持します。これは、OpenAI の [コンピュータ操作ガイド](https://developers.openai.com/api/docs/guides/tools-computer-use/) で説明されているプラットフォーム移行を反映しています:
|
||||
|
||||
- モデル: `computer-use-preview` -> `gpt-5.4`
|
||||
- モデル: `computer-use-preview` -> `gpt-5.5`
|
||||
- ツールセレクター: `computer_use_preview` -> `computer`
|
||||
- Computer 呼び出し形状: `computer_call` あたり 1 つの `action` -> `computer_call` 上のバッチ `actions[]`
|
||||
- Truncation: プレビューパスでは `ModelSettings(truncation="auto")` が必須 -> GA パスでは不要
|
||||
- コンピュータ呼び出しの形状: `computer_call` ごとに 1 つの `action` -> `computer_call` 上のバッチ化された `actions[]`
|
||||
- 切り詰め: プレビュー経路では `ModelSettings(truncation="auto")` が必要 -> GA 経路では不要
|
||||
|
||||
SDK は、実際の Responses リクエスト上の有効モデルから wire 形状を選択します。プロンプトテンプレートを使い、プロンプト側が `model` を所有するためリクエストに `model` がない場合、`model="gpt-5.4"` を明示するか、`ModelSettings(tool_choice="computer")` または `ModelSettings(tool_choice="computer_use")` で GA セレクターを強制しない限り、 SDK はプレビュー互換 computer ペイロードを維持します。
|
||||
SDK は、実際の Responses リクエストで有効なモデルに基づいて、そのワイヤ形式を選択します。プロンプトテンプレートを使用していて、プロンプト側がモデルを保持しているためにリクエストで `model` が省略される場合、`model="gpt-5.5"` を明示的に保持するか、`ModelSettings(tool_choice="computer")` または `ModelSettings(tool_choice="computer_use")` で GA セレクターを強制しない限り、SDK はプレビュー互換のコンピュータペイロードを維持します。
|
||||
|
||||
[`ComputerTool`][agents.tool.ComputerTool] が存在する場合、`tool_choice="computer"`、`"computer_use"`、`"computer_use_preview"` はすべて受け入れられ、有効リクエストモデルに一致する組み込みセレクターへ正規化されます。`ComputerTool` がない場合、これらの文字列は通常の関数名として動作します。
|
||||
[`ComputerTool`][agents.tool.ComputerTool] が存在する場合、`tool_choice="computer"`、`"computer_use"`、`"computer_use_preview"` はすべて受け付けられ、有効なリクエストモデルに一致する組み込みセレクターに正規化されます。`ComputerTool` がない場合、これらの文字列は引き続き通常の関数名のように動作します。
|
||||
|
||||
この違いは、`ComputerTool` が [`ComputerProvider`][agents.tool.ComputerProvider] ファクトリーに支えられている場合に重要です。GA の `computer` ペイロードはシリアライズ時に `environment` や寸法を必要としないため、未解決ファクトリーでも問題ありません。プレビュー互換シリアライズでは、 SDK が `environment`、`display_width`、`display_height` を送るため、解決済みの `Computer` または `AsyncComputer` インスタンスが依然必要です。
|
||||
この違いは、`ComputerTool` が [`ComputerProvider`][agents.tool.ComputerProvider] ファクトリによって支えられている場合に重要です。GA の `computer` ペイロードはシリアライズ時に `environment` や寸法を必要としないため、未解決のファクトリでも問題ありません。プレビュー互換のシリアライズでは、SDK が `environment`、`display_width`、`display_height` を送信できるように、解決済みの `Computer` または `AsyncComputer` インスタンスが引き続き必要です。
|
||||
|
||||
実行時は、どちらのパスも同じローカルハーネスを使います。プレビュー応答は単一 `action` の `computer_call` アイテムを出力し、`gpt-5.4` はバッチ `actions[]` を出力でき、 SDK は `computer_call_output` スクリーンショットアイテムを生成する前に順番に実行します。実行可能な Playwright ベースのハーネスは `examples/tools/computer_use.py` を参照してください。
|
||||
ランタイムでは、どちらの経路も同じローカルハーネスを使用します。プレビューレスポンスは単一の `action` を持つ `computer_call` 項目を出力します。`gpt-5.5` はバッチ化された `actions[]` を出力でき、SDK は `computer_call_output` スクリーンショット項目を生成する前に、それらを順番に実行します。実行可能な Playwright ベースのハーネスについては、`examples/tools/computer_use.py` を参照してください。
|
||||
|
||||
```python
|
||||
from agents import Agent, ApplyPatchTool, ShellTool
|
||||
@@ -247,16 +247,16 @@ agent = Agent(
|
||||
|
||||
## 関数ツール
|
||||
|
||||
任意の Python 関数をツールとして使えます。 Agents SDK が自動的にツールを設定します。
|
||||
任意の Python 関数をツールとして使用できます。Agents SDK がツールを自動的に設定します:
|
||||
|
||||
- ツール名は Python 関数名になります (または名前を提供できます)
|
||||
- ツール説明は関数の docstring から取得されます (または説明を提供できます)
|
||||
- 関数入力のスキーマは、関数引数から自動生成されます
|
||||
- 各入力の説明は、無効化しない限り関数の docstring から取得されます
|
||||
- ツール名は Python 関数の名前になります (または名前を指定できます)
|
||||
- ツールの説明は関数の docstring から取得されます (または説明を指定できます)
|
||||
- 関数入力のスキーマは、関数の引数から自動的に作成されます
|
||||
- 各入力の説明は、無効化されていない限り、関数の docstring から取得されます
|
||||
|
||||
関数シグネチャ抽出には Python の `inspect` モジュールを使用し、docstring 解析には [`griffe`](https://mkdocstrings.github.io/griffe/)、スキーマ作成には `pydantic` を使用します。
|
||||
関数シグネチャを抽出するために Python の `inspect` モジュールを使用し、docstring の解析に [`griffe`](https://mkdocstrings.github.io/griffe/) を、スキーマ作成に `pydantic` を併用します。
|
||||
|
||||
OpenAI Responses モデルを使用している場合、`@function_tool(defer_loading=True)` は `ToolSearchTool()` が読み込むまで関数ツールを非表示にします。[`tool_namespace()`][agents.tool.tool_namespace] で関連関数ツールをグループ化することもできます。完全な設定と制約は [Hosted tool search](#hosted-tool-search) を参照してください。
|
||||
OpenAI Responses モデルを使用している場合、`@function_tool(defer_loading=True)` は `ToolSearchTool()` が読み込むまで関数ツールを非表示にします。関連する関数ツールを [`tool_namespace()`][agents.tool.tool_namespace] でグループ化することもできます。完全な設定と制約については、[ホスト型ツール検索](#hosted-tool-search) を参照してください。
|
||||
|
||||
```python
|
||||
import json
|
||||
@@ -308,12 +308,12 @@ for tool in agent.tools:
|
||||
|
||||
```
|
||||
|
||||
1. 関数引数には任意の Python 型を使用でき、関数は sync / async どちらでも構いません。
|
||||
2. docstring がある場合、説明と引数説明の取得に使用されます。
|
||||
3. 関数は任意で `context` を受け取れます (最初の引数である必要があります)。ツール名、説明、使用する docstring スタイルなどのオーバーライドも設定できます。
|
||||
4. デコレートした関数をツールリストに渡せます。
|
||||
1. 関数の引数には任意の Python 型を使用でき、関数は同期でも非同期でも構いません。
|
||||
2. docstring が存在する場合、説明と引数の説明を取得するために使用されます
|
||||
3. 関数は必要に応じて `context` を受け取ることができます (最初の引数である必要があります)。ツール名、説明、使用する docstring スタイルなどのオーバーライドも設定できます。
|
||||
4. デコレートされた関数をツールのリストに渡すことができます。
|
||||
|
||||
??? note "出力を表示"
|
||||
??? note "出力を表示するには展開"
|
||||
|
||||
```
|
||||
fetch_weather
|
||||
@@ -385,20 +385,20 @@ for tool in agent.tools:
|
||||
|
||||
### 関数ツールからの画像またはファイルの返却
|
||||
|
||||
テキスト出力の返却に加えて、関数ツールの出力として 1 つ以上の画像またはファイルを返せます。そのためには、次のいずれかを返します。
|
||||
テキスト出力を返すことに加えて、関数ツールの出力として 1 つまたは複数の画像やファイルを返すことができます。そのためには、次のいずれかを返すことができます:
|
||||
|
||||
- 画像: [`ToolOutputImage`][agents.tool.ToolOutputImage] (または TypedDict 版の [`ToolOutputImageDict`][agents.tool.ToolOutputImageDict])
|
||||
- ファイル: [`ToolOutputFileContent`][agents.tool.ToolOutputFileContent] (または TypedDict 版の [`ToolOutputFileContentDict`][agents.tool.ToolOutputFileContentDict])
|
||||
- テキスト: 文字列、文字列化可能オブジェクト、または [`ToolOutputText`][agents.tool.ToolOutputText] (または TypedDict 版の [`ToolOutputTextDict`][agents.tool.ToolOutputTextDict])
|
||||
- 画像: [`ToolOutputImage`][agents.tool.ToolOutputImage] (または TypedDict バージョンの [`ToolOutputImageDict`][agents.tool.ToolOutputImageDict])
|
||||
- ファイル: [`ToolOutputFileContent`][agents.tool.ToolOutputFileContent] (または TypedDict バージョンの [`ToolOutputFileContentDict`][agents.tool.ToolOutputFileContentDict])
|
||||
- テキスト: 文字列または文字列化可能なオブジェクト、または [`ToolOutputText`][agents.tool.ToolOutputText] (または TypedDict バージョンの [`ToolOutputTextDict`][agents.tool.ToolOutputTextDict])
|
||||
|
||||
### カスタム関数ツール
|
||||
|
||||
場合によっては、 Python 関数をツールとして使いたくないことがあります。その場合は、必要に応じて [`FunctionTool`][agents.tool.FunctionTool] を直接作成できます。必要なものは次のとおりです。
|
||||
場合によっては、Python 関数をツールとして使用したくないことがあります。その場合は、必要に応じて [`FunctionTool`][agents.tool.FunctionTool] を直接作成できます。次を提供する必要があります:
|
||||
|
||||
- `name`
|
||||
- `description`
|
||||
- `params_json_schema` (引数の JSON スキーマ)
|
||||
- `on_invoke_tool` ( [`ToolContext`][agents.tool_context.ToolContext] と JSON 文字列としての引数を受け取り、ツール出力 (たとえばテキスト、構造化ツール出力オブジェクト、または出力リスト) を返す async 関数)
|
||||
- `params_json_schema`: 引数用の JSON スキーマです
|
||||
- `on_invoke_tool`: [`ToolContext`][agents.tool_context.ToolContext] と JSON 文字列としての引数を受け取り、ツール出力 (例えば、テキスト、構造化されたツール出力オブジェクト、または出力のリスト) を返す async 関数です。
|
||||
|
||||
```python
|
||||
from typing import Any
|
||||
@@ -433,16 +433,16 @@ tool = FunctionTool(
|
||||
|
||||
### 引数と docstring の自動解析
|
||||
|
||||
前述のとおり、ツール用スキーマ抽出のために関数シグネチャを自動解析し、ツール説明と個別引数説明抽出のために docstring を解析します。注意点は次のとおりです。
|
||||
前述のとおり、ツールのスキーマを抽出するために関数シグネチャを自動的に解析し、ツールと個々の引数の説明を抽出するために docstring を解析します。いくつか注意点があります:
|
||||
|
||||
1. シグネチャ解析は `inspect` モジュールで行います。引数型の理解には型アノテーションを使い、全体スキーマを表す Pydantic モデルを動的に構築します。 Python プリミティブ、Pydantic モデル、TypedDict などを含む、ほとんどの型をサポートします。
|
||||
2. docstring 解析には `griffe` を使用します。サポートされる docstring 形式は `google`、`sphinx`、`numpy` です。docstring 形式は自動検出を試みますが、これはベストエフォートであり、`function_tool` 呼び出し時に明示設定できます。`use_docstring_info` を `False` に設定して docstring 解析を無効化することもできます。
|
||||
1. シグネチャ解析は `inspect` モジュールを通じて行われます。型アノテーションを使用して引数の型を理解し、全体のスキーマを表す Pydantic モデルを動的に構築します。Python の基本型、Pydantic モデル、TypedDict など、ほとんどの型をサポートしています。
|
||||
2. docstring の解析には `griffe` を使用します。サポートされる docstring 形式は `google`、`sphinx`、`numpy` です。docstring 形式の自動検出を試みますが、これはベストエフォートであり、`function_tool` を呼び出す際に明示的に設定できます。`use_docstring_info` を `False` に設定して、docstring 解析を無効にすることもできます。
|
||||
|
||||
スキーマ抽出コードは [`agents.function_schema`][] にあります。
|
||||
スキーマ抽出のコードは [`agents.function_schema`][] にあります。
|
||||
|
||||
### Pydantic Field による引数制約と説明
|
||||
### Pydantic Field による引数の制約と説明
|
||||
|
||||
Pydantic の [`Field`](https://docs.pydantic.dev/latest/concepts/fields/) を使うと、ツール引数に制約 (例: 数値の最小 / 最大、文字列の長さやパターン) と説明を追加できます。Pydantic と同様に、デフォルトベース (`arg: int = Field(..., ge=1)`) と `Annotated` (`arg: Annotated[int, Field(..., ge=1)]`) の両形式をサポートします。生成される JSON スキーマとバリデーションには、これらの制約が含まれます。
|
||||
Pydantic の [`Field`](https://docs.pydantic.dev/latest/concepts/fields/) を使用して、制約 (例: 数値の min / max、文字列の長さやパターン) と説明をツール引数に追加できます。Pydantic と同様に、デフォルトベース (`arg: int = Field(..., ge=1)`) と `Annotated` (`arg: Annotated[int, Field(..., ge=1)]`) の両方の形式がサポートされています。生成される JSON スキーマとバリデーションには、これらの制約が含まれます。
|
||||
|
||||
```python
|
||||
from typing import Annotated
|
||||
@@ -482,13 +482,13 @@ agent = Agent(
|
||||
)
|
||||
```
|
||||
|
||||
タイムアウトに達した場合、デフォルト動作は `timeout_behavior="error_as_result"` で、モデル可視のタイムアウトメッセージ (例: `Tool 'slow_lookup' timed out after 2 seconds.`) を送信します。
|
||||
タイムアウトに達した場合、デフォルトの動作は `timeout_behavior="error_as_result"` で、モデルから見えるタイムアウトメッセージ (例えば、`Tool 'slow_lookup' timed out after 2 seconds.`) を送信します。
|
||||
|
||||
タイムアウト処理は次のように制御できます。
|
||||
タイムアウト処理を制御できます:
|
||||
|
||||
- `timeout_behavior="error_as_result"` (デフォルト): タイムアウトメッセージをモデルに返し、復旧できるようにします。
|
||||
- `timeout_behavior="raise_exception"`: [`ToolTimeoutError`][agents.exceptions.ToolTimeoutError] を発生させ、 run を失敗させます。
|
||||
- `timeout_error_function=...`: `error_as_result` 使用時のタイムアウトメッセージをカスタマイズします。
|
||||
- `timeout_behavior="error_as_result"` (デフォルト): タイムアウトメッセージをモデルに返し、モデルが回復できるようにします。
|
||||
- `timeout_behavior="raise_exception"`: [`ToolTimeoutError`][agents.exceptions.ToolTimeoutError] を送出し、実行を失敗させます。
|
||||
- `timeout_error_function=...`: `error_as_result` を使用する場合のタイムアウトメッセージをカスタマイズします。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -511,15 +511,15 @@ except ToolTimeoutError as e:
|
||||
|
||||
!!! note
|
||||
|
||||
タイムアウト設定は async `@function_tool` ハンドラーでのみサポートされます。
|
||||
タイムアウト設定は、async `@function_tool` ハンドラーでのみサポートされています。
|
||||
|
||||
### 関数ツールでのエラー処理
|
||||
|
||||
`@function_tool` で関数ツールを作成する際、`failure_error_function` を渡せます。これは、ツール呼び出しがクラッシュしたときに LLM へ返すエラー応答を提供する関数です。
|
||||
`@function_tool` を通じて関数ツールを作成する場合、`failure_error_function` を渡すことができます。これは、ツール呼び出しがクラッシュした場合に LLM へエラーレスポンスを提供する関数です。
|
||||
|
||||
- デフォルト (何も渡さない場合) では、エラー発生を LLM に伝える `default_tool_error_function` が実行されます。
|
||||
- 独自のエラー関数を渡すと、代わりにそれが実行され、その応答が LLM に送られます。
|
||||
- 明示的に `None` を渡すと、ツール呼び出しエラーはあなたが処理できるよう再送出されます。これはモデルが無効 JSON を生成した場合の `ModelBehaviorError` や、コードがクラッシュした場合の `UserError` などです。
|
||||
- デフォルトでは (つまり、何も渡さない場合)、`default_tool_error_function` が実行され、エラーが発生したことを LLM に伝えます。
|
||||
- 独自のエラー関数を渡した場合は、それが代わりに実行され、そのレスポンスが LLM に送信されます。
|
||||
- 明示的に `None` を渡した場合、ツール呼び出しエラーは再送出され、ユーザー側で処理できます。これは、モデルが無効な JSON を生成した場合の `ModelBehaviorError` や、コードがクラッシュした場合の `UserError` などである可能性があります。
|
||||
|
||||
```python
|
||||
from agents import function_tool, RunContextWrapper
|
||||
@@ -542,11 +542,11 @@ def get_user_profile(user_id: str) -> str:
|
||||
|
||||
```
|
||||
|
||||
`FunctionTool` オブジェクトを手動作成する場合は、`on_invoke_tool` 関数内でエラーを処理する必要があります。
|
||||
`FunctionTool` オブジェクトを手動で作成している場合は、`on_invoke_tool` 関数内でエラーを処理する必要があります。
|
||||
|
||||
## Agents as tools
|
||||
|
||||
一部のワークフローでは、制御をハンドオフする代わりに、中央エージェントで専門エージェントのネットワークをエージェントオーケストレーションしたい場合があります。これは、エージェントをツールとしてモデル化することで実現できます。
|
||||
一部のワークフローでは、制御をハンドオフするのではなく、中央エージェントに専門エージェントのネットワークをオーケストレーションさせたい場合があります。これは、エージェントを Agents as tools としてモデル化することで実現できます。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
@@ -565,7 +565,7 @@ french_agent = Agent(
|
||||
orchestrator_agent = Agent(
|
||||
name="orchestrator_agent",
|
||||
instructions=(
|
||||
"You are a translation agent. You use the tools given to you to translate."
|
||||
"You are a translation agent. You use the tools given to you to translate. "
|
||||
"If asked for multiple translations, you call the relevant tools."
|
||||
),
|
||||
tools=[
|
||||
@@ -587,7 +587,7 @@ async def main():
|
||||
|
||||
### ツールエージェントのカスタマイズ
|
||||
|
||||
`agent.as_tool` 関数は、エージェントをツールに変換しやすくするための便利メソッドです。`max_turns`、`run_config`、`hooks`、`previous_response_id`、`conversation_id`、`session`、`needs_approval` などの一般的なランタイムオプションをサポートします。さらに、`parameters`、`input_builder`、`include_input_schema` による構造化入力もサポートします。高度なオーケストレーション (例: 条件付きリトライ、フォールバック動作、複数エージェント呼び出しの連鎖) では、ツール実装内で `Runner.run` を直接使用してください。
|
||||
`agent.as_tool` 関数は、エージェントをツールに簡単に変換するための便利なメソッドです。`max_turns`、`run_config`、`hooks`、`previous_response_id`、`conversation_id`、`session`、`needs_approval` などの一般的なランタイムオプションをサポートしています。また、`parameters`、`input_builder`、`include_input_schema` による構造化入力もサポートしています。高度なオーケストレーション (例えば、条件付きリトライ、フォールバック動作、複数のエージェント呼び出しのチェーン) には、ツール実装内で `Runner.run` を直接使用してください:
|
||||
|
||||
```python
|
||||
@function_tool
|
||||
@@ -608,13 +608,13 @@ async def run_my_agent() -> str:
|
||||
|
||||
### ツールエージェントの構造化入力
|
||||
|
||||
デフォルトでは、`Agent.as_tool()` は単一文字列入力 (`{"input": "..."}`) を想定しますが、`parameters` (Pydantic モデルまたは dataclass 型) を渡すことで構造化スキーマを公開できます。
|
||||
デフォルトでは、`Agent.as_tool()` は単一の文字列入力 (`{"input": "..."}`) を想定していますが、`parameters` (Pydantic モデルまたは dataclass 型) を渡すことで構造化スキーマを公開できます。
|
||||
|
||||
追加オプション:
|
||||
|
||||
- `include_input_schema=True` は、生成されるネスト入力に完全な JSON Schema を含めます。
|
||||
- `input_builder=...` は、構造化ツール引数をネストエージェント入力に変換する方法を完全にカスタマイズできます。
|
||||
- `RunContextWrapper.tool_input` は、ネスト run コンテキスト内に解析済み構造化ペイロードを保持します。
|
||||
- `include_input_schema=True` は、生成されるネストされた入力に完全な JSON Schema を含めます。
|
||||
- `input_builder=...` により、構造化ツール引数をネストされたエージェント入力に変換する方法を完全にカスタマイズできます。
|
||||
- `RunContextWrapper.tool_input` には、ネストされた実行コンテキスト内の解析済み構造化ペイロードが含まれます。
|
||||
|
||||
```python
|
||||
from pydantic import BaseModel, Field
|
||||
@@ -634,21 +634,21 @@ translator_tool = translator_agent.as_tool(
|
||||
)
|
||||
```
|
||||
|
||||
完全に実行可能な例は `examples/agent_patterns/agents_as_tools_structured.py` を参照してください。
|
||||
完全に実行可能なコード例については、`examples/agent_patterns/agents_as_tools_structured.py` を参照してください。
|
||||
|
||||
### ツールエージェントの承認ゲート
|
||||
|
||||
`Agent.as_tool(..., needs_approval=...)` は `function_tool` と同じ承認フローを使用します。承認が必要な場合、 run は一時停止し、保留中アイテムは `result.interruptions` に表示されます。次に `result.to_state()` を使用し、`state.approve(...)` または `state.reject(...)` 呼び出し後に再開します。完全な一時停止 / 再開パターンは [Human-in-the-loop guide](human_in_the_loop.md) を参照してください。
|
||||
`Agent.as_tool(..., needs_approval=...)` は、`function_tool` と同じ承認フローを使用します。承認が必要な場合、実行は一時停止し、保留中の項目が `result.interruptions` に表示されます。その後、`result.to_state()` を使用し、`state.approve(...)` または `state.reject(...)` を呼び出した後に再開します。完全な一時停止 / 再開パターンについては、[Human-in-the-loop ガイド](human_in_the_loop.md) を参照してください。
|
||||
|
||||
### カスタム出力抽出
|
||||
|
||||
特定のケースでは、中央エージェントに返す前にツールエージェントの出力を変更したいことがあります。これは次のような場合に有用です。
|
||||
特定のケースでは、中央エージェントに返す前にツールエージェントの出力を変更したい場合があります。これは、次のような場合に便利です:
|
||||
|
||||
- サブエージェントのチャット履歴から特定情報 (例: JSON ペイロード) を抽出する。
|
||||
- エージェントの最終回答を変換または再整形する (例: Markdown をプレーンテキストや CSV に変換)。
|
||||
- 出力を検証する、またはエージェント応答が欠落 / 不正形式の場合にフォールバック値を提供する。
|
||||
- サブエージェントのチャット履歴から特定の情報 (例: JSON ペイロード) を抽出する。
|
||||
- エージェントの最終回答を変換または再フォーマットする (例: Markdown をプレーンテキストまたは CSV に変換する)。
|
||||
- 出力を検証する、またはエージェントのレスポンスが欠落しているか不正な形式の場合にフォールバック値を提供する。
|
||||
|
||||
これは、`as_tool` メソッドに `custom_output_extractor` 引数を渡すことで実現できます。
|
||||
これは、`as_tool` メソッドに `custom_output_extractor` 引数を指定することで実現できます:
|
||||
|
||||
```python
|
||||
async def extract_json_payload(run_result: RunResult) -> str:
|
||||
@@ -667,14 +667,11 @@ json_tool = data_agent.as_tool(
|
||||
)
|
||||
```
|
||||
|
||||
カスタム抽出器内では、ネストされた [`RunResult`][agents.result.RunResult] は
|
||||
[`agent_tool_invocation`][agents.result.RunResultBase.agent_tool_invocation] も公開します。これは
|
||||
ネスト結果の後処理中に、外側ツール名、呼び出し ID、または raw 引数が必要な場合に有用です。
|
||||
[Results guide](results.md#agent-as-tool-metadata) も参照してください。
|
||||
カスタム抽出器内では、ネストされた [`RunResult`][agents.result.RunResult] から [`agent_tool_invocation`][agents.result.RunResultBase.agent_tool_invocation] も参照できます。これは、ネストされた実行結果を後処理する際に、外側のツール名、呼び出し ID、または raw 引数が必要な場合に便利です。[実行結果ガイド](results.md#agent-as-tool-metadata) を参照してください。
|
||||
|
||||
### ネストされたエージェント run のストリーミング
|
||||
### ネストされたエージェント実行のストリーミング
|
||||
|
||||
`as_tool` に `on_stream` コールバックを渡すと、ストリーム完了後に最終出力を返しつつ、ネストエージェントが出力するストリーミングイベントを監視できます。
|
||||
`on_stream` コールバックを `as_tool` に渡すと、ネストされたエージェントが出力するストリーミングイベントをリッスンしつつ、ストリーム完了後にその最終出力を返せます。
|
||||
|
||||
```python
|
||||
from agents import AgentToolStreamEvent
|
||||
@@ -692,17 +689,17 @@ billing_agent_tool = billing_agent.as_tool(
|
||||
)
|
||||
```
|
||||
|
||||
想定される挙動:
|
||||
想定されること:
|
||||
|
||||
- イベント型は `StreamEvent["type"]` を反映します: `raw_response_event`、`run_item_stream_event`、`agent_updated_stream_event`。
|
||||
- `on_stream` を提供すると、ネストエージェントは自動的にストリーミングモードで実行され、最終出力返却前にストリームがドレインされます。
|
||||
- ハンドラーは同期または非同期にでき、各イベントは到着順で配信されます。
|
||||
- `tool_call` は、モデルのツール呼び出し経由でツールが呼ばれた場合に存在します。直接呼び出しでは `None` のままの場合があります。
|
||||
- 完全に実行可能なサンプルは `examples/agent_patterns/agents_as_tools_streaming.py` を参照してください。
|
||||
- イベントタイプは `StreamEvent["type"]` を反映します: `raw_response_event`、`run_item_stream_event`、`agent_updated_stream_event`。
|
||||
- `on_stream` を指定すると、ネストされたエージェントが自動的にストリーミングモードで実行され、最終出力を返す前にストリームが drain されます。
|
||||
- ハンドラーは同期でも非同期でも構いません。各イベントは到着順に配信されます。
|
||||
- `tool_call` は、モデルのツール呼び出しを通じてツールが呼び出された場合に存在します。直接呼び出しでは `None` のままになることがあります。
|
||||
- 完全に実行可能なサンプルについては、`examples/agent_patterns/agents_as_tools_streaming.py` を参照してください。
|
||||
|
||||
### 条件付きツール有効化
|
||||
|
||||
`is_enabled` パラメーターを使うと、実行時にエージェントツールを条件付きで有効 / 無効にできます。これにより、コンテキスト、ユーザー設定、またはランタイム条件に基づいて、 LLM が利用可能なツールを動的にフィルタリングできます。
|
||||
`is_enabled` パラメーターを使用して、ランタイムでエージェントツールを条件付きで有効化または無効化できます。これにより、コンテキスト、ユーザー設定、またはランタイム条件に基づいて、LLM で利用できるツールを動的にフィルタリングできます。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -757,24 +754,24 @@ async def main():
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
`is_enabled` パラメーターは次を受け付けます。
|
||||
`is_enabled` パラメーターは次を受け付けます:
|
||||
|
||||
- **ブール値**: `True` (常に有効) または `False` (常に無効)
|
||||
- **呼び出し可能関数**: `(context, agent)` を受け取りブール値を返す関数
|
||||
- **非同期関数**: 複雑な条件ロジック向けの async 関数
|
||||
- **Callable 関数**: `(context, agent)` を受け取り、ブール値を返す関数
|
||||
- **Async 関数**: 複雑な条件ロジック用の async 関数
|
||||
|
||||
無効化されたツールは実行時に LLM から完全に隠されるため、次の用途に有効です。
|
||||
無効化されたツールはランタイムで LLM から完全に隠されるため、次のような用途に便利です:
|
||||
|
||||
- ユーザー権限に基づく機能ゲート
|
||||
- 環境別ツール可用性 ( dev vs prod )
|
||||
- 異なるツール構成の A/B テスト
|
||||
- ユーザー権限に基づく機能ゲーティング
|
||||
- 環境固有のツール可用性 (dev と prod)
|
||||
- 異なるツール設定の A/B テスト
|
||||
- ランタイム状態に基づく動的ツールフィルタリング
|
||||
|
||||
## Experimental: Codex tool
|
||||
## 実験的: Codex ツール
|
||||
|
||||
`codex_tool` は Codex CLI をラップし、エージェントがツール呼び出し中にワークスペーススコープのタスク ( shell、ファイル編集、 MCP ツール ) を実行できるようにします。この面は実験的であり、変更される可能性があります。
|
||||
`codex_tool` は Codex CLI をラップし、ツール呼び出し中にエージェントがワークスペーススコープのタスク (シェル、ファイル編集、MCP ツール) を実行できるようにします。このサーフェスは実験的であり、変更される可能性があります。
|
||||
|
||||
現在の run を離れずに、メインエージェントから Codex に境界付きワークスペースタスクを委譲したい場合に使用します。デフォルトのツール名は `codex` です。カスタム名を設定する場合、それは `codex` であるか `codex_` で始まる必要があります。エージェントに複数の Codex ツールがある場合、それぞれが一意名である必要があります。
|
||||
メインエージェントに、現在の実行を離れずに境界付けられたワークスペースタスクを Codex へ委任させたい場合に使用してください。デフォルトでは、ツール名は `codex` です。カスタム名を設定する場合は、`codex` であるか、`codex_` で始まる必要があります。エージェントに複数の Codex ツールが含まれる場合、それぞれ一意の名前を使用する必要があります。
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
@@ -788,7 +785,7 @@ agent = Agent(
|
||||
sandbox_mode="workspace-write",
|
||||
working_directory="/path/to/repo",
|
||||
default_thread_options=ThreadOptions(
|
||||
model="gpt-5.4",
|
||||
model="gpt-5.5",
|
||||
model_reasoning_effort="low",
|
||||
network_access_enabled=True,
|
||||
web_search_mode="disabled",
|
||||
@@ -803,33 +800,33 @@ agent = Agent(
|
||||
)
|
||||
```
|
||||
|
||||
まず次のオプショングループから始めてください。
|
||||
まず、次のオプショングループを確認してください:
|
||||
|
||||
- 実行面: `sandbox_mode` と `working_directory` は Codex が操作できる場所を定義します。これらは組み合わせて設定し、作業ディレクトリが Git リポジトリ内にない場合は `skip_git_repo_check=True` を設定してください。
|
||||
- スレッドデフォルト: `default_thread_options=ThreadOptions(...)` は、モデル、推論努力、承認ポリシー、追加ディレクトリ、ネットワークアクセス、 Web 検索モードを設定します。レガシーの `web_search_enabled` より `web_search_mode` を優先してください。
|
||||
- ターンデフォルト: `default_turn_options=TurnOptions(...)` は、`idle_timeout_seconds` や任意のキャンセル `signal` など、ターンごとの動作を設定します。
|
||||
- ツール I/O: ツール呼び出しには、`{ "type": "text", "text": ... }` または `{ "type": "local_image", "path": ... }` を持つ `inputs` アイテムを少なくとも 1 つ含める必要があります。`output_schema` により構造化 Codex 応答を必須にできます。
|
||||
- 実行サーフェス: `sandbox_mode` と `working_directory` は、Codex が操作できる場所を定義します。これらは併用し、作業ディレクトリが Git リポジトリ内にない場合は `skip_git_repo_check=True` を設定してください。
|
||||
- スレッドのデフォルト: `default_thread_options=ThreadOptions(...)` は、モデル、推論エフォート、承認ポリシー、追加ディレクトリ、ネットワークアクセス、Web 検索モードを設定します。レガシーな `web_search_enabled` よりも `web_search_mode` を優先してください。
|
||||
- ターンのデフォルト: `default_turn_options=TurnOptions(...)` は、`idle_timeout_seconds` やオプションのキャンセル `signal` など、ターンごとの動作を設定します。
|
||||
- ツール I/O: ツール呼び出しには、`{ "type": "text", "text": ... }` または `{ "type": "local_image", "path": ... }` を持つ `inputs` 項目を少なくとも 1 つ含める必要があります。`output_schema` により、構造化された Codex レスポンスを必須にできます。
|
||||
|
||||
スレッド再利用と永続化は別々の制御です。
|
||||
スレッドの再利用と永続化は別々の制御です:
|
||||
|
||||
- `persist_session=True` は、同一ツールインスタンスへの繰り返し呼び出しで 1 つの Codex スレッドを再利用します。
|
||||
- `use_run_context_thread_id=True` は、同じ可変コンテキストオブジェクトを共有する run 間で、 run コンテキスト内にスレッド ID を保存して再利用します。
|
||||
- スレッド ID の優先順位は、呼び出しごとの `thread_id`、次に ( 有効時 ) run-context スレッド ID、次に設定済み `thread_id` オプションです。
|
||||
- デフォルト run-context キーは、`name="codex"` では `codex_thread_id`、`name="codex_<suffix>"` では `codex_thread_id_<suffix>` です。`run_context_thread_id_key` で上書きできます。
|
||||
- `persist_session=True` は、同じツールインスタンスへの繰り返し呼び出しで 1 つの Codex スレッドを再利用します。
|
||||
- `use_run_context_thread_id=True` は、同じ可変コンテキストオブジェクトを共有する複数の実行にわたって、実行コンテキスト内にスレッド ID を保存して再利用します。
|
||||
- スレッド ID の優先順位は、呼び出しごとの `thread_id`、次に実行コンテキストのスレッド ID (有効な場合)、次に設定された `thread_id` オプションです。
|
||||
- デフォルトの実行コンテキストキーは、`name="codex"` の場合は `codex_thread_id`、`name="codex_<suffix>"` の場合は `codex_thread_id_<suffix>` です。`run_context_thread_id_key` で上書きできます。
|
||||
|
||||
ランタイム設定:
|
||||
|
||||
- 認証: `CODEX_API_KEY` (推奨) または `OPENAI_API_KEY` を設定するか、`codex_options={"api_key": "..."}` を渡します。
|
||||
- ランタイム: `codex_options.base_url` は CLI の base URL を上書きします。
|
||||
- バイナリ解決: CLI パスを固定するには `codex_options.codex_path_override` (または `CODEX_PATH`) を設定します。設定しない場合、 SDK は `PATH` から `codex` を解決し、その後バンドル済み vendor バイナリへフォールバックします。
|
||||
- 環境: `codex_options.env` はサブプロセス環境を完全に制御します。これを指定すると、サブプロセスは `os.environ` を継承しません。
|
||||
- ストリーム制限: `codex_options.codex_subprocess_stream_limit_bytes` (または `OPENAI_AGENTS_CODEX_SUBPROCESS_STREAM_LIMIT_BYTES`) は stdout / stderr リーダー制限を制御します。有効範囲は `65536` から `67108864`、デフォルトは `8388608` です。
|
||||
- ストリーミング: `on_stream` はスレッド / ターンのライフサイクルイベントとアイテムイベント (`reasoning`、`command_execution`、`mcp_tool_call`、`file_change`、`web_search`、`todo_list`、`error` のアイテム更新) を受け取ります。
|
||||
- ランタイム: `codex_options.base_url` は CLI ベース URL を上書きします。
|
||||
- バイナリ解決: CLI パスを固定するには、`codex_options.codex_path_override` (または `CODEX_PATH`) を設定します。それ以外の場合、SDK は `PATH` から `codex` を解決し、その後、同梱された vendor バイナリにフォールバックします。
|
||||
- 環境: `codex_options.env` はサブプロセス環境を完全に制御します。これが指定されている場合、サブプロセスは `os.environ` を継承しません。
|
||||
- ストリーム制限: `codex_options.codex_subprocess_stream_limit_bytes` (または `OPENAI_AGENTS_CODEX_SUBPROCESS_STREAM_LIMIT_BYTES`) は stdout / stderr リーダー制限を制御します。有効範囲は `65536` から `67108864` で、デフォルトは `8388608` です。
|
||||
- ストリーミング: `on_stream` は、スレッド / ターンのライフサイクルイベントと、項目イベント (`reasoning`、`command_execution`、`mcp_tool_call`、`file_change`、`web_search`、`todo_list`、および `error` 項目の更新) を受け取ります。
|
||||
- 出力: 結果には `response`、`usage`、`thread_id` が含まれます。usage は `RunContextWrapper.usage` に追加されます。
|
||||
|
||||
参照:
|
||||
リファレンス:
|
||||
|
||||
- [Codex tool API reference](ref/extensions/experimental/codex/codex_tool.md)
|
||||
- [ThreadOptions reference](ref/extensions/experimental/codex/thread_options.md)
|
||||
- [TurnOptions reference](ref/extensions/experimental/codex/turn_options.md)
|
||||
- 完全に実行可能なサンプルは `examples/tools/codex.py` と `examples/tools/codex_same_thread.py` を参照してください。
|
||||
- [Codex ツール API リファレンス](ref/extensions/experimental/codex/codex_tool.md)
|
||||
- [ThreadOptions リファレンス](ref/extensions/experimental/codex/thread_options.md)
|
||||
- [TurnOptions リファレンス](ref/extensions/experimental/codex/turn_options.md)
|
||||
- 完全に実行可能なサンプルについては、`examples/tools/codex.py` と `examples/tools/codex_same_thread.py` を参照してください。
|
||||
+56
-54
@@ -4,55 +4,55 @@ search:
|
||||
---
|
||||
# トレーシング
|
||||
|
||||
Agents SDK には組み込みのトレーシングが含まれており、エージェント実行中のイベントを包括的に記録します。これには、LLM の生成、ツール呼び出し、ハンドオフ、ガードレール、さらに発生したカスタムイベントも含まれます。[Traces ダッシュボード](https://platform.openai.com/traces) を使用すると、開発中および本番環境でワークフローをデバッグ、可視化、監視できます。
|
||||
Agents SDK には組み込みのトレーシングが含まれており、エージェント実行中のイベント(LLM 生成、ツール呼び出し、ハンドオフ、ガードレール、さらには発生したカスタムイベントまで)を包括的に収集します。[Traces ダッシュボード](https://platform.openai.com/traces)を使用すると、開発中および本番環境でワークフローをデバッグ、可視化、監視できます。
|
||||
|
||||
!!!note
|
||||
|
||||
トレーシングはデフォルトで有効です。無効にする一般的な方法は 3 つあります。
|
||||
トレーシングはデフォルトで有効です。一般的には次の 3 つの方法で無効化できます:
|
||||
|
||||
1. 環境変数 `OPENAI_AGENTS_DISABLE_TRACING=1` を設定して、グローバルにトレーシングを無効化できます
|
||||
2. [`set_tracing_disabled(True)`][agents.set_tracing_disabled] を使って、コード内でグローバルにトレーシングを無効化できます
|
||||
3. [`agents.run.RunConfig.tracing_disabled`][] を `True` に設定して、単一の実行に対してトレーシングを無効化できます
|
||||
1. 環境変数 `OPENAI_AGENTS_DISABLE_TRACING=1` を設定することで、トレーシングをグローバルに無効化できます
|
||||
2. コード内で [`set_tracing_disabled(True)`][agents.set_tracing_disabled] を使用して、トレーシングをグローバルに無効化できます
|
||||
3. [`agents.run.RunConfig.tracing_disabled`][] を `True` に設定することで、単一の実行に対してトレーシングを無効化できます
|
||||
|
||||
***OpenAI の API を使用し、Zero Data Retention ( ZDR ) ポリシーのもとで運用している組織では、トレーシングは利用できません。***
|
||||
***OpenAI の API を使用し、Zero Data Retention (ZDR) ポリシーの下で運用している組織では、トレーシングは利用できません。***
|
||||
|
||||
## トレースとスパン
|
||||
|
||||
- **トレース** は、1 つの「ワークフロー」における単一のエンドツーエンド操作を表します。トレースは Span で構成されます。トレースには次のプロパティがあります。
|
||||
- `workflow_name`: 論理的なワークフローまたはアプリです。たとえば、「Code generation」や「Customer service」などです。
|
||||
- **トレース** は、「ワークフロー」の単一のエンドツーエンド操作を表します。トレースはスパンで構成されます。トレースには次のプロパティがあります:
|
||||
- `workflow_name`: 論理的なワークフローまたはアプリです。たとえば「コード生成」や「カスタマーサービス」です。
|
||||
- `trace_id`: トレースの一意な ID です。指定しない場合は自動生成されます。形式は `trace_<32_alphanumeric>` である必要があります。
|
||||
- `group_id`: オプションのグループ ID で、同じ会話内の複数のトレースを関連付けるために使用します。たとえば、チャットスレッド ID を使用できます。
|
||||
- `group_id`: 任意のグループ ID で、同じ会話からの複数のトレースを関連付けるために使用します。たとえば、チャットスレッド ID を使用できます。
|
||||
- `disabled`: True の場合、トレースは記録されません。
|
||||
- `metadata`: トレースのオプションのメタデータです。
|
||||
- **スパン** は、開始時刻と終了時刻を持つ操作を表します。スパンには次のものがあります。
|
||||
- `metadata`: トレースの任意のメタデータです。
|
||||
- **スパン** は、開始時刻と終了時刻を持つ操作を表します。スパンには次のものがあります:
|
||||
- `started_at` と `ended_at` のタイムスタンプ。
|
||||
- `trace_id`: そのスパンが属するトレースを表します
|
||||
- `parent_id`: このスパンの親 Span を指します(存在する場合)
|
||||
- `span_data`: Span に関する情報です。たとえば、`AgentSpanData` には Agent に関する情報が、`GenerationSpanData` には LLM 生成に関する情報が含まれます。
|
||||
- `trace_id`: 所属するトレースを表します
|
||||
- `parent_id`: このスパンの親スパン(存在する場合)を指します
|
||||
- `span_data`: スパンに関する情報です。たとえば、`AgentSpanData` にはエージェントに関する情報が含まれ、`GenerationSpanData` には LLM 生成に関する情報が含まれます。
|
||||
|
||||
## デフォルトのトレーシング
|
||||
|
||||
デフォルトでは、SDK は次のものをトレースします。
|
||||
デフォルトでは、SDK は次をトレースします:
|
||||
|
||||
- `Runner.{run, run_sync, run_streamed}()` 全体が `trace()` でラップされます。
|
||||
- `Runner.{run, run_sync, run_streamed}()` 全体は `trace()` でラップされます。
|
||||
- エージェントが実行されるたびに、`agent_span()` でラップされます
|
||||
- LLM の生成は `generation_span()` でラップされます
|
||||
- 関数ツールの各呼び出しは `function_span()` でラップされます
|
||||
- LLM 生成は `generation_span()` でラップされます
|
||||
- 各関数ツール呼び出しは `function_span()` でラップされます
|
||||
- ガードレールは `guardrail_span()` でラップされます
|
||||
- ハンドオフは `handoff_span()` でラップされます
|
||||
- 音声入力( speech-to-text )は `transcription_span()` でラップされます
|
||||
- 音声出力( text-to-speech )は `speech_span()` でラップされます
|
||||
- 関連する音声スパンは `speech_group_span()` の配下になる場合があります
|
||||
- 音声入力(音声からテキストへの変換)は `transcription_span()` でラップされます
|
||||
- 音声出力(テキストから音声への変換)は `speech_span()` でラップされます
|
||||
- 関連する音声スパンは、`speech_group_span()` の下に親子関係として配置される場合があります
|
||||
|
||||
デフォルトでは、トレース名は「Agent workflow」です。`trace` を使用する場合はこの名前を設定できます。また、[`RunConfig`][agents.run.RunConfig] を使って名前やその他のプロパティを設定することもできます。
|
||||
デフォルトでは、トレース名は「Agent workflow」です。`trace` を使用する場合はこの名前を設定できます。また、[`RunConfig`][agents.run.RunConfig] で名前やその他のプロパティを構成できます。
|
||||
|
||||
さらに、[カスタムトレースプロセッサー](#custom-tracing-processors) を設定して、トレースを他の送信先へ送ることもできます(置き換え先または補助的な送信先として)。
|
||||
さらに、[カスタムトレーシングプロセッサー](#custom-tracing-processors)を設定して、トレースを他の送信先へ送ることもできます(置き換え先または副次的な送信先として)。
|
||||
|
||||
## 長時間実行ワーカーと即時エクスポート
|
||||
## 長時間稼働するワーカーと即時エクスポート
|
||||
|
||||
デフォルトの [`BatchTraceProcessor`][agents.tracing.processors.BatchTraceProcessor] は、数秒ごとにバックグラウンドでトレースをエクスポートします。あるいは、インメモリキューがサイズのしきい値に達した場合はそれより早くエクスポートし、さらにプロセス終了時には最終フラッシュも実行します。Celery、RQ、Dramatiq、FastAPI のバックグラウンドタスクなどの長時間実行ワーカーでは、通常は追加コードなしでトレースが自動的にエクスポートされますが、各ジョブの完了直後には Traces ダッシュボードに表示されないことがあります。
|
||||
デフォルトの [`BatchTraceProcessor`][agents.tracing.processors.BatchTraceProcessor] は、数秒ごとにバックグラウンドでトレースをエクスポートします。メモリ内キューがサイズトリガーに達した場合はそれより早くエクスポートし、プロセス終了時には最終フラッシュも行います。Celery、RQ、Dramatiq、FastAPI のバックグラウンドタスクなどの長時間稼働するワーカーでは、通常、追加のコードなしでトレースが自動的にエクスポートされますが、各ジョブ完了直後に Traces ダッシュボードに表示されない場合があります。
|
||||
|
||||
作業単位の終了時に即時配信を保証したい場合は、トレースコンテキストを抜けた後で [`flush_traces()`][agents.tracing.flush_traces] を呼び出してください。
|
||||
作業単位の終了時に即時配信を保証する必要がある場合は、トレースコンテキストが終了した後に [`flush_traces()`][agents.tracing.flush_traces] を呼び出してください。
|
||||
|
||||
```python
|
||||
from agents import Runner, flush_traces, trace
|
||||
@@ -89,11 +89,11 @@ async def run(prompt: str, background_tasks: BackgroundTasks):
|
||||
return {"status": "queued"}
|
||||
```
|
||||
|
||||
[`flush_traces()`][agents.tracing.flush_traces] は、現在バッファされているトレースとスパンがエクスポートされるまでブロックするため、不完全なトレースをフラッシュしないよう、`trace()` が閉じた後に呼び出してください。デフォルトのエクスポート遅延で問題ない場合は、この呼び出しは省略できます。
|
||||
[`flush_traces()`][agents.tracing.flush_traces] は、現在バッファーされているトレースとスパンがエクスポートされるまでブロックします。そのため、部分的に構築されたトレースをフラッシュしないよう、`trace()` が閉じた後に呼び出してください。デフォルトのエクスポート遅延で問題ない場合は、この呼び出しを省略できます。
|
||||
|
||||
## 上位レベルのトレース
|
||||
|
||||
複数の `run()` 呼び出しを 1 つのトレースに含めたい場合があります。その場合は、コード全体を `trace()` でラップできます。
|
||||
場合によっては、`run()` への複数回の呼び出しを単一のトレースの一部にしたいことがあります。コード全体を `trace()` でラップすることで実現できます。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, trace
|
||||
@@ -108,49 +108,49 @@ async def main():
|
||||
print(f"Rating: {second_result.final_output}")
|
||||
```
|
||||
|
||||
1. 2 回の `Runner.run` 呼び出しは `with trace()` でラップされているため、個別に 2 つのトレースを作成するのではなく、全体のトレースの一部になります。
|
||||
1. 2 回の `Runner.run` 呼び出しが `with trace()` でラップされているため、個々の実行は 2 つのトレースを作成するのではなく、全体のトレースの一部になります。
|
||||
|
||||
## トレースの作成
|
||||
|
||||
[`trace()`][agents.tracing.trace] 関数を使用してトレースを作成できます。トレースは開始と終了が必要です。その方法は 2 つあります。
|
||||
[`trace()`][agents.tracing.trace] 関数を使用してトレースを作成できます。トレースは開始し、終了する必要があります。その方法には 2 つの選択肢があります:
|
||||
|
||||
1. **推奨**: トレースをコンテキストマネージャーとして使用します。つまり、`with trace(...) as my_trace` のように使います。これにより、適切なタイミングでトレースが自動的に開始および終了されます。
|
||||
1. **推奨**: トレースをコンテキストマネージャーとして使用します。つまり、`with trace(...) as my_trace` のように使用します。これにより、適切なタイミングでトレースが自動的に開始および終了されます。
|
||||
2. [`trace.start()`][agents.tracing.Trace.start] と [`trace.finish()`][agents.tracing.Trace.finish] を手動で呼び出すこともできます。
|
||||
|
||||
現在のトレースは、Python の [`contextvar`](https://docs.python.org/3/library/contextvars.html) を通じて追跡されます。これは、並行実行でも自動的に動作することを意味します。トレースを手動で開始または終了する場合は、現在のトレースを更新するために `start()` / `finish()` に `mark_as_current` と `reset_current` を渡す必要があります。
|
||||
現在のトレースは、Python の [`contextvar`](https://docs.python.org/3/library/contextvars.html) を介して追跡されます。これは、並行処理でも自動的に機能することを意味します。手動でトレースを開始/終了する場合は、現在のトレースを更新するために、`start()`/`finish()` に `mark_as_current` と `reset_current` を渡す必要があります。
|
||||
|
||||
## スパンの作成
|
||||
|
||||
各種 [`*_span()`][agents.tracing.create] メソッドを使用してスパンを作成できます。一般に、スパンを手動で作成する必要はありません。カスタムのスパン情報を追跡するために [`custom_span()`][agents.tracing.custom_span] 関数も利用できます。
|
||||
各種 [`*_span()`][agents.tracing.create] メソッドを使用してスパンを作成できます。通常、手動でスパンを作成する必要はありません。カスタムスパン情報を追跡するために、[`custom_span()`][agents.tracing.custom_span] 関数を利用できます。
|
||||
|
||||
スパンは自動的に現在のトレースの一部となり、最も近い現在のスパンの配下にネストされます。これは Python の [`contextvar`](https://docs.python.org/3/library/contextvars.html) によって追跡されます。
|
||||
スパンは自動的に現在のトレースの一部となり、最も近い現在のスパンの下にネストされます。この現在のスパンは、Python の [`contextvar`](https://docs.python.org/3/library/contextvars.html) を介して追跡されます。
|
||||
|
||||
## 機微データ
|
||||
## 機密データ
|
||||
|
||||
一部のスパンでは、機微データとなり得る情報を取得する場合があります。
|
||||
一部のスパンは、潜在的に機密性の高いデータを取得する場合があります。
|
||||
|
||||
`generation_span()` は LLM 生成の入出力を保存し、`function_span()` は関数呼び出しの入出力を保存します。これらには機微データが含まれる可能性があるため、[`RunConfig.trace_include_sensitive_data`][agents.run.RunConfig.trace_include_sensitive_data] によってそのデータの取得を無効化できます。
|
||||
`generation_span()` は LLM 生成の入出力を保存し、`function_span()` は関数呼び出しの入出力を保存します。これらには機密データが含まれる可能性があるため、[`RunConfig.trace_include_sensitive_data`][agents.run.RunConfig.trace_include_sensitive_data] を使用してそのデータの取得を無効化できます。
|
||||
|
||||
同様に、音声スパンにはデフォルトで入力音声と出力音声の base64 エンコード済み PCM データが含まれます。[`VoicePipelineConfig.trace_include_sensitive_audio_data`][agents.voice.pipeline_config.VoicePipelineConfig.trace_include_sensitive_audio_data] を設定することで、この音声データの取得を無効化できます。
|
||||
同様に、音声スパンにはデフォルトで、入力音声と出力音声の base64 エンコードされた PCM データが含まれます。この音声データの取得は、[`VoicePipelineConfig.trace_include_sensitive_audio_data`][agents.voice.pipeline_config.VoicePipelineConfig.trace_include_sensitive_audio_data] を構成することで無効化できます。
|
||||
|
||||
デフォルトでは、`trace_include_sensitive_data` は `True` です。コードを書かずにデフォルト値を設定するには、アプリの実行前に環境変数 `OPENAI_AGENTS_TRACE_INCLUDE_SENSITIVE_DATA` を `true/1` または `false/0` に設定してください。
|
||||
デフォルトでは、`trace_include_sensitive_data` は `True` です。アプリを実行する前に `OPENAI_AGENTS_TRACE_INCLUDE_SENSITIVE_DATA` 環境変数を `true/1` または `false/0` にエクスポートすることで、コードを書かずにデフォルトを設定できます。
|
||||
|
||||
## カスタムトレーシングプロセッサー
|
||||
|
||||
トレーシングの高レベルなアーキテクチャは次のとおりです。
|
||||
トレーシングの高レベルのアーキテクチャは次のとおりです:
|
||||
|
||||
- 初期化時に、トレースの作成を担当するグローバルな [`TraceProvider`][agents.tracing.setup.TraceProvider] を作成します。
|
||||
- `TraceProvider` を [`BatchTraceProcessor`][agents.tracing.processors.BatchTraceProcessor] で設定します。このプロセッサーは、トレース / スパンをバッチで [`BackendSpanExporter`][agents.tracing.processors.BackendSpanExporter] に送信し、これがスパンとトレースをバッチで OpenAI バックエンドへエクスポートします。
|
||||
- 初期化時に、グローバルな [`TraceProvider`][agents.tracing.setup.TraceProvider] を作成します。これはトレースの作成を担います。
|
||||
- `TraceProvider` は、トレース/スパンをバッチ単位で [`BackendSpanExporter`][agents.tracing.processors.BackendSpanExporter] に送信する [`BatchTraceProcessor`][agents.tracing.processors.BatchTraceProcessor] で構成します。`BackendSpanExporter` は、スパンとトレースをバッチ単位で OpenAI バックエンドへエクスポートします。
|
||||
|
||||
このデフォルト設定をカスタマイズして、別のバックエンドまたは追加のバックエンドにトレースを送信したり、エクスポーターの動作を変更したりするには、2 つの方法があります。
|
||||
このデフォルト設定をカスタマイズし、代替または追加のバックエンドへトレースを送信したり、エクスポーターの動作を変更したりするには、2 つの選択肢があります:
|
||||
|
||||
1. [`add_trace_processor()`][agents.tracing.add_trace_processor] を使うと、準備が整ったトレースとスパンを受け取る **追加の** トレースプロセッサーを追加できます。これにより、トレースを OpenAI のバックエンドへ送信することに加えて、独自の処理も行えます。
|
||||
2. [`set_trace_processors()`][agents.tracing.set_trace_processors] を使うと、デフォルトのプロセッサーを独自のトレースプロセッサーで **置き換え** できます。これは、そうした処理を行う `TracingProcessor` を含めない限り、トレースが OpenAI バックエンドに送信されないことを意味します。
|
||||
1. [`add_trace_processor()`][agents.tracing.add_trace_processor] を使用すると、トレースやスパンが準備でき次第それらを受け取る **追加の** トレースプロセッサーを追加できます。これにより、OpenAI のバックエンドへトレースを送信することに加えて、独自の処理を行えます。
|
||||
2. [`set_trace_processors()`][agents.tracing.set_trace_processors] を使用すると、デフォルトのプロセッサーを独自のトレースプロセッサーで **置き換える** ことができます。つまり、その処理を行う `TracingProcessor` を含めない限り、トレースは OpenAI のバックエンドに送信されません。
|
||||
|
||||
|
||||
## non-OpenAI モデルでのトレーシング
|
||||
## OpenAI 以外のモデルでのトレーシング
|
||||
|
||||
OpenAI 以外のモデルでも、OpenAI API キーを使用することで、トレーシングを無効化することなく OpenAI Traces ダッシュボードで無料のトレーシングを有効にできます。アダプターの選択と設定上の注意点については、Models ガイドの [Third-party adapters](models/index.md#third-party-adapters) セクションを参照してください。
|
||||
OpenAI 以外のモデルで OpenAI API キーを使用すると、トレーシングを無効化する必要なく、OpenAI Traces ダッシュボードで無料のトレーシングを有効にできます。アダプターの選択とセットアップ上の注意事項については、Models ガイドの[サードパーティアダプター](models/index.md#third-party-adapters)セクションを参照してください。
|
||||
|
||||
```python
|
||||
import os
|
||||
@@ -171,7 +171,7 @@ agent = Agent(
|
||||
)
|
||||
```
|
||||
|
||||
単一の実行に対してのみ別のトレーシングキーが必要な場合は、グローバルエクスポーターを変更するのではなく、`RunConfig` 経由で渡してください。
|
||||
単一の実行に対してのみ別のトレーシングキーが必要な場合は、グローバルエクスポーターを変更する代わりに `RunConfig` 経由で渡してください。
|
||||
|
||||
```python
|
||||
from agents import Runner, RunConfig
|
||||
@@ -183,21 +183,21 @@ await Runner.run(
|
||||
)
|
||||
```
|
||||
|
||||
## 追加の注記
|
||||
- Openai Traces ダッシュボードで無料トレースを表示できます。
|
||||
## 追加の注意事項
|
||||
- OpenAI Traces ダッシュボードで無料のトレースを表示できます。
|
||||
|
||||
|
||||
## エコシステム統合
|
||||
|
||||
以下のコミュニティおよびベンダー統合は、OpenAI Agents SDK のトレーシング機能をサポートしています。
|
||||
次のコミュニティおよびベンダー統合は、OpenAI Agents SDK のトレーシングインターフェイスをサポートしています。
|
||||
|
||||
### 外部トレーシングプロセッサー一覧
|
||||
|
||||
- [Weights & Biases](https://weave-docs.wandb.ai/guides/integrations/openai_agents)
|
||||
- [Arize-Phoenix](https://docs.arize.com/phoenix/tracing/integrations-tracing/openai-agents-sdk)
|
||||
- [Future AGI](https://docs.futureagi.com/future-agi/products/observability/auto-instrumentation/openai_agents)
|
||||
- [MLflow (self-hosted/OSS)](https://mlflow.org/docs/latest/tracing/integrations/openai-agent)
|
||||
- [MLflow (Databricks hosted)](https://docs.databricks.com/aws/en/mlflow/mlflow-tracing#-automatic-tracing)
|
||||
- [MLflow (セルフホスト/OSS)](https://mlflow.org/docs/latest/tracing/integrations/openai-agent)
|
||||
- [MLflow (Databricks ホスト)](https://docs.databricks.com/aws/en/mlflow/mlflow-tracing#-automatic-tracing)
|
||||
- [Braintrust](https://braintrust.dev/docs/guides/traces/integrations#openai-agents-sdk)
|
||||
- [Pydantic Logfire](https://logfire.pydantic.dev/docs/integrations/llms/openai/#openai-agents)
|
||||
- [AgentOps](https://docs.agentops.ai/v1/integrations/agentssdk)
|
||||
@@ -218,4 +218,6 @@ await Runner.run(
|
||||
- [PromptLayer](https://docs.promptlayer.com/languages/integrations#openai-agents-sdk)
|
||||
- [HoneyHive](https://docs.honeyhive.ai/v2/integrations/openai-agents)
|
||||
- [Asqav](https://www.asqav.com/docs/integrations#openai-agents)
|
||||
- [Datadog](https://docs.datadoghq.com/llm_observability/instrumentation/auto_instrumentation/?tab=python#openai-agents)
|
||||
- [Datadog](https://docs.datadoghq.com/llm_observability/instrumentation/auto_instrumentation/?tab=python#openai-agents)
|
||||
- [Latitude](https://docs.latitude.so/telemetry/frameworks/openai-agents)
|
||||
- [DProvenanceKit](https://dprovenance.dev/openai-agents/)
|
||||
+21
-21
@@ -2,24 +2,24 @@
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# 使用方法
|
||||
# 使用量
|
||||
|
||||
Agents SDK は、すべての実行についてトークン使用量を自動的に追跡します。実行コンテキストからこれにアクセスし、コストの監視、制限の適用、または分析の記録に使用できます。
|
||||
Agents SDK は、すべての実行についてトークン使用量を自動的に追跡します。使用量には実行コンテキストからアクセスでき、コストの監視、制限の適用、分析の記録に利用できます。
|
||||
|
||||
## 追跡対象
|
||||
|
||||
- **requests**: 実行された LLM API 呼び出し回数
|
||||
- **requests**: 実行された LLM API 呼び出しの数
|
||||
- **input_tokens**: 送信された入力トークンの合計
|
||||
- **output_tokens**: 受信した出力トークンの合計
|
||||
- **output_tokens**: 受信された出力トークンの合計
|
||||
- **total_tokens**: 入力 + 出力
|
||||
- **request_usage_entries**: リクエストごとの使用量内訳の一覧
|
||||
- **request_usage_entries**: リクエストごとの使用量内訳のリスト
|
||||
- **details**:
|
||||
- `input_tokens_details.cached_tokens`
|
||||
- `output_tokens_details.reasoning_tokens`
|
||||
|
||||
## 実行からの使用量アクセス
|
||||
## 実行からの使用量へのアクセス
|
||||
|
||||
`Runner.run(...)` の後、`result.context_wrapper.usage` 経由で使用量にアクセスします。
|
||||
`Runner.run(...)` の後、使用量には `result.context_wrapper.usage` 経由でアクセスします。
|
||||
|
||||
```python
|
||||
result = await Runner.run(agent, "What's the weather in Tokyo?")
|
||||
@@ -31,20 +31,20 @@ print("Output tokens:", usage.output_tokens)
|
||||
print("Total tokens:", usage.total_tokens)
|
||||
```
|
||||
|
||||
使用量は、実行中のすべてのモデル呼び出し(ツール呼び出しとハンドオフを含む)にわたって集計されます。
|
||||
使用量は、この実行中のすべてのモデル呼び出し(ツール呼び出しやハンドオフを含む)にわたって集計されます。
|
||||
|
||||
### サードパーティアダプターでの使用量有効化
|
||||
### サードパーティ製アダプターでの使用量の有効化
|
||||
|
||||
使用量レポートは、サードパーティアダプターおよびプロバイダーバックエンドによって異なります。アダプター経由のモデルに依存し、正確な `result.context_wrapper.usage` の値が必要な場合:
|
||||
使用量のレポートは、サードパーティ製アダプターやプロバイダーバックエンドによって異なります。アダプター経由のモデルに依存しており、正確な `result.context_wrapper.usage` の値が必要な場合は:
|
||||
|
||||
- `AnyLLMModel` では、上流プロバイダーが使用量を返すと自動的に伝播されます。ストリーミング Chat Completions バックエンドでは、使用量チャンクが出力される前に `ModelSettings(include_usage=True)` が必要な場合があります。
|
||||
- `LitellmModel` では、一部のプロバイダーバックエンドは既定で使用量をレポートしないため、`ModelSettings(include_usage=True)` が必要になることがよくあります。
|
||||
- `AnyLLMModel` では、上流プロバイダーが使用量を返す場合、使用量は自動的に伝播されます。ストリーミングされた Chat Completions バックエンドでは、使用量チャンクが出力される前に `ModelSettings(include_usage=True)` が必要になる場合があります。
|
||||
- `LitellmModel` では、一部のプロバイダーバックエンドはデフォルトで使用量を報告しないため、`ModelSettings(include_usage=True)` が必要になることがよくあります。
|
||||
|
||||
Models ガイドの [Third-party adapters](models/index.md#third-party-adapters) セクションにあるアダプター固有の注意事項を確認し、デプロイ予定の正確なプロバイダーバックエンドを検証してください。
|
||||
Models ガイドの [サードパーティ製アダプター](models/index.md#third-party-adapters) セクションにあるアダプター固有の注記を確認し、デプロイ予定の正確なプロバイダーバックエンドを検証してください。
|
||||
|
||||
## リクエストごとの使用量追跡
|
||||
|
||||
SDK は、各 API リクエストの使用量を `request_usage_entries` で自動追跡します。これは、詳細なコスト計算やコンテキストウィンドウ消費の監視に役立ちます。
|
||||
SDK は、各 API リクエストの使用量を `request_usage_entries` で自動的に追跡します。これは、詳細なコスト計算やコンテキストウィンドウ消費量の監視に役立ちます。
|
||||
|
||||
```python
|
||||
result = await Runner.run(agent, "What's the weather in Tokyo?")
|
||||
@@ -53,9 +53,9 @@ for i, request in enumerate(result.context_wrapper.usage.request_usage_entries):
|
||||
print(f"Request {i + 1}: {request.input_tokens} in, {request.output_tokens} out")
|
||||
```
|
||||
|
||||
## セッションでの使用量アクセス
|
||||
## セッションでの使用量へのアクセス
|
||||
|
||||
`Session`(例: `SQLiteSession`)を使用する場合、`Runner.run(...)` の各呼び出しは、その特定の実行の使用量を返します。セッションはコンテキスト用に会話履歴を維持しますが、各実行の使用量は独立しています。
|
||||
`Session`(例: `SQLiteSession`)を使用する場合、`Runner.run(...)` の各呼び出しは、その特定の実行の使用量を返します。セッションはコンテキスト用に会話履歴を保持しますが、各実行の使用量は独立しています。
|
||||
|
||||
```python
|
||||
session = SQLiteSession("my_conversation")
|
||||
@@ -67,11 +67,11 @@ second = await Runner.run(agent, "Can you elaborate?", session=session)
|
||||
print(second.context_wrapper.usage.total_tokens) # Usage for second run
|
||||
```
|
||||
|
||||
セッションは実行間で会話コンテキストを保持しますが、各 `Runner.run()` 呼び出しで返される使用量メトリクスは、その特定の実行のみを表す点に注意してください。セッションでは、前のメッセージが各実行の入力として再投入される場合があり、これが後続ターンの入力トークン数に影響します。
|
||||
セッションは実行間の会話コンテキストを保持しますが、各 `Runner.run()` 呼び出しから返される使用量メトリクスは、その特定の実行のみを表します。セッションでは、以前のメッセージが各実行に入力として再投入される場合があり、これにより以降のターンにおける入力トークン数に影響します。
|
||||
|
||||
## フックでの使用量活用
|
||||
## フックでの使用量の利用
|
||||
|
||||
`RunHooks` を使用している場合、各フックに渡される `context` オブジェクトには `usage` が含まれます。これにより、ライフサイクルの重要なタイミングで使用量をログ記録できます。
|
||||
`RunHooks` を使用している場合、各フックに渡される `context` オブジェクトには `usage` が含まれます。これにより、主要なライフサイクルのタイミングで使用量をログ記録できます。
|
||||
|
||||
```python
|
||||
class MyHooks(RunHooks):
|
||||
@@ -82,9 +82,9 @@ class MyHooks(RunHooks):
|
||||
|
||||
## API リファレンス
|
||||
|
||||
詳細な API ドキュメントは以下を参照してください。
|
||||
詳細な API ドキュメントについては、以下を参照してください:
|
||||
|
||||
- [`Usage`][agents.usage.Usage] - 使用量追跡データ構造
|
||||
- [`RequestUsage`][agents.usage.RequestUsage] - リクエストごとの使用量詳細
|
||||
- [`RunContextWrapper`][agents.run.RunContextWrapper] - 実行コンテキストから使用量にアクセス
|
||||
- [`RunContextWrapper`][agents.run.RunContextWrapper] - 実行コンテキストからの使用量へのアクセス
|
||||
- [`RunHooks`][agents.run.RunHooks] - 使用量追跡ライフサイクルへのフック
|
||||
+25
-24
@@ -2,25 +2,25 @@
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# エージェント可視化
|
||||
# エージェントの可視化
|
||||
|
||||
エージェント可視化では、 **Graphviz** を使用して、エージェントとその関係を構造化されたグラフィカル表現として生成できます。これは、アプリケーション内でエージェント、ツール、ハンドオフがどのように相互作用するかを理解するのに役立ちます。
|
||||
エージェントの可視化では、 **Graphviz** を使用してエージェントとその関係を構造化されたグラフィカルな表現として生成できます。これは、アプリケーション内でエージェント、ツール、ハンドオフがどのように相互作用するかを理解するのに役立ちます。
|
||||
|
||||
## インストール
|
||||
|
||||
オプションの `viz` 依存関係グループをインストールします。
|
||||
オプションの `viz` 依存関係グループをインストールします:
|
||||
|
||||
```bash
|
||||
pip install "openai-agents[viz]"
|
||||
```
|
||||
|
||||
## グラフ生成
|
||||
## グラフの生成
|
||||
|
||||
`draw_graph` 関数を使用してエージェント可視化を生成できます。この関数は、以下の構成を持つ有向グラフを作成します。
|
||||
`draw_graph` 関数を使用して、エージェントの可視化を生成できます。この関数は、次のように表現される有向グラフを作成します。
|
||||
|
||||
- **エージェント** は黄色のボックスとして表現されます。
|
||||
- **MCP サーバー** は灰色のボックスとして表現されます。
|
||||
- **ツール** は緑色の楕円として表現されます。
|
||||
- **エージェント** は黄色のボックスとして表されます。
|
||||
- **MCP サーバー** は灰色のボックスとして表されます。
|
||||
- **ツール** は緑色の楕円として表されます。
|
||||
- **ハンドオフ** は、あるエージェントから別のエージェントへの有向エッジです。
|
||||
|
||||
### 使用例
|
||||
@@ -67,37 +67,38 @@ triage_agent = Agent(
|
||||
draw_graph(triage_agent)
|
||||
```
|
||||
|
||||

|
||||

|
||||
|
||||
これにより、 **トリアージエージェント** の構造と、サブエージェントおよびツールへの接続を視覚的に表すグラフが生成されます。
|
||||
|
||||
これにより、 **triage agent** の構造と、サブエージェントおよびツールへの接続を視覚的に表すグラフが生成されます。
|
||||
|
||||
## 可視化の理解
|
||||
|
||||
生成されるグラフには以下が含まれます。
|
||||
生成されたグラフには次のものが含まれます。
|
||||
|
||||
- エントリーポイントを示す **開始ノード** (`__start__`)。
|
||||
- 黄色で塗りつぶされた **長方形** として表現されるエージェント。
|
||||
- 緑色で塗りつぶされた **楕円** として表現されるツール。
|
||||
- 灰色で塗りつぶされた **長方形** として表現される MCP サーバー。
|
||||
- エントリーポイントを示す **開始ノード** ( `__start__` )。
|
||||
- 黄色で塗りつぶされた **長方形** として表されるエージェント。
|
||||
- 緑色で塗りつぶされた **楕円** として表されるツール。
|
||||
- 灰色で塗りつぶされた **長方形** として表される MCP サーバー。
|
||||
- 相互作用を示す有向エッジ:
|
||||
- エージェント間ハンドオフには **実線矢印**。
|
||||
- ツール呼び出しには **点線矢印**。
|
||||
- MCP サーバー呼び出しには **破線矢印**。
|
||||
- 実行が終了する位置を示す **終了ノード** (`__end__`)。
|
||||
- エージェント間ハンドオフを示す **実線の矢印** 。
|
||||
- ツール呼び出しを示す **点線の矢印** 。
|
||||
- MCP サーバー呼び出しを示す **破線の矢印** 。
|
||||
- 実行が終了する場所を示す **終了ノード** ( `__end__` )。
|
||||
|
||||
**注:** MCP サーバーは `agents` パッケージの最近のバージョン ( **v0.2.8** で確認済み ) で描画されます。可視化に MCP ボックスが表示されない場合は、最新リリースにアップグレードしてください。
|
||||
**注:** MCP サーバーは、最近のバージョンの `agents` パッケージで描画されます( **v0.2.8** で確認済み)。可視化で MCP ボックスが表示されない場合は、最新リリースにアップグレードしてください。
|
||||
|
||||
## グラフのカスタマイズ
|
||||
|
||||
### グラフ表示
|
||||
デフォルトでは、 `draw_graph` はグラフをインライン表示します。グラフを別ウィンドウで表示するには、次のように記述します。
|
||||
### グラフの表示
|
||||
デフォルトでは、 `draw_graph` はグラフをインラインで表示します。グラフを別ウィンドウで表示するには、次のように記述します:
|
||||
|
||||
```python
|
||||
draw_graph(triage_agent).view()
|
||||
```
|
||||
|
||||
### グラフ保存
|
||||
デフォルトでは、 `draw_graph` はグラフをインライン表示します。ファイルとして保存するには、ファイル名を指定します。
|
||||
### グラフの保存
|
||||
デフォルトでは、 `draw_graph` はグラフをインラインで表示します。ファイルとして保存するには、ファイル名を指定します:
|
||||
|
||||
```python
|
||||
draw_graph(triage_agent, filename="agent_graph")
|
||||
|
||||
+20
-18
@@ -4,7 +4,7 @@ search:
|
||||
---
|
||||
# パイプラインとワークフロー
|
||||
|
||||
[`VoicePipeline`][agents.voice.pipeline.VoicePipeline] は、エージェントのワークフローを音声アプリに簡単に変換できるクラスです。実行するワークフローを渡すと、パイプラインが入力音声の文字起こし、音声終了の検出、適切なタイミングでのワークフロー呼び出し、そしてワークフロー出力の音声への変換を担います。
|
||||
[`VoicePipeline`][agents.voice.pipeline.VoicePipeline] は、エージェント型ワークフローを音声アプリに変換しやすくするクラスです。実行するワークフローを渡すと、パイプラインが入力音声の文字起こし、音声の終了検出、適切なタイミングでのワークフロー呼び出し、ワークフロー出力の音声への変換を処理します。
|
||||
|
||||
```mermaid
|
||||
graph LR
|
||||
@@ -34,29 +34,29 @@ graph LR
|
||||
|
||||
## パイプラインの設定
|
||||
|
||||
パイプラインを作成する際、いくつかの項目を設定できます。
|
||||
パイプラインを作成するときに、いくつかの項目を設定できます。
|
||||
|
||||
1. [`workflow`][agents.voice.workflow.VoiceWorkflowBase]:新しい音声が文字起こしされるたびに実行されるコードです。
|
||||
2. 使用する [`speech-to-text`][agents.voice.model.STTModel] および [`text-to-speech`][agents.voice.model.TTSModel] モデル
|
||||
3. [`config`][agents.voice.pipeline_config.VoicePipelineConfig]:次のような項目を設定できます。
|
||||
- モデルプロバイダー(モデル名をモデルにマッピングできます)
|
||||
- トレーシング(トレーシングを無効化するかどうか、音声ファイルをアップロードするかどうか、ワークフロー名、トレース ID など)
|
||||
- TTS および STT モデルの設定(プロンプト、言語、使用するデータ型など)
|
||||
1. [`workflow`][agents.voice.workflow.VoiceWorkflowBase]。新しい音声が文字起こしされるたびに実行されるコードです。
|
||||
2. 使用される [`speech-to-text`][agents.voice.model.STTModel] モデルと [`text-to-speech`][agents.voice.model.TTSModel] モデル
|
||||
3. [`config`][agents.voice.pipeline_config.VoicePipelineConfig]。次のような項目を設定できます。
|
||||
- モデル名をモデルにマッピングできるモデルプロバイダー
|
||||
- トレーシング。トレーシングを無効にするか、音声ファイルをアップロードするか、ワークフロー名、トレース ID などを含みます。
|
||||
- TTS モデルと STT モデルの設定。プロンプト、言語、使用するデータ型などです。
|
||||
|
||||
## パイプラインの実行
|
||||
|
||||
パイプラインは [`run()`][agents.voice.pipeline.VoicePipeline.run] メソッドで実行でき、音声入力を 2 つの形式で渡せます。
|
||||
[`run()`][agents.voice.pipeline.VoicePipeline.run] メソッドを通じてパイプラインを実行できます。このメソッドでは、音声入力を 2 つの形式で渡せます。
|
||||
|
||||
1. [`AudioInput`][agents.voice.input.AudioInput]:音声の全文書き起こしがすでにあり、それに対する結果だけを生成したい場合に使用します。話者が話し終えたタイミングを検出する必要がないケースで有用です。たとえば、事前録音の音声がある場合や、ユーザーが話し終えたことが明確な push-to-talk アプリなどです。
|
||||
2. [`StreamedAudioInput`][agents.voice.input.StreamedAudioInput]:ユーザーが話し終えたことを検出する必要がある可能性がある場合に使用します。検出された音声チャンクを順次プッシュでき、音声パイプラインは「activity detection」と呼ばれるプロセスにより、適切なタイミングで自動的にエージェントのワークフローを実行します。
|
||||
1. [`AudioInput`][agents.voice.input.AudioInput] は、完全な音声入力があり、それに対する実行結果だけを生成したい場合に使用します。これは、話者が話し終えたタイミングを検出する必要がない場合に便利です。たとえば、事前に録音された音声がある場合や、ユーザーが話し終えたことが明確なプッシュ・トゥ・トークアプリの場合です。
|
||||
2. [`StreamedAudioInput`][agents.voice.input.StreamedAudioInput] は、ユーザーが話し終えたタイミングを検出する必要がある場合に使用します。検出された音声チャンクをプッシュでき、音声パイプラインは「アクティビティ検出」と呼ばれるプロセスを通じて、適切なタイミングでエージェントワークフローを自動的に実行します。
|
||||
|
||||
## 結果
|
||||
## 実行結果
|
||||
|
||||
音声パイプライン実行の結果は [`StreamedAudioResult`][agents.voice.result.StreamedAudioResult] です。これは、発生したイベントをストリーミングできるオブジェクトです。[`VoiceStreamEvent`][agents.voice.events.VoiceStreamEvent] にはいくつかの種類があり、たとえば次のものがあります。
|
||||
音声パイプライン実行の実行結果は [`StreamedAudioResult`][agents.voice.result.StreamedAudioResult] です。これは、イベントが発生したときにストリーミングできるオブジェクトです。[`VoiceStreamEvent`][agents.voice.events.VoiceStreamEvent] には、次のような種類があります。
|
||||
|
||||
1. [`VoiceStreamEventAudio`][agents.voice.events.VoiceStreamEventAudio]:音声のチャンクを含みます。
|
||||
2. [`VoiceStreamEventLifecycle`][agents.voice.events.VoiceStreamEventLifecycle]:ターンの開始や終了などのライフサイクルイベントを通知します。
|
||||
3. [`VoiceStreamEventError`][agents.voice.events.VoiceStreamEventError]:エラーイベントです。
|
||||
1. [`VoiceStreamEventAudio`][agents.voice.events.VoiceStreamEventAudio]。音声チャンクを含みます。
|
||||
2. [`VoiceStreamEventLifecycle`][agents.voice.events.VoiceStreamEventLifecycle]。ターンの開始や終了などのライフサイクルイベントを通知します。
|
||||
3. [`VoiceStreamEventError`][agents.voice.events.VoiceStreamEventError]。エラーイベントです。
|
||||
|
||||
```python
|
||||
|
||||
@@ -65,15 +65,17 @@ result = await pipeline.run(input)
|
||||
async for event in result.stream():
|
||||
if event.type == "voice_stream_event_audio":
|
||||
# play audio
|
||||
pass
|
||||
elif event.type == "voice_stream_event_lifecycle":
|
||||
# lifecycle
|
||||
pass
|
||||
elif event.type == "voice_stream_event_error":
|
||||
# error
|
||||
...
|
||||
pass
|
||||
```
|
||||
|
||||
## ベストプラクティス
|
||||
|
||||
### 割り込み
|
||||
|
||||
Agents SDK は現在、[`StreamedAudioInput`][agents.voice.input.StreamedAudioInput] に対する組み込みの割り込みサポートを提供していません。代わりに、検出された各ターンごとに、ワークフローの別個の実行がトリガーされます。アプリケーション内で割り込みを扱いたい場合は、[`VoiceStreamEventLifecycle`][agents.voice.events.VoiceStreamEventLifecycle] イベントをリッスンできます。`turn_started` は、新しいターンが文字起こしされて処理が開始されたことを示します。`turn_ended` は、該当ターンのすべての音声がディスパッチされた後にトリガーされます。これらのイベントを使って、モデルがターンを開始したときに話者のマイクをミュートし、ターンに関連する音声をすべてフラッシュした後にミュート解除するといった実装が可能です。
|
||||
Agents SDK は現在、[`StreamedAudioInput`][agents.voice.input.StreamedAudioInput] 向けの組み込みの割り込み処理を提供していません。代わりに、検出された各ターンがワークフローの個別の実行をトリガーします。アプリケーション内で割り込みを処理したい場合は、[`VoiceStreamEventLifecycle`][agents.voice.events.VoiceStreamEventLifecycle] イベントをリッスンできます。`turn_started` は、新しいターンが文字起こしされ、処理が開始されることを示します。`turn_ended` は、対応するターンのすべての音声が送出された後にトリガーされます。これらのイベントを使用して、モデルがターンを開始したときに話者のマイクをミュートし、そのターンに関連するすべての音声をフラッシュした後にミュートを解除できます。
|
||||
+17
-17
@@ -6,7 +6,7 @@ search:
|
||||
|
||||
## 前提条件
|
||||
|
||||
Agents SDK の基本的な [クイックスタート手順](../quickstart.md) に従い、仮想環境をセットアップしていることを確認してください。次に、 SDK からオプションの音声依存関係をインストールします。
|
||||
Agents SDK の基本の [クイックスタート手順](../quickstart.md) に従い、仮想環境をセットアップ済みであることを確認してください。その後、SDK のオプションの音声依存関係をインストールします:
|
||||
|
||||
```bash
|
||||
pip install 'openai-agents[voice]'
|
||||
@@ -14,11 +14,11 @@ pip install 'openai-agents[voice]'
|
||||
|
||||
## 概念
|
||||
|
||||
主に理解しておくべき概念は [`VoicePipeline`][agents.voice.pipeline.VoicePipeline] で、これは 3 ステップのプロセスです。
|
||||
知っておくべき主な概念は [`VoicePipeline`][agents.voice.pipeline.VoicePipeline] です。これは 3 ステップのプロセスです:
|
||||
|
||||
1. 音声認識モデルを実行して、音声をテキストに変換します。
|
||||
2. コード(通常はエージェントオーケストレーションのワークフロー)を実行して、結果を生成します。
|
||||
3. 音声合成モデルを実行して、結果のテキストを音声に戻します。
|
||||
1. 音声をテキストに変換するために、speech-to-text モデルを実行します。
|
||||
2. 通常はエージェントを用いたワークフローであるコードを実行して、結果を生成します。
|
||||
3. text-to-speech モデルを実行して、結果のテキストを音声に戻します。
|
||||
|
||||
```mermaid
|
||||
graph LR
|
||||
@@ -48,7 +48,7 @@ graph LR
|
||||
|
||||
## エージェント
|
||||
|
||||
まず、いくつかの Agents をセットアップしましょう。この SDK でエージェントを構築したことがあれば、ここは馴染みのある内容です。複数の Agents と、ハンドオフ、ツールを用意します。
|
||||
まず、いくつかのエージェントをセットアップしましょう。この SDK でエージェントを構築したことがあれば、なじみのある内容のはずです。ここでは、いくつかのエージェント、ハンドオフ、ツールを用意します。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -72,19 +72,19 @@ def get_weather(city: str) -> str:
|
||||
|
||||
spanish_agent = Agent(
|
||||
name="Spanish",
|
||||
handoff_description="A spanish speaking agent.",
|
||||
handoff_description="A Spanish-speaking agent.",
|
||||
instructions=prompt_with_handoff_instructions(
|
||||
"You're speaking to a human, so be polite and concise. Speak in Spanish.",
|
||||
),
|
||||
model="gpt-5.4",
|
||||
model="gpt-5.5",
|
||||
)
|
||||
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
instructions=prompt_with_handoff_instructions(
|
||||
"You're speaking to a human, so be polite and concise. If the user speaks in Spanish, handoff to the spanish agent.",
|
||||
"You're speaking to a human, so be polite and concise. If the user speaks in Spanish, hand off to the Spanish agent.",
|
||||
),
|
||||
model="gpt-5.4",
|
||||
model="gpt-5.5",
|
||||
handoffs=[spanish_agent],
|
||||
tools=[get_weather],
|
||||
)
|
||||
@@ -92,14 +92,14 @@ agent = Agent(
|
||||
|
||||
## 音声パイプライン
|
||||
|
||||
ワークフローとして [`SingleAgentVoiceWorkflow`][agents.voice.workflow.SingleAgentVoiceWorkflow] を使い、シンプルな音声パイプラインをセットアップします。
|
||||
ワークフローとして [`SingleAgentVoiceWorkflow`][agents.voice.workflow.SingleAgentVoiceWorkflow] を使用して、シンプルな音声パイプラインをセットアップします。
|
||||
|
||||
```python
|
||||
from agents.voice import SingleAgentVoiceWorkflow, VoicePipeline
|
||||
pipeline = VoicePipeline(workflow=SingleAgentVoiceWorkflow(agent))
|
||||
```
|
||||
|
||||
## パイプライン実行
|
||||
## パイプラインの実行
|
||||
|
||||
```python
|
||||
import numpy as np
|
||||
@@ -156,19 +156,19 @@ def get_weather(city: str) -> str:
|
||||
|
||||
spanish_agent = Agent(
|
||||
name="Spanish",
|
||||
handoff_description="A spanish speaking agent.",
|
||||
handoff_description="A Spanish-speaking agent.",
|
||||
instructions=prompt_with_handoff_instructions(
|
||||
"You're speaking to a human, so be polite and concise. Speak in Spanish.",
|
||||
),
|
||||
model="gpt-5.4",
|
||||
model="gpt-5.5",
|
||||
)
|
||||
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
instructions=prompt_with_handoff_instructions(
|
||||
"You're speaking to a human, so be polite and concise. If the user speaks in Spanish, handoff to the spanish agent.",
|
||||
"You're speaking to a human, so be polite and concise. If the user speaks in Spanish, hand off to the Spanish agent.",
|
||||
),
|
||||
model="gpt-5.4",
|
||||
model="gpt-5.5",
|
||||
handoffs=[spanish_agent],
|
||||
tools=[get_weather],
|
||||
)
|
||||
@@ -195,4 +195,4 @@ if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
この example を実行すると、エージェントがあなたに話しかけます。[examples/voice/static](https://github.com/openai/openai-agents-python/tree/main/examples/voice/static) の example では、自分でエージェントに話しかけられるデモを確認できます。
|
||||
この例を実行すると、エージェントが話しかけてきます! エージェントに自分で話しかけられるデモについては、[examples/voice/static](https://github.com/openai/openai-agents-python/tree/main/examples/voice/static) の例を確認してください。
|
||||
@@ -4,15 +4,15 @@ search:
|
||||
---
|
||||
# トレーシング
|
||||
|
||||
[エージェントがトレーシングされる](../tracing.md)のと同様に、音声パイプラインも自動的にトレーシングされます。
|
||||
[エージェントのトレーシング](../tracing.md)と同様に、音声パイプラインも自動的にトレーシングされます。
|
||||
|
||||
基本的なトレーシング情報については上記のトレーシングドキュメントを参照できますが、[`VoicePipelineConfig`][agents.voice.pipeline_config.VoicePipelineConfig] を介してパイプラインのトレーシングを追加で設定することもできます。
|
||||
基本的なトレーシング情報については上記のトレーシングドキュメントを参照できますが、さらに [`VoicePipelineConfig`][agents.voice.pipeline_config.VoicePipelineConfig] を通じてパイプラインのトレーシングを設定できます。
|
||||
|
||||
トレーシングに関連する主要なフィールドは次のとおりです。
|
||||
トレーシングに関連する主なフィールドは次のとおりです。
|
||||
|
||||
- [`tracing_disabled`][agents.voice.pipeline_config.VoicePipelineConfig.tracing_disabled]: トレーシングを無効化するかどうかを制御します。デフォルトでは、トレーシングは有効です。
|
||||
- [`trace_include_sensitive_data`][agents.voice.pipeline_config.VoicePipelineConfig.trace_include_sensitive_data]: トレースに、音声文字起こしのような潜在的に機微なデータを含めるかどうかを制御します。これは音声パイプライン専用であり、Workflow 内で行われるものには適用されません。
|
||||
- [`tracing_disabled`][agents.voice.pipeline_config.VoicePipelineConfig.tracing_disabled]: トレーシングを無効にするかどうかを制御します。デフォルトでは、トレーシングは有効です。
|
||||
- [`trace_include_sensitive_data`][agents.voice.pipeline_config.VoicePipelineConfig.trace_include_sensitive_data]: トレースに音声文字起こしなど、潜在的に機微なデータを含めるかどうかを制御します。これは音声パイプライン専用であり、Workflow 内で行われる処理には適用されません。
|
||||
- [`trace_include_sensitive_audio_data`][agents.voice.pipeline_config.VoicePipelineConfig.trace_include_sensitive_audio_data]: トレースに音声データを含めるかどうかを制御します。
|
||||
- [`workflow_name`][agents.voice.pipeline_config.VoicePipelineConfig.workflow_name]: トレース Workflow の名前です。
|
||||
- [`group_id`][agents.voice.pipeline_config.VoicePipelineConfig.group_id]: トレースの `group_id` で、複数のトレースを関連付けられます。
|
||||
- [`workflow_name`][agents.voice.pipeline_config.VoicePipelineConfig.workflow_name]: トレースワークフローの名前です。
|
||||
- [`group_id`][agents.voice.pipeline_config.VoicePipelineConfig.group_id]: トレースの `group_id` で、複数のトレースを関連付けることができます。
|
||||
- [`trace_metadata`][agents.voice.pipeline_config.VoicePipelineConfig.trace_metadata]: トレースに含める追加のメタデータです。
|
||||
+79
-80
@@ -4,49 +4,49 @@ search:
|
||||
---
|
||||
# 에이전트
|
||||
|
||||
에이전트는 앱의 핵심 구성 요소입니다. 에이전트는 instructions, tools, 그리고 handoffs, 가드레일, structured outputs 같은 선택적 런타임 동작으로 구성된 대규모 언어 모델(LLM)입니다
|
||||
에이전트는 앱의 핵심 구성 요소입니다. 에이전트는 instructions, tools와 핸드오프, 가드레일, structured outputs 같은 선택적 런타임 동작으로 구성된 LLM입니다.
|
||||
|
||||
이 페이지는 단일 일반 `Agent`를 정의하거나 커스터마이즈하려는 경우에 사용합니다. 여러 에이전트가 어떻게 협업해야 하는지 결정하려면 [에이전트 오케스트레이션](multi_agent.md)을 읽어보세요. 에이전트가 manifest로 정의된 파일과 샌드박스 네이티브 기능을 갖춘 격리된 워크스페이스 내부에서 실행되어야 한다면 [Sandbox agent concepts](sandbox/guide.md)를 읽어보세요
|
||||
하나의 일반 `Agent`를 정의하거나 사용자 지정하려는 경우 이 페이지를 사용하세요. 여러 에이전트가 어떻게 협업해야 할지 결정하는 중이라면 [에이전트 오케스트레이션](multi_agent.md)을 읽어보세요. 에이전트가 매니페스트에 정의된 파일과 샌드박스 네이티브 기능을 갖춘 격리된 워크스페이스 내부에서 실행되어야 한다면 [Sandbox 에이전트 개념](sandbox/guide.md)을 읽어보세요.
|
||||
|
||||
SDK는 OpenAI 모델에 기본적으로 Responses API를 사용하지만, 여기서의 차이는 오케스트레이션입니다: `Agent`와 `Runner`를 함께 사용하면 SDK가 턴, 도구, 가드레일, 핸드오프, 세션을 대신 관리합니다. 이 루프를 직접 제어하고 싶다면 Responses API를 직접 사용하세요
|
||||
SDK는 OpenAI 모델에 기본적으로 Responses API를 사용하지만, 여기서의 차이는 오케스트레이션입니다. `Agent`와 `Runner`를 함께 사용하면 SDK가 턴, 도구, 가드레일, 핸드오프, 세션을 대신 관리해 줍니다. 해당 루프를 직접 제어하고 싶다면 대신 Responses API를 직접 사용하세요.
|
||||
|
||||
## 다음 가이드 선택
|
||||
|
||||
이 페이지를 에이전트 정의의 허브로 사용하세요. 다음에 내려야 할 결정에 맞는 인접 가이드로 이동하세요
|
||||
에이전트 정의의 허브로 이 페이지를 사용하세요. 다음에 내려야 할 결정에 맞는 인접 가이드로 이동하세요.
|
||||
|
||||
| 다음을 원한다면... | 다음 읽기 |
|
||||
| 원하는 작업 | 다음 읽기 |
|
||||
| --- | --- |
|
||||
| 모델 또는 provider 설정 선택 | [모델](models/index.md) |
|
||||
| 모델 또는 프로바이더 설정 선택 | [모델](models/index.md) |
|
||||
| 에이전트에 기능 추가 | [도구](tools.md) |
|
||||
| 실제 repo, 문서 번들 또는 격리된 워크스페이스에 대해 에이전트 실행 | [Sandbox agents quickstart](sandbox_agents.md) |
|
||||
| 관리자 스타일 오케스트레이션과 핸드오프 중 선택 | [에이전트 오케스트레이션](multi_agent.md) |
|
||||
| 실제 리포지토리, 문서 번들 또는 격리된 워크스페이스에서 에이전트 실행 | [Sandbox 에이전트 빠른 시작](sandbox_agents.md) |
|
||||
| 매니저 스타일 오케스트레이션과 핸드오프 중 선택 | [에이전트 오케스트레이션](multi_agent.md) |
|
||||
| 핸드오프 동작 구성 | [핸드오프](handoffs.md) |
|
||||
| 턴 실행, 이벤트 스트리밍, 대화 상태 관리 | [에이전트 실행](running_agents.md) |
|
||||
| 최종 출력, 실행 항목, 재개 가능한 상태 점검 | [결과](results.md) |
|
||||
| 로컬 의존성 및 런타임 상태 공유 | [컨텍스트 관리](context.md) |
|
||||
| 턴 실행, 이벤트 스트리밍 또는 대화 상태 관리 | [에이전트 실행](running_agents.md) |
|
||||
| 최종 출력, 실행 항목 또는 재개 가능한 상태 검사 | [결과](results.md) |
|
||||
| 로컬 의존성과 런타임 상태 공유 | [컨텍스트 관리](context.md) |
|
||||
|
||||
## 기본 구성
|
||||
|
||||
에이전트의 가장 일반적인 속성은 다음과 같습니다
|
||||
에이전트에서 가장 일반적인 속성은 다음과 같습니다.
|
||||
|
||||
| 속성 | 필수 | 설명 |
|
||||
| 속성 | 필수 여부 | 설명 |
|
||||
| --- | --- | --- |
|
||||
| `name` | 예 | 사람이 읽기 쉬운 에이전트 이름 |
|
||||
| `instructions` | 예 | 시스템 프롬프트 또는 동적 instructions 콜백. [동적 instructions](#dynamic-instructions) 참고 |
|
||||
| `prompt` | 아니요 | OpenAI Responses API 프롬프트 구성. 정적 프롬프트 객체 또는 함수를 허용합니다. [프롬프트 템플릿](#prompt-templates) 참고 |
|
||||
| `handoff_description` | 아니요 | 이 에이전트가 핸드오프 대상으로 제공될 때 노출되는 짧은 설명 |
|
||||
| `handoffs` | 아니요 | 대화를 전문 에이전트에 위임합니다. [handoffs](handoffs.md) 참고 |
|
||||
| `model` | 아니요 | 사용할 LLM. [모델](models/index.md) 참고 |
|
||||
| `model_settings` | 아니요 | `temperature`, `top_p`, `tool_choice` 같은 모델 튜닝 매개변수 |
|
||||
| `tools` | 아니요 | 에이전트가 호출할 수 있는 도구. [도구](tools.md) 참고 |
|
||||
| `mcp_servers` | 아니요 | 에이전트를 위한 MCP 기반 도구. [MCP 가이드](mcp.md) 참고 |
|
||||
| `mcp_config` | 아니요 | strict schema conversion, MCP failure formatting 등 MCP 도구 준비 방식을 세부 조정합니다. [MCP 가이드](mcp.md#agent-level-mcp-configuration) 참고 |
|
||||
| `input_guardrails` | 아니요 | 이 에이전트 체인의 첫 사용자 입력에서 실행되는 가드레일. [가드레일](guardrails.md) 참고 |
|
||||
| `output_guardrails` | 아니요 | 이 에이전트의 최종 출력에서 실행되는 가드레일. [가드레일](guardrails.md) 참고 |
|
||||
| `output_type` | 아니요 | 일반 텍스트 대신 structured output 타입. [출력 타입](#output-types) 참고 |
|
||||
| `hooks` | 아니요 | 에이전트 범위의 lifecycle 콜백. [라이프사이클 이벤트(hooks)](#lifecycle-events-hooks) 참고 |
|
||||
| `tool_use_behavior` | 아니요 | 도구 결과를 모델로 다시 보낼지 실행을 종료할지 제어합니다. [도구 사용 동작](#tool-use-behavior) 참고 |
|
||||
| `reset_tool_choice` | 아니요 | 도구 호출 후 `tool_choice`를 재설정(기본값: `True`)하여 도구 사용 루프를 방지합니다. [도구 사용 강제](#forcing-tool-use) 참고 |
|
||||
| `name` | 예 | 사람이 읽을 수 있는 에이전트 이름입니다. |
|
||||
| `instructions` | 아니요 | 시스템 프롬프트 또는 동적 instructions 콜백입니다. 강력히 권장됩니다. [동적 instructions](#dynamic-instructions)를 참고하세요. |
|
||||
| `prompt` | 아니요 | OpenAI Responses API 프롬프트 구성입니다. 정적 프롬프트 객체 또는 함수를 허용합니다. [프롬프트 템플릿](#prompt-templates)을 참고하세요. |
|
||||
| `handoff_description` | 아니요 | 이 에이전트가 핸드오프 대상으로 제공될 때 노출되는 짧은 설명입니다. |
|
||||
| `handoffs` | 아니요 | 대화를 전문 에이전트에 위임합니다. [핸드오프](handoffs.md)를 참고하세요. |
|
||||
| `model` | 아니요 | 사용할 LLM입니다. [모델](models/index.md)을 참고하세요. |
|
||||
| `model_settings` | 아니요 | `temperature`, `top_p`, `tool_choice` 같은 모델 튜닝 매개변수입니다. |
|
||||
| `tools` | 아니요 | 에이전트가 호출할 수 있는 도구입니다. [도구](tools.md)를 참고하세요. |
|
||||
| `mcp_servers` | 아니요 | 에이전트를 위한 MCP 기반 도구입니다. [MCP 가이드](mcp.md)를 참고하세요. |
|
||||
| `mcp_config` | 아니요 | 엄격한 스키마 변환 및 MCP 실패 형식화 등 MCP 도구가 준비되는 방식을 세부 조정합니다. [MCP 가이드](mcp.md#agent-level-mcp-configuration)를 참고하세요. |
|
||||
| `input_guardrails` | 아니요 | 이 에이전트 체인의 첫 번째 사용자 입력에 대해 실행되는 가드레일입니다. [가드레일](guardrails.md)을 참고하세요. |
|
||||
| `output_guardrails` | 아니요 | 이 에이전트의 최종 출력에 대해 실행되는 가드레일입니다. [가드레일](guardrails.md)을 참고하세요. |
|
||||
| `output_type` | 아니요 | 일반 텍스트 대신 structured outputs 타입입니다. [출력 타입](#output-types)을 참고하세요. |
|
||||
| `hooks` | 아니요 | 에이전트 범위 생명주기 콜백입니다. [생명주기 이벤트(훅)](#lifecycle-events-hooks)를 참고하세요. |
|
||||
| `tool_use_behavior` | 아니요 | 도구 결과를 모델로 다시 전달할지, 실행을 종료할지 제어합니다. [도구 사용 동작](#tool-use-behavior)을 참고하세요. |
|
||||
| `reset_tool_choice` | 아니요 | 도구 호출 후 `tool_choice`를 재설정(기본값: `True`)하여 도구 사용 루프를 방지합니다. [도구 사용 강제](#forcing-tool-use)를 참고하세요. |
|
||||
|
||||
```python
|
||||
from agents import Agent, ModelSettings, function_tool
|
||||
@@ -64,23 +64,23 @@ agent = Agent(
|
||||
)
|
||||
```
|
||||
|
||||
이 섹션의 모든 내용은 `Agent`에 적용됩니다. `SandboxAgent`는 같은 개념을 기반으로 하고, 워크스페이스 범위 실행을 위해 `default_manifest`, `base_instructions`, `capabilities`, `run_as`를 추가합니다. [Sandbox agent concepts](sandbox/guide.md) 참고
|
||||
`Agent`에는 이 섹션의 모든 내용이 적용됩니다. `SandboxAgent`는 같은 아이디어를 기반으로 하며, 여기에 워크스페이스 범위 실행을 위한 `default_manifest`, `base_instructions`, `capabilities`, `run_as`를 추가합니다. [Sandbox 에이전트 개념](sandbox/guide.md)을 참고하세요.
|
||||
|
||||
## 프롬프트 템플릿
|
||||
|
||||
`prompt`를 설정하여 OpenAI 플랫폼에서 생성한 프롬프트 템플릿을 참조할 수 있습니다. 이는 Responses API를 사용하는 OpenAI 모델에서 동작합니다
|
||||
`prompt`를 설정하여 OpenAI 플랫폼에서 생성한 프롬프트 템플릿을 참조할 수 있습니다. 이는 Responses API를 사용하는 OpenAI 모델에서 작동합니다.
|
||||
|
||||
사용 방법은 다음과 같습니다:
|
||||
사용하려면 다음을 수행하세요.
|
||||
|
||||
1. https://platform.openai.com/playground/prompts 로 이동합니다
|
||||
2. 새 프롬프트 변수 `poem_style`를 생성합니다
|
||||
3. 다음 내용으로 시스템 프롬프트를 생성합니다:
|
||||
1. https://platform.openai.com/playground/prompts로 이동합니다.
|
||||
2. 새 프롬프트 변수 `poem_style`를 생성합니다.
|
||||
3. 다음 내용으로 시스템 프롬프트를 생성합니다.
|
||||
|
||||
```
|
||||
Write a poem in {{poem_style}}
|
||||
```
|
||||
|
||||
4. `--prompt-id` 플래그로 예제를 실행합니다
|
||||
4. `--prompt-id` 플래그로 예제를 실행합니다.
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
@@ -95,7 +95,7 @@ agent = Agent(
|
||||
)
|
||||
```
|
||||
|
||||
실행 시점에 프롬프트를 동적으로 생성할 수도 있습니다:
|
||||
런타임에 프롬프트를 동적으로 생성할 수도 있습니다.
|
||||
|
||||
```python
|
||||
from dataclasses import dataclass
|
||||
@@ -127,9 +127,9 @@ result = await Runner.run(
|
||||
|
||||
## 컨텍스트
|
||||
|
||||
에이전트는 `context` 타입에 대해 제네릭입니다. 컨텍스트는 의존성 주입 도구입니다: 사용자가 생성해 `Runner.run()`에 전달하는 객체로, 모든 에이전트, 도구, 핸드오프 등에 전달되며 에이전트 실행에 필요한 의존성과 상태를 담는 바구니 역할을 합니다. 컨텍스트로는 어떤 Python 객체든 제공할 수 있습니다
|
||||
에이전트는 `context` 타입에 대해 제네릭입니다. 컨텍스트는 의존성 주입 도구입니다. 사용자가 생성해 `Runner.run()`에 전달하는 객체이며, 모든 에이전트, 도구, 핸드오프 등에 전달되고 에이전트 실행에 필요한 의존성과 상태를 담는 용도로 사용됩니다. 어떤 Python 객체든 컨텍스트로 제공할 수 있습니다.
|
||||
|
||||
전체 `RunContextWrapper` 표면, 공유 사용량 추적, 중첩 `tool_input`, 직렬화 관련 주의사항은 [컨텍스트 가이드](context.md)를 읽어보세요
|
||||
전체 `RunContextWrapper` 인터페이스, 공유 사용량 추적, 중첩된 `tool_input`, 직렬화 시 주의사항은 [컨텍스트 가이드](context.md)를 읽어보세요.
|
||||
|
||||
```python
|
||||
@dataclass
|
||||
@@ -148,7 +148,7 @@ agent = Agent[UserContext](
|
||||
|
||||
## 출력 타입
|
||||
|
||||
기본적으로 에이전트는 일반 텍스트(즉 `str`) 출력을 생성합니다. 에이전트가 특정 타입의 출력을 생성하게 하려면 `output_type` 매개변수를 사용할 수 있습니다. 일반적인 선택지는 [Pydantic](https://docs.pydantic.dev/) 객체지만, Pydantic [TypeAdapter](https://docs.pydantic.dev/latest/api/type_adapter/)로 래핑할 수 있는 타입이라면 모두 지원합니다 - dataclasses, lists, TypedDict 등
|
||||
기본적으로 에이전트는 일반 텍스트(즉, `str`) 출력을 생성합니다. 에이전트가 특정 출력 타입을 생성하도록 하려면 `output_type` 매개변수를 사용할 수 있습니다. 일반적으로 [Pydantic](https://docs.pydantic.dev/) 객체를 사용하는 경우가 많지만, Pydantic [TypeAdapter](https://docs.pydantic.dev/latest/api/type_adapter/)로 감쌀 수 있는 모든 타입(dataclasses, lists, TypedDict 등)을 지원합니다.
|
||||
|
||||
```python
|
||||
from pydantic import BaseModel
|
||||
@@ -169,20 +169,20 @@ agent = Agent(
|
||||
|
||||
!!! note
|
||||
|
||||
`output_type`를 전달하면, 모델은 일반 일반 텍스트 응답 대신 [structured outputs](https://platform.openai.com/docs/guides/structured-outputs)를 사용합니다
|
||||
`output_type`를 전달하면, 이는 모델에 일반 텍스트 응답 대신 [structured outputs](https://platform.openai.com/docs/guides/structured-outputs)를 사용하라고 지시합니다.
|
||||
|
||||
## 멀티 에이전트 시스템 설계 패턴
|
||||
|
||||
멀티 에이전트 시스템을 설계하는 방법은 많지만, 일반적으로 널리 적용 가능한 두 가지 패턴이 자주 사용됩니다:
|
||||
멀티 에이전트 시스템을 설계하는 방법은 많지만, 일반적으로 다음 두 가지 광범위하게 적용 가능한 패턴을 자주 봅니다.
|
||||
|
||||
1. 매니저(Agents as tools): 중앙 매니저/오케스트레이터가 전문 하위 에이전트를 도구로 호출하고 대화 제어를 유지합니다
|
||||
2. 핸드오프: 동등한 에이전트가 대화를 인계받아 처리할 전문 에이전트로 제어를 넘깁니다. 이는 분산형 패턴입니다
|
||||
1. 매니저 (agents as tools): 중앙 매니저/오케스트레이터가 전문 하위 에이전트를 도구로 호출하고 대화 제어권을 유지합니다.
|
||||
2. 핸드오프: 피어 에이전트가 대화의 제어권을 넘겨받을 전문 에이전트에게 제어권을 핸드오프합니다. 이는 분산형 방식입니다.
|
||||
|
||||
자세한 내용은 [our practical guide to building agents](https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf)를 참고하세요
|
||||
자세한 내용은 [에이전트 구축 실전 가이드](https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf)를 참고하세요.
|
||||
|
||||
### 매니저(Agents as tools)
|
||||
### 매니저 (agents as tools)
|
||||
|
||||
`customer_facing_agent`는 모든 사용자 상호작용을 처리하고, 도구로 노출된 전문 하위 에이전트를 호출합니다. 자세한 내용은 [tools](tools.md#agents-as-tools) 문서를 참고하세요
|
||||
`customer_facing_agent`는 모든 사용자 상호작용을 처리하고, 도구로 노출된 전문 하위 에이전트를 호출합니다. 자세한 내용은 [도구](tools.md#agents-as-tools) 문서를 읽어보세요.
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
@@ -211,7 +211,7 @@ customer_facing_agent = Agent(
|
||||
|
||||
### 핸드오프
|
||||
|
||||
핸드오프는 에이전트가 위임할 수 있는 하위 에이전트입니다. 핸드오프가 발생하면 위임된 에이전트가 대화 기록을 전달받아 대화를 이어받습니다. 이 패턴은 단일 작업에 뛰어난 모듈형 전문 에이전트를 가능하게 합니다. 자세한 내용은 [handoffs](handoffs.md) 문서를 참고하세요
|
||||
핸드오프는 에이전트가 위임할 수 있는 하위 에이전트입니다. 핸드오프가 발생하면 위임된 에이전트가 대화 기록을 받아 대화를 이어받습니다. 이 패턴을 통해 하나의 작업에 뛰어난 모듈식 전문 에이전트를 만들 수 있습니다. 자세한 내용은 [핸드오프](handoffs.md) 문서를 읽어보세요.
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
@@ -232,7 +232,7 @@ triage_agent = Agent(
|
||||
|
||||
## 동적 instructions
|
||||
|
||||
대부분의 경우 에이전트를 생성할 때 instructions를 제공하면 됩니다. 하지만 함수를 통해 동적 instructions를 제공할 수도 있습니다. 함수는 에이전트와 컨텍스트를 전달받고 프롬프트를 반환해야 합니다. 일반 함수와 `async` 함수 모두 허용됩니다
|
||||
대부분의 경우 에이전트를 생성할 때 instructions를 제공할 수 있습니다. 그러나 함수를 통해 동적 instructions를 제공할 수도 있습니다. 이 함수는 에이전트와 컨텍스트를 받고, 프롬프트를 반환해야 합니다. 일반 함수와 `async` 함수 모두 허용됩니다.
|
||||
|
||||
```python
|
||||
def dynamic_instructions(
|
||||
@@ -247,29 +247,28 @@ agent = Agent[UserContext](
|
||||
)
|
||||
```
|
||||
|
||||
## 라이프사이클 이벤트(hooks)
|
||||
## 생명주기 이벤트(훅)
|
||||
|
||||
때로는 에이전트의 라이프사이클을 관찰하고 싶을 수 있습니다. 예를 들어 이벤트를 로깅하거나, 데이터를 사전 로드하거나, 특정 이벤트 발생 시 사용량을 기록할 수 있습니다
|
||||
때로는 에이전트의 생명주기를 관찰하고 싶을 수 있습니다. 예를 들어 특정 이벤트가 발생할 때 이벤트를 기록하거나, 데이터를 미리 가져오거나, 사용량을 기록할 수 있습니다.
|
||||
|
||||
hook 범위는 두 가지입니다:
|
||||
훅 범위는 두 가지입니다.
|
||||
|
||||
- [`RunHooks`][agents.lifecycle.RunHooks]는 다른 에이전트로의 핸드오프를 포함해 전체 `Runner.run(...)` 호출을 관찰합니다
|
||||
- [`AgentHooks`][agents.lifecycle.AgentHooks]는 `agent.hooks`를 통해 특정 에이전트 인스턴스에 연결됩니다
|
||||
- [`RunHooks`][agents.lifecycle.RunHooks]는 다른 에이전트로의 핸드오프를 포함해 전체 `Runner.run(...)` 호출을 관찰합니다.
|
||||
- [`AgentHooks`][agents.lifecycle.AgentHooks]는 `agent.hooks`를 통해 특정 에이전트 인스턴스에 연결됩니다.
|
||||
|
||||
이벤트에 따라 콜백 컨텍스트도 달라집니다:
|
||||
콜백 컨텍스트도 이벤트에 따라 달라집니다.
|
||||
|
||||
- 에이전트 시작/종료 hook은 [`AgentHookContext`][agents.run_context.AgentHookContext]를 받으며, 이는 원래 컨텍스트를 래핑하고 공유 실행 사용량 상태를 포함합니다
|
||||
- LLM, 도구, 핸드오프 hook은 [`RunContextWrapper`][agents.run_context.RunContextWrapper]를 받습니다
|
||||
- 에이전트 시작/종료 훅은 [`AgentHookContext`][agents.run_context.AgentHookContext]를 받으며, 이는 원래 컨텍스트를 감싸고 공유 실행 사용량 상태를 포함합니다.
|
||||
- LLM, 도구, 핸드오프 훅은 [`RunContextWrapper`][agents.run_context.RunContextWrapper]를 받습니다.
|
||||
|
||||
일반적인 hook 타이밍:
|
||||
일반적인 훅 호출 시점은 다음과 같습니다.
|
||||
|
||||
- `on_agent_start` / `on_agent_end`: 특정 에이전트가 최종 출력을 생성하기 시작하거나 끝낼 때
|
||||
- `on_llm_start` / `on_llm_end`: 각 모델 호출 직전/직후
|
||||
- `on_tool_start` / `on_tool_end`: 각 로컬 도구 호출 전후
|
||||
함수 도구의 경우 hook `context`는 보통 `ToolContext`이므로 `tool_call_id` 같은 도구 호출 메타데이터를 확인할 수 있습니다
|
||||
- `on_handoff`: 제어가 한 에이전트에서 다른 에이전트로 이동할 때
|
||||
- `on_agent_start` / `on_agent_end`: 특정 에이전트가 최종 출력을 생성하기 시작하거나 완료할 때입니다.
|
||||
- `on_llm_start` / `on_llm_end`: 각 모델 호출 직전과 직후입니다.
|
||||
- `on_tool_start` / `on_tool_end`: 각 로컬 도구 호출 전후입니다. 함수 도구의 경우 훅 `context`는 일반적으로 `ToolContext`이므로 `tool_call_id` 같은 도구 호출 메타데이터를 검사할 수 있습니다.
|
||||
- `on_handoff`: 제어권이 한 에이전트에서 다른 에이전트로 이동할 때입니다.
|
||||
|
||||
전체 워크플로에 대해 단일 관찰자가 필요하면 `RunHooks`를, 하나의 에이전트에 맞춤 부수 효과가 필요하면 `AgentHooks`를 사용하세요
|
||||
`RunHooks`는 전체 워크플로에 대한 단일 관찰자가 필요할 때 사용하고, `AgentHooks`는 특정 에이전트에 사용자 지정 부수 효과가 필요할 때 사용하세요.
|
||||
|
||||
```python
|
||||
from agents import Agent, RunHooks, Runner
|
||||
@@ -291,21 +290,21 @@ result = await Runner.run(agent, "Explain quines", hooks=LoggingHooks())
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
전체 콜백 표면은 [Lifecycle API reference](ref/lifecycle.md)를 참고하세요
|
||||
전체 콜백 인터페이스는 [생명주기 API 참조](ref/lifecycle.md)를 참고하세요.
|
||||
|
||||
## 가드레일
|
||||
|
||||
가드레일을 사용하면 에이전트 실행과 병렬로 사용자 입력에 대한 검사/검증을 수행하고, 에이전트 출력이 생성된 뒤 해당 출력에 대해서도 검사/검증을 수행할 수 있습니다. 예를 들어 사용자 입력과 에이전트 출력의 관련성을 점검할 수 있습니다. 자세한 내용은 [guardrails](guardrails.md) 문서를 참고하세요
|
||||
가드레일을 사용하면 에이전트 실행과 병렬로 사용자 입력에 대한 검사/검증을 실행하고, 에이전트 출력이 생성된 뒤 그 출력에 대해서도 검사/검증을 실행할 수 있습니다. 예를 들어 사용자 입력과 에이전트 출력의 관련성을 선별할 수 있습니다. 자세한 내용은 [가드레일](guardrails.md) 문서를 참고하세요.
|
||||
|
||||
## 에이전트 복제/복사
|
||||
|
||||
에이전트에서 `clone()` 메서드를 사용하면 Agent를 복제하고, 원하면 어떤 속성이든 변경할 수 있습니다
|
||||
에이전트에서 `clone()` 메서드를 사용하면 에이전트를 복제하고, 선택적으로 원하는 속성을 변경할 수 있습니다.
|
||||
|
||||
```python
|
||||
pirate_agent = Agent(
|
||||
name="Pirate",
|
||||
instructions="Write like a pirate",
|
||||
model="gpt-5.4",
|
||||
model="gpt-5.5",
|
||||
)
|
||||
|
||||
robot_agent = pirate_agent.clone(
|
||||
@@ -316,14 +315,14 @@ robot_agent = pirate_agent.clone(
|
||||
|
||||
## 도구 사용 강제
|
||||
|
||||
도구 목록을 제공한다고 해서 항상 LLM이 도구를 사용하는 것은 아닙니다. [`ModelSettings.tool_choice`][agents.model_settings.ModelSettings.tool_choice]를 설정해 도구 사용을 강제할 수 있습니다. 유효한 값은 다음과 같습니다:
|
||||
도구 목록을 제공한다고 해서 항상 LLM이 도구를 사용한다는 의미는 아닙니다. [`ModelSettings.tool_choice`][agents.model_settings.ModelSettings.tool_choice]를 설정하여 도구 사용을 강제할 수 있습니다. 유효한 값은 다음과 같습니다.
|
||||
|
||||
1. `auto`: LLM이 도구 사용 여부를 결정하도록 허용
|
||||
2. `required`: LLM이 도구를 반드시 사용(단, 어떤 도구를 쓸지는 지능적으로 결정 가능)
|
||||
3. `none`: LLM이 도구를 _사용하지 않도록_ 강제
|
||||
4. 특정 문자열(예: `my_tool`) 설정: LLM이 해당 특정 도구를 사용하도록 강제
|
||||
1. `auto`: LLM이 도구를 사용할지 여부를 결정하도록 허용합니다.
|
||||
2. `required`: LLM이 도구를 사용하도록 요구합니다(하지만 어떤 도구를 사용할지는 지능적으로 결정할 수 있습니다).
|
||||
3. `none`: LLM이 도구를 _사용하지 않도록_ 요구합니다.
|
||||
4. `my_tool` 같은 특정 문자열을 설정하면 LLM이 해당 특정 도구를 사용하도록 요구합니다.
|
||||
|
||||
OpenAI Responses 도구 검색을 사용할 때는 이름 기반 도구 선택이 더 제한됩니다: `tool_choice`로는 네임스페이스 이름만 있는 도구나 deferred-only 도구를 지정할 수 없고, `tool_choice="tool_search"`는 [`ToolSearchTool`][agents.tool.ToolSearchTool]을 대상으로 하지 않습니다. 이런 경우 `auto` 또는 `required`를 권장합니다. Responses 전용 제약 사항은 [Hosted tool search](tools.md#hosted-tool-search)를 참고하세요
|
||||
OpenAI Responses 도구 검색을 사용하는 경우, 이름이 지정된 도구 선택에는 더 많은 제한이 있습니다. `tool_choice`로 단독 네임스페이스 이름이나 deferred-only 도구를 대상으로 지정할 수 없으며, `tool_choice="tool_search"`는 [`ToolSearchTool`][agents.tool.ToolSearchTool]를 대상으로 지정하지 않습니다. 이러한 경우에는 `auto` 또는 `required`를 선호하세요. Responses에 특화된 제약 사항은 [호스티드 툴 검색](tools.md#hosted-tool-search)을 참고하세요.
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, function_tool, ModelSettings
|
||||
@@ -343,10 +342,10 @@ agent = Agent(
|
||||
|
||||
## 도구 사용 동작
|
||||
|
||||
`Agent` 구성의 `tool_use_behavior` 매개변수는 도구 출력 처리 방식을 제어합니다:
|
||||
`Agent` 구성의 `tool_use_behavior` 매개변수는 도구 출력이 처리되는 방식을 제어합니다.
|
||||
|
||||
- `"run_llm_again"`: 기본값입니다. 도구를 실행하고 LLM이 결과를 처리해 최종 응답을 생성합니다
|
||||
- `"stop_on_first_tool"`: 추가 LLM 처리 없이 첫 번째 도구 호출의 출력을 최종 응답으로 사용합니다
|
||||
- `"run_llm_again"`: 기본값입니다. 도구가 실행되고, LLM이 결과를 처리하여 최종 응답을 생성합니다.
|
||||
- `"stop_on_first_tool"`: 첫 번째 도구 호출의 출력이 추가 LLM 처리 없이 최종 응답으로 사용됩니다.
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, function_tool, ModelSettings
|
||||
@@ -364,7 +363,7 @@ agent = Agent(
|
||||
)
|
||||
```
|
||||
|
||||
- `StopAtTools(stop_at_tool_names=[...])`: 지정된 도구 중 하나라도 호출되면 중지하고 해당 출력을 최종 응답으로 사용합니다
|
||||
- `StopAtTools(stop_at_tool_names=[...])`: 지정된 도구 중 하나가 호출되면 중지하고, 해당 출력을 최종 응답으로 사용합니다.
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, function_tool
|
||||
@@ -388,7 +387,7 @@ agent = Agent(
|
||||
)
|
||||
```
|
||||
|
||||
- `ToolsToFinalOutputFunction`: 도구 결과를 처리하고 LLM으로 계속할지 중지할지 결정하는 사용자 정의 함수입니다
|
||||
- `ToolsToFinalOutputFunction`: 도구 결과를 처리하고 LLM으로 계속 진행할지 중지할지 결정하는 사용자 지정 함수입니다.
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, function_tool, FunctionToolResult, RunContextWrapper
|
||||
@@ -426,4 +425,4 @@ agent = Agent(
|
||||
|
||||
!!! note
|
||||
|
||||
무한 루프를 방지하기 위해 프레임워크는 도구 호출 후 `tool_choice`를 자동으로 "auto"로 재설정합니다. 이 동작은 [`agent.reset_tool_choice`][agents.agent.Agent.reset_tool_choice]로 구성할 수 있습니다. 무한 루프가 생기는 이유는 도구 결과가 LLM으로 전달되고, LLM이 `tool_choice` 때문에 또 다른 도구 호출을 생성하는 과정이 무한 반복되기 때문입니다
|
||||
무한 루프를 방지하기 위해 프레임워크는 도구 호출 후 `tool_choice`를 자동으로 "auto"로 재설정합니다. 이 동작은 [`agent.reset_tool_choice`][agents.agent.Agent.reset_tool_choice]를 통해 구성할 수 있습니다. 무한 루프는 도구 결과가 LLM으로 전송되고, 이후 LLM이 `tool_choice` 때문에 또 다른 도구 호출을 생성하는 일이 끝없이 반복되기 때문에 발생합니다.
|
||||
+57
-25
@@ -4,21 +4,21 @@ search:
|
||||
---
|
||||
# 구성
|
||||
|
||||
이 페이지에서는 기본 OpenAI 키 또는 클라이언트, 기본 OpenAI API 형태, 트레이싱 내보내기 기본값, 로깅 동작처럼 보통 애플리케이션 시작 시 한 번 설정하는 SDK 전역 기본값을 다룹니다
|
||||
이 페이지에서는 애플리케이션 시작 시 일반적으로 한 번 설정하는 SDK 전체 기본값을 다룹니다. 예를 들어 기본 OpenAI 키 또는 클라이언트, 기본 OpenAI API 형태, 트레이싱 내보내기 기본값, 로깅 동작 등이 있습니다.
|
||||
|
||||
이러한 기본값은 샌드박스 기반 워크플로에도 계속 적용되지만, 샌드박스 워크스페이스, 샌드박스 클라이언트, 세션 재사용은 별도로 구성합니다
|
||||
이러한 기본값은 샌드박스 기반 워크플로에도 계속 적용되지만, 샌드박스 워크스페이스, 샌드박스 클라이언트, 세션 재사용은 별도로 구성합니다.
|
||||
|
||||
대신 특정 에이전트 또는 실행을 구성해야 한다면, 다음부터 시작하세요:
|
||||
대신 특정 에이전트 또는 실행을 구성해야 한다면 다음부터 시작하세요:
|
||||
|
||||
- 일반 `Agent`의 instructions, tools, 출력 타입, 핸드오프, 가드레일은 [Agents](agents.md)
|
||||
- `RunConfig`, 세션, 대화 상태 옵션은 [에이전트 실행](running_agents.md)
|
||||
- `SandboxRunConfig`, 매니페스트, 기능, 샌드박스 클라이언트 전용 워크스페이스 설정은 [샌드박스 에이전트](sandbox/guide.md)
|
||||
- 모델 선택 및 프로바이더 구성은 [모델](models/index.md)
|
||||
- 실행별 트레이싱 메타데이터와 사용자 지정 트레이스 프로세서는 [트레이싱](tracing.md)
|
||||
- [에이전트](agents.md): 일반 `Agent`의 instructions, tools, 출력 유형, 핸드오프, 가드레일
|
||||
- [에이전트 실행](running_agents.md): `RunConfig`, 세션, 대화 상태 옵션
|
||||
- [샌드박스 에이전트](sandbox/guide.md): `SandboxRunConfig`, 매니페스트, 기능, 샌드박스 클라이언트별 워크스페이스 설정
|
||||
- [모델](models/index.md): 모델 선택 및 제공자 구성
|
||||
- [트레이싱](tracing.md): 실행별 트레이싱 메타데이터 및 사용자 지정 트레이스 프로세서
|
||||
|
||||
## API 키와 클라이언트
|
||||
## API 키 및 클라이언트
|
||||
|
||||
기본적으로 SDK는 LLM 요청과 트레이싱에 `OPENAI_API_KEY` 환경 변수를 사용합니다. 이 키는 SDK가 처음 OpenAI 클라이언트를 생성할 때(지연 초기화) 확인되므로, 첫 모델 호출 전에 환경 변수를 설정하세요. 앱 시작 전에 해당 환경 변수를 설정할 수 없다면 [set_default_openai_key()][agents.set_default_openai_key] 함수를 사용해 키를 설정할 수 있습니다.
|
||||
기본적으로 SDK는 LLM 요청과 트레이싱에 `OPENAI_API_KEY` 환경 변수를 사용합니다. 키는 SDK가 처음으로 OpenAI 클라이언트를 생성할 때 확인됩니다(지연 초기화). 따라서 첫 모델 호출 전에 환경 변수를 설정하세요. 앱 시작 전에 해당 환경 변수를 설정할 수 없다면 [set_default_openai_key()][agents.set_default_openai_key] 함수를 사용해 키를 설정할 수 있습니다.
|
||||
|
||||
```python
|
||||
from agents import set_default_openai_key
|
||||
@@ -26,7 +26,7 @@ from agents import set_default_openai_key
|
||||
set_default_openai_key("sk-...")
|
||||
```
|
||||
|
||||
또는 사용할 OpenAI 클라이언트를 구성할 수도 있습니다. 기본적으로 SDK는 환경 변수의 API 키 또는 위에서 설정한 기본 키를 사용해 `AsyncOpenAI` 인스턴스를 생성합니다. [set_default_openai_client()][agents.set_default_openai_client] 함수를 사용해 이를 변경할 수 있습니다.
|
||||
또는 사용할 OpenAI 클라이언트를 구성할 수도 있습니다. 기본적으로 SDK는 환경 변수의 API 키나 위에서 설정한 기본 키를 사용하여 `AsyncOpenAI` 인스턴스를 생성합니다. [set_default_openai_client()][agents.set_default_openai_client] 함수를 사용해 이를 변경할 수 있습니다.
|
||||
|
||||
```python
|
||||
from openai import AsyncOpenAI
|
||||
@@ -36,14 +36,14 @@ custom_client = AsyncOpenAI(base_url="...", api_key="...")
|
||||
set_default_openai_client(custom_client)
|
||||
```
|
||||
|
||||
환경 기반 엔드포인트 구성을 선호한다면, 기본 OpenAI 프로바이더는 `OPENAI_BASE_URL`도 읽습니다. Responses websocket 전송을 활성화하면 websocket `/responses` 엔드포인트에 `OPENAI_WEBSOCKET_BASE_URL`도 읽습니다.
|
||||
환경 변수 기반 엔드포인트 구성을 선호한다면, 기본 OpenAI 제공자는 `OPENAI_BASE_URL`도 읽습니다. Responses WebSocket 전송을 활성화하면 WebSocket `/responses` 엔드포인트용 `OPENAI_WEBSOCKET_BASE_URL`도 읽습니다.
|
||||
|
||||
```bash
|
||||
export OPENAI_BASE_URL="https://your-openai-compatible-endpoint.example/v1"
|
||||
export OPENAI_WEBSOCKET_BASE_URL="wss://your-openai-compatible-endpoint.example/v1"
|
||||
```
|
||||
|
||||
마지막으로, 사용되는 OpenAI API를 사용자 지정할 수도 있습니다. 기본적으로는 OpenAI Responses API를 사용합니다. [set_default_openai_api()][agents.set_default_openai_api] 함수를 사용하면 이를 재정의해 Chat Completions API를 사용할 수 있습니다.
|
||||
마지막으로 사용할 OpenAI API를 사용자 지정할 수도 있습니다. 기본적으로 OpenAI Responses API를 사용합니다. [set_default_openai_api()][agents.set_default_openai_api] 함수를 사용해 이를 재정의하여 Chat Completions API를 사용할 수 있습니다.
|
||||
|
||||
```python
|
||||
from agents import set_default_openai_api
|
||||
@@ -51,9 +51,41 @@ from agents import set_default_openai_api
|
||||
set_default_openai_api("chat_completions")
|
||||
```
|
||||
|
||||
## OpenAI 제공자 기본값
|
||||
|
||||
OpenAI 기반 제공자는 모델 이름을 확인할 때도 SDK 전체 기본값을 읽습니다. OpenAI Responses 모델이 기본적으로 WebSocket 전송을 사용하도록 하려면 [`set_default_openai_responses_transport()`][agents.set_default_openai_responses_transport]를 사용하세요:
|
||||
|
||||
```python
|
||||
from agents import set_default_openai_responses_transport
|
||||
|
||||
set_default_openai_responses_transport("websocket")
|
||||
```
|
||||
|
||||
이는 기본 OpenAI 제공자가 확인한 OpenAI Responses 모델에 영향을 줍니다. 제공자 수준 설정, 연결 재사용, keepalive 옵션, 사용자 지정 WebSocket 엔드포인트는 [Responses WebSocket 전송](models/index.md#responses-websocket-transport)을 참조하세요.
|
||||
|
||||
OpenAI 설정에서 제공자 수준 에이전트 등록 메타데이터를 기대하는 경우, 시작 시 기본 harness ID를 한 번 구성하세요:
|
||||
|
||||
```python
|
||||
from agents import set_default_openai_harness
|
||||
|
||||
set_default_openai_harness("your-harness-id")
|
||||
```
|
||||
|
||||
전체 등록 객체를 전달할 수도 있습니다:
|
||||
|
||||
```python
|
||||
from agents import OpenAIAgentRegistrationConfig, set_default_openai_agent_registration
|
||||
|
||||
set_default_openai_agent_registration(
|
||||
OpenAIAgentRegistrationConfig(harness_id="your-harness-id")
|
||||
)
|
||||
```
|
||||
|
||||
SDK 기본값이 설정되어 있지 않으면 OpenAI 기반 제공자는 `OPENAI_AGENT_HARNESS_ID` 환경 변수로 폴백합니다. harness ID가 구성되어 있으면, SDK는 `RunConfig.trace_metadata`에 해당 키가 이미 있는 경우를 제외하고 이를 `agent_harness_id`로 트레이스 메타데이터에 추가합니다.
|
||||
|
||||
## 트레이싱
|
||||
|
||||
트레이싱은 기본적으로 활성화되어 있습니다. 기본적으로 위 섹션의 모델 요청과 동일한 OpenAI API 키(즉, 환경 변수 또는 설정한 기본 키)를 사용합니다. [`set_tracing_export_api_key`][agents.set_tracing_export_api_key] 함수를 사용해 트레이싱에 사용할 API 키를 별도로 설정할 수 있습니다.
|
||||
트레이싱은 기본적으로 활성화되어 있습니다. 기본적으로 위 섹션의 모델 요청과 동일한 OpenAI API 키(즉, 환경 변수 또는 설정한 기본 키)를 사용합니다. [`set_tracing_export_api_key`][agents.set_tracing_export_api_key] 함수를 사용하여 트레이싱에 사용할 API 키를 별도로 설정할 수 있습니다.
|
||||
|
||||
```python
|
||||
from agents import set_tracing_export_api_key
|
||||
@@ -61,7 +93,7 @@ from agents import set_tracing_export_api_key
|
||||
set_tracing_export_api_key("sk-...")
|
||||
```
|
||||
|
||||
모델 트래픽은 하나의 키 또는 클라이언트를 사용하지만 트레이싱은 다른 OpenAI 키를 사용해야 한다면, 기본 키 또는 클라이언트를 설정할 때 `use_for_tracing=False`를 전달한 다음 트레이싱을 별도로 구성하세요. 사용자 지정 클라이언트를 사용하지 않는 경우 [`set_default_openai_key()`][agents.set_default_openai_key]에도 같은 패턴을 적용할 수 있습니다.
|
||||
모델 트래픽에는 한 키나 클라이언트를 사용하지만 트레이싱에는 다른 OpenAI 키를 사용해야 한다면, 기본 키나 클라이언트를 설정할 때 `use_for_tracing=False`를 전달한 다음 트레이싱을 별도로 구성하세요. 사용자 지정 클라이언트를 사용하지 않는 경우 [`set_default_openai_key()`][agents.set_default_openai_key]에서도 동일한 패턴을 사용할 수 있습니다.
|
||||
|
||||
```python
|
||||
from openai import AsyncOpenAI
|
||||
@@ -76,14 +108,14 @@ set_default_openai_client(custom_client, use_for_tracing=False)
|
||||
set_tracing_export_api_key("sk-tracing")
|
||||
```
|
||||
|
||||
기본 내보내기를 사용할 때 트레이스를 특정 조직 또는 프로젝트에 귀속해야 한다면, 앱 시작 전에 다음 환경 변수를 설정하세요:
|
||||
기본 익스포터를 사용할 때 트레이스를 특정 조직 또는 프로젝트에 귀속해야 한다면 앱 시작 전에 다음 환경 변수를 설정하세요:
|
||||
|
||||
```bash
|
||||
export OPENAI_ORG_ID="org_..."
|
||||
export OPENAI_PROJECT_ID="proj_..."
|
||||
```
|
||||
|
||||
전역 내보내기를 변경하지 않고 실행별로 트레이싱 API 키를 설정할 수도 있습니다.
|
||||
전역 익스포터를 변경하지 않고 실행별로 트레이싱 API 키를 설정할 수도 있습니다.
|
||||
|
||||
```python
|
||||
from agents import Runner, RunConfig
|
||||
@@ -95,7 +127,7 @@ await Runner.run(
|
||||
)
|
||||
```
|
||||
|
||||
[`set_tracing_disabled()`][agents.set_tracing_disabled] 함수를 사용해 트레이싱을 완전히 비활성화할 수도 있습니다.
|
||||
[`set_tracing_disabled()`][agents.set_tracing_disabled] 함수를 사용하여 트레이싱을 완전히 비활성화할 수도 있습니다.
|
||||
|
||||
```python
|
||||
from agents import set_tracing_disabled
|
||||
@@ -103,7 +135,7 @@ from agents import set_tracing_disabled
|
||||
set_tracing_disabled(True)
|
||||
```
|
||||
|
||||
트레이싱은 활성화한 채로 유지하되 트레이스 페이로드에서 잠재적으로 민감한 입력/출력을 제외하려면 [`RunConfig.trace_include_sensitive_data`][agents.run.RunConfig.trace_include_sensitive_data]를 `False`로 설정하세요:
|
||||
트레이싱은 활성화한 상태로 유지하되 트레이스 페이로드에서 민감할 수 있는 입력/출력을 제외하려면 [`RunConfig.trace_include_sensitive_data`][agents.run.RunConfig.trace_include_sensitive_data]를 `False`로 설정하세요:
|
||||
|
||||
```python
|
||||
from agents import Runner, RunConfig
|
||||
@@ -115,17 +147,17 @@ await Runner.run(
|
||||
)
|
||||
```
|
||||
|
||||
코드 없이 기본값을 변경하려면 앱 시작 전에 이 환경 변수를 설정할 수도 있습니다:
|
||||
앱 시작 전에 다음 환경 변수를 설정하면 코드 없이도 기본값을 변경할 수 있습니다:
|
||||
|
||||
```bash
|
||||
export OPENAI_AGENTS_TRACE_INCLUDE_SENSITIVE_DATA=0
|
||||
```
|
||||
|
||||
전체 트레이싱 제어는 [트레이싱 가이드](tracing.md)를 참고하세요.
|
||||
전체 트레이싱 제어 옵션은 [트레이싱 가이드](tracing.md)를 참조하세요.
|
||||
|
||||
## 디버그 로깅
|
||||
|
||||
SDK는 두 개의 Python 로거(`openai.agents` 및 `openai.agents.tracing`)를 정의하며 기본적으로 핸들러를 연결하지 않습니다. 로그는 애플리케이션의 Python 로깅 구성 설정을 따릅니다.
|
||||
SDK는 두 개의 Python 로거(`openai.agents` 및 `openai.agents.tracing`)를 정의하며, 기본적으로 핸들러를 연결하지 않습니다. 로그는 애플리케이션의 Python 로깅 구성을 따릅니다.
|
||||
|
||||
상세 로깅을 활성화하려면 [`enable_verbose_stdout_logging()`][agents.enable_verbose_stdout_logging] 함수를 사용하세요.
|
||||
|
||||
@@ -135,7 +167,7 @@ from agents import enable_verbose_stdout_logging
|
||||
enable_verbose_stdout_logging()
|
||||
```
|
||||
|
||||
또는 핸들러, 필터, 포매터 등을 추가해 로그를 사용자 지정할 수 있습니다. 자세한 내용은 [Python 로깅 가이드](https://docs.python.org/3/howto/logging.html)를 참고하세요.
|
||||
또는 핸들러, 필터, 포매터 등을 추가하여 로그를 사용자 지정할 수 있습니다. 자세한 내용은 [Python 로깅 가이드](https://docs.python.org/3/howto/logging.html)를 참조하세요.
|
||||
|
||||
```python
|
||||
import logging
|
||||
@@ -158,14 +190,14 @@ logger.addHandler(logging.StreamHandler())
|
||||
|
||||
일부 로그에는 민감한 데이터(예: 사용자 데이터)가 포함될 수 있습니다.
|
||||
|
||||
기본적으로 SDK는 LLM 입력/출력이나 도구 입력/출력을 로깅하지 **않습니다**. 이러한 보호는 다음으로 제어됩니다:
|
||||
기본적으로 SDK는 LLM 입력/출력이나 도구 입력/출력을 로그로 기록하지 **않습니다**. 이러한 보호 기능은 다음으로 제어됩니다:
|
||||
|
||||
```bash
|
||||
OPENAI_AGENTS_DONT_LOG_MODEL_DATA=1
|
||||
OPENAI_AGENTS_DONT_LOG_TOOL_DATA=1
|
||||
```
|
||||
|
||||
디버깅을 위해 이 데이터를 일시적으로 포함해야 한다면 앱 시작 전에 변수 중 하나를 `0`(또는 `false`)으로 설정하세요:
|
||||
디버깅을 위해 이 데이터를 일시적으로 포함해야 한다면, 앱 시작 전에 둘 중 하나의 변수를 `0`(또는 `false`)으로 설정하세요:
|
||||
|
||||
```bash
|
||||
export OPENAI_AGENTS_DONT_LOG_MODEL_DATA=0
|
||||
|
||||
+43
-43
@@ -4,49 +4,49 @@ search:
|
||||
---
|
||||
# 컨텍스트 관리
|
||||
|
||||
컨텍스트는 중의적으로 사용되는 용어입니다. 보통 신경 써야 할 컨텍스트는 두 가지 주요 범주가 있습니다
|
||||
컨텍스트는 여러 의미로 쓰이는 용어입니다. 주로 고려해야 할 컨텍스트에는 두 가지 주요 유형이 있습니다.
|
||||
|
||||
1. 코드에서 로컬로 사용할 수 있는 컨텍스트: 도구 함수 실행 시, `on_handoff` 같은 콜백, 라이프사이클 훅 등에서 필요할 수 있는 데이터와 의존성입니다
|
||||
2. LLM에서 사용할 수 있는 컨텍스트: LLM이 응답을 생성할 때 보는 데이터입니다
|
||||
1. 코드에서 로컬로 사용할 수 있는 컨텍스트: 도구 함수가 실행될 때, `on_handoff` 같은 콜백 중에, 생명주기 훅 등에서 필요할 수 있는 데이터와 의존성입니다.
|
||||
2. LLM이 사용할 수 있는 컨텍스트: LLM이 응답을 생성할 때 보는 데이터입니다.
|
||||
|
||||
## 로컬 컨텍스트
|
||||
|
||||
이는 [`RunContextWrapper`][agents.run_context.RunContextWrapper] 클래스와 그 안의 [`context`][agents.run_context.RunContextWrapper.context] 속성으로 표현됩니다. 동작 방식은 다음과 같습니다
|
||||
이는 [`RunContextWrapper`][agents.run_context.RunContextWrapper] 클래스와 그 안의 [`context`][agents.run_context.RunContextWrapper.context] 속성으로 표현됩니다. 작동 방식은 다음과 같습니다.
|
||||
|
||||
1. 원하는 Python 객체를 생성합니다. 일반적으로 dataclass 또는 Pydantic 객체를 사용합니다
|
||||
2. 해당 객체를 다양한 run 메서드에 전달합니다(예: `Runner.run(..., context=whatever)`)
|
||||
3. 모든 도구 호출, 라이프사이클 훅 등은 `RunContextWrapper[T]` 래퍼 객체를 전달받으며, 여기서 `T`는 `wrapper.context`로 접근 가능한 컨텍스트 객체 타입입니다
|
||||
1. 원하는 Python 객체를 만듭니다. 일반적인 패턴은 dataclass나 Pydantic 객체를 사용하는 것입니다.
|
||||
2. 해당 객체를 다양한 run 메서드에 전달합니다(예: `Runner.run(..., context=whatever)`).
|
||||
3. 모든 도구 호출, 생명주기 훅 등에는 래퍼 객체인 `RunContextWrapper[T]`가 전달되며, 여기서 `T`는 `wrapper.context`를 통해 접근할 수 있는 컨텍스트 객체 타입을 나타냅니다.
|
||||
|
||||
일부 런타임 전용 콜백에서는 SDK가 `RunContextWrapper[T]`의 더 특화된 하위 클래스를 전달할 수 있습니다. 예를 들어, 함수 도구 라이프사이클 훅은 보통 `ToolContext`를 받으며, 이는 `tool_call_id`, `tool_name`, `tool_arguments` 같은 도구 호출 메타데이터도 제공합니다
|
||||
일부 런타임별 콜백의 경우 SDK가 `RunContextWrapper[T]`의 더 특화된 서브클래스를 전달할 수 있습니다. 예를 들어 함수 도구 생명주기 훅은 일반적으로 `ToolContext`를 받으며, 이는 `tool_call_id`, `tool_name`, `tool_arguments` 같은 도구 호출 메타데이터도 노출합니다.
|
||||
|
||||
가장 **중요한** 점은 다음과 같습니다: 특정 에이전트 실행에서 모든 에이전트, 도구 함수, 라이프사이클 등은 동일한 컨텍스트 _타입_ 을 사용해야 합니다
|
||||
알아두어야 할 **가장 중요한** 점은 특정 에이전트 실행에 포함되는 모든 에이전트, 도구 함수, 생명주기 등은 동일한 컨텍스트 _타입_을 사용해야 한다는 것입니다.
|
||||
|
||||
컨텍스트는 다음과 같은 용도로 사용할 수 있습니다
|
||||
컨텍스트는 다음과 같은 용도로 사용할 수 있습니다.
|
||||
|
||||
- 실행에 대한 맥락 데이터(예: 사용자 이름/uid 또는 사용자에 관한 기타 정보)
|
||||
- 의존성(예: logger 객체, 데이터 fetcher 등)
|
||||
- 헬퍼 함수
|
||||
- 실행을 위한 컨텍스트 데이터(예: 사용자 이름/uid 또는 사용자에 대한 기타 정보)
|
||||
- 의존성(예: 로거 객체, 데이터 페처 등)
|
||||
- 헬퍼 함수
|
||||
|
||||
!!! danger "참고"
|
||||
|
||||
컨텍스트 객체는 LLM으로 전송되지 **않습니다**. 이는 순수하게 로컬 객체이며, 읽고 쓰고 메서드를 호출할 수 있습니다
|
||||
컨텍스트 객체는 LLM으로 전송되지 **않습니다**. 이는 오직 로컬 객체이며, 읽고 쓰거나 해당 객체의 메서드를 호출할 수 있습니다.
|
||||
|
||||
단일 run 내에서 파생 래퍼는 동일한 기본 앱 컨텍스트, 승인 상태, 사용량 추적을 공유합니다. 중첩된 [`Agent.as_tool()`][agents.agent.Agent.as_tool] run은 다른 `tool_input`을 연결할 수 있지만, 기본적으로 앱 상태의 격리된 복사본을 받지는 않습니다
|
||||
단일 실행 내에서 파생된 래퍼들은 동일한 기본 앱 컨텍스트, 승인 상태, 사용량 추적을 공유합니다. 중첩된 [`Agent.as_tool()`][agents.agent.Agent.as_tool] 실행은 다른 `tool_input`을 붙일 수 있지만, 기본적으로 앱 상태의 격리된 복사본을 받지는 않습니다.
|
||||
|
||||
### `RunContextWrapper` 노출 항목
|
||||
### `RunContextWrapper`의 노출 항목
|
||||
|
||||
[`RunContextWrapper`][agents.run_context.RunContextWrapper]는 앱에서 정의한 컨텍스트 객체를 감싸는 래퍼입니다. 실제로는 주로 다음을 사용합니다
|
||||
[`RunContextWrapper`][agents.run_context.RunContextWrapper]는 앱에서 정의한 컨텍스트 객체를 감싸는 래퍼입니다. 실제로는 대부분 다음을 사용하게 됩니다.
|
||||
|
||||
- 자체 변경 가능한 앱 상태와 의존성을 위한 [`wrapper.context`][agents.run_context.RunContextWrapper.context]
|
||||
- 현재 run 전체의 요청/토큰 사용량 집계를 위한 [`wrapper.usage`][agents.run_context.RunContextWrapper.usage]
|
||||
- 현재 run이 [`Agent.as_tool()`][agents.agent.Agent.as_tool] 내부에서 실행 중일 때 구조화된 입력을 위한 [`wrapper.tool_input`][agents.run_context.RunContextWrapper.tool_input]
|
||||
- 승인 상태를 프로그래밍 방식으로 업데이트해야 할 때 [`wrapper.approve_tool(...)`][agents.run_context.RunContextWrapper.approve_tool] / [`wrapper.reject_tool(...)`][agents.run_context.RunContextWrapper.reject_tool]
|
||||
- [`wrapper.context`][agents.run_context.RunContextWrapper.context]: 직접 사용하는 변경 가능한 앱 상태와 의존성
|
||||
- [`wrapper.usage`][agents.run_context.RunContextWrapper.usage]: 현재 실행 전반의 집계된 요청 및 토큰 사용량
|
||||
- [`wrapper.tool_input`][agents.run_context.RunContextWrapper.tool_input]: 현재 실행이 [`Agent.as_tool()`][agents.agent.Agent.as_tool] 안에서 실행 중일 때의 구조화된 입력
|
||||
- [`wrapper.approve_tool(...)`][agents.run_context.RunContextWrapper.approve_tool] / [`wrapper.reject_tool(...)`][agents.run_context.RunContextWrapper.reject_tool]: 승인 상태를 프로그래밍 방식으로 업데이트해야 할 때 사용
|
||||
|
||||
`wrapper.context`만 앱에서 정의한 객체입니다. 나머지 필드는 SDK가 관리하는 런타임 메타데이터입니다
|
||||
`wrapper.context`만 앱에서 정의한 객체입니다. 다른 필드는 SDK가 관리하는 런타임 메타데이터입니다.
|
||||
|
||||
나중에 휴먼인더루프 (HITL) 또는 내구성 있는 작업 워크플로를 위해 [`RunState`][agents.run_state.RunState]를 직렬화하면, 해당 런타임 메타데이터도 상태와 함께 저장됩니다. 직렬화된 상태를 저장하거나 전송할 계획이라면 [`RunContextWrapper.context`][agents.run_context.RunContextWrapper.context]에 비밀 정보를 넣지 마세요
|
||||
나중에 휴먼인더루프 (HITL) 또는 내구성 있는 작업 워크플로를 위해 [`RunState`][agents.run_state.RunState]를 직렬화하면, 해당 런타임 메타데이터가 상태와 함께 저장됩니다. 직렬화된 상태를 영속화하거나 전송할 계획이라면 [`RunContextWrapper.context`][agents.run_context.RunContextWrapper.context]에 비밀 정보를 넣지 마세요.
|
||||
|
||||
대화 상태는 별도의 관심사입니다. 턴을 어떻게 이어갈지에 따라 `result.to_input_list()`, `session`, `conversation_id`, 또는 `previous_response_id`를 사용하세요. 이 결정은 [결과](results.md), [에이전트 실행](running_agents.md), [세션](sessions/index.md)을 참고하세요
|
||||
대화 상태는 별개의 문제입니다. 턴을 이어가는 방식에 따라 `result.to_input_list()`, `session`, `conversation_id`, 또는 `previous_response_id`를 사용하세요. 이 결정에 대해서는 [결과](results.md), [에이전트 실행](running_agents.md), [세션](sessions/index.md)을 참고하세요.
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -85,18 +85,18 @@ if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
1. 이것이 컨텍스트 객체입니다. 여기서는 dataclass를 사용했지만 어떤 타입이든 사용할 수 있습니다
|
||||
2. 이것은 도구입니다. `RunContextWrapper[UserInfo]`를 받는 것을 볼 수 있습니다. 도구 구현은 컨텍스트에서 값을 읽습니다
|
||||
3. 타입 체커가 오류를 잡을 수 있도록(예: 다른 컨텍스트 타입을 받는 도구를 전달하려는 경우) 에이전트에 제네릭 `UserInfo`를 표시합니다
|
||||
4. 컨텍스트는 `run` 함수에 전달됩니다
|
||||
5. 에이전트가 도구를 올바르게 호출하고 나이를 가져옵니다
|
||||
1. 이것이 컨텍스트 객체입니다. 여기서는 dataclass를 사용했지만, 어떤 타입이든 사용할 수 있습니다.
|
||||
2. 이것은 도구입니다. `RunContextWrapper[UserInfo]`를 받는 것을 볼 수 있습니다. 도구 구현은 컨텍스트에서 읽습니다.
|
||||
3. 에이전트에 제네릭 `UserInfo`를 표시하여, 타입 검사기가 오류를 잡을 수 있도록 합니다(예를 들어 다른 컨텍스트 타입을 받는 도구를 전달하려고 한 경우).
|
||||
4. 컨텍스트가 `run` 함수에 전달됩니다.
|
||||
5. 에이전트가 도구를 올바르게 호출하고 나이를 가져옵니다.
|
||||
|
||||
---
|
||||
|
||||
### 고급: `ToolContext`
|
||||
|
||||
경우에 따라 실행 중인 도구에 대한 추가 메타데이터(예: 이름, 호출 ID, 원시 인자 문자열)에 접근하고 싶을 수 있습니다
|
||||
이때는 `RunContextWrapper`를 확장한 [`ToolContext`][agents.tool_context.ToolContext] 클래스를 사용할 수 있습니다
|
||||
어떤 경우에는 실행 중인 도구에 대한 추가 메타데이터(예: 이름, 호출 ID, 원문 인수 문자열)에 접근하고 싶을 수 있습니다.
|
||||
이를 위해 `RunContextWrapper`를 확장하는 [`ToolContext`][agents.tool_context.ToolContext] 클래스를 사용할 수 있습니다.
|
||||
|
||||
```python
|
||||
from typing import Annotated
|
||||
@@ -124,25 +124,25 @@ agent = Agent(
|
||||
)
|
||||
```
|
||||
|
||||
`ToolContext`는 `RunContextWrapper`와 동일한 `.context` 속성을 제공하며
|
||||
현재 도구 호출에 특화된 추가 필드도 제공합니다
|
||||
`ToolContext`는 `RunContextWrapper`와 동일한 `.context` 속성을 제공하며,
|
||||
현재 도구 호출에 특화된 추가 필드도 제공합니다.
|
||||
|
||||
- `tool_name` – 호출되는 도구의 이름
|
||||
- `tool_call_id` – 이 도구 호출의 고유 식별자
|
||||
- `tool_arguments` – 도구에 전달된 원시 인자 문자열
|
||||
- `tool_namespace` – 도구가 `tool_namespace()` 또는 다른 네임스페이스 표면을 통해 로드된 경우, 도구 호출의 Responses 네임스페이스
|
||||
- `qualified_tool_name` – 네임스페이스가 있을 때 네임스페이스가 포함된 도구 이름
|
||||
- `tool_arguments` – 도구에 전달된 원문 인수 문자열
|
||||
- `tool_namespace` – 도구가 `tool_namespace()` 또는 다른 네임스페이스가 지정된 표면을 통해 로드된 경우, 도구 호출의 Responses 네임스페이스
|
||||
- `qualified_tool_name` – 네임스페이스가 있을 때 해당 네임스페이스로 한정된 도구 이름
|
||||
|
||||
실행 중 도구 수준 메타데이터가 필요할 때 `ToolContext`를 사용하세요
|
||||
에이전트와 도구 간의 일반적인 컨텍스트 공유에는 `RunContextWrapper`로 충분합니다. `ToolContext`는 `RunContextWrapper`를 확장하므로, 중첩된 `Agent.as_tool()` run이 구조화된 입력을 제공한 경우 `.tool_input`도 노출할 수 있습니다
|
||||
실행 중에 도구 수준 메타데이터가 필요할 때 `ToolContext`를 사용하세요.
|
||||
에이전트와 도구 간의 일반적인 컨텍스트 공유에는 `RunContextWrapper`로 충분합니다. `ToolContext`는 `RunContextWrapper`를 확장하므로, 중첩된 `Agent.as_tool()` 실행이 구조화된 입력을 제공한 경우 `.tool_input`도 노출할 수 있습니다.
|
||||
|
||||
---
|
||||
|
||||
## 에이전트/LLM 컨텍스트
|
||||
|
||||
LLM이 호출될 때 LLM이 볼 수 있는 데이터는 대화 기록뿐입니다. 즉, LLM에서 새로운 데이터를 사용할 수 있게 하려면 해당 기록에서 접근 가능하도록 만들어야 합니다. 방법은 몇 가지가 있습니다
|
||||
LLM이 호출될 때 LLM이 볼 수 있는 **유일한** 데이터는 대화 기록에 있는 데이터입니다. 즉, LLM이 어떤 새 데이터를 사용할 수 있게 하려면 해당 기록에서 사용할 수 있는 방식으로 제공해야 합니다. 이를 수행하는 방법은 몇 가지가 있습니다.
|
||||
|
||||
1. Agent `instructions`에 추가할 수 있습니다. 이는 "시스템 프롬프트" 또는 "개발자 메시지"라고도 합니다. 시스템 프롬프트는 정적 문자열일 수도 있고, 컨텍스트를 받아 문자열을 출력하는 동적 함수일 수도 있습니다. 이는 항상 유용한 정보(예: 사용자 이름 또는 현재 날짜)에 자주 쓰이는 방법입니다
|
||||
2. `Runner.run` 함수를 호출할 때 `input`에 추가합니다. 이는 `instructions` 방식과 유사하지만, [명령 체계](https://cdn.openai.com/spec/model-spec-2024-05-08.html#follow-the-chain-of-command)에서 더 낮은 우선순위의 메시지를 둘 수 있게 해줍니다
|
||||
3. 함수 도구를 통해 노출합니다. 이는 _온디맨드_ 컨텍스트에 유용합니다. LLM이 어떤 데이터가 필요할 때를 스스로 결정하고, 그 데이터를 가져오기 위해 도구를 호출할 수 있습니다
|
||||
4. retrieval 또는 웹 검색을 사용합니다. 이는 파일이나 데이터베이스(retrieval), 또는 웹(웹 검색)에서 관련 데이터를 가져올 수 있는 특수 도구입니다. 이는 관련 컨텍스트 데이터에 응답을 "grounding"하는 데 유용합니다
|
||||
1. Agent `instructions`에 추가할 수 있습니다. 이는 "시스템 프롬프트" 또는 "개발자 메시지"라고도 합니다. 시스템 프롬프트는 정적 문자열일 수도 있고, 컨텍스트를 받아 문자열을 출력하는 동적 함수일 수도 있습니다. 이는 항상 유용한 정보(예: 사용자의 이름 또는 현재 날짜)에 흔히 사용하는 전략입니다.
|
||||
2. `Runner.run` 함수를 호출할 때 `input`에 추가합니다. 이는 `instructions` 전략과 유사하지만, [명령 체계](https://cdn.openai.com/spec/model-spec-2024-05-08.html#follow-the-chain-of-command)에서 더 낮은 위치의 메시지를 사용할 수 있게 해줍니다.
|
||||
3. 함수 도구를 통해 노출합니다. 이는 _온디맨드_ 컨텍스트에 유용합니다. LLM이 어떤 데이터가 필요한 시점을 결정하고, 해당 데이터를 가져오기 위해 도구를 호출할 수 있습니다.
|
||||
4. 검색 또는 웹 검색을 사용합니다. 이는 파일이나 데이터베이스에서 관련 데이터를 가져올 수 있는 특수 도구(검색) 또는 웹에서 가져올 수 있는 특수 도구(웹 검색)입니다. 이는 응답을 관련 컨텍스트 데이터에 "근거화"하는 데 유용합니다.
|
||||
+81
-95
@@ -4,139 +4,125 @@ search:
|
||||
---
|
||||
# 코드 예제
|
||||
|
||||
[repo](https://github.com/openai/openai-agents-python/tree/main/examples)의 examples 섹션에서 SDK의 다양한 샘플 구현을 확인해 보세요. examples는 서로 다른 패턴과 기능을 보여주는 여러 카테고리로 구성되어 있습니다.
|
||||
[리포지토리](https://github.com/openai/openai-agents-python/tree/main/examples)의 examples 섹션에서 SDK의 다양한 샘플 구현을 확인해 보세요. 코드 예제는 서로 다른 패턴과 기능을 보여주는 여러 카테고리로 구성되어 있습니다.
|
||||
|
||||
## 카테고리
|
||||
|
||||
- **[agent_patterns](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns):**
|
||||
이 카테고리의 예제는 다음과 같은 일반적인 에이전트 설계 패턴을 보여줍니다
|
||||
- **[agent_patterns](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns):** 이 카테고리의 코드 예제는 다음과 같은 일반적인 에이전트 설계 패턴을 보여줍니다
|
||||
|
||||
- 결정론적 워크플로
|
||||
- Agents as tools
|
||||
- 스트리밍 이벤트를 포함한 Agents as tools (`examples/agent_patterns/agents_as_tools_streaming.py`)
|
||||
- 구조화된 입력 매개변수를 포함한 Agents as tools (`examples/agent_patterns/agents_as_tools_structured.py`)
|
||||
- 스트리밍 이벤트가 포함된 Agents as tools(`examples/agent_patterns/agents_as_tools_streaming.py`)
|
||||
- 구조화된 입력 매개변수가 포함된 Agents as tools(`examples/agent_patterns/agents_as_tools_structured.py`)
|
||||
- 병렬 에이전트 실행
|
||||
- 조건부 도구 사용
|
||||
- 서로 다른 동작으로 도구 사용 강제 (`examples/agent_patterns/forcing_tool_use.py`)
|
||||
- 입출력 가드레일
|
||||
- 심판 역할의 LLM
|
||||
- 서로 다른 동작으로 도구 사용 강제(`examples/agent_patterns/forcing_tool_use.py`)
|
||||
- 입력/출력 가드레일
|
||||
- 판정자로서의 LLM
|
||||
- 라우팅
|
||||
- 스트리밍 가드레일
|
||||
- 도구 승인 및 상태 직렬화를 포함한 휴먼인더루프 (HITL) (`examples/agent_patterns/human_in_the_loop.py`)
|
||||
- 스트리밍을 포함한 휴먼인더루프 (HITL) (`examples/agent_patterns/human_in_the_loop_stream.py`)
|
||||
- 승인 플로를 위한 사용자 지정 거절 메시지 (`examples/agent_patterns/human_in_the_loop_custom_rejection.py`)
|
||||
- 도구 승인 및 상태 직렬화를 사용하는 휴먼인더루프 (HITL)(`examples/agent_patterns/human_in_the_loop.py`)
|
||||
- 스트리밍을 사용하는 휴먼인더루프 (HITL)(`examples/agent_patterns/human_in_the_loop_stream.py`)
|
||||
- 승인 플로를 위한 사용자 지정 거부 메시지(`examples/agent_patterns/human_in_the_loop_custom_rejection.py`)
|
||||
|
||||
- **[basic](https://github.com/openai/openai-agents-python/tree/main/examples/basic):**
|
||||
이 예제들은 다음과 같은 SDK의 기본 기능을 보여줍니다
|
||||
- **[basic](https://github.com/openai/openai-agents-python/tree/main/examples/basic):** 이 코드 예제는 SDK의 기반 기능을 보여줍니다
|
||||
|
||||
- Hello World 예제 (기본 모델, GPT-5, 오픈 웨이트 모델)
|
||||
- 에이전트 라이프사이클 관리
|
||||
- 실행 훅 및 에이전트 훅 라이프사이클 예제 (`examples/basic/lifecycle_example.py`)
|
||||
- Hello world 코드 예제(기본 모델, GPT-5, 오픈 웨이트 모델)
|
||||
- 에이전트 수명 주기 관리
|
||||
- 실행 훅 및 에이전트 훅 수명 주기 예제(`examples/basic/lifecycle_example.py`)
|
||||
- 동적 시스템 프롬프트
|
||||
- 기본 도구 사용 (`examples/basic/tools.py`)
|
||||
- 도구 입출력 가드레일 (`examples/basic/tool_guardrails.py`)
|
||||
- 이미지 도구 출력 (`examples/basic/image_tool_output.py`)
|
||||
- 스트리밍 출력 (텍스트, 항목, 함수 호출 인자)
|
||||
- 턴 간 공유 세션 헬퍼를 사용하는 Responses websocket 전송 (`examples/basic/stream_ws.py`)
|
||||
- 기본 도구 사용(`examples/basic/tools.py`)
|
||||
- 도구 입력/출력 가드레일(`examples/basic/tool_guardrails.py`)
|
||||
- 이미지 도구 출력(`examples/basic/image_tool_output.py`)
|
||||
- 스트리밍 출력(텍스트, 항목, 함수 호출 인수)
|
||||
- 턴 전반에서 공유 세션 헬퍼를 사용하는 Responses WebSocket 전송(`examples/basic/stream_ws.py`)
|
||||
- 프롬프트 템플릿
|
||||
- 파일 처리 (로컬 및 원격, 이미지 및 PDF)
|
||||
- 파일 처리(로컬 및 원격, 이미지 및 PDF)
|
||||
- 사용량 추적
|
||||
- Runner 관리 재시도 설정 (`examples/basic/retry.py`)
|
||||
- 서드파티 어댑터를 통한 Runner 관리 재시도 (`examples/basic/retry_litellm.py`)
|
||||
- 비엄격 출력 타입
|
||||
- Runner 관리 재시도 설정(`examples/basic/retry.py`)
|
||||
- 서드파티 어댑터를 통한 Runner 관리 재시도(`examples/basic/retry_litellm.py`)
|
||||
- 비엄격 출력 유형
|
||||
- 이전 응답 ID 사용
|
||||
|
||||
- **[customer_service](https://github.com/openai/openai-agents-python/tree/main/examples/customer_service):**
|
||||
항공사를 위한 고객 서비스 시스템 예제입니다
|
||||
- **[customer_service](https://github.com/openai/openai-agents-python/tree/main/examples/customer_service):** 항공사를 위한 고객 서비스 시스템 예제입니다.
|
||||
|
||||
- **[financial_research_agent](https://github.com/openai/openai-agents-python/tree/main/examples/financial_research_agent):**
|
||||
금융 데이터 분석을 위한 에이전트와 도구를 사용한 구조화된 리서치 워크플로를 보여주는 금융 리서치 에이전트입니다
|
||||
- **[financial_research_agent](https://github.com/openai/openai-agents-python/tree/main/examples/financial_research_agent):** 금융 데이터 분석을 위한 에이전트와 도구를 사용하여 구조화된 리서치 워크플로를 보여주는 금융 리서치 에이전트입니다.
|
||||
|
||||
- **[handoffs](https://github.com/openai/openai-agents-python/tree/main/examples/handoffs):**
|
||||
메시지 필터링을 포함한 에이전트 핸드오프의 실용적인 예제:
|
||||
- **[handoffs](https://github.com/openai/openai-agents-python/tree/main/examples/handoffs):** 메시지 필터링을 포함한 에이전트 핸드오프의 실용적인 코드 예제입니다. 포함 항목:
|
||||
|
||||
- 메시지 필터 예제 (`examples/handoffs/message_filter.py`)
|
||||
- 스트리밍을 포함한 메시지 필터 (`examples/handoffs/message_filter_streaming.py`)
|
||||
- 메시지 필터 예제(`examples/handoffs/message_filter.py`)
|
||||
- 스트리밍을 사용하는 메시지 필터(`examples/handoffs/message_filter_streaming.py`)
|
||||
|
||||
- **[hosted_mcp](https://github.com/openai/openai-agents-python/tree/main/examples/hosted_mcp):**
|
||||
OpenAI Responses API와 함께 호스티드 MCP (Model context protocol)를 사용하는 방법을 보여주는 예제:
|
||||
- **[hosted_mcp](https://github.com/openai/openai-agents-python/tree/main/examples/hosted_mcp):** OpenAI Responses API와 함께 호스티드 MCP (Model Context Protocol)를 사용하는 방법을 보여주는 코드 예제입니다. 포함 항목:
|
||||
|
||||
- 승인 없는 간단한 호스티드 MCP (`examples/hosted_mcp/simple.py`)
|
||||
- Google Calendar 같은 MCP 커넥터 (`examples/hosted_mcp/connectors.py`)
|
||||
- 인터럽션(중단 처리) 기반 승인을 포함한 휴먼인더루프 (HITL) (`examples/hosted_mcp/human_in_the_loop.py`)
|
||||
- MCP 도구 호출용 승인 시 콜백 (`examples/hosted_mcp/on_approval.py`)
|
||||
- 승인 없는 간단한 호스티드 MCP(`examples/hosted_mcp/simple.py`)
|
||||
- Google Calendar와 같은 MCP 커넥터(`examples/hosted_mcp/connectors.py`)
|
||||
- 인터럽션(중단 처리) 기반 승인을 사용하는 휴먼인더루프 (HITL)(`examples/hosted_mcp/human_in_the_loop.py`)
|
||||
- MCP 도구 호출을 위한 승인 시 콜백(`examples/hosted_mcp/on_approval.py`)
|
||||
|
||||
- **[mcp](https://github.com/openai/openai-agents-python/tree/main/examples/mcp):**
|
||||
MCP (Model context protocol)로 에이전트를 구축하는 방법을 알아보세요:
|
||||
- **[mcp](https://github.com/openai/openai-agents-python/tree/main/examples/mcp):** MCP (Model Context Protocol)로 에이전트를 구축하는 방법을 알아보세요. 포함 항목:
|
||||
|
||||
- 파일시스템 예제
|
||||
- Git 예제
|
||||
- MCP 프롬프트 서버 예제
|
||||
- SSE (Server-Sent Events) 예제
|
||||
- SSE 원격 서버 연결 (`examples/mcp/sse_remote_example`)
|
||||
- Streamable HTTP 예제
|
||||
- Streamable HTTP 원격 연결 (`examples/mcp/streamable_http_remote_example`)
|
||||
- Streamable HTTP용 사용자 지정 HTTP 클라이언트 팩토리 (`examples/mcp/streamablehttp_custom_client_example`)
|
||||
- `MCPUtil.get_all_function_tools`를 사용한 모든 MCP 도구 프리패칭 (`examples/mcp/get_all_mcp_tools_example`)
|
||||
- FastAPI를 사용하는 MCPServerManager (`examples/mcp/manager_example`)
|
||||
- MCP 도구 필터링 (`examples/mcp/tool_filter_example`)
|
||||
- 파일시스템 코드 예제
|
||||
- Git 코드 예제
|
||||
- MCP 프롬프트 서버 코드 예제
|
||||
- SSE (Server-Sent Events) 코드 예제
|
||||
- SSE 원격 서버 연결(`examples/mcp/sse_remote_example`)
|
||||
- Streamable HTTP 코드 예제
|
||||
- Streamable HTTP 원격 연결(`examples/mcp/streamable_http_remote_example`)
|
||||
- Streamable HTTP용 사용자 지정 HTTP 클라이언트 팩토리(`examples/mcp/streamablehttp_custom_client_example`)
|
||||
- `MCPUtil.get_all_function_tools`로 모든 MCP 도구 미리 가져오기(`examples/mcp/get_all_mcp_tools_example`)
|
||||
- FastAPI와 함께 사용하는 MCPServerManager(`examples/mcp/manager_example`)
|
||||
- MCP 도구 필터링(`examples/mcp/tool_filter_example`)
|
||||
|
||||
- **[memory](https://github.com/openai/openai-agents-python/tree/main/examples/memory):**
|
||||
에이전트를 위한 다양한 메모리 구현 예제:
|
||||
- **[memory](https://github.com/openai/openai-agents-python/tree/main/examples/memory):** 에이전트를 위한 다양한 메모리 구현 코드 예제입니다. 포함 항목:
|
||||
|
||||
- SQLite 세션 저장소
|
||||
- 고급 SQLite 세션 저장소
|
||||
- Redis 세션 저장소
|
||||
- SQLAlchemy 세션 저장소
|
||||
- Dapr 상태 저장소 세션 저장소
|
||||
- 암호화된 세션 저장소
|
||||
- OpenAI Conversations 세션 저장소
|
||||
- Responses 컴팩션 세션 저장소
|
||||
- `ModelSettings(store=False)`를 사용한 상태 비저장 Responses 컴팩션 (`examples/memory/compaction_session_stateless_example.py`)
|
||||
- 파일 기반 세션 저장소 (`examples/memory/file_session.py`)
|
||||
- 휴먼인더루프 (HITL)를 포함한 파일 기반 세션 (`examples/memory/file_hitl_example.py`)
|
||||
- 휴먼인더루프 (HITL)를 포함한 SQLite 인메모리 세션 (`examples/memory/memory_session_hitl_example.py`)
|
||||
- 휴먼인더루프 (HITL)를 포함한 OpenAI Conversations 세션 (`examples/memory/openai_session_hitl_example.py`)
|
||||
- 세션 전반의 HITL 승인/거절 시나리오 (`examples/memory/hitl_session_scenario.py`)
|
||||
- SQLite 세션 스토리지
|
||||
- 고급 SQLite 세션 스토리지
|
||||
- Redis 세션 스토리지
|
||||
- SQLAlchemy 세션 스토리지
|
||||
- Dapr 상태 저장소 세션 스토리지
|
||||
- 암호화된 세션 스토리지
|
||||
- OpenAI Conversations 세션 스토리지
|
||||
- Responses 압축 세션 스토리지
|
||||
- `ModelSettings(store=False)`를 사용하는 상태 비저장 Responses 압축(`examples/memory/compaction_session_stateless_example.py`)
|
||||
- 파일 기반 세션 스토리지(`examples/memory/file_session.py`)
|
||||
- 휴먼인더루프 (HITL)를 사용하는 파일 기반 세션(`examples/memory/file_hitl_example.py`)
|
||||
- 휴먼인더루프 (HITL)를 사용하는 SQLite 인메모리 세션(`examples/memory/memory_session_hitl_example.py`)
|
||||
- 휴먼인더루프 (HITL)를 사용하는 OpenAI Conversations 세션(`examples/memory/openai_session_hitl_example.py`)
|
||||
- 세션 전반의 HITL 승인/거부 시나리오(`examples/memory/hitl_session_scenario.py`)
|
||||
|
||||
- **[model_providers](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers):**
|
||||
사용자 지정 프로바이더와 서드파티 어댑터를 포함해 SDK에서 OpenAI 이외 모델을 사용하는 방법을 살펴보세요
|
||||
- **[model_providers](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers):** 사용자 지정 제공자 및 서드파티 어댑터를 포함하여 SDK에서 OpenAI가 아닌 모델을 사용하는 방법을 살펴보세요.
|
||||
|
||||
- **[realtime](https://github.com/openai/openai-agents-python/tree/main/examples/realtime):**
|
||||
SDK를 사용해 실시간 경험을 구축하는 방법을 보여주는 예제:
|
||||
- **[realtime](https://github.com/openai/openai-agents-python/tree/main/examples/realtime):** SDK를 사용하여 실시간 경험을 구축하는 방법을 보여주는 코드 예제입니다. 포함 항목:
|
||||
|
||||
- 구조화된 텍스트 및 이미지 메시지를 사용하는 웹 애플리케이션 패턴
|
||||
- 커맨드라인 오디오 루프 및 재생 처리
|
||||
- 명령줄 오디오 루프 및 재생 처리
|
||||
- WebSocket을 통한 Twilio Media Streams 통합
|
||||
- Realtime Calls API attach 플로를 사용하는 Twilio SIP 통합
|
||||
|
||||
- **[reasoning_content](https://github.com/openai/openai-agents-python/tree/main/examples/reasoning_content):**
|
||||
reasoning content를 다루는 방법을 보여주는 예제:
|
||||
- **[reasoning_content](https://github.com/openai/openai-agents-python/tree/main/examples/reasoning_content):** 추론 콘텐츠를 다루는 방법을 보여주는 코드 예제입니다. 포함 항목:
|
||||
|
||||
- Runner API의 reasoning content, 스트리밍 및 비스트리밍 (`examples/reasoning_content/runner_example.py`)
|
||||
- OpenRouter를 통한 OSS 모델의 reasoning content (`examples/reasoning_content/gpt_oss_stream.py`)
|
||||
- 기본 reasoning content 예제 (`examples/reasoning_content/main.py`)
|
||||
- Runner API, 스트리밍 및 비스트리밍을 사용하는 추론 콘텐츠(`examples/reasoning_content/runner_example.py`)
|
||||
- OpenRouter를 통해 OSS 모델을 사용하는 추론 콘텐츠(`examples/reasoning_content/gpt_oss_stream.py`)
|
||||
- 기본 추론 콘텐츠 예제(`examples/reasoning_content/main.py`)
|
||||
|
||||
- **[research_bot](https://github.com/openai/openai-agents-python/tree/main/examples/research_bot):**
|
||||
복잡한 멀티 에이전트 리서치 워크플로를 보여주는 간단한 딥 리서치 클론입니다
|
||||
- **[research_bot](https://github.com/openai/openai-agents-python/tree/main/examples/research_bot):** 복잡한 다중 에이전트 리서치 워크플로를 보여주는 간단한 딥 리서치 클론입니다.
|
||||
|
||||
- **[tools](https://github.com/openai/openai-agents-python/tree/main/examples/tools):**
|
||||
다음과 같은 OpenAI 호스트하는 도구 및 실험적 Codex 도구 기능을 구현하는 방법을 알아보세요:
|
||||
- **[tools](https://github.com/openai/openai-agents-python/tree/main/examples/tools):** 다음과 같은 OpenAI 호스트하는 도구와 실험적 Codex 도구를 구현하는 방법을 알아보세요
|
||||
|
||||
- 웹 검색 및 필터를 포함한 웹 검색
|
||||
- 웹 검색 및 필터가 있는 웹 검색
|
||||
- 파일 검색
|
||||
- Code Interpreter
|
||||
- 파일 편집 및 승인을 포함한 패치 적용 도구 (`examples/tools/apply_patch.py`)
|
||||
- 승인 콜백을 포함한 셸 도구 실행 (`examples/tools/shell.py`)
|
||||
- 휴먼인더루프 (HITL) 인터럽션(중단 처리) 기반 승인을 포함한 셸 도구 (`examples/tools/shell_human_in_the_loop.py`)
|
||||
- 인라인 스킬을 포함한 호스티드 컨테이너 셸 (`examples/tools/container_shell_inline_skill.py`)
|
||||
- 스킬 참조를 포함한 호스티드 컨테이너 셸 (`examples/tools/container_shell_skill_reference.py`)
|
||||
- 로컬 스킬을 포함한 로컬 셸 (`examples/tools/local_shell_skill.py`)
|
||||
- 네임스페이스 및 지연 도구를 사용하는 도구 검색 (`examples/tools/tool_search.py`)
|
||||
- Code interpreter
|
||||
- 파일 편집 및 승인을 사용하는 패치 적용 도구(`examples/tools/apply_patch.py`)
|
||||
- 승인 콜백을 사용하는 셸 도구 실행(`examples/tools/shell.py`)
|
||||
- 휴먼인더루프 (HITL) 인터럽션(중단 처리) 기반 승인을 사용하는 셸 도구(`examples/tools/shell_human_in_the_loop.py`)
|
||||
- 인라인 스킬을 사용하는 호스티드 컨테이너 셸(`examples/tools/container_shell_inline_skill.py`)
|
||||
- 스킬 참조를 사용하는 호스티드 컨테이너 셸(`examples/tools/container_shell_skill_reference.py`)
|
||||
- 로컬 스킬을 사용하는 로컬 셸(`examples/tools/local_shell_skill.py`)
|
||||
- 네임스페이스와 지연된 도구를 사용하는 도구 검색(`examples/tools/tool_search.py`)
|
||||
- 컴퓨터 사용
|
||||
- 이미지 생성
|
||||
- 실험적 Codex 도구 워크플로 (`examples/tools/codex.py`)
|
||||
- 실험적 Codex 동일 스레드 워크플로 (`examples/tools/codex_same_thread.py`)
|
||||
- 실험적 Codex 도구 워크플로(`examples/tools/codex.py`)
|
||||
- 실험적 Codex 동일 스레드 워크플로(`examples/tools/codex_same_thread.py`)
|
||||
|
||||
- **[voice](https://github.com/openai/openai-agents-python/tree/main/examples/voice):**
|
||||
스트리밍 음성 예제를 포함해 TTS 및 STT 모델을 사용하는 음성 에이전트 예제를 확인해 보세요
|
||||
- **[voice](https://github.com/openai/openai-agents-python/tree/main/examples/voice):** OpenAI의 TTS 및 STT 모델을 사용하는 음성 에이전트 코드 예제를 확인하세요. 스트리밍 음성 코드 예제가 포함됩니다.
|
||||
+35
-34
@@ -4,74 +4,75 @@ search:
|
||||
---
|
||||
# 가드레일
|
||||
|
||||
가드레일을 사용하면 사용자 입력과 에이전트 출력에 대한 검사 및 검증을 수행할 수 있습니다. 예를 들어, 고객 요청을 돕기 위해 매우 똑똑한(따라서 느리고/비싼) 모델을 사용하는 에이전트가 있다고 가정해 보겠습니다. 악의적인 사용자가 그 모델에게 수학 숙제를 도와달라고 요청하게 두고 싶지는 않을 것입니다. 따라서 빠르고/저렴한 모델로 가드레일을 실행할 수 있습니다. 가드레일이 악의적인 사용을 감지하면 즉시 오류를 발생시켜 비싼 모델의 실행을 막을 수 있어 시간과 비용을 절약할 수 있습니다(**blocking guardrails를 사용할 때; parallel guardrails의 경우 가드레일이 완료되기 전에 비싼 모델이 이미 실행을 시작했을 수 있습니다. 자세한 내용은 아래의 "Execution modes"를 참고하세요**).
|
||||
가드레일을 사용하면 사용자 입력과 에이전트 출력에 대한 검사 및 검증을 수행할 수 있습니다. 예를 들어, 고객 요청을 돕기 위해 매우 똑똑한(따라서 느리고 비용이 많이 드는) 모델을 사용하는 에이전트가 있다고 가정해 보겠습니다. 악의적인 사용자가 모델에게 수학 숙제를 도와달라고 요청하는 것은 원치 않을 것입니다. 따라서 빠르고 저렴한 모델로 가드레일을 실행할 수 있습니다. 가드레일이 악의적 사용을 감지하면 즉시 오류를 발생시켜 비용이 많이 드는 모델이 실행되지 않도록 하여 시간과 비용을 절약할 수 있습니다(**차단형 가드레일을 사용할 때입니다. 병렬 가드레일의 경우, 가드레일이 완료되기 전에 비용이 많이 드는 모델이 이미 실행을 시작했을 수 있습니다. 자세한 내용은 아래의 "실행 모드"를 참조하세요**).
|
||||
|
||||
가드레일에는 두 가지 종류가 있습니다:
|
||||
|
||||
1. 입력 가드레일은 초기 사용자 입력에서 실행됩니다
|
||||
1. 입력 가드레일은 최초 사용자 입력에서 실행됩니다
|
||||
2. 출력 가드레일은 최종 에이전트 출력에서 실행됩니다
|
||||
|
||||
## 워크플로 경계
|
||||
|
||||
가드레일은 에이전트와 도구에 연결되지만, 워크플로의 동일한 지점에서 모두 실행되지는 않습니다:
|
||||
가드레일은 에이전트와 도구에 연결되지만, 워크플로의 모든 지점에서 실행되는 것은 아닙니다:
|
||||
|
||||
- **입력 가드레일**은 체인의 첫 번째 에이전트에 대해서만 실행됩니다
|
||||
- **출력 가드레일**은 최종 출력을 생성하는 에이전트에 대해서만 실행됩니다
|
||||
- **도구 가드레일**은 모든 커스텀 함수 도구 호출에서 실행되며, 실행 전에는 입력 가드레일이, 실행 후에는 출력 가드레일이 실행됩니다
|
||||
- **입력 가드레일**은 체인의 첫 번째 에이전트에 대해서만 실행됩니다.
|
||||
- **출력 가드레일**은 최종 출력을 생성하는 에이전트에 대해서만 실행됩니다.
|
||||
- **도구 가드레일**은 모든 사용자 지정 함수 도구 호출마다 실행되며, 실행 전에는 입력 가드레일이, 실행 후에는 출력 가드레일이 실행됩니다.
|
||||
|
||||
매니저, 핸드오프 또는 위임된 전문 에이전트가 포함된 워크플로에서 각 커스텀 함수 도구 호출마다 검사가 필요하다면, 에이전트 수준의 입력/출력 가드레일에만 의존하지 말고 도구 가드레일을 사용하세요.
|
||||
매니저, 핸드오프 또는 위임된 전문 에이전트를 포함하는 워크플로에서 각 사용자 지정 함수 도구 호출 전후로 검사가 필요하다면, 에이전트 수준의 입력/출력 가드레일에만 의존하지 말고 도구 가드레일을 사용하세요.
|
||||
|
||||
## 입력 가드레일
|
||||
|
||||
입력 가드레일은 3단계로 실행됩니다:
|
||||
|
||||
1. 먼저, 가드레일은 에이전트에 전달된 것과 동일한 입력을 받습니다
|
||||
2. 다음으로, 가드레일 함수가 실행되어 [`GuardrailFunctionOutput`][agents.guardrail.GuardrailFunctionOutput]을 생성하고, 이는 [`InputGuardrailResult`][agents.guardrail.InputGuardrailResult]로 래핑됩니다
|
||||
3. 마지막으로, [`.tripwire_triggered`][agents.guardrail.GuardrailFunctionOutput.tripwire_triggered]가 true인지 확인합니다. true이면 [`InputGuardrailTripwireTriggered`][agents.exceptions.InputGuardrailTripwireTriggered] 예외가 발생하므로, 사용자에게 적절히 응답하거나 예외를 처리할 수 있습니다
|
||||
1. 먼저 가드레일은 에이전트에 전달된 것과 동일한 입력을 받습니다.
|
||||
2. 다음으로 가드레일 함수가 실행되어 [`GuardrailFunctionOutput`][agents.guardrail.GuardrailFunctionOutput]을 생성하고, 이는 다시 [`InputGuardrailResult`][agents.guardrail.InputGuardrailResult]로 래핑됩니다
|
||||
3. 마지막으로 [`.tripwire_triggered`][agents.guardrail.GuardrailFunctionOutput.tripwire_triggered]가 true인지 확인합니다. true이면 [`InputGuardrailTripwireTriggered`][agents.exceptions.InputGuardrailTripwireTriggered] 예외가 발생하므로, 사용자에게 적절히 응답하거나 예외를 처리할 수 있습니다.
|
||||
|
||||
!!! Note
|
||||
|
||||
입력 가드레일은 사용자 입력에서 실행되도록 설계되었으므로, 에이전트의 가드레일은 해당 에이전트가 *첫 번째* 에이전트일 때만 실행됩니다. 그렇다면 왜 가드레일을 `Runner.run`에 전달하지 않고 에이전트의 `guardrails` 속성에 두는지 궁금할 수 있습니다. 이는 가드레일이 실제 Agent와 관련되는 경향이 있기 때문입니다. 에이전트마다 다른 가드레일을 실행하게 되므로 코드를 함께 배치하면 가독성에 유리합니다.
|
||||
입력 가드레일은 사용자 입력에서 실행되도록 설계되었으므로, 에이전트의 가드레일은 해당 에이전트가 *첫 번째* 에이전트인 경우에만 실행됩니다. 가드레일을 `Runner.run`에 전달하지 않고 에이전트의 `guardrails` 속성에 두는 이유가 궁금할 수 있습니다. 이는 가드레일이 실제 에이전트와 관련되는 경우가 많기 때문입니다. 에이전트마다 서로 다른 가드레일을 실행하게 되므로, 코드를 한곳에 배치하는 것이 가독성에 유용합니다.
|
||||
|
||||
### 실행 모드
|
||||
|
||||
입력 가드레일은 두 가지 실행 모드를 지원합니다:
|
||||
|
||||
- **병렬 실행**(기본값, `run_in_parallel=True`): 가드레일이 에이전트 실행과 동시에 실행됩니다. 둘 다 같은 시점에 시작되므로 지연 시간 측면에서 가장 유리합니다. 하지만 가드레일이 실패하면, 취소되기 전에 에이전트가 이미 토큰을 소비하고 도구를 실행했을 수 있습니다
|
||||
- **병렬 실행**(기본값, `run_in_parallel=True`): 가드레일은 에이전트 실행과 동시에 실행됩니다. 둘 다 같은 시점에 시작하므로 지연 시간이 가장 짧습니다. 그러나 가드레일이 실패하면, 취소되기 전에 에이전트가 이미 토큰을 소비하고 도구를 실행했을 수 있습니다.
|
||||
|
||||
- **차단 실행**(`run_in_parallel=False`): 에이전트가 시작되기 *전에* 가드레일이 실행되고 완료됩니다. 가드레일 트립와이어가 트리거되면 에이전트는 전혀 실행되지 않아 토큰 소비와 도구 실행을 방지합니다. 비용 최적화가 중요하고 도구 호출로 인한 잠재적 부작용을 피하고 싶을 때 이상적입니다
|
||||
- **차단 실행**(`run_in_parallel=False`): 가드레일은 에이전트가 시작되기 *전에* 실행되어 완료됩니다. 가드레일 트립와이어가 트리거되면 에이전트는 전혀 실행되지 않아 토큰 소비와 도구 실행을 방지합니다. 이는 비용 최적화에 이상적이며 도구 호출의 잠재적 부작용을 피하고 싶을 때 적합합니다.
|
||||
|
||||
## 출력 가드레일
|
||||
|
||||
출력 가드레일은 3단계로 실행됩니다:
|
||||
|
||||
1. 먼저, 가드레일은 에이전트가 생성한 출력을 받습니다
|
||||
2. 다음으로, 가드레일 함수가 실행되어 [`GuardrailFunctionOutput`][agents.guardrail.GuardrailFunctionOutput]을 생성하고, 이는 [`OutputGuardrailResult`][agents.guardrail.OutputGuardrailResult]로 래핑됩니다
|
||||
3. 마지막으로, [`.tripwire_triggered`][agents.guardrail.GuardrailFunctionOutput.tripwire_triggered]가 true인지 확인합니다. true이면 [`OutputGuardrailTripwireTriggered`][agents.exceptions.OutputGuardrailTripwireTriggered] 예외가 발생하므로, 사용자에게 적절히 응답하거나 예외를 처리할 수 있습니다
|
||||
1. 먼저 가드레일은 에이전트가 생성한 출력을 받습니다.
|
||||
2. 다음으로 가드레일 함수가 실행되어 [`GuardrailFunctionOutput`][agents.guardrail.GuardrailFunctionOutput]을 생성하고, 이는 다시 [`OutputGuardrailResult`][agents.guardrail.OutputGuardrailResult]로 래핑됩니다
|
||||
3. 마지막으로 [`.tripwire_triggered`][agents.guardrail.GuardrailFunctionOutput.tripwire_triggered]가 true인지 확인합니다. true이면 [`OutputGuardrailTripwireTriggered`][agents.exceptions.OutputGuardrailTripwireTriggered] 예외가 발생하므로, 사용자에게 적절히 응답하거나 예외를 처리할 수 있습니다.
|
||||
|
||||
!!! Note
|
||||
|
||||
출력 가드레일은 최종 에이전트 출력에서 실행되도록 설계되었으므로, 에이전트의 가드레일은 해당 에이전트가 *마지막* 에이전트일 때만 실행됩니다. 입력 가드레일과 마찬가지로 이렇게 하는 이유는 가드레일이 실제 Agent와 관련되는 경향이 있기 때문입니다. 에이전트마다 다른 가드레일을 실행하게 되므로 코드를 함께 배치하면 가독성에 유리합니다.
|
||||
출력 가드레일은 최종 에이전트 출력에서 실행되도록 설계되었으므로, 에이전트의 가드레일은 해당 에이전트가 *마지막* 에이전트인 경우에만 실행됩니다. 입력 가드레일과 마찬가지로, 이는 가드레일이 실제 에이전트와 관련되는 경우가 많기 때문입니다. 에이전트마다 서로 다른 가드레일을 실행하게 되므로, 코드를 한곳에 배치하는 것이 가독성에 유용합니다.
|
||||
|
||||
출력 가드레일은 항상 에이전트 완료 후 실행되므로 `run_in_parallel` 매개변수를 지원하지 않습니다.
|
||||
출력 가드레일은 에이전트가 완료된 후에 항상 실행되므로 `run_in_parallel` 매개변수를 지원하지 않습니다.
|
||||
|
||||
## 도구 가드레일
|
||||
|
||||
도구 가드레일은 **함수 도구**를 감싸서 실행 전후에 도구 호출을 검증하거나 차단할 수 있게 합니다. 도구 자체에 구성되며 해당 도구가 호출될 때마다 실행됩니다.
|
||||
도구 가드레일은 **함수 도구**를 감싸며 실행 전후에 도구 호출을 검증하거나 차단할 수 있게 합니다. 도구 자체에 구성되며 해당 도구가 호출될 때마다 실행됩니다.
|
||||
|
||||
- 입력 도구 가드레일은 도구 실행 전에 실행되며 호출 건너뛰기, 메시지로 출력 대체, 또는 트립와이어 발생을 수행할 수 있습니다
|
||||
- 출력 도구 가드레일은 도구 실행 후에 실행되며 출력 대체 또는 트립와이어 발생을 수행할 수 있습니다
|
||||
- 도구 가드레일은 [`function_tool`][agents.tool.function_tool]로 생성된 함수 도구에만 적용됩니다. 핸드오프는 일반 함수 도구 파이프라인이 아닌 SDK의 핸드오프 파이프라인을 통해 실행되므로, 핸드오프 호출 자체에는 도구 가드레일이 적용되지 않습니다. Hosted tools(`WebSearchTool`, `FileSearchTool`, `HostedMCPTool`, `CodeInterpreterTool`, `ImageGenerationTool`) 및 내장 실행 도구(`ComputerTool`, `ShellTool`, `ApplyPatchTool`, `LocalShellTool`)도 이 가드레일 파이프라인을 사용하지 않으며, [`Agent.as_tool()`][agents.agent.Agent.as_tool]은 현재 도구 가드레일 옵션을 직접 노출하지 않습니다
|
||||
- 입력 도구 가드레일은 도구 실행 전에 실행되며 호출을 건너뛰거나, 출력을 메시지로 대체하거나, 트립와이어를 트리거할 수 있습니다.
|
||||
- 출력 도구 가드레일은 도구 실행 후에 실행되며 출력을 대체하거나 트립와이어를 트리거할 수 있습니다.
|
||||
- 함수 도구에 승인이 필요한 경우, 입력 도구 가드레일은 일반적으로 승인 후 실행 직전에 실행됩니다. 해당 입력 검사를 승인 대기 인터럽션(중단 처리)이 발생하기 전에 실행하려면 [`RunConfig.tool_execution`][agents.run.RunConfig.tool_execution]을 [`ToolExecutionConfig(pre_approval_tool_input_guardrails=True)`][agents.run.ToolExecutionConfig]로 설정하세요. 이 승인 전 검사를 통과한 호출도 승인 이후 도구가 실행되기 전에 다시 검사됩니다.
|
||||
- 도구 가드레일은 [`function_tool`][agents.tool.function_tool]로 생성된 함수 도구에만 적용됩니다. 핸드오프는 일반 함수 도구 파이프라인이 아니라 SDK의 핸드오프 파이프라인을 통해 실행되므로, 도구 가드레일은 핸드오프 호출 자체에는 적용되지 않습니다. 호스티드 툴(`WebSearchTool`, `FileSearchTool`, `HostedMCPTool`, `CodeInterpreterTool`, `ImageGenerationTool`)과 기본 제공 실행 도구(`ComputerTool`, `ShellTool`, `ApplyPatchTool`, `LocalShellTool`)도 이 가드레일 파이프라인을 사용하지 않으며, [`Agent.as_tool()`][agents.agent.Agent.as_tool]은 현재 도구 가드레일 옵션을 직접 노출하지 않습니다.
|
||||
|
||||
자세한 내용은 아래 코드 스니펫을 참고하세요.
|
||||
자세한 내용은 아래 코드 스니펫을 참조하세요.
|
||||
|
||||
## 트립와이어
|
||||
|
||||
입력 또는 출력이 가드레일 검사를 통과하지 못하면, Guardrail은 트립와이어로 이를 신호할 수 있습니다. 트립와이어가 트리거된 가드레일을 확인하는 즉시 `{Input,Output}GuardrailTripwireTriggered` 예외를 발생시키고 Agent 실행을 중단합니다.
|
||||
입력 또는 출력이 가드레일 검사를 통과하지 못하면, 가드레일은 이를 트립와이어로 신호할 수 있습니다. 트립와이어를 트리거한 가드레일이 확인되는 즉시, `{Input,Output}GuardrailTripwireTriggered` 예외를 발생시키고 에이전트 실행을 중단합니다.
|
||||
|
||||
## 가드레일 구현
|
||||
|
||||
입력을 받아 [`GuardrailFunctionOutput`][agents.guardrail.GuardrailFunctionOutput]을 반환하는 함수를 제공해야 합니다. 이 예제에서는 내부적으로 에이전트를 실행하는 방식으로 이를 수행합니다.
|
||||
입력을 받아 [`GuardrailFunctionOutput`][agents.guardrail.GuardrailFunctionOutput]을 반환하는 함수를 제공해야 합니다. 이 예제에서는 내부적으로 에이전트를 실행해 이를 수행합니다.
|
||||
|
||||
```python
|
||||
from pydantic import BaseModel
|
||||
@@ -124,10 +125,10 @@ async def main():
|
||||
print("Math homework guardrail tripped")
|
||||
```
|
||||
|
||||
1. 가드레일 함수에서 이 에이전트를 사용합니다
|
||||
2. 에이전트의 입력/컨텍스트를 받아 결과를 반환하는 가드레일 함수입니다
|
||||
3. 가드레일 결과에 추가 정보를 포함할 수 있습니다
|
||||
4. 워크플로를 정의하는 실제 에이전트입니다
|
||||
1. 이 에이전트를 가드레일 함수에서 사용합니다.
|
||||
2. 이것은 에이전트의 입력/컨텍스트를 받아 결과를 반환하는 가드레일 함수입니다.
|
||||
3. 가드레일 결과에 추가 정보를 포함할 수 있습니다.
|
||||
4. 이것은 워크플로를 정의하는 실제 에이전트입니다.
|
||||
|
||||
출력 가드레일도 유사합니다.
|
||||
|
||||
@@ -182,12 +183,12 @@ async def main():
|
||||
print("Math output guardrail tripped")
|
||||
```
|
||||
|
||||
1. 실제 에이전트의 출력 타입입니다
|
||||
2. 가드레일의 출력 타입입니다
|
||||
3. 에이전트의 출력을 받아 결과를 반환하는 가드레일 함수입니다
|
||||
4. 워크플로를 정의하는 실제 에이전트입니다
|
||||
1. 이것은 실제 에이전트의 출력 타입입니다.
|
||||
2. 이것은 가드레일의 출력 타입입니다.
|
||||
3. 이것은 에이전트의 출력을 받아 결과를 반환하는 가드레일 함수입니다.
|
||||
4. 이것은 워크플로를 정의하는 실제 에이전트입니다.
|
||||
|
||||
마지막으로, 다음은 도구 가드레일 예시입니다.
|
||||
마지막으로, 다음은 도구 가드레일의 코드 예제입니다.
|
||||
|
||||
```python
|
||||
import json
|
||||
|
||||
+36
-36
@@ -4,17 +4,17 @@ search:
|
||||
---
|
||||
# 핸드오프
|
||||
|
||||
핸드오프를 사용하면 한 에이전트가 다른 에이전트에 작업을 위임할 수 있습니다. 이는 서로 다른 에이전트가 각기 다른 영역을 전문으로 하는 시나리오에서 특히 유용합니다. 예를 들어 고객 지원 앱에는 주문 상태, 환불, FAQ 등의 작업을 각각 전담하는 에이전트가 있을 수 있습니다.
|
||||
핸드오프를 사용하면 에이전트가 다른 에이전트에게 작업을 위임할 수 있습니다. 이는 서로 다른 에이전트가 각기 다른 영역을 전문으로 하는 시나리오에서 특히 유용합니다. 예를 들어 고객 지원 앱에는 주문 상태, 환불, FAQ 등의 작업을 각각 전담하는 에이전트가 있을 수 있습니다.
|
||||
|
||||
핸드오프는 LLM에 도구로 표현됩니다. 따라서 `Refund Agent`라는 이름의 에이전트로 핸드오프가 있으면 도구 이름은 `transfer_to_refund_agent`가 됩니다.
|
||||
핸드오프는 LLM에 도구로 표시됩니다. 따라서 `Refund Agent`라는 에이전트로 핸드오프가 있으면 도구는 `transfer_to_refund_agent`라고 호출됩니다.
|
||||
|
||||
## 핸드오프 생성
|
||||
|
||||
모든 에이전트에는 [`handoffs`][agents.agent.Agent.handoffs] 매개변수가 있으며, 여기에 `Agent`를 직접 전달하거나 핸드오프를 사용자 지정하는 `Handoff` 객체를 전달할 수 있습니다.
|
||||
모든 에이전트에는 [`handoffs`][agents.agent.Agent.handoffs] 매개변수가 있으며, 이 매개변수는 `Agent`를 직접 받거나 핸드오프를 사용자 지정하는 `Handoff` 객체를 받을 수 있습니다.
|
||||
|
||||
일반 `Agent` 인스턴스를 전달하면 해당 [`handoff_description`][agents.agent.Agent.handoff_description] (설정된 경우)이 기본 도구 설명에 추가됩니다. 전체 `handoff()` 객체를 작성하지 않고도 모델이 해당 핸드오프를 선택해야 하는 시점을 힌트로 제공할 때 사용하세요.
|
||||
일반 `Agent` 인스턴스를 전달하면 해당 [`handoff_description`][agents.agent.Agent.handoff_description](설정된 경우)이 기본 도구 설명에 추가됩니다. 전체 `handoff()` 객체를 작성하지 않고도 모델이 해당 핸드오프를 선택해야 하는 시점을 힌트로 제공하는 데 사용하세요.
|
||||
|
||||
Agents SDK가 제공하는 [`handoff()`][agents.handoffs.handoff] 함수를 사용해 핸드오프를 만들 수 있습니다. 이 함수로 핸드오프 대상 에이전트와 선택적 재정의 및 입력 필터를 지정할 수 있습니다.
|
||||
Agents SDK에서 제공하는 [`handoff()`][agents.handoffs.handoff] 함수를 사용하여 핸드오프를 만들 수 있습니다. 이 함수로 핸드오프할 에이전트와 선택적 재정의 및 입력 필터를 지정할 수 있습니다.
|
||||
|
||||
### 기본 사용법
|
||||
|
||||
@@ -30,22 +30,22 @@ refund_agent = Agent(name="Refund agent")
|
||||
triage_agent = Agent(name="Triage agent", handoffs=[billing_agent, handoff(refund_agent)])
|
||||
```
|
||||
|
||||
1. 에이전트를 직접 사용할 수 있고(`billing_agent`처럼), 또는 `handoff()` 함수를 사용할 수 있습니다.
|
||||
1. 에이전트를 직접 사용할 수도 있고(`billing_agent`처럼), `handoff()` 함수를 사용할 수도 있습니다.
|
||||
|
||||
### `handoff()` 함수로 핸드오프 사용자 지정
|
||||
### `handoff()` 함수를 통한 핸드오프 사용자 지정
|
||||
|
||||
[`handoff()`][agents.handoffs.handoff] 함수로 여러 항목을 사용자 지정할 수 있습니다.
|
||||
[`handoff()`][agents.handoffs.handoff] 함수를 사용하면 여러 항목을 사용자 지정할 수 있습니다.
|
||||
|
||||
- `agent`: 핸드오프 대상 에이전트입니다.
|
||||
- `tool_name_override`: 기본적으로 `Handoff.default_tool_name()` 함수가 사용되며, `transfer_to_<agent_name>`으로 해석됩니다. 이를 재정의할 수 있습니다.
|
||||
- `tool_description_override`: `Handoff.default_tool_description()`의 기본 도구 설명을 재정의합니다
|
||||
- `on_handoff`: 핸드오프가 호출될 때 실행되는 콜백 함수입니다. 핸드오프 호출이 확정되는 즉시 데이터 페칭을 시작하는 등의 용도에 유용합니다. 이 함수는 에이전트 컨텍스트를 받으며, 선택적으로 LLM이 생성한 입력도 받을 수 있습니다. 입력 데이터는 `input_type` 매개변수로 제어됩니다.
|
||||
- `input_type`: 핸드오프 도구 호출 인자의 스키마입니다. 설정하면 파싱된 페이로드가 `on_handoff`로 전달됩니다.
|
||||
- `input_filter`: 다음 에이전트가 받는 입력을 필터링할 수 있습니다. 자세한 내용은 아래를 참고하세요.
|
||||
- `is_enabled`: 핸드오프 활성화 여부입니다. 불리언 또는 불리언을 반환하는 함수가 될 수 있어 런타임에 동적으로 핸드오프를 활성화/비활성화할 수 있습니다.
|
||||
- `nest_handoff_history`: RunConfig 수준의 `nest_handoff_history` 설정에 대한 선택적 호출별 재정의입니다. `None`이면 활성 run 설정에 정의된 값을 대신 사용합니다.
|
||||
- `tool_name_override`: 기본적으로 `Handoff.default_tool_name()` 함수가 사용되며, 이는 `transfer_to_<agent_name>`으로 해석됩니다. 이를 재정의할 수 있습니다.
|
||||
- `tool_description_override`: `Handoff.default_tool_description()`의 기본 도구 설명을 재정의합니다.
|
||||
- `on_handoff`: 핸드오프가 호출될 때 실행되는 콜백 함수입니다. 핸드오프가 호출된다는 사실을 알게 되는 즉시 일부 데이터 가져오기를 시작하는 등의 작업에 유용합니다. 이 함수는 에이전트 컨텍스트를 받으며, 선택적으로 LLM이 생성한 입력도 받을 수 있습니다. 입력 데이터는 `input_type` 매개변수로 제어됩니다.
|
||||
- `input_type`: 핸드오프 도구 호출 인수의 스키마입니다. 설정하면 파싱된 페이로드가 `on_handoff`에 전달됩니다.
|
||||
- `input_filter`: 이를 통해 다음 에이전트가 받는 입력을 필터링할 수 있습니다. 자세한 내용은 아래를 참고하세요.
|
||||
- `is_enabled`: 핸드오프가 활성화되어 있는지 여부입니다. 불리언이거나 불리언을 반환하는 함수일 수 있으며, 런타임에 핸드오프를 동적으로 활성화하거나 비활성화할 수 있습니다.
|
||||
- `nest_handoff_history`: RunConfig 수준의 `nest_handoff_history` 설정에 대한 호출별 선택적 재정의입니다. `None`이면 활성 실행 구성에 정의된 값이 대신 사용됩니다.
|
||||
|
||||
[`handoff()`][agents.handoffs.handoff] 헬퍼는 항상 전달한 특정 `agent`로 제어를 넘깁니다. 가능한 대상이 여러 개라면 대상마다 하나의 핸드오프를 등록하고 모델이 그중에서 선택하게 하세요. 호출 시점에 어떤 에이전트를 반환할지 직접 핸드오프 코드에서 결정해야 할 때만 사용자 지정 [`Handoff`][agents.handoffs.Handoff]를 사용하세요.
|
||||
[`handoff()`][agents.handoffs.handoff] 헬퍼는 항상 전달한 특정 `agent`로 제어권을 넘깁니다. 가능한 목적지가 여러 개라면 목적지마다 하나의 핸드오프를 등록하고 모델이 그중에서 선택하도록 하세요. 자체 핸드오프 코드가 호출 시점에 어떤 에이전트를 반환할지 결정해야 하는 경우에만 사용자 지정 [`Handoff`][agents.handoffs.Handoff]를 사용하세요.
|
||||
|
||||
```python
|
||||
from agents import Agent, handoff, RunContextWrapper
|
||||
@@ -65,7 +65,7 @@ handoff_obj = handoff(
|
||||
|
||||
## 핸드오프 입력
|
||||
|
||||
특정 상황에서는 핸드오프를 호출할 때 LLM이 일부 데이터를 제공하도록 하고 싶을 수 있습니다. 예를 들어 "Escalation agent"로 핸드오프한다고 가정해 보겠습니다. 이때 기록을 남기기 위해 사유를 함께 받도록 할 수 있습니다.
|
||||
특정 상황에서는 LLM이 핸드오프를 호출할 때 일부 데이터를 제공하도록 하고 싶을 수 있습니다. 예를 들어 "에스컬레이션 에이전트"로 핸드오프한다고 가정해 보겠습니다. 로그로 남길 수 있도록 모델이 사유를 제공하길 원할 수 있습니다.
|
||||
|
||||
```python
|
||||
from pydantic import BaseModel
|
||||
@@ -87,44 +87,44 @@ handoff_obj = handoff(
|
||||
)
|
||||
```
|
||||
|
||||
`input_type`은 핸드오프 도구 호출 자체의 인자를 설명합니다. SDK는 그 스키마를 핸드오프 도구의 `parameters`로 모델에 노출하고, 반환된 JSON을 로컬에서 검증한 뒤, 파싱된 값을 `on_handoff`에 전달합니다.
|
||||
`input_type`은 핸드오프 도구 호출 자체의 인수를 설명합니다. SDK는 해당 스키마를 핸드오프 도구의 `parameters`로 모델에 노출하고, 반환된 JSON을 로컬에서 검증한 뒤 파싱된 값을 `on_handoff`에 전달합니다.
|
||||
|
||||
이는 다음 에이전트의 기본 입력을 대체하지 않으며, 다른 목적지를 선택하지도 않습니다. [`handoff()`][agents.handoffs.handoff] 헬퍼는 여전히 래핑한 특정 에이전트로 전송하며, 수신 에이전트는 [`input_filter`][agents.handoffs.Handoff.input_filter] 또는 중첩 핸드오프 기록 설정으로 변경하지 않는 한 대화 기록을 계속 확인합니다.
|
||||
이는 다음 에이전트의 기본 입력을 대체하지 않으며, 다른 목적지를 선택하지도 않습니다. [`handoff()`][agents.handoffs.handoff] 헬퍼는 여전히 래핑한 특정 에이전트로 전달하며, [`input_filter`][agents.handoffs.Handoff.input_filter] 또는 중첩 핸드오프 기록 설정으로 변경하지 않는 한 수신 에이전트는 여전히 대화 기록을 보게 됩니다.
|
||||
|
||||
`input_type`은 [`RunContextWrapper.context`][agents.run_context.RunContextWrapper.context]와도 별개입니다. 이미 로컬에 있는 애플리케이션 상태나 의존성이 아니라, 모델이 핸드오프 시점에 결정하는 메타데이터에 `input_type`을 사용하세요.
|
||||
`input_type`은 [`RunContextWrapper.context`][agents.run_context.RunContextWrapper.context]와도 별개입니다. 로컬에 이미 있는 애플리케이션 상태나 의존성이 아니라, 핸드오프 시점에 모델이 결정하는 메타데이터에 `input_type`을 사용하세요.
|
||||
|
||||
### `input_type` 사용 시점
|
||||
|
||||
핸드오프에 `reason`, `language`, `priority`, `summary` 같은 모델 생성 메타데이터의 작은 조각이 필요할 때 `input_type`을 사용하세요. 예를 들어 트리아지 에이전트는 `{ "reason": "duplicate_charge", "priority": "high" }`와 함께 환불 에이전트로 핸드오프할 수 있으며, `on_handoff`는 환불 에이전트가 이어받기 전에 해당 메타데이터를 기록하거나 저장할 수 있습니다.
|
||||
핸드오프에 `reason`, `language`, `priority`, `summary`와 같은 작은 규모의 모델 생성 메타데이터가 필요할 때 `input_type`을 사용하세요. 예를 들어 분류 에이전트는 `{ "reason": "duplicate_charge", "priority": "high" }`와 함께 환불 에이전트로 핸드오프할 수 있으며, `on_handoff`는 환불 에이전트가 이어받기 전에 해당 메타데이터를 로그로 남기거나 영속화할 수 있습니다.
|
||||
|
||||
목적이 다르면 다른 메커니즘을 선택하세요:
|
||||
목표가 다를 경우에는 다른 메커니즘을 선택하세요:
|
||||
|
||||
- 기존 애플리케이션 상태와 의존성은 [`RunContextWrapper.context`][agents.run_context.RunContextWrapper.context]에 넣으세요. [컨텍스트 가이드](context.md)를 참고하세요.
|
||||
- 수신 에이전트가 보는 기록을 바꾸려면 [`input_filter`][agents.handoffs.Handoff.input_filter], [`RunConfig.nest_handoff_history`][agents.run.RunConfig.nest_handoff_history], 또는 [`RunConfig.handoff_history_mapper`][agents.run.RunConfig.handoff_history_mapper]를 사용하세요.
|
||||
- 가능한 전문 에이전트 대상이 여러 개라면 대상마다 하나의 핸드오프를 등록하세요. `input_type`은 선택된 핸드오프에 메타데이터를 추가할 수는 있지만, 대상 간 디스패치를 수행하지는 않습니다.
|
||||
- 대화를 전송하지 않고 중첩 전문 에이전트에 구조화된 입력을 주고 싶다면 [`Agent.as_tool(parameters=...)`][agents.agent.Agent.as_tool]을 우선 사용하세요. [도구](tools.md#structured-input-for-tool-agents)를 참고하세요.
|
||||
- 수신 에이전트가 보게 되는 기록을 변경하려면 [`input_filter`][agents.handoffs.Handoff.input_filter], [`RunConfig.nest_handoff_history`][agents.run.RunConfig.nest_handoff_history] 또는 [`RunConfig.handoff_history_mapper`][agents.run.RunConfig.handoff_history_mapper]를 사용하세요.
|
||||
- 가능한 전문 에이전트가 여러 개라면 목적지마다 하나의 핸드오프를 등록하세요. `input_type`은 선택된 핸드오프에 메타데이터를 추가할 수 있지만, 목적지 간 라우팅을 수행하지는 않습니다.
|
||||
- 대화를 이전하지 않고 중첩된 전문 에이전트에 구조화된 입력을 제공하려면 [`Agent.as_tool(parameters=...)`][agents.agent.Agent.as_tool]를 우선 사용하세요. [도구](tools.md#structured-input-for-tool-agents)를 참고하세요.
|
||||
|
||||
## 입력 필터
|
||||
|
||||
핸드오프가 발생하면 새 에이전트가 대화를 이어받아 이전 전체 대화 기록을 보는 것과 같습니다. 이를 변경하려면 [`input_filter`][agents.handoffs.Handoff.input_filter]를 설정할 수 있습니다. 입력 필터는 [`HandoffInputData`][agents.handoffs.HandoffInputData]를 통해 기존 입력을 받고, 새로운 `HandoffInputData`를 반환해야 하는 함수입니다.
|
||||
핸드오프가 발생하면 새 에이전트가 대화를 이어받는 것과 같으며, 이전 대화 기록 전체를 볼 수 있습니다. 이를 변경하려면 [`input_filter`][agents.handoffs.Handoff.input_filter]를 설정할 수 있습니다. 입력 필터는 [`HandoffInputData`][agents.handoffs.HandoffInputData]를 통해 기존 입력을 받는 함수이며, 새 `HandoffInputData`를 반환해야 합니다.
|
||||
|
||||
[`HandoffInputData`][agents.handoffs.HandoffInputData]에는 다음이 포함됩니다:
|
||||
|
||||
- `input_history`: `Runner.run(...)` 시작 전의 입력 기록
|
||||
- `pre_handoff_items`: 핸드오프가 호출된 에이전트 턴 이전에 생성된 항목
|
||||
- `new_items`: 핸드오프 호출 및 핸드오프 출력 항목을 포함해 현재 턴에서 생성된 항목
|
||||
- `input_items`: `new_items` 대신 다음 에이전트로 전달할 선택적 항목으로, 세션 기록용 `new_items`는 유지하면서 모델 입력을 필터링할 수 있게 해줍니다
|
||||
- `run_context`: 핸드오프 호출 시점의 활성 [`RunContextWrapper`][agents.run_context.RunContextWrapper]
|
||||
- `input_history`: `Runner.run(...)`이 시작되기 전의 입력 기록입니다.
|
||||
- `pre_handoff_items`: 핸드오프가 호출된 에이전트 턴 이전에 생성된 항목입니다.
|
||||
- `new_items`: 현재 턴 중 생성된 항목이며, 핸드오프 호출과 핸드오프 출력 항목을 포함합니다.
|
||||
- `input_items`: `new_items` 대신 다음 에이전트에 전달할 선택적 항목입니다. 이를 통해 세션 기록용으로 `new_items`는 그대로 유지하면서 모델 입력을 필터링할 수 있습니다.
|
||||
- `run_context`: 핸드오프가 호출된 시점의 활성 [`RunContextWrapper`][agents.run_context.RunContextWrapper]입니다.
|
||||
|
||||
중첩 핸드오프는 옵트인 베타로 제공되며 안정화 중이므로 기본적으로 비활성화되어 있습니다. [`RunConfig.nest_handoff_history`][agents.run.RunConfig.nest_handoff_history]를 활성화하면 러너는 이전 전사를 단일 어시스턴트 요약 메시지로 축약하고, 동일 run에서 여러 핸드오프가 발생할 때 새 턴이 계속 추가되도록 `<CONVERSATION HISTORY>` 블록으로 감쌉니다. 전체 `input_filter`를 작성하지 않고 생성된 메시지를 대체하려면 [`RunConfig.handoff_history_mapper`][agents.run.RunConfig.handoff_history_mapper]를 통해 자체 매핑 함수를 제공할 수 있습니다. 이 옵트인은 핸드오프와 run 어느 쪽에서도 명시적 `input_filter`를 제공하지 않을 때만 적용되므로, 이미 페이로드를 사용자 지정하는 기존 코드(이 저장소의 예제 포함)는 변경 없이 현재 동작을 유지합니다. [`handoff(...)`][agents.handoffs.handoff]에 `nest_handoff_history=True` 또는 `False`를 전달해 단일 핸드오프의 중첩 동작을 재정의할 수 있으며, 이는 [`Handoff.nest_handoff_history`][agents.handoffs.Handoff.nest_handoff_history]를 설정합니다. 생성된 요약의 래퍼 텍스트만 바꾸면 된다면 에이전트를 실행하기 전에 [`set_conversation_history_wrappers`][agents.handoffs.set_conversation_history_wrappers] (및 선택적으로 [`reset_conversation_history_wrappers`][agents.handoffs.reset_conversation_history_wrappers])를 호출하세요.
|
||||
중첩 핸드오프는 명시적으로 활성화해야 하는 베타 기능으로 제공되며, 안정화하는 동안 기본적으로 비활성화되어 있습니다. [`RunConfig.nest_handoff_history`][agents.run.RunConfig.nest_handoff_history]를 활성화하면 러너는 이전 대화 기록을 하나의 어시스턴트 요약 메시지로 압축하고, 동일한 실행 중 여러 핸드오프가 발생할 때 새 턴을 계속 추가하는 `<CONVERSATION HISTORY>` 블록으로 감쌉니다. 전체 `input_filter`를 작성하지 않고도 [`RunConfig.handoff_history_mapper`][agents.run.RunConfig.handoff_history_mapper]를 통해 자체 매핑 함수를 제공하여 생성된 메시지를 대체할 수 있습니다. 이 명시적 활성화는 핸드오프와 실행 모두 명시적 `input_filter`를 제공하지 않는 경우에만 적용되므로, 이미 페이로드를 사용자 지정하는 기존 코드(이 저장소의 코드 예제를 포함)는 변경 없이 현재 동작을 유지합니다. 단일 핸드오프에 대해서는 [`handoff(...)`][agents.handoffs.handoff]에 `nest_handoff_history=True` 또는 `False`를 전달하여 중첩 동작을 재정의할 수 있으며, 이는 [`Handoff.nest_handoff_history`][agents.handoffs.Handoff.nest_handoff_history]를 설정합니다. 생성된 요약의 래퍼 텍스트만 변경하면 된다면, 에이전트를 실행하기 전에 [`set_conversation_history_wrappers`][agents.handoffs.set_conversation_history_wrappers]를 호출하세요(그리고 선택적으로 [`reset_conversation_history_wrappers`][agents.handoffs.reset_conversation_history_wrappers]도 호출할 수 있습니다).
|
||||
|
||||
핸드오프와 활성 [`RunConfig.handoff_input_filter`][agents.run.RunConfig.handoff_input_filter] 양쪽 모두 필터를 정의한 경우, 해당 핸드오프에는 핸드오프별 [`input_filter`][agents.handoffs.Handoff.input_filter]가 우선 적용됩니다.
|
||||
핸드오프와 활성 [`RunConfig.handoff_input_filter`][agents.run.RunConfig.handoff_input_filter]가 모두 필터를 정의하는 경우, 해당 특정 핸드오프에는 핸드오프별 [`input_filter`][agents.handoffs.Handoff.input_filter]가 우선합니다.
|
||||
|
||||
!!! note
|
||||
|
||||
핸드오프는 단일 run 내에서만 유지됩니다. 입력 가드레일은 체인의 첫 번째 에이전트에만 계속 적용되고, 출력 가드레일은 최종 출력을 생성하는 에이전트에만 적용됩니다. 워크플로 내 각 사용자 지정 함수 도구 호출 주변에서 검사가 필요하다면 도구 가드레일을 사용하세요.
|
||||
핸드오프는 단일 실행 내에 머뭅니다. 입력 가드레일은 여전히 체인의 첫 번째 에이전트에만 적용되고, 출력 가드레일은 최종 출력을 생성하는 에이전트에만 적용됩니다. 워크플로 내부의 각 사용자 지정 함수 도구 호출에 대한 검사가 필요할 때는 도구 가드레일을 사용하세요.
|
||||
|
||||
일부 일반 패턴(예: 기록에서 모든 도구 호출 제거)은 [`agents.extensions.handoff_filters`][]에 구현되어 있습니다
|
||||
몇 가지 일반적인 패턴(예: 기록에서 모든 도구 호출 제거)은 [`agents.extensions.handoff_filters`][]에 구현되어 있습니다
|
||||
|
||||
```python
|
||||
from agents import Agent, handoff
|
||||
@@ -142,7 +142,7 @@ handoff_obj = handoff(
|
||||
|
||||
## 권장 프롬프트
|
||||
|
||||
LLM이 핸드오프를 올바르게 이해하도록 하려면, 에이전트에 핸드오프 관련 정보를 포함할 것을 권장합니다. [`agents.extensions.handoff_prompt.RECOMMENDED_PROMPT_PREFIX`][]에 권장 접두사가 있으며, [`agents.extensions.handoff_prompt.prompt_with_handoff_instructions`][]를 호출해 프롬프트에 권장 데이터를 자동으로 추가할 수도 있습니다.
|
||||
LLM이 핸드오프를 올바르게 이해하도록 하려면, 에이전트에 핸드오프 관련 정보를 포함하는 것을 권장합니다. 제안된 접두사는 [`agents.extensions.handoff_prompt.RECOMMENDED_PROMPT_PREFIX`][]에 있으며, 또는 [`agents.extensions.handoff_prompt.prompt_with_handoff_instructions`][]를 호출하여 프롬프트에 권장 데이터를 자동으로 추가할 수 있습니다.
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
|
||||
@@ -4,17 +4,17 @@ search:
|
||||
---
|
||||
# 휴먼인더루프 (HITL)
|
||||
|
||||
휴먼인더루프 (HITL) 흐름을 사용해 민감한 도구 호출을 사람이 승인하거나 거절할 때까지 에이전트 실행을 일시 중지할 수 있습니다. 도구는 승인 필요 여부를 선언하고, 실행 결과는 대기 중인 승인을 인터럽션으로 노출하며, `RunState`를 통해 결정 이후 실행을 직렬화하고 재개할 수 있습니다
|
||||
휴먼인더루프 (HITL) 흐름을 사용해 사람이 민감한 도구 호출을 승인하거나 거부할 때까지 에이전트 실행을 일시 중지합니다. 도구는 승인이 필요한 시점을 선언하고, 실행 결과는 보류 중인 승인을 인터럽션(중단 처리)으로 노출하며, `RunState`를 사용하면 결정이 내려진 뒤 실행을 직렬화하고 재개할 수 있습니다.
|
||||
|
||||
이 승인 표면은 현재 최상위 에이전트로 제한되지 않고 실행 전체에 적용됩니다. 동일한 패턴은 도구가 현재 에이전트에 속한 경우, 핸드오프를 통해 도달한 에이전트에 속한 경우, 또는 중첩된 [`Agent.as_tool()`][agents.agent.Agent.as_tool] 실행에 속한 경우에도 적용됩니다. 중첩된 `Agent.as_tool()`의 경우에도 인터럽션은 바깥 실행에 나타나므로, 바깥 `RunState`에서 승인 또는 거절하고 원래 최상위 실행을 재개합니다
|
||||
이 승인 처리는 실행 전체 범위에 적용되며, 현재 최상위 에이전트로 제한되지 않습니다. 도구가 현재 에이전트에 속한 경우, 핸드오프로 도달한 에이전트에 속한 경우, 또는 중첩된 [`Agent.as_tool()`][agents.agent.Agent.as_tool] 실행에 속한 경우에도 같은 패턴이 적용됩니다. 중첩된 `Agent.as_tool()`의 경우에도 인터럽션(중단 처리)은 외부 실행에 노출되므로, 외부 `RunState`에서 승인하거나 거부한 뒤 원래 최상위 실행을 재개합니다.
|
||||
|
||||
`Agent.as_tool()`에서는 서로 다른 두 계층에서 승인이 발생할 수 있습니다: 에이전트 도구 자체가 `Agent.as_tool(..., needs_approval=...)`를 통해 승인을 요구할 수 있고, 중첩된 실행이 시작된 뒤에는 중첩 에이전트 내부 도구가 자체 승인을 다시 요청할 수 있습니다. 둘 다 동일한 바깥 실행 인터럽션 흐름으로 처리됩니다
|
||||
`Agent.as_tool()`에서는 두 가지 계층에서 승인이 발생할 수 있습니다. 에이전트 도구 자체가 `Agent.as_tool(..., needs_approval=...)`를 통해 승인을 요구할 수 있고, 중첩된 에이전트 내부의 도구가 중첩 실행이 시작된 뒤 자체 승인 요청을 나중에 발생시킬 수 있습니다. 둘 다 동일한 외부 실행 인터럽션(중단 처리) 흐름을 통해 처리됩니다.
|
||||
|
||||
이 페이지는 `interruptions`를 통한 수동 승인 흐름에 중점을 둡니다. 앱에서 코드로 판단할 수 있다면, 일부 도구 유형은 프로그래매틱 승인 콜백도 지원하므로 실행을 멈추지 않고 계속할 수 있습니다
|
||||
이 페이지는 `interruptions`를 통한 수동 승인 흐름에 중점을 둡니다. 앱이 코드로 결정을 내릴 수 있다면, 일부 도구 유형은 실행을 일시 중지하지 않고 계속 진행할 수 있도록 프로그래밍 방식 승인 콜백도 지원합니다.
|
||||
|
||||
## 승인 필요 도구 표시
|
||||
## 승인이 필요한 도구 표시
|
||||
|
||||
항상 승인을 요구하려면 `needs_approval`를 `True`로 설정하거나, 호출별로 판단하는 비동기 함수를 제공하세요. 호출 가능 객체는 실행 컨텍스트, 파싱된 도구 매개변수, 도구 호출 ID를 받습니다
|
||||
항상 승인을 요구하려면 `needs_approval`을 `True`로 설정하거나, 호출마다 결정하는 비동기 함수를 제공합니다. 호출 가능한 함수는 실행 컨텍스트, 파싱된 도구 매개변수, 도구 호출 ID를 받습니다.
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, function_tool
|
||||
@@ -41,28 +41,28 @@ agent = Agent(
|
||||
)
|
||||
```
|
||||
|
||||
`needs_approval`는 [`function_tool`][agents.tool.function_tool], [`Agent.as_tool`][agents.agent.Agent.as_tool], [`ShellTool`][agents.tool.ShellTool], [`ApplyPatchTool`][agents.tool.ApplyPatchTool]에서 사용할 수 있습니다. 로컬 MCP 서버도 [`MCPServerStdio`][agents.mcp.server.MCPServerStdio], [`MCPServerSse`][agents.mcp.server.MCPServerSse], [`MCPServerStreamableHttp`][agents.mcp.server.MCPServerStreamableHttp]의 `require_approval`를 통해 승인을 지원합니다. 호스티드 MCP 서버는 [`HostedMCPTool`][agents.tool.HostedMCPTool]에서 `tool_config={"require_approval": "always"}`와 선택적 `on_approval_request` 콜백으로 승인을 지원합니다. Shell 및 apply_patch 도구는 인터럽션을 노출하지 않고 자동 승인 또는 자동 거절하려는 경우 `on_approval` 콜백을 받을 수 있습니다
|
||||
`needs_approval`은 [`function_tool`][agents.tool.function_tool], [`Agent.as_tool`][agents.agent.Agent.as_tool], [`ShellTool`][agents.tool.ShellTool], [`ApplyPatchTool`][agents.tool.ApplyPatchTool]에서 사용할 수 있습니다. 로컬 MCP 서버도 [`MCPServerStdio`][agents.mcp.server.MCPServerStdio], [`MCPServerSse`][agents.mcp.server.MCPServerSse], [`MCPServerStreamableHttp`][agents.mcp.server.MCPServerStreamableHttp]의 `require_approval`을 통해 승인을 지원합니다. 호스티드 MCP 서버는 `tool_config={"require_approval": "always"}`와 선택적 `on_approval_request` 콜백을 사용해 [`HostedMCPTool`][agents.tool.HostedMCPTool]을 통한 승인을 지원합니다. 셸 및 apply_patch 도구는 인터럽션(중단 처리)을 노출하지 않고 자동 승인 또는 자동 거부하려는 경우 `on_approval` 콜백을 허용합니다.
|
||||
|
||||
## 승인 흐름 작동 방식
|
||||
## 승인 흐름의 작동 방식
|
||||
|
||||
1. 모델이 도구 호출을 생성하면 러너는 해당 도구의 승인 규칙(`needs_approval`, `require_approval`, 또는 호스티드 MCP 동등 설정)을 평가합니다
|
||||
2. 해당 도구 호출에 대한 승인 결정이 이미 [`RunContextWrapper`][agents.run_context.RunContextWrapper]에 저장되어 있으면, 러너는 추가 확인 없이 진행합니다. 호출별 승인은 특정 호출 ID 범위에만 적용됩니다. 실행의 나머지 동안 같은 도구의 향후 호출에도 동일한 결정을 유지하려면 `always_approve=True` 또는 `always_reject=True`를 전달하세요
|
||||
3. 그렇지 않으면 실행이 일시 중지되고 `RunResult.interruptions`(또는 `RunResultStreaming.interruptions`)에 `agent.name`, `tool_name`, `arguments` 같은 세부 정보를 담은 [`ToolApprovalItem`][agents.items.ToolApprovalItem] 항목이 포함됩니다. 여기에는 핸드오프 이후 또는 중첩 `Agent.as_tool()` 실행 내부에서 발생한 승인도 포함됩니다
|
||||
4. `result.to_state()`로 결과를 `RunState`로 변환하고, `state.approve(...)` 또는 `state.reject(...)`를 호출한 뒤, `Runner.run(agent, state)` 또는 `Runner.run_streamed(agent, state)`로 재개하세요. 여기서 `agent`는 해당 실행의 원래 최상위 에이전트입니다
|
||||
5. 재개된 실행은 중단된 지점부터 계속되며, 새 승인이 필요하면 이 흐름으로 다시 진입합니다
|
||||
1. 모델이 도구 호출을 내보내면, 러너는 해당 승인 규칙(`needs_approval`, `require_approval` 또는 호스티드 MCP 대응 기능)을 평가합니다.
|
||||
2. 해당 도구 호출에 대한 승인 결정이 이미 [`RunContextWrapper`][agents.run_context.RunContextWrapper]에 저장되어 있으면, 러너는 프롬프트를 표시하지 않고 진행합니다. 호출별 승인은 특정 호출 ID 범위로 한정됩니다. 실행의 나머지 동안 해당 도구에 대한 이후 호출에도 같은 결정을 유지하려면 `always_approve=True` 또는 `always_reject=True`를 전달합니다.
|
||||
3. 그렇지 않으면 실행이 일시 중지되고 `RunResult.interruptions`(또는 `RunResultStreaming.interruptions`)에 `agent.name`, `tool_name`, `arguments` 같은 세부 정보가 포함된 [`ToolApprovalItem`][agents.items.ToolApprovalItem] 항목이 들어갑니다. 여기에는 핸드오프 이후 또는 중첩된 `Agent.as_tool()` 실행 내부에서 발생한 승인도 포함됩니다.
|
||||
4. `result.to_state()`로 결과를 `RunState`로 변환하고, `state.approve(...)` 또는 `state.reject(...)`를 호출한 다음, `Runner.run(agent, state)` 또는 `Runner.run_streamed(agent, state)`로 재개합니다. 여기서 `agent`는 해당 실행의 원래 최상위 에이전트입니다.
|
||||
5. 재개된 실행은 중단된 지점부터 계속 진행되며, 새 승인이 필요하면 이 흐름으로 다시 들어갑니다.
|
||||
|
||||
`always_approve=True` 또는 `always_reject=True`로 생성된 고정 결정은 실행 상태에 저장되므로, 나중에 동일한 일시 중지 실행을 재개할 때 `state.to_string()` / `RunState.from_string(...)` 및 `state.to_json()` / `RunState.from_json(...)`을 거쳐도 유지됩니다
|
||||
`always_approve=True` 또는 `always_reject=True`로 생성된 고정 결정은 실행 상태에 저장되므로, 나중에 같은 일시 중지된 실행을 재개할 때 `state.to_string()` / `RunState.from_string(...)` 및 `state.to_json()` / `RunState.from_json(...)` 이후에도 유지됩니다.
|
||||
|
||||
같은 패스에서 모든 대기 중 승인을 처리할 필요는 없습니다. `interruptions`에는 일반 함수 도구, 호스티드 MCP 승인, 중첩 `Agent.as_tool()` 승인이 혼합되어 있을 수 있습니다. 일부 항목만 승인 또는 거절한 뒤 다시 실행하면, 해결된 호출은 계속 진행되고 미해결 항목은 `interruptions`에 남아 실행을 다시 일시 중지합니다
|
||||
보류 중인 모든 승인을 같은 단계에서 해결할 필요는 없습니다. `interruptions`에는 일반 함수 도구, 호스티드 MCP 승인, 중첩된 `Agent.as_tool()` 승인이 섞여 있을 수 있습니다. 일부 항목만 승인하거나 거부한 뒤 다시 실행하면, 해결된 호출은 계속 진행될 수 있고 해결되지 않은 호출은 `interruptions`에 남아 실행을 다시 일시 중지합니다.
|
||||
|
||||
## 사용자 지정 거절 메시지
|
||||
## 사용자 지정 거부 메시지
|
||||
|
||||
기본적으로 거절된 도구 호출은 SDK의 표준 거절 텍스트를 실행으로 다시 반환합니다. 이 메시지는 두 계층에서 사용자 지정할 수 있습니다
|
||||
기본적으로 거부된 도구 호출은 SDK의 표준 거부 텍스트를 실행으로 다시 반환합니다. 이 메시지는 두 계층에서 사용자 지정할 수 있습니다.
|
||||
|
||||
- 실행 전체 대체값: [`RunConfig.tool_error_formatter`][agents.run.RunConfig.tool_error_formatter]를 설정해 실행 전체의 승인 거절에 대한 기본 모델 표시 메시지를 제어합니다
|
||||
- 호출별 재정의: 특정 거절 도구 호출에 다른 메시지를 노출하려면 `state.reject(...)`에 `rejection_message=...`를 전달합니다
|
||||
- 실행 전체 폴백: 전체 실행에서 승인 거부에 대해 모델에 표시되는 기본 메시지를 제어하려면 [`RunConfig.tool_error_formatter`][agents.run.RunConfig.tool_error_formatter]를 설정합니다.
|
||||
- 호출별 재정의: 특정 거부된 도구 호출 하나에 다른 메시지를 노출하려면 `state.reject(...)`에 `rejection_message=...`를 전달합니다.
|
||||
|
||||
둘 다 제공되면 호출별 `rejection_message`가 실행 전체 포매터보다 우선합니다
|
||||
둘 다 제공되면 호출별 `rejection_message`가 실행 전체 포매터보다 우선합니다.
|
||||
|
||||
```python
|
||||
from agents import RunConfig, ToolErrorFormatterArgs
|
||||
@@ -83,27 +83,27 @@ state.reject(
|
||||
)
|
||||
```
|
||||
|
||||
두 계층을 함께 보여주는 완전한 예시는 [`examples/agent_patterns/human_in_the_loop_custom_rejection.py`](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns/human_in_the_loop_custom_rejection.py)를 참조하세요
|
||||
두 계층을 함께 보여 주는 전체 코드 예제는 [`examples/agent_patterns/human_in_the_loop_custom_rejection.py`](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns/human_in_the_loop_custom_rejection.py)를 참고하세요.
|
||||
|
||||
## 자동 승인 결정
|
||||
|
||||
수동 `interruptions`가 가장 일반적인 패턴이지만 유일한 방법은 아닙니다
|
||||
수동 `interruptions`가 가장 일반적인 패턴이지만, 유일한 방식은 아닙니다.
|
||||
|
||||
- 로컬 [`ShellTool`][agents.tool.ShellTool] 및 [`ApplyPatchTool`][agents.tool.ApplyPatchTool]은 `on_approval`을 사용해 코드에서 즉시 승인 또는 거절할 수 있습니다
|
||||
- [`HostedMCPTool`][agents.tool.HostedMCPTool]은 `tool_config={"require_approval": "always"}`와 `on_approval_request`를 함께 사용해 같은 유형의 프로그래매틱 결정을 내릴 수 있습니다
|
||||
- 일반 [`function_tool`][agents.tool.function_tool] 도구와 [`Agent.as_tool()`][agents.agent.Agent.as_tool]은 이 페이지의 수동 인터럽션 흐름을 사용합니다
|
||||
- 로컬 [`ShellTool`][agents.tool.ShellTool] 및 [`ApplyPatchTool`][agents.tool.ApplyPatchTool]은 코드에서 즉시 승인하거나 거부하기 위해 `on_approval`을 사용할 수 있습니다.
|
||||
- [`HostedMCPTool`][agents.tool.HostedMCPTool]은 같은 유형의 프로그래밍 방식 결정을 위해 `tool_config={"require_approval": "always"}`를 `on_approval_request`와 함께 사용할 수 있습니다.
|
||||
- 일반 [`function_tool`][agents.tool.function_tool] 도구와 [`Agent.as_tool()`][agents.agent.Agent.as_tool]는 이 페이지의 수동 인터럽션(중단 처리) 흐름을 사용합니다.
|
||||
|
||||
이 콜백들이 결정을 반환하면 실행은 사람 응답을 기다리며 멈추지 않고 계속됩니다. Realtime 및 음성 세션 API의 경우 [Realtime 가이드](realtime/guide.md)의 승인 흐름을 참조하세요
|
||||
이러한 콜백이 결정을 반환하면, 실행은 사람의 응답을 기다리기 위해 일시 중지하지 않고 계속됩니다. Realtime 및 음성 세션 API의 경우 [Realtime 가이드](realtime/guide.md)의 승인 흐름을 참고하세요.
|
||||
|
||||
## 스트리밍 및 세션
|
||||
## 스트리밍과 세션
|
||||
|
||||
동일한 인터럽션 흐름은 스트리밍 실행에서도 동작합니다. 스트리밍 실행이 일시 중지된 뒤에는 반복자가 끝날 때까지 [`RunResultStreaming.stream_events()`][agents.result.RunResultStreaming.stream_events]를 계속 소비하고, [`RunResultStreaming.interruptions`][agents.result.RunResultStreaming.interruptions]를 확인해 해결한 다음, 재개 출력도 계속 스트리밍하려면 [`Runner.run_streamed(...)`][agents.run.Runner.run_streamed]로 재개하세요. 이 패턴의 스트리밍 버전은 [스트리밍](streaming.md)을 참조하세요
|
||||
동일한 인터럽션(중단 처리) 흐름은 스트리밍 실행에서도 작동합니다. 스트리밍 실행이 일시 중지된 뒤에는 반복자가 끝날 때까지 [`RunResultStreaming.stream_events()`][agents.result.RunResultStreaming.stream_events]를 계속 소비하고, [`RunResultStreaming.interruptions`][agents.result.RunResultStreaming.interruptions]를 검사해 해결한 다음, 재개된 출력도 계속 스트리밍되게 하려면 [`Runner.run_streamed(...)`][agents.run.Runner.run_streamed]로 재개합니다. 이 패턴의 스트리밍 버전은 [스트리밍](streaming.md)을 참고하세요.
|
||||
|
||||
세션도 함께 사용 중이라면 `RunState`에서 재개할 때 동일한 세션 인스턴스를 계속 전달하거나, 같은 백엔드 스토어를 가리키는 다른 세션 객체를 전달하세요. 그러면 재개된 턴이 같은 저장 대화 기록에 추가됩니다. 세션 수명주기 상세는 [세션](sessions/index.md)을 참조하세요
|
||||
세션도 함께 사용하는 경우 `RunState`에서 재개할 때 같은 세션 인스턴스를 계속 전달하거나, 동일한 백킹 스토어를 가리키는 다른 세션 객체를 전달합니다. 그러면 재개된 턴이 동일하게 저장된 대화 기록에 추가됩니다. 세션 수명 주기 세부 정보는 [세션](sessions/index.md)을 참고하세요.
|
||||
|
||||
## 예시: 일시 중지, 승인, 재개
|
||||
## 예제: 일시 중지, 승인, 재개
|
||||
|
||||
아래 스니펫은 JavaScript HITL 가이드를 반영합니다: 도구에 승인이 필요하면 일시 중지하고, 상태를 디스크에 저장했다가, 다시 불러와 결정 수집 후 재개합니다
|
||||
아래 스니펫은 JavaScript HITL 가이드와 동일한 흐름을 따릅니다. 도구에 승인이 필요할 때 일시 중지하고, 상태를 디스크에 저장한 뒤, 다시 로드하고, 결정을 수집한 후 재개합니다.
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -167,35 +167,35 @@ if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
이 예시에서 `prompt_approval`는 `input()`을 사용하고 `run_in_executor(...)`로 실행되므로 동기식입니다. 승인 소스가 이미 비동기(예: HTTP 요청 또는 비동기 데이터베이스 쿼리)라면 `async def` 함수를 사용해 대신 직접 `await`할 수 있습니다
|
||||
이 예제에서 `prompt_approval`은 `input()`을 사용하고 `run_in_executor(...)`로 실행되기 때문에 동기 함수입니다. 승인 소스가 이미 비동기인 경우(예: HTTP 요청 또는 비동기 데이터베이스 쿼리), 대신 `async def` 함수를 사용하고 직접 `await`할 수 있습니다.
|
||||
|
||||
승인 대기 중에도 출력을 스트리밍하려면 `Runner.run_streamed`를 호출하고, `result.stream_events()`를 완료될 때까지 소비한 다음, 위에 나온 동일한 `result.to_state()` 및 재개 단계를 따르세요
|
||||
승인을 기다리는 동안 출력을 스트리밍하려면 `Runner.run_streamed`를 호출하고, 완료될 때까지 `result.stream_events()`를 소비한 다음, 위에 표시된 것과 동일한 `result.to_state()` 및 재개 단계를 따릅니다.
|
||||
|
||||
## 저장소 패턴 및 예제
|
||||
## 리포지토리 패턴과 코드 예제
|
||||
|
||||
- **스트리밍 승인**: `examples/agent_patterns/human_in_the_loop_stream.py`는 `stream_events()`를 모두 소비한 뒤 대기 중인 도구 호출을 승인하고 `Runner.run_streamed(agent, state)`로 재개하는 방법을 보여줍니다
|
||||
- **사용자 지정 거절 텍스트**: `examples/agent_patterns/human_in_the_loop_custom_rejection.py`는 승인이 거절될 때 실행 수준 `tool_error_formatter`와 호출별 `rejection_message` 재정의를 결합하는 방법을 보여줍니다
|
||||
- **도구로서의 에이전트 승인**: `Agent.as_tool(..., needs_approval=...)`는 위임된 에이전트 작업에 검토가 필요할 때 동일한 인터럽션 흐름을 적용합니다. 중첩 인터럽션도 바깥 실행에 노출되므로 중첩 에이전트가 아니라 원래 최상위 에이전트를 재개하세요
|
||||
- **로컬 shell 및 apply_patch 도구**: `ShellTool`과 `ApplyPatchTool`도 `needs_approval`를 지원합니다. 향후 호출에 대한 결정을 캐시하려면 `state.approve(interruption, always_approve=True)` 또는 `state.reject(..., always_reject=True)`를 사용하세요. 자동 결정을 위해서는 `on_approval`를 제공하고(`examples/tools/shell.py` 참조), 수동 결정을 위해서는 인터럽션을 처리하세요(`examples/tools/shell_human_in_the_loop.py` 참조). 호스티드 shell 환경은 `needs_approval` 또는 `on_approval`를 지원하지 않습니다. [도구 가이드](tools.md)를 참조하세요
|
||||
- **로컬 MCP 서버**: MCP 도구 호출을 제어하려면 `MCPServerStdio` / `MCPServerSse` / `MCPServerStreamableHttp`에서 `require_approval`를 사용하세요(`examples/mcp/get_all_mcp_tools_example/main.py`, `examples/mcp/tool_filter_example/main.py` 참조)
|
||||
- **호스티드 MCP 서버**: HITL을 강제하려면 `HostedMCPTool`에서 `require_approval`를 `"always"`로 설정하고, 필요 시 `on_approval_request`를 제공해 자동 승인 또는 거절할 수 있습니다(`examples/hosted_mcp/human_in_the_loop.py`, `examples/hosted_mcp/on_approval.py` 참조). 신뢰 가능한 서버에는 `"never"`를 사용하세요(`examples/hosted_mcp/simple.py`)
|
||||
- **세션 및 메모리**: 승인과 대화 기록이 여러 턴에 걸쳐 유지되도록 `Runner.run`에 세션을 전달하세요. SQLite 및 OpenAI Conversations 세션 변형은 `examples/memory/memory_session_hitl_example.py`와 `examples/memory/openai_session_hitl_example.py`에 있습니다
|
||||
- **실시간 에이전트**: realtime 데모는 `RealtimeSession`의 `approve_tool_call` / `reject_tool_call`을 통해 도구 호출을 승인 또는 거절하는 WebSocket 메시지를 노출합니다(서버 측 핸들러는 `examples/realtime/app/server.py`, API 표면은 [Realtime 가이드](realtime/guide.md#tool-approvals) 참조)
|
||||
- **스트리밍 승인**: `examples/agent_patterns/human_in_the_loop_stream.py`는 `stream_events()`를 모두 소비한 다음, `Runner.run_streamed(agent, state)`로 재개하기 전에 보류 중인 도구 호출을 승인하는 방법을 보여 줍니다.
|
||||
- **사용자 지정 거부 텍스트**: `examples/agent_patterns/human_in_the_loop_custom_rejection.py`는 승인이 거부될 때 실행 수준 `tool_error_formatter`와 호출별 `rejection_message` 재정의를 결합하는 방법을 보여 줍니다.
|
||||
- **도구로 사용하는 에이전트 승인**: `Agent.as_tool(..., needs_approval=...)`는 위임된 에이전트 작업에 검토가 필요할 때 동일한 인터럽션(중단 처리) 흐름을 적용합니다. 중첩된 인터럽션(중단 처리)은 여전히 외부 실행에 노출되므로, 중첩된 에이전트가 아니라 원래 최상위 에이전트를 재개합니다.
|
||||
- **로컬 셸 및 apply_patch 도구**: `ShellTool` 및 `ApplyPatchTool`도 `needs_approval`을 지원합니다. 향후 호출에 대한 결정을 캐시하려면 `state.approve(interruption, always_approve=True)` 또는 `state.reject(..., always_reject=True)`를 사용합니다. 자동 결정을 위해서는 `on_approval`을 제공하세요(`examples/tools/shell.py` 참고). 수동 결정을 위해서는 인터럽션(중단 처리)을 처리하세요(`examples/tools/shell_human_in_the_loop.py` 참고). 호스티드 셸 환경은 `needs_approval` 또는 `on_approval`을 지원하지 않습니다. [도구 가이드](tools.md)를 참고하세요.
|
||||
- **로컬 MCP 서버**: MCP 도구 호출을 제한하려면 `MCPServerStdio` / `MCPServerSse` / `MCPServerStreamableHttp`에서 `require_approval`을 사용합니다(`examples/mcp/get_all_mcp_tools_example/main.py` 및 `examples/mcp/tool_filter_example/main.py` 참고).
|
||||
- **호스티드 MCP 서버**: HITL을 강제하려면 `HostedMCPTool`에서 `require_approval`을 `"always"`로 설정하고, 선택적으로 자동 승인 또는 거부를 위해 `on_approval_request`를 제공합니다(`examples/hosted_mcp/human_in_the_loop.py` 및 `examples/hosted_mcp/on_approval.py` 참고). 신뢰할 수 있는 서버에는 `"never"`를 사용합니다(`examples/hosted_mcp/simple.py`).
|
||||
- **세션과 메모리**: 승인과 대화 기록이 여러 턴 동안 유지되도록 `Runner.run`에 세션을 전달합니다. SQLite 및 OpenAI Conversations 세션 변형은 `examples/memory/memory_session_hitl_example.py` 및 `examples/memory/openai_session_hitl_example.py`에 있습니다.
|
||||
- **실시간 에이전트**: 실시간 데모는 `RealtimeSession`의 `approve_tool_call` / `reject_tool_call`을 통해 도구 호출을 승인하거나 거부하는 WebSocket 메시지를 노출합니다. 서버 측 핸들러는 `examples/realtime/app/server.py`를, API 인터페이스는 [Realtime 가이드](realtime/guide.md#tool-approvals)를 참고하세요.
|
||||
|
||||
## 장기 실행 승인
|
||||
|
||||
`RunState`는 내구성을 고려해 설계되었습니다. 대기 작업을 데이터베이스나 큐에 저장하려면 `state.to_json()` 또는 `state.to_string()`을 사용하고, 나중에 `RunState.from_json(...)` 또는 `RunState.from_string(...)`으로 다시 생성하세요
|
||||
`RunState`는 내구성을 갖도록 설계되었습니다. `state.to_json()` 또는 `state.to_string()`을 사용해 보류 중인 작업을 데이터베이스나 큐에 저장하고, 나중에 `RunState.from_json(...)` 또는 `RunState.from_string(...)`으로 다시 생성합니다.
|
||||
|
||||
유용한 직렬화 옵션:
|
||||
|
||||
- `context_serializer`: 매핑이 아닌 컨텍스트 객체를 직렬화하는 방식을 사용자 지정합니다
|
||||
- `context_deserializer`: `RunState.from_json(...)` 또는 `RunState.from_string(...)`으로 상태를 불러올 때 매핑이 아닌 컨텍스트 객체를 재구성합니다
|
||||
- `strict_context=True`: 컨텍스트가 이미 매핑이거나 적절한 serializer/deserializer를 제공하지 않으면 직렬화 또는 역직렬화를 실패시킵니다
|
||||
- `context_override`: 상태를 불러올 때 직렬화된 컨텍스트를 대체합니다. 원래 컨텍스트 객체를 복원하지 않으려는 경우 유용하지만, 이미 직렬화된 페이로드에서 해당 컨텍스트를 제거하지는 않습니다
|
||||
- `include_tracing_api_key=True`: 재개된 작업이 동일한 자격 증명으로 트레이스를 계속 내보내야 할 때 직렬화된 트레이스 페이로드에 트레이싱 API 키를 포함합니다
|
||||
- `context_serializer`: 비매핑 컨텍스트 객체가 직렬화되는 방식을 사용자 지정합니다.
|
||||
- `context_deserializer`: `RunState.from_json(...)` 또는 `RunState.from_string(...)`으로 상태를 로드할 때 비매핑 컨텍스트 객체를 다시 빌드합니다.
|
||||
- `strict_context=True`: 컨텍스트가 이미 매핑이거나 적절한 serializer/deserializer를 제공한 경우가 아니면 직렬화 또는 역직렬화에 실패합니다.
|
||||
- `context_override`: 상태를 로드할 때 직렬화된 컨텍스트를 대체합니다. 원래 컨텍스트 객체를 복원하고 싶지 않을 때 유용하지만, 이미 직렬화된 페이로드에서 해당 컨텍스트를 제거하지는 않습니다.
|
||||
- `include_tracing_api_key=True`: 재개된 작업이 동일한 자격 증명으로 트레이스를 계속 내보내야 하는 경우, 직렬화된 트레이스 페이로드에 트레이싱 API 키를 포함합니다.
|
||||
|
||||
직렬화된 실행 상태에는 앱 컨텍스트와 함께 승인, 사용량, 직렬화된 `tool_input`, 중첩 에이전트-as-tool 재개, 트레이스 메타데이터, 서버 관리 대화 설정 같은 SDK 관리 런타임 메타데이터가 포함됩니다. 직렬화된 상태를 저장하거나 전송할 계획이라면 `RunContextWrapper.context`를 영속 데이터로 취급하고, 상태와 함께 이동시키려는 의도가 없는 한 비밀 정보를 그 안에 두지 마세요
|
||||
직렬화된 실행 상태에는 앱 컨텍스트와 함께 승인, 사용량, 직렬화된 `tool_input`, 중첩된 agent-as-tool 재개, 트레이스 메타데이터, 서버 관리 대화 설정 같은 SDK 관리 런타임 메타데이터가 포함됩니다. 직렬화된 상태를 저장하거나 전송하려는 경우 `RunContextWrapper.context`를 영속화된 데이터로 취급하고, 상태와 함께 이동하기를 의도한 경우가 아니라면 그 안에 비밀 정보를 두지 마세요.
|
||||
|
||||
## 대기 작업 버전 관리
|
||||
## 보류 중인 작업 버전 관리
|
||||
|
||||
승인이 한동안 대기 상태로 있을 수 있다면, 직렬화된 상태와 함께 에이전트 정의 또는 SDK의 버전 마커를 저장하세요. 그러면 모델, 프롬프트 또는 도구 정의가 바뀔 때 발생할 수 있는 비호환성을 피하기 위해 역직렬화를 일치하는 코드 경로로 라우팅할 수 있습니다
|
||||
승인이 한동안 대기할 수 있다면, 에이전트 정의 또는 SDK의 버전 표시자를 직렬화된 상태와 함께 저장하세요. 그러면 모델, 프롬프트 또는 도구 정의가 변경될 때 비호환성을 피하기 위해 역직렬화를 일치하는 코드 경로로 라우팅할 수 있습니다.
|
||||
+43
-43
@@ -4,51 +4,51 @@ search:
|
||||
---
|
||||
# OpenAI Agents SDK
|
||||
|
||||
[OpenAI Agents SDK](https://github.com/openai/openai-agents-python)는 매우 적은 추상화만으로 에이전트형 AI 앱을 가볍고 사용하기 쉬운 패키지로 구축할 수 있게 해줍니다. 이는 이전의 에이전트 실험용 프레임워크인 [Swarm](https://github.com/openai/swarm/tree/main)을 프로덕션 준비 수준으로 확장한 것입니다. Agents SDK는 매우 작은 기본 구성 요소 집합을 제공합니다.
|
||||
[OpenAI Agents SDK](https://github.com/openai/openai-agents-python)를 사용하면 매우 적은 추상화만으로 가볍고 사용하기 쉬운 패키지에서 에이전트형 AI 앱을 구축할 수 있습니다. 이는 이전 에이전트 실험 프로젝트인 [Swarm](https://github.com/openai/swarm/tree/main)을 프로덕션에 바로 사용할 수 있도록 업그레이드한 것입니다. Agents SDK는 매우 작은 기본 구성 요소 집합을 갖습니다.
|
||||
|
||||
- **에이전트**: instructions와 tools를 갖춘 LLM
|
||||
- **Agents as tools / 핸드오프**: 에이전트가 특정 작업을 위해 다른 에이전트에 위임할 수 있게 해주는 기능
|
||||
- **가드레일**: 에이전트 입력과 출력을 검증할 수 있게 해주는 기능
|
||||
- **Agents as tools / 핸드오프**: 에이전트가 특정 작업을 다른 에이전트에 위임할 수 있게 하는 기능
|
||||
- **가드레일**: 에이전트 입력과 출력을 검증할 수 있게 하는 기능
|
||||
|
||||
이러한 기본 구성 요소는 Python과 결합될 때 도구와 에이전트 간의 복잡한 관계를 표현할 수 있을 만큼 강력하며, 가파른 학습 곡선 없이도 실제 애플리케이션을 구축할 수 있게 해줍니다. 또한 SDK에는 에이전트형 흐름을 시각화하고 디버그할 수 있을 뿐만 아니라 이를 평가하고 애플리케이션에 맞게 모델을 파인튜닝할 수 있도록 해주는 내장 **트레이싱**도 포함되어 있습니다.
|
||||
Python과 결합하면 이러한 기본 구성 요소만으로도 도구와 에이전트 간의 복잡한 관계를 표현하기에 충분히 강력하며, 가파른 학습 곡선 없이 실제 애플리케이션을 구축할 수 있습니다. 또한 SDK에는 에이전트형 흐름을 시각화하고 디버깅하며, 이를 평가하고 애플리케이션에 맞게 모델을 파인튜닝할 수 있는 내장 **트레이싱** 기능이 포함되어 있습니다.
|
||||
|
||||
## Agents SDK 사용 이유
|
||||
## Agents SDK를 사용하는 이유
|
||||
|
||||
SDK에는 두 가지 핵심 설계 원칙이 있습니다.
|
||||
|
||||
1. 사용할 가치가 있을 만큼 충분한 기능을 제공하면서도, 빠르게 익힐 수 있을 만큼 기본 구성 요소 수는 적게 유지합니다
|
||||
2. 기본 상태로도 훌륭하게 동작하지만, 정확히 어떤 일이 일어날지 세밀하게 사용자 지정할 수 있습니다
|
||||
1. 사용할 가치가 있을 만큼 충분한 기능을 제공하되, 빠르게 배울 수 있을 만큼 기본 구성 요소는 적게 유지합니다.
|
||||
2. 기본 설정만으로도 잘 작동하지만, 어떤 일이 일어나는지는 정확하게 사용자 지정할 수 있습니다.
|
||||
|
||||
다음은 SDK의 주요 기능입니다.
|
||||
SDK의 주요 기능은 다음과 같습니다.
|
||||
|
||||
- **에이전트 루프**: 도구 호출을 처리하고, 결과를 LLM에 다시 전달하며, 작업이 완료될 때까지 계속하는 내장 에이전트 루프
|
||||
- **파이썬 우선**: 새로운 추상화를 배울 필요 없이, 내장 언어 기능을 사용해 에이전트를 오케스트레이션하고 연결합니다
|
||||
- **Agents as tools / 핸드오프**: 여러 에이전트에 걸쳐 작업을 조율하고 위임하기 위한 강력한 메커니즘
|
||||
- **샌드박스 에이전트**: 매니페스트로 정의된 파일, 샌드박스 클라이언트 선택, 재개 가능한 샌드박스 세션을 갖춘 실제 격리 작업공간 안에서 전문 에이전트를 실행합니다
|
||||
- **가드레일**: 에이전트 실행과 병렬로 입력 검증 및 안전성 검사를 수행하고, 검사를 통과하지 못하면 즉시 실패 처리합니다
|
||||
- **함수 도구**: 자동 스키마 생성과 Pydantic 기반 검증을 통해 모든 Python 함수를 도구로 변환합니다
|
||||
- **에이전트 루프**: 도구 호출을 처리하고, 결과를 LLM에 다시 보내며, 작업이 완료될 때까지 계속 실행하는 내장 에이전트 루프
|
||||
- **파이썬 우선**: 새로운 추상화를 배울 필요 없이, 내장 언어 기능을 사용해 에이전트를 오케스트레이션하고 체인으로 연결
|
||||
- **Agents as tools / 핸드오프**: 여러 에이전트 간 작업을 조율하고 위임하기 위한 강력한 메커니즘
|
||||
- **샌드박스 에이전트**: 매니페스트로 정의된 파일, 샌드박스 클라이언트 선택, 재개 가능한 샌드박스 세션을 통해 실제 격리된 워크스페이스 안에서 전문가 실행
|
||||
- **가드레일**: 에이전트 실행과 병렬로 입력 검증 및 안전성 검사를 실행하고, 검사를 통과하지 못하면 빠르게 실패 처리
|
||||
- **함수 도구**: 자동 스키마 생성 및 Pydantic 기반 검증을 통해 모든 Python 함수를 도구로 변환
|
||||
- **MCP 서버 도구 호출**: 함수 도구와 동일한 방식으로 작동하는 내장 MCP 서버 도구 통합
|
||||
- **세션**: 에이전트 루프 내에서 작업 컨텍스트를 유지하기 위한 지속형 메모리 계층
|
||||
- **휴먼인더루프 (HITL)**: 에이전트 실행 전반에 걸쳐 사람이 개입할 수 있도록 하는 내장 메커니즘
|
||||
- **트레이싱**: 워크플로를 시각화, 디버그, 모니터링하기 위한 내장 트레이싱으로, OpenAI의 평가, 파인튜닝, 증류 도구 모음을 지원합니다
|
||||
- **실시간 에이전트**: `gpt-realtime-1.5`와 자동 인터럽션(중단 처리) 감지, 컨텍스트 관리, 가드레일 등을 사용해 강력한 음성 에이전트를 구축합니다
|
||||
- **세션**: 에이전트 루프 내에서 작업 컨텍스트를 유지하기 위한 영속 메모리 계층
|
||||
- **휴먼인더루프 (HITL)**: 에이전트 실행 전반에 사람을 참여시키기 위한 내장 메커니즘
|
||||
- **트레이싱**: OpenAI의 평가, 파인튜닝, 증류 도구 모음 지원과 함께 워크플로를 시각화, 디버깅, 모니터링하기 위한 내장 트레이싱
|
||||
- **실시간 에이전트**: `gpt-realtime-2`, 자동 인터럽션(중단 처리) 감지, 컨텍스트 관리, 가드레일 등을 활용해 강력한 음성 에이전트 구축
|
||||
|
||||
## Agents SDK 또는 Responses API
|
||||
|
||||
SDK는 OpenAI 모델에 대해 기본적으로 Responses API를 사용하지만, 모델 호출 위에 더 높은 수준의 런타임을 추가로 제공합니다.
|
||||
SDK는 OpenAI 모델에 기본적으로 Responses API를 사용하지만, 모델 호출 주변에 더 높은 수준의 런타임을 추가합니다.
|
||||
|
||||
다음과 같은 경우에는 Responses API를 직접 사용하세요.
|
||||
다음과 같은 경우 Responses API를 직접 사용하세요.
|
||||
|
||||
- 루프, 도구 디스패치, 상태 처리를 직접 관리하고 싶은 경우
|
||||
- 워크플로가 짧게 유지되며 주로 모델의 응답을 반환하는 것이 목적일 경우
|
||||
- 루프, 도구 디스패치, 상태 처리를 직접 관리하려는 경우
|
||||
- 워크플로가 짧게 실행되며 주로 모델의 응답을 반환하는 것이 목적인 경우
|
||||
|
||||
다음과 같은 경우에는 Agents SDK를 사용하세요.
|
||||
다음과 같은 경우 Agents SDK를 사용하세요.
|
||||
|
||||
- 런타임이 턴, 도구 실행, 가드레일, 핸드오프 또는 세션을 관리하길 원하는 경우
|
||||
- 에이전트가 아티팩트를 생성하거나 여러 조정된 단계에 걸쳐 작업해야 하는 경우
|
||||
- [샌드박스 에이전트](sandbox_agents.md)를 통해 실제 작업공간이나 재개 가능한 실행이 필요한 경우
|
||||
- 런타임이 턴, 도구 실행, 가드레일, 핸드오프 또는 세션을 관리하기를 원하는 경우
|
||||
- 에이전트가 아티팩트를 생성하거나 여러 조율된 단계에 걸쳐 동작해야 하는 경우
|
||||
- 실제 워크스페이스나 [샌드박스 에이전트](sandbox_agents.md)를 통한 재개 가능한 실행이 필요한 경우
|
||||
|
||||
둘 중 하나를 전역적으로 선택할 필요는 없습니다. 많은 애플리케이션이 관리형 워크플로에는 SDK를 사용하고, 더 낮은 수준의 경로에는 Responses API를 직접 호출합니다.
|
||||
둘 중 하나를 전역적으로 선택할 필요는 없습니다. 많은 애플리케이션은 관리형 워크플로에는 SDK를 사용하고, 더 낮은 수준의 경로에는 Responses API를 직접 호출합니다.
|
||||
|
||||
## 설치
|
||||
|
||||
@@ -56,7 +56,7 @@ SDK는 OpenAI 모델에 대해 기본적으로 Responses API를 사용하지만,
|
||||
pip install openai-agents
|
||||
```
|
||||
|
||||
## Hello World 예제
|
||||
## Hello world 예제
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
@@ -71,7 +71,7 @@ print(result.final_output)
|
||||
# Infinite loop's dance.
|
||||
```
|
||||
|
||||
(_이를 실행하려면 `OPENAI_API_KEY` 환경 변수를 설정했는지 확인하세요_)
|
||||
(_이를 실행하는 경우 `OPENAI_API_KEY` 환경 변수를 설정했는지 확인하세요_)
|
||||
|
||||
```bash
|
||||
export OPENAI_API_KEY=sk-...
|
||||
@@ -79,23 +79,23 @@ export OPENAI_API_KEY=sk-...
|
||||
|
||||
## 시작 지점
|
||||
|
||||
- [Quickstart](quickstart.md)로 첫 번째 텍스트 기반 에이전트를 구축하세요
|
||||
- 그런 다음 [에이전트 실행](running_agents.md#choose-a-memory-strategy)에서 턴 간 상태를 어떻게 유지할지 결정하세요
|
||||
- 작업이 실제 파일, 저장소 또는 에이전트별로 격리된 작업공간 상태에 의존한다면 [샌드박스 에이전트 빠른 시작](sandbox_agents.md)을 읽어보세요
|
||||
- 핸드오프와 관리자 스타일 오케스트레이션 중 무엇을 선택할지 결정하고 있다면 [에이전트 오케스트레이션](multi_agent.md)을 읽어보세요
|
||||
- [빠른 시작](quickstart.md)으로 첫 텍스트 기반 에이전트를 구축하세요.
|
||||
- 그런 다음 [에이전트 실행](running_agents.md#choose-a-memory-strategy)에서 턴 간 상태를 어떻게 유지할지 결정하세요.
|
||||
- 작업이 실제 파일, 리포지토리 또는 에이전트별 격리된 워크스페이스 상태에 의존한다면 [샌드박스 에이전트 빠른 시작](sandbox_agents.md)을 읽어 보세요.
|
||||
- 핸드오프와 매니저 스타일 오케스트레이션 중에서 결정하는 중이라면 [에이전트 오케스트레이션](multi_agent.md)을 읽어 보세요.
|
||||
|
||||
## 경로 선택
|
||||
|
||||
원하는 작업은 알고 있지만 어떤 페이지가 이를 설명하는지 모를 때 이 표를 사용하세요.
|
||||
하려는 작업은 알고 있지만 어느 페이지에서 설명하는지 모를 때 이 표를 사용하세요.
|
||||
|
||||
| 목표 | 시작 지점 |
|
||||
| --- | --- |
|
||||
| 첫 번째 텍스트 에이전트를 만들고 하나의 전체 실행을 확인하기 | [Quickstart](quickstart.md) |
|
||||
| 함수 도구, 호스티드 툴 또는 Agents as tools 추가하기 | [도구](tools.md) |
|
||||
| 실제 격리 작업공간 안에서 코딩, 리뷰 또는 문서 에이전트 실행하기 | [샌드박스 에이전트 빠른 시작](sandbox_agents.md) 및 [샌드박스 클라이언트](sandbox/clients.md) |
|
||||
| 핸드오프와 관리자 스타일 오케스트레이션 중 선택하기 | [에이전트 오케스트레이션](multi_agent.md) |
|
||||
| 턴 간 메모리 유지하기 | [에이전트 실행](running_agents.md#choose-a-memory-strategy) 및 [세션](sessions/index.md) |
|
||||
| OpenAI 모델, websocket 전송 또는 OpenAI가 아닌 제공자 사용하기 | [모델](models/index.md) |
|
||||
| 출력, 실행 항목, 인터럽션(중단 처리), 재개 상태 검토하기 | [결과](results.md) |
|
||||
| `gpt-realtime-1.5`로 저지연 음성 에이전트 구축하기 | [실시간 에이전트 빠른 시작](realtime/quickstart.md) 및 [실시간 전송](realtime/transport.md) |
|
||||
| speech-to-text / 에이전트 / text-to-speech 파이프라인 구축하기 | [음성 파이프라인 빠른 시작](voice/quickstart.md) |
|
||||
| 첫 텍스트 에이전트를 만들고 전체 실행 한 번 확인 | [빠른 시작](quickstart.md) |
|
||||
| 함수 도구, 호스티드 툴 또는 agents as tools 추가 | [도구](tools.md) |
|
||||
| 실제 격리된 워크스페이스 안에서 코딩, 리뷰 또는 문서 에이전트 실행 | [샌드박스 에이전트 빠른 시작](sandbox_agents.md) 및 [샌드박스 클라이언트](sandbox/clients.md) |
|
||||
| 핸드오프와 매니저 스타일 오케스트레이션 중에서 결정 | [에이전트 오케스트레이션](multi_agent.md) |
|
||||
| 턴 간 메모리 유지 | [에이전트 실행](running_agents.md#choose-a-memory-strategy) 및 [세션](sessions/index.md) |
|
||||
| OpenAI 모델, 웹소켓 전송 또는 비 OpenAI 제공자 사용 | [모델](models/index.md) |
|
||||
| 출력, 실행 항목, 인터럽션(중단 처리), 재개 상태 검토 | [결과](results.md) |
|
||||
| `gpt-realtime-2`로 지연 시간이 낮은 음성 에이전트 구축 | [실시간 에이전트 빠른 시작](realtime/quickstart.md) 및 [실시간 전송](realtime/transport.md) |
|
||||
| 음성-텍스트 변환 / 에이전트 / 텍스트-음성 변환 파이프라인 구축 | [음성 파이프라인 빠른 시작](voice/quickstart.md) |
|
||||
+97
-107
@@ -4,32 +4,31 @@ search:
|
||||
---
|
||||
# Model context protocol (MCP)
|
||||
|
||||
[Model context protocol](https://modelcontextprotocol.io/introduction)(MCP)은 애플리케이션이 언어 모델에 도구와 컨텍스트를 노출하는 방식을 표준화합니다. 공식 문서에서 다음과 같이 설명합니다:
|
||||
[Model context protocol](https://modelcontextprotocol.io/introduction) (MCP)는 애플리케이션이 도구와
|
||||
컨텍스트를 언어 모델에 노출하는 방식을 표준화합니다. 공식 문서에 따르면 다음과 같습니다.
|
||||
|
||||
> MCP는 애플리케이션이 LLM에 컨텍스트를 제공하는 방식을 표준화하는 개방형 프로토콜입니다. MCP를 AI 애플리케이션용 USB-C 포트라고 생각해 보세요
|
||||
> USB-C가 다양한 주변기기 및 액세서리에 기기를 연결하는 표준화된 방법을 제공하듯, MCP는
|
||||
> AI 모델을 서로 다른 데이터 소스 및 도구에 연결하는 표준화된 방법을 제공합니다
|
||||
> MCP는 애플리케이션이 LLMs에 컨텍스트를 제공하는 방식을 표준화하는 개방형 프로토콜입니다. MCP를 AI
|
||||
> 애플리케이션을 위한 USB-C 포트처럼 생각해 보세요. USB-C가 기기를 다양한 주변 장치와 액세서리에 연결하는 표준화된 방식을 제공하듯이, MCP는
|
||||
> AI 모델을 다양한 데이터 소스와 도구에 연결하는 표준화된 방식을 제공합니다.
|
||||
|
||||
Agents Python SDK는 여러 MCP 전송 방식을 이해합니다. 이를 통해 기존 MCP 서버를 재사용하거나 직접 구축하여
|
||||
파일시스템, HTTP 또는 커넥터 기반 도구를 에이전트에 노출할 수 있습니다.
|
||||
Agents Python SDK는 여러 MCP 전송 방식을 지원합니다. 이를 통해 기존 MCP 서버를 재사용하거나, 파일 시스템, HTTP 또는 커넥터 기반 도구를 에이전트에 노출하도록 직접 빌드할 수 있습니다.
|
||||
|
||||
## MCP 통합 선택
|
||||
|
||||
에이전트에 MCP 서버를 연결하기 전에 도구 호출이 어디에서 실행되어야 하는지, 어떤 전송 방식에 도달할 수 있는지 결정하세요. 아래
|
||||
매트릭스는 Python SDK가 지원하는 옵션을 요약합니다.
|
||||
MCP 서버를 에이전트에 연결하기 전에 도구 호출을 어디서 실행해야 하는지, 어떤 전송 방식에 접근할 수 있는지 결정하세요. 아래 표는 Python SDK가 지원하는 옵션을 요약합니다.
|
||||
|
||||
| 필요한 항목 | 권장 옵션 |
|
||||
| 필요한 사항 | 권장 옵션 |
|
||||
| ------------------------------------------------------------------------------------ | ----------------------------------------------------- |
|
||||
| OpenAI의 Responses API가 모델을 대신해 공개적으로 접근 가능한 MCP 서버를 호출하도록 하기| [`HostedMCPTool`][agents.tool.HostedMCPTool]을 통한 **호스티드 MCP 서버 도구** |
|
||||
| 로컬 또는 원격에서 실행하는 Streamable HTTP 서버에 연결 | [`MCPServerStreamableHttp`][agents.mcp.server.MCPServerStreamableHttp]를 통한 **Streamable HTTP MCP 서버** |
|
||||
| Server-Sent Events를 사용하는 HTTP를 구현한 서버와 통신 | [`MCPServerSse`][agents.mcp.server.MCPServerSse]를 통한 **SSE 기반 HTTP MCP 서버** |
|
||||
| 로컬 프로세스를 실행하고 stdin/stdout으로 통신 | [`MCPServerStdio`][agents.mcp.server.MCPServerStdio]를 통한 **stdio MCP 서버** |
|
||||
| 로컬 또는 원격에서 실행하는 Streamable HTTP 서버에 연결하기 | [`MCPServerStreamableHttp`][agents.mcp.server.MCPServerStreamableHttp]를 통한 **Streamable HTTP MCP 서버** |
|
||||
| Server-Sent Events가 포함된 HTTP를 구현하는 서버와 통신하기 | [`MCPServerSse`][agents.mcp.server.MCPServerSse]를 통한 **HTTP with SSE MCP 서버** |
|
||||
| 로컬 프로세스를 시작하고 stdin/stdout을 통해 통신하기 | [`MCPServerStdio`][agents.mcp.server.MCPServerStdio]를 통한 **stdio MCP 서버** |
|
||||
|
||||
아래 섹션에서는 각 옵션, 구성 방법, 그리고 어떤 전송 방식을 선호해야 하는지를 안내합니다.
|
||||
아래 섹션에서는 각 옵션, 구성 방법, 그리고 어떤 경우에 한 전송 방식을 다른 전송 방식보다 선호해야 하는지 살펴봅니다.
|
||||
|
||||
## 에이전트 수준 MCP 구성
|
||||
|
||||
전송 방식 선택 외에도 `Agent.mcp_config`를 설정하여 MCP 도구 준비 방식을 조정할 수 있습니다.
|
||||
전송 방식을 선택하는 것 외에도 `Agent.mcp_config`를 설정하여 MCP 도구가 준비되는 방식을 조정할 수 있습니다.
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
@@ -43,38 +42,39 @@ agent = Agent(
|
||||
# If None, MCP tool failures are raised as exceptions instead of
|
||||
# returning model-visible error text.
|
||||
"failure_error_function": None,
|
||||
# Prefix local MCP tool names with their server name.
|
||||
"include_server_in_tool_names": True,
|
||||
},
|
||||
)
|
||||
```
|
||||
|
||||
참고:
|
||||
|
||||
- `convert_schemas_to_strict`는 최선의 노력 방식입니다. 스키마를 변환할 수 없으면 원래 스키마를 사용합니다
|
||||
- `failure_error_function`은 MCP 도구 호출 실패가 모델에 어떻게 표시될지 제어합니다
|
||||
- `failure_error_function`이 설정되지 않으면 SDK는 기본 도구 오류 포매터를 사용합니다
|
||||
- 서버 수준 `failure_error_function`은 해당 서버에서 `Agent.mcp_config["failure_error_function"]`보다 우선합니다
|
||||
- `convert_schemas_to_strict`는 최선 노력 방식입니다. 스키마를 변환할 수 없으면 원래 스키마가 사용됩니다.
|
||||
- `failure_error_function`은 MCP 도구 호출 실패가 모델에 어떻게 표시되는지 제어합니다.
|
||||
- `failure_error_function`이 설정되지 않은 경우 SDK는 기본 도구 오류 포매터를 사용합니다.
|
||||
- 서버 수준의 `failure_error_function`은 해당 서버에 대해 `Agent.mcp_config["failure_error_function"]`을 재정의합니다.
|
||||
- `include_server_in_tool_names`는 옵트인 방식입니다. 활성화하면 각 로컬 MCP 도구가 결정적인 서버 접두사 이름으로 모델에 노출되어, 여러 MCP 서버가 같은 이름의 도구를 게시할 때 충돌을 피하는 데 도움이 됩니다. 생성된 이름은 ASCII-safe이고, 함수 도구 이름 길이 제한 내에 있으며, 동일한 에이전트에 있는 기존 로컬 함수 도구 이름 및 활성화된 핸드오프 이름과 충돌하지 않습니다. SDK는 여전히 원래 서버에서 원래 MCP 도구 이름을 호출합니다.
|
||||
|
||||
## 전송 방식 전반의 공통 패턴
|
||||
|
||||
전송 방식을 선택한 뒤에는 대부분의 통합에서 동일한 후속 결정을 해야 합니다:
|
||||
전송 방식을 선택한 뒤에는 대부분의 통합에서 다음과 같은 후속 결정이 필요합니다.
|
||||
|
||||
- 도구의 일부만 노출하는 방법([도구 필터링](#tool-filtering))
|
||||
- 서버가 재사용 가능한 프롬프트도 제공하는지 여부([프롬프트](#prompts))
|
||||
- `list_tools()`를 캐시해야 하는지 여부([캐싱](#caching))
|
||||
- MCP 활동이 트레이스에 어떻게 표시되는지([트레이싱](#tracing))
|
||||
- 도구의 일부만 노출하는 방법([도구 필터링](#tool-filtering)).
|
||||
- 서버가 재사용 가능한 프롬프트도 제공하는지 여부([프롬프트](#prompts)).
|
||||
- `list_tools()`를 캐시해야 하는지 여부([캐싱](#caching)).
|
||||
- MCP 활동이 트레이스에 표시되는 방식([트레이싱](#tracing)).
|
||||
|
||||
로컬 MCP 서버(`MCPServerStdio`, `MCPServerSse`, `MCPServerStreamableHttp`)의 경우 승인 정책과 호출별 `_meta` 페이로드도 공통 개념입니다. Streamable HTTP 섹션에 가장 완전한 예제가 있으며, 동일한 패턴이 다른 로컬 전송 방식에도 적용됩니다.
|
||||
로컬 MCP 서버(`MCPServerStdio`, `MCPServerSse`, `MCPServerStreamableHttp`)의 경우 승인 정책과 호출별 `_meta` 페이로드도 공통 개념입니다. Streamable HTTP 섹션은 가장 완전한 예를 보여주며, 동일한 패턴이 다른 로컬 전송 방식에도 적용됩니다.
|
||||
|
||||
## 1. 호스티드 MCP 서버 도구
|
||||
|
||||
호스티드 도구는 도구 라운드트립 전체를 OpenAI 인프라로 이동시킵니다. 코드가 도구를 나열하고 호출하는 대신
|
||||
[`HostedMCPTool`][agents.tool.HostedMCPTool]이 서버 레이블(및 선택적 커넥터 메타데이터)을 Responses API로 전달합니다. 모델은
|
||||
원격 서버의 도구를 나열하고 Python 프로세스에 추가 콜백 없이 이를 호출합니다. 현재 호스티드 도구는 Responses API의 호스티드 MCP 통합을 지원하는 OpenAI 모델에서 동작합니다.
|
||||
호스티드 툴은 전체 도구 왕복 과정을 OpenAI 인프라로 보냅니다. 코드가 도구를 나열하고 호출하는 대신, [`HostedMCPTool`][agents.tool.HostedMCPTool]은 서버 레이블(및 선택적 커넥터 메타데이터)을 Responses API로 전달합니다. 모델은 Python 프로세스에 대한 추가 콜백 없이 원격 서버의 도구를 나열하고 호출합니다. 현재 호스티드 툴은 Responses API의 호스티드 MCP 통합을 지원하는 OpenAI 모델에서 작동합니다.
|
||||
|
||||
### 기본 호스티드 MCP 도구
|
||||
|
||||
에이전트의 `tools` 목록에 [`HostedMCPTool`][agents.tool.HostedMCPTool]을 추가하여 호스티드 도구를 생성합니다. `tool_config`
|
||||
딕셔너리는 REST API로 보내는 JSON을 반영합니다:
|
||||
에이전트의 `tools` 목록에 [`HostedMCPTool`][agents.tool.HostedMCPTool]을 추가하여 호스티드 툴을 만듭니다. `tool_config`
|
||||
dict는 REST API로 보낼 JSON과 동일한 구조입니다.
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -84,32 +84,36 @@ from agents import Agent, HostedMCPTool, Runner
|
||||
async def main() -> None:
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
instructions="Use the DeepWiki hosted MCP server to inspect openai/openai-agents-python.",
|
||||
tools=[
|
||||
HostedMCPTool(
|
||||
tool_config={
|
||||
"type": "mcp",
|
||||
"server_label": "gitmcp",
|
||||
"server_url": "https://gitmcp.io/openai/codex",
|
||||
"server_label": "deepwiki",
|
||||
"server_url": "https://mcp.deepwiki.com/mcp",
|
||||
"require_approval": "never",
|
||||
}
|
||||
)
|
||||
],
|
||||
)
|
||||
|
||||
result = await Runner.run(agent, "Which language is this repository written in?")
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Which language is the repository openai/openai-agents-python written in?",
|
||||
)
|
||||
print(result.final_output)
|
||||
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
호스티드 서버는 도구를 자동으로 노출하므로 `mcp_servers`에 추가할 필요가 없습니다.
|
||||
호스티드 서버는 자체 도구를 자동으로 노출하므로, 이를 `mcp_servers`에 추가하지 않습니다.
|
||||
|
||||
호스티드 도구 검색에서 호스티드 MCP 서버를 지연 로드하려면 `tool_config["defer_loading"] = True`로 설정하고 에이전트에 [`ToolSearchTool`][agents.tool.ToolSearchTool]을 추가하세요. 이는 OpenAI Responses 모델에서만 지원됩니다. 전체 도구 검색 설정과 제약 사항은 [도구](tools.md#hosted-tool-search)를 참고하세요.
|
||||
호스티드 툴 검색이 호스티드 MCP 서버를 지연 로드하도록 하려면 `tool_config["defer_loading"] = True`를 설정하고 [`ToolSearchTool`][agents.tool.ToolSearchTool]을 에이전트에 추가하세요. 이는 OpenAI Responses 모델에서만 지원됩니다. 전체 도구 검색 설정 및 제약 사항은 [도구](tools.md#hosted-tool-search)를 참조하세요.
|
||||
|
||||
### 호스티드 MCP 결과 스트리밍
|
||||
|
||||
호스티드 도구는 함수 도구와 정확히 동일한 방식으로 결과 스트리밍을 지원합니다. `Runner.run_streamed`를 사용해
|
||||
모델이 아직 작업 중일 때 점진적인 MCP 출력을 소비하세요:
|
||||
호스티드 툴은 함수 도구와 정확히 같은 방식으로 스트리밍 결과를 지원합니다. 모델이 계속 작업하는 동안
|
||||
증분 MCP 출력을 소비하려면 `Runner.run_streamed`를 사용하세요.
|
||||
|
||||
```python
|
||||
result = Runner.run_streamed(agent, "Summarise this repository's top languages")
|
||||
@@ -121,13 +125,12 @@ print(result.final_output)
|
||||
|
||||
### 선택적 승인 흐름
|
||||
|
||||
서버가 민감한 작업을 수행할 수 있다면 각 도구 실행 전에 사람 또는 프로그래매틱 승인을 요구할 수 있습니다. `tool_config`에서
|
||||
`require_approval`을 단일 정책(`"always"`, `"never"`) 또는 도구 이름별 정책 딕셔너리로 구성하세요. Python 내부에서 결정을 내리려면 `on_approval_request` 콜백을 제공하세요.
|
||||
서버가 민감한 작업을 수행할 수 있다면 각 도구 실행 전에 사람 또는 프로그램 방식의 승인을 요구할 수 있습니다. `tool_config`에서 `require_approval`을 단일 정책(`"always"`, `"never"`) 또는 도구 이름을 정책에 매핑하는 dict로 구성하세요. Python 내부에서 결정을 내리려면 `on_approval_request` 콜백을 제공하세요.
|
||||
|
||||
```python
|
||||
from agents import MCPToolApprovalFunctionResult, MCPToolApprovalRequest
|
||||
|
||||
SAFE_TOOLS = {"read_project_metadata"}
|
||||
SAFE_TOOLS = {"read_wiki_structure", "read_wiki_contents", "ask_question"}
|
||||
|
||||
def approve_tool(request: MCPToolApprovalRequest) -> MCPToolApprovalFunctionResult:
|
||||
if request.data.name in SAFE_TOOLS:
|
||||
@@ -140,8 +143,8 @@ agent = Agent(
|
||||
HostedMCPTool(
|
||||
tool_config={
|
||||
"type": "mcp",
|
||||
"server_label": "gitmcp",
|
||||
"server_url": "https://gitmcp.io/openai/codex",
|
||||
"server_label": "deepwiki",
|
||||
"server_url": "https://mcp.deepwiki.com/mcp",
|
||||
"require_approval": "always",
|
||||
},
|
||||
on_approval_request=approve_tool,
|
||||
@@ -150,12 +153,11 @@ agent = Agent(
|
||||
)
|
||||
```
|
||||
|
||||
콜백은 동기 또는 비동기일 수 있으며, 모델이 실행을 계속하기 위해 승인 데이터가 필요할 때마다 호출됩니다.
|
||||
콜백은 동기 또는 비동기일 수 있으며, 모델이 계속 실행하기 위해 승인 데이터가 필요할 때마다 호출됩니다.
|
||||
|
||||
### 커넥터 기반 호스티드 서버
|
||||
|
||||
호스티드 MCP는 OpenAI 커넥터도 지원합니다. `server_url`을 지정하는 대신 `connector_id`와 액세스 토큰을 제공하세요. Responses
|
||||
API가 인증을 처리하고 호스티드 서버가 커넥터의 도구를 노출합니다.
|
||||
호스티드 MCP는 OpenAI 커넥터도 지원합니다. `server_url`을 지정하는 대신 `connector_id`와 액세스 토큰을 제공하세요. Responses API가 인증을 처리하고 호스티드 서버가 커넥터의 도구를 노출합니다.
|
||||
|
||||
```python
|
||||
import os
|
||||
@@ -171,14 +173,11 @@ HostedMCPTool(
|
||||
)
|
||||
```
|
||||
|
||||
스트리밍, 승인, 커넥터를 포함한 완전한 동작 예제는
|
||||
[`examples/hosted_mcp`](https://github.com/openai/openai-agents-python/tree/main/examples/hosted_mcp)에 있습니다.
|
||||
스트리밍, 승인, 커넥터를 포함해 완전히 동작하는 호스티드 툴 샘플은 [`examples/hosted_mcp`](https://github.com/openai/openai-agents-python/tree/main/examples/hosted_mcp)에 있습니다.
|
||||
|
||||
## 2. Streamable HTTP MCP 서버
|
||||
|
||||
네트워크 연결을 직접 관리하려면
|
||||
[`MCPServerStreamableHttp`][agents.mcp.server.MCPServerStreamableHttp]를 사용하세요. Streamable HTTP 서버는 전송 계층을 제어하거나
|
||||
지연 시간을 낮게 유지하면서 자체 인프라에서 서버를 실행하려는 경우에 이상적입니다.
|
||||
네트워크 연결을 직접 관리하려면 [`MCPServerStreamableHttp`][agents.mcp.server.MCPServerStreamableHttp]를 사용하세요. Streamable HTTP 서버는 전송 방식을 제어하거나, 지연 시간을 낮게 유지하면서 자체 인프라 내부에서 서버를 실행하려는 경우에 적합합니다.
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -213,27 +212,26 @@ async def main() -> None:
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
생성자는 다음과 같은 추가 옵션을 받습니다:
|
||||
생성자는 추가 옵션을 받습니다.
|
||||
|
||||
- `client_session_timeout_seconds`는 HTTP 읽기 타임아웃을 제어합니다
|
||||
- `use_structured_content`는 텍스트 출력보다 `tool_result.structured_content`를 우선할지 전환합니다
|
||||
- `max_retry_attempts`와 `retry_backoff_seconds_base`는 `list_tools()`와 `call_tool()`에 자동 재시도를 추가합니다
|
||||
- `tool_filter`는 도구 일부만 노출할 수 있게 합니다([도구 필터링](#tool-filtering) 참조)
|
||||
- `require_approval`은 로컬 MCP 도구에 휴먼인더루프 (HITL) 승인 정책을 활성화합니다
|
||||
- `failure_error_function`은 모델에 표시되는 MCP 도구 실패 메시지를 사용자 지정합니다. 대신 오류를 발생시키려면 `None`으로 설정하세요
|
||||
- `tool_meta_resolver`는 `call_tool()` 전에 호출별 MCP `_meta` 페이로드를 주입합니다
|
||||
- `client_session_timeout_seconds`는 HTTP 읽기 타임아웃을 제어합니다.
|
||||
- `use_structured_content`는 텍스트 출력보다 `tool_result.structured_content`를 선호할지 여부를 전환합니다.
|
||||
- `max_retry_attempts`와 `retry_backoff_seconds_base`는 `list_tools()` 및 `call_tool()`에 대한 자동 재시도를 추가합니다.
|
||||
- `tool_filter`를 사용하면 도구의 일부만 노출할 수 있습니다([도구 필터링](#tool-filtering) 참조).
|
||||
- `require_approval`은 로컬 MCP 도구에서 휴먼인더루프 (HITL) 승인 정책을 활성화합니다.
|
||||
- `failure_error_function`은 모델에 표시되는 MCP 도구 실패 메시지를 사용자 지정합니다. 오류를 대신 발생시키려면 이를 `None`으로 설정하세요.
|
||||
- `tool_meta_resolver`는 `call_tool()` 전에 호출별 MCP `_meta` 페이로드를 주입합니다.
|
||||
|
||||
### 로컬 MCP 서버용 승인 정책
|
||||
### 로컬 MCP 서버의 승인 정책
|
||||
|
||||
`MCPServerStdio`, `MCPServerSse`, `MCPServerStreamableHttp`는 모두 `require_approval`을 지원합니다.
|
||||
`MCPServerStdio`, `MCPServerSse`, `MCPServerStreamableHttp`는 모두 `require_approval`을 받습니다.
|
||||
|
||||
지원 형식:
|
||||
지원되는 형식:
|
||||
|
||||
- 모든 도구에 대해 `"always"` 또는 `"never"`
|
||||
- `True` / `False`(always/never와 동일)
|
||||
- 도구별 맵(예: `{"delete_file": "always", "read_file": "never"}`)
|
||||
- 그룹 객체:
|
||||
`{"always": {"tool_names": [...]}, "never": {"tool_names": [...]}}`
|
||||
- `True` / `False`(`always`/`never`와 동일)
|
||||
- 도구별 맵, 예: `{"delete_file": "always", "read_file": "never"}`
|
||||
- 그룹화된 객체: `{"always": {"tool_names": [...]}, "never": {"tool_names": [...]}}`
|
||||
|
||||
```python
|
||||
async with MCPServerStreamableHttp(
|
||||
@@ -244,11 +242,11 @@ async with MCPServerStreamableHttp(
|
||||
...
|
||||
```
|
||||
|
||||
전체 일시정지/재개 흐름은 [휴먼인더루프](human_in_the_loop.md) 및 `examples/mcp/get_all_mcp_tools_example/main.py`를 참고하세요.
|
||||
전체 일시 중지/재개 흐름은 [휴먼인더루프 (HITL)](human_in_the_loop.md)와 `examples/mcp/get_all_mcp_tools_example/main.py`를 참조하세요.
|
||||
|
||||
### `tool_meta_resolver`를 사용한 호출별 메타데이터
|
||||
### 호출별 메타데이터와 `tool_meta_resolver`
|
||||
|
||||
MCP 서버가 `_meta`에 요청 메타데이터(예: 테넌트 ID 또는 트레이스 컨텍스트)를 기대한다면 `tool_meta_resolver`를 사용하세요. 아래 예제는 `Runner.run(...)`에 `context`로 `dict`를 전달한다고 가정합니다.
|
||||
MCP 서버가 `_meta`에서 요청 메타데이터(예: 테넌트 ID 또는 트레이스 컨텍스트)를 기대하는 경우 `tool_meta_resolver`를 사용하세요. 아래 예시는 `Runner.run(...)`에 `context`로 `dict`를 전달한다고 가정합니다.
|
||||
|
||||
```python
|
||||
from agents.mcp import MCPServerStreamableHttp, MCPToolMetaContext
|
||||
@@ -269,20 +267,19 @@ server = MCPServerStreamableHttp(
|
||||
)
|
||||
```
|
||||
|
||||
실행 컨텍스트가 Pydantic 모델, dataclass 또는 사용자 정의 클래스라면 대신 속성 접근으로 테넌트 ID를 읽으세요.
|
||||
실행 컨텍스트가 Pydantic 모델, dataclass 또는 사용자 지정 클래스라면 속성 접근으로 테넌트 ID를 읽으세요.
|
||||
|
||||
### MCP 도구 출력: 텍스트 및 이미지
|
||||
### MCP 도구 출력: 텍스트와 이미지
|
||||
|
||||
MCP 도구가 이미지 콘텐츠를 반환하면 SDK가 이를 이미지 도구 출력 항목으로 자동 매핑합니다. 텍스트/이미지 혼합 응답은 출력 항목 목록으로 전달되므로 에이전트는 일반 함수 도구의 이미지 출력과 동일한 방식으로 MCP 이미지 결과를 소비할 수 있습니다.
|
||||
MCP 도구가 이미지 콘텐츠를 반환하면 SDK는 이를 이미지 도구 출력 항목으로 자동 매핑합니다. 텍스트/이미지 혼합 응답은 출력 항목 목록으로 전달되므로, 에이전트는 일반 함수 도구의 이미지 출력을 소비하는 것과 같은 방식으로 MCP 이미지 결과를 소비할 수 있습니다.
|
||||
|
||||
## 3. SSE 기반 HTTP MCP 서버
|
||||
## 3. HTTP with SSE MCP 서버
|
||||
|
||||
!!! warning
|
||||
|
||||
MCP 프로젝트는 Server-Sent Events 전송 방식을 더 이상 권장하지 않습니다. 새 통합에는 Streamable HTTP 또는 stdio를 우선 사용하고, SSE는 레거시 서버에만 유지하세요
|
||||
MCP 프로젝트는 Server-Sent Events 전송 방식을 deprecated 처리했습니다. 새 통합에는 Streamable HTTP 또는 stdio를 선호하고 SSE는 레거시 서버에만 유지하세요.
|
||||
|
||||
MCP 서버가 SSE 기반 HTTP 전송 방식을 구현한 경우
|
||||
[`MCPServerSse`][agents.mcp.server.MCPServerSse]를 인스턴스화하세요. 전송 방식 외에는 API가 Streamable HTTP 서버와 동일합니다.
|
||||
MCP 서버가 HTTP with SSE 전송 방식을 구현하는 경우 [`MCPServerSse`][agents.mcp.server.MCPServerSse]를 인스턴스화하세요. 전송 방식을 제외하면 API는 Streamable HTTP 서버와 동일합니다.
|
||||
|
||||
```python
|
||||
|
||||
@@ -311,8 +308,7 @@ async with MCPServerSse(
|
||||
|
||||
## 4. stdio MCP 서버
|
||||
|
||||
로컬 서브프로세스로 실행되는 MCP 서버에는 [`MCPServerStdio`][agents.mcp.server.MCPServerStdio]를 사용하세요. SDK가 프로세스를 생성하고
|
||||
파이프를 열린 상태로 유지하며, 컨텍스트 매니저가 종료되면 자동으로 닫습니다. 이 옵션은 빠른 개념 검증이나 서버가 명령줄 엔트리 포인트만 노출할 때 유용합니다.
|
||||
로컬 하위 프로세스로 실행되는 MCP 서버에는 [`MCPServerStdio`][agents.mcp.server.MCPServerStdio]를 사용하세요. SDK는 프로세스를 생성하고 파이프를 열린 상태로 유지하며, 컨텍스트 관리자가 종료될 때 자동으로 닫습니다. 이 옵션은 빠른 개념 증명이나 서버가 명령줄 엔트리 포인트만 노출하는 경우에 유용합니다.
|
||||
|
||||
```python
|
||||
from pathlib import Path
|
||||
@@ -338,10 +334,9 @@ async with MCPServerStdio(
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
## 5. MCP 서버 매니저
|
||||
## 5. MCP 서버 관리자
|
||||
|
||||
여러 MCP 서버가 있는 경우 `MCPServerManager`를 사용해 미리 연결하고, 연결된 하위 집합을 에이전트에 노출하세요.
|
||||
생성자 옵션과 재연결 동작은 [MCPServerManager API 참조](ref/mcp/manager.md)를 참고하세요.
|
||||
MCP 서버가 여러 개 있다면 `MCPServerManager`를 사용해 미리 연결하고 연결된 하위 집합을 에이전트에 노출하세요. 생성자 옵션 및 재연결 동작은 [MCPServerManager API 참조](ref/mcp/manager.md)를 참조하세요.
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
@@ -362,26 +357,25 @@ async with MCPServerManager(servers) as manager:
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
핵심 동작:
|
||||
주요 동작:
|
||||
|
||||
- `drop_failed_servers=True`(기본값)일 때 `active_servers`에는 연결에 성공한 서버만 포함됩니다
|
||||
- 실패는 `failed_servers`와 `errors`에 추적됩니다
|
||||
- 첫 연결 실패에서 예외를 발생시키려면 `strict=True`로 설정하세요
|
||||
- 실패한 서버만 재시도하려면 `reconnect(failed_only=True)`, 모든 서버를 재시작하려면 `reconnect(failed_only=False)`를 호출하세요
|
||||
- 라이프사이클 동작을 조정하려면 `connect_timeout_seconds`, `cleanup_timeout_seconds`, `connect_in_parallel`을 사용하세요
|
||||
- `drop_failed_servers=True`(기본값)인 경우 `active_servers`에는 성공적으로 연결된 서버만 포함됩니다.
|
||||
- 실패는 `failed_servers`와 `errors`에 추적됩니다.
|
||||
- 첫 번째 연결 실패 시 오류를 발생시키려면 `strict=True`를 설정하세요.
|
||||
- 실패한 서버를 다시 시도하려면 `reconnect(failed_only=True)`를 호출하고, 모든 서버를 다시 시작하려면 `reconnect(failed_only=False)`를 호출하세요.
|
||||
- 수명 주기 동작을 조정하려면 `connect_timeout_seconds`, `cleanup_timeout_seconds`, `connect_in_parallel`을 사용하세요.
|
||||
|
||||
## 공통 서버 기능
|
||||
|
||||
아래 섹션은 MCP 서버 전송 방식 전반에 적용됩니다(API 표면은 서버 클래스에 따라 정확히 달라질 수 있음).
|
||||
아래 섹션은 MCP 서버 전송 방식 전반에 적용됩니다(정확한 API 범위는 서버 클래스에 따라 달라짐).
|
||||
|
||||
## 도구 필터링
|
||||
|
||||
각 MCP 서버는 도구 필터를 지원하므로 에이전트에 필요한 함수만 노출할 수 있습니다. 필터링은
|
||||
생성 시점이나 실행별 동적으로 수행할 수 있습니다.
|
||||
각 MCP 서버는 도구 필터를 지원하므로 에이전트에 필요한 함수만 노출할 수 있습니다. 필터링은 생성 시점에 수행하거나 실행별로 동적으로 수행할 수 있습니다.
|
||||
|
||||
### 정적 도구 필터링
|
||||
|
||||
간단한 허용/차단 목록을 구성하려면 [`create_static_tool_filter`][agents.mcp.create_static_tool_filter]를 사용하세요:
|
||||
간단한 허용/차단 목록을 구성하려면 [`create_static_tool_filter`][agents.mcp.create_static_tool_filter]를 사용하세요.
|
||||
|
||||
```python
|
||||
from pathlib import Path
|
||||
@@ -399,13 +393,11 @@ filesystem_server = MCPServerStdio(
|
||||
)
|
||||
```
|
||||
|
||||
`allowed_tool_names`와 `blocked_tool_names`가 모두 제공되면 SDK는 먼저 허용 목록을 적용한 뒤, 남은 집합에서
|
||||
차단된 도구를 제거합니다.
|
||||
`allowed_tool_names`와 `blocked_tool_names`가 모두 제공되면 SDK는 허용 목록을 먼저 적용한 뒤 남은 집합에서 차단된 도구를 제거합니다.
|
||||
|
||||
### 동적 도구 필터링
|
||||
|
||||
더 정교한 로직이 필요하면 [`ToolFilterContext`][agents.mcp.ToolFilterContext]를 받는 callable을 전달하세요. 해당 callable은
|
||||
동기 또는 비동기일 수 있으며, 도구를 노출해야 하면 `True`를 반환합니다.
|
||||
더 복잡한 로직의 경우 [`ToolFilterContext`][agents.mcp.ToolFilterContext]를 받는 호출 가능 객체를 전달하세요. 호출 가능 객체는 동기 또는 비동기일 수 있으며, 도구를 노출해야 할 때 `True`를 반환합니다.
|
||||
|
||||
```python
|
||||
from pathlib import Path
|
||||
@@ -429,15 +421,15 @@ async with MCPServerStdio(
|
||||
...
|
||||
```
|
||||
|
||||
필터 컨텍스트는 활성 `run_context`, 도구를 요청하는 `agent`, `server_name`을 노출합니다.
|
||||
필터 컨텍스트는 활성 `run_context`, 도구를 요청하는 `agent`, 그리고 `server_name`을 노출합니다.
|
||||
|
||||
## 프롬프트
|
||||
|
||||
MCP 서버는 에이전트 instructions를 동적으로 생성하는 프롬프트도 제공할 수 있습니다. 프롬프트를 지원하는 서버는 두 가지
|
||||
메서드를 노출합니다:
|
||||
MCP 서버는 에이전트 지침을 동적으로 생성하는 프롬프트도 제공할 수 있습니다. 프롬프트를 지원하는 서버는 두 가지
|
||||
메서드를 노출합니다.
|
||||
|
||||
- `list_prompts()`는 사용 가능한 프롬프트 템플릿을 열거합니다
|
||||
- `get_prompt(name, arguments)`는 선택적으로 매개변수와 함께 구체적인 프롬프트를 가져옵니다
|
||||
- `list_prompts()`는 사용 가능한 프롬프트 템플릿을 열거합니다.
|
||||
- `get_prompt(name, arguments)`는 매개변수를 선택적으로 포함하여 구체적인 프롬프트를 가져옵니다.
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
@@ -457,21 +449,19 @@ agent = Agent(
|
||||
|
||||
## 캐싱
|
||||
|
||||
모든 에이전트 실행은 각 MCP 서버에서 `list_tools()`를 호출합니다. 원격 서버는 눈에 띄는 지연 시간을 유발할 수 있으므로 모든 MCP
|
||||
서버 클래스는 `cache_tools_list` 옵션을 노출합니다. 도구 정의가 자주
|
||||
변경되지 않는다고 확신할 때만 이를 `True`로 설정하세요. 나중에 최신 목록을 강제로 가져오려면 서버 인스턴스에서 `invalidate_tools_cache()`를 호출하세요.
|
||||
각 에이전트 실행은 모든 MCP 서버에서 `list_tools()`를 호출합니다. 원격 서버는 눈에 띄는 지연 시간을 유발할 수 있으므로, 모든 MCP 서버 클래스는 `cache_tools_list` 옵션을 노출합니다. 도구 정의가 자주 변경되지 않는다고 확신하는 경우에만 이를 `True`로 설정하세요. 나중에 최신 목록을 강제로 가져오려면 서버 인스턴스에서 `invalidate_tools_cache()`를 호출하세요.
|
||||
|
||||
## 트레이싱
|
||||
|
||||
[트레이싱](./tracing.md)은 다음을 포함해 MCP 활동을 자동으로 수집합니다:
|
||||
[트레이싱](./tracing.md)은 다음을 포함한 MCP 활동을 자동으로 캡처합니다.
|
||||
|
||||
1. 도구 목록 조회를 위한 MCP 서버 호출
|
||||
1. 도구 목록을 나열하기 위한 MCP 서버 호출
|
||||
2. 도구 호출의 MCP 관련 정보
|
||||
|
||||

|
||||

|
||||
|
||||
## 추가 읽을거리
|
||||
## 추가 자료
|
||||
|
||||
- [Model Context Protocol](https://modelcontextprotocol.io/) – 명세 및 설계 가이드
|
||||
- [examples/mcp](https://github.com/openai/openai-agents-python/tree/main/examples/mcp) – 실행 가능한 stdio, SSE, Streamable HTTP 샘플
|
||||
- [Model Context Protocol](https://modelcontextprotocol.io/) – 사양 및 설계 가이드
|
||||
- [examples/mcp](https://github.com/openai/openai-agents-python/tree/main/examples/mcp) – 실행 가능한 stdio, SSE 및 Streamable HTTP 샘플
|
||||
- [examples/hosted_mcp](https://github.com/openai/openai-agents-python/tree/main/examples/hosted_mcp) – 승인 및 커넥터를 포함한 완전한 호스티드 MCP 데모
|
||||
+180
-137
@@ -4,42 +4,42 @@ search:
|
||||
---
|
||||
# 모델
|
||||
|
||||
Agents SDK는 OpenAI 모델을 두 가지 방식으로 즉시 사용할 수 있도록 지원합니다:
|
||||
Agents SDK는 두 가지 방식의 OpenAI 모델을 기본 지원합니다:
|
||||
|
||||
- **권장**: 새 [Responses API](https://platform.openai.com/docs/api-reference/responses)를 사용해 OpenAI API를 호출하는 [`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel]
|
||||
- **권장**: 새로운 [Responses API](https://platform.openai.com/docs/api-reference/responses)를 사용해 OpenAI API를 호출하는 [`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel]
|
||||
- [Chat Completions API](https://platform.openai.com/docs/api-reference/chat)를 사용해 OpenAI API를 호출하는 [`OpenAIChatCompletionsModel`][agents.models.openai_chatcompletions.OpenAIChatCompletionsModel]
|
||||
|
||||
## 모델 설정 선택
|
||||
|
||||
현재 설정에 맞는 가장 단순한 경로부터 시작하세요:
|
||||
설정에 맞는 가장 단순한 경로부터 시작하세요:
|
||||
|
||||
| 다음을 하려는 경우 | 권장 경로 | 자세히 보기 |
|
||||
| 하려는 작업 | 권장 경로 | 더 읽기 |
|
||||
| --- | --- | --- |
|
||||
| OpenAI 모델만 사용 | 기본 OpenAI provider와 Responses 모델 경로 사용 | [OpenAI 모델](#openai-models) |
|
||||
| websocket 전송으로 OpenAI Responses API 사용 | Responses 모델 경로를 유지하고 websocket 전송 활성화 | [Responses WebSocket 전송](#responses-websocket-transport) |
|
||||
| OpenAI가 아닌 provider 하나 사용 | 내장 provider 통합 지점부터 시작 | [OpenAI가 아닌 모델](#non-openai-models) |
|
||||
| 에이전트 전반에서 모델 또는 provider 혼합 | 실행(run)별 또는 에이전트별로 provider 선택 후 기능 차이 검토 | [하나의 워크플로에서 모델 혼합](#mixing-models-in-one-workflow) 및 [provider 간 모델 혼합](#mixing-models-across-providers) |
|
||||
| OpenAI 모델만 사용 | 기본 OpenAI 제공자를 Responses 모델 경로와 함께 사용 | [OpenAI 모델](#openai-models) |
|
||||
| 웹소켓 전송을 통해 OpenAI Responses API 사용 | Responses 모델 경로를 유지하고 웹소켓 전송 활성화 | [Responses WebSocket 전송](#responses-websocket-transport) |
|
||||
| 단일 non-OpenAI 제공자 사용 | 내장 제공자 통합 지점부터 시작 | [non-OpenAI 모델](#non-openai-models) |
|
||||
| 에이전트 간 모델 또는 제공자 혼합 | 실행별 또는 에이전트별로 제공자를 선택하고 기능 차이 검토 | [단일 워크플로 내 모델 혼합](#mixing-models-in-one-workflow) 및 [제공자 간 모델 혼합](#mixing-models-across-providers) |
|
||||
| 고급 OpenAI Responses 요청 설정 조정 | OpenAI Responses 경로에서 `ModelSettings` 사용 | [고급 OpenAI Responses 설정](#advanced-openai-responses-settings) |
|
||||
| OpenAI가 아닌 또는 혼합 provider 라우팅에 서드파티 어댑터 사용 | 지원되는 베타 어댑터를 비교하고 출시할 provider 경로 검증 | [서드파티 어댑터](#third-party-adapters) |
|
||||
| non-OpenAI 또는 혼합 제공자 라우팅을 위한 서드 파티 어댑터 사용 | 지원되는 베타 어댑터를 비교하고 출시하려는 제공자 경로 검증 | [서드 파티 어댑터](#third-party-adapters) |
|
||||
|
||||
## OpenAI 모델
|
||||
|
||||
대부분의 OpenAI 전용 앱에서는 기본 OpenAI provider와 문자열 모델 이름을 사용하고, Responses 모델 경로를 유지하는 것이 권장됩니다
|
||||
대부분의 OpenAI 전용 앱에서는 기본 OpenAI 제공자와 문자열 모델 이름을 사용하고 Responses 모델 경로를 유지하는 것이 권장 경로입니다.
|
||||
|
||||
`Agent` 초기화 시 모델을 지정하지 않으면 기본 모델이 사용됩니다. 현재 기본값은 호환성과 낮은 지연 시간을 위해 [`gpt-4.1`](https://developers.openai.com/api/docs/models/gpt-4.1)입니다. 접근 권한이 있다면, 명시적인 `model_settings`를 유지하면서 더 높은 품질을 위해 에이전트를 [`gpt-5.4`](https://developers.openai.com/api/docs/models/gpt-5.4)로 설정하는 것을 권장합니다
|
||||
`Agent`를 초기화할 때 모델을 지정하지 않으면 기본 모델이 사용됩니다. 현재 기본값은 저지연 에이전트 워크플로를 위해 `reasoning.effort="none"` 및 `verbosity="low"`가 설정된 [`gpt-5.4-mini`](https://developers.openai.com/api/docs/models/gpt-5.4-mini)입니다. 액세스 권한이 있다면 명시적인 `model_settings`를 유지하면서 더 높은 품질을 위해 에이전트를 [`gpt-5.5`](https://developers.openai.com/api/docs/models/gpt-5.5)로 설정하는 것을 권장합니다.
|
||||
|
||||
[`gpt-5.4`](https://developers.openai.com/api/docs/models/gpt-5.4) 같은 다른 모델로 전환하려면 에이전트를 구성하는 방법이 두 가지 있습니다
|
||||
[`gpt-5.5`](https://developers.openai.com/api/docs/models/gpt-5.5) 같은 다른 모델로 전환하려면 에이전트를 구성하는 방법이 두 가지 있습니다.
|
||||
|
||||
### 기본 모델
|
||||
|
||||
첫째, 사용자 지정 모델을 설정하지 않은 모든 에이전트에서 특정 모델을 일관되게 사용하려면, 에이전트를 실행하기 전에 `OPENAI_DEFAULT_MODEL` 환경 변수를 설정하세요
|
||||
첫째, 사용자 지정 모델을 설정하지 않는 모든 에이전트에 특정 모델을 일관되게 사용하려면 에이전트를 실행하기 전에 `OPENAI_DEFAULT_MODEL` 환경 변수를 설정하세요.
|
||||
|
||||
```bash
|
||||
export OPENAI_DEFAULT_MODEL=gpt-5.4
|
||||
export OPENAI_DEFAULT_MODEL=gpt-5.5
|
||||
python3 my_awesome_agent.py
|
||||
```
|
||||
|
||||
둘째, `RunConfig`를 통해 실행(run) 단위 기본 모델을 설정할 수 있습니다. 에이전트에 모델을 설정하지 않으면 이 실행의 모델이 사용됩니다
|
||||
둘째, `RunConfig`를 통해 실행의 기본 모델을 설정할 수 있습니다. 에이전트에 모델을 설정하지 않으면 이 실행의 모델이 사용됩니다.
|
||||
|
||||
```python
|
||||
from agents import Agent, RunConfig, Runner
|
||||
@@ -52,13 +52,13 @@ agent = Agent(
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Hello",
|
||||
run_config=RunConfig(model="gpt-5.4"),
|
||||
run_config=RunConfig(model="gpt-5.5"),
|
||||
)
|
||||
```
|
||||
|
||||
#### GPT-5 모델
|
||||
|
||||
이 방식으로 [`gpt-5.4`](https://developers.openai.com/api/docs/models/gpt-5.4) 같은 GPT-5 모델을 사용할 때 SDK는 기본 `ModelSettings`를 적용합니다. 대부분의 사용 사례에서 가장 잘 동작하는 설정이 적용됩니다. 기본 모델의 reasoning effort를 조정하려면 사용자 지정 `ModelSettings`를 전달하세요:
|
||||
이 방식으로 [`gpt-5.5`](https://developers.openai.com/api/docs/models/gpt-5.5) 같은 GPT-5 모델을 사용하는 경우 SDK는 기본 `ModelSettings`를 적용합니다. 대부분의 사용 사례에 가장 잘 맞는 값으로 설정됩니다. 기본 모델의 reasoning effort를 조정하려면 자체 `ModelSettings`를 전달하세요:
|
||||
|
||||
```python
|
||||
from openai.types.shared import Reasoning
|
||||
@@ -67,28 +67,28 @@ from agents import Agent, ModelSettings
|
||||
my_agent = Agent(
|
||||
name="My Agent",
|
||||
instructions="You're a helpful agent.",
|
||||
# If OPENAI_DEFAULT_MODEL=gpt-5.4 is set, passing only model_settings works.
|
||||
# If OPENAI_DEFAULT_MODEL=gpt-5.5 is set, passing only model_settings works.
|
||||
# It's also fine to pass a GPT-5 model name explicitly:
|
||||
model="gpt-5.4",
|
||||
model="gpt-5.5",
|
||||
model_settings=ModelSettings(reasoning=Reasoning(effort="high"), verbosity="low")
|
||||
)
|
||||
```
|
||||
|
||||
더 낮은 지연 시간을 위해 `gpt-5.4`에서 `reasoning.effort="none"` 사용을 권장합니다. gpt-4.1 계열( mini 및 nano 변형 포함)도 인터랙티브 에이전트 앱 구축에 여전히 좋은 선택입니다
|
||||
지연 시간을 줄이려면 GPT-5 모델에서 `reasoning.effort="none"`을 사용하는 것이 권장됩니다.
|
||||
|
||||
#### ComputerTool 모델 선택
|
||||
|
||||
에이전트에 [`ComputerTool`][agents.tool.ComputerTool]이 포함된 경우, 실제 Responses 요청의 유효 모델이 SDK가 전송할 컴퓨터 도구 페이로드를 결정합니다. 명시적인 `gpt-5.4` 요청은 GA 내장 `computer` 도구를 사용하고, 명시적인 `computer-use-preview` 요청은 기존 `computer_use_preview` 페이로드를 유지합니다
|
||||
에이전트에 [`ComputerTool`][agents.tool.ComputerTool]이 포함된 경우, 실제 Responses 요청에서 적용되는 모델이 SDK가 보내는 컴퓨터 도구 페이로드를 결정합니다. 명시적인 `gpt-5.5` 요청은 정식 출시(GA) 내장 `computer` 도구를 사용하고, 명시적인 `computer-use-preview` 요청은 기존 `computer_use_preview` 페이로드를 유지합니다.
|
||||
|
||||
주요 예외는 프롬프트 관리 호출입니다. 프롬프트 템플릿이 모델을 소유하고 SDK가 요청에서 `model`을 생략하면, SDK는 프롬프트가 어떤 모델을 고정했는지 추측하지 않기 위해 preview 호환 컴퓨터 페이로드를 기본으로 사용합니다. 이 흐름에서 GA 경로를 유지하려면 요청에 `model="gpt-5.4"`를 명시하거나 `ModelSettings(tool_choice="computer")` 또는 `ModelSettings(tool_choice="computer_use")`로 GA 선택기를 강제하세요
|
||||
프롬프트 관리형 호출이 주요 예외입니다. 프롬프트 템플릿이 모델을 소유하고 SDK가 요청에서 `model`을 생략하는 경우, SDK는 프롬프트가 어떤 모델에 고정되어 있는지 추측하지 않도록 프리뷰 호환 컴퓨터 페이로드를 기본값으로 사용합니다. 해당 흐름에서 GA 경로를 유지하려면 요청에 `model="gpt-5.5"`를 명시하거나 `ModelSettings(tool_choice="computer")` 또는 `ModelSettings(tool_choice="computer_use")`로 GA 선택기를 강제하세요.
|
||||
|
||||
[`ComputerTool`][agents.tool.ComputerTool]이 등록된 상태에서는 `tool_choice="computer"`, `"computer_use"`, `"computer_use_preview"`가 유효 요청 모델에 맞는 내장 선택기로 정규화됩니다. `ComputerTool`이 등록되지 않은 경우에는 이러한 문자열이 일반 함수 이름처럼 계속 동작합니다
|
||||
등록된 [`ComputerTool`][agents.tool.ComputerTool]이 있는 경우 `tool_choice="computer"`, `"computer_use"`, `"computer_use_preview"`는 실제 요청 모델과 일치하는 내장 선택기로 정규화됩니다. 등록된 `ComputerTool`이 없으면 이러한 문자열은 일반 함수 이름처럼 계속 동작합니다.
|
||||
|
||||
preview 호환 요청은 `environment`와 디스플레이 크기를 사전에 직렬화해야 하므로, [`ComputerProvider`][agents.tool.ComputerProvider] 팩토리를 사용하는 프롬프트 관리 흐름은 구체적인 `Computer` 또는 `AsyncComputer` 인스턴스를 전달하거나 요청 전 GA 선택기를 강제해야 합니다. 전체 마이그레이션 세부 사항은 [Tools](../tools.md#computertool-and-the-responses-computer-tool)를 참고하세요
|
||||
프리뷰 호환 요청은 `environment`와 표시 크기를 사전에 직렬화해야 하므로, [`ComputerProvider`][agents.tool.ComputerProvider] 팩터리를 사용하는 프롬프트 관리형 흐름은 구체적인 `Computer` 또는 `AsyncComputer` 인스턴스를 전달하거나 요청을 보내기 전에 GA 선택기를 강제해야 합니다. 전체 마이그레이션 세부 사항은 [도구](../tools.md#computertool-and-the-responses-computer-tool)를 참조하세요.
|
||||
|
||||
#### GPT-5가 아닌 모델
|
||||
|
||||
사용자 지정 `model_settings` 없이 GPT-5가 아닌 모델 이름을 전달하면 SDK는 모든 모델과 호환되는 일반 `ModelSettings`로 되돌아갑니다
|
||||
사용자 지정 `model_settings` 없이 GPT-5가 아닌 모델 이름을 전달하면 SDK는 모든 모델과 호환되는 일반 `ModelSettings`로 되돌아갑니다.
|
||||
|
||||
### Responses 전용 도구 검색 기능
|
||||
|
||||
@@ -98,11 +98,11 @@ preview 호환 요청은 `environment`와 디스플레이 크기를 사전에
|
||||
- [`tool_namespace()`][agents.tool.tool_namespace]
|
||||
- `@function_tool(defer_loading=True)` 및 기타 지연 로딩 Responses 도구 표면
|
||||
|
||||
이 기능들은 Chat Completions 모델과 non-Responses 백엔드에서는 거부됩니다. 지연 로딩 도구를 사용할 때는 에이전트에 `ToolSearchTool()`을 추가하고, 네임스페이스 이름이나 지연 전용 함수 이름을 강제하는 대신 모델이 `auto` 또는 `required` tool choice를 통해 도구를 로드하도록 하세요. 설정 세부 사항과 현재 제약은 [Tools](../tools.md#hosted-tool-search)를 참고하세요
|
||||
이 기능들은 Chat Completions 모델 및 Responses가 아닌 백엔드에서 거부됩니다. 지연 로딩 도구를 사용할 때는 에이전트에 `ToolSearchTool()`을 추가하고, 네임스페이스 이름만 또는 지연 전용 함수 이름을 강제하는 대신 모델이 `auto` 또는 `required` tool choice를 통해 도구를 로드하도록 하세요. 설정 세부 사항과 현재 제약 사항은 [도구](../tools.md#hosted-tool-search)를 참조하세요.
|
||||
|
||||
### Responses WebSocket 전송
|
||||
|
||||
기본적으로 OpenAI Responses API 요청은 HTTP 전송을 사용합니다. OpenAI 기반 모델 사용 시 websocket 전송을 선택할 수 있습니다
|
||||
기본적으로 OpenAI Responses API 요청은 HTTP 전송을 사용합니다. OpenAI 기반 모델을 사용할 때 웹소켓 전송을 선택할 수 있습니다.
|
||||
|
||||
#### 기본 설정
|
||||
|
||||
@@ -112,13 +112,13 @@ from agents import set_default_openai_responses_transport
|
||||
set_default_openai_responses_transport("websocket")
|
||||
```
|
||||
|
||||
이 설정은 기본 OpenAI provider가 해석하는 OpenAI Responses 모델(`"gpt-5.4"` 같은 문자열 모델 이름 포함)에 영향을 줍니다
|
||||
이는 기본 OpenAI 제공자가 해석하는 OpenAI Responses 모델(예: `"gpt-5.5"` 같은 문자열 모델 이름)에 영향을 줍니다.
|
||||
|
||||
전송 방식 선택은 SDK가 모델 이름을 모델 인스턴스로 해석할 때 발생합니다. 구체적인 [`Model`][agents.models.interface.Model] 객체를 전달하면 전송 방식은 이미 고정됩니다: [`OpenAIResponsesWSModel`][agents.models.openai_responses.OpenAIResponsesWSModel]은 websocket, [`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel]은 HTTP, [`OpenAIChatCompletionsModel`][agents.models.openai_chatcompletions.OpenAIChatCompletionsModel]은 Chat Completions를 유지합니다. `RunConfig(model_provider=...)`를 전달하면 전역 기본값 대신 해당 provider가 전송 선택을 제어합니다
|
||||
전송 선택은 SDK가 모델 이름을 모델 인스턴스로 해석할 때 이루어집니다. 구체적인 [`Model`][agents.models.interface.Model] 객체를 전달하면 해당 전송은 이미 고정되어 있습니다. [`OpenAIResponsesWSModel`][agents.models.openai_responses.OpenAIResponsesWSModel]은 웹소켓을 사용하고, [`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel]은 HTTP를 사용하며, [`OpenAIChatCompletionsModel`][agents.models.openai_chatcompletions.OpenAIChatCompletionsModel]은 Chat Completions에 유지됩니다. `RunConfig(model_provider=...)`를 전달하면 전역 기본값 대신 해당 제공자가 전송 선택을 제어합니다.
|
||||
|
||||
#### provider 또는 실행(run) 수준 설정
|
||||
#### 제공자 또는 실행 수준 설정
|
||||
|
||||
websocket 전송은 provider별 또는 실행(run)별로도 설정할 수 있습니다:
|
||||
제공자별 또는 실행별로도 웹소켓 전송을 구성할 수 있습니다:
|
||||
|
||||
```python
|
||||
from agents import Agent, OpenAIProvider, RunConfig, Runner
|
||||
@@ -127,6 +127,8 @@ provider = OpenAIProvider(
|
||||
use_responses_websocket=True,
|
||||
# Optional; if omitted, OPENAI_WEBSOCKET_BASE_URL is used when set.
|
||||
websocket_base_url="wss://your-proxy.example/v1",
|
||||
# Optional low-level websocket keepalive settings.
|
||||
responses_websocket_options={"ping_interval": 20.0, "ping_timeout": 60.0},
|
||||
)
|
||||
|
||||
agent = Agent(name="Assistant")
|
||||
@@ -137,7 +139,7 @@ result = await Runner.run(
|
||||
)
|
||||
```
|
||||
|
||||
OpenAI 기반 provider는 선택적 에이전트 등록 설정도 허용합니다. 이는 OpenAI 설정이 harness ID 같은 provider 수준 등록 메타데이터를 기대하는 경우를 위한 고급 옵션입니다
|
||||
OpenAI 기반 제공자는 선택적 에이전트 등록 구성도 받습니다. 이는 OpenAI 설정이 harness ID 같은 제공자 수준 등록 메타데이터를 기대하는 경우를 위한 고급 옵션입니다.
|
||||
|
||||
```python
|
||||
from agents import (
|
||||
@@ -163,14 +165,14 @@ result = await Runner.run(
|
||||
|
||||
#### `MultiProvider`를 사용한 고급 라우팅
|
||||
|
||||
접두사 기반 모델 라우팅(예: 한 실행에서 `openai/...`와 `any-llm/...` 모델 이름 혼합)이 필요하면 [`MultiProvider`][agents.MultiProvider]를 사용하고 거기서 `openai_use_responses_websocket=True`를 설정하세요
|
||||
접두사 기반 모델 라우팅이 필요한 경우(예: 한 실행에서 `openai/...` 및 `any-llm/...` 모델 이름을 혼합) [`MultiProvider`][agents.MultiProvider]를 사용하고 거기에 `openai_use_responses_websocket=True`를 설정하세요.
|
||||
|
||||
`MultiProvider`는 두 가지 기존 기본값을 유지합니다:
|
||||
|
||||
- `openai/...`는 OpenAI provider의 별칭으로 처리되어, `openai/gpt-4.1`은 모델 `gpt-4.1`로 라우팅됩니다
|
||||
- 알 수 없는 접두사는 통과되지 않고 `UserError`를 발생시킵니다
|
||||
- `openai/...`는 OpenAI 제공자의 별칭으로 처리되므로 `openai/gpt-4.1`은 모델 `gpt-4.1`로 라우팅됩니다.
|
||||
- 알 수 없는 접두사는 그대로 전달되는 대신 `UserError`를 발생시킵니다.
|
||||
|
||||
OpenAI provider를 문자 그대로의 네임스페이스 모델 ID를 기대하는 OpenAI 호환 엔드포인트에 연결하는 경우, 명시적으로 pass-through 동작을 활성화하세요. websocket 활성 설정에서도 `MultiProvider`에 `openai_use_responses_websocket=True`를 유지하세요:
|
||||
OpenAI 제공자를 리터럴 네임스페이스 모델 ID를 기대하는 OpenAI 호환 엔드포인트로 지정할 때는 패스스루(pass-through) 동작을 명시적으로 선택하세요. 웹소켓 활성화 설정에서는 `MultiProvider`에도 `openai_use_responses_websocket=True`를 유지하세요:
|
||||
|
||||
```python
|
||||
from agents import Agent, MultiProvider, RunConfig, Runner
|
||||
@@ -196,38 +198,40 @@ result = await Runner.run(
|
||||
)
|
||||
```
|
||||
|
||||
백엔드가 문자 그대로 `openai/...` 문자열을 기대하면 `openai_prefix_mode="model_id"`를 사용하세요. 백엔드가 `openrouter/openai/gpt-4.1-mini` 같은 다른 네임스페이스 모델 ID를 기대하면 `unknown_prefix_mode="model_id"`를 사용하세요. 이 옵션들은 websocket 전송 외부의 `MultiProvider`에서도 동작합니다; 이 예제는 이 섹션에서 설명한 전송 설정의 일부이므로 websocket을 활성화한 상태를 유지합니다. 동일한 옵션은 [`responses_websocket_session()`][agents.responses_websocket_session]에서도 사용할 수 있습니다
|
||||
백엔드가 리터럴 `openai/...` 문자열을 기대할 때는 `openai_prefix_mode="model_id"`를 사용하세요. 백엔드가 `openrouter/openai/gpt-4.1-mini` 같은 다른 네임스페이스 모델 ID를 기대할 때는 `unknown_prefix_mode="model_id"`를 사용하세요. 이러한 옵션은 웹소켓 전송 외부의 `MultiProvider`에서도 동작합니다. 이 예제에서는 이 섹션에서 설명하는 전송 설정의 일부이므로 웹소켓을 활성화된 상태로 유지합니다. 동일한 옵션은 [`responses_websocket_session()`][agents.responses_websocket_session]에서도 사용할 수 있습니다.
|
||||
|
||||
`MultiProvider`를 통해 라우팅하면서 동일한 provider 수준 등록 메타데이터가 필요하면 `openai_agent_registration=OpenAIAgentRegistrationConfig(...)`를 전달하면 하위 OpenAI provider로 전달됩니다
|
||||
`MultiProvider`를 통해 라우팅하면서 동일한 제공자 수준 등록 메타데이터가 필요하다면 `openai_agent_registration=OpenAIAgentRegistrationConfig(...)`를 전달하세요. 그러면 기본 OpenAI 제공자로 전달됩니다.
|
||||
|
||||
사용자 지정 OpenAI 호환 엔드포인트 또는 프록시를 사용하는 경우 websocket 전송에도 호환되는 websocket `/responses` 엔드포인트가 필요합니다. 이러한 설정에서는 `websocket_base_url`을 명시적으로 설정해야 할 수 있습니다
|
||||
사용자 지정 OpenAI 호환 엔드포인트 또는 프록시를 사용하는 경우, 웹소켓 전송에는 호환되는 웹소켓 `/responses` 엔드포인트도 필요합니다. 이러한 설정에서는 `websocket_base_url`을 명시적으로 설정해야 할 수 있습니다.
|
||||
|
||||
#### 참고
|
||||
#### 참고 사항
|
||||
|
||||
- 이는 websocket 전송 기반 Responses API이며 [Realtime API](../realtime/guide.md)가 아닙니다. Chat Completions나 Responses websocket `/responses` 엔드포인트를 지원하지 않는 OpenAI가 아닌 provider에는 적용되지 않습니다
|
||||
- 환경에 `websockets` 패키지가 없다면 설치하세요
|
||||
- websocket 전송 활성화 후 [`Runner.run_streamed()`][agents.run.Runner.run_streamed]를 바로 사용할 수 있습니다. 여러 턴 워크플로에서 턴 간(및 중첩된 agent-as-tool 호출 간) 동일 websocket 연결을 재사용하려면 [`responses_websocket_session()`][agents.responses_websocket_session] 헬퍼를 권장합니다. [Running agents](../running_agents.md) 가이드와 [`examples/basic/stream_ws.py`](https://github.com/openai/openai-agents-python/tree/main/examples/basic/stream_ws.py)를 참고하세요
|
||||
- 이는 웹소켓 전송을 통한 Responses API이지 [Realtime API](../realtime/guide.md)가 아닙니다. Responses 웹소켓 `/responses` 엔드포인트를 지원하지 않는 한 Chat Completions 또는 non-OpenAI 제공자에는 적용되지 않습니다.
|
||||
- 환경에서 아직 사용할 수 없다면 `websockets` 패키지를 설치하세요.
|
||||
- 웹소켓 전송을 활성화한 후 [`Runner.run_streamed()`][agents.run.Runner.run_streamed]를 직접 사용할 수 있습니다. 동일한 웹소켓 연결을 여러 턴(및 중첩 agent-as-tool 호출)에서 재사용하려는 멀티턴 워크플로에는 [`responses_websocket_session()`][agents.responses_websocket_session] 헬퍼가 권장됩니다. [에이전트 실행](../running_agents.md) 가이드와 [`examples/basic/stream_ws.py`](https://github.com/openai/openai-agents-python/tree/main/examples/basic/stream_ws.py)를 참조하세요.
|
||||
- 긴 reasoning 턴이나 지연 시간 급증이 있는 네트워크에서는 `responses_websocket_options`로 웹소켓 keepalive 동작을 사용자 지정하세요. 지연된 pong 프레임을 허용하려면 `ping_timeout`을 늘리거나, ping은 활성화한 상태로 하트비트 타임아웃을 비활성화하려면 `ping_timeout=None`을 설정하세요. 신뢰성이 웹소켓 지연 시간보다 중요할 때는 HTTP/SSE 전송을 선호하세요.
|
||||
- 기본적으로 SDK는 수신 메시지 크기 제한(`max_size=None`)을 비활성화합니다. 프록시 뒤에서 장기간 실행되는 에이전트 프로세스나 메모리 제약이 있는 컨테이너에서는 메시지당 메모리 사용량을 제한하기 위해 `responses_websocket_options={"max_size": 8 * 1024 * 1024}`를 설정하세요.
|
||||
|
||||
## OpenAI가 아닌 모델
|
||||
## non-OpenAI 모델
|
||||
|
||||
OpenAI가 아닌 provider가 필요하면 SDK의 내장 provider 통합 지점부터 시작하세요. 많은 설정에서는 서드파티 어댑터를 추가하지 않아도 충분합니다. 각 패턴의 예제는 [examples/model_providers](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/)에 있습니다
|
||||
non-OpenAI 제공자가 필요하다면 SDK의 내장 제공자 통합 지점부터 시작하세요. 많은 설정에서는 서드 파티 어댑터를 추가하지 않아도 이것으로 충분합니다. 각 패턴의 예제는 [examples/model_providers](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/)에 있습니다.
|
||||
|
||||
### OpenAI가 아닌 provider 통합 방식
|
||||
### non-OpenAI 제공자 통합 방식
|
||||
|
||||
| 접근 방식 | 사용 시점 | 범위 |
|
||||
| --- | --- | --- |
|
||||
| [`set_default_openai_client`][agents.set_default_openai_client] | OpenAI 호환 엔드포인트 하나를 대부분 또는 전체 에이전트의 기본값으로 사용해야 할 때 | 전역 기본값 |
|
||||
| [`ModelProvider`][agents.models.interface.ModelProvider] | 사용자 지정 provider 하나를 단일 실행(run)에 적용해야 할 때 | 실행(run)별 |
|
||||
| [`Agent.model`][agents.agent.Agent.model] | 서로 다른 에이전트에 서로 다른 provider 또는 구체적인 모델 객체가 필요할 때 | 에이전트별 |
|
||||
| 서드파티 어댑터 | 내장 경로가 제공하지 않는 어댑터 관리 provider 범위 또는 라우팅이 필요할 때 | [서드파티 어댑터](#third-party-adapters) 참고 |
|
||||
| [`set_default_openai_client`][agents.set_default_openai_client] | 하나의 OpenAI 호환 엔드포인트가 대부분 또는 모든 에이전트의 기본값이어야 할 때 | 전역 기본값 |
|
||||
| [`ModelProvider`][agents.models.interface.ModelProvider] | 하나의 사용자 지정 제공자를 단일 실행에 적용해야 할 때 | 실행별 |
|
||||
| [`Agent.model`][agents.agent.Agent.model] | 에이전트마다 서로 다른 제공자 또는 구체적인 모델 객체가 필요할 때 | 에이전트별 |
|
||||
| 서드 파티 어댑터 | 내장 경로가 제공하지 않는 어댑터 관리형 제공자 커버리지 또는 라우팅이 필요할 때 | [서드 파티 어댑터](#third-party-adapters) 참조 |
|
||||
|
||||
다음 내장 경로로 다른 LLM provider를 통합할 수 있습니다:
|
||||
다음 내장 경로를 사용해 다른 LLM 제공자를 통합할 수 있습니다:
|
||||
|
||||
1. [`set_default_openai_client`][agents.set_default_openai_client]는 `AsyncOpenAI` 인스턴스를 LLM 클라이언트로 전역 사용하려는 경우에 유용합니다. LLM provider가 OpenAI 호환 API 엔드포인트를 제공하고 `base_url` 및 `api_key`를 설정할 수 있는 경우를 위한 방식입니다. 설정 가능한 예제는 [examples/model_providers/custom_example_global.py](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/custom_example_global.py)를 참고하세요
|
||||
2. [`ModelProvider`][agents.models.interface.ModelProvider]는 `Runner.run` 수준에서 동작합니다. 이를 통해 "이 실행의 모든 에이전트에 사용자 지정 모델 provider를 사용"이라고 지정할 수 있습니다. 설정 가능한 예제는 [examples/model_providers/custom_example_provider.py](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/custom_example_provider.py)를 참고하세요
|
||||
3. [`Agent.model`][agents.agent.Agent.model]은 특정 Agent 인스턴스에 모델을 지정할 수 있게 해줍니다. 이를 통해 서로 다른 에이전트에 서로 다른 provider를 혼합해 사용할 수 있습니다. 설정 가능한 예제는 [examples/model_providers/custom_example_agent.py](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/custom_example_agent.py)를 참고하세요
|
||||
1. [`set_default_openai_client`][agents.set_default_openai_client]는 `AsyncOpenAI` 인스턴스를 LLM 클라이언트로 전역적으로 사용하려는 경우에 유용합니다. 이는 LLM 제공자가 OpenAI 호환 API 엔드포인트를 가지고 있고 `base_url`과 `api_key`를 설정할 수 있는 경우를 위한 것입니다. 구성 가능한 예제는 [examples/model_providers/custom_example_global.py](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/custom_example_global.py)를 참조하세요.
|
||||
2. [`ModelProvider`][agents.models.interface.ModelProvider]는 `Runner.run` 수준입니다. 이를 통해 "이 실행의 모든 에이전트에 사용자 지정 모델 제공자를 사용"하도록 지정할 수 있습니다. 구성 가능한 예제는 [examples/model_providers/custom_example_provider.py](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/custom_example_provider.py)를 참조하세요.
|
||||
3. [`Agent.model`][agents.agent.Agent.model]을 사용하면 특정 Agent 인스턴스에서 모델을 지정할 수 있습니다. 이를 통해 서로 다른 에이전트에 대해 서로 다른 제공자를 혼합해 사용할 수 있습니다. 구성 가능한 예제는 [examples/model_providers/custom_example_agent.py](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/custom_example_agent.py)를 참조하세요.
|
||||
|
||||
`platform.openai.com`의 API 키가 없는 경우 [`set_tracing_disabled()`]로 트레이싱을 비활성화하거나 [다른 트레이싱 프로세서](../tracing.md)를 설정하는 것을 권장합니다
|
||||
`platform.openai.com`의 API 키가 없는 경우 `set_tracing_disabled()`를 통해 트레이싱을 비활성화하거나 [다른 트레이싱 프로세서](../tracing.md)를 설정하는 것을 권장합니다.
|
||||
|
||||
``` python
|
||||
from agents import Agent, AsyncOpenAI, OpenAIChatCompletionsModel, set_tracing_disabled
|
||||
@@ -242,19 +246,19 @@ agent= Agent(name="Helping Agent", instructions="You are a Helping Agent", model
|
||||
|
||||
!!! note
|
||||
|
||||
이 예제들에서는 많은 LLM provider가 아직 Responses API를 지원하지 않기 때문에 Chat Completions API/모델을 사용합니다. LLM provider가 이를 지원한다면 Responses 사용을 권장합니다
|
||||
이 예제들에서는 Chat Completions API/모델을 사용합니다. 아직 많은 LLM 제공자가 Responses API를 지원하지 않기 때문입니다. LLM 제공자가 이를 지원한다면 Responses를 사용하는 것을 권장합니다.
|
||||
|
||||
## 하나의 워크플로에서 모델 혼합
|
||||
## 단일 워크플로 내 모델 혼합
|
||||
|
||||
단일 워크플로 내에서 에이전트마다 서로 다른 모델을 사용하고 싶을 수 있습니다. 예를 들어 분류(triage)에는 더 작고 빠른 모델을, 복잡한 작업에는 더 크고 성능이 높은 모델을 사용할 수 있습니다. [`Agent`][agents.Agent]를 구성할 때 다음 중 하나로 특정 모델을 선택할 수 있습니다:
|
||||
단일 워크플로 안에서 각 에이전트마다 서로 다른 모델을 사용하고 싶을 수 있습니다. 예를 들어 더 작고 빠른 모델을 분류에 사용하고, 더 크고 성능이 뛰어난 모델을 복잡한 작업에 사용할 수 있습니다. [`Agent`][agents.Agent]를 구성할 때 다음 중 하나로 특정 모델을 선택할 수 있습니다:
|
||||
|
||||
1. 모델 이름 전달
|
||||
2. 모델 이름 + 해당 이름을 Model 인스턴스로 매핑할 수 있는 [`ModelProvider`][agents.models.interface.ModelProvider] 전달
|
||||
3. [`Model`][agents.models.interface.Model] 구현을 직접 제공
|
||||
2. 해당 이름을 Model 인스턴스로 매핑할 수 있는 임의의 모델 이름 + [`ModelProvider`][agents.models.interface.ModelProvider] 전달
|
||||
3. [`Model`][agents.models.interface.Model] 구현 직접 제공
|
||||
|
||||
!!! note
|
||||
|
||||
SDK는 [`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel]과 [`OpenAIChatCompletionsModel`][agents.models.openai_chatcompletions.OpenAIChatCompletionsModel] 형태를 모두 지원하지만, 두 형태는 지원 기능과 도구 집합이 다르므로 워크플로마다 단일 모델 형태를 사용하는 것을 권장합니다. 워크플로에서 모델 형태를 혼합해야 한다면 사용하는 모든 기능이 양쪽에서 모두 가능한지 확인하세요
|
||||
SDK는 [`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel] 및 [`OpenAIChatCompletionsModel`][agents.models.openai_chatcompletions.OpenAIChatCompletionsModel] 형태를 모두 지원하지만, 두 형태가 서로 다른 기능과 도구 집합을 지원하므로 각 워크플로에는 단일 모델 형태를 사용하는 것을 권장합니다. 워크플로에서 모델 형태를 섞어야 한다면 사용하는 모든 기능이 양쪽 모두에서 제공되는지 확인하세요.
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, AsyncOpenAI, OpenAIChatCompletionsModel
|
||||
@@ -279,7 +283,7 @@ triage_agent = Agent(
|
||||
name="Triage agent",
|
||||
instructions="Handoff to the appropriate agent based on the language of the request.",
|
||||
handoffs=[spanish_agent, english_agent],
|
||||
model="gpt-5.4",
|
||||
model="gpt-5.5",
|
||||
)
|
||||
|
||||
async def main():
|
||||
@@ -287,10 +291,10 @@ async def main():
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
1. OpenAI 모델 이름을 직접 설정합니다
|
||||
2. [`Model`][agents.models.interface.Model] 구현을 제공합니다
|
||||
1. OpenAI 모델 이름을 직접 설정합니다.
|
||||
2. [`Model`][agents.models.interface.Model] 구현을 제공합니다.
|
||||
|
||||
에이전트가 사용할 모델을 추가로 구성하려면 temperature 같은 선택적 모델 구성 매개변수를 제공하는 [`ModelSettings`][agents.models.interface.ModelSettings]를 전달할 수 있습니다
|
||||
에이전트에 사용되는 모델을 추가로 구성하려면 [`ModelSettings`][agents.models.interface.ModelSettings]를 전달할 수 있으며, 이는 temperature 같은 선택적 모델 구성 매개변수를 제공합니다.
|
||||
|
||||
```python
|
||||
from agents import Agent, ModelSettings
|
||||
@@ -305,30 +309,32 @@ english_agent = Agent(
|
||||
|
||||
## 고급 OpenAI Responses 설정
|
||||
|
||||
OpenAI Responses 경로에서 더 많은 제어가 필요하면 `ModelSettings`부터 시작하세요
|
||||
OpenAI Responses 경로에서 더 많은 제어가 필요할 때는 `ModelSettings`부터 시작하세요.
|
||||
|
||||
### 일반적인 고급 `ModelSettings` 옵션
|
||||
|
||||
OpenAI Responses API를 사용할 때는 여러 요청 필드에 이미 직접 대응되는 `ModelSettings` 필드가 있으므로 `extra_args`가 필요하지 않습니다
|
||||
OpenAI Responses API를 사용할 때 여러 요청 필드는 이미 직접적인 `ModelSettings` 필드가 있으므로, 해당 필드에 `extra_args`를 사용할 필요가 없습니다.
|
||||
|
||||
- `parallel_tool_calls`: 같은 턴에서 여러 도구 호출 허용 또는 금지
|
||||
- `truncation`: 컨텍스트가 넘칠 때 실패 대신 가장 오래된 대화 항목을 Responses API가 삭제하도록 `"auto"` 설정
|
||||
- `store`: 생성된 응답을 나중에 조회할 수 있도록 서버 측에 저장할지 제어. 응답 ID에 의존하는 후속 워크플로와 `store=False`일 때 로컬 입력으로 폴백이 필요할 수 있는 세션 압축 흐름에 중요합니다
|
||||
- `prompt_cache_retention`: 예를 들어 `"24h"`처럼 캐시된 프롬프트 접두사를 더 오래 유지
|
||||
- `response_include`: `web_search_call.action.sources`, `file_search_call.results`, `reasoning.encrypted_content` 같은 더 풍부한 응답 페이로드 요청
|
||||
- `top_logprobs`: 출력 텍스트에 대한 상위 토큰 로그확률 요청. SDK는 `message.output_text.logprobs`도 자동 추가합니다
|
||||
- `retry`: 모델 호출에 대해 runner 관리 재시도 설정 사용. [Runner 관리 재시도](#runner-managed-retries) 참고
|
||||
- `parallel_tool_calls`: 같은 턴에서 여러 도구 호출을 허용하거나 금지합니다.
|
||||
- `truncation`: 컨텍스트가 넘칠 때 실패하는 대신 Responses API가 가장 오래된 대화 항목을 삭제하도록 `"auto"`를 설정합니다.
|
||||
- `store`: 생성된 응답을 나중에 검색할 수 있도록 서버 측에 저장할지 제어합니다. 이는 응답 ID에 의존하는 후속 워크플로와, `store=False`일 때 로컬 입력으로 폴백해야 할 수 있는 세션 압축 흐름에 중요합니다.
|
||||
- `context_management`: `compact_threshold`가 있는 Responses 압축 등 서버 측 컨텍스트 처리를 구성합니다.
|
||||
- `prompt_cache_retention`: 예를 들어 `"24h"`로 캐시된 프롬프트 접두사를 더 오래 유지합니다.
|
||||
- `response_include`: `web_search_call.action.sources`, `file_search_call.results`, 또는 `reasoning.encrypted_content` 같은 더 풍부한 응답 페이로드를 요청합니다.
|
||||
- `top_logprobs`: 출력 텍스트에 대한 상위 토큰 logprobs를 요청합니다. SDK는 `message.output_text.logprobs`도 자동으로 추가합니다.
|
||||
- `retry`: 모델 호출에 runner 관리 재시도 설정을 선택 적용합니다. [Runner 관리 재시도](#runner-managed-retries)를 참조하세요.
|
||||
|
||||
```python
|
||||
from agents import Agent, ModelSettings
|
||||
|
||||
research_agent = Agent(
|
||||
name="Research agent",
|
||||
model="gpt-5.4",
|
||||
model="gpt-5.5",
|
||||
model_settings=ModelSettings(
|
||||
parallel_tool_calls=False,
|
||||
truncation="auto",
|
||||
store=True,
|
||||
context_management=[{"type": "compaction", "compact_threshold": 200000}],
|
||||
prompt_cache_retention="24h",
|
||||
response_include=["web_search_call.action.sources"],
|
||||
top_logprobs=5,
|
||||
@@ -336,13 +342,15 @@ research_agent = Agent(
|
||||
)
|
||||
```
|
||||
|
||||
`store=False`로 설정하면 Responses API는 해당 응답을 이후 서버 측 조회용으로 유지하지 않습니다. 이는 무상태 또는 zero-data-retention 스타일 흐름에 유용하지만, 응답 ID를 재사용하던 기능은 대신 로컬 관리 상태에 의존해야 함을 의미합니다. 예를 들어 [`OpenAIResponsesCompactionSession`][agents.memory.openai_responses_compaction_session.OpenAIResponsesCompactionSession]은 마지막 응답이 저장되지 않았을 때 기본 `"auto"` 압축 경로를 입력 기반 압축으로 전환합니다. [Sessions 가이드](../sessions/index.md#openai-responses-compaction-sessions)를 참고하세요
|
||||
`store=False`를 설정하면 Responses API는 해당 응답을 나중에 서버 측에서 검색할 수 있도록 보관하지 않습니다. 이는 상태 비저장 또는 무데이터 보존 방식의 흐름에 유용하지만, 응답 ID를 재사용하던 기능은 대신 로컬에서 관리되는 상태에 의존해야 함을 의미합니다. 예를 들어 [`OpenAIResponsesCompactionSession`][agents.memory.openai_responses_compaction_session.OpenAIResponsesCompactionSession]은 마지막 응답이 저장되지 않은 경우 기본 `"auto"` 압축 경로를 입력 기반 압축으로 전환합니다. [세션 가이드](../sessions/index.md#openai-responses-compaction-sessions)를 참조하세요.
|
||||
|
||||
서버 측 압축은 [`OpenAIResponsesCompactionSession`][agents.memory.openai_responses_compaction_session.OpenAIResponsesCompactionSession]과 다릅니다. `context_management=[{"type": "compaction", "compact_threshold": ...}]`는 각 Responses API 요청과 함께 전송되며, 렌더링된 컨텍스트가 임계값을 넘을 때 API가 응답의 일부로 압축 항목을 내보낼 수 있습니다. `OpenAIResponsesCompactionSession`은 턴 사이에 독립 실행형 `responses.compact` 엔드포인트를 호출하고 로컬 세션 기록을 다시 작성합니다.
|
||||
|
||||
### `extra_args` 전달
|
||||
|
||||
SDK가 아직 최상위에서 직접 노출하지 않는 provider별 또는 최신 요청 필드가 필요할 때 `extra_args`를 사용하세요
|
||||
SDK가 아직 최상위 수준에서 직접 노출하지 않는 제공자별 요청 필드나 더 새로운 요청 필드가 필요할 때 `extra_args`를 사용하세요.
|
||||
|
||||
또한 OpenAI Responses API 사용 시 [몇 가지 추가 선택적 매개변수](https://platform.openai.com/docs/api-reference/responses/create)(예: `user`, `service_tier` 등)가 있습니다. 최상위에 없으면 `extra_args`로 전달할 수 있습니다
|
||||
또한 OpenAI의 Responses API를 사용할 때 [몇 가지 다른 선택적 매개변수](https://platform.openai.com/docs/api-reference/responses/create)(예: `user`, `service_tier` 등)가 있습니다. 최상위 수준에서 사용할 수 없다면 `extra_args`로 전달할 수 있습니다. 동일한 요청 필드를 직접 `ModelSettings` 필드를 통해서도 설정하지 마세요.
|
||||
|
||||
```python
|
||||
from agents import Agent, ModelSettings
|
||||
@@ -360,14 +368,14 @@ english_agent = Agent(
|
||||
|
||||
## Runner 관리 재시도
|
||||
|
||||
재시도는 런타임 전용이며 옵트인입니다. SDK는 `ModelSettings(retry=...)`를 설정하고 재시도 정책이 재시도를 선택한 경우를 제외하면 일반 모델 요청을 재시도하지 않습니다
|
||||
재시도는 런타임 전용이며 명시적으로 선택해야 합니다. `ModelSettings(retry=...)`를 설정하고 재시도 정책이 재시도를 선택하지 않는 한, SDK는 일반 모델 요청을 재시도하지 않습니다.
|
||||
|
||||
```python
|
||||
from agents import Agent, ModelRetrySettings, ModelSettings, retry_policies
|
||||
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
model="gpt-5.4",
|
||||
model="gpt-5.5",
|
||||
model_settings=ModelSettings(
|
||||
retry=ModelRetrySettings(
|
||||
max_retries=4,
|
||||
@@ -388,85 +396,120 @@ agent = Agent(
|
||||
)
|
||||
```
|
||||
|
||||
`ModelRetrySettings`에는 세 필드가 있습니다:
|
||||
`ModelRetrySettings`에는 세 가지 필드가 있습니다:
|
||||
|
||||
<div class="field-table" markdown="1">
|
||||
|
||||
| 필드 | 타입 | 참고 |
|
||||
| 필드 | 타입 | 참고 사항 |
|
||||
| --- | --- | --- |
|
||||
| `max_retries` | `int | None` | 초기 요청 이후 허용되는 재시도 횟수 |
|
||||
| `backoff` | `ModelRetryBackoffSettings | dict | None` | 정책이 명시적 지연을 반환하지 않고 재시도할 때의 기본 지연 전략 |
|
||||
| `policy` | `RetryPolicy | None` | 재시도 여부를 결정하는 콜백. 이 필드는 런타임 전용이며 직렬화되지 않습니다 |
|
||||
| `max_retries` | `int | None` | 초기 요청 이후 허용되는 재시도 시도 횟수입니다. |
|
||||
| `backoff` | `ModelRetryBackoffSettings | dict | None` | 정책이 명시적 지연을 반환하지 않고 재시도할 때 사용하는 기본 지연 전략입니다. `backoff.max_delay`는 계산된 이 백오프 지연만 제한합니다. 정책이 반환한 명시적 지연이나 retry-after 힌트는 제한하지 않습니다. |
|
||||
| `policy` | `RetryPolicy | None` | 재시도 여부를 결정하는 콜백입니다. 이 필드는 런타임 전용이며 직렬화되지 않습니다. |
|
||||
|
||||
</div>
|
||||
|
||||
재시도 정책은 [`RetryPolicyContext`][agents.retry.RetryPolicyContext]를 전달받으며 다음을 포함합니다:
|
||||
재시도 정책은 다음을 포함하는 [`RetryPolicyContext`][agents.retry.RetryPolicyContext]를 받습니다:
|
||||
|
||||
- `attempt`, `max_retries`: 시도 횟수 인지 기반 의사결정 가능
|
||||
- `stream`: 스트리밍/비스트리밍 동작 분기 가능
|
||||
- `error`: 원시 검사
|
||||
- `normalized`: `status_code`, `retry_after`, `error_code`, `is_network_error`, `is_timeout`, `is_abort` 같은 정보
|
||||
- `provider_advice`: 하위 모델 어댑터가 재시도 가이드를 제공할 수 있을 때의 정보
|
||||
- `attempt` 및 `max_retries`: 시도 횟수를 고려한 결정을 내릴 수 있습니다.
|
||||
- `stream`: 스트리밍 동작과 비스트리밍 동작을 분기할 수 있습니다.
|
||||
- `error`: 원문 검사에 사용합니다.
|
||||
- `normalized`: `status_code`, `retry_after`, `error_code`, `is_network_error`, `is_timeout`, `is_abort` 같은 정규화된 정보입니다.
|
||||
- `provider_advice`: 기본 모델 어댑터가 재시도 지침을 제공할 수 있는 경우 사용됩니다.
|
||||
|
||||
정책은 다음 중 하나를 반환할 수 있습니다:
|
||||
|
||||
- 단순 재시도 결정을 위한 `True` / `False`
|
||||
- 지연 재정의 또는 진단 사유 첨부가 필요한 경우 [`RetryDecision`][agents.retry.RetryDecision]
|
||||
- 간단한 재시도 결정을 위한 `True` / `False`
|
||||
- 지연을 재정의하거나 진단 사유를 첨부하려는 경우 [`RetryDecision`][agents.retry.RetryDecision]
|
||||
|
||||
SDK는 `retry_policies`에 즉시 사용 가능한 헬퍼를 제공합니다:
|
||||
SDK는 `retry_policies`에 준비된 헬퍼를 내보냅니다:
|
||||
|
||||
| 헬퍼 | 동작 |
|
||||
| --- | --- |
|
||||
| `retry_policies.never()` | 항상 재시도하지 않음 |
|
||||
| `retry_policies.provider_suggested()` | 가능할 때 provider 재시도 권고를 따름 |
|
||||
| `retry_policies.network_error()` | 일시적 전송/타임아웃 실패에 일치 |
|
||||
| `retry_policies.http_status([...])` | 선택한 HTTP 상태 코드에 일치 |
|
||||
| `retry_policies.retry_after()` | retry-after 힌트가 있을 때만 해당 지연으로 재시도 |
|
||||
| `retry_policies.any(...)` | 중첩 정책 중 하나라도 선택하면 재시도 |
|
||||
| `retry_policies.all(...)` | 중첩 정책 모두 선택할 때만 재시도 |
|
||||
| `retry_policies.never()` | 항상 선택 해제합니다. |
|
||||
| `retry_policies.provider_suggested()` | 사용 가능한 경우 제공자의 재시도 조언을 따릅니다. |
|
||||
| `retry_policies.network_error()` | 일시적인 전송 및 타임아웃 실패와 일치합니다. |
|
||||
| `retry_policies.http_status([...])` | 선택한 HTTP 상태 코드와 일치합니다. |
|
||||
| `retry_policies.retry_after()` | retry-after 힌트를 사용할 수 있을 때만 해당 지연을 사용해 재시도합니다. 이 헬퍼는 retry-after 값을 명시적 정책 지연으로 처리하므로 `backoff.max_delay`가 이를 제한하지 않습니다. |
|
||||
| `retry_policies.any(...)` | 중첩된 정책 중 하나라도 선택하면 재시도합니다. |
|
||||
| `retry_policies.all(...)` | 모든 중첩된 정책이 선택할 때만 재시도합니다. |
|
||||
|
||||
정책을 조합할 때 `provider_suggested()`는 provider가 이를 구분할 수 있을 때 provider veto와 replay 안전 승인(replay-safety approvals)을 보존하므로 가장 안전한 첫 구성 요소입니다
|
||||
정책을 조합할 때 `provider_suggested()`는 가장 안전한 첫 번째 구성 요소입니다. 제공자가 이를 구분할 수 있을 때 제공자 거부와 replay-safety 승인을 보존하기 때문입니다.
|
||||
|
||||
##### 안전 경계
|
||||
|
||||
일부 실패는 자동으로 재시도되지 않습니다:
|
||||
|
||||
- Abort 오류
|
||||
- provider 권고가 replay를 안전하지 않다고 표시한 요청
|
||||
- 출력이 이미 시작되어 replay가 안전하지 않게 되는 스트리밍 실행
|
||||
- 중단 오류
|
||||
- 제공자 조언이 replay를 안전하지 않은 것으로 표시한 요청
|
||||
- replay를 안전하지 않게 만들 방식으로 출력이 이미 시작된 후의 스트리밍 실행
|
||||
|
||||
`previous_response_id` 또는 `conversation_id`를 사용하는 상태 기반 후속 요청도 더 보수적으로 처리됩니다. 이런 요청에서는 `network_error()`나 `http_status([500])` 같은 non-provider 조건만으로는 충분하지 않습니다. 재시도 정책에 보통 `retry_policies.provider_suggested()`를 통한 provider의 replay-safe 승인이 포함되어야 합니다
|
||||
`previous_response_id` 또는 `conversation_id`를 사용하는 상태 저장 후속 요청도 더 보수적으로 처리됩니다. 이러한 요청의 경우 `network_error()` 또는 `http_status([500])` 같은 제공자 외부 조건만으로는 충분하지 않습니다. 재시도 정책에는 일반적으로 `retry_policies.provider_suggested()`를 통해 제공자의 replay-safe 승인이 포함되어야 합니다.
|
||||
|
||||
##### Runner와 에이전트 병합 동작
|
||||
##### Runner 및 에이전트 병합 동작
|
||||
|
||||
`retry`는 runner 수준과 에이전트 수준 `ModelSettings` 사이에서 깊은 병합(deep-merge)됩니다:
|
||||
`retry`는 runner 수준과 에이전트 수준 `ModelSettings` 사이에서 깊게 병합됩니다:
|
||||
|
||||
- 에이전트는 `retry.max_retries`만 재정의하고 runner의 `policy`는 상속할 수 있습니다
|
||||
- 에이전트는 `retry.backoff`의 일부만 재정의하고 runner의 형제 backoff 필드는 유지할 수 있습니다
|
||||
- `policy`는 런타임 전용이므로 직렬화된 `ModelSettings`에는 `max_retries`와 `backoff`는 남고 콜백 자체는 제외됩니다
|
||||
- 에이전트가 `retry.max_retries`만 재정의해도 runner의 `policy`를 계속 상속할 수 있습니다.
|
||||
- 에이전트가 `retry.backoff`의 일부만 재정의하고 runner의 형제 backoff 필드를 유지할 수 있습니다.
|
||||
- `policy`는 런타임 전용이므로 직렬화된 `ModelSettings`는 `max_retries`와 `backoff`를 유지하지만 콜백 자체는 생략합니다.
|
||||
|
||||
더 자세한 예제는 [`examples/basic/retry.py`](https://github.com/openai/openai-agents-python/tree/main/examples/basic/retry.py) 및 [어댑터 기반 재시도 예제](https://github.com/openai/openai-agents-python/tree/main/examples/basic/retry_litellm.py)를 참고하세요
|
||||
더 완전한 예제는 [`examples/basic/retry.py`](https://github.com/openai/openai-agents-python/tree/main/examples/basic/retry.py) 및 [어댑터 기반 재시도 예제](https://github.com/openai/openai-agents-python/tree/main/examples/basic/retry_litellm.py)를 참조하세요.
|
||||
|
||||
## OpenAI가 아닌 provider 문제 해결
|
||||
## non-OpenAI 제공자 문제 해결
|
||||
|
||||
### 트레이싱 클라이언트 오류 401
|
||||
|
||||
트레이싱 관련 오류가 발생하면, trace가 OpenAI 서버로 업로드되는데 OpenAI API 키가 없기 때문입니다. 해결 방법은 세 가지입니다:
|
||||
트레이싱 관련 오류가 발생한다면 이는 트레이스가 OpenAI 서버로 업로드되는데 OpenAI API 키가 없기 때문입니다. 이를 해결할 수 있는 옵션은 세 가지입니다:
|
||||
|
||||
1. 트레이싱 완전 비활성화: [`set_tracing_disabled(True)`][agents.set_tracing_disabled]
|
||||
2. 트레이싱용 OpenAI 키 설정: [`set_tracing_export_api_key(...)`][agents.set_tracing_export_api_key]. 이 API 키는 trace 업로드에만 사용되며 [platform.openai.com](https://platform.openai.com/) 발급 키여야 합니다
|
||||
3. OpenAI가 아닌 트레이스 프로세서 사용. [트레이싱 문서](../tracing.md#custom-tracing-processors) 참고
|
||||
1. 트레이싱을 완전히 비활성화합니다: [`set_tracing_disabled(True)`][agents.set_tracing_disabled]
|
||||
2. 트레이싱용 OpenAI 키를 설정합니다: [`set_tracing_export_api_key(...)`][agents.set_tracing_export_api_key]. 이 API 키는 트레이스 업로드에만 사용되며 [platform.openai.com](https://platform.openai.com/)에서 발급된 것이어야 합니다.
|
||||
3. non-OpenAI 트레이스 프로세서를 사용합니다. [트레이싱 문서](../tracing.md#custom-tracing-processors)를 참조하세요.
|
||||
|
||||
### Responses API 지원
|
||||
|
||||
SDK는 기본적으로 Responses API를 사용하지만, 다른 많은 LLM provider는 아직 이를 지원하지 않습니다. 그 결과 404 또는 유사한 문제가 발생할 수 있습니다. 해결하려면 두 가지 옵션이 있습니다:
|
||||
SDK는 기본적으로 Responses API를 사용하지만, 여전히 많은 다른 LLM 제공자는 이를 지원하지 않습니다. 그 결과 404 또는 유사한 문제가 나타날 수 있습니다. 해결하려면 두 가지 옵션이 있습니다:
|
||||
|
||||
1. [`set_default_openai_api("chat_completions")`][agents.set_default_openai_api] 호출. 환경 변수로 `OPENAI_API_KEY`와 `OPENAI_BASE_URL`을 설정하는 경우 동작합니다
|
||||
2. [`OpenAIChatCompletionsModel`][agents.models.openai_chatcompletions.OpenAIChatCompletionsModel] 사용. 예제는 [여기](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/)에 있습니다
|
||||
1. [`set_default_openai_api("chat_completions")`][agents.set_default_openai_api]를 호출합니다. 이는 환경 변수를 통해 `OPENAI_API_KEY`와 `OPENAI_BASE_URL`을 설정하는 경우에 동작합니다.
|
||||
2. [`OpenAIChatCompletionsModel`][agents.models.openai_chatcompletions.OpenAIChatCompletionsModel]을 사용합니다. 예제는 [여기](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/)에 있습니다.
|
||||
|
||||
### structured outputs 지원
|
||||
### Chat Completions 호환성 옵션
|
||||
|
||||
일부 모델 provider는 [structured outputs](https://platform.openai.com/docs/guides/structured-outputs)를 지원하지 않습니다. 이 경우 아래와 같은 오류가 발생할 수 있습니다:
|
||||
Chat Completions를 통해 라우팅할 때, SDK는 Chat Completions가 보낼 수 없는 Responses 전용 필드(예: `previous_response_id`, `conversation_id`, 프롬프트, 텍스트 전용이 아닌 도구 출력)를 조용히 삭제하여 호환성을 유지합니다. 개발 중 이런 불일치를 빠르게 실패하게 만들고 싶다면 OpenAI 제공자에서 엄격한 기능 검증을 활성화하세요:
|
||||
|
||||
```python
|
||||
from agents import Agent, OpenAIProvider, RunConfig, Runner
|
||||
|
||||
provider = OpenAIProvider(
|
||||
use_responses=False,
|
||||
strict_feature_validation=True,
|
||||
)
|
||||
|
||||
agent = Agent(name="Assistant")
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Hello",
|
||||
run_config=RunConfig(model_provider=provider),
|
||||
)
|
||||
```
|
||||
|
||||
[`MultiProvider`][agents.MultiProvider]를 사용하는 경우 대신 `openai_strict_feature_validation=True`를 전달하세요.
|
||||
|
||||
일부 OpenAI 호환 Chat Completions 제공자는 증분 SDK 처리에 충분히 신뢰할 수 없는 청크로 tool-call delta를 스트리밍합니다. 이 경우 SDK가 제공자 스트림이 끝난 후에만 도구 호출을 내보내도록 스트리밍 tool-call 버퍼링을 활성화하세요:
|
||||
|
||||
```python
|
||||
from agents import OpenAIProvider
|
||||
|
||||
provider = OpenAIProvider(
|
||||
use_responses=False,
|
||||
buffer_streamed_tool_calls=True,
|
||||
)
|
||||
```
|
||||
|
||||
[`MultiProvider`][agents.MultiProvider]의 경우 `openai_buffer_streamed_tool_calls=True`를 사용하세요.
|
||||
|
||||
### Structured outputs 지원
|
||||
|
||||
일부 모델 제공자는 [structured outputs](https://platform.openai.com/docs/guides/structured-outputs)를 지원하지 않습니다. 이로 인해 때때로 다음과 유사한 오류가 발생합니다:
|
||||
|
||||
```
|
||||
|
||||
@@ -474,34 +517,34 @@ BadRequestError: Error code: 400 - {'error': {'message': "'response_format.type'
|
||||
|
||||
```
|
||||
|
||||
이는 일부 모델 provider의 한계입니다 - JSON 출력은 지원하지만 출력에 사용할 `json_schema` 지정은 허용하지 않습니다. 현재 수정 작업 중이지만, 그렇지 않으면 잘못된 JSON으로 앱이 자주 깨질 수 있으므로 JSON schema 출력을 지원하는 provider 사용을 권장합니다
|
||||
이는 일부 모델 제공자의 한계입니다. JSON 출력은 지원하지만 출력에 사용할 `json_schema`를 지정하도록 허용하지는 않습니다. 이 문제를 수정하기 위해 작업 중이지만, JSON 스키마 출력을 지원하는 제공자에 의존하는 것을 권장합니다. 그렇지 않으면 잘못된 형식의 JSON 때문에 앱이 자주 중단될 수 있습니다.
|
||||
|
||||
## provider 간 모델 혼합
|
||||
## 제공자 간 모델 혼합
|
||||
|
||||
모델 provider 간 기능 차이를 인지해야 하며, 그렇지 않으면 오류가 발생할 수 있습니다. 예를 들어 OpenAI는 structured outputs, 멀티모달 입력, 호스티드 파일 검색 및 웹 검색을 지원하지만 다른 많은 provider는 이를 지원하지 않습니다. 다음 제한을 유의하세요:
|
||||
모델 제공자 간 기능 차이를 알고 있어야 하며, 그렇지 않으면 오류가 발생할 수 있습니다. 예를 들어 OpenAI는 structured outputs, 멀티모달 입력, 호스티드 파일 검색 및 웹 검색을 지원하지만 다른 많은 제공자는 이러한 기능을 지원하지 않습니다. 다음 제한 사항에 유의하세요:
|
||||
|
||||
- 지원하지 않는 provider에 지원되지 않는 `tools`를 보내지 마세요
|
||||
- 텍스트 전용 모델 호출 전 멀티모달 입력을 필터링하세요
|
||||
- structured JSON 출력을 지원하지 않는 provider는 가끔 유효하지 않은 JSON을 생성할 수 있음을 유의하세요
|
||||
- 지원되지 않는 `tools`를 이해하지 못하는 제공자에 보내지 마세요
|
||||
- 텍스트 전용 모델을 호출하기 전에 멀티모달 입력을 필터링하세요
|
||||
- 구조화된 JSON 출력을 지원하지 않는 제공자는 때때로 유효하지 않은 JSON을 생성할 수 있습니다.
|
||||
|
||||
## 서드파티 어댑터
|
||||
## 서드 파티 어댑터
|
||||
|
||||
SDK의 내장 provider 통합 지점만으로 부족할 때만 서드파티 어댑터를 사용하세요. 이 SDK로 OpenAI 모델만 사용하는 경우 Any-LLM이나 LiteLLM 대신 내장 [`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel] 경로를 우선하세요. 서드파티 어댑터는 OpenAI 모델과 OpenAI가 아닌 provider를 결합해야 하거나, 내장 경로가 제공하지 않는 어댑터 관리 provider 범위/라우팅이 필요할 때를 위한 것입니다. 어댑터는 SDK와 상위 모델 provider 사이에 또 하나의 호환 계층을 추가하므로 기능 지원과 요청 의미론은 provider마다 다를 수 있습니다. SDK는 현재 Any-LLM과 LiteLLM을 best-effort 베타 어댑터 통합으로 포함합니다
|
||||
SDK의 내장 제공자 통합 지점만으로 충분하지 않을 때만 서드 파티 어댑터를 사용하세요. 이 SDK로 OpenAI 모델만 사용하는 경우 Any-LLM 또는 LiteLLM 대신 내장 [`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel] 경로를 선호하세요. 서드 파티 어댑터는 OpenAI 모델을 non-OpenAI 제공자와 결합해야 하거나, 내장 경로가 제공하지 않는 어댑터 관리형 제공자 커버리지 또는 라우팅이 필요한 경우를 위한 것입니다. 어댑터는 SDK와 업스트림 모델 제공자 사이에 또 다른 호환성 계층을 추가하므로, 기능 지원과 요청 의미 체계는 제공자마다 달라질 수 있습니다. SDK는 현재 최선 노력(best-effort) 방식의 베타 어댑터 통합으로 Any-LLM과 LiteLLM을 포함합니다.
|
||||
|
||||
### Any-LLM
|
||||
|
||||
Any-LLM 지원은 Any-LLM 관리 provider 범위 또는 라우팅이 필요한 경우를 위해 best-effort 베타로 포함됩니다
|
||||
Any-LLM 지원은 Any-LLM 관리형 제공자 커버리지 또는 라우팅이 필요한 경우를 위해 최선 노력 기반의 베타로 포함되어 있습니다.
|
||||
|
||||
상위 provider 경로에 따라 Any-LLM은 Responses API, Chat Completions 호환 API 또는 provider별 호환 계층을 사용할 수 있습니다
|
||||
업스트림 제공자 경로에 따라 Any-LLM은 Responses API, Chat Completions 호환 API 또는 제공자별 호환성 계층을 사용할 수 있습니다.
|
||||
|
||||
Any-LLM이 필요하면 `openai-agents[any-llm]`을 설치한 뒤 [`examples/model_providers/any_llm_auto.py`](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/any_llm_auto.py) 또는 [`examples/model_providers/any_llm_provider.py`](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/any_llm_provider.py)부터 시작하세요. [`MultiProvider`][agents.MultiProvider]와 함께 `any-llm/...` 모델 이름을 사용하거나, `AnyLLMModel`을 직접 인스턴스화하거나, 실행 범위에서 `AnyLLMProvider`를 사용할 수 있습니다. 모델 표면을 명시적으로 고정해야 하면 `AnyLLMModel` 생성 시 `api="responses"` 또는 `api="chat_completions"`를 전달하세요
|
||||
Any-LLM이 필요하다면 `openai-agents[any-llm]`을 설치한 다음 [`examples/model_providers/any_llm_auto.py`](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/any_llm_auto.py) 또는 [`examples/model_providers/any_llm_provider.py`](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/any_llm_provider.py)부터 시작하세요. [`MultiProvider`][agents.MultiProvider]와 함께 `any-llm/...` 모델 이름을 사용하거나, `AnyLLMModel`을 직접 인스턴스화하거나, 실행 범위에서 `AnyLLMProvider`를 사용할 수 있습니다. 모델 API 표면을 명시적으로 고정해야 한다면 `AnyLLMModel`을 생성할 때 `api="responses"` 또는 `api="chat_completions"`를 전달하세요.
|
||||
|
||||
Any-LLM은 서드파티 어댑터 계층이므로 provider 의존성과 기능 격차는 SDK가 아니라 Any-LLM 상위 계층에서 정의됩니다. 사용량 메트릭은 상위 provider가 반환하면 자동 전파되지만, 스트리밍 Chat Completions 백엔드는 사용량 청크를 내보내기 전에 `ModelSettings(include_usage=True)`가 필요할 수 있습니다. structured outputs, 도구 호출, 사용량 보고, Responses 전용 동작에 의존한다면 배포 예정 provider 백엔드를 정확히 검증하세요
|
||||
Any-LLM은 여전히 서드 파티 어댑터 계층이므로, 제공자 의존성과 기능 격차는 SDK가 아니라 Any-LLM에 의해 업스트림에서 정의됩니다. 업스트림 제공자가 사용량 지표를 반환하면 해당 지표는 자동으로 전파되지만, 스트리밍 Chat Completions 백엔드는 사용량 청크를 내보내기 전에 `ModelSettings(include_usage=True)`가 필요할 수 있습니다. structured outputs, 도구 호출, 사용량 보고 또는 Responses별 동작에 의존한다면 배포하려는 정확한 제공자 백엔드를 검증하세요.
|
||||
|
||||
### LiteLLM
|
||||
|
||||
LiteLLM 지원은 LiteLLM 전용 provider 범위 또는 라우팅이 필요한 경우를 위해 best-effort 베타로 포함됩니다
|
||||
LiteLLM 지원은 LiteLLM별 제공자 커버리지 또는 라우팅이 필요한 경우를 위해 최선 노력 기반의 베타로 포함되어 있습니다.
|
||||
|
||||
LiteLLM이 필요하면 `openai-agents[litellm]`을 설치한 뒤 [`examples/model_providers/litellm_auto.py`](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/litellm_auto.py) 또는 [`examples/model_providers/litellm_provider.py`](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/litellm_provider.py)부터 시작하세요. `litellm/...` 모델 이름을 사용하거나 [`LitellmModel`][agents.extensions.models.litellm_model.LitellmModel]을 직접 인스턴스화할 수 있습니다
|
||||
LiteLLM이 필요하다면 `openai-agents[litellm]`을 설치한 다음 [`examples/model_providers/litellm_auto.py`](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/litellm_auto.py) 또는 [`examples/model_providers/litellm_provider.py`](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/litellm_provider.py)부터 시작하세요. `litellm/...` 모델 이름을 사용하거나 [`LitellmModel`][agents.extensions.models.litellm_model.LitellmModel]을 직접 인스턴스화할 수 있습니다.
|
||||
|
||||
일부 LiteLLM 기반 provider는 기본적으로 SDK 사용량 메트릭을 채우지 않습니다. 사용량 보고가 필요하면 `ModelSettings(include_usage=True)`를 전달하고, structured outputs, 도구 호출, 사용량 보고, 어댑터별 라우팅 동작에 의존한다면 배포 예정 provider 백엔드를 정확히 검증하세요
|
||||
일부 LiteLLM 기반 제공자는 기본적으로 SDK 사용량 지표를 채우지 않습니다. 사용량 보고가 필요하다면 `ModelSettings(include_usage=True)`를 전달하고, structured outputs, 도구 호출, 사용량 보고 또는 어댑터별 라우팅 동작에 의존한다면 배포하려는 정확한 제공자 백엔드를 검증하세요.
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user