main
1429 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
35c6b880f7 |
Python: Preserve Mistral prompt-cache usage details (#7597)
* Fix Mistral cached token usage Map prompt cache hits from Mistral chat usage into the standard usage details. Add regression coverage for regular and streaming responses. * Validate Mistral cached token usage * fix(mistral): satisfy strict cached token typing Narrow prompt token details before reading cached_tokens so the Mistral package passes strict Pyright without changing runtime validation.\n\nAddresses https://github.com/microsoft/agent-framework/pull/7597#discussion_r3750712320 |
||
|
|
3a5d00be54 |
Python: add checkpointing support to AgentFrameworkWorkflow.run() in agent-framework-ag-ui (#6646)
* Python: add checkpointing support to AgentFrameworkWorkflow.run() in ag-ui The ag-ui AgentFrameworkWorkflow.run() previously accepted only a RunAgentInput payload and exposed no way to use the core workflow's checkpointing/state-persistence, unlike the core agent-framework workflow implementations. This left ag-ui workflows without resumable execution. Add optional checkpoint_storage and checkpoint_id keyword arguments to run(), threaded through run_workflow_stream() into the core Workflow.run(). This delegates to the existing core capability instead of reinventing it and keeps the public surface consistent with Workflow.run(): - checkpoint_storage enables checkpoint creation at each superstep boundary. - checkpoint_id resumes a run from a persisted checkpoint; incoming messages are forwarded only as request-info responses (never as a new start-executor message) to honor the core's message/checkpoint_id mutual exclusivity, and responses + checkpoint_id performs a restore-then-send in one call. Both can also be supplied via the input_data keys __ag_ui_checkpoint_storage and __ag_ui_checkpoint_id so the FastAPI endpoint (which calls run(input_data) positionally) can opt in without changing its call site; explicit keyword arguments take precedence. Checkpoint resume bypasses the AG-UI thread snapshot hydration early-returns so it always reaches the core restore path. Backward compatible: run(input_data) keeps working unchanged, and the non-checkpoint path still calls run_workflow_stream(input_data, workflow) with its original two-argument convention. Adds focused tests covering checkpoint creation, resume-from-checkpoint, input-data-keyed params, and the unchanged default path. Fixes #6632. * Import Executor from the public agent_framework API in ag-ui workflow test * Fix ag-ui checkpoint resume: preserve thread snapshot, coerce resume responses; fix CI lint/typing A checkpoint-only resume no longer clobbers the stored AG-UI thread snapshot: the snapshot builder is seeded with the prior stored history so the saved snapshot keeps the earlier replayable transcript plus the newly produced output. Resume responses are now coerced against the post-restore pending requests on a checkpoint restore, so a JSON function_approval_response resumes through AG-UI after a cold restore instead of failing with a response-type mismatch. Also update the test-double workflow run() overrides to match the new keyword-only parent signature and re-sort the workflow test imports so ruff and the typing checkers pass. * Coerce ag-ui resume responses without a second checkpoint restore Reading pending request_info events for resume-response coercion previously restored the checkpoint into the live workflow, which invoked every executor's on_checkpoint_restore hook. workflow.run(checkpoint_id=...) then restored again, running those hooks a second time. Custom restore hooks are not required to be idempotent, so this could duplicate restoration work or break workflows that expect exactly one restore per resume. Load the persisted WorkflowCheckpoint directly from storage (runtime override or the workflow's build-time context storage) and read its pending_request_info_events instead. This exposes the same post-restore pending set for the resume contract and response coercion without mutating workflow state or running any restore hook, leaving workflow.run(checkpoint_id=...) as the single restore per resume. Add a regression test asserting on_checkpoint_restore runs exactly once on a checkpointed ag-ui resume. * Python: rework AG-UI workflow checkpointing onto public configuration surfaces Checkpoint storage is now configured on AgentFrameworkWorkflow (or the FastAPI endpoint) instead of being smuggled through input_data keys, and a run resumes by supplying its checkpoint id in the AG-UI forwarded props. With storage always in hand, resume-response coercion reads the pending request set straight from the persisted checkpoint via the public CheckpointStorage.load(), replacing the private runner-context fallback, and the core run call forwards checkpoint arguments directly, relying on core validation for conflicting parameters. Requesting a resume without configured storage now fails with a clear error. * Assign endpoint checkpoint storage in a single place The raw-workflow branch assigned checkpoint_storage at construction and the wiring block assigned it again. Construct the wrapper bare and let the wiring block own the assignment; the existing-storage guard keeps allowing a pre-wrapped runner without storage to adopt the endpoint's. --------- Co-authored-by: Evan Mattson <35585003+moonbox3@users.noreply.github.com> Co-authored-by: Evan Mattson <evan.mattson@microsoft.com> |
||
|
|
27d82b1567 |
Python: Ignore non-project workspace glob matches (#7509)
* Python: Ignore non-project workspace glob matches * test: collect workspace script tests * test: remove standalone script test --------- Co-authored-by: Luis Rodriguez <25299418+luisangelrod@users.noreply.github.com> Co-authored-by: Luis Rodriguez <luis.rodriguez@bcpos.com> |
||
|
|
5e52c6a718 |
Python: Fix ClaudeAgent reusing one SDK client across distinct fresh sessions (#7404)
* Python: Fix ClaudeAgent reusing one SDK client across distinct fresh sessions RawClaudeAgent kept a single mutable ClaudeSDKClient on the agent instance and reused it across distinct fresh AgentSession objects, because a fresh session passes session_id=None and the old reuse check treated that as "keep the current client". Two independent fresh sessions on one shared agent instance therefore shared a single provider conversation, so the second session continued the first session's conversation. Treat a fresh (None) continuation id as always requiring a new client, so an unbound session never inherits an existing provider conversation. Legitimate continuity is preserved: once a session runs, its service_session_id is written back, so later runs pass a real id and resume correctly. Guard client selection/creation with an asyncio.Lock so concurrent runs cannot race between the check and the client assignment. Add regression tests asserting two fresh sessions produce two clients and that an explicit continuation id still resumes the existing client. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 598a9fe1-28c5-4db1-88fd-e14acd9340af * Python: Bind Claude SDK client ownership to each run Replace the single mutable ClaudeSDKClient stored on the agent with a per-run client. Because a ClaudeSDKClient represents exactly one provider conversation, sharing one across distinct sessions collapsed them onto the same conversation and, for concurrent runs, let a fresh session disconnect a client another run was still streaming from. _acquire_client now returns a per-run client (owned) that resumes the framework session's provider conversation when one exists, and _get_stream releases it in a finally once the run completes. An injected client is reused verbatim and left to the caller. The streaming loop moves into _stream_run so the client is a local per-run value rather than shared agent state, which keeps distinct sessions isolated even under concurrency. Continuity is preserved: a session's service_session_id is written back after each run and forwarded as the resume id on subsequent runs. Replace the client-lifecycle tests with per-run ownership and end-to-end isolation tests (two fresh sessions get two separate clients, each disconnected). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 598a9fe1-28c5-4db1-88fd-e14acd9340af * Python: Close remaining Claude session-isolation gaps Address three shared-state gaps in the Claude adapter surfaced in review: - Run-scope structured output: carry the run's structured_output through a per-run state holder and a per-run finalizer instead of storing it on the agent, so a concurrent run cannot overwrite another run's value before its finalizer reads it. - Bind an injected client to one session: an injected ClaudeSDKClient is a single Claude conversation, so bind it to the first session that uses it and raise AgentInvalidRequestException if a different session tries to reuse it. A no-session run reuses the bound session so multi-turn continuity still works; multi-session callers must omit client= or use one agent per session. - Serialize the injected-client path with an asyncio.Lock so concurrent runs cannot race its connect or interleave queries on the one shared client. Owned per-run clients stay lock-free. Update and extend the tests accordingly. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 598a9fe1-28c5-4db1-88fd-e14acd9340af * Python: Bind injected Claude client on provider conversation identity Compare an injected client's binding on the session's service_session_id (the Claude conversation identity) rather than the framework-local session_id, falling back to session_id only when the incoming session has no provider id yet. A reconstructed session from get_session(service_session_id=...) carries a fresh session_id but the same provider conversation, so it now continues the bound conversation instead of raising. Sessions targeting a different conversation are still rejected. Add regression tests for reconstructed-same-conversation continuation and different-conversation rejection. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 598a9fe1-28c5-4db1-88fd-e14acd9340af --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 598a9fe1-28c5-4db1-88fd-e14acd9340af |
||
|
|
30996433ac |
Python: Restore Gemini thought_signature on approval replays (#7546)
Gemini 3.x rejects a request whose functionCall parts lack a thought_signature. The signature was carried as base64 protected_data on a text_reasoning content and re-attached by adjacency, which requires the carrier to immediately precede its call. An approval round trip replays the call with no carrier at all, so the next turn failed with a 400. Track signatures in a bounded per-client call_id map populated at parse time from the resolved call_id, and backfill only when the emitted part has no signature. Also stop clearing the held signature on contents that emit no Part, so an approval response or an unsigned thought summary between the carrier and its call no longer drops it. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: dd0909cd-c7c3-42cb-aef1-1e9a3e64d917 |
||
|
|
e85b3c8ba8 |
Python: Fix FHA session ID translation (#7608)
* Fix FHA session ID traslation * Fix tests * Address comments and fix tests * Fix typing * Show how to use user created sessions * Update README |
||
|
|
db979b616a |
Python: Improve Json parsing for declarative workflow (#7550)
* Json parsing improvement * Fix PR comments * Address PR comments. |
||
|
|
d0a4165f17 |
[BREAKING] Python: Migrate FHA to responses==2.0.0b1 and add Foundry state store (#7533)
* Migrate FHA to responses==2.0.0b1 and add Foundry state store * Fix session id error * Fix tests * Improve tests * Fix copilot comments * Address comments * Revert sample changes * Address comments * Add ContextScopedStoreProvider * Fix type check * Fix type check * Export ContextScopedStoreProvider |
||
|
|
4357ff5742 |
Bump postcss (#7529)
Bumps [postcss](https://github.com/postcss/postcss) from 8.5.22 to 8.5.25. - [Release notes](https://github.com/postcss/postcss/releases) - [Changelog](https://github.com/postcss/postcss/blob/main/CHANGELOG.md) - [Commits](https://github.com/postcss/postcss/compare/8.5.22...8.5.25) --- updated-dependencies: - dependency-name: postcss dependency-version: 8.5.25 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
221f4b6df1 |
Bump postcss from 8.5.15 to 8.5.25 in /python/packages/devui/frontend (#7493)
Bumps [postcss](https://github.com/postcss/postcss) from 8.5.15 to 8.5.25. - [Release notes](https://github.com/postcss/postcss/releases) - [Changelog](https://github.com/postcss/postcss/blob/main/CHANGELOG.md) - [Commits](https://github.com/postcss/postcss/compare/8.5.15...8.5.25) --- updated-dependencies: - dependency-name: postcss dependency-version: 8.5.25 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
a9f7b2b788 |
Bump pyrefly from 1.1.1 to 1.2.0 in /python (#7541)
Bumps [pyrefly](https://github.com/facebook/pyrefly) from 1.1.1 to 1.2.0. - [Release notes](https://github.com/facebook/pyrefly/releases) - [Commits](https://github.com/facebook/pyrefly/compare/1.1.1...1.2.0) --- updated-dependencies: - dependency-name: pyrefly dependency-version: 1.2.0 dependency-type: direct:development update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
adcc3de654 |
Bump js-yaml from 4.3.0 to 4.3.1 in /python/packages/devui/frontend (#7554)
Bumps [js-yaml](https://github.com/nodeca/js-yaml) from 4.3.0 to 4.3.1. - [Changelog](https://github.com/nodeca/js-yaml/blob/4.3.1/CHANGELOG.md) - [Commits](https://github.com/nodeca/js-yaml/compare/4.3.0...4.3.1) --- updated-dependencies: - dependency-name: js-yaml dependency-version: 4.3.1 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
034f5fa119 |
Bump zuban from 0.9.0 to 0.9.1 in /python (#7545)
Bumps [zuban](https://github.com/zubanls/zubanls-python) from 0.9.0 to 0.9.1. - [Release notes](https://github.com/zubanls/zubanls-python/releases) - [Commits](https://github.com/zubanls/zubanls-python/compare/v0.9.0...v0.9.1) --- updated-dependencies: - dependency-name: zuban dependency-version: 0.9.1 dependency-type: direct:development update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
48e547506b |
Python: Make encrypted reasoning opt-in for Foundry chat (#7536)
* Python: Make Foundry encrypted reasoning opt-in * Python: Opt hosted replay test into encrypted reasoning |
||
|
|
4b1afd9052 |
Python: surface Gemini thought summaries as reasoning content (#7488)
Gemini thought-summary parts (part.thought=True) were dropped in _parse_parts, so reasoning never reached ChatResponse.contents. Emit them as text_reasoning content instead, matching OpenAIResponsesClient. Round-trip is safe: _convert_message_contents never re-emits reasoning text as a Part. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: b8aa906e-1408-40c1-9a45-6deb40dc36f8 |
||
|
|
45c515b8a7 |
Python: fix CopilotStudioAgent LineTooLong on large activities (#7417)
* Python: fix CopilotStudioAgent LineTooLong on large activities Bump microsoft-agents-copilotstudio-client to >=1.2.0,<2 and forward a configurable read_bufsize (default 1 MiB) to the underlying aiohttp ClientSession via ConnectionSettings.client_session_settings. Copilot Studio streams each activity as a single SSE data line, so activities larger than aiohttp's 512 KB per-line limit previously raised aiohttp.http_exceptions.LineTooLong. Adds a client_session_settings parameter to CopilotStudioAgent and unit tests covering the default, override, and partial-settings cases. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 2766dc09-ab5f-4adc-8627-98361d7ccef0 * Python: apply read_bufsize default to supplied CopilotStudio settings Address review feedback on the LineTooLong fix: when a user supplies their own ConnectionSettings but no client, inject the read_bufsize default so activities larger than aiohttp's 512 KB per-line limit still stream. Document configuring read_bufsize on the explicit pre-built-client path in the package and sample READMEs and the explicit-settings sample. Add unit tests covering the supplied-settings path. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 2766dc09-ab5f-4adc-8627-98361d7ccef0 --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 2766dc09-ab5f-4adc-8627-98361d7ccef0 |
||
|
|
b2a2fcbd87 |
Python: Add response/request customization hooks to OpenAIChatCompletionClient (#7028)
* Python: Fix reasoning content parsing in OpenAIChatCompletionClient Fix two issues with reasoning content handling in the Chat Completions client: 1. (#6979) reasoning_details plaintext buried as encrypted data: The client dumped the entire reasoning_details array into Content.protected_data without setting Content.text, causing AG-UI to emit ReasoningEncryptedValueEvent instead of visible ReasoningMessageContentEvent for plaintext reasoning providers (e.g. OpenRouter). Now extracts readable text from reasoning_details entries into Content.text while preserving protected_data for round-trip fidelity. 2. (#6978) Mistral list content causes crash: Mistral reasoning models return content as a list of typed chunks ([{"type": "thinking", ...}, {"type": "text", ...}]) instead of a plain string. _parse_text_from_openai assumed content was always a string, causing a Pydantic ValidationError downstream. Now detects list content and parses thinking chunks as Content.from_text_reasoning and text chunks as Content.from_text. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix pyright strict-mode type errors and handle content-as-string shape - Use cast() for proper type narrowing in _extract_reasoning_text and _parse_chunked_content to satisfy pyright strict mode - Handle {"content": "..."} string shape in _extract_reasoning_text (addresses review comment about missing format coverage) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix mypy errors: cast list content to Any in tests model_construct bypasses Pydantic runtime validation but mypy still checks declared types. Use cast(Any, ...) for the list content args. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address review comments: summary field, reasoning field, and round-trip - Add 'summary' field extraction in _extract_reasoning_text for reasoning.summary entries from OpenRouter - Handle message.reasoning and message.reasoning_content top-level fields (plaintext reasoning without reasoning_details) in both streaming and non-streaming paths - reasoning_details takes priority when both fields are present - Preserve original Mistral chunk list in additional_properties ('_source_content_list') so _prepare_message_for_openai can reconstruct the structured list content for multi-turn reasoning - Add 5 new tests covering all new behaviors Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix ruff used-dummy-variable: rename _skip_structured_siblings Remove leading underscore from _skip_structured_siblings variable since it is accessed (not a dummy variable). Ruff's used-dummy-variable rule flags variables with leading underscores that are read. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix missing newline at end of test file Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address review: type-agnostic chunk round-trip and reasoning field echo-back - Honor the _source_content_list marker regardless of the first emitted content's type by handling it before the type match, so a chunk list beginning with a text chunk still round-trips as one structured message (addresses github-actions review comment on results[0]). - Tag every chunked-content item with a shared _structured_content_group id and skip only exact group siblings during serialization, instead of suppressing all later text/reasoning content. - Record provenance of top-level reasoning/reasoning_content fields in _reasoning_source_field and echo the value back under the same key on the next request, which providers such as vLLM require (addresses Kimahriman review comment). Replaces the prior behavior that replayed surfaced reasoning as visible answer text. - Factor the duplicated reasoning parsing into _parse_reasoning_content. - Add tests for provenance capture, reasoning/reasoning_content round-trip, reasoning-only messages, text-first chunk round-trip, and unrelated sibling preservation. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 2f3c0308-51bf-4b66-8b53-87a8546743f5 * Replace provider-specific reasoning logic with configurable parse/prepare hooks Following review feedback (#7028), keep OpenAIChatCompletionClient free of provider-specific quirks for 'almost OpenAI-compatible' endpoints. Instead of branching in core for OpenRouter/vLLM/Mistral, expose two optional callables so callers adapt the client themselves: - response_parser (OpenAIChatResponseContentsParser): post-processes the Content list parsed from each response choice/streaming delta, to surface non-standard fields (e.g. reasoning/reasoning_content/reasoning_details) for display. - message_preparer (OpenAIChatMessagePreparer): post-processes the outgoing request message dicts built from each framework Message, to echo provider-specific fields back on later turns (e.g. vLLM reasoning) for multi-turn continuity. Both default to None (no-op; byte-identical stock OpenAI behavior). This reverts the provider-specific reasoning/chunked-content parsing and round-trip markers previously added to core; Mistral chunked content is now handled by agent-framework-mistral. - Add the two callables to RawOpenAIChatCompletionClient / OpenAIChatCompletionClient constructors and invoke them at the parse and prepare seams. - Export the type aliases from the package and the core lazy openai namespace (+ .pyi). - Replace the removed-behavior tests with tests for the two hooks. - Document the hooks in packages/openai/AGENTS.md. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 2f3c0308-51bf-4b66-8b53-87a8546743f5 * Skip non-string content in default text parsing Structured list content (e.g. Mistral reasoning models returning content as a list of chunks) was wrapped verbatim into a text Content, producing a malformed Content whose text is a list that crashes downstream (issue #6978). Default text parsing now skips non-string content so a configured response_parser receives a clean slate to expand it. Applies to both streaming and non-streaming paths. Add tests for the skip and for a response_parser expanding chunked content. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 2f3c0308-51bf-4b66-8b53-87a8546743f5 * Address review: hook signature, per-role preparer, robust round-trip - response_parser now receives the already-selected ChatCompletionMessage / ChoiceDelta instead of Choice | ChunkChoice, so callers no longer duplicate the streaming dispatch (removes the Any/hasattr pattern from tests). The client owns the dispatch; parsers read provider fields directly. - message_preparer now runs once per Message for every role: the build logic moved to _build_openai_messages and the hook is applied at a single exit point in _prepare_message_for_openai, so system/developer messages no longer bypass it. - Round-trip example/test now correlates surfaced reasoning via an additional_properties marker on message.contents with bounded, order-aware, one-to-one dict removal, instead of fragile request-string matching. Adds a test proving an answer whose text equals the reasoning text is no longer dropped. - Update packages/openai/AGENTS.md for the new parser signature and guidance. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 2f3c0308-51bf-4b66-8b53-87a8546743f5 --------- Co-authored-by: Copilot <copilot@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 2f3c0308-51bf-4b66-8b53-87a8546743f5 |
||
|
|
7302d0bf23 |
Python: agent-hooks interception contract as a first-class experimental core feature (#7515)
* feat(python): add agent-hooks middleware as experimental core feature Implement the AGENT-HOOKS-0.1 interception contract as a first-class experimental feature in agent_framework core. - Single public factory agent_hooks_middleware() returning a private agent/chat/function middleware trio (one object per middleware category); partial or stacked installs fail closed with loud errors. - All eight interception points: input/output at the agent seam, pre/post_model_call at the chat seam, pre/post_tool_call at the function seam, agent_startup/agent_shutdown bracketing each run. - Fail-closed enforcement throughout: transforms write back into the native contexts (messages, arguments, results) or raise; content is preserved as Content objects; MiddlewareTermination short-circuits are guarded at every seam; enforcement-layer failures halt the run; interceptor crashes surface as host_error denies. - Streaming is fully buffered per spec buffered_output semantics: no update egresses before the post_model_call/output verdicts; a deny at pull time releases zero updates; run state stays active across lazy pulls with cleanup on every exit path. - Session scoping: per-run by default (startup/shutdown bracket each run) or host-owned via emitter/builder parameters for one session spanning multiple runs. - agent-hooks-sdk is an opt-in agent-hooks extra (not in all), lazy-imported per the _mcp.py pattern; core imports cleanly without it and the factory raises a clear ModuleNotFoundError. - ExperimentalFeature.AGENT_HOOKS + @experimental decorator, lazy root export, typing surface, PACKAGE_STATUS.md entry. - 55 tests built on real Agent/mock-client flows covering deny-before- execution, transform write-back, rich-content preservation, complete streaming ordering, error cleanup, concurrency isolation, nested agents, and importability without the optional SDK. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> * style(python): unquote ResponseStream annotation per pyupgrade The pre-commit pyupgrade hook rewrites the quoted forward reference; ResponseStream is imported at runtime in this module, so the quotes were unnecessary. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> * refactor(python): address agent-hooks review feedback Reworks the agent-hooks feature per PR review: - Verdicts now precede durability: a run-scoped persistence gate (_sessions.py) defers per-service-call history persistence and after-run provider work until the covering post_model_call/output verdict permits; denied content never persists, transforms persist post-write-back. Unhooked runs are unchanged (verified against an instrumented baseline). - ResponseStream.buffered_and_gated: a buffered-gate combinator that applies the run's pending stream hooks before the gate, then seals the stream, so no middleware can rewrite egress after the output verdict. Replaces the hand-rolled replay iterator. - MiddlewareBundle (public, _middleware.py): the factory returns an indivisible bundle categorize_middleware splits, making partial installs impossible by construction; members are validated at construction. Bare (non-sequence) middleware at agent construction is now normalized instead of silently dropped, and unrecognized middleware logs a warning instead of vanishing. - Factory split and rename: create_agent_hooks_middleware (per-run sessions) and create_agent_hooks_middleware_from_emitter (host-owned); the sentinel parameter-diffing is gone. - Wire conversions live in per-point codec classes owning to_wire and write_back. Fixes in that code: tool-call name transforms apply or raise; non-object args transforms raise; argument write-back merges only changed keys (original values, including bytes, preserved by identity); message-list write-back matches by identity, not index. - function_approval_request objects on the normal return path pass through un-emitted, preserving the human approval pause. - Hosted (service-executed) tool calls surface in the post_model_call content projection; the tool-seam limitation is documented. - Import probe covers the full SDK surface and re-raises as missing-extra only for the agent_hooks module; module logger added; _json_safe replaced by make_json_safe (which gained bytes support); tools_registered uses normalize_tools; dependency-pyright analyzes the module again via the test dependency-group. - Tests: 75 in the feature suite (persistence gating, stream-hook sealing, approval passthrough, codec units, bundle validation, bare-bundle installs), full core suite green. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> * refactor(python): second review round for agent-hooks Addresses the second review round on the agent-hooks feature: - Nested-run persistence ownership: RawAgent.run stamps a run identity over the run's dynamic extent (including streaming pulls and result hooks); the persistence gate binds to its owning run via an offer/adopt handshake keyed to the agent instance and accepts only its owner's persists — nested runs persist inline regardless of how they were started (tool calls, middleware, custom run loops). The tool-seam suspension remains for custom-loop sub-agents invoked as tools; the one residual case (custom loop nested in a custom loop off the tool path) is fail-closed and documented. Fixes a latent pre-existing re-deferral: flush() now drains with the gate context suspended, so a nested hooked run's permitted after-run persistence no longer re-defers into an enclosing gate. - as_tool stream_callback consumes the released (verdicted) stream; observers cannot see denied or pre-transform content. Both directions are regression-tested. - categorize_middleware gained supported_categories: a bundle member landing in a category a call site cannot install raises; bare middleware warns like _add_middleware. Wired at the chat-client sites and the provider seam. - ResponseStream.buffered_and_gated owns the re-derivation rule via a rederive callable (gates cannot choose released updates) and is marked experimental. - Wire codecs compare with bool-aware equality (Python == equates 1 == True, which made bool/number transforms look untouched and get dropped) and _ToolResultCodec.write_back owns the untouched-wire rule via the before value. - middleware parameters accept a bare middleware or bundle everywhere the runtime does (constructors, run overloads, as_agent, telemetry and harness layers, foundry); the bare-source rule has a single owner in categorize_middleware; bare middleware assigned to the attribute now executes (documented behavior change). - MiddlewareBundle is experimental and validates members; approval passthrough, typing-check fixes (ty ignores mypy-coded ignore comments), logging, and documentation updates per review. Test count: 85 feature tests plus 12 new this round across sessions, middleware, agents; full core suite green; typing checked under mypy, pyrefly, ty, zuban, and pyright. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> * docs(python): drop previous-behavior notes from middleware docstrings Per review: docstrings describe current behavior only. The bare-middleware behavior change stays recorded in the PR description and commit history. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> * fix(python): gate ownership survives retrying middleware A retry or fallback middleware issuing a second call_next() gave the new attempt a fresh run identity that the persistence gate's first-bind-wins ownership rejected, so the retried attempt's history persisted inline before the output verdict — a denied response became durable again. The gate now accumulates every identity adopted through its own offer ticket: all attempts' persistence stays behind the one final verdict (deny drops all of it, allow flushes all of it). Accumulation over rebind-replace is deliberate: rebinding would flip an earlier attempt's still-running background work from deferred to inline, which is the fail-open direction. A foreign agent still cannot bind: tickets are minted only by the covered pipeline's final handler and adoption is instance-keyed. Also consolidates the bare-middleware-source rule into a single _as_middleware_list owner used by every interpretation site (the harness merge, BaseAgent.__init__, categorize_middleware, both client-kwargs merges, get_response, SessionContext.extend_middleware), including the str/bytes exclusion the stray copies missed. The constructor now stores a copy of the caller's sequence; assign to the middleware attribute for post-construction changes. Retry regression tests cover denied and allowed retried runs in both stream modes and fail with first-bind-wins restored. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> * fix(python): streaming seam runs pipeline descent inside the gate The streaming agent seam ran call_next() outside the persistence gate (only _consume entered it later), so a retry middleware that drained a successful attempt with get_final_response() and discarded it persisted that attempt's exchange before any verdict existed; a later deny dropped only the retry attempt's deferred work. The descent is now wrapped in the gate exactly like the non-streaming seam: attempt identities adopted during descent are accepted owners, so in-pipeline draining defers, deny drops every attempt, and a middleware that raises after draining strands the pending persists unexecuted. The bind_owner docstring now states the actual soundness invariant covering both bind sites: every bind comes from a run inside the covered pipeline. New tests cover drained-and-discarded attempts (deny and allow, both stream modes) and a sub-agent tool inside a drained attempt; the streaming deny variant fails with the gate wrap reverted. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> * fix(python): flush deferred persistence on streaming no-result termination With the pipeline descent now running inside the persistence gate, a middleware that drains a successful attempt and then terminates without a result left that attempt's deferred persistence stranded: the streaming no-result termination path raised before any flush, so history of exchanges that really happened and passed their own verdicts quietly vanished (streaming only; non-streaming already flushes before its re-raise). The path now flushes before re-raising the termination, with a state.halted guard first so an enforcement failure during the drained attempt still strands pending fail-closed and surfaces the halt, mirroring the non-streaming ordering exactly. The regression test covers both seams; the streaming variant fails without the fix. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> --------- Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> |
||
|
|
422160eabe |
Python: Add windows junction detection for skills (#7507)
* Add windows junction detection for skills * Address PR comment |
||
|
|
5a1d96df67 |
Python: Separate mem0 storage and search scopes (#7531)
* Separate mem0 storage and search scopes * Apply suggestions from code review Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> |
||
|
|
594954700a |
Python: Fix AG-UI conversation correlation across runs (#7430)
* Add single agent AGUI sample * Fix AG-UI conversation correlation across runs * Address PR review and code quality feedback * Correlate AG-UI chat spans across runs --------- Co-authored-by: Tao Chen <taochen@microsoft.com> |
||
|
|
5f3ca8f93c |
Python: Fix AG-UI approval resume at the protocol boundary (#7480)
* Python: Fix Ollama approval resume message handling * Python: Reject empty Ollama approval resume payload * Python: Keep AG-UI approval controls out of provider input * Python: Do not trust pending AG-UI tool results |
||
|
|
4d3c7844d6 |
Python: Bound tool result compaction summaries (#7396)
* Python: bound tool result compaction summaries Keep ToolResultCompactionStrategy from re-inserting oversized tool result payloads through the synthetic summary message by bounding the generated digest text. Add regression coverage proving a large tool result is not embedded verbatim, keeps a bounded prefix, and marks truncation. * Python: keep excluded tool results out of compaction digests Build ToolResultCompactionStrategy digest content from messages still included in the group so a summary cannot restore payloads that an earlier compaction already excluded. Use the strategy cap constant in the large-payload regression and add coverage for already-excluded tool results. Validation: uv run pytest packages/core/tests/core/test_compaction.py -q -k 'tool_result_compaction'; uv run ruff check packages/core/agent_framework/_compaction.py packages/core/tests/core/test_compaction.py; uv run ruff format --check packages/core/agent_framework/_compaction.py packages/core/tests/core/test_compaction.py; uv run poe test -P core; uv run poe build -P core; env HOME=/tmp/sds-home XDG_CACHE_HOME=/tmp/sds-cache uv run poe test -A. * Python: align compaction digest review cleanup Align ToolResultCompactionStrategy's included-message filter with the module's existing EXCLUDED_KEY boolean semantics. Make the large-payload regression size scale from _SUMMARY_MAX_CHARS so it continues to exercise truncation if the digest cap changes. Validation: uv run pytest packages/core/tests/core/test_compaction.py -q -k 'tool_result_compaction'; uv run ruff check packages/core/agent_framework/_compaction.py packages/core/tests/core/test_compaction.py; uv run ruff format --check packages/core/agent_framework/_compaction.py packages/core/tests/core/test_compaction.py; uv run poe test -P core; uv run poe build -P core. * Python: collapse tool result digest scan --------- Co-authored-by: Eduard van Valkenburg <eavanvalkenburg@users.noreply.github.com> |
||
|
|
07511b80c9 |
Python: Prevent orphaned local approval responses (#7462)
* Python: Prevent orphaned local approval responses Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 0efaca91-a0a7-46f4-9b81-022385607fe4 * Python: Clarify approval serialization boundaries Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 0efaca91-a0a7-46f4-9b81-022385607fe4 --------- Copilot-Session: 0efaca91-a0a7-46f4-9b81-022385607fe4 |
||
|
|
e84b5a07c1 |
Python: fix LocalEvaluator reporting zero-check items as passed (#7399)
LocalEvaluator.evaluate initialized item_passed to True and only ever cleared it inside the loop over check results. With no checks configured the loop never runs, so an item with zero scores was recorded as passed: result_counts reported one pass, all_passed was True, and raise_for_status() did not raise. Initialize item_passed from bool(check_results) so an item with no evaluated checks fails closed. This matches the .NET contract in this repository, where AgentEvaluationResults.ItemPassed ends with 'return result.Metrics.Count > 0' and is pinned by LocalEvaluator_WithZeroChecks_ItemsHaveZeroMetricsAndFailAsync. Add a focused regression covering the counts, all_passed, the empty score list, and raise_for_status(). Update the LocalEvaluator class and evaluate() docstrings, which previously described the pass rule without the zero-check case. Fixes #7397 |
||
|
|
84d5a5eec1 |
Consolidate Dependabot dependency updates (#7445)
* Bump AgentMemory from 1.2.0 to 1.3.0 --- updated-dependencies: - dependency-name: AgentMemory dependency-version: 1.3.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> * .NET: consolidate #7280 AgentMemory.AgentFramework 1.3.0 * Bump github/codeql-action/init from 4.37.0 to 4.37.3 Bumps [github/codeql-action/init](https://github.com/github/codeql-action) from 4.37.0 to 4.37.3. - [Release notes](https://github.com/github/codeql-action/releases) - [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md) - [Commits](https://github.com/github/codeql-action/compare/99df26d4f13ea111d4ec1a7dddef6063f76b97e9...e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81) --- updated-dependencies: - dependency-name: github/codeql-action/init dependency-version: 4.37.3 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> * Bump astral-sh/setup-uv from 8.3.2 to 9.0.0 Bumps [astral-sh/setup-uv](https://github.com/astral-sh/setup-uv) from 8.3.2 to 9.0.0. - [Release notes](https://github.com/astral-sh/setup-uv/releases) - [Commits](https://github.com/astral-sh/setup-uv/compare/11f9893b081a58869d3b5fccaea48c9e9e46f990...c771a70e6277c0a99b617c7a806ffedaca235ff9) --- updated-dependencies: - dependency-name: astral-sh/setup-uv dependency-version: 9.0.0 dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> * Bump github/codeql-action/analyze from 4.37.0 to 4.37.3 Bumps [github/codeql-action/analyze](https://github.com/github/codeql-action) from 4.37.0 to 4.37.3. - [Release notes](https://github.com/github/codeql-action/releases) - [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md) - [Commits](https://github.com/github/codeql-action/compare/99df26d4f13ea111d4ec1a7dddef6063f76b97e9...e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81) --- updated-dependencies: - dependency-name: github/codeql-action/analyze dependency-version: 4.37.3 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> * Bump actions/cache from 5.0.5 to 6.1.0 Bumps [actions/cache](https://github.com/actions/cache) from 5.0.5 to 6.1.0. - [Release notes](https://github.com/actions/cache/releases) - [Changelog](https://github.com/actions/cache/blob/main/RELEASES.md) - [Commits](https://github.com/actions/cache/compare/27d5ce7f107fe9357f9df03efb73ab90386fccae...55cc8345863c7cc4c66a329aec7e433d2d1c52a9) --- updated-dependencies: - dependency-name: actions/cache dependency-version: 6.1.0 dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> * Bump actions/checkout from 6.0.2 to 7.0.1 Bumps [actions/checkout](https://github.com/actions/checkout) from 6.0.2 to 7.0.1. - [Release notes](https://github.com/actions/checkout/releases) - [Changelog](https://github.com/actions/checkout/blob/main/CHANGELOG.md) - [Commits](https://github.com/actions/checkout/compare/de0fac2e4500dabe0009e67214ff5f5447ce83dd...3d3c42e5aac5ba805825da76410c181273ba90b1) --- updated-dependencies: - dependency-name: actions/checkout dependency-version: 7.0.1 dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> * Bump astral-sh/setup-uv in /.github/actions/python-setup Bumps [astral-sh/setup-uv](https://github.com/astral-sh/setup-uv) from 8.3.2 to 9.0.0. - [Release notes](https://github.com/astral-sh/setup-uv/releases) - [Commits](https://github.com/astral-sh/setup-uv/compare/11f9893b081a58869d3b5fccaea48c9e9e46f990...c771a70e6277c0a99b617c7a806ffedaca235ff9) --- updated-dependencies: - dependency-name: astral-sh/setup-uv dependency-version: 9.0.0 dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> * Bump ty from 0.0.60 to 0.0.64 in /python Bumps [ty](https://github.com/astral-sh/ty) from 0.0.60 to 0.0.64. - [Release notes](https://github.com/astral-sh/ty/releases) - [Changelog](https://github.com/astral-sh/ty/blob/main/CHANGELOG.md) - [Commits](https://github.com/astral-sh/ty/compare/0.0.60...0.0.64) --- updated-dependencies: - dependency-name: ty dependency-version: 0.0.65 dependency-type: direct:development update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> * Bump prek from 0.4.10 to 0.4.11 in /python Bumps [prek](https://github.com/j178/prek) from 0.4.10 to 0.4.11. - [Release notes](https://github.com/j178/prek/releases) - [Changelog](https://github.com/j178/prek/blob/master/CHANGELOG.md) - [Commits](https://github.com/j178/prek/compare/v0.4.10...v0.4.11) --- updated-dependencies: - dependency-name: prek dependency-version: 0.4.11 dependency-type: direct:development update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> * Bump uv from 0.11.29 to 0.11.32 in /python Bumps [uv](https://github.com/astral-sh/uv) from 0.11.29 to 0.11.32. - [Release notes](https://github.com/astral-sh/uv/releases) - [Changelog](https://github.com/astral-sh/uv/blob/main/CHANGELOG.md) - [Commits](https://github.com/astral-sh/uv/compare/0.11.29...0.11.32) --- updated-dependencies: - dependency-name: uv dependency-version: 0.12.0 dependency-type: direct:development update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> * Bump ruff from 0.15.22 to 0.16.0 in /python Bumps [ruff](https://github.com/astral-sh/ruff) from 0.15.22 to 0.16.0. - [Release notes](https://github.com/astral-sh/ruff/releases) - [Changelog](https://github.com/astral-sh/ruff/blob/main/CHANGELOG.md) - [Commits](https://github.com/astral-sh/ruff/compare/0.15.22...0.16.0) --- updated-dependencies: - dependency-name: ruff dependency-version: 0.16.0 dependency-type: direct:development update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> * update uv-build requirement in /python --- updated-dependencies: - dependency-name: uv-build dependency-version: 0.12.0 dependency-type: direct:development ... Signed-off-by: dependabot[bot] <support@github.com> * Python: align workspace pins for #7436-#7439 * Python: support ty 0.0.64 diagnostics for #7436 * Python: apply Ruff 0.16 formatting for #7439 * Update workflow action version annotations --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
8d379168b2 |
Python: Improve python sample validation workflow (#7350)
* Add skill to replace hardcoded foundry project endpoint and model * Include more samples and fix migration samples part 1 * Fix migration samples * Replace Foundry hosted agent validation skill * Fix hosted agent file sample * Fix agent result format * Reorganize jobs * Update discovery heuristic for apps * Split agents into even more jobs * Add toolbox endpoint * Add more pre configured resources * Fix using deployed agent sample * Add sample status * Add playbook * Exclude hidden folder in sample discovery * Install autogen dependencies * Grant azure search RBAC role * Increase timeout for magentic * Build search resouce id deterministically * Remove grant in the workflow * Move azure cli login closer to when the sample actually runs * Refactor playbook * Fix using deployed agent sample * Actually save the playbooks * Fix action syntax error * Fix magentic sample * Address copilot comments * Fix link inspection * Address comments * Correct README * Fix playbook path * Remove trailing space |
||
|
|
18997c2fde |
Python: Give the AG-UI Thread Snapshot lifecycle a single owner module (#7479)
* Python: Give the AG-UI Thread Snapshot lifecycle a single owner module Both the agent and workflow runners independently implemented the thread snapshot lifecycle: hydration replay, the load-once stored read, resume message seeding, the stored/request/deferred-default state overlay, and the save whose storage failures must never surface on an already-streamed run. The two copies had already drifted in small ways (one hydrate helper re-checked a store the caller had verified; the two cancelled-resume-id helpers differed on missing-id handling). Introduce ThreadSnapshotSession in _snapshot_session.py as the one owner of that lifecycle, opened once per run and inert when no store or scope is configured so callers stop branching on configuration. Rewire both runners onto it, consolidate _cancelled_resume_interrupt_ids in _run_common (defensive variant) and _event_messages_to_snapshot_dicts in the new module, and delete the superseded per-runner copies. The session interface is covered by dedicated tests; existing suites pin runner behavior. Public exports are unchanged. * Python: Narrow AG-UI event types in snapshot session tests The hydration test accessed run_id, snapshot, and messages on values typed as BaseEvent, which fails the tests/samples type checkers. Narrow each event with isinstance assertions before reading its fields. |
||
|
|
5cc1b8e3c3 |
Python: Add hosted agent sample for the agent harness (#7010)
* Python: Add hosted agent sample for the agent harness * Disable file providers and fix call_server usage in hosted harness sample Addresses PR review: disable the harness file-memory and file-access providers so the headless sample doesn't expose file tools or write outside storage/, and correct the app.py docstring to match call_server.py (which takes no prompt argument). * Python: update hosted harness sample for current APIs --------- Co-authored-by: Evan Mattson <evan.mattson@microsoft.com> |
||
|
|
9ce55cae00 |
Python: Remove dead AG-UI orchestration helpers and flatten subpackage (#7426)
The _orchestration/_helpers module had no production callers; its only importer was its own test file. It also carried a stale fork of the live metadata sanitization in _agent_run.py: the dead copy truncated oversized values, behavior the live copy deliberately replaced with drop-plus-warning because truncation can produce invalid JSON. Move _tooling.py and _predictive_state.py to the package root and remove the now-empty _orchestration subpackage. Public exports are unchanged. |
||
|
|
06c0fc2b10 |
Python: forward Azure AI Search query-source identity (#7278)
* Forward Azure AI Search query-source identity * Address query source credential review feedback |
||
|
|
a74811edec | fix(python): preserve falsey EditTableV2 items (#7380) | ||
|
|
f5dfb1413e |
Python: Add Mistral chat client (#7392)
* feat(python): add Mistral chat client Implements native Mistral support (#7366) with streaming, tool calling, and structured output. Talks to the REST API directly over httpx: the mistralai SDK's pinned OpenTelemetry deps conflict with the workspace. * refactor(python): simplify Mistral client per review Drop the streamed tool-call accumulator and multi-choice parsing in favor of the framework's built-in fragment merging, mark n unsupported, omit unset strict from json_schema, and leave CI secret wiring to maintainers. * test(python): drop n forwarding assertion n is typed as unsupported on MistralChatOptions; the option-mapping test still passed n, failing pyrefly/ty/zuban/mypy in CI. * refactor(python): drop n from MistralChatOptions n is not part of the base ChatOptions, so removing the key rejects it without an explicit None override. * feat(python): mark Mistral feature usage Both clients flip the shared FeatureIndex.MISTRAL bit before each request, matching the feature-usage telemetry other providers emit. * fix(python): key streamed tool calls by index Mistral omits the tool call id on continuation fragments, and the framework only coalesces empty-id fragments into the immediately preceding call, so interleaved parallel calls merged into the wrong call with corrupted arguments. Accumulate fragments per (choice, index) and emit each call only once complete. * fix(python): restore Mistral SDK client injection Dropping the mistralai dependency turned the embedding client's client= parameter into a breaking change for injected SDK clients. Add http_client= for httpx.AsyncClient and keep client= working: httpx goes to the REST path, a duck-typed mistralai.Mistral goes through the legacy SDK path with a DeprecationWarning until the next major release. * chore(python): tidy Mistral sample header |
||
|
|
43309018be |
.NET and Python: Extract Durable Task and Azure Functions integrations (#7465)
* Extract Durable Task and Azure Functions integrations Remove the migrated implementations, samples, tests, documentation, and repository wiring now owned by microsoft/agent-framework-durable-extension. Preserve Python compatibility through the agent_framework.azure shim and agent-framework-core[all], and leave customer-facing redirects to the new repository. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6181dcf9-857b-43ea-9fd2-fcd6b175ffdd * Fix feature registry validation after extraction Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6181dcf9-857b-43ea-9fd2-fcd6b175ffdd * Narrow external feature package paths Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 6181dcf9-857b-43ea-9fd2-fcd6b175ffdd --------- Copilot-Session: 6181dcf9-857b-43ea-9fd2-fcd6b175ffdd |
||
|
|
10fe3c4c72 |
Python: Ignore excluded tool results during compaction (#7391)
* Python: Ignore excluded tool results during compaction * fix: avoid extra compaction message pass |
||
|
|
e39a8a2e79 |
Python: Bump Python package versions for 1.13.0 release (#7443)
CodeQL / Analyze (csharp) (push) Has been cancelled
CodeQL / Analyze (python) (push) Has been cancelled
dotnet-build-and-test / paths-filter (push) Has been cancelled
dotnet-build-and-test / dotnet-build-and-test-check (push) Has been cancelled
dotnet-build-and-test / dotnet-build (Debug, windows-latest, net9.0) (push) Has been cancelled
dotnet-build-and-test / dotnet-build (Release, ubuntu-latest, net10.0) (push) Has been cancelled
dotnet-build-and-test / dotnet-build (Release, ubuntu-latest, net8.0) (push) Has been cancelled
dotnet-build-and-test / dotnet-build (Release, windows-latest, net472) (push) Has been cancelled
dotnet-build-and-test / dotnet-test (Release, integration, true, ubuntu-latest, net10.0) (push) Has been cancelled
dotnet-build-and-test / dotnet-test (Release, integration, true, windows-latest, net472) (push) Has been cancelled
dotnet-build-and-test / dotnet-foundry-hosted-it (push) Has been cancelled
dotnet-build-and-test / dotnet-test-functions (push) Has been cancelled
dotnet-build-and-test / Integration Test Report (push) Has been cancelled
* Bump Python package versions for 1.13.0 release Bump all 37 Python package projects because the CHANGELOG-driven release includes cross-package feature-usage telemetry, with core and root advancing to 1.13.0, OpenAI to 1.12.0, patch bumps for other stable packages, and 260730 stamps for alpha and beta packages. No optional beta cohort bump was applied; every prerelease package changed. Raise core floors conservatively across co-released packages. Copilot-Session: e234a28b-c2fd-4ff4-a51d-3d8917936541 * Align co-released Python package dependencies Update the four hosting adapter pins to the co-released agent-framework-hosting alpha and raise the Azure Functions Durable Task floor to the co-released beta. Copilot-Session: e234a28b-c2fd-4ff4-a51d-3d8917936541 * Minimize Python release lockfile updates Regenerate uv.lock with the pre-commit hook pinned uv version so the release changes only workspace package versions while preserving platform markers and agentlightning 0.3.0. Copilot-Session: e234a28b-c2fd-4ff4-a51d-3d8917936541 --------- Copilot-Session: e234a28b-c2fd-4ff4-a51d-3d8917936541 |
||
|
|
25ec4c3b5c |
Python: Support archive-type MCP skills (source, toolbox, sample) (#7121)
* Python: Support archive-type MCP skills in MCPSkillsSource Add `archive`-type skill support to `MCPSkillsSource` so an MCP server can advertise packaged skills (ZIP / TAR / gzip-compressed TAR) that are downloaded, safely unpacked to a local directory, and served like file-based skills, while keeping the guarantee that MCP-delivered scripts are never executed. - Dispatch `skill://index.json` entries by `type`: `skill-md` (existing, fetched on demand) and `archive` (new). Unknown types are skipped. - `_ArchiveEntryLoader` downloads, extracts, and prunes archive skills and delegates discovery to an internal `FileSkillsSource` created with no script extensions and no runner, so bundled scripts surface as read-only resources only. - Hardened stdlib extraction: path-traversal (zip-slip) guard, non-regular TAR member skipping, and file-count / uncompressed-size / download-size limits. - Configure via `archive_*` constructor kwargs (no options object, per Python conventions); use `CachingSkillsSource` for refresh rather than a source level refresh interval. - Fix `FileSkillsSource` to treat `None` extensions as "use defaults" and an empty tuple as "discover none" (an empty tuple previously fell back to defaults). Port of .NET PR #6631. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9e358a4e-538f-46be-8c58-128b6182352d * Propagate non-not-found archive download errors in MCPSkillsSource Only swallow "resource not found" MCP errors when downloading an archive resource; re-raise every other error (auth failure, INTERNAL_ERROR, connection drop, timeout) so a transient transport failure is not silently turned into a missing skill. This matches the existing failure model used by `_try_read_index` and `MCPSkill.get_resource`, and avoids a failed `CachingSkillsSource` refresh overwriting a previously cached list with a partial result. Add tests asserting archive-download INTERNAL_ERROR and ConnectionError propagate out of `get_skills`. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9e358a4e-538f-46be-8c58-128b6182352d * Python: Expose archive skill options on FoundryToolbox and demo in sample - FoundryToolbox.as_skills_provider() now forwards the MCPSkillsSource archive options (archive_skills_directory, archive_resource_extensions, archive_resource_search_depth, archive_max_file_count, archive_max_size_bytes, archive_max_uncompressed_size_bytes). Only explicitly-set options are forwarded so unset ones keep the MCPSkillsSource defaults. This lets a hosted toolbox agent redirect archive extraction to a writable directory (the default is under the cwd, which may be read-only in a container). - Add unit tests covering default (no options forwarded) and override forwarding. - Update the 12_foundry_toolbox_mcp_skills sample to demonstrate all three progressive-disclosure stages with an archive skill: escalation-policy now ships a references/refund-matrix.md resource and is uploaded as a ZIP archive; main.py disables load_skill and read_skill_resource approval and points archive extraction at a temp directory. README, toolbox.yaml, and ignore files updated accordingly. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 4f14f83d-1868-45c1-be1a-12f49a58ac36 * Python: Fix ty type error in toolbox archive-option test Cast provider._source to _FoundryToolboxSkillsSource before accessing the private _archive_options, so the ty checker (which runs over tests) resolves the concrete type instead of the SkillsSource base. Replaces the mypy-style type: ignore that ty did not honor. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 4f14f83d-1868-45c1-be1a-12f49a58ac36 * Rework archive-type skill support in MCPSkillsSource to unpack archives entirely in memory instead of extracting them to a local directory, and apply reviewer feedback. * Python: Raise on archive member path-traversal (zip-slip) Treat a `..` path-traversal member in an archive skill as a hostile archive and reject the whole skill, matching how the file-count and uncompressed-size limits reject a malformed archive (previously the member was silently skipped while the rest of the skill still loaded). - `_normalize_archive_member_name` now raises `ValueError` on a `..` escape; benign degenerate entries (empty, `.`, `/`) still return None (skipped) and absolute paths are still neutralized to relative. The raise propagates to `_ArchiveEntryLoader._build_skill`, which already skips the skill on error. - Update tests: traversal cases now assert a raise, and add an end-to-end test that a zip-slip archive drops the whole skill. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9e358a4e-538f-46be-8c58-128b6182352d * Python: Revert archive skill demo in toolbox MCP skills sample Restore the 12_foundry_toolbox_mcp_skills sample to its pre-PR, skill-md-only form (matching the .NET Agent_Step26_FoundryToolboxMcpSkills sample, which uses skill-md and no ZIP archive): - Revert main.py, toolbox.yaml, README.md, .azdignore, .dockerignore, and escalation-policy/SKILL.md to the single-file SKILL.md version. - Remove the archive demo files added by this PR (.gitignore and escalation-policy/references/refund-matrix.md). - Soften two README notes so they no longer claim archive skills are unsupported/silently dropped (this PR adds archive support); instead frame single-file SKILL.md as a focus choice and point to the archive_* options. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9e358a4e-538f-46be-8c58-128b6182352d * Python: Clarify archive framing in mcp_based_skill sample README The mcp_based_skill sample is a generic MCP consumer that discovers whatever the server advertises; it does not itself demonstrate archive skills. Reword the archive note so it reads as an MCPSkillsSource capability rather than a sample feature, and fix the stale "unpacked to a local directory" claim to "unpacked in memory" (matching the in-memory extraction implementation). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9e358a4e-538f-46be-8c58-128b6182352d --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9e358a4e-538f-46be-8c58-128b6182352d Copilot-Session: 4f14f83d-1868-45c1-be1a-12f49a58ac36 |
||
|
|
28389df805 |
Python: Move SessionStore to core and persist Foundry Responses sessions (#7306)
* Python: Move session persistence into core Move SessionStore and durable msgspec-backed storage into core, restore sessions in Foundry Responses hosting with per-user isolation, and document the serialization design. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3e5c81ad-75e8-4e92-a883-8bbd676c6c8c * Python: Address session persistence review feedback Harden scoped file paths and corruption recovery, preserve session serialization compatibility, clarify dependency placement, and add reproducible benchmark evidence. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3e5c81ad-75e8-4e92-a883-8bbd676c6c8c * Python: Preserve session snapshot compatibility Deep-copy in-memory session writes and retain existing Telegram session keys so stored conversations continue resolving. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3e5c81ad-75e8-4e92-a883-8bbd676c6c8c * Python: Simplify Foundry session isolation Add experimental FoundrySessionStore backed by Agent Server request context, remove resolver plumbing, and centralize v2 user isolation for sessions, checkpoints, and approvals. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3e5c81ad-75e8-4e92-a883-8bbd676c6c8c * Python: Reduce Foundry session helper layering Inline the single-use request user accessor while keeping separate context validation, fingerprint, and directory helpers for their distinct callers. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3e5c81ad-75e8-4e92-a883-8bbd676c6c8c * Python: Clarify Foundry request context validation Separate fail-fast request validation from context retrieval so Responses no longer appears to discard a returned context. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3e5c81ad-75e8-4e92-a883-8bbd676c6c8c * Python: Share Foundry request context helpers Move protocol validation and user-scope derivation into a dedicated request-context module, leaving the session-store module focused on storage. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3e5c81ad-75e8-4e92-a883-8bbd676c6c8c * Restore Foundry checkpoint storage paths Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3e5c81ad-75e8-4e92-a883-8bbd676c6c8c * Simplify Foundry session storage paths Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3e5c81ad-75e8-4e92-a883-8bbd676c6c8c * Persist Foundry sessions under hosted home Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3e5c81ad-75e8-4e92-a883-8bbd676c6c8c * Make hosted path test platform independent Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3e5c81ad-75e8-4e92-a883-8bbd676c6c8c * Address session persistence review feedback Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3e5c81ad-75e8-4e92-a883-8bbd676c6c8c * Isolate Foundry session path handling Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3e5c81ad-75e8-4e92-a883-8bbd676c6c8c * Clarify Foundry session path terminology Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3e5c81ad-75e8-4e92-a883-8bbd676c6c8c * Align Foundry sessions with Responses continuity Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3e5c81ad-75e8-4e92-a883-8bbd676c6c8c * Finalize Foundry Responses session persistence Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3e5c81ad-75e8-4e92-a883-8bbd676c6c8c * Add session store feature usage telemetry Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3e5c81ad-75e8-4e92-a883-8bbd676c6c8c * Fix hosted per-call history persistence Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3e5c81ad-75e8-4e92-a883-8bbd676c6c8c --------- Copilot-Session: 3e5c81ad-75e8-4e92-a883-8bbd676c6c8c |
||
|
|
143386fecc |
Python: fix(core): restrict unpickler module-prefix allowlist to types only (#5923)
* fix(core): harden restricted pickle attribute resolution * fix(core): validate nested pickle types against allowlist --------- Co-authored-by: White-Mouse <15983334+White-Mouse@users.noreply.github.com> Co-authored-by: Evan Mattson <evan.mattson@microsoft.com> Co-authored-by: Evan Mattson <35585003+moonbox3@users.noreply.github.com> |
||
|
|
47c8e29b64 |
Python: Add FileMemoryProvider context provider sample (#7428)
* Add FileMemoryProvider sample * Address PR comments |
||
|
|
12b2893bac |
Python: Apply header_provider headers to the MCP initialize handshake and other ambient requests (#7305)
* Python: Apply header_provider headers to ambient MCP requests
MCPStreamableHTTPTool.header_provider was only invoked from call_tool(),
so the initialize handshake, load_tools/load_prompts discovery, and
background pings all went out with no headers. MCP servers that require
auth on initialize (e.g. Azure AI Search knowledge-base MCP endpoints)
therefore returned 401 before any tool call could run.
Add an ambient fallback in the _inject_headers httpx request hook: when
neither the per-call ContextVar nor the active-call snapshot is set, the
hook invokes header_provider({}) so every ambient request is
authenticated. Providers that require per-call kwargs raise on the empty
dict; that is caught, logged, and the request proceeds unauthenticated,
preserving prior behavior. Calling the provider on demand also keeps
dynamic token refresh working for post-connect requests.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: cf0c1dbf-99bc-4f3f-bcf4-7791ce7dbe6a
* Python: address review - distinguish unset vs empty headers, warn once
Review feedback on the ambient header_provider fallback:
- Distinguish 'unset' (no active call) from 'set but empty' (call_tool
produced no headers). Use _mcp_call_headers.get(None) and the None-ness
of the snapshot instead of a truthiness check, so a provider that
legitimately returns {} during a real call is no longer re-invoked by
the ambient fallback mid-call.
- A kwargs-dependent provider raises on every ambient request (initialize,
discovery, recurring pings). Warn once per tool instance with a
traceback via _ambient_header_warning_emitted and drop subsequent
occurrences to DEBUG to avoid log spam.
Add regression tests for both behaviors.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: cf0c1dbf-99bc-4f3f-bcf4-7791ce7dbe6a
* Python: narrow ambient header_provider catch to KeyError
Only the missing-per-call-kwargs case (KeyError, e.g. the
mcp_api_key_auth.py sample indexing kwargs['mcp_api_key']) is tolerated
during ambient requests. Any other exception - a token-refresh failure
or a provider bug - now propagates instead of being silently converted
into unauthenticated traffic, matching the call_tool path which does not
catch header_provider exceptions.
Add a regression test asserting a non-KeyError provider failure
surfaces from the request hook.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: cf0c1dbf-99bc-4f3f-bcf4-7791ce7dbe6a
* Python: address review - raise instead of assert, simplify ambient logging
- Reword the ambient-fallback comment to describe the kwargs-dependent
provider pattern generically instead of naming a sample file, which
would go stale if the sample is renamed (also in a test docstring).
- Replace the type-narrowing assert with a RuntimeError carrying a
concise message for the unreachable no-provider state.
- Drop the warn-once/_ambient_header_warning_emitted machinery; the
KeyError ambient case is expected and benign, so log a single DEBUG
line and proceed without headers.
Update the corresponding test to assert behavior (request proceeds
without an Authorization header and no WARNING is emitted) instead of
log-count.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: cf0c1dbf-99bc-4f3f-bcf4-7791ce7dbe6a
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Evan Mattson <35585003+moonbox3@users.noreply.github.com>
Copilot-Session: cf0c1dbf-99bc-4f3f-bcf4-7791ce7dbe6a
|
||
|
|
93e8cb2de3 |
Python: make SerializationMixin.from_dict enforce the documented type check (#7256)
from_dict resolved the expected type identifier from the payload itself
(_get_type_identifier(value) prefers value["type"]), so the mismatch
guard could never fire: any supplied 'type' matched itself, and a payload
like {"type": "function_tool", ...} silently deserialized into a Message,
getting its type rewritten on the next to_dict. The docstring has always
promised a ValueError on mismatch.
Resolve the identifier from the class instead, matching what to_dict
emits, so a mismatched or foreign 'type' now raises as documented.
Payloads without a 'type' field and dependency-injection lookups are
unchanged: in every previously valid case the class-resolved identifier
is the same string the payload carried.
Fixes #7255
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
Co-authored-by: Eduard van Valkenburg <eavanvalkenburg@users.noreply.github.com>
|
||
|
|
b64a2e2f82 |
Python: add feature-usage User-Agent telemetry (#7420)
* Python: add first-pass feature usage telemetry Add the 128-bit feature accumulator, package-local indexes, activation markers, and destination-scoped User-Agent emission for the initial Python implementation slice. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 346bf168-b668-4c4a-a8db-a67282ee5e5f * Python: track declarative feature usage Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 346bf168-b668-4c4a-a8db-a67282ee5e5f * Python: complete feature usage telemetry Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 346bf168-b668-4c4a-a8db-a67282ee5e5f * Python: report core version in User-Agent Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 346bf168-b668-4c4a-a8db-a67282ee5e5f * Python: configure Lab telemetry import path Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 346bf168-b668-4c4a-a8db-a67282ee5e5f * Python: preserve telemetry transport behavior Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 346bf168-b668-4c4a-a8db-a67282ee5e5f * Python: preserve caller-owned Foundry transports Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 346bf168-b668-4c4a-a8db-a67282ee5e5f * Python: remove stale Anthropic test import Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 346bf168-b668-4c4a-a8db-a67282ee5e5f --------- Copilot-Session: 346bf168-b668-4c4a-a8db-a67282ee5e5f |
||
|
|
962b86ddbb |
Python: preserve model emission order in AG-UI MESSAGES_SNAPSHOT (#7239)
* Preserve model emission order in AG-UI messages snapshot * Address moonbox3's review: cover the remaining snapshot gaps - Preopened message ids (tool-only path) now open a text segment when the first text arrives, so their content can't drop out of the snapshot. - A tool result closes the current tool-call segment, so call A -> result A -> call B snapshots as two pairs in stream order. - emitted_call_ids only marks calls actually emitted, keeping stale segment ids eligible for the leftover fallback. - The leftover path carries its tool results too instead of dropping them. * Python: narrow leftover tool-call ids so pyright accepts the update --------- Co-authored-by: Evan Mattson <35585003+moonbox3@users.noreply.github.com> |
||
|
|
80eb2570c7 | Remove indices in FHA sample names (#7405) | ||
|
|
d42c78c8cf |
Python: Add GitHub Copilot BYOK sample (#7336)
* Python: Add GitHub Copilot BYOK sample Demonstrates routing GitHubCopilotAgent requests through a custom OpenAI-compatible endpoint via ProviderConfig instead of the default GitHub Copilot backend. * Potential fix for pull request finding Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> * Python: Address BYOK sample review feedback - Make the provider type configurable via BYOK_PROVIDER_TYPE (default "openai") instead of hardcoding "openai" — a partial autofix commit had already updated the docstring to document this env var but left the code hardcoded, which this finishes. - Stop calling the endpoint "OpenAI-compatible" everywhere; Anthropic isn't OpenAI-wire- compatible, so reword to "your own endpoint" and list the actual supported providers (mirrors the equivalent .NET sample fix). --------- Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> |
||
|
|
d9e1990484 |
Bump postcss (#7315)
Bumps [postcss](https://github.com/postcss/postcss) from 8.5.15 to 8.5.23. - [Release notes](https://github.com/postcss/postcss/releases) - [Changelog](https://github.com/postcss/postcss/blob/main/CHANGELOG.md) - [Commits](https://github.com/postcss/postcss/compare/8.5.15...8.5.23) --- updated-dependencies: - dependency-name: postcss dependency-version: 8.5.23 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
fa7cc021c6 |
Bump postcss (#7314)
Bumps [postcss](https://github.com/postcss/postcss) from 8.5.15 to 8.5.23. - [Release notes](https://github.com/postcss/postcss/releases) - [Changelog](https://github.com/postcss/postcss/blob/main/CHANGELOG.md) - [Commits](https://github.com/postcss/postcss/compare/8.5.15...8.5.23) --- updated-dependencies: - dependency-name: postcss dependency-version: 8.5.23 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
4d67eefa5f |
Python: Fix FoundryAgent inheriting OPENAI_CHAT_MODEL for agent-reference requests (#7283)
* Python: Fix FoundryAgent inheriting OPENAI_CHAT_MODEL for agent-reference requests (#7272) * Python: Fix FoundryAgent inheriting OPENAI_CHAT_MODEL for agent-reference requests * fix(foundry): update test typing annotations to pass mypy, pyrefly, and ty --------- Co-authored-by: Eduard van Valkenburg <eavanvalkenburg@users.noreply.github.com> |
||
|
|
32928e645b |
Python: Bound summarization input before provider call (#7375)
* Bound summarization input before provider call SummarizationStrategy now selects complete message groups that fit a configurable summary input token budget before calling the summary client. Only messages actually sent to the summarizer are annotated and excluded, leaving oversized later groups for a later compaction pass instead of shipping the whole transcript unbounded. Validation: uv run pytest packages/core/tests/core/test_compaction.py -k bounds_summary_input -m "not integration" failed before the implementation and passed after it; uv run pytest packages/core/tests/core/test_compaction.py -m "not integration" passed; uv run poe test -P core passed; uv run poe install completed; uv run poe check -P core passed. * Handle oversized leading summary groups Skip individually over-budget leading groups when selecting summarization input so a large early transcript item does not prevent later compactable groups from being summarized. Validation: uv run pytest packages/core/tests/core/test_compaction.py -k skips_oversized_first_group -q; uv run pytest packages/core/tests/core/test_compaction.py -q; uv run poe check -P core. * Escalate repeated summary failures Track consecutive SummarizationStrategy failures and emit a single error once the strategy has failed three times without a successful summary. Reset the escalation state after a successful summary so only persistent failures become loud. Validation: uv run pytest packages/core/tests/core/test_compaction.py -k 'repeated_summary_failures or resets_failure_escalation' -q; uv run pytest packages/core/tests/core/test_compaction.py -q; uv run poe check -P core. * Refine summary input selection Avoid rebuilding and re-tokenizing the full selected summary transcript on every candidate group while preserving complete-group selection and oversized leading group skipping. Tighten the scripted summarizer test helper to expected Exception failures instead of BaseException. Verification: uv run pytest packages/core/tests/core/test_compaction.py -q; uv run poe syntax -P core. --------- Co-authored-by: Eduard van Valkenburg <eavanvalkenburg@users.noreply.github.com> |