07ddabbef72677cc44ec484cb2e84aefdd78b4e9
960 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7e5ba70884 |
Python: Stop swallowing skill script and resource errors so the model can self-correct (#6755)
* Python: Add include_detailed_errors option for skill script execution Port the .NET fix from #6680. SkillsProvider previously swallowed exceptions from skill script execution and resource reading, returning a generic error string so the model could not self-correct. - Add an include_detailed_errors option to SkillsProvider.__init__ and from_paths. When True, script-execution failures return an error string with the exception message appended; when False (default), the exception is logged and re-raised, delegating to the function-invocation pipeline's own include_detailed_errors policy. - _read_skill_resource now logs and re-raises instead of returning a generic error string. Resources take no model arguments, so a swallowed generic error is not actionable by the model. - Update and add tests covering the new propagation and detailed-error behavior. Fixes #6681 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Re-raise skill script/resource errors instead of adding a provider option Address PR review: returning a plain error string from the skill provider bypassed the shared tool-error contract (no exception metadata, not counted toward consecutive-error limits), risking infinite retries. Instead of porting the .NET provider-level IncludeDetailedErrors option, _run_skill_script and _read_skill_resource now always log and re-raise on failure. This delegates error handling to the function-invocation pipeline, whose existing include_detailed_errors policy is the Python equivalent of .NET's FunctionInvokingChatClient.IncludeDetailedErrors and correctly preserves exception metadata and consecutive-error counting. Validation failures (empty/unknown skill, script, or resource names) still return user-facing error strings. Tests updated accordingly. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
6dd30950c1 |
Python: Fix background agent telemetry context error (#6764)
* Fix issue when using background agents with telemetry * Address PR comments * Fix issue when using background agents with telemetry * Address PR comments * Fix uv.lock * Remove unecessary comments and commit hook reformatted code |
||
|
|
c89a539c02 |
Python: [BREAKING] Make all SkillsProvider tools require approval by default (#6754)
* Python: [BREAKING] Make all SkillsProvider tools require approval by default All tools exposed by SkillsProvider (load_skill, read_skill_resource, run_skill_script) now require approval by default. Previously only run_skill_script could be gated, and only when require_script_approval=True. - Register all three tools with approval_mode="always_require" - Add read_only_tools_auto_approval_rule and all_tools_auto_approval_rule static rules plus tool-name constants (mirrors FileAccessProvider) - Remove the require_script_approval option from __init__ and from_paths - Add skills_auto_approval sample; update script_approval sample/docs Closes #6728 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address PR review: batch skill approval responses and tidy sample - Collect a response for every approval request and send them in a single agent.run so the approval loop always makes progress (no infinite loop when a request lacks a function_call); reject non-function requests instead of skipping them. Applied to both the skills_auto_approval and script_approval samples. - Extract ToolApprovalMiddleware into a local variable in skills_auto_approval for readability. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address PR review: add approval handling to remaining skills samples The secure-by-default change makes all SkillsProvider tools require approval, which left the other skills samples emitting approval requests instead of the documented answers. Add ToolApprovalMiddleware with the all-tools auto-approval rule (and a session, which the middleware requires) so these samples run unattended as before: - code_defined_skill, file_based_skill, class_based_skill, mixed_skills, skill_filtering, mcp_based_skill - providers/foundry/foundry_chat_client_with_toolbox_skills The dedicated script_approval (manual) and skills_auto_approval (selective) samples continue to demonstrate interactive approval handling. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address PR review: simplify "host approval" wording to "approval" Apply maintainer suggestions dropping "host" from the skill-approval docstrings, and align the matching SkillsProvider docstring/AGENTS.md note for consistency. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
6dfcbc5c62 |
Python: support stable + preview Azure AI Search (Foundry IQ) API versions (#6603)
Update agent-framework-azure-ai-search to work across the stable/GA azure-search-documents SDK (12.0.0, api-version 2026-04-01) and the preview SDK (12.1.0b1, api-version 2026-05-01-preview) for both semantic and agentic modes. - Bump the dependency to azure-search-documents>=12.0.0,<13 and the package to 1.0.0b260618. - Add an api_version parameter (threaded into SearchClient, SearchIndexClient, and KnowledgeBaseRetrievalClient) plus STABLE_API_VERSION/PREVIEW_API_VERSION constants, re-exported from agent_framework.azure. - Auto-detect preview-only agentic features (output mode, low/medium reasoning effort) via _preview_features_active(), which requires both the preview SDK and a preview api-version; defaults (extractive + minimal) work on both channels and preview-only options raise an actionable error otherwise. - Make knowledge-base imports SDK-version resilient and fix the 12.x surface (k -> k_nearest_neighbors, defensive additional_properties). - Update tests (pass on both SDKs), docs, samples, CHANGELOG, and uv.lock. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
4272d90051 |
Python/.Net: Agent Harness blog post accompanying samples part 2 (#6692)
* Add samples for the harness blog part 2 * Address PR comments * Fix blog links. * Address PR comments * Fix bug where mode was incorrectly defaulted when reading the mode before the first run. * Add reference to new sample readme |
||
|
|
9a565f2bf8 |
Python: convert Pydantic model class response_format to JSON schema in OllamaChatClient (#6782)
Ollama's `format` param only accepts '', 'json', or a JSON-schema dict, so passing a Pydantic model class (the form OpenAIChatClient/FoundryChatClient and create_harness_agent plan mode use) raised a ValidationError while building the request. Convert a model class to its JSON schema when mapping response_format -> format, keeping the original class for typed response parsing. |
||
|
|
9fd3d29e09 |
Python: Fix FunctionShellTool throw and empty streaming shell command (#6763)
* Fix shell tool bug * Address PR feedback * Fix uv.lock changes * Update uv.lock |
||
|
|
6968a7fc59 | Updating background agent loop to resolve provider automatically and add feedback message builder. (#6735) | ||
|
|
4d4db7f501 |
Python: create_harness_agent skills_paths accepts str | Path | Sequence[str | Path] | None (#6717)
* fix: skills_paths accepts str | Path | Sequence[str | Path] | None * fix: update _assemble_context_providers skills_paths annotation to match public signature --------- Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> |
||
|
|
730bcee9ea |
Python: Autolabelling MCP servers based on hints and Github MCP server ifc labels (#6171)
* Python: add GitHub MCP security label sample * modified samples to create devui auth token, support debugging with security, and change context label only using the labels of unhidden result from tools * FIDES: secure MCP labeling, _meta IFC parsing, and docs updates * FIDES: secure MCP labeling, _meta IFC parsing, and docs updates * modified docs * fixed PR comments, simplified github_mcp example * commented github_mcp example * remove the parse_github_mcp_labels and fix the user_identity label propogation * fix: use standard GitHub MCP endpoint with X-MCP-Features: ifc_labels instead of /insiders - Switch MCP_URL from /mcp/insiders to /mcp/ in github_mcp_example.py - Add MCP_HEADERS constant with X-MCP-Features: ifc_labels to opt-in to server-side IFC label emission in _meta payloads - Fix SecureMCPToolProxy to pass headers via httpx.AsyncClient so they are included on session.initialize(), not just on tool calls (was causing 401 to silently surface as anyio cancel-scope CancelledError) - Update README, FIDES_DEVELOPER_GUIDE, FIDES_IMPLEMENTATION_SUMMARY, and 0024-prompt-injection-defense.md to remove all /insiders references * address PR comments * Simplify GitHub MCP security sample to DevUI-only; document SecureAgentConfig quarantine client global behavior * minor PR comments * fixing failed checks * fixing failed checks --------- Co-authored-by: Eduard van Valkenburg <eavanvalkenburg@users.noreply.github.com> |
||
|
|
a1c37b69e0 |
[Generated by SRE Agent] Clarify identifier security guidance (#6510)
Co-authored-by: Azure SRE Agent <noreply@microsoft.com> Co-authored-by: Evan Mattson <35585003+moonbox3@users.noreply.github.com> |
||
|
|
f1d838fc5e |
Python: bump package versions for 1.10.0 release (#6753)
* Python: bump package versions for 1.10.0 release - Released cohort (core, openai, foundry, root): 1.9.0/1.8.2 -> 1.10.0 - agent-framework-ag-ui: rc5 -> rc6 (tool history replay fix) - Beta/alpha packages with changes: anthropic, azurefunctions, bedrock, durabletask, hyperlight, purview, foundry-hosting, gemini, hosting, hosting-responses, hosting-telegram, tools bumped to new date stamp (260625) - Inter-package dependency bounds updated for changed packages - CHANGELOG.md updated with [1.10.0] section and compare links Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix: update stale hosting dependency pins in hosting-responses and hosting-telegram Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * CI: cap xdist workers at 4 for Azure OpenAI and Functions integration jobs The Azure OpenAI and Functions+Durable Task integration jobs ran with `-n logical` (~20 workers on the hosted runner), oversubscribing the box and collapsing the whole pytest session (all workers reporting `node down: Not properly terminated`) in the merge queue. Pin these two jobs to `-n 4` in python-merge-tests.yml and python-integration-tests.yml to remove the oversubscription while keeping full coverage. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * test: temporarily skip flaky Python integration tests crashing the merge queue Revert the `-n 4` xdist experiment (it did not prevent the runner crash) and instead skip the integration tests that collapse the pytest-xdist runner in the merge queue (all workers report `node down: Not properly terminated`): - Azure OpenAI: flip the per-file `skip_if_azure_openai_integration_tests_disabled` guard to an unconditional skip (integration tests only; unit tests still run). - Azure Functions / Durable Task: skip the four specific failing tests (test_weather_agent, test_parallel_workflow_end_to_end, test_weather_agent_with_tool, test_conditional_branching). Tracked for re-enablement in #6777. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * test: skip flaky test_math_agent_with_tool (durabletask integration) Same empty-AgentResponse flakiness as test_weather_agent_with_tool in the same file (AssertionError: assert 0 > 0 / empty .text). Skip it in the merge queue. Tracked in #6777. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
3c3feb8705 |
Python: Refactor runner/workflow responsibilities and fix checkpoint ancestry bug (#6695)
* Refactor runner/workflow responsibilities, add concurrency guards, and fix checkpoint ancestry bug Move runner-state ownership out of Workflow into Runner for clearer responsibilities. Add a weakref-based concurrent-run guard in Workflow and fix the stream-drop race in run_until_convergence. Fix the checkpoint ancestry bug by tracking the previous checkpoint id as runner instance state so parent pointers persist across resumed runs. Move Runner to a deprecated lazy __getattr__ export (backward-compatible with DeprecationWarning) and export CheckpointID. * Scope runtime checkpoint storage to its owning run Close the stream-drop race where a dropped run's deferred async-generator finalizer could leave a runtime checkpoint storage override set (inherited by a new run) or clear a successor run's storage. run() now defensively clears any stale override before starting, and _run_core only clears the override if this run still owns it (mirroring the _active_run ownership guard). Adds regression tests for both the inheritance and clobber cases. * Collapse runtime-storage ownership into the active-run weakref _runtime_storage_owner always held the same weakref as _active_run, so the two ownership conditions were equivalent. Derive ownership from a single owns_run = (_active_run is my_active_run) captured before the active-run clear, and remove the redundant field. No behavior change. * Nest runtime-storage clear under the owns_run guard Both the active-run release and the runtime-storage clear are gated on owns_run, so fold the storage clear inside the if owns_run block. No behavior change. * Reset resume flag in a finally so it can't leak across runs _resumed_from_checkpoint was only cleared on the success path of run_until_convergence, so a failure during a resumed run (e.g. executor failure) left it True. The next fresh run then skipped the superstep-0 checkpoint and parented later checkpoints to the stale resume point. Move the reset into a finally. Add a regression test that fails a resumed run via an executor error and asserts the next fresh run creates the superstep-0 checkpoint. * Fix tests and formatting * Fix formatting * Address comments * Update type ignore statements |
||
|
|
d75f2286f4 |
Python: Add Telegram channel for agent-framework-hosting (#6698)
* Python: Add Telegram channel for agent-framework-hosting - Add agent-framework-hosting-telegram package with TelegramChannel supporting polling and webhook transports, streaming edits with Telegram Bot API rate limiting, per-chat serial workers, and multi-modal inbound/outbound (text, photo, document, voice) - Add local_telegram sample demonstrating multi-channel hosting with a TelegramChannel alongside ResponsesChannel, using per-chat FileHistoryProvider and a run_hook for Telegram persona temperature - Fix test layout: move tests to tests/hosting_telegram/ (no __init__.py) - Remove old [tool.mypy] section and mypy poe task; source type-checking is handled by pyright via shared_tasks - Update uv.lock, pyproject.toml workspace sources, and PACKAGE_STATUS.md Fixes #6588 Refs #6265 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Address Telegram channel CI failures and review feedback - Fix webhook secret validation to use constant-time compare_digest - Harden webhook update parsing: require integer chat IDs and guard slash-only commands - Fix streaming edge cases in TelegramChannel: - prevent edit worker deadlocks when text exceeds 4096 chars - prevent deadlock when placeholder send fails (message_id stays None) - enforce edit throttling with minimum interval sleep - honor send_typing_action=False in streaming mode - always forward final multimodal output (e.g. images), while avoiding duplicate text sends - Expand Telegram tests for slash-only command handling, non-int chat IDs, and streaming behavior (long text, final images, typing toggle) - Fix sample/docs feedback: - rename sample package to agent-framework-hosting-sample-local-telegram - switch sample uv.sources from feature branch to main - align docs/tool names with lookup_weather - fix broken links and server run instructions in README/call_server.py - align local_telegram app docstrings with reasoning hook behavior and strip model in responses_hook Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Fix TelegramChannel streaming to iterate contents for multimodal support - Remove stale PR reference from module docstring - Add Google-style docstring to TelegramChannel.__init__ documenting all keyword args - Fix _stream_to_chat to iterate update.contents instead of using getattr(update, 'text', None); text chunks are extracted from Content items with type='text', non-text content in updates is correctly ignored (images etc. are forwarded via the final response) - Update _FakeStreamUpdate test helper to use contents list matching the real AgentResponseUpdate API; add from_text/from_image class methods - Update _FakeResponseStream to accept _FakeStreamUpdate objects directly - Add test verifying multimodal stream updates don't corrupt text accumulator Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Split local_telegram into simple Telegram-only and new multi-channel sample local_telegram is now a focused Telegram-only sample: - Removes ResponsesChannel and all responses_hook code - Removes call_server.py (no HTTP endpoint to call) - Uses a deterministic lookup_weather tool (hash-based, not random) - Single run_hook that strips model and raises reasoning effort - Drops agent-framework-hosting-responses dependency New local_multi_channel sample shows running both channels at once: - ResponsesChannel + TelegramChannel sharing a FileHistoryProvider - Cross-channel session resumption via previous_response_id - call_server.py moved here (the Responses endpoint lives here now) - Demonstrates the multi-channel coordination story Update README table to list both samples with clear descriptions. Also delete personal_assistant/.venv which was not tracked but caused pyright to crawl the entire installed venv (thousands of files), making sample pyright checks hang indefinitely. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Fallback when Telegram final edit fails - only mark final edit as sent after a confirmed 2xx edit response - fall back to sendMessage when final edit returns a non-success status - add regression test covering failed final edit fallback behavior Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Fix optional await_args typing in telegram test - assert await_args is not None before reading kwargs in streaming fallback test - resolves test-typing failures across mypy/pyright/ty/zuban for hosting-telegram Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
ce74c84bdb |
Python: Preserve OTel parent context for deferred streams (#6709)
* Python: Preserve OTel parent context for deferred streams - capture current OTel context when opening host-managed streaming runs - re-activate captured context during deferred stream pulls and finalization - add host-level regression coverage for deferred stream parent-span linkage - add Responses channel integration coverage for request parent span propagation Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Capture OTel stream context before target.run - capture OTel context snapshot before invoking target.run in _invoke_stream - add regression test guarding capture-before-run evaluation order Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
4fb1fb615a |
Python: Fix Hyperlight CodeAct span parenting (#6712)
* Fix Hyperlight CodeAct span parenting Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix Hyperlight test OTEL fixture Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix Hyperlight test typing annotations Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
9f1ee23a4b |
Python: [BREAKING] Refactor FileSkillsSource for depth-based discovery and predicate filters (#6488)
* Python: [Breaking] Refactor FileSkillsSource for depth-based discovery and predicate filters Refactors FileSkillsSource to make script and resource discovery more flexible. ## Changes - **Drops** resource_directories / script_directories options (preconfigured directory whitelists). - **Adds** search_depth option (>= 1, default 2): controls how deep the recursive scan goes within each skill directory. - **Adds** script_filter / resource_filter predicate options that receive a FileSkillFilterContext (skill_name + relative_file_path), allowing whitelist/blacklist filtering by file path. - **Adds** FileSkillFilterContext class exported from agent_framework. ## Notes - The Skills API is marked @experimental -- the option removals are intentional breaking changes within the experimental surface. - Security checks (path containment, symlink detection) are preserved and continue to use the skill root directory as the trusted boundary. - Ports the same refactoring from .NET PR #6109 while following Python conventions (instance methods, Callable type hints, __slots__). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address PR feedback: clarify depth constants and skip nested skill directories - Add clarifying comments distinguishing MAX_SEARCH_DEPTH (SKILL.md discovery) from DEFAULT_SEARCH_DEPTH (per-skill resource/script scanning). - Stop recursing into subdirectories that contain their own SKILL.md, preventing child skill files from being attached to the parent skill. - Add test verifying nested skill boundary is respected. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Remove __slots__ from FileSkillFilterContext and add type-ignore comments - Remove __slots__ from FileSkillFilterContext per reviewer feedback — the optimization is negligible and inconsistent with sibling classes. - Add type: ignore[attr-defined] / ty: ignore[unresolved-attribute] comments to test lines accessing private _resources/_scripts attributes, matching the convention established on main. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Simplify filter predicates: remove FileSkillFilterContext, use Callable[[str, str], bool] Address reviewer feedback: - Remove FileSkillFilterContext class — a dedicated class for two strings is overkill in Python. Filters now receive (skill_name, relative_file_path) directly as positional args. - Update docstrings to describe behavior instead of referencing private instance attributes. - Remove FileSkillFilterContext from exports and __all__. - Update all test lambdas and remove TestFileSkillFilterContext class. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Use DEFAULT_SEARCH_DEPTH as default argument directly Instead of accepting int | None and resolving None to the default internally, use DEFAULT_SEARCH_DEPTH as the parameter default value on both FileSkillsSource.__init__() and SkillsProvider.from_paths(). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
1df47667ea |
Python: surface cache and reasoning token counts for the Bedrock and Gemini connectors (#6640)
* Python: surface Gemini cached and thinking token counts in usage details * Python: surface Bedrock cache token counts in usage details * Python: surface Gemini cached and thinking token counts in usage details * Python: surface Bedrock cache token counts in usage details * Return None from Bedrock _parse_usage when no token counts are present Matches the UsageDetails | None return annotation and the Gemini connector's behavior, so a usage payload with no recognized keys no longer propagates an empty mapping. Adds a regression test. |
||
|
|
91f639a694 |
Python: Explicitly emit available_resources and available_scripts in skill content (#6694)
Skill content now always emits <available_resources> and <available_scripts> blocks, using self-closing elements when empty, so models receive an authoritative list per category and do not hallucinate resource/script names. FileSkill now also emits its resources block. Closes #6348 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
d5c15f2fe1 |
.NET/Python: Purview: prefer token principal for user identity (#6693)
* Purview: prefer token principal for user identity Align Purview middleware identity resolution so user-token principals are preferred before supplied message identities, while app-token flows continue to use validated fallback user IDs. Also fix the content activities user route and add regression coverage for identity precedence and route construction. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * .NET: Fix user ID resolution logic in ScopedContentProcessor and add unit test for empty token user ID --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
acb28a63b5 |
Python: Fix MCP metadata and tool name handling (#6656)
* Fix MCP metadata and tool name handling Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address MCP review feedback Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
f2d02e58b3 |
Python: Add hosting core and Responses channel (#6580)
* Add Python hosting core and Responses channel Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address hosting core review feedback Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Adopt source pyright typing setup for hosting packages Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Cover ResponsesChannel custom path routing Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Align hosting tests with package layout Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix hosting workflow fixture imports in aggregate tests Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Apply useful Responses channel hardening Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix hosting package typing checks Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix hosting pyright under Python 3.11 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Avoid static diskcache dependency in hosting Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix aggregate typing and Docker test resilience Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Simplify local Responses workflow sample Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Clarify generic hosting is not Foundry hosting Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Revert "Clarify generic hosting is not Foundry hosting" This reverts commit 73b584d919053bed43a258d75dc2b76406e9c181. * Clarify isolation key source flexibility Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Clarify isolation header reuse boundary Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Support multimodal Responses channel outputs Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Preserve multimodal streaming Responses output Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Stream Responses output items from updates Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Improve Responses streaming output handling * Tighten Responses channel default option handling - Restore full option parsing in parse_responses_request: known fields are remapped (max_output_tokens→max_tokens, parallel_tool_calls→ allow_multiple_tool_calls), transport/session keys excluded, None values dropped, everything else forwarded as-is so run_hook can inspect the full set. - Add a default _strip_options_hook on ResponsesChannel that removes all parsed options before reaching the agent. Callers cannot inject generation params (temperature, instructions, tools, …) unless the host explicitly allows it. - A custom run_hook replaces the default entirely and receives the full ChannelRequest.options plus the raw protocol_request. - Update tests to cover remap, default-strip, and custom-hook paths. - Clarify host debug-log docstring to match new option flow. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
36420c515e |
Python: Align serialized tool format to OTel GenAI tool def format (#6556)
* Align serialized tool format to OTel GenAI tool def format * Cache serialized tools |
||
|
|
7051a4920d |
Python: Add MCP as a hard dep in Foundry Hosting (#6634)
* Add MCP as a hard dep in Foundry Hosting * Pin GitHub SDK * Fix formatting * Fix formatting |
||
|
|
a2018b40f9 |
Python: [BREAKING] Require approval for file-access tools with read-only auto-approval (#6599)
* Require approvals for file-access and expose auto approval funcs for it * Scope file-access auto-approval rules to local tools; fix base-Agent sample Address PR #6599 review feedback: - read_only/all_tools auto-approval rules now reject any call carrying a server_label so they stay scoped to FileAccessProvider's local tools and never auto-approve a same-named hosted tool. - Expand the FileAccessProvider docstring to explain the runtime effect of approval_mode="always_require" and point to ToolApprovalMiddleware / create_harness_agent. - Fix the base-Agent file_access_data_processing sample, which would otherwise stop executing file tools under the new always_require defaults, by adding ToolApprovalMiddleware with all_tools_auto_approval_rule. - Add tests covering hosted (server_label) calls and update docs. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Clean up comments * Update sample after merge --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
7f2e19ca2f |
Python: Ensure spans created inside sync preparations in streaming call are correctly nested (#6552)
* Make sure spans created inside sync ops in streaming path are correctly nested * Add tests * Fix comments * Fix typing |
||
|
|
7b6f582b13 |
Python: Agent Harness blog post accompanying samples part 1 (#6605)
* Add samples for harness blog post part 1 * Add readme for python samples * Update python instructions to match dotnet instructions * Address PR comments * Add link to blog posts * Fix blog post naming. * Add more blog post links |
||
|
|
d108d4b549 |
Python: [BREAKING] Integrate looping into HarnessAgent (#6607)
* Integrate looping into harness * Address PR comments * Address PR comments. * Fix typing error |
||
|
|
fd160a7782 |
Python: fix dependency maintenance cutoff (#6658)
* Python: fix dependency maintenance cutoff Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: fix Hyperlight output dir typing Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
fc3111c391 |
Python: Add FoundryAgent conversation session helper (#6623)
* Add FoundryAgent conversation session helper Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Simplify Foundry conversation session helper Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Rename Foundry conversation helper Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * use named kw --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
7435dd48d0 |
Python: harden Hyperlight output capture against symlinks (#6601)
* Python: harden Hyperlight output capture against symlinks Mirror the input-staging symlink hardening on the output-capture path of HyperlightExecuteCodeTool. Output discovery now walks via the symlink-safe _iter_real_entries instead of rglob, per-file collection validates that no path component is a symlink and the final entry is a regular file, and file reads use os.O_NOFOLLOW. Adds regression tests for the output path. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address review: reject traversal, fix listing test, harden read - _is_safe_output_file now rejects '.'/'..' components (lexical relative_to could otherwise escape root without a symlink) - _read_output_file_bytes adds a cross-platform TOCTOU guard (lstat/fstat st_dev+st_ino identity check) since O_NOFOLLOW is absent on Windows - fix intermediate-dir-symlink test to use a relative listing path so it exercises normalization + validation; add a parent-traversal unit test Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
148f57020a |
Python: host MAF workflows on a standalone Durable Task worker (#6418)
* feat(durabletask): host MAF workflows on a standalone Durable Task worker
Add a host-agnostic workflow execution engine to agent-framework-durabletask so a MAF Workflow can run as a durable orchestration outside Azure Functions:
- WorkflowOrchestrationContext protocol + DurableTaskWorkflowContext adapter, the superstep orchestrator, serialization helpers, capturing runner context, and the shared non-agent activity body (including the yield-output classifier so intermediate executors are not surfaced as final outputs).
- DurableAIAgentWorker.configure_workflow auto-registers agent executors as entities, non-agent executors as activities, and the workflow orchestrator.
- plan_workflow_registration centralizes the 'what to register' decision so it can be shared across hosts.
- run_agent_coroutine runs all agent coroutines on one persistent event loop, fixing a cross-loop hang when shared chat clients/credentials bind their asyncio primitives to a dead loop.
- DurableWorkflowClient (start/await workflow + HITL discover/respond); DurableAIAgentClient stays agent-only.
* refactor(azurefunctions): delegate workflow execution to agent-framework-durabletask
AgentFunctionApp now reuses the shared orchestrator, activity body, and registration planner from agent_framework_durabletask instead of maintaining its own copies; _workflow.py becomes a thin host-specific adapter (AzureFunctionsWorkflowContext).
- Run agent entity coroutines on the shared persistent event loop, fixing the cross-loop hang.
- Relocate state-diff unit tests to the durabletask package; update entity loop tests.
* feat(core): expose durabletask workflow symbols via agent_framework.azure
Lazily re-export WORKFLOW_ORCHESTRATOR_NAME and DurableWorkflowClient from the agent_framework.azure namespace so standalone hosts can import them without depending on internal module paths.
* docs(samples): add standalone durabletask workflow and HITL samples
Add two samples under samples/04-hosting/durabletask demonstrating MAF workflows on a standalone Durable Task worker (no Azure Functions):
- 08_workflow: conditional spam-detection workflow started via DurableWorkflowClient.start_workflow / await_workflow_output.
- 09_workflow_hitl: content-moderation workflow that pauses with ctx.request_info and is resumed via DurableWorkflowClient.get_pending_hitl_requests / send_hitl_response.
Also add the durabletask workflow integration test (test_08_dt_workflow).
* fix: address PR review feedback
- Sanitize HITL external-event responses with strip_pickle_markers in the orchestrator (defense-in-depth for callers that bypass DurableWorkflowClient).
- Raise WorkflowConvergenceException when max_iterations is reached with pending messages, matching the core WorkflowRunner instead of silently returning partial output.
- Route falsy 'sent' messages (use 'is not None' instead of truthiness).
- Normalize None shared_state_snapshot/source_executor_ids in execute_workflow_activity.
- Cast Any returns in AzureFunctionsWorkflowContext to satisfy mypy/pyright.
- Fix sample docstrings to reference DurableWorkflowClient.
* fix: resolve pyright Package Checks errors
- Use typed locals instead of cast in AzureFunctionsWorkflowContext (mypy sees Any, pyright sees concrete types -> avoid reportUnnecessaryCast).
- Annotate shared_state_snapshot and cast partially-typed durabletask SDK returns / HITL custom-status parsing to satisfy reportUnknownVariableType/reportUnknownMemberType.
- Drop the dead deserialize/serialize re-export in _workflow.py and mark the intentional private _extract_message_content re-export.
* fix(durabletask): agent-executor identity and typed workflow input
Register each workflow agent entity under the executor id that the orchestrator dispatches to (instead of the agent name), so AgentExecutor(agent, id=...) works when the id differs from agent.name. The azure-functions host mirrors this.
Reconstruct the start executor declared input type from the workflow initial JSON payload in the shared engine (mirroring in-process delivery) instead of string-coercing it per host. Untrusted input is stripped of pickle markers before reconstruction to prevent deserialization RCE.
* fix(samples): type durable workflow start executors for reconstructed input
The HITL and parallel workflow samples no longer hand-parse a JSON string. Their start executors now declare their real input type (ContentSubmission / DocumentInput), which the durable engine reconstructs from the client payload before delivery.
* test(durabletask): unit coverage for registration, client, worker, and input coercion
Add unit tests for plan_workflow_registration, DurableWorkflowClient, the agent-executor identity registration (entity keyed by executor id), and the typed initial-input coercion including pickle-marker neutralization.
* test(durabletask): HITL and parallel durable workflow integration tests
Add an integration test for the standalone durabletask HITL workflow sample via a new workflow_client fixture. Re-enable the Azure Functions parallel workflow test, consolidated into one end-to-end case so the work-stealing xdist scheduler cannot spawn multiple func hosts for this sample.
* refactor(durabletask): group workflow modules into a _workflows subpackage
Move the eight workflow modules into a private _workflows/ subpackage and drop the redundant _workflow_ prefix (orchestrator.py, registration.py, activity.py, client.py, context.py, dt_context.py, runner_context.py, serialization.py). The public API and __all__ are unchanged; only direct internal-module imports were repointed (package __init__, the worker, the azure-functions shared shim, and the affected unit tests).
* fix(durabletask): harden workflow type resolution and HITL response handling
- resolve_type returns only real classes (avoids issubclass TypeError in reconstruct_to_type)
- re-wait on HITL responses rejected by pickle-marker sanitization instead of dropping the request and losing the run
- American spelling in strip_pickle_markers docstring
- unit tests for resolve_type
* fix(durabletask): treat async edge conditions as not-matched on the synchronous host
The durabletask orchestrator evaluates edge conditions synchronously and does not support async edge conditions. Such an edge is now treated as not matched (the edge is not traversed) rather than assuming a result. Adds unit coverage; full async-condition support will be handled separately.
* fix(durabletask): reconstruct typed workflow outputs at the host boundary
await_workflow_output and the Azure Functions status endpoint now decode the checkpoint-encoded outputs the shared activity produces, via a shared deserialize_workflow_output helper. The client returns the original objects; the AF endpoint emits clean domain JSON instead of checkpoint-marker dicts, keeping the two hosts consistent.
* fix(durabletask): address review findings on workflow hosting
- AF: register workflow agents through add_agent(entity_id=...) so they remain tracked in app.agents / get_agent() (restores documented behavior) while keying by the executor id the orchestrator dispatches to; mirrors DurableAIAgentWorker.add_agent.
- async bridge: treat the shared loop as reusable only while its backing thread is alive, so a dead loop thread is replaced instead of hanging future.result() forever.
- client: add get_runtime_status; the standalone HITL sample now stops polling and reports the real terminal state instead of a generic timeout.
- tests: guard send_hitl_response pickle-marker stripping and add get_runtime_status coverage.
* fix(durabletask): wait indefinitely for HITL responses, matching core
The durable workflow host previously raced HITL responses against a 72h timer and failed the orchestration on elapse. MAF core's request_info has no timeout concept (it waits for the response), and the .NET durable host waits too, so the durable Python host now does the same: it stays paused until a response arrives. Removes the hitl_timeout_hours parameter and DEFAULT_HITL_TIMEOUT_HOURS constant from both hosts. A configurable timeout can be added later once core defines the contract (what happens on elapse).
* feat(durabletask): typed workflow event streaming and async client API
Add a brokerless workflow event stream to the durable host. Each non-agent executor runs inside a durable activity that captures its real WorkflowEvents (with data payloads); the orchestrator replays them into the orchestration custom status after each superstep, and the client streams them back as typed WorkflowEvent objects with reconstructed data. Agent executors contribute synthesized invoked/completed lifecycle events.
Add async client methods run_workflow (start with optional wait) and stream_workflow (typed event iterator), plus is_replaying plumbing through the orchestration context protocol and both host adapters so live status is published only on non-replay execution.
* docs(samples): standalone durabletask workflow streaming sample
Add sample 10_workflow_streaming demonstrating the async DurableWorkflowClient API on a standalone Durable Task worker: run_workflow(wait=False) to start without blocking, then stream_workflow to consume typed WorkflowEvent objects as a WriterAgent -> ReviewerAgent -> publish pipeline runs.
* refactor(durabletask): internal-only checkpoint codec and host-scoped workflow event streaming
Two related hardening changes to the durable workflow hosting layer, plus a
rebase-restored improvement.
Internal-only serialization codec (MSRC follow-up):
- Rename serialize_value/deserialize_value -> _serialize_value/_deserialize_value
in the shared durabletask serialization module and update all call sites, so the
pickle-backed checkpoint codec is unambiguously framework-internal. Untrusted
input is still neutralized with strip_pickle_markers at the HTTP boundary.
- Remove the duplicate agent_framework_azurefunctions._serialization module and
import strip_pickle_markers from the shared durabletask module instead. Move its
unique serialization/strip-marker tests into the durabletask test suite.
Scope workflow event streaming to hosts that can carry it:
- Add WorkflowOrchestrationContext.supports_event_streaming. The standalone
DurableTask host returns True (no custom-status size cap, has a stream_workflow
consumer); the Azure Functions host returns False.
- The orchestrator now accumulates and publishes the WorkflowEvent timeline to the
orchestration custom status only when the host supports streaming. On Azure
Functions the custom status returns to its pre-streaming shape
({state[, pending_requests]}), which fixes orchestrator failures with
"The size of the JSON-serialized payload must not exceed 16 KB" and stops leaking
pickle markers into the HTTP status response. The Azure Functions status endpoint
never consumed the event stream.
Workflow start endpoint:
- Accept text/plain raw request bodies (fall back from get_json to the raw body),
restoring an improvement from main that the rebase conflict resolution dropped.
* fix(azurefunctions): scope workflow status/respond endpoints to the workflow orchestrator
The workflow/status/{instanceId} and workflow/respond/{instanceId}/{requestId}
HTTP endpoints resolved durable instances by ID only. The durable client looks up
IDs across every orchestration in the task hub (agent entities, any
user-registered orchestrations, and other apps sharing the hub), so a caller
holding one instance ID could read another orchestration's status -- including
pending HITL request payloads -- or inject external events into it.
Add AgentFunctionApp._is_workflow_orchestration() and gate both endpoints on it:
an instance whose orchestration name is not WORKFLOW_ORCHESTRATOR_NAME now returns
404 instead of leaking state or accepting events. send_hitl_response now fetches
the orchestration status and validates ownership before raising the external
event. Legitimate workflow instances are unaffected.
Mirrors the .NET fix in PR #6608.
* fix(durabletask): resolve CI typing failures
- serialization: rename _serialize_value/_deserialize_value back to
serialize_value/deserialize_value to follow the package convention for
cross-module internal helpers (matches strip_pickle_markers, resolve_type).
The leading underscore tripped pyright reportPrivateUsage on cross-module
imports under the strict source gate; internal-only status is preserved by
not exporting them from the public API.
- Remove type-ignore comments pyright flags as unnecessary
(reportUnnecessaryTypeIgnoreComment) in _worker.py, orchestrator.py,
serialization.py.
- test_08_dt_workflow: add AgentClientFactoryProtocol and annotate the
agent_client_factory fixture as type[AgentClientFactoryProtocol] (matching
test_01-07) so mypy/ty stop reporting "type has no attribute create".
- samples (08_workflow, 09_workflow_hitl): pass structured output via
FoundryChatOptions[Any](response_format=...) instead of a plain dict so the
samples pyright (basic) config accepts default_options.
---------
Co-authored-by: Gavin Aguiar <80794152+gavin-aguiar@users.noreply.github.com>
|
||
|
|
2d0555c537 |
Python: re-role trailing assistant message to user for Anthropic compatibility (fixes #5008) (#6207)
* Fix auto function calling stripping explicit null arguments (fixes #5934) * fix: re-role trailing assistant message to user for Anthropic (fixes #5008) * fix: address Copilot review feedback (exclude_unset, test coverage, synthetic user turn) * fix: update docstring and extend exclude_unset to auto_invoke_function * revert: remove unrelated core _tools.py changes from Anthropic PR The exclude_none/exclude_unset changes in the core package are out of scope for this Anthropic-specific fix. This PR now only contains the Anthropic chat client docstring fix and the synthetic user turn append. * fix: avoid appending user turn after Anthropic tool use * Fix Anthropic tool-use type narrowing Use object-typed content narrowing before checking Anthropic tool-use block types so strict Pyright no longer treats dynamic message content as Unknown. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Evan Mattson <evan.mattson@microsoft.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
5145d50be8 |
Python: Fix AG-UI tool history replay sanitization (#6581)
* Python: Fix AG-UI tool history replay sanitization * Python: Address AG-UI replay review comments Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
54a30571aa |
Dotnet - Add support for Foundry Adaptive evals (#6267)
* .NET: feat(evals): RubricScore type + EvalScoreResult.Dimensions Adds the core rubric-evaluator surface that mirrors the Python work in PR #6101 (commit e45b934cc). Provider-agnostic types only — no Foundry coupling. Subsequent commits will wire these into FoundryEvals. - RubricScore: per-dimension score record (Id, Score?, Applicable, Weight, Reason). - EvalScoreResult.Dimensions: optional init-only list of RubricScore. Null for non-rubric (built-in) evaluators. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * .NET: feat(evals): GeneratedEvaluatorRef + assertion helpers Adds the provider-agnostic surface for referencing a pre-existing rubric evaluator and gating CI on per-item / per-dimension thresholds. Mirrors Python PR #6101 commits e5830dd7f (ref type) and 4bc60462d (asserts). - GeneratedEvaluatorRef: name + optional version/display-name, plus a Latest(name) factory for versionless refs (discouraged for CI; consumers should warn at run time). - AgentEvaluationResults.AssertScoreAtLeast: walks DetailedItems[].Scores, optionally filtered by evaluator name, recurses into SubResults. - AgentEvaluationResults.AssertDimensionScoreAtLeast: walks each score's Dimensions list, skips non-applicable dimensions by default, supports requireApplicable to flip that, recurses into SubResults. - AgentEvaluationResults.AssertNoFailedItems: walks DetailedItems for fail/error statuses, recurses into SubResults. All helpers throw InvalidOperationException (matches existing AssertAllPassed). Truncates offender lists to the first 5 with a '+N more' suffix to keep CI output readable, mirroring the Python helpers. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * .NET: feat(foundry-evals): accept GeneratedEvaluatorRef in evaluators= Adds FoundryEvaluatorSpec, a readonly-struct union with implicit conversions from both string and GeneratedEvaluatorRef so call sites can mix built-in evaluator names with rubric evaluator references: var evals = new FoundryEvals( projectClient, model, new GeneratedEvaluatorRef("policy-rubric", "3"), FoundryEvals.Relevance, FoundryEvals.Coherence); FoundryEvals constructors (3 overloads), EvaluateTracesAsync, and EvaluateFoundryTargetAsync now take FoundryEvaluatorSpec[]/params instead of string[]/params. Existing call sites using string literals or string[] keep working unchanged via implicit conversion. FoundryEvalConverter.BuildTestingCriteria emits the documented Foundry wire format for rubric refs: { "type": "azure_ai_evaluator", "name": <DisplayName ?? Name>, "evaluator_name": <Name>, "evaluator_version": <Version>, // omitted when null "initialization_parameters": { "deployment_name": <model> }, "data_mapping": { conversation arrays, optional tool_definitions } } WireTestingCriterion gains an optional EvaluatorVersion field. Rubric refs are preserved through FilterToolEvaluators (tool-aware but not tool-required) and ignored by FindMissingGroundTruthEvaluators. A versionless ref emits a Trace.TraceWarning at criterion-build time so CI authors notice the floating version (mirrors the Python warning). Adds 6 new Foundry unit tests (3 BuildTestingCriteria rubric paths, 1 FindMissingGroundTruthEvaluators, 1 FilterToolEvaluators preservation, 1 mixed-order). 369/369 Foundry tests pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * .NET: feat(foundry-evals): parse rubric dimension_scores into RubricScore Adds FoundryEvals.ParseRubricScores, called per result inside ParseDetailedItem. Each EvalScoreResult now populates Dimensions when the evaluator's sample carries a rubric breakdown. Accepts three shapes for forward compatibility with provider SDK iterations: 1. sample.properties.dimension_scores (canonical Foundry runtime shape) 2. sample.properties.rubric_scores (preview/legacy key) 3. top-level sample.dimension_scores / sample.rubric_scores (defensive fallback) Entries missing 'id', 'weight', or 'applicable' are skipped without invalidating well-formed siblings. Non-applicable dimensions may omit 'score' (parsed as null). Adds 6 unit tests covering canonical and legacy keys, top-level fallback, no-match returns null, malformed-entry skipping, and the non-applicable null-score path. 375/375 Foundry tests pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * .NET: feat(samples): Evaluation_FoundryRubric end-to-end sample Adds dotnet/samples/05-end-to-end/Evaluation/Evaluation_FoundryRubric mirroring the Python evaluate_with_rubric_sample.py: - Fetches a pre-existing Foundry agent via AgentAdministrationClient (GetAgentAsync for latest, GetAgentVersionAsync when FOUNDRY_AGENT_VERSION is pinned). - References a rubric evaluator by GeneratedEvaluatorRef(name, version); falls back to GeneratedEvaluatorRef.Latest(name) with the documented floating-version warning. - Mixes the rubric with FoundryEvals.Relevance and FoundryEvals.Coherence in a single FoundryEvals run (implicit string-and-ref conversion). - Prints per-dimension breakdowns from EvalScoreResult.Dimensions for each item. - Demonstrates a CI quality gate with AssertDimensionScoreAtLeast("general_quality", 3.0). Documents the FOUNDRY_PROJECT_ENDPOINT footgun (must be project-scoped URL .../api/projects/<project>, not the bare Azure OpenAI endpoint) and the Eval-Definition-vs-Rubric-Evaluator distinction in the README. Ships a .env.example with the FOUNDRY_* variables. Registers the project in agent-framework-dotnet.slnx and cross-links from the sibling Evaluation_Multimodal / Evaluation_ExpectedOutputs READMEs. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(foundry-evals): harden FoundryEvals public surface for review Address PR #6267 review comments on the .NET FoundryEvals integration: - Add source-compat overloads accepting `string[] evaluators` for `FoundryEvals` ctor, `EvaluateTracesAsync`, and `EvaluateFoundryTargetAsync` so existing callers passing string arrays keep compiling unchanged. New overloads forward via a private `ToSpecs` helper that wraps each name through the implicit `string -> FoundryEvaluatorSpec` conversion. - Guard against `default(FoundryEvaluatorSpec)` entries (both `BuiltinName` and `GeneratedRef` null) that would NRE the downstream converter. Adds `FoundryEvaluatorSpec.IsValid` / `EnsureValid` plus an internal `EnsureAllSpecsValid` helper, wired into the main ctor and both static evaluation entry points. - Add 6 unit tests covering the new validation surface. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(sample): set ExitCode=1 when rubric dimension gate trips PR #6267 review comment: the FoundryRubric sample swallowed the AssertDimensionScoreAtLeast failure, so a CI run that included it as a quality gate would still exit 0 even when the rubric regressed. Set `System.Environment.ExitCode = 1` in the catch so CI fails while still letting the rest of the sample's logging complete cleanly. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(foundry-evals): search typed Sample directly for rubric scores PR #6267 review comment: `_extract_rubric_scores` only searched the `properties` dict when the sample exposed one. When the Azure AI Projects typed SDK returns a Sample object that puts `dimension_scores` / `rubric_scores` directly on the instance (no `properties` wrapper), we missed them and surfaced no per-dimension scores. Add an `else: containers.append(sample)` branch so non-dict typed samples are also inspected for the score keys. Covered by two new tests: one with `dimension_scores` directly on a typed Sample without a `properties` wrapper, and one with the legacy `rubric_scores` key in the same shape. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * test(evals): cover assert_score_at_least and assert_no_failed_items PR #6267 review comments: both assertion helpers shipped without unit tests. Add `TestAssertScoreAtLeast` (above threshold, below w/ offenders, evaluator filter, sub_results recursion) and `TestAssertNoFailedItems` (all passing, failed/errored statuses, sub_results recursion) with a shared `_score_results` fixture builder. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(samples): remove dead rubric-evaluator doc link from FoundryRubric sample The Azure AI Foundry rubric evaluator concept doc page has not yet been published, so the link in the sample README and Program.cs comment 404s. Drop the references until the upstream doc is live. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Potential fix for pull request finding Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> * Address PR 6267 review nits Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Ben Thomas <25218250+alliscode@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> |
||
|
|
dc445592ed |
Python: [BREAKING] Port FileMemoryProvider and integrate FileMemoryProvider & FileAccess into the harness agent (#6547)
* Port FileMemoryProvider to python and integrate it and FileAccessProvider into the harness * Address PR comments * Address PR comments * Create FileSystemAgentFileStore root lazily on first write Construction no longer calls mkdir, so building a store (and therefore a default create_harness_agent, which wires default file-memory and file-access stores under the CWD) performs no filesystem writes and does not fail in read-only working directories. The root directory is created on the first write_file / create_directory call; all read/list/search operations already tolerate a missing root. Updates docstrings and adds a regression test. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix typing * Fixing typing errors --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
6e95517659 |
Python: Split type checkers by target (pyright source, 5 checkers on tests/samples) (#6443)
* Python: Split type checkers by target (pyright source, 5 checkers on tests/samples) Rework the typing setup along the lines of the 'too many type checkers' approach: - Pyright (strict) is now the sole source-code type checker; mypy is removed from source and its [tool.mypy] block becomes a relaxed profile used only for tests/samples. - Tests are checked by all five checkers (pyright relaxed, mypy, pyrefly, ty, zuban); samples by pyright, pyrefly, and ty. All run in a relaxed/ basic profile so authors aren't forced into over-annotation. - Add pyrightconfig.tests.json and bump sample pyright configs to basic. - Unify test/sample typing onto the same parallel fan-out used by source pyright via run_command_items in task_runner.py. - Make version-conditional imports symmetric: keep or drop the '# type: ignore' on both branches so results match across interpreter versions (local vs CI). - Update SKILL.md, DEV_SETUP.md, and CODING_STANDARD.md for the five gating checkers and pyright on source+tests+samples. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Fix merge regressions from main (typing + runtime) Merging main into the type-checker split branch surfaced regressions that the new five-checker test suite and unit tests caught: Runtime fixes: - anthropic: restore the dropped `cache_read_input_token_count` mapping in _parse_usage_from_anthropic (lost during merge conflict resolution). - gemini: _get_function_calling_mode test helper returned str(enum) ('FunctionCallingConfigMode.AUTO') instead of the enum value ('AUTO'). - openai: _response_id_from_token test helper was an infinite self-recursion; return token['response_id']. - orchestrations: reset output_events per approval iteration so the terminal output assertion counts only the final run. - core: drop a stale duplicate harness test whose message ('non-negative') contradicted the source ('positive'). - purview: import PolicyLocation/PolicyScope/ProtectionScopeActivities/ ExecutionMode used by the processor tests. Type-checker fixes (tests, relaxed profile): - core: pyright/mypy/pyrefly/ty/zuban green-ups across the harness, MCP, observability and types tests. - anthropic/openai: route provider-namespaced UsageDetails keys through a dict cast (extra_items TypedDict unsupported by mypy/ty). - purview: typed model constructors and cache-mock casts. - ag-ui: annotate WorkflowContext[Any, Any] so yield_output accepts test payloads, guard Optional forwarded_props, and ty-ignore intentional bad args. Source pyright (sole source checker) flagged unnecessary ignores newly introduced by merged code in core _tools.py and declarative _declarative_base.py. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Isolate per-package mypy cache in test-typing fan-out The parallel test-typing fan-out runs many mypy processes concurrently, all defaulting to a single shared ./.mypy_cache. Concurrent writes corrupt the cache and mypy aborts with INTERNAL ERROR (intermittently, depending on worker timing) -- which is why CI's Test Typing job failed on a shifting set of packages while a single-package run was fine. Give each mypy invocation an isolated cache dir keyed by its target paths so incremental caching still works per package without races. Other checkers (zuban/pyrefly/ty/pyright) maintain their own caches and are unaffected. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Make lab pyright-only on source (drop source mypy) Lab was the last package still running mypy on its source code, requiring mypy-only `# type: ignore` comments that pyright (the sole source checker everywhere else) flags as unnecessary. Align lab with the rest of the monorepo: - Remove the lab source mypy poe tasks (mypy-gaia/lightning/tau2) and the now-dead strict [tool.mypy] config block. - Drop the 'Run lab mypy' CI step; lab source is type-checked by pyright only. Lab tests remain covered by the workspace test-typing fan-out (mypy, pyrefly, ty, zuban, pyright over tests using the relaxed root config). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Fix test-typing regressions from latest main merge A fresh merge from main brought in new test code never run under the five-checker test-typing suite. Green up across the affected packages: - core: narrow Optional span.attributes with 'and' guards in span filters and assert+cast the json.loads(...attributes[...]) reads (test_observability); match the existing as_agent ignore on the protocol-typed fixture (test_clients). - openai: align new streaming tests with the established chat_options dict pattern (ChatOptions TypedDict isn't assignable to dict), route Optional .annotations[0] access through a small _first_annotation helper (mirrors the file's assert-not-None convention), and annotate a mapped ResponseStream. - foundry_hosting: annotate error: dict[str, Any] = body.get(...) or {} (zuban needs the annotation). - foundry: narrow ignores for the live AIProjectClient credential arg (pyrefly) and connections.get_default (zuban) SDK type gaps. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * updated pyright version * pyright fix * Python: Fix source typing for pyright 1.1.410 Pyright 1.1.410 tightened several checks. Apply the same source fixes as upstream PR #6275: - anthropic: import AsyncAnthropicBedrock from anthropic.lib.bedrock and AsyncAnthropicVertex from anthropic.lib.vertex (no longer re-exported from the anthropic top-level package -> reportPrivateImportUsage). - core _types.py: cast the transform-hook result to UpdateT (reportAssignmentType). - core _workflows/_events.py: annotate the @contextmanager helper as Generator[None] instead of Iterator[None] (reportDeprecated). - redis: build the combined filter expression with an explicit loop instead of reduce(and_, ...), which pyright could no longer fully type (drops the now unused functools.reduce / operator.and_ imports). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Accept plain-text body in Azure Functions workflow/run endpoint The workflow_orchestrator already accepts plain strings as well as JSON objects via context.get_input(), but the start_workflow_orchestration HTTP handler only accepted JSON and returned 400 for any non-JSON body. This made the functions integration tests that POST text/plain to /api/workflow/run (e.g. test_09_workflow_shared_state) fail consistently with 400 != 202. Fall back to the raw request body (decoded as UTF-8) when the body is not JSON, rejecting only a truly empty body. The JSON path is unchanged. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
c22fc8d653 |
Build(deps): Bump esbuild, @tailwindcss/vite, @vitejs/plugin-react and vite (#6503)
Removes [esbuild](https://github.com/evanw/esbuild). It's no longer used after updating ancestor dependencies [esbuild](https://github.com/evanw/esbuild), [@tailwindcss/vite](https://github.com/tailwindlabs/tailwindcss/tree/HEAD/packages/@tailwindcss-vite), [@vitejs/plugin-react](https://github.com/vitejs/vite-plugin-react/tree/HEAD/packages/plugin-react) and [vite](https://github.com/vitejs/vite/tree/HEAD/packages/vite). These dependencies need to be updated together. Removes `esbuild` Updates `@tailwindcss/vite` from 4.1.12 to 4.3.1 - [Release notes](https://github.com/tailwindlabs/tailwindcss/releases) - [Changelog](https://github.com/tailwindlabs/tailwindcss/blob/main/CHANGELOG.md) - [Commits](https://github.com/tailwindlabs/tailwindcss/commits/v4.3.1/packages/@tailwindcss-vite) Updates `@vitejs/plugin-react` from 5.0.1 to 5.2.0 - [Release notes](https://github.com/vitejs/vite-plugin-react/releases) - [Changelog](https://github.com/vitejs/vite-plugin-react/blob/plugin-react@5.2.0/packages/plugin-react/CHANGELOG.md) - [Commits](https://github.com/vitejs/vite-plugin-react/commits/plugin-react@5.2.0/packages/plugin-react) Updates `vite` from 7.3.2 to 8.0.16 - [Release notes](https://github.com/vitejs/vite/releases) - [Changelog](https://github.com/vitejs/vite/blob/main/packages/vite/CHANGELOG.md) - [Commits](https://github.com/vitejs/vite/commits/v8.0.16/packages/vite) --- updated-dependencies: - dependency-name: esbuild dependency-version: dependency-type: indirect - dependency-name: "@tailwindcss/vite" dependency-version: 4.3.1 dependency-type: direct:production - dependency-name: "@vitejs/plugin-react" dependency-version: 5.2.0 dependency-type: direct:development - dependency-name: vite dependency-version: 8.0.16 dependency-type: direct:development ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
3d46595111 |
Python: Bump prek from 0.4.3 to 0.4.5 in /python (#6527)
* Bump prek from 0.4.3 to 0.4.5 in /python Bumps [prek](https://github.com/j178/prek) from 0.4.3 to 0.4.5. - [Release notes](https://github.com/j178/prek/releases) - [Changelog](https://github.com/j178/prek/blob/master/CHANGELOG.md) - [Commits](https://github.com/j178/prek/compare/v0.4.3...v0.4.5) --- updated-dependencies: - dependency-name: prek dependency-version: 0.4.5 dependency-type: direct:development update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> * fix python workspace prek pin mismatch --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> |
||
|
|
205f7bcca8 |
Python: Bump pytest from 9.0.3 to 9.1.0 across /python workspace (#6524)
* Bump pytest from 9.0.3 to 9.1.0 in /python Bumps [pytest](https://github.com/pytest-dev/pytest) from 9.0.3 to 9.1.0. - [Release notes](https://github.com/pytest-dev/pytest/releases) - [Changelog](https://github.com/pytest-dev/pytest/blob/main/CHANGELOG.rst) - [Commits](https://github.com/pytest-dev/pytest/compare/9.0.3...9.1.0) --- updated-dependencies: - dependency-name: pytest dependency-version: 9.1.0 dependency-type: direct:development update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> * Fix Python workspace pytest pin mismatch --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> |
||
|
|
b55992bb67 |
Bump Python package versions for 1.9.0 release (#6583)
Selective, CHANGELOG-driven version bumps for the 2026-06-18 release. Released tier: agent-framework-core and the root agent-framework go to 1.9.0 (minor). Core ships new public APIs (agent-loop middleware, tool-approval middleware and harness integration, shell-tool harness integration, AG-UI thread snapshot persistence, context-provider telemetry) plus two behavioral breaking changes on evolving surfaces: MCP sampling now denies server-initiated requests by default, and the FileAccess tools were aligned with the .NET implementation. These are treated as within-1.x changes because every package caps core at <2; a major bump would require rewriting those caps. The foundry and openai packages go to 1.8.2 (patch, bug fixes only). The root agent-framework-core[all] pin was moved to 1.9.0 in lockstep with core. Release-candidate tier: ag-ui to 1.0.0rc5 and declarative to 1.0.0rc2 for their respective changes. orchestrations is promoted to stable 1.0.0; PACKAGE_STATUS and the README install hint were updated accordingly. Prerelease tier (new Pacific date stamp 260618): anthropic (beta), azure-contentunderstanding (alpha) and foundry-hosting (alpha). No beta cohort bump was applied; only packages with changes this cycle were stamped. Dependency floors: following the established convention, the core floor was raised to >=1.9.0 on every non-core package bumped this cycle, preserving the existing <2 upper bound. Also resolves two pre-existing failures in the dependency-bounds validator that are unrelated to the version bumps. Hosted-environment detection now catches a bare ImportError so optional Foundry hosting probing cannot crash user-agent setup. The harness shell-tool integration, which lazily imports the separate agent-framework-tools package to avoid a circular runtime dependency, is now type-checked and tested in isolated environments via a core dev dependency-group, with the shell-tool tests guarded to skip when that package is absent. |
||
|
|
d7e63d7d0e |
Fix Foundry aiohttp dependency (#6567)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
f59d5c67d8 |
Python: Adopt azure-ai-contentunderstanding to_llm_input in CU context provider (#5796)
* Refactor DocumentEntry model and update result handling - Changed the type of `result` in DocumentEntry from dict to str to store LLM-ready text. - Introduced `search_payload` in DocumentEntry for optional alternate rendering. - Updated FileSearchConfig to include `include_fields` option for vector store uploads. - Modified tests to reflect changes in DocumentEntry and FileSearchConfig. - Adjusted integration tests to validate new result structure and rendering. - Removed legacy format_result tests as rendering is now handled by the SDK. * Potential fix for pull request finding Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> * Add test to ensure page markers are preserved in LLM input Co-authored-by: Copilot <copilot@github.com> * fix(cu-context-provider): scope LLMStats telemetry filter to rai_warnings block Address PR #5796 review comment: the previous defensive scrubber ran a global regex substitution over the full rendered string, so any markdown body bullet shaped like '- LLMStats: ...' would also be silently deleted. Add a _strip_rai_telemetry helper that confines the substitution to the front-matter rai_warnings: YAML sub-block, leaving the body verbatim. Cover the new behavior with three tests (scoped strip, body preservation, and no-op branches). * Sync uv.lock with azure-ai-contentunderstanding>=1.2.0b1 dependency bump * Python: Drop search_payload/include_fields, single to_llm_input rendering (CU context provider) Address PR #5796 review: remove the redundant search_payload field and _render_search_payload helper, drop the include_fields opt-in (already covered by output_sections), rename _resolve_pending_tokens -> _resolve_pending_analysis, and have _upload_to_vector_store read entry['result'] directly. * Python: Adopt SDK 1.2.0b2 LLMStats filtering, drop local workaround (CU context provider) azure-ai-contentunderstanding 1.2.0b2 filters LLMStats telemetry from rai_warnings and emits InputPageNumber page markers in to_llm_input, so the provider's local defense is redundant. - Bump dependency to azure-ai-contentunderstanding>=1.2.0b2 (re-lock uv.lock) - Remove _strip_rai_telemetry and its two regexes; _render_for_llm now returns to_llm_input(...) directly - Delete 4 workaround unit tests for the removed helper --------- Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> Co-authored-by: Copilot <copilot@github.com> Co-authored-by: changjian-wang <v-changjwang@microsoft.com> Co-authored-by: aluneth <wangchangjian1130@163.com> |
||
|
|
4ff952e100 |
Python: Capture context provider instructions in agent telemetry (#6515)
* Fix agent instructions telemetry Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Simplify agent instructions telemetry guard Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix observability mypy cast Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
8e10c0399a |
Python: Remove unsupported as_agent function_invocation_configuration (#6520)
* Remove unsupported as_agent config parameter Fixes #6313 Remove the unsupported function_invocation_configuration parameter from BaseChatClient.as_agent(), which currently forwards an invalid kwarg into Agent.__init__(). This keeps the existing TypeError behavior for callers but changes the error source to the public API boundary, which we do not consider a breaking change. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix sample --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
bce2757477 | Foundry hosted agent responses emit failed events (#6502) | ||
|
|
0db9305625 |
Python: Integrate tool approval into the harness (#6522)
* Integrate auto tool-approval feature into harness * Potential fix for pull request finding Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> * Rename disable_tool_approval to disable_tool_auto_approval Addresses PR review feedback that the parameter name was unclear. The flag toggles the auto/standing tool-approval middleware. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
571cae426c |
Python: Fix Azure AI Search citation URLs (#6453)
* Fix Azure AI Search citation URLs Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Enrich MCP search citation metadata Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix Azure AI Search citation enrichment follow-ups Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address PR #6453 review comments Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * updated filter for paths * also updated python paths * reverted dotnet-format change --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: westey <164392973+westey-m@users.noreply.github.com> |
||
|
|
d7e8d2206d |
Python: Fix Python OTel usage detail attributes (#6493)
* fix python otel usage detail attributes Map cached/read/reasoning usage detail fields to standard OTel GenAI attributes while preserving provider-specific legacy keys. Add focused coverage for direct response spans, aggregated agent spans, and provider usage parsing. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * address usage detail review feedback Omit missing OpenAI Responses usage detail counts while preserving zero-valued counts. Record zero-valued token usage in OTel histograms and add regression coverage. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
d7027fc1f9 |
Python: [BREAKING] Align FileAccess tools with .NET — directory discovery and recursive search (#6476)
* Align FileAccess tools with .Net; add directory discovery and recursive search * Fix choices field description: spacing, line length, grammar Addresses PR review: separate concatenated string literals with proper spacing/newlines, wrap lines under the 120-char Ruff limit, and fix "doesn't" -> "don't". Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address PR comments --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |