Promote the GitHub Copilot package from release candidate (1.0.0rc4) to released (1.0.0): bump the version, switch the classifier to Production/Stable, update PACKAGE_STATUS.md, and drop the --pre install flag from the package and sample READMEs. Add a github-copilot-1.0.0 CHANGELOG section covering the promotion and the input-attachment forwarding shipped in this release. No core/root bump: this is a standalone package promotion and the core[all] extra references the package without a version pin.
Copilot-Session: f523064c-60b4-4d18-bf95-c16c5fda9126
* Python: Support prompt cache breakpoints for GPT-5.6 models in OpenAI clients
Add request-level prompt_cache_options to OpenAIChatOptions and
OpenAIChatCompletionOptions, and forward a per-part prompt_cache_breakpoint
from Content.additional_properties onto the content blocks each API supports.
Text parts that carry a breakpoint keep typed list content, since the
plain-string form cannot hold one; without a breakpoint the existing string
forms are unchanged.
* Clarify system-message content-shape comment
* Address review: SDK prompt cache types, private helper, add sample
Replace the custom PromptCacheOptions TypedDict with the openai SDK's own
types for each API, which raises the openai floor to 2.45.0 where those
types were introduced. Make the breakpoint helper private to the two chat
clients. Add a prompt caching sample with a README entry, and unquote the
helper's Content annotation so the pyupgrade hook passes.
* Guard the prompt cache options import for older openai versions
The SDK's PromptCacheOptions types only exist in openai 2.45.0 and
later, so each client falls back to a local mirror when the import
fails and the dependency floor stays at 2.25.0. A TYPE_CHECKING-only
import is not enough because the options classes are introspected with
get_type_hints() at runtime. Verified against openai 2.25.0: the
package imports, the fallback resolves, and part-level breakpoints
still work; sending the option itself requires 2.45.0, which the field
docstrings now note.
* Make the old-openai fallback for PromptCacheOptions deliberately empty
Assigning None instead, as suggested in review, trips pyright's
reportInvalidTypeForm on the field annotation (the symbol becomes
type | None after the try/except). An empty TypedDict gives the same
effect for users on older openai versions: any content they put in
prompt_cache_options is flagged by their type checker, since the
option cannot be sent on those versions anyway, while
get_type_hints() on the options classes keeps working at runtime.
* Guard prompt_cache_options at runtime instead of via an empty fallback type
The empty-TypedDict fallback flagged valid `prompt_cache_options` usage under
pyright on every openai version — including this PR's own
`client_prompt_caching.py` sample (`poe check -S`) — because pyright resolves the
try/except symbol to the fallback shape regardless of the installed openai, while
mypy/ty resolve the failed import to `Any` and never warn. So a type-only "warn on
old openai" signal is not achievable cleanly across type checkers.
Restore the faithful fallback (mirrors the SDK's `mode`/`ttl` shape) so the option
type-checks identically on every supported openai version, and add a runtime guard:
setting `prompt_cache_options` on openai < 2.45 now raises a clear
ChatClientInvalidRequestException instead of forwarding an unusable option to the
SDK. This keeps the option non-silent for all users regardless of type checker,
without forcing an openai upgrade. Adds tests covering the guard for both clients.
* Gate system/developer breakpoint shape on a real mapping value
The system/developer branch switched to list-form content whenever
prompt_cache_breakpoint was set to any non-None value, but the option is
only attached when the value is a mapping. A malformed value (e.g. a
string) therefore changed the message shape without adding a breakpoint.
Decide the shape from the built part instead, matching the user-role path.
* Bump Python package versions for 1.12.0 release
Bump packages represented in the 1.12.0 changelog, promote Foundry Hosting, Azure Content Understanding, Gemini, Mistral, Monty, and Tools to beta, and apply the requested beta cohort date stamp. Root and core move to 1.12.0, released and RC packages use their selected increments, alpha packages including Hosting MCP use the 260721 stamp, and core floors are raised only for proven consumers.
Copilot-Session: 2dd9980a-b869-4c16-8642-75b7a6d6ebdf
* fix version in readme
* Add Responses conversation ID changes to release notes
Include the breaking Hosting Responses conversation ID helper changes from #7234 in the Python 1.12.0 changelog.
Copilot-Session: 2dd9980a-b869-4c16-8642-75b7a6d6ebdf
* Python: feat: cross-session origin attribution on context messages
Add an optional origin_session_id parameter to SessionContext.extend_messages
that propagates into the existing _attribution payload on
Message.additional_properties. Downstream context observers can use it to
detect when a provider injects content stored under a different session than
the requesting one.
Populate the field from the harness memory consolidation pipeline
(_harness/_memory.py) when injected topics include contributions from
sessions other than the current one. Add a self-contained sample observer
under samples/02-agents/context_providers/cross_session_observer.py
demonstrating how to subscribe to the signal.
Backward-compatible: omitting the parameter preserves the existing
attribution shape exactly. Tests added in test_sessions.py and
test_harness_memory.py cover the new parameter, the harness cross-session
case, and the same-session case.
Motivated by Dai et al., Stateful Agent Backdoor (arXiv:2605.06158, May
2026), which specifically surveys MAF in section 6.1 / Table 10.
See #5914 for design discussion.
Surfaced during independent audit conducted by @finnoybu (Ken Tannenbaum, AEGIS Initiative); [MEDIUM, python/packages/core].
* Address cross-session attribution review feedback
* Address follow-up review feedback
* Address follow-up review comments
* Address origin attribution review feedback
---------
Co-authored-by: finnoybu <21694570+finnoybu@users.noreply.github.com>
* Python: Remove experimental marker from Skills API
Promote the Skills feature from experimental to stable, mirroring
.NET PR #6861. Removes the @experimental(SKILLS) decorators from the
skills APIs and the SKILLS ExperimentalFeature enum member, updates
tests and samples accordingly. MCP skills (MCP_SKILLS) remain
experimental, matching the .NET change.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Add experimental-stage assertions for MCP skills types
Guard MCPSkill, MCPSkillResource, and MCPSkillsSource against accidental
promotion by asserting their docstring warning block and
__feature_stage__/__feature_id__ metadata remain experimental (MCP_SKILLS).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Remove redundant stable-stage test for Skills API
Drop TestSkillsStableStage: asserting the absence of experimental
markers on a released API is not meaningful, and the feature-stage
decorator machinery is already covered by test_feature_stage.py.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Python: add ATR validation FunctionMiddleware sample (execution-boundary validation, #5366)
Adds python/samples/02-agents/middleware/atr_validation_middleware.py: a
FunctionMiddleware that validates tool arguments at the execution boundary and
raises MiddlewareTermination before call_next() when they match an attack
pattern, so the tool never runs. This is the deterministic, single-enforcement-
point pattern named in #5366 and answers its open follow-up about a recommended
validation-at-execution-boundary sample.
The check is a small self-contained deny-list mirroring Agent Threat Rules (ATR)
intent (prompt injection, exfiltration, credential access in tool args); a
docstring notes how to swap in the full open ruleset via pyatr. No external
dependency, so the sample stays import-clean.
Updates the middleware README Files table.
Signed-off-by: Adam Lin <adam@agentthreatrule.org>
* Python: Samples: run the real ATR engine in atr_validation_middleware
Address review on #6528:
- Load and run the real ATR ruleset via pyatr (ATREngine + AgentEvent
tool_call event) instead of re-implementing a regex deny-list; the
built-in deny-list is now only a fallback when pyatr is not installed.
- Add re.DOTALL (and a whole-text scan) to the fallback patterns so
multiline injection payloads are not missed.
- Move load_dotenv() into main() so importing the module has no side
effects.
- Route the middleware block/allow messages through a module logger
instead of print().
- Include the matched ATR rule id in the log and in the
MiddlewareTermination message for auditability.
- Update the middleware README entry to match.
* fix(samples): make ATR validation middleware pass ty/pyrefly typing CI
Resolve the three type-checker errors flagged on the samples typing jobs
(ty + pyrefly, reportMissingImports/reportAttributeAccessIssue via pyright):
- pyatr is an optional, unstubbed runtime dependency that is not installed
in the typing CI env; mark its imports with `# type: ignore` so the
unresolved-import error is suppressed while keeping the graceful
ImportError -> deny-list fallback intact.
- Replace the function-attribute engine cache
(`_detect_with_atr._engine`), which ty/pyrefly reject, with a clean
`functools.lru_cache`-backed `_load_atr_engine()` loader.
- Type the argument-scanning helpers to accept the real
`FunctionInvocationContext.arguments` type (`BaseModel | Mapping[str, Any]`)
and normalise a pydantic model via `model_dump()` before scanning, fixing
the invalid-argument-type error.
ty / pyrefly / pyright (samples config) / ruff check + format all clean on
the file; runtime block/allow behaviour verified for both dict and BaseModel
arguments.
* Python: Samples: simplify ATR middleware to plain pyatr import
Address review feedback (@eavanvalkenburg): now that the sample runs the
real pyatr engine, drop the optional-import scaffolding.
- Add a dependency header declaring pyatr (pip install pyatr).
- Switch to a plain top-level `import pyatr` and remove the
try/except ImportError fallback path.
- Remove the regex deny-list (_FALLBACK_PATTERNS, _detect_with_fallback);
keep 2-3 representative pattern shapes inline as a reference comment so
readers still see the kind of rules ATR encodes. Detection is now a
single straight-line engine call.
- Keep the prior typing fixes: `# type: ignore` on the pyatr import
(unstubbed, absent in the typing CI env), the functools.lru_cache
engine loader, and the BaseModel | Mapping[str, Any] signatures.
* fix: use PEP 723 inline script metadata for sample dependencies
---------
Signed-off-by: Adam Lin <adam@agentthreatrule.org>
Co-authored-by: eeee2345 <eeee2345@users.noreply.github.com>
Co-authored-by: Eduard van Valkenburg <eavanvalkenburg@users.noreply.github.com>
* Python: Add SkillsSourceContext to SkillsSource.get_skills
Thread an invocation context (agent + optional session) through the skill
source pipeline so sources and decorators can make context-aware decisions.
- Add frozen, experimental SkillsSourceContext(agent, session).
- Change SkillsSource.get_skills and all sources/decorators to accept and
forward the context.
- Make FilteringSkillsSource predicate context-aware: (skill, context) -> bool.
- Add optional cache_isolation_key_selector to CachingSkillsSource for
per-key cache isolation (None keeps the shared-bucket behavior).
- Build the context in SkillsProvider from before_run agent/session.
- Update foundry_hosting toolbox source, exports, tests, and docs.
Python port of .NET PR #6797.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Python: Clarify skills source docstring examples
Address PR review: docstring examples referenced `context` without
constructing it. Add a `SkillsSourceContext` construction line (with a
placeholder agent) to each source example and a note that the provider
normally supplies it. Use `source_context` in the FilteringSkillsSource
example to avoid clashing with the predicate's `context` parameter.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Python: Fix CI type errors and skill_filtering sample predicate
Address CI failures from the SkillsSourceContext change:
- Update the skill_filtering sample to the 2-arg predicate signature
(skill, context); the old 1-arg lambda would fail at runtime.
- Replace ad-hoc _StubAgent test stubs with the shared MockAgent /
MockAgentSession from conftest so all type checkers (incl. ty) accept
the SupportsAgentRun-typed agent. Add a small _NamedMockAgent subclass
for tests needing distinct agent names, and drop now-unnecessary
attr-defined ignores.
- Use cast(SupportsAgentRun, ...) in foundry_hosting tests, which have no
shared mock infrastructure.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Python: Make SkillsProvider caching safe-by-default; clarify context docstrings
Address PR review comments:
- Do not auto-wrap a caller-supplied SkillsSource in the provider's default
CachingSkillsSource. A shared, unkeyed cache around a context-aware source
replays the first invocation's skills for later SkillsSourceContexts,
leaking skills across agents/tenants. Default caching now applies only to
the built-in, context-independent file/in-memory leaf sources
(Deduplicating(Caching(leaf))), matching the .NET provider. Callers who
want caching on a custom pipeline compose CachingSkillsSource (optionally
with a cache_isolation_key_selector) themselves. disable_caching now only
affects the built-in leaves. Adds a leak-prevention test.
- Reword the misleading "Unused by this source" context docstrings on the
File/InMemory/MCP sources: the param is part of the get_skills contract;
these sources just return the same skills regardless of context.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Python: align GitHub Copilot approval to SDK on_pre_tool_use hook
Replace the bespoke on_function_approval enforcement in the GitHub Copilot provider with the Copilot SDK's native on_pre_tool_use hook. When no caller hook is supplied, a default hook returns 'ask' for approval_mode='always_require' tools (routed to on_permission_request) and defers others; a caller-supplied on_pre_tool_use takes precedence and logs a warning for any unenforced approval tool.
Fixes#6746
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Fix type-checker errors and restore load_dotenv in sample
Use a complete PreToolUseHookInput in on_pre_tool_use hook tests so pyright/pyrefly/ty/zuban no longer report missing required TypedDict keys. Restore load_dotenv() in the function-approval sample for consistency with the other GitHub Copilot samples (PR review feedback).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Deprecate on_function_approval instead of removing it
Per PR review feedback, keep the on_function_approval callback working (still enforced in the tool handler for approval_mode='always_require' tools) but emit a DeprecationWarning at construction, so existing users get a signal rather than a silent behavior change. The default on_pre_tool_use ask-hook is not installed when on_function_approval is set, avoiding double-gating. Precedence: user on_pre_tool_use > on_function_approval > default ask-hook. Adds tests for the deprecated path and documents it in the package README.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Make on_function_approval and on_pre_tool_use mutually exclusive
Per automated review feedback, instead of a precedence ordering between the deprecated on_function_approval callback and the new on_pre_tool_use hook (which silently double-gated when both were set), raise ValueError if both are supplied - at construction (both in default_options) or per run (per-run on_pre_tool_use with a construction-time on_function_approval). This matches the repo convention for deprecated-vs-new params (see _workflows/_workflow.py) and removes the flag-threading. Updates tests and the package README.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Python: [BREAKING] Make all SkillsProvider tools require approval by default
All tools exposed by SkillsProvider (load_skill, read_skill_resource,
run_skill_script) now require approval by default. Previously only
run_skill_script could be gated, and only when require_script_approval=True.
- Register all three tools with approval_mode="always_require"
- Add read_only_tools_auto_approval_rule and all_tools_auto_approval_rule
static rules plus tool-name constants (mirrors FileAccessProvider)
- Remove the require_script_approval option from __init__ and from_paths
- Add skills_auto_approval sample; update script_approval sample/docs
Closes#6728
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Address PR review: batch skill approval responses and tidy sample
- Collect a response for every approval request and send them in a single
agent.run so the approval loop always makes progress (no infinite loop when
a request lacks a function_call); reject non-function requests instead of
skipping them. Applied to both the skills_auto_approval and script_approval
samples.
- Extract ToolApprovalMiddleware into a local variable in skills_auto_approval
for readability.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Address PR review: add approval handling to remaining skills samples
The secure-by-default change makes all SkillsProvider tools require approval,
which left the other skills samples emitting approval requests instead of the
documented answers. Add ToolApprovalMiddleware with the all-tools auto-approval
rule (and a session, which the middleware requires) so these samples run
unattended as before:
- code_defined_skill, file_based_skill, class_based_skill, mixed_skills,
skill_filtering, mcp_based_skill
- providers/foundry/foundry_chat_client_with_toolbox_skills
The dedicated script_approval (manual) and skills_auto_approval (selective)
samples continue to demonstrate interactive approval handling.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Address PR review: simplify "host approval" wording to "approval"
Apply maintainer suggestions dropping "host" from the skill-approval
docstrings, and align the matching SkillsProvider docstring/AGENTS.md note for
consistency.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Update agent-framework-azure-ai-search to work across the stable/GA azure-search-documents SDK (12.0.0, api-version 2026-04-01) and the preview SDK (12.1.0b1, api-version 2026-05-01-preview) for both semantic and agentic modes.
- Bump the dependency to azure-search-documents>=12.0.0,<13 and the package to 1.0.0b260618.
- Add an api_version parameter (threaded into SearchClient, SearchIndexClient, and KnowledgeBaseRetrievalClient) plus STABLE_API_VERSION/PREVIEW_API_VERSION constants, re-exported from agent_framework.azure.
- Auto-detect preview-only agentic features (output mode, low/medium reasoning effort) via _preview_features_active(), which requires both the preview SDK and a preview api-version; defaults (extractive + minimal) work on both channels and preview-only options raise an actionable error otherwise.
- Make knowledge-base imports SDK-version resilient and fix the 12.x surface (k -> k_nearest_neighbors, defensive additional_properties).
- Update tests (pass on both SDKs), docs, samples, CHANGELOG, and uv.lock.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Add samples for the harness blog part 2
* Address PR comments
* Fix blog links.
* Address PR comments
* Fix bug where mode was incorrectly defaulted when reading the mode before the first run.
* Add reference to new sample readme
* Python: add GitHub MCP security label sample
* modified samples to create devui auth token, support debugging with security, and change context label only using the labels of unhidden result from tools
* FIDES: secure MCP labeling, _meta IFC parsing, and docs updates
* FIDES: secure MCP labeling, _meta IFC parsing, and docs updates
* modified docs
* fixed PR comments, simplified github_mcp example
* commented github_mcp example
* remove the parse_github_mcp_labels and fix the user_identity label propogation
* fix: use standard GitHub MCP endpoint with X-MCP-Features: ifc_labels instead of /insiders
- Switch MCP_URL from /mcp/insiders to /mcp/ in github_mcp_example.py
- Add MCP_HEADERS constant with X-MCP-Features: ifc_labels to opt-in to
server-side IFC label emission in _meta payloads
- Fix SecureMCPToolProxy to pass headers via httpx.AsyncClient so they are
included on session.initialize(), not just on tool calls (was causing 401
to silently surface as anyio cancel-scope CancelledError)
- Update README, FIDES_DEVELOPER_GUIDE, FIDES_IMPLEMENTATION_SUMMARY, and
0024-prompt-injection-defense.md to remove all /insiders references
* address PR comments
* Simplify GitHub MCP security sample to DevUI-only; document SecureAgentConfig quarantine client global behavior
* minor PR comments
* fixing failed checks
* fixing failed checks
---------
Co-authored-by: Eduard van Valkenburg <eavanvalkenburg@users.noreply.github.com>
* Require approvals for file-access and expose auto approval funcs for it
* Scope file-access auto-approval rules to local tools; fix base-Agent sample
Address PR #6599 review feedback:
- read_only/all_tools auto-approval rules now reject any call carrying a
server_label so they stay scoped to FileAccessProvider's local tools and
never auto-approve a same-named hosted tool.
- Expand the FileAccessProvider docstring to explain the runtime effect of
approval_mode="always_require" and point to ToolApprovalMiddleware /
create_harness_agent.
- Fix the base-Agent file_access_data_processing sample, which would otherwise
stop executing file tools under the new always_require defaults, by adding
ToolApprovalMiddleware with all_tools_auto_approval_rule.
- Add tests covering hosted (server_label) calls and update docs.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Clean up comments
* Update sample after merge
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Add samples for harness blog post part 1
* Add readme for python samples
* Update python instructions to match dotnet instructions
* Address PR comments
* Add link to blog posts
* Fix blog post naming.
* Add more blog post links
* Port FileMemoryProvider to python and integrate it and FileAccessProvider into the harness
* Address PR comments
* Address PR comments
* Create FileSystemAgentFileStore root lazily on first write
Construction no longer calls mkdir, so building a store (and therefore a
default create_harness_agent, which wires default file-memory and file-access
stores under the CWD) performs no filesystem writes and does not fail in
read-only working directories. The root directory is created on the first
write_file / create_directory call; all read/list/search operations already
tolerate a missing root. Updates docstrings and adds a regression test.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Fix typing
* Fixing typing errors
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Python: Split type checkers by target (pyright source, 5 checkers on tests/samples)
Rework the typing setup along the lines of the 'too many type checkers'
approach:
- Pyright (strict) is now the sole source-code type checker; mypy is
removed from source and its [tool.mypy] block becomes a relaxed profile
used only for tests/samples.
- Tests are checked by all five checkers (pyright relaxed, mypy, pyrefly,
ty, zuban); samples by pyright, pyrefly, and ty. All run in a relaxed/
basic profile so authors aren't forced into over-annotation.
- Add pyrightconfig.tests.json and bump sample pyright configs to basic.
- Unify test/sample typing onto the same parallel fan-out used by source
pyright via run_command_items in task_runner.py.
- Make version-conditional imports symmetric: keep or drop the
'# type: ignore' on both branches so results match across interpreter
versions (local vs CI).
- Update SKILL.md, DEV_SETUP.md, and CODING_STANDARD.md for the five
gating checkers and pyright on source+tests+samples.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Python: Fix merge regressions from main (typing + runtime)
Merging main into the type-checker split branch surfaced regressions that
the new five-checker test suite and unit tests caught:
Runtime fixes:
- anthropic: restore the dropped `cache_read_input_token_count` mapping in
_parse_usage_from_anthropic (lost during merge conflict resolution).
- gemini: _get_function_calling_mode test helper returned str(enum)
('FunctionCallingConfigMode.AUTO') instead of the enum value ('AUTO').
- openai: _response_id_from_token test helper was an infinite self-recursion;
return token['response_id'].
- orchestrations: reset output_events per approval iteration so the terminal
output assertion counts only the final run.
- core: drop a stale duplicate harness test whose message ('non-negative')
contradicted the source ('positive').
- purview: import PolicyLocation/PolicyScope/ProtectionScopeActivities/
ExecutionMode used by the processor tests.
Type-checker fixes (tests, relaxed profile):
- core: pyright/mypy/pyrefly/ty/zuban green-ups across the harness, MCP,
observability and types tests.
- anthropic/openai: route provider-namespaced UsageDetails keys through a
dict cast (extra_items TypedDict unsupported by mypy/ty).
- purview: typed model constructors and cache-mock casts.
- ag-ui: annotate WorkflowContext[Any, Any] so yield_output accepts test
payloads, guard Optional forwarded_props, and ty-ignore intentional bad args.
Source pyright (sole source checker) flagged unnecessary ignores newly
introduced by merged code in core _tools.py and declarative _declarative_base.py.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Python: Isolate per-package mypy cache in test-typing fan-out
The parallel test-typing fan-out runs many mypy processes concurrently,
all defaulting to a single shared ./.mypy_cache. Concurrent writes corrupt
the cache and mypy aborts with INTERNAL ERROR (intermittently, depending on
worker timing) -- which is why CI's Test Typing job failed on a shifting set
of packages while a single-package run was fine.
Give each mypy invocation an isolated cache dir keyed by its target paths so
incremental caching still works per package without races. Other checkers
(zuban/pyrefly/ty/pyright) maintain their own caches and are unaffected.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Python: Make lab pyright-only on source (drop source mypy)
Lab was the last package still running mypy on its source code, requiring
mypy-only `# type: ignore` comments that pyright (the sole source checker
everywhere else) flags as unnecessary. Align lab with the rest of the
monorepo:
- Remove the lab source mypy poe tasks (mypy-gaia/lightning/tau2) and the
now-dead strict [tool.mypy] config block.
- Drop the 'Run lab mypy' CI step; lab source is type-checked by pyright only.
Lab tests remain covered by the workspace test-typing fan-out (mypy, pyrefly,
ty, zuban, pyright over tests using the relaxed root config).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Python: Fix test-typing regressions from latest main merge
A fresh merge from main brought in new test code never run under the
five-checker test-typing suite. Green up across the affected packages:
- core: narrow Optional span.attributes with 'and' guards in span filters
and assert+cast the json.loads(...attributes[...]) reads (test_observability);
match the existing as_agent ignore on the protocol-typed fixture (test_clients).
- openai: align new streaming tests with the established chat_options dict
pattern (ChatOptions TypedDict isn't assignable to dict), route Optional
.annotations[0] access through a small _first_annotation helper (mirrors the
file's assert-not-None convention), and annotate a mapped ResponseStream.
- foundry_hosting: annotate error: dict[str, Any] = body.get(...) or {}
(zuban needs the annotation).
- foundry: narrow ignores for the live AIProjectClient credential arg (pyrefly)
and connections.get_default (zuban) SDK type gaps.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* updated pyright version
* pyright fix
* Python: Fix source typing for pyright 1.1.410
Pyright 1.1.410 tightened several checks. Apply the same source fixes as
upstream PR #6275:
- anthropic: import AsyncAnthropicBedrock from anthropic.lib.bedrock and
AsyncAnthropicVertex from anthropic.lib.vertex (no longer re-exported from
the anthropic top-level package -> reportPrivateImportUsage).
- core _types.py: cast the transform-hook result to UpdateT (reportAssignmentType).
- core _workflows/_events.py: annotate the @contextmanager helper as
Generator[None] instead of Iterator[None] (reportDeprecated).
- redis: build the combined filter expression with an explicit loop instead of
reduce(and_, ...), which pyright could no longer fully type (drops the now
unused functools.reduce / operator.and_ imports).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Python: Accept plain-text body in Azure Functions workflow/run endpoint
The workflow_orchestrator already accepts plain strings as well as JSON
objects via context.get_input(), but the start_workflow_orchestration HTTP
handler only accepted JSON and returned 400 for any non-JSON body. This made
the functions integration tests that POST text/plain to /api/workflow/run
(e.g. test_09_workflow_shared_state) fail consistently with 400 != 202.
Fall back to the raw request body (decoded as UTF-8) when the body is not
JSON, rejecting only a truly empty body. The JSON path is unchanged.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Adding an observer to the python harness for web search tools
* Escape dynamic strings with rich.markup.escape() in WebSearchDisplayObserver
Apply rich.markup.escape() to all user/tool-provided strings (queries, URLs,
titles, patterns) before interpolation into Rich-markup-enabled output. This
prevents characters like '['/']' from being interpreted as Rich markup tags.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Fix render issue for tools that are streamed in parts.
* Address PR review: missing call_id fallback, empty-mapping args, _is_complete perf
- Print call_id-less function calls as-is instead of merging under a name-derived
key (which could drop distinct unnamed calls).
- Preserve an empty {} mapping rather than coercing it to None.
- Add a structural bracket-balance gate before json.loads in _is_complete to
avoid O(n^2) re-parsing of growing streamed arguments.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Remove unsupported as_agent config parameter
Fixes#6313
Remove the unsupported function_invocation_configuration parameter from BaseChatClient.as_agent(), which currently forwards an invalid kwarg into Agent.__init__(). This keeps the existing TypeError behavior for callers but changes the error source to the public API boundary, which we do not consider a breaking change.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix sample
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Fix ollama_chat_client.py sample: pass tools via options dict
The sample was passing tools as a direct keyword argument to
get_response(), which caused a TypeError. The tools parameter
must be passed inside the options dict per the SupportsChatGetResponse
protocol.
Fixes#6411
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Wrap tools in a list as expected by OllamaChatClient
_prepare_tools_for_ollama iterates the tools value, so it must be a
list rather than a bare FunctionTool instance.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Python: Add AgentLoopMiddleware for re-running agents in a loop
Add `AgentLoopMiddleware`, an `AgentMiddleware` that re-runs the wrapped
agent in a loop. A single configurable class covers three common patterns,
each with a convenience classmethod factory:
- Ralph loop (`.ralph(...)`): no exit criteria, with feedback tracking
(`record_feedback`/`progress`), progress injection (`inject_progress`),
optional fresh context per iteration (`fresh_context`), and an early-stop
completion signal (`is_complete`).
- Predicate (`.with_predicate(...)`): loop while a `should_continue` callable
returns True (e.g. paired with `todos_remaining`/`background_tasks_running`).
- Judge (`.with_judge(...)`): a second chat client decides whether the original
request was answered, using a `JudgeVerdict` structured-output response.
The loop also auto-resolves pending function-approval / user-input requests via
an `on_approval_request` callable (bounded by `max_approval_rounds`), and the
next iteration's input is controlled by `next_message`. Supports both streaming
and non-streaming runs.
Exports `AgentLoopMiddleware`, `JudgeVerdict`, `todos_remaining`, and
`background_tasks_running`. Adds tests, a sample, and docs.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Python: Refine AgentLoopMiddleware API and sample
- with_judge: add criteria list with {{criteria}} templating into judge
instructions plus an agent-side instruction; add fresh_context, additional
judge feedback relay; default judge max_iterations.
- should_continue is now required and positional; supports (bool, str|None)
feedback tuples surfaced to next_message/record_feedback via feedback kwarg.
- Judge forwards full multi-modal request and response messages.
- Default max_iterations=10 (explicit None = unbounded); removed is_complete and
Ralph terminology; ShouldContinueResult is a real TypeAlias.
- Sample: stream all loops, print iteration counts via injected user-block
boundaries (robust to function calling), <role>: content formatting, per-method
expected output, and a looping todo sample.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Python: Fix CI checks for AgentLoopMiddleware
- Resolve pyright errors in _loop.py: drop the always-true final_result None
check (the while loop always assigns it) and cast finish_reason to the
AgentResponse constructor's expected type.
- Apply pyupgrade --py310-plus: import TypeAlias from typing.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Python: Resolve mypy/pyright disagreement on finish_reason
pyright infers AgentResponse.finish_reason as including str and rejects the
direct assignment, while mypy considers a cast redundant. Drop the cast and
suppress only pyright with a targeted reportArgumentType ignore, satisfying
both type checkers.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Python: Add todo+judge AgentLoopMiddleware sample
Add a second AgentLoopMiddleware sample that composes two criteria in one
should_continue predicate: a TodoProvider check (evaluated first) and a
report-style judge chat client (evaluated once todos are complete) that grades
the assembled report against shared requirements. Register it in the middleware
samples README.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Python: Compose todo+judge loops as two middleware
Rework the todo+judge sample to compose two AgentLoopMiddleware on the agent
itself (middleware=[judge_loop, todo_loop]) instead of a single hand-written
predicate. The inner todos_remaining loop drafts the report todo-by-todo and the
outer with_judge loop re-runs it until an editor chat client judges the report
publication-ready, reusing the built-in helpers.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Reset session for fresh_context loops via snapshot/restore
AgentLoopMiddleware.fresh_context previously only reset context.messages,
so with an attached session each iteration still reloaded the local
transcript or re-threaded the service-side conversation id and the model
saw the accumulated history. Snapshot the session once before the loop
(via to_dict) and restore it (from_dict + field copy) between iterations,
so every pass starts from the pre-loop baseline. The final iteration's
pass is persisted (no restore after the terminating iteration), so a
subsequent agent.run continues from there.
Removed the obsolete warning, updated docstrings and core AGENTS.md, and
added tests: a snapshot/restore round-trip, a session-reset
streaming x fresh_context x inject_progress x store matrix across multiple
runs and loop iterations, and response_format parsing across the loop.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Updated samples and docstrings
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Integrate shell tool into AgentHarness
* Validate shell_executor exposes as_function() with a clear TypeError
Addresses PR review feedback: a public factory should fail fast with an
actionable error rather than a cryptic AttributeError when an incompatible
shell_executor is supplied. Validation happens upfront, regardless of whether
the client supports shell tools.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Type shell harness params via TYPE_CHECKING import
Addresses PR review feedback: type shell_executor and
shell_environment_provider_options instead of Any, using a TYPE_CHECKING
import from agent_framework_tools.shell. The import never executes at
runtime, so there is no circular dependency, and the lazy runtime import of
ShellEnvironmentProvider is retained. Since ShellExecutor is a protocol
without as_function(), the validated getattr result is invoked directly.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Parse structuredContent from MCP CallToolResult (#3313)
The _parse_tool_result_from_mcp method only iterated over the content
field from CallToolResult, ignoring the structuredContent field entirely.
MCP servers that return JSON data via structuredContent (e.g., Power BI
MCP) appeared to return None.
Add handling for structuredContent: when present, serialize it as JSON
text and append it to the result list. This preserves the data for the
LLM while maintaining backward compatibility with existing behavior.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Python: Parse MCP CallToolResult.structuredContent field to prevent tool results returning None
Fixes#3313
* Address review feedback: add default=str to json.dumps and remove .checkpoints/
- Add default=str to json.dumps for structuredContent serialization so
non-JSON-serializable values (e.g. bytes) degrade gracefully instead
of raising TypeError
- Remove all .checkpoints/ runtime artifacts from the repository
- Add **/.checkpoints/ to .gitignore to prevent future accidental commits
- Add test for non-serializable structuredContent values
Fixes#3313
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Address review feedback for #3313: Python: MCP CallToolResult.structuredContent field is not parsed, causing tool results to return None
---------
Co-authored-by: Copilot <copilot@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Add sampling guardrails to MCP tools
Add approval, token, and request-count controls to the MCP sampling
callback used when an MCPTool is configured with a chat client.
- Add `sampling_approval_callback`, `sampling_max_tokens`, and
`sampling_max_requests` parameters to `MCPTool` and its
`MCPStdioTool`, `MCPStreamableHTTPTool`, and `MCPWebsocketTool`
subclasses, positioned directly after `client`.
- Gate each server-initiated `sampling/createMessage` request behind the
approval callback, which denies by default when no callback is provided.
- Clamp the requested `maxTokens` to `sampling_max_tokens` and enforce a
per-session request count via `sampling_max_requests`.
- Log incoming sampling requests at WARNING level (counts only).
- Export `SamplingApprovalCallback` from the public API.
- Add tests, a sample, and documentation updates.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Make sampling denial message context-aware
Distinguish the deny-by-default case (no approval callback configured)
from an explicit denial by a configured `sampling_approval_callback`, so
the returned ErrorData message is accurate for callback-driven denials
and exceptions.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* MCP long-running task support in Python
* Fix pyupgrade and AGENTS.md reconnect description
- pyupgrade: drop forward-reference string annotations in _mcp.py (Python 3.10+ resolves them natively now that MCPTaskOptions is defined before use).
- AGENTS.md: align reconnect description with current behavior. Phase 1 (initial tools/call) does NOT retry on connection loss; raises 'connection lost; task state unknown' instead, so a server that accepted the request but lost the response cannot start the operation twice. Phase 2 (tasks/get / tasks/result) still reconnects once against the same task_id.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Fix bandit nosec marker for CI pipeline
* Address PR feedbacks
* Clarifiied comments and addressed more PR feedbacks.
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>