* Python: Split type checkers by target (pyright source, 5 checkers on tests/samples) Rework the typing setup along the lines of the 'too many type checkers' approach: - Pyright (strict) is now the sole source-code type checker; mypy is removed from source and its [tool.mypy] block becomes a relaxed profile used only for tests/samples. - Tests are checked by all five checkers (pyright relaxed, mypy, pyrefly, ty, zuban); samples by pyright, pyrefly, and ty. All run in a relaxed/ basic profile so authors aren't forced into over-annotation. - Add pyrightconfig.tests.json and bump sample pyright configs to basic. - Unify test/sample typing onto the same parallel fan-out used by source pyright via run_command_items in task_runner.py. - Make version-conditional imports symmetric: keep or drop the '# type: ignore' on both branches so results match across interpreter versions (local vs CI). - Update SKILL.md, DEV_SETUP.md, and CODING_STANDARD.md for the five gating checkers and pyright on source+tests+samples. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Fix merge regressions from main (typing + runtime) Merging main into the type-checker split branch surfaced regressions that the new five-checker test suite and unit tests caught: Runtime fixes: - anthropic: restore the dropped `cache_read_input_token_count` mapping in _parse_usage_from_anthropic (lost during merge conflict resolution). - gemini: _get_function_calling_mode test helper returned str(enum) ('FunctionCallingConfigMode.AUTO') instead of the enum value ('AUTO'). - openai: _response_id_from_token test helper was an infinite self-recursion; return token['response_id']. - orchestrations: reset output_events per approval iteration so the terminal output assertion counts only the final run. - core: drop a stale duplicate harness test whose message ('non-negative') contradicted the source ('positive'). - purview: import PolicyLocation/PolicyScope/ProtectionScopeActivities/ ExecutionMode used by the processor tests. Type-checker fixes (tests, relaxed profile): - core: pyright/mypy/pyrefly/ty/zuban green-ups across the harness, MCP, observability and types tests. - anthropic/openai: route provider-namespaced UsageDetails keys through a dict cast (extra_items TypedDict unsupported by mypy/ty). - purview: typed model constructors and cache-mock casts. - ag-ui: annotate WorkflowContext[Any, Any] so yield_output accepts test payloads, guard Optional forwarded_props, and ty-ignore intentional bad args. Source pyright (sole source checker) flagged unnecessary ignores newly introduced by merged code in core _tools.py and declarative _declarative_base.py. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Isolate per-package mypy cache in test-typing fan-out The parallel test-typing fan-out runs many mypy processes concurrently, all defaulting to a single shared ./.mypy_cache. Concurrent writes corrupt the cache and mypy aborts with INTERNAL ERROR (intermittently, depending on worker timing) -- which is why CI's Test Typing job failed on a shifting set of packages while a single-package run was fine. Give each mypy invocation an isolated cache dir keyed by its target paths so incremental caching still works per package without races. Other checkers (zuban/pyrefly/ty/pyright) maintain their own caches and are unaffected. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Make lab pyright-only on source (drop source mypy) Lab was the last package still running mypy on its source code, requiring mypy-only `# type: ignore` comments that pyright (the sole source checker everywhere else) flags as unnecessary. Align lab with the rest of the monorepo: - Remove the lab source mypy poe tasks (mypy-gaia/lightning/tau2) and the now-dead strict [tool.mypy] config block. - Drop the 'Run lab mypy' CI step; lab source is type-checked by pyright only. Lab tests remain covered by the workspace test-typing fan-out (mypy, pyrefly, ty, zuban, pyright over tests using the relaxed root config). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Fix test-typing regressions from latest main merge A fresh merge from main brought in new test code never run under the five-checker test-typing suite. Green up across the affected packages: - core: narrow Optional span.attributes with 'and' guards in span filters and assert+cast the json.loads(...attributes[...]) reads (test_observability); match the existing as_agent ignore on the protocol-typed fixture (test_clients). - openai: align new streaming tests with the established chat_options dict pattern (ChatOptions TypedDict isn't assignable to dict), route Optional .annotations[0] access through a small _first_annotation helper (mirrors the file's assert-not-None convention), and annotate a mapped ResponseStream. - foundry_hosting: annotate error: dict[str, Any] = body.get(...) or {} (zuban needs the annotation). - foundry: narrow ignores for the live AIProjectClient credential arg (pyrefly) and connections.get_default (zuban) SDK type gaps. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * updated pyright version * pyright fix * Python: Fix source typing for pyright 1.1.410 Pyright 1.1.410 tightened several checks. Apply the same source fixes as upstream PR #6275: - anthropic: import AsyncAnthropicBedrock from anthropic.lib.bedrock and AsyncAnthropicVertex from anthropic.lib.vertex (no longer re-exported from the anthropic top-level package -> reportPrivateImportUsage). - core _types.py: cast the transform-hook result to UpdateT (reportAssignmentType). - core _workflows/_events.py: annotate the @contextmanager helper as Generator[None] instead of Iterator[None] (reportDeprecated). - redis: build the combined filter expression with an explicit loop instead of reduce(and_, ...), which pyright could no longer fully type (drops the now unused functools.reduce / operator.and_ imports). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Python: Accept plain-text body in Azure Functions workflow/run endpoint The workflow_orchestrator already accepts plain strings as well as JSON objects via context.get_input(), but the start_workflow_orchestration HTTP handler only accepted JSON and returned 400 for any non-JSON body. This made the functions integration tests that POST text/plain to /api/workflow/run (e.g. test_09_workflow_shared_state) fail consistently with 400 != 202. Fall back to the raw request body (decoded as UTF-8) when the body is not JSON, rejecting only a truly empty body. The JSON path is unchanged. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Tools
Samples that show how to define, configure, and control function tools for an agent — from basic declarations to approvals, invocation limits, session injection, and dynamic (progressive) tool exposure.
Function tools
| File | Demonstrates |
|---|---|
function_tool_with_explicit_schema.py |
Defining a tool with an explicit JSON schema. |
function_tool_declaration_only.py |
A declaration-only tool (schema without a local implementation). |
function_tool_with_kwargs.py |
Passing extra keyword arguments into a tool. |
function_tool_from_dict_with_dependency_injection.py |
Dependency injection into a tool defined from a dict. |
function_tool_with_session_injection.py |
Injecting the session into a tool. |
tool_in_class.py |
Using a method on a class as a tool. |
agent_as_tool_with_session_propagation.py |
Exposing an agent as a tool with session propagation. |
Approvals & invocation control
| File | Demonstrates |
|---|---|
function_tool_with_approval.py |
Requiring human approval before a tool runs. |
function_tool_with_approval_and_sessions.py |
Tool approvals combined with sessions. |
tool_approval_middleware.py |
Session-backed approval coordination, mixed-batch approvals, and "always approve" rules. |
function_invocation_configuration.py |
Configuring function-invocation settings (e.g. max iterations). |
control_total_tool_executions.py |
All the ways to cap how many times tools run. |
function_tool_with_max_invocations.py |
Limiting the number of invocations per tool. |
function_tool_with_max_exceptions.py |
Limiting the number of exceptions a tool may raise. |
function_tool_recover_from_failures.py |
Returning errors so the agent can recover from tool failures. |
Progressive tool exposure (dynamic loading)
| File | Demonstrates |
|---|---|
dynamic_tool_exposure.py |
A "loader" tool that adds more tools at runtime via FunctionInvocationContext. |
Frontloading a model with hundreds of tools hurts tool-selection accuracy,
bloats context, and raises cost. Instead, start with a small set of loader
tools and let the model pull in more on demand. Inside a tool, the injected
ctx: FunctionInvocationContext exposes a live ctx.tools list plus
ctx.add_tools(...) / ctx.remove_tools(...) helpers. Tools added or removed
take effect on the next iteration of the function-calling loop.
Note
Progressive tool exposure applies to the standard function-calling loop. It does not apply to CodeAct providers (
agent-framework-monty,agent-framework-hyperlight). In CodeAct the model only sees a singleexecute_codetool, and host tools are exposed inside the sandbox as typed Python functions rather than as model tool-schemas. Host tools there are invoked without aFunctionInvocationContext, soctx.add_tools()is not available; the helpers fail fast with a clearRuntimeErrorinstead of silently doing nothing. To change a CodeAct agent's tool set, use the provider's ownadd_tools/remove_tool/clear_toolsmethods (applied between runs). The recommended provider-driven path for Monty and Hyperlight is shown in../context_providers/code_act/(code_act.pyfor Hyperlight,monty_code_act.pyfor Monty).
Local shell & code interpreters
| Path | Demonstrates |
|---|---|
local_shell_with_allowlist.py |
LocalShellTool restricted by a strict command allow-list. |
local_shell_with_environment_provider.py |
LocalShellTool wired with a ShellEnvironmentProvider. |
local_code_interpreter/ |
Hyperlight-backed sandboxed code interpreter (standalone tool — extra pattern). |
monty_code_interpreter/ |
Monty-backed sandboxed code interpreter (standalone tool — extra pattern). |
Tip
The
local_code_interpreter/andmonty_code_interpreter/samples show the standalone-tool wiring and are provided as extra reference. For most Monty/Hyperlight use cases the recommended path is the provider-driven CodeAct setup in../context_providers/code_act/, which adds dynamic tool / capability management.