* docs: ADR-0027 feature-usage bitmask in the User-Agent Add an ADR, design spec, and per-language bit registry for a lightweight feature-usage signal: a 64-bit mask, emitted as a `(feat=vN.<hex>)` User-Agent comment, stamped per request on first-party (Azure/Foundry) clients only. - docs/decisions/0027-feature-usage-bitmask-user-agent.md — ADR (options-first, with Limitations, Open Questions, and v1->v2 migration) - docs/specs/002-feature-usage-telemetry.md — design spec + implementation plan - docs/specs/feature-usage-bit-registry.md — per-language bit tables + governance Granularity is per package with core broken out per feature (each orchestration pattern and built-in context/history provider). Registries are per language (decoder selects by the language already in the UA). OpenTelemetry emission is deferred (privacy). Docs only; no code changes. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs: fix dead links to removed registry JSON in ADR-0027 The registry JSON was consolidated into feature-usage-bit-registry.md; point the ADR's two remaining links at the markdown instead of the deleted file (fixes markdown-link-check 404s). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs: address review — drop JSON-parity wording, clarify per-language decode - ADR option J: the parity test compares the enum against the per-language table in the registry doc, not a (now-removed) JSON file. - Spec .NET mapping: the wire format is shared, but the mask is decoded per-language (select the table via the UA product token) — fixes the "decoded numbers mean the same thing in both SDKs" wording that conflicted with the per-language, non-synchronized bit indexes. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs: add dedicated mask-only opt-out env var (AGENT_FRAMEWORK_FEATURE_MASK_DISABLED) Re-introduce a dedicated opt-out that disables only the feature mask while keeping the base agent-framework-<lang>/{version} User-Agent, alongside the existing AGENT_FRAMEWORK_USER_AGENT_DISABLED (whole UA). Updates the spec accumulator gate, API surface, opt-out table and examples; the registry opt-out section; and the ADR (decision outcome, consequences, open questions -> decided). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs: add prior-art comparison (AWS botocore m/, Stainless, Azure, etc.) Add a Prior art section to ADR-0027 surveying how comparable SDKs encode identity/usage in the User-Agent or sidecar headers, with citations: - AWS botocore `m/` feature-code list — the direct analog (per-request, usage-based feature flags in the UA); contrasts short-code set vs our hex bitmask. - OpenAI/Anthropic Stainless `X-Stainless-*` headers (static identity). - Azure azure-core UserAgentPolicy + AZURE_TELEMETRY_DISABLED. - Google x-goog-api-client; LangSmith version token + tracing opt-in. Also add an Open Question on honoring the cross-tool DO_NOT_TRACK convention. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs: fold in botocore lessons; record accumulation-scope decision botocore's m/ feature list scopes features to a per-request contextvars set that resets between calls — clean per-call attribution, but it assumes every feature lives inside a service request. That holds for an SDK natively bound to its own services; it does not for us, where many features (agent/workflow/provider construction, session setup) are not bound to any request. - ADR: add Accumulation scope options — P (process-global monotonic, chosen) vs Q (botocore per-request set, rejected) with the request-binding rationale; reference P in the decision; reframe the "no per-call attribution" limitation as a deliberate scope choice. - ADR Prior art: bitmask gives bounded token size for free (vs botocore's 1024-byte cap + truncation); mechanism is private, wire format is the contract; fix a duplicated phrase. - Spec: note the mask is process-global, monotonic, never reset (intentional, lock/Interlocked.Or-safe), the token is safe-by-construction (no sanitization), and the helpers are private API. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs: update feature mask ADR Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs: refresh feature usage telemetry design Rebase the proposal on current main, renumber it to ADR-0033/SPEC-004, and reconcile the registry and implementation notes with current Python and .NET surfaces. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs: expand feature usage mask to 128 bits Repartition the v1 registries with additional skill categories, define the bit-allocation tenet, and document the two-lane .NET accumulator and 128-bit decoder contract. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 346bf168-b668-4c4a-a8db-a67282ee5e5f * docs: tighten feature telemetry activation and scoping Require approved pipeline and actual-origin classification, preserve OpenAI transport defaults, use activation-based marking, and move index ownership into packages with parity and no-overlap validation. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 346bf168-b668-4c4a-a8db-a67282ee5e5f * docs: preserve SDK transport defaults for telemetry Record the transport-preservation requirement at the ADR decision level. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 346bf168-b668-4c4a-a8db-a67282ee5e5f * docs: split declarative agent and workflow usage Allocate separate adjacent v1 indexes for declarative agents and declarative workflows in Python and .NET, shifting later unreleased rows. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 346bf168-b668-4c4a-a8db-a67282ee5e5f * docs: accept feature usage telemetry ADR Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 346bf168-b668-4c4a-a8db-a67282ee5e5f * docs: record feature telemetry ADR participants Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 346bf168-b668-4c4a-a8db-a67282ee5e5f * docs: expand feature telemetry ADR consultation Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 346bf168-b668-4c4a-a8db-a67282ee5e5f * docs: clarify feature telemetry semantics Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 346bf168-b668-4c4a-a8db-a67282ee5e5f --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 346bf168-b668-4c4a-a8db-a67282ee5e5f
24 KiB
status, contact, date, deciders, consulted, informed
| status | contact | date | deciders | consulted | informed |
|---|---|---|---|---|---|
| proposed | eavanvalkenburg | 2026-07-22 | eavanvalkenburg |
Feature-usage telemetry via an accumulating bitmask
Companion design for ADR-0033. The per-language bit tables, encoding, opt-out, and governance live in feature-usage-bit-registry.md. The registry allocates indexes; package-local
FeatureIndexdeclarations implement them.
What is the goal of this feature?
Give the Agent Framework team a lightweight signal about which framework features are actually exercised at runtime (not merely installed), so we can prioritise investment based on real usage. We emit a single small number — a feature mask — on the User-Agent that already goes out with each request.
Reach is deliberately bounded. The mask accumulates from all feature usage,
but the feat= token is only stamped through an explicit allowlist of
first-party Azure/Foundry client pipelines whose User-Agent telemetry the
team can ingest (initially Foundry/Azure OpenAI). We do not send the token to
third-party providers (OpenAI direct, Anthropic, Bedrock, Gemini, Ollama,
Mistral), or to an Azure service merely because its hostname is first-party;
doing so would leak a deployment fingerprint into logs we cannot read (see
Emission).
The current candidate uses package-level bits plus selected major capabilities: one bit per orchestration pattern (sequential / concurrent / group-chat / magentic / handoff), one bit per built-in context/history provider, selected skill source types, and separate Foundry chat/agent/memory/evals/toolbox bits (plus embedding in Python). See the registry. ADR-0033 still leaves final v1 granularity open. The refreshed candidate assigns 63 Python indexes and 52 .NET indexes. V1 uses 128 bits, leaving 65 Python and 76 .NET positions for additive growth.
Success metric: within one release after rollout, ≥80% of eligible, framework-created first-party (Foundry) requests carry a non-empty feature token whose mask reflects features activated after client construction (i.e. the token is live, not frozen — see the request-time stamping requirement below). This measures transport coverage, not feature invocation volume. Secondary: ability to describe which process-lifetime feature bits are observed together in eligible traffic (e.g. "requests observed from processes that have used workflows"). Repeated requests carrying a bit are not additional uses.
This is done transparently: the bit registry is public, the emitted value is
human-decodable, and a dedicated AGENT_FRAMEWORK_FEATURE_MASK_DISABLED
disables the mask while preserving the base User-Agent. Python's existing
AGENT_FRAMEWORK_USER_AGENT_DISABLED continues to suppress its entire
User-Agent contribution, mask included.
What is the problem being solved?
Today we only know which packages are installed (from package telemetry) or
that some Agent Framework call happened (the existing
agent-framework-python/{version} User-Agent). We have no usage-based signal
about feature combinations, and no way to tell that, say, a process uses
workflows + MCP + Foundry together. Collecting this through bespoke events would
add cost and new data flows; folding a tiny accumulating integer into telemetry
we already send is far cheaper and easier to reason about for privacy.
Mechanism
Process-global accumulator in core
The accumulator and its helpers live in the existing
agent_framework/_telemetry.py (alongside get_user_agent() /
prepend_agent_framework_to_user_agent()), so the User-Agent machinery stays in
one module. It owns a process-global 128-bit accumulator. Python's arbitrary-size
int stores it directly. A dedicated
AGENT_FRAMEWORK_FEATURE_MASK_DISABLED that drops only the feature mask
while keeping the base agent-framework-python/{version} User-Agent is
introduced by this design. The existing Python
AGENT_FRAMEWORK_USER_AGENT_DISABLED continues to drop the whole User-Agent
contribution, mask included:
# agent_framework/_telemetry.py (same module as get_user_agent)
# IS_TELEMETRY_ENABLED already defined here (AGENT_FRAMEWORK_USER_AGENT_DISABLED)
FEATURE_MASK_DISABLED_ENV_VAR = "AGENT_FRAMEWORK_FEATURE_MASK_DISABLED"
REGISTRY_VERSION = 1
_feature_mask = 0
_feature_mask_lock = threading.Lock()
def _feature_mask_enabled() -> bool:
"""Mask is on unless the UA is disabled or the dedicated flag is set."""
if not IS_TELEMETRY_ENABLED:
return False
return os.environ.get(FEATURE_MASK_DISABLED_ENV_VAR, "false").lower() not in ("true", "1")
def mark_feature_used(index: int) -> None:
"""OR a feature bit into the process-global mask.
Called the first time a feature is exercised. Cheap and idempotent;
a no-op when the feature mask is disabled.
"""
global _feature_mask
if not _feature_mask_enabled():
return
if not 0 <= index < 128:
raise ValueError(f"Feature index must be in range 0..127, got {index}")
with _feature_mask_lock:
_feature_mask |= 1 << index
def get_feature_token() -> str | None:
"""Return ``v<version>.<hex_mask>`` for the accumulated mask, or None."""
if not _feature_mask_enabled() or _feature_mask == 0:
return None
return f"v{REGISTRY_VERSION}.{_feature_mask:x}"
- Per package/feature, usage-based:
mark_feature_used()is called at the feature's first meaningful activation, never at import/install time. For operational clients, tools, providers, and hosts, activation is the first public operation that exercises the capability. Construction is a valid mark point only when construction itself performs the capability (for example, registering/starting runtime resources), not merely because a DI container instantiated an otherwise-unused object. - Process-global and monotonic — intentionally never reset. Unlike a
per-request scheme (e.g. botocore's
contextvarsfeature set that resets between calls), our mask spans the whole process because many features are not bound to any service request — an agent or workflow may first run, a provider may first participate in a session, and a host may start serving independently of the later request that emits the token. The single global mask is the only scope that can represent them, and its monotonic "usage so far" growth is the intended semantic, not a bleed bug. Concurrency-safe via the module lock (Python) / two atomic 64-bit lanes in .NET. - Binary and non-countable. A set bit means "this feature was observed at least once in this process before this request." Repeating that bit on every later eligible request does not represent additional uses and must not be interpreted as request, invocation, agent, user, or tenant counts.
- No scoped enable/disable bookkeeping. Making the mask exact per operation would add hot-path state changes, context propagation, and reset/error-path handling. It would also produce a more detailed behavioral trace and therefore increase privacy sensitivity. V1 deliberately keeps the coarser process-level Boolean.
- Token is safe by construction. The emitted value is
v{int}.{hex}— characters limited to[0-9a-fv.]— so no header-injection sanitization is required. A 128-bit mask is at most 32 hex characters (contrast botocore, which must sanitize and cap arbitrary component strings). - Private API.
mark_feature_used,get_feature_token,apply_feature_tokenand the mask itself are internal helpers; only the emitted token and the per-language registry tables are the stable, decodable contract. - No import cycles: the accumulator lives in core, while each package owns private index constants for its own features and calls the core marker. Core never imports optional packages.
Interpretation contract
At time 1, Agent A in a worker can use MCP and a Foundry chat client. At time 2, Agent B in the same worker can make a normal Foundry chat call without MCP. The time-2 request still carries the MCP bit because MCP was observed earlier in the process.
That request means only "this process has used MCP." It does not mean Agent B used MCP, that MCP was used on the time-2 request, or that two requests carrying the bit equal two MCP uses. Without a separate stable process identifier, the signal also cannot produce unique-process counts. Supported analysis is limited to coarse observed-feature prevalence and feature co-occurrence, with the request-weighting limitation called out explicitly.
Bit constants
The registry is the allocation authority. Each package defines a private,
hand-written FeatureIndex IntEnum (or equivalent constants) containing only
the rows it owns. Core owns core indexes plus the accumulator; optional packages
can allocate and ship new indexes without requiring a core release after the
marker API exists.
# agent_framework_foundry/_feature_usage.py
from enum import IntEnum
from agent_framework._telemetry import mark_feature_used # pyright: ignore[reportAttributeAccessIssue]
class FeatureIndex(IntEnum):
FOUNDRY_CHAT_CLIENT = 48
class RawFoundryChatClient:
async def _send_request(self) -> None:
mark_feature_used(FeatureIndex.FOUNDRY_CHAT_CLIENT)
...
A repository validation test reads every package-local declaration and the
matching language/version table. It fails when an index is out of range, missing
from the registry, duplicated/overlapping across packages, or mapped to the wrong
id. For reference, in v1 FoundryChatClient → index 48,
FoundryAgent → index 49, Foundry memory → index 50.
Usage activation points
- Clients/embeddings/evals: first outbound operation.
- Tools/MCP: first connection, discovery, or invocation that exercises the tool surface.
- Context/history providers: first provider hook or load/save operation, not constructor-only registration.
- Agents/workflows/orchestrations: first run/build/start operation that activates the defined runtime.
- Hosting: first serve/start/route activation.
- Constructor marking: allowed only when construction itself performs one of those activations or acquires/registers the runtime resource.
Emission
One path in v1: the User-Agent feat= token, stamped at request time on an
explicit allowlist of first-party Azure/Foundry client pipelines only.
Marking (mark_feature_used) is universal — every feature sets its index
regardless of provider. Only emission is scoped. A user who never calls a
first-party endpoint emits no token; this is the honest, intended behaviour (no
third-party leakage, no signal we couldn't read anyway).
The existing base User-Agent behavior (agent-framework-python/{version} plus
any dynamically detected hosting prefix) is unchanged; packages continue using
their current default_headers, user_agent, suffix, or policy mechanisms.
get_user_agent() stays base-only (no feat=). The feat= token is
separate, added only by eligible Azure/Foundry clients, and
re-evaluated on each request so it reflects the mask accumulated so far. A
helper stamps it:
This request-time read does not make the signal request-scoped. The payload remains the process-global Boolean history described above.
# agent_framework/_telemetry.py
def apply_feature_token(user_agent: str) -> str:
"""Append/refresh the live ``(feat=v<ver>.<hex>)`` comment on a UA string.
Re-reads the current mask on every call, so newly accumulated bits are
reflected immediately. Idempotent: replaces an existing ``(feat=...)``
comment rather than appending a second.
"""
token = get_feature_token() # None when disabled or mask == 0
base = _strip_feature_comment(user_agent)
return f"{base} (feat={token})" if token else base
Emission requires both:
- an explicitly approved framework client/pipeline family; and
- the actual request's normalized HTTPS origin matching that family's reviewed first-party origin allowlist.
Credentials, use_azure, or an Azure-named setting alone do not approve a
destination. Approval depends on the resolved origin: customer-specific
subdomains on reviewed Azure/Foundry suffixes remain eligible even when supplied
through base_url / AZURE_OPENAI_BASE_URL, while customer gateways and unknown
OpenAI-compatible origins are denied by default. The check runs on every actual
request, including redirect hops; a cross-origin or otherwise unapproved redirect
removes (feat=...) before sending.
Eligible first-party clients install a request hook that performs this
classification and calls apply_feature_token():
- OpenAI-SDK clients created by Agent Framework: construct the underlying
client with
http_client=DefaultAsyncHttpxClient(event_hooks={"request": [_stamp_feat_hook]}). Using OpenAI'sDefaultAsyncHttpxClientpreserves the SDK's redirect, connection-limit, and timeout defaults; a plainhttpx.AsyncClientmust not replace them. The hook adds or removes the token based on the approved pipeline plus actual-origin classification. Caller-supplied clients/transports are not replaced or patched. - azure-core pipeline clients: start with
AIProjectClientpaths whose telemetry is confirmed ingestible. When Agent Framework constructs/configures an approved pipeline, add a separate per-callSansIOHTTPPolicywhoseon_requestperforms the same actual-origin check and callsapply_feature_token()onrequest.http_request.headers["User-Agent"]. Do not stampSearchClient,CosmosClient, or another Azure client merely because it is first-party; add it to the allowlist only after confirming the data path. This mirrors .NET's request-timePipelinePolicyexactly.
This fixes the frozen-at-construction problem: the token is materialised at send time, not client-init time, so it carries features activated after the client was created. It also confines the token to first-party endpoints. Caller-owned clients are not patched, and toolkit-owned clients without a supported public hook are outside v1 coverage.
Encoding uses the RFC 7231 comment form (feat=v1.<hex>) (metadata, not a
product token), placed after the agent-framework product token, e.g.:
foundry-hosting/agent-framework-python/1.2.3 (feat=v1.2a)
OpenTelemetry — not in v1
An OTel span attribute carrying the same value was considered but deferred — primarily for privacy, not complexity. Unlike the first-party-only UA token, a span attribute broadcasts the feature-combination fingerprint into the user's general telemetry pipeline, which is commonly exported to third-party APM vendors (Datadog, Honeycomb, …) — re-introducing exactly the leakage the first-party scoping was chosen to avoid. (It also carries a cardinality footgun: a monotonically-growing, combinatorial value must never become a metric dimension.) The version prefix leaves the door open to add it later if the User-Agent path cannot answer a concrete query and there is an acceptable scoped/redacted variant; v1 ships the UA path only. See ADR-0033 → option C.
API Changes
New internal cross-package surface in
agent_framework._telemetry (not exported from agent_framework):
mark_feature_used(index: int) -> Noneget_feature_token() -> str | None— returnsv<ver>.<hex>orNone.apply_feature_token(user_agent: str) -> str— live, idempotent UA stamper used by first-party request hooks.FEATURE_MASK_DISABLED_ENV_VARconstant — the dedicated mask-only opt-out env var name (AGENT_FRAMEWORK_FEATURE_MASK_DISABLED).
Each package also adds a private package-local FeatureIndex declaration for
the rows it owns. The dedicated mask-only opt-out and Python's existing
whole-User-Agent opt-out gate the Python mask; see Opt-out.
Behavioural change to existing API:
get_user_agent()/prepend_agent_framework_to_user_agent()are unchanged — they keep returning the base UA with nofeat=token. The token is added only by first-party request hooks viaapply_feature_token().
No breaking changes: when the mask is empty or disabled, for any non-first-party client, or for an injected client outside the supported-hook set, output is byte-for-byte identical to today.
Opt-out
The dedicated mask-only opt-out is shared by both SDKs. Python also retains its pre-existing whole-User-Agent opt-out:
| Env var | SDKs | Effect |
|---|---|---|
AGENT_FRAMEWORK_FEATURE_MASK_DISABLED |
Python and .NET | disables only the feature mask; the base agent-framework-<lang>/{version} User-Agent is still sent |
AGENT_FRAMEWORK_USER_AGENT_DISABLED |
Python (existing behavior) | disables the entire Python AF User-Agent contribution, mask included |
The flags accept true/1 (case-insensitive). The dedicated flag lets a
privacy-conscious user keep contributing the SDK identity/version (useful for
support and compat triage) while withholding the feature-usage signal. The mask
is also disabled implicitly whenever Python's whole User-Agent is disabled. A
new whole-User-Agent opt-out for .NET is outside this design.
E2E example
from agent_framework import Agent
from agent_framework_foundry import FoundryChatClient
from agent_framework_openai import OpenAIChatClient
# First-party (Foundry) client: request hook stamps the live feat token.
agent = Agent(client=FoundryChatClient(...), instructions="...")
# Agent use marks bit 0; FoundryChatClient marks bit 48
await agent.run("Hello")
# Outgoing request to Foundry carries:
# User-Agent: agent-framework-python/1.2.3 (feat=v1.<mask-at-send-time>)
# Third-party client: NO feat token is added (no first-party hook).
other = Agent(client=OpenAIChatClient(...), instructions="...")
await other.run("Hi")
# Outgoing request to OpenAI carries only:
# User-Agent: agent-framework-python/1.2.3
Drop only the feature mask (keep the base User-Agent):
AGENT_FRAMEWORK_FEATURE_MASK_DISABLED=true python app.py
# Foundry request User-Agent: agent-framework-python/1.2.3 (no (feat=...) comment)
Python only: use the existing flag to drop its entire User-Agent contribution (mask included):
AGENT_FRAMEWORK_USER_AGENT_DISABLED=true python app.py
.NET mapping
- Core owns
FeatureUsage.MarkUsed(int index)plus the core package's private index declaration. Each optional assembly owns a privateFeatureIndexenum containing only its allocated rows. These are index positions0..127, not[Flags]values;MarkUsedperforms the shift. - Store the 128-bit mask as two
longlanes (lowfor bits 0–63,highfor 64–127). Marking touches one lane withInterlocked.Orwhere available and a smallInterlocked.CompareExchangeloop onnetstandard2.0/net472. Read each lane atomically. Since bits only move from zero to one, a concurrent two-lane snapshot may miss a just-added bit but can never invent or clear one; the next request includes it. - Format without depending on
UInt128: ifhigh == 0, emitlowas lowercase hex; otherwise emithighwithout leading zeros followed bylow:x16. Cast each signed lane toulongbefore formatting so bits 63 and 127 are preserved. Reject indexes outside0..127. - Emission is stamped at request time and first-party-scoped, matching
Python. The
existing
AgentFrameworkUserAgentPolicy/HostedAgentUserAgentPolicypipeline policies already run per request — extend them to apply the same approved-pipeline + actual-origin classifier, append/refresh the(feat=...)comment only for approved destinations, and remove it on unapproved redirect hops. Do not register it on third-partyIChatClients. - Same wire format (
v<version>.<hex>comment, hex encoding) and the same dedicated mask-only opt-out (AGENT_FRAMEWORK_FEATURE_MASK_DISABLED). The mask is decoded per language: indexes are not shared, so a decoder must read the language from the UA product token and select that language's table before decoding. (.NET's policy was already request-time, so there is no Python/.NET timing asymmetry.) Adding a .NET whole-User-Agent opt-out is outside this design.
Keeping the bitmap in sync
feature-usage-bit-registry.md is the published
allocation contract. Package-local FeatureIndex declarations are the runtime
implementation. There is deliberately no shared numbering across languages
and no machine-readable registry file.
One repository validation test gathers every package-local declaration for one language/version and parses the matching Markdown table. It asserts:
- every declared index is within
0..127; - every
(index, id)exactly matches one registry row; - the union of declarations has no duplicate/overlapping indexes;
- every non-reserved registry row is declared exactly once.
Adding an optional-package feature therefore changes that package and the registry, not core. If a programmatic decoder is built later, export the table to JSON then.
Decoding
UA: agent-framework-python/1.2.3 (feat=v1.2a)
│ │ └ hex mask
│ └ version
└ language → pick the Python table (version 1)
Read language → pick the table; read vN → pick that version; AND the hex mask
against each bit. Unknown bits (from a newer SDK than the decoder's copy of the
table) are ignored.
Implementation plan (post-approval)
- Privacy approval — confirm the first-party-only feature-combination signal, retention, access, allowed queries, and opt-out behavior before code ships.
- Core accumulator — in
agent_framework/_telemetry.pyadd the 128-bit mask, lock,mark_feature_used(index),get_feature_token, andapply_feature_token;get_user_agent()stays base-only. - Package-local indexes + validation — add private
FeatureIndexdeclarations to packages and a repository test for exact registry parity, complete coverage, range, and zero overlap. - First-party request-time hooks — use OpenAI's
DefaultAsyncHttpxClientfor framework-created clients and the separate azure-coreSansIOHTTPPolicy. Require approved pipeline and approved actual origin on every request/redirect hop. Verify custom origins and cross-origin redirects never carry the token. - Mark feature usage — call
mark_feature_used(FeatureIndex.X)at the first meaningful activation. Operational clients/providers/tools mark on their first real operation; build/start points mark compositional features. Constructor-only marking requires construction itself to exercise the capability. - .NET parity — package-local index enums plus the two atomic 64-bit lanes
with
Interlocked.Or/ compare-exchange fallback; extend existing request-time Foundry UA policies through the shared destination classifier and formatter. - Docs & tests — update package
AGENTS.md/skills; tests for both Python opt-out paths (dedicated mask-only and existing whole-UA), the dedicated .NET mask-only opt-out, first-party scoping, and the live (non-frozen) UA.
Limitations & open questions
The decision-level limitations and unresolved trade-offs — reach, per-process (not per-call) attribution, v1 granularity, fingerprinting residue, and the OTel question — are owned by the ADR (the dedicated mask-only opt-out is now decided and included). See ADR-0033 → Limitations and Open Questions. This spec is the implementation reference; it does not re-litigate those choices.
Implementation-only note:
- Per-request hook overhead is negligible (a flag check, one Python integer snapshot or two atomic .NET lane reads, and a string concat per first-party request), but benchmark the hot path once if a high-QPS Foundry scenario is in scope.