54a30571aa
* .NET: feat(evals): RubricScore type + EvalScoreResult.Dimensions Adds the core rubric-evaluator surface that mirrors the Python work in PR #6101 (commit e45b934cc). Provider-agnostic types only — no Foundry coupling. Subsequent commits will wire these into FoundryEvals. - RubricScore: per-dimension score record (Id, Score?, Applicable, Weight, Reason). - EvalScoreResult.Dimensions: optional init-only list of RubricScore. Null for non-rubric (built-in) evaluators. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * .NET: feat(evals): GeneratedEvaluatorRef + assertion helpers Adds the provider-agnostic surface for referencing a pre-existing rubric evaluator and gating CI on per-item / per-dimension thresholds. Mirrors Python PR #6101 commits e5830dd7f (ref type) and 4bc60462d (asserts). - GeneratedEvaluatorRef: name + optional version/display-name, plus a Latest(name) factory for versionless refs (discouraged for CI; consumers should warn at run time). - AgentEvaluationResults.AssertScoreAtLeast: walks DetailedItems[].Scores, optionally filtered by evaluator name, recurses into SubResults. - AgentEvaluationResults.AssertDimensionScoreAtLeast: walks each score's Dimensions list, skips non-applicable dimensions by default, supports requireApplicable to flip that, recurses into SubResults. - AgentEvaluationResults.AssertNoFailedItems: walks DetailedItems for fail/error statuses, recurses into SubResults. All helpers throw InvalidOperationException (matches existing AssertAllPassed). Truncates offender lists to the first 5 with a '+N more' suffix to keep CI output readable, mirroring the Python helpers. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * .NET: feat(foundry-evals): accept GeneratedEvaluatorRef in evaluators= Adds FoundryEvaluatorSpec, a readonly-struct union with implicit conversions from both string and GeneratedEvaluatorRef so call sites can mix built-in evaluator names with rubric evaluator references: var evals = new FoundryEvals( projectClient, model, new GeneratedEvaluatorRef("policy-rubric", "3"), FoundryEvals.Relevance, FoundryEvals.Coherence); FoundryEvals constructors (3 overloads), EvaluateTracesAsync, and EvaluateFoundryTargetAsync now take FoundryEvaluatorSpec[]/params instead of string[]/params. Existing call sites using string literals or string[] keep working unchanged via implicit conversion. FoundryEvalConverter.BuildTestingCriteria emits the documented Foundry wire format for rubric refs: { "type": "azure_ai_evaluator", "name": <DisplayName ?? Name>, "evaluator_name": <Name>, "evaluator_version": <Version>, // omitted when null "initialization_parameters": { "deployment_name": <model> }, "data_mapping": { conversation arrays, optional tool_definitions } } WireTestingCriterion gains an optional EvaluatorVersion field. Rubric refs are preserved through FilterToolEvaluators (tool-aware but not tool-required) and ignored by FindMissingGroundTruthEvaluators. A versionless ref emits a Trace.TraceWarning at criterion-build time so CI authors notice the floating version (mirrors the Python warning). Adds 6 new Foundry unit tests (3 BuildTestingCriteria rubric paths, 1 FindMissingGroundTruthEvaluators, 1 FilterToolEvaluators preservation, 1 mixed-order). 369/369 Foundry tests pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * .NET: feat(foundry-evals): parse rubric dimension_scores into RubricScore Adds FoundryEvals.ParseRubricScores, called per result inside ParseDetailedItem. Each EvalScoreResult now populates Dimensions when the evaluator's sample carries a rubric breakdown. Accepts three shapes for forward compatibility with provider SDK iterations: 1. sample.properties.dimension_scores (canonical Foundry runtime shape) 2. sample.properties.rubric_scores (preview/legacy key) 3. top-level sample.dimension_scores / sample.rubric_scores (defensive fallback) Entries missing 'id', 'weight', or 'applicable' are skipped without invalidating well-formed siblings. Non-applicable dimensions may omit 'score' (parsed as null). Adds 6 unit tests covering canonical and legacy keys, top-level fallback, no-match returns null, malformed-entry skipping, and the non-applicable null-score path. 375/375 Foundry tests pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * .NET: feat(samples): Evaluation_FoundryRubric end-to-end sample Adds dotnet/samples/05-end-to-end/Evaluation/Evaluation_FoundryRubric mirroring the Python evaluate_with_rubric_sample.py: - Fetches a pre-existing Foundry agent via AgentAdministrationClient (GetAgentAsync for latest, GetAgentVersionAsync when FOUNDRY_AGENT_VERSION is pinned). - References a rubric evaluator by GeneratedEvaluatorRef(name, version); falls back to GeneratedEvaluatorRef.Latest(name) with the documented floating-version warning. - Mixes the rubric with FoundryEvals.Relevance and FoundryEvals.Coherence in a single FoundryEvals run (implicit string-and-ref conversion). - Prints per-dimension breakdowns from EvalScoreResult.Dimensions for each item. - Demonstrates a CI quality gate with AssertDimensionScoreAtLeast("general_quality", 3.0). Documents the FOUNDRY_PROJECT_ENDPOINT footgun (must be project-scoped URL .../api/projects/<project>, not the bare Azure OpenAI endpoint) and the Eval-Definition-vs-Rubric-Evaluator distinction in the README. Ships a .env.example with the FOUNDRY_* variables. Registers the project in agent-framework-dotnet.slnx and cross-links from the sibling Evaluation_Multimodal / Evaluation_ExpectedOutputs READMEs. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(foundry-evals): harden FoundryEvals public surface for review Address PR #6267 review comments on the .NET FoundryEvals integration: - Add source-compat overloads accepting `string[] evaluators` for `FoundryEvals` ctor, `EvaluateTracesAsync`, and `EvaluateFoundryTargetAsync` so existing callers passing string arrays keep compiling unchanged. New overloads forward via a private `ToSpecs` helper that wraps each name through the implicit `string -> FoundryEvaluatorSpec` conversion. - Guard against `default(FoundryEvaluatorSpec)` entries (both `BuiltinName` and `GeneratedRef` null) that would NRE the downstream converter. Adds `FoundryEvaluatorSpec.IsValid` / `EnsureValid` plus an internal `EnsureAllSpecsValid` helper, wired into the main ctor and both static evaluation entry points. - Add 6 unit tests covering the new validation surface. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(sample): set ExitCode=1 when rubric dimension gate trips PR #6267 review comment: the FoundryRubric sample swallowed the AssertDimensionScoreAtLeast failure, so a CI run that included it as a quality gate would still exit 0 even when the rubric regressed. Set `System.Environment.ExitCode = 1` in the catch so CI fails while still letting the rest of the sample's logging complete cleanly. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * fix(foundry-evals): search typed Sample directly for rubric scores PR #6267 review comment: `_extract_rubric_scores` only searched the `properties` dict when the sample exposed one. When the Azure AI Projects typed SDK returns a Sample object that puts `dimension_scores` / `rubric_scores` directly on the instance (no `properties` wrapper), we missed them and surfaced no per-dimension scores. Add an `else: containers.append(sample)` branch so non-dict typed samples are also inspected for the score keys. Covered by two new tests: one with `dimension_scores` directly on a typed Sample without a `properties` wrapper, and one with the legacy `rubric_scores` key in the same shape. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * test(evals): cover assert_score_at_least and assert_no_failed_items PR #6267 review comments: both assertion helpers shipped without unit tests. Add `TestAssertScoreAtLeast` (above threshold, below w/ offenders, evaluator filter, sub_results recursion) and `TestAssertNoFailedItems` (all passing, failed/errored statuses, sub_results recursion) with a shared `_score_results` fixture builder. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * docs(samples): remove dead rubric-evaluator doc link from FoundryRubric sample The Azure AI Foundry rubric evaluator concept doc page has not yet been published, so the link in the sample README and Program.cs comment 404s. Drop the references until the upstream doc is live. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Potential fix for pull request finding Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> * Address PR 6267 review nits Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Ben Thomas <25218250+alliscode@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Getting started
The getting started samples demonstrate the fundamental concepts and functionality of the agent framework.
Samples
| Sample | Description |
|---|---|
| Agents | Step-by-step instructions for getting started with agents |
| Foundry | Foundry agent samples using FoundryAgent and AIProjectClient.AsAIAgent(...) |
| Agent Providers | Getting started with creating agents using various providers |
| Agents With Retrieval Augmented Generation (RAG) | Adding Retrieval Augmented Generation (RAG) capabilities to your agents |
| Agents With Memory | Adding memory capabilities to your agents |
| Agents With CodeAct (Hyperlight) | Enabling sandboxed code execution (CodeAct) for your agents via Hyperlight |
| Agent Open Telemetry | Getting started with OpenTelemetry for agents |
| OpenAI | Using OpenAI exchange types with agents |
| Anthropic | Getting started with agents using Anthropic Claude |
| Model Context Protocol | Getting started with Model Context Protocol |
| Agent Skills | Getting started with Agent Skills |
| Agent Harness with built-in tools | Demonstrating how to build an Agent Harness with built-in planning, todo, and mode management tooling |
| Declarative Agents | Loading and executing AI agents from YAML configuration files |
| AG-UI | Getting started with AG-UI (Agent UI Protocol) servers and clients |
| Dev UI | Interactive web interface for testing and debugging AI agents during development |
| A2A Agents | Working with Agent-to-Agent (A2A) specific features |