AgentEval 0.43.0-beta
Prefix Reserveddotnet add package AgentEval --version 0.43.0-beta
NuGet\Install-Package AgentEval -Version 0.43.0-beta
<PackageReference Include="AgentEval" Version="0.43.0-beta" />
<PackageVersion Include="AgentEval" Version="0.43.0-beta" />
<PackageReference Include="AgentEval" />
paket add AgentEval --version 0.43.0-beta
#r "nuget: AgentEval, 0.43.0-beta"
#:package AgentEval@0.43.0-beta
#addin nuget:?package=AgentEval&version=0.43.0-beta&prerelease
#tool nuget:?package=AgentEval&version=0.43.0-beta&prerelease
AgentEval
The .NET Evaluation Toolkit for AI Agents
Built first for Microsoft Agent Framework (MAF) and Microsoft.Extensions.AI. What RAGAS and DeepEval do for Python, AgentEval does for .NET.
Preview. AgentEval is experimental: APIs and behavior may change without notice, and breaking changes are listed in the CHANGELOG. Only the Gatekeeper public surface is frozen by a test. Do not use it in production or safety-critical systems without your own review and testing.
Features
- π― Tool Tracking β Monitor tool/function calls with timing, arguments, and ordering
- β
Fluent Assertions β Expressive assertions with rich failure messages,
becausereasons, and assertion scopes - π Performance Metrics β TTFT, latency, tokens, cost estimation for 8+ models
- π¬ RAG Metrics β Faithfulness, relevance, context precision/recall, answer correctness
- π‘οΈ Red Team Security β 14 attack types, 264 probes, full OWASP LLM Top 10 coverage
- πͺ Gatekeeper β Fail-closed runtime enforcement: block forbidden/poisoned tool calls before they run, red-team probes as live guards, and human-in-the-loop approval
- βοΈ Responsible AI β Toxicity, bias, and misinformation detection metrics
- π Stochastic Evaluation β Statistical model comparison with multi-run analysis
- π Trace Record & Replay β Deterministic CI testing without LLM calls
- π― Calibrated Judge β Several LLM judges score the same reply and vote, with their agreement reported
- π Extensible β Adapter pattern for any agent framework
Quick Start
using AgentEval.Assertions;
using AgentEval.Core;
using AgentEval.MAF;
using AgentEval.Models;
// Create evaluation harness (evaluatorClient: the IChatClient that grades replies)
var harness = new MAFEvaluationHarness(evaluatorClient);
// Wrap your Microsoft Agent Framework AIAgent for evaluation
var agent = new MAFAgentAdapter(aiAgent);
// Run evaluation with tool tracking. ModelName selects the price used to estimate cost;
// without a name found in the price table, no cost is estimated.
var result = await harness.RunEvaluationAsync(agent, new TestCase
{
Name = "Feature Planning Test",
Input = "Plan a user authentication feature",
EvaluationCriteria = ["Should include security considerations"]
}, new EvaluationOptions { ModelName = "gpt-4o" });
// Assert tool usage with "because" reasons
result.ToolUsage!
.Should()
.HaveCalledTool("SecurityTool", because: "auth features require security review")
.BeforeTool("FeatureTool")
.WithoutError()
.And()
.HaveNoErrors();
// Assert performance. A metric that was not captured (such as cost with no ModelName) cannot
// fail its check: inside an AgentEvalScope it is recorded as inconclusive, outside one it is skipped.
result.Performance!
.Should()
.HaveTotalDurationUnder(TimeSpan.FromSeconds(10))
.HaveEstimatedCostUnder(0.10m);
Red Team Security Scanning
using AgentEval.RedTeam;
using AgentEval.RedTeam.Reporting;
var result = await AttackPipeline.Create()
.WithAllAttacks()
.ScanAsync(agent);
result.Should()
.HavePassed() // fails on any compromised probe, and on a scan too inconclusive to trust
.HaveMinimumScore(85); // percentage of probes resisted
await new SarifReportExporter().ExportToFileAsync(result, "security-report.sarif");
Trace Record & Replay
Capture agent executions for deterministic replay β no LLM calls needed in CI:
using AgentEval.Tracing;
// Record
await using var recorder = new TraceRecordingAgent(realAgent, "weather_test");
var response = await recorder.InvokeAsync("What's the weather?");
await recorder.SaveAsync("trace.json");
// Replay (deterministic, free)
var trace = await TraceSerializer.LoadFromFileAsync("trace.json");
var replayer = new TraceReplayingAgent(trace);
var replayed = await replayer.InvokeAsync("What's the weather?");
Model Comparison
using AgentEval.Comparison;
var comparer = new ModelComparer(new StochasticRunner(harness));
// CreateAgent(deployment) is your code: it returns an IEvaluableAgent for that model
var results = await comparer.CompareModelsAsync(
factories: new IAgentFactory[]
{
new DelegateAgentFactory("gpt-4o", "GPT-4o", () => CreateAgent("gpt-4o")),
new DelegateAgentFactory("gpt-4o-mini", "GPT-4o Mini", () => CreateAgent("gpt-4o-mini"))
},
testCases: testSuite,
options: new ModelComparisonOptions(RunsPerModel: 5));
Console.WriteLine(results.ToMarkdown());
The quality, speed, cost and reliability scores rank the models against each other (best 100, worst 0), and cost is priced from a single model name. See Model Comparison for how the scores are computed.
Quality Assurance
- The test suite runs in CI on .NET 8, 9 and 10; the build status shows the latest result
Installation
dotnet add package AgentEval --prerelease
Single package, modular internals β the AgentEval package embeds these assemblies; none of them is published as a separate package:
AgentEval.Abstractionsβ Public contracts and interfacesAgentEval.Coreβ Metrics, assertions, comparison, tracingAgentEval.DataLoadersβ Data loading and export (JSON, YAML, CSV, JSONL)AgentEval.MAFβ Microsoft Agent Framework integration, including GatekeeperAgentEval.Memoryβ Memory evaluation, benchmarks, LongMemEval, HTML reportingAgentEval.RedTeamandAgentEval.RedTeam.Gatekeeperβ Security testing, and red-team evaluators as runtime gatesAgentEval.Evals.AgenticandAgentEval.Evals.Performanceβ The agentic evaluator suite and the performance benchmarksAgentEval.Compliance.Core,AgentEval.Compliance.GdprandAgentEval.Compliance.EuAiActβ Compliance benchmarksAgentEval.Rendering.Pdfβ PDF report rendering
The Copilot Studio add-on (AgentEval.MAF.CopilotStudio) is not part of this package and is not on NuGet; build it from the repository if you need it.
Service Registration
// Register all services at once (recommended):
services.AddAgentEvalAll();
// Or register selectively:
services.AddAgentEval(); // Core services only
services.AddAgentEvalDataLoaders(); // DataLoaders + Exporters
services.AddAgentEvalRedTeam(); // Red Team security testing
services.AddAgentEvalMemory(); // Memory evaluation (AgentEval.Memory.Extensions)
services.AddAgentEvalAgentic(); // Agentic evaluator options (AgentEval.Evals.Agentic)
Documentation
- Getting Started
- Fluent Assertions
- Metrics Reference
- Red Team Security
- Gatekeeper (Runtime Enforcement)
- Trace Record & Replay
- Stochastic Evaluation
- Model Comparison
- Architecture
License
MIT License β See LICENSE for details.
Deterministic evals
A deterministic eval is one measurement of an agent run, computed in code β no model, no cost, same
answer every time. Register one with AgentEvalBuilder.AddEval(eval, floor): the door takes the eval
and the chance floor it is judged against, because a score you cannot compare to luck is not a
measurement. See Deterministic evals
for the contract β what the eval sees, what null versus [] tool calls mean, and how to say "this
could not be measured" without saying "this scored zero".
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net8.0 is compatible. net8.0-android was computed. net8.0-browser was computed. net8.0-ios was computed. net8.0-maccatalyst was computed. net8.0-macos was computed. net8.0-tvos was computed. net8.0-windows was computed. net9.0 is compatible. net9.0-android was computed. net9.0-browser was computed. net9.0-ios was computed. net9.0-maccatalyst was computed. net9.0-macos was computed. net9.0-tvos was computed. net9.0-windows was computed. net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
-
net10.0
- JsonSchema.Net (>= 7.3.4)
- Microsoft.Agents.AI (>= 1.23.0)
- Microsoft.Agents.AI.Workflows (>= 1.23.0)
- Microsoft.Extensions.AI (>= 10.10.0)
- Microsoft.Extensions.AI.Evaluation.Quality (>= 10.10.0)
- Microsoft.Extensions.DependencyInjection (>= 10.0.8)
- OpenTelemetry.Api (>= 1.18.0)
- PdfSharp-MigraDoc (>= 6.2.4)
- QuestPDF (>= 2026.2.4)
- System.Numerics.Tensors (>= 10.0.12)
- YamlDotNet (>= 16.3.0)
-
net8.0
- JsonSchema.Net (>= 7.3.4)
- Microsoft.Agents.AI (>= 1.23.0)
- Microsoft.Agents.AI.Workflows (>= 1.23.0)
- Microsoft.Extensions.AI (>= 10.10.0)
- Microsoft.Extensions.AI.Evaluation.Quality (>= 10.10.0)
- Microsoft.Extensions.DependencyInjection (>= 10.0.8)
- OpenTelemetry.Api (>= 1.18.0)
- PdfSharp-MigraDoc (>= 6.2.4)
- QuestPDF (>= 2026.2.4)
- System.Numerics.Tensors (>= 10.0.12)
- YamlDotNet (>= 16.3.0)
-
net9.0
- JsonSchema.Net (>= 7.3.4)
- Microsoft.Agents.AI (>= 1.23.0)
- Microsoft.Agents.AI.Workflows (>= 1.23.0)
- Microsoft.Extensions.AI (>= 10.10.0)
- Microsoft.Extensions.AI.Evaluation.Quality (>= 10.10.0)
- Microsoft.Extensions.DependencyInjection (>= 10.0.8)
- OpenTelemetry.Api (>= 1.18.0)
- PdfSharp-MigraDoc (>= 6.2.4)
- QuestPDF (>= 2026.2.4)
- System.Numerics.Tensors (>= 10.0.12)
- YamlDotNet (>= 16.3.0)
NuGet packages
This package is not used by any NuGet packages.
GitHub repositories
This package is not used by any popular GitHub repositories.
| Version | Downloads | Last Updated |
|---|---|---|
| 0.43.0-beta | 86 | 10/6/2026 |
| 0.42.0-beta | 213 | 10/1/2026 |
| 0.41.0-beta | 4,866 | 9/21/2026 |
| 0.40.0-beta | 73 | 9/20/2026 |
| 0.39.0-beta | 77 | 9/16/2026 |
| 0.38.0-beta | 248 | 9/15/2026 |
| 0.37.0-beta | 77 | 9/14/2026 |
| 0.36.0-beta | 102 | 9/13/2026 |
| 0.35.0-beta | 1,212 | 9/8/2026 |
| 0.34.0-beta | 96 | 9/4/2026 |
| 0.33.0-beta | 122 | 9/2/2026 |
| 0.32.0-beta | 98 | 9/1/2026 |
| 0.31.0-beta | 144 | 8/30/2026 |
| 0.30.0-beta | 91 | 8/29/2026 |
| 0.29.0-beta | 146 | 8/28/2026 |
| 0.28.0-beta | 99 | 8/22/2026 |
| 0.27.0-beta | 83 | 8/22/2026 |
| 0.26.0-beta | 113 | 8/19/2026 |
| 0.25.0-beta | 98 | 8/17/2026 |
| 0.24.0-beta | 95 | 8/17/2026 |
Per-version release notes: https://github.com/AgentEvalHQ/AgentEval/blob/main/CHANGELOG.md