From Edge to Enterprise Cloud: Part 4 — GenAIOps, Automated Evaluations and Observability with Microsoft Foundry
In our previous posts, From Edge to Enterprise-Cloud: Part 1: Building a Cost-Free Local Loop with Foundry Local, From Edge to Enterprise Cloud: Part 2: Tool Orchestration with the Model Context Protocol and From Edge to Enterprise Cloud: Part 3: Cloud Migration and Enterprise Identity with Microsoft Foundry Agent Service, we built a hardware-accelerated local loop, measured how schema hardening changes tool calling reliability on Small Language Models and moved the agent into Microsoft Foundry with Bicep and Microsoft Entra ID.
Part 3 ended with a question: how do we keep the agent correct when the model, the system prompt or a tool schema changes? Part 2 already showed why that question is not academic. A bigger model from the same family, qwen2.5-7b, dropped the name field of a nested object in ten out of ten calls, while its smaller sibling got most of them right. Nothing crashed. Nothing in a standard CI pipeline would have noticed.
Executive summary: An AI agent can get worse without a single failing unit test. This final part adds the missing safety net. A golden dataset, built from the Part 2 study, runs on every pull request. A deterministic check that costs nothing blocks the merge when tool calls start failing. An optional LLM judge adds a second opinion. The observability infrastructure is provisioned with Bicep and uses no keys and no stored secrets. Every run can be traced from a developer laptop to Application Insights. For decision makers this means model upgrades become a measured, reviewable change instead of a leap of faith.
All code in this post lives in the part04-observability-evals folder of our companion repository, Frezz146/from-edge-to-enterprise-cloud, in Python and C#.
The Three Pillars: Evaluate, Trace and Monitor
Operating agents means applying DevOps discipline to software that is not deterministic. Microsoft frames observability for Foundry agents in three pillars. They map cleanly onto this part:
| Pillar | Question it answers | What this part builds |
|---|---|---|
| Evaluate | Is this change good enough to ship? | A golden dataset and a quality gate in GitHub Actions |
| Trace | What exactly happened in this one run? | OpenTelemetry spans to an Aspire Dashboard or Application Insights |
| Monitor | Is quality drifting over time in production? | The infrastructure for it, plus an outlook at the end |
One principle carries over from the whole series: measured, not assumed. Part 4 does not introduce a new test suite written for the occasion. It reuses the exact cases, the exact MCP server and the exact error taxonomy that produced the Part 2 numbers. The quality gate imports that code instead of copying it, so the study and the gate can never drift apart.

Infrastructure as Code: Extending the Part 3 Bicep
Observability infrastructure deserves the same rigor as the agent itself, so it is Bicep again. The Part 4 template does not recreate anything from Part 3. It references the Foundry account, the project and the managed identity as existing resources and derives their names with the same uniqueString() formula Part 3 used. The only value to keep in sync is namePrefix.
It adds four blocks:
- Log Analytics and Application Insights with
DisableLocalAuth: true. Ingestion keys stop working. Every span has to arrive with a Microsoft Entra ID token. - The Application Insights connection on the Foundry account and project. This switches on tracing for agents Foundry runs. It is also where our client code looks up the connection string.
- Two OIDC federated credentials on the Part 3 identity: one for pull requests and one for
main. GitHub Actions signs in as that identity without any stored secret. - Least-privilege role assignments, one per job to be done.
📄 Bicep Infrastructure Template (part04-observability-evals/infra/main.bicep)
“ApiKey” on a Keyless Resource?
This one surprised me while building it. It is worth a paragraph because it looks like a contradiction. My first draft declared the connection with Entra ID authentication. A look into the SDK showed why that cannot work: when a client asks the project for its connection string, azure-ai-projects only accepts an ApiKey style connection and raises otherwise. Microsoft’s own foundry-samples template uses exactly the shape shown above.
The resolution is DisableLocalAuth. With local authentication disabled, the connection string is only an address. It names the resource, but Application Insights rejects any ingestion that does not carry an Entra ID token. So the connection tells the client where to send data. The identity plus its role decide whether it may.
CAUTION: For agents that Foundry itself runs (published prompt agents, hosted agents), server-side trace ingestion into an Entra-only Application Insights resource requires the connection’s auth type Project managed identity. That option is currently in public preview. The template already grants the project identity the role it needs. Switch the auth type in the portal if you use that path. Alternatively, deploy with
disableLocalAuth=false.
Who Gets Which Role
Instead of one broad role per identity, every assignment answers one question:
| Identity | Role | Scope | Why |
|---|---|---|---|
| Agent identity (also CI via OIDC) | Monitoring Metrics Publisher | Application Insights | Write spans with an Entra ID token |
| Agent identity | Cognitive Services OpenAI User | Foundry account | Cloud gate and LLM judge |
| Project managed identity | Monitoring Metrics Publisher, Log Analytics Reader, Privileged Monitoring Data Reader | Application Insights | Foundry’s own traces and evaluations |
| You (optional parameter) | Monitoring Metrics Publisher, Log Analytics Reader, Cognitive Services OpenAI User | Application Insights, account | Run everything locally and browse traces |
Note what is missing: the CI identity can write telemetry but not read it. A pipeline that only writes has no reason to see production traces, which may contain customer content.
az deployment group create \
--resource-group <your-resource-group> \
--template-file infra/main.bicep \
--parameters infra/main.bicepparam \
--parameters developerPrincipalId=$(az ad signed-in-user show --query id -o tsv)
The outputs are the client ID, tenant ID, subscription ID and the Azure OpenAI endpoint. None of them is a secret. They go into GitHub as repository variables, which is the whole point.
The Golden Dataset: Reusing the Schema Hardening Scenarios
A golden dataset is only as useful as the failures it is known to catch. Ours has a head start: the ten booking scenarios from Part 2 already caught a real regression. They move from a Python list into golden_dataset.json, next to the thresholds the gate enforces. Both language tracks read the same file.
{
"tool": "book_appointment_strict",
"quality_gate": {
"min_clean_call_rate": 0.8,
"min_tool_call_accuracy": 4.0,
"min_judge_pass_rate": 0.8
},
"cases": [
{
"id": "case_08_standard_priority",
"query": "Schedule a standard-priority 30-minute sync with Riley Nguyen (riley.nguyen@example.com) covering the Q4 planning doc.",
"expected_tool": "book_appointment_strict",
"expected_parameters": {
"attendee": { "name": "Riley Nguyen", "email": "riley.nguyen@example.com" },
"topics": ["Q4 planning doc"],
"priority": "medium",
"duration_minutes": 30
}
}
]
}
Two design decisions matter here. The gate targets only the strict schema, because that is the one you ship after reading Part 2. And the thresholds live with the data, not in code: tightening the gate is a reviewed one-line change to a JSON file.
A practical tip on thresholds: do not guess them. Run the gate against your current model and anchor the threshold to that measured baseline, as shown in the baseline section further down. A gate that fails on the current main teaches the team to ignore it.
Evaluators: A Deterministic Taxonomy First, an LLM Judge Second
For every case the pipeline does four things:
- It sends the query and the
book_appointment_strictschema to the model under test, either a local Foundry Local model or thegpt-5-minideployment from Part 3. - It executes the returned call against Part 2’s own MCP server.
SERVER_REJECTEDtherefore means exactly what it meant in Part 2: Pydantic refused the arguments before the handler ran. - It classifies the outcome with Part 2’s error taxonomy (
OK,NO_TOOL_CALL,WRONG_TOOL,MALFORMED_JSON,SERVER_REJECTED,WRONG_NAME,WRONG_EMAIL,WRONG_PRIORITY,WRONG_DURATION,MISSING_TOPIC). - With
--judge, it asksToolCallAccuracyEvaluatorfromazure-ai-evaluationfor a score between 1 and 5.
Why this order? The taxonomy is deterministic, runs in milliseconds and costs nothing. The same model answer always gets the same verdict, which is the property you want from something that blocks a merge. The judge is a model: it catches problems no rule anticipated, but it costs tokens and its scores vary slightly from run to run. So the taxonomy is the gate and the judge is the second opinion, applied to the average over the dataset rather than to single cases.
🐍 Python Evaluation Pipeline (part04-observability-evals/python/eval_pipeline.py)
# Imported, not copied: this is the code that produced the Part 2 numbers.
from schema_hardening_eval import (
SYSTEM_PROMPT,
Case,
classify_semantic,
load_local_model,
mcp_tool_to_foundry_tool,
parse_fallback_tool_call,
)
async def run_case(complete_chat, session, tool, case) -> CaseResult:
tool_name = tool["function"]["name"]
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": case["query"]},
]
message = complete_chat(messages, tools=[tool]).choices[0].message
if message.tool_calls:
call = message.tool_calls[0]
call_name = call.function.name
try:
args = json.loads(call.function.arguments)
except json.JSONDecodeError:
return CaseResult(case["id"], ["MALFORMED_JSON"])
else:
# Part 2's fallback for models that write tool calls as tagged text
properties = tool["function"]["parameters"].get("properties", {})
param_types = {key: prop.get("type") for key, prop in properties.items()}
fallback = parse_fallback_tool_call(message.content or "", param_types)
if fallback is None:
return CaseResult(case["id"], ["NO_TOOL_CALL"])
call_name, args = fallback
tool_call = {"name": call_name, "arguments": args}
if call_name != tool_name:
return CaseResult(case["id"], ["WRONG_TOOL"], tool_call)
result = await session.call_tool(tool_name, arguments=args) # Part 2's MCP server
if result.isError:
return CaseResult(case["id"], ["SERVER_REJECTED"], tool_call, ...)
errors = classify_semantic(
tool_name, args, to_part2_case(case)
) # Part 2's taxonomy
return CaseResult(case["id"], errors or ["OK"], tool_call)
The judge is authenticated like everything else in this series, with DefaultAzureCredential instead of an API key:
def build_judge(threshold: float):
deployment = os.environ.get("JUDGE_MODEL", "gpt-5-mini")
model_config = {
"azure_endpoint": require_env("AZURE_OPENAI_ENDPOINT"),
"azure_deployment": deployment,
"api_version": "2025-04-01-preview", # the evaluator default predates reasoning models
}
return ToolCallAccuracyEvaluator(
model_config=model_config,
credential=DefaultAzureCredential(), # not inside model_config, see below
threshold=threshold,
# gpt-5 family and o-series reject the sampling settings in the judge prompt
is_reasoning_model=deployment.startswith(("gpt-5", "o1", "o3", "o4")),
)
The credential placement is not a style choice. The SDK’s AzureOpenAIModelConfiguration type lists a credential key, but in azure-ai-evaluation 1.18 a config that contains it fails the SDK’s own validation with “Model config validation failed”. My first verification run stopped exactly there. Passed as the evaluator’s own credential argument, it works.
Two more details are easy to miss. A case without a parseable tool call gets the rubric’s floor score of 1, so a model that simply stops calling tools cannot raise the average. And ideally the judge is not the model under test. The demo uses one deployment for both to keep costs down. In a real setup, deploy a second model for the judge.
C# Parity with Microsoft.Extensions.AI.Evaluation
The C# track (csharp/eval-runner) mirrors the Python gate: same golden dataset, same taxonomy, Part 2’s C# MCP server spawned over stdio. The judge comes from Microsoft.Extensions.AI.Evaluation.Quality, which ships a ToolCallAccuracyEvaluator of its own.
One difference matters when you compare numbers across languages: the .NET evaluator returns a BooleanMetric, accurate or not, while the Python evaluator returns a score from 1 to 5. A single threshold cannot serve both, which is why the dataset carries min_tool_call_accuracy for Python and min_judge_pass_rate for .NET. After Part 2 found that “strict” was not equally strict in both languages, this is the same lesson at the evaluation layer: parity has to be verified, not assumed.
The verification run then reproduced the other open question from Part 2. The C# runner uses the hand-authored strict schema from Part 2’s C# track. Against it, qwen2.5-1.5b scored 0/10: it dropped attendee.name in every single call and durationMinutes in three of them. The same weights score 8/10 against the Python server’s Pydantic schema. Part 2 called the cause an open hypothesis about schema presentation. Part 4 now shows the same effect with a completely different harness. The C# runner therefore defaults to qwen3.5-4b, which scores 9/10 in C#, matching its Part 2 result. A gate is only useful with a model that passes it on main.
CAUTION, experimental APIs (AIEVAL001, OPENAI001): the quality evaluators in
Microsoft.Extensions.AI.Evaluation.Qualityare marked[Experimental("AIEVAL001")]. The first real build also failed onOPENAI001: the OpenAIChatClientconstructor that accepts anAuthenticationPolicy, which is how the judge signs in with Entra ID instead of a key, is experimental as well. The project acknowledges both once with<NoWarn>$(NoWarn);AIEVAL001;OPENAI001</NoWarn>in the.csprojinstead of scattering pragmas through the code. Treat them as preview components in a production pipeline.
var openAiChatClient = new global::OpenAI.Chat.ChatClient(
judgeDeployment,
new BearerTokenPolicy(new DefaultAzureCredential(), "https://ai.azure.com/.default"),
new global::OpenAI.OpenAIClientOptions { Endpoint = new Uri($"{endpoint.TrimEnd('/')}/openai/v1/") });
judgeConfiguration = new ChatConfiguration(
new ReasoningModelCompatibleChatClient(AI.OpenAIClientExtensions.AsIChatClient(openAiChatClient)));
judge = new ToolCallAccuracyEvaluator();
judgeContext = new ToolCallAccuracyEvaluatorContext(
AI.AIFunctionFactory.CreateDeclaration(
strictTool.Function!.Name!,
strictTool.Function.Description,
JsonSerializer.SerializeToElement(strictTool.Function.Parameters)));
// Per case: wrap the local model's call as a FunctionCallContent and let the judge decide
var modelResponse = new AI.ChatResponse(new AI.ChatMessage(
AI.ChatRole.Assistant,
[new AI.FunctionCallContent($"call_{goldenCase.Id}", result.CallName, result.Arguments)]));
var evaluation = await judge.EvaluateAsync(conversation, modelResponse, judgeConfiguration, [judgeContext], ct);
bool? accurate = evaluation.Get<booleanmetric>(ToolCallAccuracyEvaluator.ToolCallAccuracyMetricName).Value;
The .NET evaluator sets Temperature, TopP and a tight output budget on its judge requests. Reasoning models such as gpt-5-mini reject non-default sampling settings. Reasoning tokens can also use up a small budget before the verdict is written. A ten-line DelegatingChatClient strips those options for the judge, the .NET counterpart to Python’s is_reasoning_model flag.
Quality Gates in GitHub Actions
The workflow runs two jobs on every pull request that touches Part 2’s Python track, Part 4 or the workflow itself:
| Job | Model under test | Needs Azure? | Purpose |
|---|---|---|---|
local-gate | Foundry Local on the runner | No | Deterministic taxonomy, free, also runs for forks |
cloud-gate | gpt-5-mini from Part 3 | Yes, via OIDC | Taxonomy plus LLM judge, spans to Application Insights |
env:
# The model under test. Changing this line in a pull request is exactly
# the kind of change the gate exists to check.
FOUNDRY_LOCAL_MODEL: ${{ inputs.local_model || 'qwen2.5-1.5b' }}
jobs:
local-gate:
runs-on: ubuntu-latest
steps:
# (checkout, setup-python and pip install omitted)
- name: Local quality gate (${{ env.FOUNDRY_LOCAL_MODEL }})
run: python eval_pipeline.py --backend local --model "$FOUNDRY_LOCAL_MODEL"
cloud-gate:
if: ${{ vars.AZURE_CLIENT_ID != '' }}
runs-on: ubuntu-latest
permissions:
contents: read
id-token: write # lets azure/login exchange the GitHub OIDC token, no client secret
steps:
# (checkout, setup-python and pip install omitted)
- uses: azure/login@v2
with:
client-id: ${{ vars.AZURE_CLIENT_ID }}
tenant-id: ${{ vars.AZURE_TENANT_ID }}
subscription-id: ${{ vars.AZURE_SUBSCRIPTION_ID }}
- name: Cloud quality gate (gpt-5-mini, LLM judge, traced to Application Insights)
run: python eval_pipeline.py --backend cloud --judge --trace azure
Everything the cloud job needs is a vars.* value, not a secrets.* value. Client ID, tenant ID and subscription ID are identifiers. The federated credential from the Bicep template is what turns them into a token. There is nothing in the repository settings an attacker could reuse outside this repository and these two triggers.
The first real cloud-gate run still failed at azure/login, with AADSTS700213: No matching federated identity record found for presented assertion subject. The error message quoted the subject GitHub had sent: repo:Frezz146@115118426/from-edge-to-enterprise-cloud@1360166622:ref:refs/heads/main. That is GitHub’s immutable subject format, which includes the numeric owner and repository IDs so that a renamed or recreated repository cannot inherit the trust. The federated credential used the familiar repo:<owner>/<name>:... format and therefore never matched. The Bicep template now takes both IDs as parameters (gh api repos/<owner>/<name> --jq ".owner.id, .id" prints them) and falls back to the legacy format when they are empty. It is a one-line mistake that no local test can catch, which is exactly why a pipeline has to run in the pipeline at least once. After redeploying the credentials both jobs passed on GitHub: local-gate with 8/10 and cloud-gate with 10/10 and a judge score of 5.00.
The local-gate job needs no Azure access at all, so it also protects pull requests from forks, which GitHub does not give an OIDC token. Foundry Local falls back to CPU execution providers on a hosted runner. That is slower than a developer machine. A self-hosted runner is faster and keeps the model cache between runs. Once both jobs are green on main, make them required status checks in a branch protection rule.
The Baseline: What qwen2.5-1.5b Scores Today
Before a gate can block anything, it needs a baseline. This is the local gate against qwen2.5-1.5b on my developer machine, server messages shortened:
=== local (qwen2.5-1.5b) | 10 golden cases | tool book_appointment_strict ===
[case_01_explicit] OK
[case_02_two_topics] OK
[case_03_informal_low_key] OK
[case_04_one_hour] OK
[case_05_three_topics] OK
[case_06_urgent] OK
[case_07_no_rush] SERVER_REJECTED
server said: ... priority Input should be 'low', 'medium' or 'high' [type=literal_error, input_value='normal', input_type=str] ...
[case_08_standard_priority] SERVER_REJECTED
server said: ... priority Input should be 'low', 'medium' or 'high' [type=literal_error, input_value='standard', input_type=str] ...
[case_09_two_low_topics] OK
[case_10_critical] OK
=== Quality gate summary ===
Backend | Clean calls | Judge | Errors
------------------------------------------------------------------------------------------
local (qwen2.5-1.5b) | 8/10 | - | SERVER_REJECTED x2
clean call rate 80% (min 80%)
Quality gate PASSED.
Both misses are the same kind of mistake. The model invented priority values the enum does not allow: “no-rush” became normal and “standard-priority” became standard. Everything else was right in all ten cases, including the nested attendee object, the topic arrays and the durations. That matches the Part 2 Python result for this model (8/10, SERVER_REJECTED x2) and Part 2’s fourth discovery: informal priority wording is the hardest field for small models. With the loose schema both values would have been booked silently as WRONG_PRIORITY. The strict enum turns them into loud rejections that a gate can count.
The LLM judge agrees. Run with --judge, it scored every clean call 5.0 and both rejected calls 2.0, for a mean of 4.40 out of 5. The cloud deployment from Part 3 sets the upper bound: gpt-5-mini passes all ten cases with a judge score of 5.00, mapping “no-rush” to low and “standard-priority” to medium as intended.
| Model under test | Track | Clean calls | Judge |
|---|---|---|---|
qwen2.5-1.5b (local) | Python | 8/10 | 4.40/5 |
gpt-5-mini (cloud) | Python | 10/10 | 5.00/5 |
qwen3.5-4b (local) | C# | 9/10 | 8 and 9 of 10 accurate in two runs |
qwen2.5-1.5b (local) | C# | 0/10 | 0 of 10 accurate |
Because the local run uses temperature 0, the verdict is repeatable. The gate therefore sits exactly at the baseline: min_clean_call_rate: 0.8 means “no worse than today”. One open question was the CI runner: it uses CPU execution providers instead of the GPU on a developer machine. A different execution provider can change greedy decoding in edge cases. The first CI run answered it. On a hosted Ubuntu runner the local gate produced the same 8/10 with the same two rejected cases, so one threshold serves both.
Live Demonstration: Catching the Part 2 Regression in CI
Now picture a reasonable pull request. A developer bumps FOUNDRY_LOCAL_MODEL from qwen2.5-1.5b to qwen2.5-7b, expecting a bigger model in the same family to be at least as good. Part 2 already told us what that model does with the strict schema: it omits attendee.name in every single call. This time nobody has to notice. Here is the real output of the exact command the local-gate job runs, executed on my machine (MCP server log lines removed, validation messages shortened):
[Local] Loading Foundry Local model 'qwen2.5-7b'...
capabilities='tool-calling' supports_tool_calling=True
=== local (qwen2.5-7b) | 10 golden cases | tool book_appointment_strict ===
[case_01_explicit] SERVER_REJECTED
server said: ... 1 validation error: attendee.name Field required [type=missing, input_value={'email': 'jamie.chen@example.com'}, input_type=dict]
[case_02_two_topics] SERVER_REJECTED
server said: ... 1 validation error: attendee.name Field required [type=missing, input_value={'email': 'alex.kim@example.com'}, input_type=dict]
[case_03_informal_low_key] SERVER_REJECTED
server said: ... 1 validation error: attendee.name Field required [type=missing, input_value={'email': 'sam.rivera@example.com'}, input_type=dict]
[case_04_one_hour] SERVER_REJECTED
server said: ... 1 validation error: attendee.name Field required [type=missing, input_value={'email': 'morgan.lee@example.com'}, input_type=dict]
[case_05_three_topics] SERVER_REJECTED
server said: ... 1 validation error: attendee.name Field required [type=missing, input_value={'email': 'taylor.brooks@example.com'}, input_type=dict]
[case_06_urgent] SERVER_REJECTED
server said: ... 1 validation error: attendee.name Field required [type=missing, input_value={'email': 'priya.patel@example.com'}, input_type=dict]
[case_07_no_rush] SERVER_REJECTED
server said: ... 1 validation error: attendee.name Field required [type=missing, input_value={'email': 'devon.clarke@example.com'}, input_type=dict]
[case_08_standard_priority] SERVER_REJECTED
server said: ... 2 validation errors: attendee.name Field required [type=missing, input_value={'email': 'riley.nguyen@example.com'}, input_type=dict] priority Input should be 'low', 'medium' or 'high' [type=literal_error, input_value='standard', input_type=str]
[case_09_two_low_topics] SERVER_REJECTED
server said: ... 1 validation error: attendee.name Field required [type=missing, input_value={'email': 'casey.morgan@example.com'}, input_type=dict]
[case_10_critical] SERVER_REJECTED
server said: ... 1 validation error: attendee.name Field required [type=missing, input_value={'email': 'jordan.ellis@example.com'}, input_type=dict]
=== Quality gate summary ===
Backend | Clean calls | Judge | Errors
------------------------------------------------------------------------------------------
local (qwen2.5-7b) | 0/10 | - | SERVER_REJECTED x10
clean call rate 0% (min 80%)
Quality gate FAILED.
Results written to evaluation_results.json
Ten out of ten calls drop attendee.name, exactly the pattern Part 2 documented. Case 8 even shows both failure modes at once: the missing name of the 7B model and the invented standard priority that already tripped up the 1.5B model. In CI the check turns red. The run summary page shows a table with every case and its outcome. evaluation_results.json is attached as an artifact with the rejected arguments and the server’s validation message. The merge is blocked before any user books a meeting with a blank name.
This is the payoff of the whole series in one screen. Part 2 found the failure with a careful study. Part 4 makes finding it the default.
Tracing Across the Hybrid Stack
Evaluations tell you whether something is wrong. Traces tell you what happened. Microsoft Foundry distinguishes two levels. Our agent needs both:
- Server-side tracing covers agents that Foundry runs, such as published prompt agents and hosted agents. Once the Application Insights connection from the Bicep template exists, Foundry traces them without any code change.
- Client-side tracing covers code that runs in your own process. The Part 3 agent’s
agent.run()loop is exactly that: Agent Framework calls the model in Foundry, but the loop, the function tools and the session live in our Python process.
Agent Framework already instruments agent runs, model calls and tool invocations with the OpenTelemetry GenAI semantic conventions. The only decision left is where the spans go. That is where the hybrid theme of the series returns.
Inner Loop: an Aspire Dashboard on Your Machine
During development you want traces the way Part 1 wanted inference: local, instant and free. The standalone .NET Aspire Dashboard accepts OpenTelemetry over OTLP and runs in a single container:
docker run --rm -it -d -p 18888:18888 -p 4317:18889 -p 4318:18890 \
-e ASPIRE_DASHBOARD_UNSECURED_ALLOW_ANONYMOUS=true \
--name aspire-dashboard mcr.microsoft.com/dotnet/aspire-dashboard:latest
python traced_agent.py --target otlp # the Part 3 agent, fully traced
python eval_pipeline.py --trace otlp # the quality gate, one span per case
Open http://localhost:18888 and you see the whole conversation as one trace: the parent span, both turns, every model call and every get_weather or calculate invocation with its duration. The first docker run pulls the dashboard image once. After that the inner loop starts in seconds.
Outer Loop: Application Insights, without a Connection String in Code
In the cloud the same spans go to the Application Insights resource from the Bicep template. Notice what is not in this code: a connection string, a key or an environment variable pointing at Application Insights.
from cloud_agent import AGENT_NAME, calculate, get_weather # Part 3, unchanged
credential = DefaultAzureCredential()
client = FoundryChatClient(credential=credential)
# Reads the connection string from the project's Application Insights connection
# (infra/main.bicep) and exports with an Entra ID token, since local auth is disabled.
await client.configure_azure_monitor(enable_sensitive_data=False, credential=credential)
agent = Agent(
client=client, name=AGENT_NAME, instructions=..., tools=[get_weather, calculate]
)
with get_tracer().start_as_current_span("part04.traced_conversation") as span:
session = agent.create_session()
for question in QUESTIONS:
result = await agent.run(question, session=session)
span.set_attribute("gen_ai.conversation.id", session.service_session_id or "")
configure_azure_monitor() asks the Foundry project for the Application Insights connection the Bicep template created. It then exports every span with the Entra ID token from DefaultAzureCredential. The script prints the trace ID at the end, so you can jump straight to it in Transaction search. In the verification run the agent answered both turns (Tokyo at 18 °C, 42 × 17 = 714, then 64.4 °F) and its spans arrived in an Application Insights resource that accepts no ingestion key at all. The quality gate does the same with --trace azure, which is how the cloud-gate job in CI shows up next to the agent’s own traces.
Data Privacy and Content Recording
Spans can carry prompts, responses and tool arguments. In enterprise environments, particularly in the DACH region and across the EU, storing that content in telemetry tables is a decision that needs an owner, not a default somebody forgot to change. The switch depends on the library that emits the spans:
| Stack | Switch | Default |
|---|---|---|
Agent Framework (traced_agent.py) | enable_sensitive_data or ENABLE_SENSITIVE_DATA=true | off |
Azure AI SDK instrumentors (azure-ai-projects, azure-ai-inference) | AZURE_TRACING_GEN_AI_CONTENT_RECORDING_ENABLED=true | off |
Our quality gate (eval_pipeline.py) | none, it never records content | off |
Watch the exact variable name for the Azure SDKs: it ends in _ENABLED. Without that suffix the SDK ignores it and records nothing, which is the safe failure but rarely the intended one. The gate itself only writes case IDs, outcomes and scores into its spans. Keep content recording for dev and staging projects where it has been agreed on. Also remember who can read it: Log Analytics Reader on the Application Insights resource or Privileged Monitoring Data Reader when the underlying tables are protected.
From Pipeline to Production: Continuous Evaluation and Alerts
GenAIOps does not stop at the merge button. Users phrase things differently than a golden dataset does. Quality can drift without any code change. The infrastructure from this part is the foundation for closing that loop: sample live conversations, score them on a schedule with the same evaluators and raise an Azure Monitor alert when the rolling average falls below a threshold. Foundry offers continuous evaluation for the agents it runs.
I deliberately left this out of the repository. A demo agent has no real traffic. An alert on synthetic traffic would teach the wrong lesson. When your agent has users, this is the next step. From then on the golden dataset grows from the production traces that surprised you.
What Building Part 4 Taught Me
Like Part 2, this part produced findings that only showed up while building and running it end to end:
- The judge is a second opinion, not a gate: Two C# runs with the same model at temperature 0 produced 8 and then 9 accurate verdicts, one case simply flipped. The deterministic taxonomy never flips. Where both looked at the same calls, they agreed on every rejected one.
- Schema presentation is real:
qwen2.5-1.5bscores 8/10 against the Python schema and 0/10 against the hand-written C# schema, dropping the same nested field every time. Part 2 raised the hypothesis. Part 4 reproduced it with a different harness. - Read the SDK, then run it:
azure-ai-evaluation1.18 rejects its own model config when it contains acredential. Its evaluators also default to an API version older than any reasoning model. Neither shows up until the first real run. - A keyless resource can still need an “ApiKey” connection: The SDK lookup for the connection string only accepts that shape.
DisableLocalAuthis what keeps it safe, because the string alone cannot ingest anything. - “The same evaluator” is not the same metric: Python scores 1 to 5, .NET returns true or false. Cross-language dashboards need a mapping, not a shared threshold.
- OIDC subjects are exact strings: GitHub sent this repository’s subject in the immutable format with numeric IDs. Only the first real CI run revealed it. The error message itself contained the fix.
- A judge needs capacity: With the gate, the cloud backend and the judge sharing one 10,000 TPM deployment, back to back runs hit
429rate limits. The SDK retried and finished, but a real pipeline deserves a separate judge deployment. - Reuse beats rewriting: Importing the Part 2 MCP server and taxonomy into the gate took a few lines. It guarantees that the regression the study found is the regression the gate catches.
The Grand Finale: The Complete 4-Part Hybrid Agent Architecture
Four parts ago we started with a 0.5B model on a laptop and a question about latency and cost. We end with an agent that runs in Microsoft Foundry, authenticates without secrets and cannot be made worse by a pull request without somebody noticing.
| Architecture layer | Part 1: Local Loop | Part 2: MCP Orchestration | Part 3: Cloud Migration | Part 4: GenAIOps & Observability |
|---|---|---|---|---|
| Execution domain | On-device (workstation, NPU/GPU) | Local stdio subprocesses | Cloud-hosted (Microsoft Foundry) | Hybrid (local and cloud) |
| Core capability | In-process streaming inference | Standardized tool transport | Scalable agent lifecycle | Quality gates and distributed tracing |
| Security & identity | Local user permissions | Local environment variables | Microsoft Entra ID, managed identity | OIDC federation, Entra-only telemetry |
| Primary technology | Foundry Local GA SDK | Model Context Protocol | Bicep, Agent Framework, Foundry Toolbox | azure-ai-evaluation, M.E.AI.Evaluation, OpenTelemetry, GitHub Actions |
| Problem solved | Cloud cost and WAN latency in the dev loop | Silent tool-call errors through schema hardening | Local process ceilings and secret sprawl | Silent regressions after model, prompt or schema changes |
Key Architectural Lessons Learned
- Local-first accelerates everything: Foundry Local made thousands of iterations free. The same local model now runs the cheapest quality gate in the pipeline.
- Schema hardening trades silent errors for loud ones: Strict schemas turn wrong data into rejected calls. Rejected calls are something a gate can count.
- Identity over secrets: From Part 3 on, nothing authenticates with a key:
DefaultAzureCredentialfor the agent, OIDC for the pipeline, Entra-only ingestion for telemetry. The repository stores no key and no client secret. - Evaluations are your compiler: A model upgrade is a code change. It deserves the same automated check before it reaches
main.
The complete, runnable source code for all four parts is available in our official companion repository: Frezz146/from-edge-to-enterprise-cloud.
Thank you for following along through all four parts. Happy enterprise coding!