Skip to content
Frezz Tech logo with a futuristic AI robot and cloud circuitry, representing artificial intelligence, Microsoft Copilot, Azure AI, and cloud innovation. Frezz Tech logo with a futuristic AI robot and cloud circuitry, representing artificial intelligence, Microsoft Copilot, Azure AI, and cloud innovation. Frezz Tech

Real-world AI, Copilot and Azure architecture

Frezz Tech logo with a futuristic AI robot and cloud circuitry, representing artificial intelligence, Microsoft Copilot, Azure AI, and cloud innovation. Frezz Tech logo with a futuristic AI robot and cloud circuitry, representing artificial intelligence, Microsoft Copilot, Azure AI, and cloud innovation. Frezz Tech

Real-world AI, Copilot and Azure architecture

  • Home
  • Blog
  • About
  • Home
  • Blog
  • About
Close

Search

  • Home
  • Blog
  • About
Frezz Tech logo with a futuristic AI robot and cloud circuitry, representing artificial intelligence, Microsoft Copilot, Azure AI, and cloud innovation. Frezz Tech logo with a futuristic AI robot and cloud circuitry, representing artificial intelligence, Microsoft Copilot, Azure AI, and cloud innovation. Frezz Tech

Real-world AI, Copilot and Azure architecture

Frezz Tech logo with a futuristic AI robot and cloud circuitry, representing artificial intelligence, Microsoft Copilot, Azure AI, and cloud innovation. Frezz Tech logo with a futuristic AI robot and cloud circuitry, representing artificial intelligence, Microsoft Copilot, Azure AI, and cloud innovation. Frezz Tech

Real-world AI, Copilot and Azure architecture

  • Home
  • Blog
  • About
  • Home
  • Blog
  • About
Close

Search

  • Home
  • Blog
  • About
Technical architecture diagram illustrating end-to-end GenAIOps pipeline, inner loop local evaluation, OIDC GitHub Actions CI/CD quality gates, and outer loop Microsoft Foundry OpenTelemetry tracing in Azure Application Insights
ArchitectureAzure AIFoundry LocalImplementation & OperationsLocal AIMicrosoft Foundry

From Edge to Enterprise Cloud: Part 4 — GenAIOps, Automated Evaluations and Observability with Microsoft Foundry

28.09.2026 20 Min Read
0

In our previous posts, From Edge to Enterprise-Cloud: Part 1: Building a Cost-Free Local Loop with Foundry Local, From Edge to Enterprise Cloud: Part 2: Tool Orchestration with the Model Context Protocol and From Edge to Enterprise Cloud: Part 3: Cloud Migration and Enterprise Identity with Microsoft Foundry Agent Service, we built a hardware-accelerated local loop, measured how schema hardening changes tool calling reliability on Small Language Models and moved the agent into Microsoft Foundry with Bicep and Microsoft Entra ID.

Part 3 ended with a question: how do we keep the agent correct when the model, the system prompt or a tool schema changes? Part 2 already showed why that question is not academic. A bigger model from the same family, qwen2.5-7b, dropped the name field of a nested object in ten out of ten calls, while its smaller sibling got most of them right. Nothing crashed. Nothing in a standard CI pipeline would have noticed.

Executive summary: An AI agent can get worse without a single failing unit test. This final part adds the missing safety net. A golden dataset, built from the Part 2 study, runs on every pull request. A deterministic check that costs nothing blocks the merge when tool calls start failing. An optional LLM judge adds a second opinion. The observability infrastructure is provisioned with Bicep and uses no keys and no stored secrets. Every run can be traced from a developer laptop to Application Insights. For decision makers this means model upgrades become a measured, reviewable change instead of a leap of faith.

All code in this post lives in the part04-observability-evals folder of our companion repository, Frezz146/from-edge-to-enterprise-cloud, in Python and C#.

The Three Pillars: Evaluate, Trace and Monitor

Operating agents means applying DevOps discipline to software that is not deterministic. Microsoft frames observability for Foundry agents in three pillars. They map cleanly onto this part:

PillarQuestion it answersWhat this part builds
EvaluateIs this change good enough to ship?A golden dataset and a quality gate in GitHub Actions
TraceWhat exactly happened in this one run?OpenTelemetry spans to an Aspire Dashboard or Application Insights
MonitorIs quality drifting over time in production?The infrastructure for it, plus an outlook at the end

One principle carries over from the whole series: measured, not assumed. Part 4 does not introduce a new test suite written for the occasion. It reuses the exact cases, the exact MCP server and the exact error taxonomy that produced the Part 2 numbers. The quality gate imports that code instead of copying it, so the study and the gate can never drift apart.

Architecture diagram: a pull request triggers two GitHub Actions quality gates that score a golden dataset. The local gate runs Foundry Local with an MCP server. The cloud gate signs in via OIDC and calls gpt-5-mini in Microsoft Foundry with an LLM judge. Traces flow to Application Insights and Log Analytics.
Every pull request passes a local and a cloud quality gate before it can merge, while OpenTelemetry traces land in Application Insights.

Infrastructure as Code: Extending the Part 3 Bicep

Observability infrastructure deserves the same rigor as the agent itself, so it is Bicep again. The Part 4 template does not recreate anything from Part 3. It references the Foundry account, the project and the managed identity as existing resources and derives their names with the same uniqueString() formula Part 3 used. The only value to keep in sync is namePrefix.

It adds four blocks:

  1. Log Analytics and Application Insights with DisableLocalAuth: true. Ingestion keys stop working. Every span has to arrive with a Microsoft Entra ID token.
  2. The Application Insights connection on the Foundry account and project. This switches on tracing for agents Foundry runs. It is also where our client code looks up the connection string.
  3. Two OIDC federated credentials on the Part 3 identity: one for pull requests and one for main. GitHub Actions signs in as that identity without any stored secret.
  4. Least-privilege role assignments, one per job to be done.
📄 Bicep Infrastructure Template (part04-observability-evals/infra/main.bicep)
bicep
// Part 3 names, same formula as Part 3: only namePrefix has to match
var uniqueSuffix = uniqueString(resourceGroup().id, namePrefix)
var accountName = '${namePrefix}-${uniqueSuffix}'
var identityName = '${namePrefix}-agent-identity'

resource account 'Microsoft.CognitiveServices/accounts@2025-06-01' existing = {
  name: accountName
}

resource appInsights 'Microsoft.Insights/components@2020-02-02' = {
  name: appInsightsName
  location: location
  kind: 'web'
  properties: {
    Application_Type: 'web'
    WorkspaceResourceId: workspace.id
    DisableLocalAuth: disableLocalAuth // default true: Entra ID tokens only
  }
}

// Same shape as Microsoft's foundry-samples template
resource projectConnection 'Microsoft.CognitiveServices/accounts/projects/connections@2025-06-01' = {
  parent: project
  name: appInsightsName
  properties: {
    category: 'AppInsights'
    target: appInsights.id
    authType: 'ApiKey'
    isSharedToAll: true
    credentials: {
      key: appInsights.properties.ConnectionString
    }
    metadata: {
      ApiType: 'Azure'
      ResourceId: appInsights.id
    }
  }
}

// GitHub's immutable OIDC subject: owner and repository with their numeric IDs
var githubSubjectRepo = '${githubOwner}@${githubOwnerId}/${githubName}@${githubRepositoryId}'

resource githubPullRequest 'Microsoft.ManagedIdentity/userAssignedIdentities/federatedIdentityCredentials@2023-01-31' = {
  parent: agentIdentity
  name: 'github-pull-request'
  properties: {
    issuer: 'https://token.actions.githubusercontent.com'
    subject: 'repo:${githubSubjectRepo}:pull_request'
    audiences: [ 'api://AzureADTokenExchange' ]
  }
}

“ApiKey” on a Keyless Resource?

This one surprised me while building it. It is worth a paragraph because it looks like a contradiction. My first draft declared the connection with Entra ID authentication. A look into the SDK showed why that cannot work: when a client asks the project for its connection string, azure-ai-projects only accepts an ApiKey style connection and raises otherwise. Microsoft’s own foundry-samples template uses exactly the shape shown above.

The resolution is DisableLocalAuth. With local authentication disabled, the connection string is only an address. It names the resource, but Application Insights rejects any ingestion that does not carry an Entra ID token. So the connection tells the client where to send data. The identity plus its role decide whether it may.

CAUTION: For agents that Foundry itself runs (published prompt agents, hosted agents), server-side trace ingestion into an Entra-only Application Insights resource requires the connection’s auth type Project managed identity. That option is currently in public preview. The template already grants the project identity the role it needs. Switch the auth type in the portal if you use that path. Alternatively, deploy with disableLocalAuth=false.

Who Gets Which Role

Instead of one broad role per identity, every assignment answers one question:

IdentityRoleScopeWhy
Agent identity (also CI via OIDC)Monitoring Metrics PublisherApplication InsightsWrite spans with an Entra ID token
Agent identityCognitive Services OpenAI UserFoundry accountCloud gate and LLM judge
Project managed identityMonitoring Metrics Publisher, Log Analytics Reader, Privileged Monitoring Data ReaderApplication InsightsFoundry’s own traces and evaluations
You (optional parameter)Monitoring Metrics Publisher, Log Analytics Reader, Cognitive Services OpenAI UserApplication Insights, accountRun everything locally and browse traces

Note what is missing: the CI identity can write telemetry but not read it. A pipeline that only writes has no reason to see production traces, which may contain customer content.

az deployment group create \
  --resource-group <your-resource-group> \
  --template-file infra/main.bicep \
  --parameters infra/main.bicepparam \
  --parameters developerPrincipalId=$(az ad signed-in-user show --query id -o tsv)

The outputs are the client ID, tenant ID, subscription ID and the Azure OpenAI endpoint. None of them is a secret. They go into GitHub as repository variables, which is the whole point.

The Golden Dataset: Reusing the Schema Hardening Scenarios

A golden dataset is only as useful as the failures it is known to catch. Ours has a head start: the ten booking scenarios from Part 2 already caught a real regression. They move from a Python list into golden_dataset.json, next to the thresholds the gate enforces. Both language tracks read the same file.

{
  "tool": "book_appointment_strict",
  "quality_gate": {
    "min_clean_call_rate": 0.8,
    "min_tool_call_accuracy": 4.0,
    "min_judge_pass_rate": 0.8
  },
  "cases": [
    {
      "id": "case_08_standard_priority",
      "query": "Schedule a standard-priority 30-minute sync with Riley Nguyen (riley.nguyen@example.com) covering the Q4 planning doc.",
      "expected_tool": "book_appointment_strict",
      "expected_parameters": {
        "attendee": { "name": "Riley Nguyen", "email": "riley.nguyen@example.com" },
        "topics": ["Q4 planning doc"],
        "priority": "medium",
        "duration_minutes": 30
      }
    }
  ]
}

Two design decisions matter here. The gate targets only the strict schema, because that is the one you ship after reading Part 2. And the thresholds live with the data, not in code: tightening the gate is a reviewed one-line change to a JSON file.

A practical tip on thresholds: do not guess them. Run the gate against your current model and anchor the threshold to that measured baseline, as shown in the baseline section further down. A gate that fails on the current main teaches the team to ignore it.

Evaluators: A Deterministic Taxonomy First, an LLM Judge Second

For every case the pipeline does four things:

  1. It sends the query and the book_appointment_strict schema to the model under test, either a local Foundry Local model or the gpt-5-mini deployment from Part 3.
  2. It executes the returned call against Part 2’s own MCP server. SERVER_REJECTED therefore means exactly what it meant in Part 2: Pydantic refused the arguments before the handler ran.
  3. It classifies the outcome with Part 2’s error taxonomy (OK, NO_TOOL_CALL, WRONG_TOOL, MALFORMED_JSON, SERVER_REJECTED, WRONG_NAME, WRONG_EMAIL, WRONG_PRIORITY, WRONG_DURATION, MISSING_TOPIC).
  4. With --judge, it asks ToolCallAccuracyEvaluator from azure-ai-evaluation for a score between 1 and 5.

Why this order? The taxonomy is deterministic, runs in milliseconds and costs nothing. The same model answer always gets the same verdict, which is the property you want from something that blocks a merge. The judge is a model: it catches problems no rule anticipated, but it costs tokens and its scores vary slightly from run to run. So the taxonomy is the gate and the judge is the second opinion, applied to the average over the dataset rather than to single cases.

🐍 Python Evaluation Pipeline (part04-observability-evals/python/eval_pipeline.py)
# Imported, not copied: this is the code that produced the Part 2 numbers.
from schema_hardening_eval import (
    SYSTEM_PROMPT,
    Case,
    classify_semantic,
    load_local_model,
    mcp_tool_to_foundry_tool,
    parse_fallback_tool_call,
)


async def run_case(complete_chat, session, tool, case) -> CaseResult:
    tool_name = tool["function"]["name"]
    messages = [
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": case["query"]},
    ]
    message = complete_chat(messages, tools=[tool]).choices[0].message

    if message.tool_calls:
        call = message.tool_calls[0]
        call_name = call.function.name
        try:
            args = json.loads(call.function.arguments)
        except json.JSONDecodeError:
            return CaseResult(case["id"], ["MALFORMED_JSON"])
    else:
        # Part 2's fallback for models that write tool calls as tagged text
        properties = tool["function"]["parameters"].get("properties", {})
        param_types = {key: prop.get("type") for key, prop in properties.items()}
        fallback = parse_fallback_tool_call(message.content or "", param_types)
        if fallback is None:
            return CaseResult(case["id"], ["NO_TOOL_CALL"])
        call_name, args = fallback

    tool_call = {"name": call_name, "arguments": args}
    if call_name != tool_name:
        return CaseResult(case["id"], ["WRONG_TOOL"], tool_call)

    result = await session.call_tool(tool_name, arguments=args)  # Part 2's MCP server
    if result.isError:
        return CaseResult(case["id"], ["SERVER_REJECTED"], tool_call, ...)

    errors = classify_semantic(
        tool_name, args, to_part2_case(case)
    )  # Part 2's taxonomy
    return CaseResult(case["id"], errors or ["OK"], tool_call)

The judge is authenticated like everything else in this series, with DefaultAzureCredential instead of an API key:

def build_judge(threshold: float):
    deployment = os.environ.get("JUDGE_MODEL", "gpt-5-mini")
    model_config = {
        "azure_endpoint": require_env("AZURE_OPENAI_ENDPOINT"),
        "azure_deployment": deployment,
        "api_version": "2025-04-01-preview",  # the evaluator default predates reasoning models
    }
    return ToolCallAccuracyEvaluator(
        model_config=model_config,
        credential=DefaultAzureCredential(),  # not inside model_config, see below
        threshold=threshold,
        # gpt-5 family and o-series reject the sampling settings in the judge prompt
        is_reasoning_model=deployment.startswith(("gpt-5", "o1", "o3", "o4")),
    )

The credential placement is not a style choice. The SDK’s AzureOpenAIModelConfiguration type lists a credential key, but in azure-ai-evaluation 1.18 a config that contains it fails the SDK’s own validation with “Model config validation failed”. My first verification run stopped exactly there. Passed as the evaluator’s own credential argument, it works.

Two more details are easy to miss. A case without a parseable tool call gets the rubric’s floor score of 1, so a model that simply stops calling tools cannot raise the average. And ideally the judge is not the model under test. The demo uses one deployment for both to keep costs down. In a real setup, deploy a second model for the judge.

C# Parity with Microsoft.Extensions.AI.Evaluation

The C# track (csharp/eval-runner) mirrors the Python gate: same golden dataset, same taxonomy, Part 2’s C# MCP server spawned over stdio. The judge comes from Microsoft.Extensions.AI.Evaluation.Quality, which ships a ToolCallAccuracyEvaluator of its own.

One difference matters when you compare numbers across languages: the .NET evaluator returns a BooleanMetric, accurate or not, while the Python evaluator returns a score from 1 to 5. A single threshold cannot serve both, which is why the dataset carries min_tool_call_accuracy for Python and min_judge_pass_rate for .NET. After Part 2 found that “strict” was not equally strict in both languages, this is the same lesson at the evaluation layer: parity has to be verified, not assumed.

The verification run then reproduced the other open question from Part 2. The C# runner uses the hand-authored strict schema from Part 2’s C# track. Against it, qwen2.5-1.5b scored 0/10: it dropped attendee.name in every single call and durationMinutes in three of them. The same weights score 8/10 against the Python server’s Pydantic schema. Part 2 called the cause an open hypothesis about schema presentation. Part 4 now shows the same effect with a completely different harness. The C# runner therefore defaults to qwen3.5-4b, which scores 9/10 in C#, matching its Part 2 result. A gate is only useful with a model that passes it on main.

CAUTION, experimental APIs (AIEVAL001, OPENAI001): the quality evaluators in Microsoft.Extensions.AI.Evaluation.Quality are marked [Experimental("AIEVAL001")]. The first real build also failed on OPENAI001: the OpenAI ChatClient constructor that accepts an AuthenticationPolicy, which is how the judge signs in with Entra ID instead of a key, is experimental as well. The project acknowledges both once with <NoWarn>$(NoWarn);AIEVAL001;OPENAI001</NoWarn> in the .csproj instead of scattering pragmas through the code. Treat them as preview components in a production pipeline.

var openAiChatClient = new global::OpenAI.Chat.ChatClient(
    judgeDeployment,
    new BearerTokenPolicy(new DefaultAzureCredential(), "https://ai.azure.com/.default"),
    new global::OpenAI.OpenAIClientOptions { Endpoint = new Uri($"{endpoint.TrimEnd('/')}/openai/v1/") });

judgeConfiguration = new ChatConfiguration(
    new ReasoningModelCompatibleChatClient(AI.OpenAIClientExtensions.AsIChatClient(openAiChatClient)));
judge = new ToolCallAccuracyEvaluator();
judgeContext = new ToolCallAccuracyEvaluatorContext(
    AI.AIFunctionFactory.CreateDeclaration(
        strictTool.Function!.Name!,
        strictTool.Function.Description,
        JsonSerializer.SerializeToElement(strictTool.Function.Parameters)));

// Per case: wrap the local model's call as a FunctionCallContent and let the judge decide
var modelResponse = new AI.ChatResponse(new AI.ChatMessage(
    AI.ChatRole.Assistant,
    [new AI.FunctionCallContent($"call_{goldenCase.Id}", result.CallName, result.Arguments)]));

var evaluation = await judge.EvaluateAsync(conversation, modelResponse, judgeConfiguration, [judgeContext], ct);
bool? accurate = evaluation.Get<booleanmetric>(ToolCallAccuracyEvaluator.ToolCallAccuracyMetricName).Value;

The .NET evaluator sets Temperature, TopP and a tight output budget on its judge requests. Reasoning models such as gpt-5-mini reject non-default sampling settings. Reasoning tokens can also use up a small budget before the verdict is written. A ten-line DelegatingChatClient strips those options for the judge, the .NET counterpart to Python’s is_reasoning_model flag.

Quality Gates in GitHub Actions

The workflow runs two jobs on every pull request that touches Part 2’s Python track, Part 4 or the workflow itself:

JobModel under testNeeds Azure?Purpose
local-gateFoundry Local on the runnerNoDeterministic taxonomy, free, also runs for forks
cloud-gategpt-5-mini from Part 3Yes, via OIDCTaxonomy plus LLM judge, spans to Application Insights
env:
  # The model under test. Changing this line in a pull request is exactly
  # the kind of change the gate exists to check.
  FOUNDRY_LOCAL_MODEL: ${{ inputs.local_model || 'qwen2.5-1.5b' }}

jobs:
  local-gate:
    runs-on: ubuntu-latest
    steps:
      # (checkout, setup-python and pip install omitted)
      - name: Local quality gate (${{ env.FOUNDRY_LOCAL_MODEL }})
        run: python eval_pipeline.py --backend local --model "$FOUNDRY_LOCAL_MODEL"

  cloud-gate:
    if: ${{ vars.AZURE_CLIENT_ID != '' }}
    runs-on: ubuntu-latest
    permissions:
      contents: read
      id-token: write # lets azure/login exchange the GitHub OIDC token, no client secret
    steps:
      # (checkout, setup-python and pip install omitted)
      - uses: azure/login@v2
        with:
          client-id: ${{ vars.AZURE_CLIENT_ID }}
          tenant-id: ${{ vars.AZURE_TENANT_ID }}
          subscription-id: ${{ vars.AZURE_SUBSCRIPTION_ID }}
      - name: Cloud quality gate (gpt-5-mini, LLM judge, traced to Application Insights)
        run: python eval_pipeline.py --backend cloud --judge --trace azure

Everything the cloud job needs is a vars.* value, not a secrets.* value. Client ID, tenant ID and subscription ID are identifiers. The federated credential from the Bicep template is what turns them into a token. There is nothing in the repository settings an attacker could reuse outside this repository and these two triggers.

The first real cloud-gate run still failed at azure/login, with AADSTS700213: No matching federated identity record found for presented assertion subject. The error message quoted the subject GitHub had sent: repo:Frezz146@115118426/from-edge-to-enterprise-cloud@1360166622:ref:refs/heads/main. That is GitHub’s immutable subject format, which includes the numeric owner and repository IDs so that a renamed or recreated repository cannot inherit the trust. The federated credential used the familiar repo:<owner>/<name>:... format and therefore never matched. The Bicep template now takes both IDs as parameters (gh api repos/<owner>/<name> --jq ".owner.id, .id" prints them) and falls back to the legacy format when they are empty. It is a one-line mistake that no local test can catch, which is exactly why a pipeline has to run in the pipeline at least once. After redeploying the credentials both jobs passed on GitHub: local-gate with 8/10 and cloud-gate with 10/10 and a judge score of 5.00.

The local-gate job needs no Azure access at all, so it also protects pull requests from forks, which GitHub does not give an OIDC token. Foundry Local falls back to CPU execution providers on a hosted runner. That is slower than a developer machine. A self-hosted runner is faster and keeps the model cache between runs. Once both jobs are green on main, make them required status checks in a branch protection rule.

The Baseline: What qwen2.5-1.5b Scores Today

Before a gate can block anything, it needs a baseline. This is the local gate against qwen2.5-1.5b on my developer machine, server messages shortened:

=== local (qwen2.5-1.5b) | 10 golden cases | tool book_appointment_strict ===
  [case_01_explicit] OK
  [case_02_two_topics] OK
  [case_03_informal_low_key] OK
  [case_04_one_hour] OK
  [case_05_three_topics] OK
  [case_06_urgent] OK
  [case_07_no_rush] SERVER_REJECTED
    server said: ... priority Input should be 'low', 'medium' or 'high' [type=literal_error, input_value='normal', input_type=str] ...
  [case_08_standard_priority] SERVER_REJECTED
    server said: ... priority Input should be 'low', 'medium' or 'high' [type=literal_error, input_value='standard', input_type=str] ...
  [case_09_two_low_topics] OK
  [case_10_critical] OK

=== Quality gate summary ===
Backend                        | Clean calls | Judge   | Errors
------------------------------------------------------------------------------------------
local (qwen2.5-1.5b)           | 8/10        | -       | SERVER_REJECTED x2
  clean call rate 80% (min 80%)

Quality gate PASSED.

Both misses are the same kind of mistake. The model invented priority values the enum does not allow: “no-rush” became normal and “standard-priority” became standard. Everything else was right in all ten cases, including the nested attendee object, the topic arrays and the durations. That matches the Part 2 Python result for this model (8/10, SERVER_REJECTED x2) and Part 2’s fourth discovery: informal priority wording is the hardest field for small models. With the loose schema both values would have been booked silently as WRONG_PRIORITY. The strict enum turns them into loud rejections that a gate can count.

The LLM judge agrees. Run with --judge, it scored every clean call 5.0 and both rejected calls 2.0, for a mean of 4.40 out of 5. The cloud deployment from Part 3 sets the upper bound: gpt-5-mini passes all ten cases with a judge score of 5.00, mapping “no-rush” to low and “standard-priority” to medium as intended.

Model under testTrackClean callsJudge
qwen2.5-1.5b (local)Python8/104.40/5
gpt-5-mini (cloud)Python10/105.00/5
qwen3.5-4b (local)C#9/108 and 9 of 10 accurate in two runs
qwen2.5-1.5b (local)C#0/100 of 10 accurate

Because the local run uses temperature 0, the verdict is repeatable. The gate therefore sits exactly at the baseline: min_clean_call_rate: 0.8 means “no worse than today”. One open question was the CI runner: it uses CPU execution providers instead of the GPU on a developer machine. A different execution provider can change greedy decoding in edge cases. The first CI run answered it. On a hosted Ubuntu runner the local gate produced the same 8/10 with the same two rejected cases, so one threshold serves both.

Live Demonstration: Catching the Part 2 Regression in CI

Now picture a reasonable pull request. A developer bumps FOUNDRY_LOCAL_MODEL from qwen2.5-1.5b to qwen2.5-7b, expecting a bigger model in the same family to be at least as good. Part 2 already told us what that model does with the strict schema: it omits attendee.name in every single call. This time nobody has to notice. Here is the real output of the exact command the local-gate job runs, executed on my machine (MCP server log lines removed, validation messages shortened):

[Local] Loading Foundry Local model 'qwen2.5-7b'...
  capabilities='tool-calling' supports_tool_calling=True

=== local (qwen2.5-7b) | 10 golden cases | tool book_appointment_strict ===
  [case_01_explicit] SERVER_REJECTED
    server said: ... 1 validation error: attendee.name Field required [type=missing, input_value={'email': 'jamie.chen@example.com'}, input_type=dict]
  [case_02_two_topics] SERVER_REJECTED
    server said: ... 1 validation error: attendee.name Field required [type=missing, input_value={'email': 'alex.kim@example.com'}, input_type=dict]
  [case_03_informal_low_key] SERVER_REJECTED
    server said: ... 1 validation error: attendee.name Field required [type=missing, input_value={'email': 'sam.rivera@example.com'}, input_type=dict]
  [case_04_one_hour] SERVER_REJECTED
    server said: ... 1 validation error: attendee.name Field required [type=missing, input_value={'email': 'morgan.lee@example.com'}, input_type=dict]
  [case_05_three_topics] SERVER_REJECTED
    server said: ... 1 validation error: attendee.name Field required [type=missing, input_value={'email': 'taylor.brooks@example.com'}, input_type=dict]
  [case_06_urgent] SERVER_REJECTED
    server said: ... 1 validation error: attendee.name Field required [type=missing, input_value={'email': 'priya.patel@example.com'}, input_type=dict]
  [case_07_no_rush] SERVER_REJECTED
    server said: ... 1 validation error: attendee.name Field required [type=missing, input_value={'email': 'devon.clarke@example.com'}, input_type=dict]
  [case_08_standard_priority] SERVER_REJECTED
    server said: ... 2 validation errors: attendee.name Field required [type=missing, input_value={'email': 'riley.nguyen@example.com'}, input_type=dict] priority Input should be 'low', 'medium' or 'high' [type=literal_error, input_value='standard', input_type=str]
  [case_09_two_low_topics] SERVER_REJECTED
    server said: ... 1 validation error: attendee.name Field required [type=missing, input_value={'email': 'casey.morgan@example.com'}, input_type=dict]
  [case_10_critical] SERVER_REJECTED
    server said: ... 1 validation error: attendee.name Field required [type=missing, input_value={'email': 'jordan.ellis@example.com'}, input_type=dict]

=== Quality gate summary ===
Backend                        | Clean calls | Judge   | Errors
------------------------------------------------------------------------------------------
local (qwen2.5-7b)             | 0/10        | -       | SERVER_REJECTED x10
  clean call rate 0% (min 80%)

Quality gate FAILED.
Results written to evaluation_results.json

Ten out of ten calls drop attendee.name, exactly the pattern Part 2 documented. Case 8 even shows both failure modes at once: the missing name of the 7B model and the invented standard priority that already tripped up the 1.5B model. In CI the check turns red. The run summary page shows a table with every case and its outcome. evaluation_results.json is attached as an artifact with the rejected arguments and the server’s validation message. The merge is blocked before any user books a meeting with a blank name.

This is the payoff of the whole series in one screen. Part 2 found the failure with a careful study. Part 4 makes finding it the default.

Tracing Across the Hybrid Stack

Evaluations tell you whether something is wrong. Traces tell you what happened. Microsoft Foundry distinguishes two levels. Our agent needs both:

  • Server-side tracing covers agents that Foundry runs, such as published prompt agents and hosted agents. Once the Application Insights connection from the Bicep template exists, Foundry traces them without any code change.
  • Client-side tracing covers code that runs in your own process. The Part 3 agent’s agent.run() loop is exactly that: Agent Framework calls the model in Foundry, but the loop, the function tools and the session live in our Python process.

Agent Framework already instruments agent runs, model calls and tool invocations with the OpenTelemetry GenAI semantic conventions. The only decision left is where the spans go. That is where the hybrid theme of the series returns.

Inner Loop: an Aspire Dashboard on Your Machine

During development you want traces the way Part 1 wanted inference: local, instant and free. The standalone .NET Aspire Dashboard accepts OpenTelemetry over OTLP and runs in a single container:

docker run --rm -it -d -p 18888:18888 -p 4317:18889 -p 4318:18890 \
  -e ASPIRE_DASHBOARD_UNSECURED_ALLOW_ANONYMOUS=true \
  --name aspire-dashboard mcr.microsoft.com/dotnet/aspire-dashboard:latest

python traced_agent.py --target otlp     # the Part 3 agent, fully traced
python eval_pipeline.py --trace otlp     # the quality gate, one span per case

Open http://localhost:18888 and you see the whole conversation as one trace: the parent span, both turns, every model call and every get_weather or calculate invocation with its duration. The first docker run pulls the dashboard image once. After that the inner loop starts in seconds.

Outer Loop: Application Insights, without a Connection String in Code

In the cloud the same spans go to the Application Insights resource from the Bicep template. Notice what is not in this code: a connection string, a key or an environment variable pointing at Application Insights.

from cloud_agent import AGENT_NAME, calculate, get_weather  # Part 3, unchanged

credential = DefaultAzureCredential()
client = FoundryChatClient(credential=credential)

# Reads the connection string from the project's Application Insights connection
# (infra/main.bicep) and exports with an Entra ID token, since local auth is disabled.
await client.configure_azure_monitor(enable_sensitive_data=False, credential=credential)

agent = Agent(
    client=client, name=AGENT_NAME, instructions=..., tools=[get_weather, calculate]
)

with get_tracer().start_as_current_span("part04.traced_conversation") as span:
    session = agent.create_session()
    for question in QUESTIONS:
        result = await agent.run(question, session=session)
    span.set_attribute("gen_ai.conversation.id", session.service_session_id or "")

configure_azure_monitor() asks the Foundry project for the Application Insights connection the Bicep template created. It then exports every span with the Entra ID token from DefaultAzureCredential. The script prints the trace ID at the end, so you can jump straight to it in Transaction search. In the verification run the agent answered both turns (Tokyo at 18 °C, 42 × 17 = 714, then 64.4 °F) and its spans arrived in an Application Insights resource that accepts no ingestion key at all. The quality gate does the same with --trace azure, which is how the cloud-gate job in CI shows up next to the agent’s own traces.

Data Privacy and Content Recording

Spans can carry prompts, responses and tool arguments. In enterprise environments, particularly in the DACH region and across the EU, storing that content in telemetry tables is a decision that needs an owner, not a default somebody forgot to change. The switch depends on the library that emits the spans:

StackSwitchDefault
Agent Framework (traced_agent.py)enable_sensitive_data or ENABLE_SENSITIVE_DATA=trueoff
Azure AI SDK instrumentors (azure-ai-projects, azure-ai-inference)AZURE_TRACING_GEN_AI_CONTENT_RECORDING_ENABLED=trueoff
Our quality gate (eval_pipeline.py)none, it never records contentoff

Watch the exact variable name for the Azure SDKs: it ends in _ENABLED. Without that suffix the SDK ignores it and records nothing, which is the safe failure but rarely the intended one. The gate itself only writes case IDs, outcomes and scores into its spans. Keep content recording for dev and staging projects where it has been agreed on. Also remember who can read it: Log Analytics Reader on the Application Insights resource or Privileged Monitoring Data Reader when the underlying tables are protected.

From Pipeline to Production: Continuous Evaluation and Alerts

GenAIOps does not stop at the merge button. Users phrase things differently than a golden dataset does. Quality can drift without any code change. The infrastructure from this part is the foundation for closing that loop: sample live conversations, score them on a schedule with the same evaluators and raise an Azure Monitor alert when the rolling average falls below a threshold. Foundry offers continuous evaluation for the agents it runs.

I deliberately left this out of the repository. A demo agent has no real traffic. An alert on synthetic traffic would teach the wrong lesson. When your agent has users, this is the next step. From then on the golden dataset grows from the production traces that surprised you.

What Building Part 4 Taught Me

Like Part 2, this part produced findings that only showed up while building and running it end to end:

  1. The judge is a second opinion, not a gate: Two C# runs with the same model at temperature 0 produced 8 and then 9 accurate verdicts, one case simply flipped. The deterministic taxonomy never flips. Where both looked at the same calls, they agreed on every rejected one.
  2. Schema presentation is real: qwen2.5-1.5b scores 8/10 against the Python schema and 0/10 against the hand-written C# schema, dropping the same nested field every time. Part 2 raised the hypothesis. Part 4 reproduced it with a different harness.
  3. Read the SDK, then run it: azure-ai-evaluation 1.18 rejects its own model config when it contains a credential. Its evaluators also default to an API version older than any reasoning model. Neither shows up until the first real run.
  4. A keyless resource can still need an “ApiKey” connection: The SDK lookup for the connection string only accepts that shape. DisableLocalAuth is what keeps it safe, because the string alone cannot ingest anything.
  5. “The same evaluator” is not the same metric: Python scores 1 to 5, .NET returns true or false. Cross-language dashboards need a mapping, not a shared threshold.
  6. OIDC subjects are exact strings: GitHub sent this repository’s subject in the immutable format with numeric IDs. Only the first real CI run revealed it. The error message itself contained the fix.
  7. A judge needs capacity: With the gate, the cloud backend and the judge sharing one 10,000 TPM deployment, back to back runs hit 429 rate limits. The SDK retried and finished, but a real pipeline deserves a separate judge deployment.
  8. Reuse beats rewriting: Importing the Part 2 MCP server and taxonomy into the gate took a few lines. It guarantees that the regression the study found is the regression the gate catches.

The Grand Finale: The Complete 4-Part Hybrid Agent Architecture

Four parts ago we started with a 0.5B model on a laptop and a question about latency and cost. We end with an agent that runs in Microsoft Foundry, authenticates without secrets and cannot be made worse by a pull request without somebody noticing.

Architecture layerPart 1: Local LoopPart 2: MCP OrchestrationPart 3: Cloud MigrationPart 4: GenAIOps & Observability
Execution domainOn-device (workstation, NPU/GPU)Local stdio subprocessesCloud-hosted (Microsoft Foundry)Hybrid (local and cloud)
Core capabilityIn-process streaming inferenceStandardized tool transportScalable agent lifecycleQuality gates and distributed tracing
Security & identityLocal user permissionsLocal environment variablesMicrosoft Entra ID, managed identityOIDC federation, Entra-only telemetry
Primary technologyFoundry Local GA SDKModel Context ProtocolBicep, Agent Framework, Foundry Toolboxazure-ai-evaluation, M.E.AI.Evaluation, OpenTelemetry, GitHub Actions
Problem solvedCloud cost and WAN latency in the dev loopSilent tool-call errors through schema hardeningLocal process ceilings and secret sprawlSilent regressions after model, prompt or schema changes

Key Architectural Lessons Learned

  1. Local-first accelerates everything: Foundry Local made thousands of iterations free. The same local model now runs the cheapest quality gate in the pipeline.
  2. Schema hardening trades silent errors for loud ones: Strict schemas turn wrong data into rejected calls. Rejected calls are something a gate can count.
  3. Identity over secrets: From Part 3 on, nothing authenticates with a key: DefaultAzureCredential for the agent, OIDC for the pipeline, Entra-only ingestion for telemetry. The repository stores no key and no client secret.
  4. Evaluations are your compiler: A model upgrade is a code change. It deserves the same automated check before it reaches main.

The complete, runnable source code for all four parts is available in our official companion repository: Frezz146/from-edge-to-enterprise-cloud.

Thank you for following along through all four parts. Happy enterprise coding!

Tags:

AzureAzure OpenAIFoundry LocalLocal AIMicrosoft Foundry
Author

Alexander Dierkes

Follow Me
Other Articles
Conceptual dark mode 16:9 illustration representing cloud migration from edge workstations to Microsoft Foundry Agent Service with Microsoft Entra ID Zero-Trust identity shields and managed toolboxes
Previous

From Edge to Enterprise Cloud: Part 3 — Cloud Migration and Enterprise Identity with Microsoft Foundry Agent Service

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest Posts

  • From Edge to Enterprise Cloud: Part 4 — GenAIOps, Automated Evaluations and Observability with Microsoft Foundry
  • From Edge to Enterprise Cloud: Part 3 — Cloud Migration and Enterprise Identity with Microsoft Foundry Agent Service
  • From Edge to Enterprise Cloud: Part 2 — Tool Orchestration with the Model Context Protocol
  • From Edge to Enterprise-Cloud: Part 1 — Building a Cost-Free Local Loop with Foundry Local
  • Hands-On with GPT-Realtime-1.5 on Azure

Azure Azure OpenAI Foundry Local Local AI Microsoft Foundry

Resources

About | Imprint

Disclaimer

Opinions expressed here are my own and may not reflect those of others. Unless I'm quoting someone, they're just my own views.

© 2026 Frezz Tech - All rights reserved by Alexander Dierkes
Independent technology blog. Not affiliated with Microsoft.