Skip to content

Build1 publisher3 min readPublished

Low scores for a Snowflake Cortex Agent trace partly to its own evaluation

A developer testing a Snowflake Cortex Agent found the app's tool-call counter recorded missing telemetry as zero tool use. The retests show a low agent score can come from gaps in the evaluation, so the test and its telemetry need checking before anyone rewrites the prompt.

The Engineer · Build desk

Illustration accompanying Low scores for a Snowflake Cortex Agent trace partly to its own evaluation

What happened

  • Some expected tool lists in the agent's tests left out prerequisites its own instructions required, and others allowed one route where several were valid.
  • When the agent wrongly reported an object as missing, a fallback instruction added to its prompt did not fix the lookup.
  • The retest recovered the object and its consumers only after guidance changed on both the agent and the semantic view Cortex Analyst uses.
  • The fixed lookup still earned partial tool-selection credit because Snowflake's check counted its extra calls against it.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint A Cortex tool-selection score cannot be reported as answer accuracy, since a lower number can mean the agent took extra steps on the way to a better answer.
  • exposure Dashboards and cost reports that count tool calls on the application side with a zero default will undercount agent activity whenever the metadata drops out.
  • decision A failing eval forces a choice among three maintained layers before any prompt edit: the agent, the evidence about its behavior, or the test judging it.
  • cost Turning production conversations into regression cases takes curation, because each case needs its prior context and a once-successful or time-sensitive answer must be rechecked first.

The counter defect is the easiest to explain. The application's tool-call counter took its figure from metadata, and when the metadata was missing it defaulted to zero. Snowflake's native traces showed tool activity the counter had missed [1]. "A missing measurement had been made to look like a measured zero," the author wrote [2]. Zero is a reasonable default for a counter until someone scores an agent with it. I'd store the field as null and have the eval decline to score any run where it is null.

Between the prompt fix that failed and the one that worked sat a wrong theory. The object the agent called nonexistent was in the metadata, recorded as a source that other views consume, not as the view the agent was searching for [5]. The next explanation was that the semantic tool could not search by source. Reading the tool's definition disproved that. The source dimension already existed, and the SQL-generation guidance steered toward searches by view name [7]. Missing data, a tool limit and a badly phrased request each lead to a different fix, the author wrote, and inspecting the definition and retesting proved more useful than accepting the first diagnosis [18].

The successful retest was scoped narrowly too. The author concluded that retrieval improved in that retest, not that every returned statement was correct or that the whole agent got better [19]. Finding downstream consumers did not establish how each upstream object was loaded, and a later correction in the conversation made that distinction explicit [17].

The case supports checking the evaluation before the prompt, and the author put a limit on doing so. "I should not make a test easier just because the agent failed it," the author wrote [11]. The rule that follows is specific: "Each change to the expected behavior needs an independent reason: a verified alternative route, a documented prerequisite, or a correction to the case itself." [12] Adding a prerequisite the instructions already require meets that bar; widening a tool list to match whatever the agent happened to call does not.

The evidence is one anonymized Cortex Agent observed through retests, and the author says it is not a controlled benchmark [4]. It does not show how often the evaluation, and not the agent, is at fault. The pattern should carry over to a team with hand-written expected tool lists [3], a tool-selection metric that charges for extra calls [9], and an application-side tool counter kept apart from the platform's traces [1].

Each low score now gets three questions from the author: whether the answer was correct, whether the path was appropriate, and whether the evaluation measured the intended behavior [14]. The author read real questions, responses and follow-ups next to native execution traces. A scheduled Cortex Code review sorted problems into instruction gaps, data gaps and tool limitations, and its findings were treated as leads to investigate [15]. Trace summaries were materialized to keep a longer investigation trail. The author notes this addressed the history available in that environment, not a universal retention limit [15].

What to watch

  • Whether Snowflake's evaluation tooling lets a test case accept more than one valid tool route; the author hit cases that allowed only one.
  • Whether the author publishes per-case retest counts, the first figure that would show how often the test, and not the agent, was wrong.
  • Whether teams move tool-call metrics off application counters and onto Cortex Agent native traces, closing the gap the zero default exposed.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories