Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished

NVIDIA took second at KDD Cup 2026 by shrinking its agent's harness to four functions

NVIDIA's KGMON team took second in the KDD Cup 2026 Data Agents competition with an agent limited to four functions. The rules fixed a small model, so all the tuning went into the harness, the same position as any team committed to a small open model.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying NVIDIA took second at KDD Cup 2026 by shrinking its agent's harness to four functions
Generated illustration

What happened

  • The competition asked agents to answer natural-language questions over databases, CSV and JSON files, prose documents, PDFs and briefing videos.
  • Middleware repaired malformed tool calls so that a single bad call did not end an attempt.
  • Final answers went through one function, write_answer(df), which writes the answer file atomically.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint The main simplification depends on staging all structured data in one queryable store before the agent runs, so teams whose data cannot be staged that way lose it.
  • decision Teams improving a small-model agent can try a read-only preflight briefing before adding another retrieval tool, the order NVIDIA's team argues for.
  • cost Reliability gains here are paid for in harness engineering time by the team building the agent, and they come back as spare turns for retries and evaluation.

The model was fixed by the rules. KDD required every team to run the same small LLM, so the harness was the main thing KGMON could optimize [4]. The tasks also asked for more than retrieval: the agent had to inspect the data, pick tools, reason across sources, write an answer file and avoid the traps that show up in analytical work [3].

All structured data sits behind one interface. KGMON converted every CSV and JSON file into a table in the SQLite database the task already had [5]. The agent reaches it through two functions, schema() and sql(query). The team built both into its own Python environment; neither comes from Python or SQLite [6]. Before the reasoning loop starts, a scouting pass reads the tables for possible join keys, duplicate names, look-alike fields, units, null patterns and row-grain problems [8]. The agent gets that briefing at the start of each task and skips a discovery turn [9]. According to NVIDIA, the single interface cut routing failures and wasted turns [7]. The scouting cut errors from wrong columns, missed joins and confusion about what an answer row should represent [17].

The rest of the action space is small. Besides the two query functions, the agent has prose_helper(), which answers from prose or extracts it into SQL, and write_answer(df), which writes the final answer atomically [10]. Every answer goes down that one output path [13]. An atomic write means the scorer sees a complete file or no file. It is the least glamorous function in the list and the one I would copy first. Middleware repairs malformed tool calls, so one bad call does not end an attempt [11]. The Python environment keeps variables between calls, so the agent can reuse intermediate results [12].

The budget all of this protects is turns. NVIDIA says short, valid attempts left more turns for additional runs and evaluation [16]. The team found its failures through execution traces, repeated attempts and trajectory inspection [15]. A small model that spends one turn finding the schema, one on a syntax error and one on a broken file write has that many fewer turns for analysis.

Treat the result as a claim about one workload. NVIDIA wrote that the playbook "isn't a general recipe for every data science agent" [14]. The harness sections of the write-up do not name the model, report a score, or measure each technique on its own, so the share of the second-place finish owed to any single change is unknown. None of the four listed functions is described as handling video, though briefing videos were among the competition's sources [20].

For the pattern to transfer, a few things have to be true in your setup. Your structured data has to load into one queryable store before the agent starts. Your model has to be fixed, or expensive to change. Your tasks have to be narrow enough that a handful of functions covers them. NVIDIA's suggested stand-ins for the temporary SQLite database are a virtualized query layer or a governed warehouse interface [18].

I think this is the right tradeoff when the model is not yours to choose. The post's own summary: "Reliability often comes less from making the model more open-ended and more from building the right harness around it." [19]

What to watch

  • Whether NVIDIA or the KDD Cup organizers publish KGMON's score or per-technique ablations, which would show which change moved the result.
  • Whether the first-place team's write-up describes a similarly narrow harness or a different approach under the same fixed model.
  • Whether the competition's fixed model is named, letting teams on that open model test the pattern directly.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence35
Adoption
Insufficient
Hype gap+10
Incentives55
Confidence40
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    The NVIDIA KGMON team placed second in the KDD Cup 2026 Data Agents competition with a system built around making an agent's harness smaller, clearer, and easier to verify.

    ReportedSupportedSource: NVIDIA developer blogView cited source
  2. [2]

    The competition asked agents to answer natural-language questions over heterogeneous data sources, including databases, CSV and JSON files, prose documents, PDFs, and briefing videos.

    ReportedSupportedSource: NVIDIA developer blogView cited source
  3. [3]

    Every task required more than retrieval: inspecting available data, choosing the right tools, reasoning across sources, producing a final answer file, and handling traps that appear in analytical workflows.

    ReportedSupportedSource: NVIDIA developer blogView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. developer.nvidia.com

    1 article · October 8, 2026

    Building Reliable Data Analytics Agents: Lessons from the KDD Cup

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Loading related stories