BuildNot yet confirmed elsewhere1 publisher3 min readPublished
NVIDIA took second at KDD Cup 2026 by shrinking its agent's harness to four functions
NVIDIA's KGMON team took second in the KDD Cup 2026 Data Agents competition with an agent limited to four functions. The rules fixed a small model, so all the tuning went into the harness, the same position as any team committed to a small open model.
The Engineer · Build desk

What happened
- The competition asked agents to answer natural-language questions over databases, CSV and JSON files, prose documents, PDFs and briefing videos.
- Middleware repaired malformed tool calls so that a single bad call did not end an attempt.
- Final answers went through one function, write_answer(df), which writes the answer file atomically.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint The main simplification depends on staging all structured data in one queryable store before the agent runs, so teams whose data cannot be staged that way lose it.
- decision Teams improving a small-model agent can try a read-only preflight briefing before adding another retrieval tool, the order NVIDIA's team argues for.
- cost Reliability gains here are paid for in harness engineering time by the team building the agent, and they come back as spare turns for retries and evaluation.
The model was fixed by the rules. KDD required every team to run the same small LLM, so the harness was the main thing KGMON could optimize [4]. The tasks also asked for more than retrieval: the agent had to inspect the data, pick tools, reason across sources, write an answer file and avoid the traps that show up in analytical work [3].
All structured data sits behind one interface. KGMON converted every CSV and JSON file into a table in the SQLite database the task already had [5]. The agent reaches it through two functions, schema() and sql(query). The team built both into its own Python environment; neither comes from Python or SQLite [6]. Before the reasoning loop starts, a scouting pass reads the tables for possible join keys, duplicate names, look-alike fields, units, null patterns and row-grain problems [8]. The agent gets that briefing at the start of each task and skips a discovery turn [9]. According to NVIDIA, the single interface cut routing failures and wasted turns [7]. The scouting cut errors from wrong columns, missed joins and confusion about what an answer row should represent [17].
The rest of the action space is small. Besides the two query functions, the agent has prose_helper(), which answers from prose or extracts it into SQL, and write_answer(df), which writes the final answer atomically [10]. Every answer goes down that one output path [13]. An atomic write means the scorer sees a complete file or no file. It is the least glamorous function in the list and the one I would copy first. Middleware repairs malformed tool calls, so one bad call does not end an attempt [11]. The Python environment keeps variables between calls, so the agent can reuse intermediate results [12].
The budget all of this protects is turns. NVIDIA says short, valid attempts left more turns for additional runs and evaluation [16]. The team found its failures through execution traces, repeated attempts and trajectory inspection [15]. A small model that spends one turn finding the schema, one on a syntax error and one on a broken file write has that many fewer turns for analysis.
Treat the result as a claim about one workload. NVIDIA wrote that the playbook "isn't a general recipe for every data science agent" [14]. The harness sections of the write-up do not name the model, report a score, or measure each technique on its own, so the share of the second-place finish owed to any single change is unknown. None of the four listed functions is described as handling video, though briefing videos were among the competition's sources [20].
For the pattern to transfer, a few things have to be true in your setup. Your structured data has to load into one queryable store before the agent starts. Your model has to be fixed, or expensive to change. Your tasks have to be narrow enough that a handful of functions covers them. NVIDIA's suggested stand-ins for the temporary SQLite database are a virtualized query layer or a governed warehouse interface [18].
I think this is the right tradeoff when the model is not yours to choose. The post's own summary: "Reliability often comes less from making the model more open-ended and more from building the right harness around it." [19]
What to watch
- Whether NVIDIA or the KDD Cup organizers publish KGMON's score or per-technique ablations, which would show which change moved the result.
- Whether the first-place team's write-up describes a similarly narrow harness or a different approach under the same fixed model.
- Whether the competition's fixed model is named, letting teams on that open model test the pattern directly.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+10
- Incentives55
- Confidence40
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The NVIDIA KGMON team placed second in the KDD Cup 2026 Data Agents competition with a system built around making an agent's harness smaller, clearer, and easier to verify.
- [2]
The competition asked agents to answer natural-language questions over heterogeneous data sources, including databases, CSV and JSON files, prose documents, PDFs, and briefing videos.
- [3]
Every task required more than retrieval: inspecting available data, choosing the right tools, reasoning across sources, producing a final answer file, and handling traps that appear in analytical workflows.
- [4]
KDD required teams to work with a small, fixed LLM to power the agent, making the harness the main optimization surface.
- [5]
KGMON converted CSV and JSON files into tables in the existing SQLite database, giving the agent one SQL interface for all structured data.
- [6]
KGMON's custom, persistent Python environment exposed schema() and sql(query) for inspecting and querying the unified data; they were built into KGMON's environment rather than provided by Python or SQLite.
- [7]
The single SQL interface reduced routing failures and wasted turns, leaving the fixed model more room to reason over the data.
- [8]
KGMON added a schema-scouting step before the main reasoning loop that inspected tables, columns, possible join keys, duplicate names, look-alike fields, units, null patterns, and row-grain issues.
- [9]
The agent received the schema context at the start of each task, saving an early discovery turn.
- [10]
The environment exposed four functions: schema() to inspect tables and columns, sql(query) to query context.db, write_answer(df) to write the final answer atomically, and prose_helper() to answer from prose or extract prose into SQL.
- [11]
Middleware repaired malformed tool calls so one bad call did not end an attempt.
- [12]
A stateful Python environment retained variables between tool calls, letting the agent reuse intermediate results.
- [13]
KGMON required final answers to follow a single output path.
- [14]
"It isn't a general recipe for every data science agent"
- [15]
Execution traces, repeated attempts, and trajectory inspection helped the team identify failures and refine the harness.
- [16]
Short, valid attempts left more turns for additional runs and evaluation.
- [17]
Schema scouting reduced errors caused by wrong columns, missed joins, or confusion about what each answer row should represent.
- [18]
NVIDIA recommends normalizing structured-data access before the agent starts via a temporary SQLite database, a virtualized query layer, or a governed warehouse interface, and says a read-only preflight briefing can be more useful than another retrieval tool.
- [19]
"Reliability often comes less from making the model more open-ended and more from building the right harness around it."
- [20]
None of the four listed functions is described as handling video, though briefing videos were among the competition's data sources.
Sources
1 independent publisher whose own reporting we read for this story.
- developer.nvidia.comBuilding Reliable Data Analytics Agents: Lessons from the KDD Cup
1 article · October 8, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.