By Jiwei Liu
Publication Date: 2026-10-08 18:30:00
The NVIDIA KGMON team placed second in the KDD Cup 2026 Data Agents competition with a system built around a simple idea of making an agent’s harness smaller, clearer, and easier to verify.
The competition asked agents to answer natural-language questions over heterogeneous data sources, including databases, CSV and JSON files, prose documents, PDFs, and briefing videos. Every task required more than retrieval, with the agent inspecting available data, choosing the right tools, reasoning across sources, producing a final answer file, and handling traps that appear in analytical workflows.
KDD also required teams to work with a small, fixed LLM to power the agent, making the harness the main optimization surface. The team aimed to make the task easier for the available model through preprocessing, constrained tools, persistent state, and evaluation. These techniques are especially useful when building reliable systems around smaller open models.
This post shares the playbook behind their result. It isn’t a general recipe for every data science agent, but the broader lesson applies to many agent systems. Reliability often comes less from making the model more open-ended and more from building the right harness around it.
Two principles behind the system
Two principles shaped the system.
First, constrain the action space. Too many ways to inspect data, call tools, write files, or recover from mistakes can lead to failures. KGMON unified data access,…


