Building Reliable Data Analytics Agents: Lessons from NVIDIA's KDD Cup Approach

· Engineer's Notes · Cem Koyluoglu

NVIDIA's second‑place KDD Cup solution shows how constrained toolsets, schema scouting, and trace logging improve reliability of data analytics agents.

What happened

NVIDIA's KGMON team placed second in the KDD Cup 2026 Data Agents competition. The competition required agents to answer natural‑language questions over heterogeneous sources—including SQL databases, CSV and JSON files, prose documents, PDFs, and briefing videos—while operating with a small, fixed large language model (LLM) and without internet access. The team built a constrained harness around the LLM, unifying all structured sources into a single SQLite database with a narrow schema and exposing only a few custom functions: schema(), sql(query), write_answer(df), and prose_helper(). A preflight schema‑scouting step supplied the agent with table definitions, join keys, and data‑quality flags before the main reasoning loop. Middleware repaired malformed tool calls, and a persistent Python environment retained intermediate results across calls. Document inspection was handled separately: targeted previews and regex searches fed prose_helper, which invoked a zero‑temperature LLM call to extract answers or tables, keeping raw prose out of the primary context. Every attempt logged prompts, tool calls, SQL queries, intermediate results, and errors, enabling a specialized inspector agent to categorize failures and guide harness improvements.

Why it matters in production

The described architecture highlights several production‑relevant considerations. Constraining the action space to a small, well‑defined toolset reduces routing failures and limits the surface for malformed inputs, directly improving reliability and lowering latency per turn. Normalizing access to structured data through a unified SQLite interface eliminates the need for the model to discover disparate query mechanisms, preserving token budget for reasoning. Preflight schema scouting eliminates early discovery turns, decreasing overall latency and the risk of incorrect joins or column selections. Separating prose handling from structured analysis prevents large documents from exhausting the LLM’s context window, which can otherwise increase compute cost and introduce latency spikes. Comprehensive trace logging creates an audit trail that can be inspected by downstream agents or humans, facilitating rapid failure diagnosis and continuous improvement without manual replay. However, repeated attempts for ensembling, while boosting coverage, increase token usage, latency, and compute expense; production pipelines must balance the marginal reliability gain against these costs. Finally, keeping humans in the loop for task definition, trace auditing, and harness refinement ensures that automated improvements do not overfit to benchmark quirks, preserving generalization to real‑world workloads.

Engineering takeaways

  • Normalize data access: Convert all structured sources into a single query surface (e.g., a temporary SQLite database) and expose only schema() and sql() functions to the agent.
  • Provide upfront schema context: Implement a read‑only scouting step that briefs the agent on tables, join keys, units, and data‑quality flags before reasoning begins.
  • Limit the toolset: Design a small, opinionated collection of helper functions and include middleware that repairs malformed calls, preserving turn efficiency.
  • Separate prose handling: Offer preview and regex‑based document tools, and delegate full‑text extraction to a dedicated zero‑temperature LLM call, keeping large documents out of the main context.
  • Log and inspect execution traces: Capture prompts, tool calls, and intermediate results for automated or human inspection, enabling rapid identification of failure patterns and targeted harness improvements.

Related work on this site

  • Agentic & LLM Automation — Multi-step LLM pipelines that run unattended: model cascades with fallbacks, validation gates, scheduled automation and alerting when something breaks.
  • RAG & Grounded LLM Systems — LLM features that answer from the right source instead of from memory: document-grounded assistants, transcript-grounded chat and quality gates around generated output.
  • YouTube AI Summarizer — Open-source Chrome extension on the Chrome Web Store: AI summaries, key points, deep analysis, a two-host AI podcast (Gemini TTS) and transcript-grounded chat for any YouTube video, using Groq or Ollama Cloud with the user’s own key.
  • Automated AI News Pipeline (this site) — The pipeline behind this site’s Tech News: scheduled GitHub Actions scrape sources, an LLM cascade on Groq rewrites and enhances articles, and layered quality gates (date integrity, language checks, instruction-leak detection, duplicate detection) decide what gets published, with Telegram alerts.

This note was drafted with AI assistance from the primary source credited on this page, and automatically checked against that source before publishing.

Source: https://developer.nvidia.com/blog/building-reliable-data-analytics-agents-lessons-from-the-kdd-cup/

← All tech news