Bar chart on a dark ground of mean end-to-end latency per query, hybrid agent in green against full-LLM agent in red on three machines: 22.91 s against 87.95 s on a 32 GB CPU, 1.53 s against 5.86 s on an RTX 4060 Ti, and 4.96 s against 23.94 s on an M5 Pro CPU.

Agentic AI · 2026

Distilled models for a faster, cheaper hybrid AI agent

A hybrid agent that answers 3.8–4.8× faster and at half the cloud cost of an all-cloud agent, keeping the data rows on-premises.

The short version

Agents that answer what-if questions on tabular data call a prediction model many times per answer. A tabular foundation model (TabPFN) takes 12.6 s per prediction on a CPU, too slow for a multi-step agent, and an agent that runs entirely on a cloud LLM sends sensitive data to an external API. The question was how much of the work can move on-premises without losing answer quality.

Role
Co-author, with Sourish Dey (SumUp)
Period
2026
PythonLangGraphLangSmithOpenAI function callingQwen2.5-3B-InstructPyTorchTabPFNknowledge distillation

Bar chart of mean cloud API cost per 1,000 queries: 0.0959 US dollars for the hybrid agent, which makes one LLM tool call and generates the answer with a local small model, against 0.2021 dollars for the full-LLM agent that runs only on the cloud LLM. One schema-constrained cloud call per query halves the token bill.

How

  • Hybrid agent: a single cloud LLM call (gpt-4o-mini) picks the tool and parses the question; a model distilled from TabPFN makes the prediction and a local Qwen2.5-3B model writes the answer, both on-premises.
  • Distilled model: a 17,059-parameter MLP that keeps 95–100% of TabPFN's accuracy across six datasets, and carries almost all of the speed-up.

Details and full results in the preprint, arXiv:2609.16091 (Sourish Dey and Aditya Kumar, 2026), not peer reviewed.