
Agentic AI · 2026
Distilled models for a faster, cheaper hybrid AI agent
A hybrid agent that answers 3.8–4.8× faster and at half the cloud cost of an all-cloud agent, keeping the data rows on-premises.
The short version
Agents that answer what-if questions on tabular data call a prediction model many times per answer. A tabular foundation model (TabPFN) takes 12.6 s per prediction on a CPU, too slow for a multi-step agent, and an agent that runs entirely on a cloud LLM sends sensitive data to an external API. The question was how much of the work can move on-premises without losing answer quality.
PythonLangGraphLangSmithOpenAI function callingQwen2.5-3B-InstructPyTorchTabPFNknowledge distillation
One schema-constrained cloud call per query halves the token bill.
How
- Hybrid agent: a single cloud LLM call (gpt-4o-mini) picks the tool and parses the question; a model distilled from TabPFN makes the prediction and a local Qwen2.5-3B model writes the answer, both on-premises.
- Distilled model: a 17,059-parameter MLP that keeps 95–100% of TabPFN's accuracy across six datasets, and carries almost all of the speed-up.
Details and full results in the preprint, arXiv:2609.16091 (Sourish Dey and Aditya Kumar, 2026), not peer reviewed.