Empowering Small LLM Agents with Hierarchical Teacher Memory
Largest accuracy jump. Closes 99% of the gap to its teacher.
The student surpasses its teacher without a single gradient update.
A 4B model reaches teacher-level accuracy on AppWorld. Zero fine-tuning, memory only.
Small agents rarely succeed enough to build useful memory, and naive teacher-memory transfer fails on the capability gap. AMD closes it, training-free, with memory at three granularities: Workflow plans, Subtask examples, and Function call guides.
Memory banks dominated by failed trajectories. Self-evolution stalls.
"Log in first" is useless if the student doesn't know how to log in.
Plans + worked examples + call guides bridge the capability gap.
Built only from successful teacher trajectories. Planning knowledge arrives before execution; corrective knowledge arrives at the moment of failure.
Verbalized strategy with typed placeholders (<ID>, <EMAIL>).
Coherent segments with the teacher's executable code + observations.
⚡ Proactive · per subtaskPer-call records: masked args, returns, and the teacher's reasoning.
🔁 Reactive · on errorAMD beats every memory baseline on every model × benchmark, while baselines often fall below zero-shot. It also cuts interaction turns toward the teacher's efficiency.
| Method | AppWorld | BFCL V3 | ToolSandbox |
|---|---|---|---|
| Teacher Agent | |||
| GPT-5-mini | 50.00 | 36.50 | 28.68 |
| Student · Qwen3-4B | |||
| Zero-shot | 14.88 | 15.50 | 16.28 |
| ReasoningBank | 10.71 | 24.25 | 16.28 |
| MemP | 16.67 | 28.25 | 14.73 |
| SASM | 15.48 | 15.25 | 18.22 |
| AMD (ours) | 49.40 | 38.50 | 20.16 |
| Δ vs. zero-shot | ▲+34.5 | ▲+23.0 | ▲+3.9 |
| Student · Gemma4-E4B | |||
| Zero-shot | 24.40 | 37.25 | 18.22 |
| ReasoningBank | 30.36 | 40.25 | 15.12 |
| MemP | 20.83 | 35.25 | 15.89 |
| SASM | 22.02 | 36.75 | 12.40 |
| AMD (ours) | 54.17 | 46.00 | 21.71 |
| Δ vs. zero-shot | ▲+29.8 | ▲+8.8 | ▲+3.5 |
| Student · Qwen3-8B | |||
| Zero-shot | 25.60 | 38.00 | 20.16 |
| AMD (ours) | 51.79 | 45.50 | 25.58 |
| Δ vs. zero-shot | ▲+26.2 | ▲+7.5 | ▲+5.4 |
| Student · Llama3.1-8B | |||
| Zero-shot | 8.93 | 9.00 | 5.43 |
| AMD (ours) | 27.38 | 14.50 | 6.20 |
| Δ vs. zero-shot | ▲+18.5 | ▲+5.5 | ▲+0.8 |
Five ablations answer: which memory matters most, which teacher is best, which student sizes benefit most, how many memories to retrieve, and how to represent them.
| Method | AppWorld | BFCL V3 | ||
|---|---|---|---|---|
| Qwen3-4B | Qwen3-8B | Qwen3-4B | Qwen3-8B | |
| Zero-shot | 14.88 | 25.60 | 15.50 | 38.00 |
| WF | 22.02 | 30.36 | 35.50 | 40.00 |
| WF + FN | 24.11 | 33.93 | 35.50 | 41.50 |
| WF + ST | 47.02 | 51.19 | 37.50 | 45.50 |
| WF + ST + FN (AMD) | 49.40 | 51.79 | 38.50 | 45.50 |
| Student Memory | 16.07 | 29.76 | 27.00 | 43.00 |
| Teacher | Teacher Acc | Qwen3-4B | Qwen3-8B |
|---|---|---|---|
| Zero-shot | n/a | 14.88 | 25.60 |
| GPT-5.5 | 91.08 | 47.02 | 58.93 |
| DeepSeek V4 Pro | 81.55 | 38.10 | 57.14 |
| GPT-5-mini | 50.00 | 49.40 | 51.79 |
| Qwen3-32B | 34.42 | 29.76 | 39.29 |
| Workflow | Subtask | Function | Accuracy |
|---|---|---|---|
| Code | Code | Code | 44.05 |
| Text | Code | Code | 49.40 |
| Text | Text | Code | 23.21 |
| Text | Code | Text | 47.62 |
| Text | Text | Text | 26.19 |
Absolute gain peaks at +34.5 %p for Qwen3-4B; larger models close the gap without help.
k = 1 is near-optimal. Subtask accuracy collapses 49.4 to 33.3 as k grows to 5. Quality beats quantity.
Under cross-split (7:3) and self-excluded retrieval, hierarchical memory still gives monotonic gains: evidence of genuine transfer, not task overlap.
One Venmo task, three cascading failures. Workflow fixes the month-scope filter, Subtask escapes a date-parsing loop, Function fixes a dict-key crash. All three were necessary: ✓ solved in 7 steps.
@misc{kim2026agentmemorydistillationempowering,
title={Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory},
author={Taeil Kim and Kangsan Kim and Sung Ju Hwang},
year={2026},
eprint={2608.07169},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2608.07169},
}