Proceedings · Published ·
Critique-Guided Distillation for Robust Reasoning via Refinement
Berkcan Kapusuzoglu, Supriyo Chakraborty, Zain Sarwar, Chia-Hsuan Lee, Sambit Sahu
International Conference on Machine Learning (ICML 2026) · Proceedings
Summary
Critique-Guided Distillation trains a student to refine flawed responses using teacher critiques as training-only supervision, without requiring critiques at inference.
Research contribution
The work introduces a training framework that separates critique consumption during fine-tuning from critique generation at inference.
Scroll the figure sideways to read its labels, or open the full-size image below.
Lead figure

Teach the student to act on feedback during training so it can answer in one pass at inference.
A student learns from teacher critiques and refined answers during training, then answers without a teacher at inference.
Source: Berkcan Kapusuzoglu et al., Critique-Guided Distillation (2026), Figure 1. License: CC BY 4.0; rasterized without scientific edits.
Open full-size figure for labels and details.
What the results show
Scroll the figure sideways to read its labels, or open the full-size image below.
Result figure

In this experiment, learning to use critiques improves accuracy beyond learning refined answers alone or learning to generate critiques.
CGD improves LLaMA3.1-8B reasoning performance compared with distilled SFT and critique fine-tuning. The experiment uses a LLaMA3.3-70B Instruct teacher and 100K WebInstruct samples.
Source: Berkcan Kapusuzoglu et al., Critique-Guided Distillation (2026), Figure 2. License: CC BY 4.0; rasterized without scientific edits.
Open full-size figure. Some numeric labels touch in the source figure. Compare the bars here and use the full-size figure for closer inspection.