Research Note

Learning from mistakes with critique-guided distillation

How training-only teacher feedback helps a student refine its reasoning, and what the reported accuracy gains establish.

What a correct answer leaves out

When a student makes the same reasoning mistake repeatedly, a correct teacher answer may not explain which assumption needs to change. Distillation can pass along the destination while leaving that repair implicit. For an engineer designing training data, the useful question is whether feedback tied to the student’s own attempt can supply a better learning signal.

Critique as training context

Critique-Guided Distillation (CGD) conditions supervised training on a prompt, a student attempt, and teacher feedback, with the teacher’s refined answer as the target. The student learns to use feedback. At deployment, it receives only the prompt.

Follow the sequence

  1. Generate an initial student response to expose its errors.
  2. Ask the teacher for a critique of that response.
  3. Generate a refined teacher answer and train the student to produce it from the augmented context.
  4. Evaluate the trained student with prompts alone. No teacher call is added at inference.

The outline changes emphasis without altering the diagram. The complete figure remains available at every step.

Scroll the figure sideways to read its labels, or open the full-size image below.

Lead figure

Three-row CGD training diagram showing the student response, teacher critique and refined answer, and student fine-tuning.: Teacher-free inference

Step 4: Teacher-free inference. Prompt only. No teacher or critique at inference. This is the deployment behavior described in Section 3.2, not a fourth row in Figure 1.

  1. Initial student response
  2. Teacher critique
  3. Refined answer and student training
  4. Teacher-free inference

Step 4 of 4: Teacher-free inference. Final frame.

Use teacher critiques as training context; deploy the student with the prompt alone.

Figure 1 contains three training rows. The sequence highlights the initial response, teacher feedback, and refinement target. The final view keeps the full diagram and explains prompt-only inference separately.

Source: Kapusuzoglu et al. (2026), Figure 1 and Section 3. License: CC BY 4.0; exact source pixels with separate CSS focus outlines.

Open full-size figure

Read the result at the right scope

Figure 2 supports a comparison of CGD with distilled SFT and critique fine-tuning for the pictured student, teacher, data, and benchmarks. It does not establish reliability on arbitrary tasks or a causal explanation of every gain.

Scroll the figure sideways to read its labels, or open the full-size image below.

Result figure

Bar chart comparing the base model and four training methods across six reasoning evaluations for a LLaMA3.1-8B Instruct student; CGD exceeds distilled SFT and critique fine-tuning.

CGD improves accuracy over distilled SFT and critique fine-tuning in this experiment.

Figure 2 compares methods for LLaMA3.1-8B Instruct using a LLaMA3.3-70B Instruct teacher and 100K WebInstruct samples. The comparisons shown are specific to this setup.

Source: Kapusuzoglu et al. (2026), Figure 2. License: CC BY 4.0; rasterized without scientific edits.

Open full-size figure. Some numeric labels touch in the source figure. Compare the bars here and use the full-size figure for closer inspection.

What to measure before adopting it

For a new application, compare prompt-only accuracy with answer distillation under matched data and training budgets. Inspect whether feedback identifies a specific error and whether the refined answer actually repairs it. Track instruction following and failures beyond the training domain, alongside response length and latency.

The paper’s limitations include dependence on informative teacher feedback, data generation cost, and evaluation focused on structured reasoning. A teacher-free inference path avoids extra critique calls; it does not guarantee shorter responses or lower end-to-end cost.

The practical lesson is to make the repair signal explicit during training and test its transfer to the deployment input. Treat that transfer as an evaluation question each time the student, teacher, or task changes.

Paper and citation

Berkcan Kapusuzoglu, Supriyo Chakraborty, Zain Sarwar, Chia-Hsuan Lee, and Sambit Sahu. Critique-Guided Distillation for Robust Reasoning via Refinement. Proceedings of the 43rd International Conference on Machine Learning, PMLR 306, 2026. Proceedings record. Public manuscript, version 4. Publication page.