Large Language Models (LLMs) are increasingly employed in structured problem-solving, yet their unreliability in producing accurate, verifiable outputs limits their applicability in mathematically grounded domains such as decision science. In this work, we introduce a general framework inspired by self-refinement techniques, which combines Retrieval-Augmented Generation (RAG) with an iterative critique mechanism to enhance the correctness and consistency of generated solutions to decision-theoretic problems. Using four openly available LLMs, DeepSeek-R1, DeepSeek-V3, DeepSeek-R1T2 Chimera, and OpenRouter’s Horizon-Beta, we evaluate our framework on core tasks in decision science. Our methodology involves three comparative settings: (i) a zero-shot baseline, (ii) a RAG-enhanced generator with task-specific retrieval, and (iii) the full pipeline loop, where the model iteratively evaluates and refines its own answers. Input–output formatting challenges using prompt engineering techniques, such as chain-of-thought reasoning, have also been addressed to guide the model through intermediate steps. Inspecting the outcomes, RAG–CRITIC reduces factual errors by up to 21 percentage points and increases precision by up to 13 times over the baseline. This work highlights the potential of loop-based reasoning architectures for trustworthy decision support and consolidates the foundations for future research in LLM-assisted judgment under constraints of rationality and optimality.

Retrieval Augmented Generation with Iterative Critique: A Framework for Structured Task Solving

Mezzina, Alessio
;
Naso, Luca;Pavone, Mario
2026-01-01

Abstract

Large Language Models (LLMs) are increasingly employed in structured problem-solving, yet their unreliability in producing accurate, verifiable outputs limits their applicability in mathematically grounded domains such as decision science. In this work, we introduce a general framework inspired by self-refinement techniques, which combines Retrieval-Augmented Generation (RAG) with an iterative critique mechanism to enhance the correctness and consistency of generated solutions to decision-theoretic problems. Using four openly available LLMs, DeepSeek-R1, DeepSeek-V3, DeepSeek-R1T2 Chimera, and OpenRouter’s Horizon-Beta, we evaluate our framework on core tasks in decision science. Our methodology involves three comparative settings: (i) a zero-shot baseline, (ii) a RAG-enhanced generator with task-specific retrieval, and (iii) the full pipeline loop, where the model iteratively evaluates and refines its own answers. Input–output formatting challenges using prompt engineering techniques, such as chain-of-thought reasoning, have also been addressed to guide the model through intermediate steps. Inspecting the outcomes, RAG–CRITIC reduces factual errors by up to 21 percentage points and increases precision by up to 13 times over the baseline. This work highlights the potential of loop-based reasoning architectures for trustworthy decision support and consolidates the foundations for future research in LLM-assisted judgment under constraints of rationality and optimality.
2026
9783032218100
9783032218117
Decision Science
Game Theory
Generative AI
Large Language Models (LLMs)
Retrieval Augmented Generation
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.11769/733609
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
social impact