BERAG: Bayesian Ensemble Retrieval-Augmented Generation for Knowledge-based Visual Question Answering

Published in arXiv preprint

A retrieval-augmented generation framework that conditions the language model on each retrieved document separately rather than concatenating them into one context, then combines the outputs as a Bayesian ensemble whose weights are document posteriors updated during generation. This avoids the lost-in-the-middle effect of long contexts, yields document-level attribution, and supports detecting insufficient grounding and pruning documents for faster decoding. Evaluated on knowledge-based VQA, document VQA, and multimodal needle-in-a-haystack benchmarks.

Paper

Recommended citation: J. Chen, J. Mei, G. Yang, B. Byrne. "BERAG: Bayesian Ensemble Retrieval-Augmented Generation for Knowledge-based Visual Question Answering." arXiv preprint arXiv:2604.22678.
Download Paper