Microsoft ends support for Internet Explorer on June 16, 2022. We recommend using one of the browsers listed below. Please contact your browser provider for download and installation instructions. June 3, 2026 Two papers authored by NTT Laboratories have been accepted at CVPR (The IEEE/CVF Conference on Computer Vision and Pattern Recognition) 2026 to be held in Denver, Colorado, USA, from June 3 to 7, 2026. This is a flagship conference on computer vision and pattern recognition where researchers seek the computational understanding, control, and generation of images and movies as well as their foundational theories. Abbreviated names of the laboratories: HI: NTT Human Informatics Laboratories CD: NTT Computer and Data Science Laboratories (The affiliations are at the time of submission.) Shin'ya Yamaguchi (CD)、Kosuke Nishida (HI)、Daiki Chijiwa (CD) Chain-of-Thought (CoT) prompting has been adapted for large vision-language models (LVLMs) to enhance multi-modal reasoning capability by generating intermediate rationales. However, our experiments reveal a key challenge: existing models often ignore the contents of these generated rationales during output generation. To address this issue, we propose Rationale-Enhanced Decoding (RED), a novel decoding strategy that requires no additional training. RED effectively harmonizes visual and rationale information by multiplying distinct image-conditional and rationale-conditional next-token distributions at the decoding time, ensuring the model's outputs are strictly grounded in the rationale. Extensive experiments demonstrate that RED consistently and significantly improves the dependency on rationales and reasoning performance across multiple benchmarks. This technology not only improves the performance of LVLM but also enhances the interpretability required for leveraging AI in critical decision-making. It is expected to find applications in various domains where humans and AI collaborate, including AI Constellation©. Shin'ya Yamaguchi (CD)、Tamao Sakao (CD)、Daiki Chijiwa (CD)、Taku Hasegawa (HI) Multi-modal in-context learning (MM-ICL) allows large vision-language models (LVLMs) to adapt to new tasks using demonstration examples. However, increasing