Microsoft ends support for Internet Explorer on June 16, 2022. We recommend using one of the browsers listed below. Please contact your browser provider for download and installation instructions. June 1, 2026 NTT, Inc. News Highlights: TOKYO — June 1, 2026 — NTT, Inc. (Headquarters: Chiyoda-ku, Tokyo; President and CEO: Akira Shimada; hereinafter "NTT") has established Rationale-Enhanced Decoding, a new inference framework designed to improve the reliability of outputs generated by multimodal foundation models that process both images and language. The technology addresses a key issue in CoT reasoning by LVLMs: the tendency to ignore self-generated rationales. Unlike conventional inference methods, the proposed approach separately performs image-based inference and rationale-based inference, then combines them through ensemble decoding. This enables the model to generate responses grounded in information derived from both visual inputs and rationales. This research will be presented at the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026(*1), one of the world's premier international conferences in the field of computer vision, to be held in Denver, Colorado, USA, from June 3 to June 7, 2026. In recent years, the development of Large Vision-Language Models (LVLMs), which integrate Large Language Models (LLMs) with pretrained image encoders, has significantly advanced multimodal reasoning capabilities. Unlike text-only LLMs, LVLMs can directly process visual inputs in addition to text, enabling their use as a foundation for complex multimodal reasoning tasks based on visual content, such as video analysis and document understanding, which are difficult to address using text alone. Similar to LLMs that operate solely on text inputs, Chain-of-Thought (CoT) reasoning has also been regarded as an effective approach for improving inference performance and enabling explainable reasoning in LVLMs. In CoT reasoning, the model first generates intermediate rationales from visual and textual inputs, then appends those rationales to the input sequence to