< The research team. From left: Sangyun Chung (KAIST, first author of the MAD study and co-first author of the DNA study); Yong Man Ro (KAIST, corresponding author); Youngjoon Yu (KAIST, co-first author of the DNA study) > Multimodal large language models (MLLMs), which process multiple types of sensory information such as text, images, and audio at the same time, are rapidly expanding the range of applications for artificial intelligence (AI). However, in real-world environments, these models can misinterpret the physical characteristics of sensors, mistakenly identify objects, or claim to hear sounds that are not actually present simply because a certain object appears in a video. These errors are known as hallucinations. A KAIST research team has developed a new technology that corrects such information confusion and physical misperceptions in AI. KAIST (President Choongsik Bae) announced on the 31st of July that a research team led by Professor Yong Man Ro from the School of Electrical Engineering has developed two core technologies that overcome the tendency of existing large language models to rely too heavily on ordinary camera (RGB) images and enable AI to suppress cross-modal hallucinations that occur when different sensory inputs become mixed. < Figure 1. Examples of vision sensor-related questions and responses by recent. It fails to understand the core principle of thermal imaging, incorrectly attributing brightness to reflected sunlight rather than emitted heat. (Courtesy of KAIST) > The first technology developed by the research team is the Diverse Negative Attributes (DNA) optimization method, which helps AI accurately understand the physical characteristics of special camera sensors such as thermal, depth, and X-ray sensors. Existing AI models often failed to understand the physical meaning of such images, for example by mistaking bright areas in thermal images for simple light reflection. The research team built VS-TDX, the first comprehensive