Abstract Vision transformers (ViTs) process images as sequences of embedded tokens, yet in existing architectures, tokenization is performed entirely in the digital domain, downstream of image capture. This separation increases energy cost and prevents token formation from occurring where visual information is physically generated. Here we introduce a physical tokenizer that performs analogue light-to-token conversion directly at the sensor level for ViTs. The prototype integrates a 32 × 32 photosensitive memory array of monolayer MoS2 floating-gate phototransistors with peripheral addressing circuitry, enabling optical images to be stored as non-volatile states and selectively combined into patch embeddings in situ, thereby eliminating separate sensing, patch division and patch embedding stages. When deployed in a standard ViT pipeline, the physical tokenizer achieves software-comparable accuracy on CIFAR-10 while reducing energy consumption by 14.3-fold relative to a digital tokenizer. These results establish physical tokenization as an energy-efficient and scalable hardware foundation for edge-intelligent, data-intensive vision systems. This is a preview of subscription content, access via your institution Access options Subscribe to this journal Receive 12 digital issues and online access to articles $119.00 per year only $9.92 per issue Buy this article - Purchase on SpringerLink - Instant access to the full article PDF. USD 39.95 Prices may be subject to local taxes which are calculated during checkout Similar content being viewed by others Subjects Data availability The data supporting the findings of this study are available from the corresponding authors upon reasonable request. Source data are provided with this paper. Code availability The code supporting the findings of this study is available from the corresponding authors upon reasonable request. References - Dosovitskiy, A. et al. An image is worth 16 × 16 words: transformers for image recognition at scale. In Proc. International Conference on Learning Representations (ICLR, 2021). - Liu, Z. et al. Swin
Direct light-to-token conversion with integrated 2D photosensitive memory | Nature Sensors
Read the original article
nature.com →