My general research interest is how the brain represents visual information for our perception and decisions. I have been investigating the computational mechanism of perceptual decisions and spatiotemporal structure of visual representation underlying ensemble perception and object/scene recognition, using a multi-faceted approach, including psychophysical experiments, computational modeling, EEG, fMRI recordings, and deep neural networks.
I am currently analyzing large-scale fMRI datasets using various brain-to-model alignment methods and computational models, with the aim of understanding how the brain represent rich information beyond what is physically present in natural scenes.
In this page, you will find an overview of both my past work and ongoing research projects, along with links to papers for further details. If any of these topics attract your interest, please feel free to contact me. I am happy to have a discussion with you.
Keywords: neural decoding / encoding model / perceptual decision making / ensemble perception / computational modeling / EEG / fMRI / artificial neural network / generative model
Contact: ryuto.yashiro (at) gmail.com
With the advent of deep neural networks, the past decade has seen substantial progress in modeling neural responses. For example, embeddings of scene captions derived from large language models (LLMs) have a high capacity to predict responses in category-selective regions to natural scenes (e.g., Doerig et al, 2025), indicating that those regions represent scene semantics beyond the mere presence of single categories, such as faces and bodies. However, because of the limited interpretability of LLM embeddings, it is hard to understand what aspects of natural scenes constitute the semantic representations.
To tackle this issue, we developed a new data-driven approach based on object co-occurrence to identify what pairs of objects drive responses in category-selective regions. Applying this framework to regions that respond strongly to human bodies (extrastriate body area; EBA), we found that responses in the EBA are modulated by objects co-occurring with human bodies in a natural scene. This framework led us to identify multiple body-related features contributing to the semantic representations in the body-selective regions, including the speed of human body motion implied in static images, number of people and body size.
This framework can be extended to any regions and categories, offering a promising approach for understanding neural representations underlying natural scene perception, even in this era of increasingly complex and difficult-to-interpret features.
Paper: Probing the content of semantic representations in body-selective regions
We defined a co-occurrence matrix that counts the pairwise frequency of object categories in a set of scene captions (A). By combining LLM-based encoding models and multiple co-occurrence matrices derived from groups of captions, major components are identified, which reflect co-occurrence patterns and their contributions to each group. This approach led us to identify three key body-related features contributing to body-selective regions (EBA and FBA), including implied body motion, number of people and body size.
The visual system is capable of computing an average of multiple pieces of visual information in the environment (ensemble perception). However, it remains unclear how and when ensemble perception is formed in the brain, because most previous studies on ensemble perception only used behavioral experiments.
To address this question, we decoded the temporal dynamics of orientation representations from EEG signals while human participants judge an average of multiple orientations. The decoded orientation representation showed that orientation ensemble perception is formed over approximately 600-700 ms after stimulus onset.
Paper: Decoding time-resolved neural representations of orientation ensemble perception
We used inverted encoding models to decode the representational strength of average orientation (central row in the colormap) from EEG signals. The average orientation was strongly represented in EEG signals from 400 to 700 ms after stimulus onset, as shown in the right panel.
Humans make decisions based on sensory information that fluctuates over time. However, we cannot focus on all pieces of information to make decisions given the limited capacity of our visual system. This naturally leads us to the following question: what information do we use for making decisions?
Using a simple perceptual decision task, we found that humans overweight outliers occurring later in time and underweight outliers occurring earlier in time. We also found that a simple evidence accumulation model (leaky integration model) can account for this tendency. Such time-dependent decision weighting can be described as "peak-at-end" rule (similar to the peak-end rule proposed in behavioral economics), which may potentially underlie the general mechanism of perceptual decisions.
Paper: Peak-at-end rule: adaptive mechanism predicts time-dependent decision weighting
If you have broad interest in the computational mechanism of perceptual decisions, these papers may also match your interest: Perception and decision mechanisms involved in average estimation of spatiotemporal ensembles
Prospective decision making for randomly moving visual stimuli
Human participants were presented with a sequence of Gabor patterns and asked to judge if the temporal average orientation was tilted clockwise or counterclockwise relative to the vertical.
A simple leaky integration model predicts time-dependent decision weighting, with outliers presented earlier underweighted and those presented later overweighted. We found that human observers weight orientation information at each temporal frame in a qualitatively similar manner to the model's prediction.