Showing posts with label vision. Show all posts
Showing posts with label vision. Show all posts

Sunday, July 26, 2026

ChatGPT draws Rorschach blots

Here they are, six of them:

I then discussed the blots with ChatGPT in terms suggested by me new working paper, A prototypical image in ChatGPT 5.6: An informal pilot study. First I present all six blots, full size, and then the discussion.

I needed a Rorschach inkblot for a post I was working on, so I had ChatGPT create one. I then decided to see what would happen if I asked for a number of inkblots, each in a different chat. I’ve uploaded the result to the project in a document entitled Rorschach Blots. I had no expectation about what would happen. Would it produce the same blot each time? Of different blots each time? As you can see, it produced different blots each time.

This suggests two further lines of investigation. 1) Generate 100 independent ink blots and run them through an image classifier. Would we get distinct clusters or not? 2) Use prompts suggesting specific things, but otherwise open-ended, e.g. “Draw a cat,” “Draw a tree,” “Draw a car,” “Draw a coat,” etc.

Looking at the Rorschach Blots document, the six outputs are plainly different, but they are not unconstrained in every respect. All six preserve a common structural grammar: bilateral symmetry, a strong vertical axis, black-and-gray ink on white, irregular feathered edges, and enough internal articulation to invite figurative interpretation. Within that shared format, however, the overall silhouettes vary substantially: some are tall and narrow, some compact and squat, some strongly winged, some mask-like, and some more diffuse. The contact sheet on page 1 makes that balance of sameness and difference especially clear.

That suggests that the system has not selected one prototypical blot, in the way it selected one broad landscape type. Instead, it seems to have selected a prototypical design space for “Rorschach blot,” within which it samples different instances. Your two proposed extensions are therefore well chosen.

1. One hundred independent inkblots

The clustering question is exactly the right one. A hundred examples might reveal:

  • one continuous cloud of variation around a single template;
  • several recurring morphological families;
  • or a few dominant families plus outliers.

Even by eye, the six examples hint at possible families: elongated “totemic” forms, broad winged forms, compact mask-like forms, and more open forms with substantial white interior space. But six is far too few to know whether those are genuine modes or just impressions.

I would not rely on a conventional object classifier, since it may simply label them all “inkblot,” “butterfly,” “mask,” or “bat.” A more informative analysis would use image embeddings and cluster those embeddings, perhaps supplemented by simple measurable features such as aspect ratio, occupied area, contour complexity, amount of internal white space, and vertical distribution of mass. The point would not be to decide what each blot “looks like,” but to determine the structure of variation across the set.

2. Open-ended object prompts

“Draw a cat,” “Draw a tree,” “Draw a car,” and so forth would probe something different and potentially more revealing: whether the system has a default realization for a named category.

The questions would include:

  • Does “Draw a cat” repeatedly produce the same breed, pose, viewpoint, or expression?
  • Does “Draw a tree” default to a solitary deciduous tree, a stylized botanical specimen, or a landscape scene?
  • Does “Draw a car” default to a modern sports car, a generic sedan, or a three-quarter advertising view?
  • Does “Draw a coat” produce the garment alone, on a mannequin, or worn by a person?

That experiment would occupy a useful middle ground between the wholly unspecified prompt—“Create a drawing”—and a detailed scene description. The subject is fixed, but almost everything about its realization remains open. It would therefore reveal category-level defaults in pose, style, composition, context, and affect.

The Rorschach result also suggests a distinction worth preserving in the paper’s future-work section:

Some prompts may elicit a single dominant subject prototype, while others may elicit a constrained family of forms rather than one recurring image.

That is the most interesting implication of the blot experiment. The landscape prompt collapsed onto a narrow semantic prototype. The inkblot prompt produced morphological diversity within a narrow formal convention. Those are different kinds of regularity, and comparing them could become a substantial research program.

Monday, June 22, 2026

Brain area specialized for visual recognition of words

Monday, April 27, 2026

The Platonic Representation Hypothesis [Not surprising]

Thursday, June 20, 2024

How an infant investigates a banana [left and right]

The infant’s mother has cut a banana in half, and then sliced one of the halves lengthwise. She then places the two slices flat-side down on the tray in front of her infant.

Three things interest me about what the infant does:

  1. While it is able to grab a banana slice and move it around on the tray, it can’t get its fingers around it in order to pick it up. So, mother turns one piece over and places it back on the tray.
  2. Now the infant manages to grab that piece with its right hand. Notice, though, when it finally manages to pick it up, it’s not looking at its hand. It appears to be looking at mama. That action is guided entirely by touch and movement.
  3. After it has brought the banana to its mouth and managed to at least taste it a bit, what does it do? After it visually inspects the banana a bit, it moves its left had toward the banana and touches the tip with its left index finger.

That last item is what interests me most. I believe that it isn’t until about six months that infants are able to coordinate the right and left halves of their body in the same space. Before that time, the left and right hands, in effect, don’t “know” about one another. So, when the infant touches the banana with its left hand, it’s making sure that the banana is at the same position in space for both hands. 

NOTE: Remember that the two halves of the body are controlled by the opposite halves of the brain. The right half of the body is controlled by the left half of the brain, and vice versa. While the two halves of the brain certainly communicate with one another through the corpus callosum, a bundle of the 200-300 million neurons, they are not a continuous field. They’ve got to ‘learn’ about one another.

Tuesday, May 14, 2024

Brain encoding models that can transfer across language and vision

Jerry Tang, Meng Du, Vy Vo, VASUDEV LAL, Alexander Huth, Brain encoding models based on multimodal transformers can transfer across language and vision, Advances in Neural Information Processing Systems 36 (NeurIPS 2023) Main Conference Track

Abstract: Encoding models have been used to assess how the human brain represents concepts in language and vision. While language and vision rely on similar concept representations, current encoding models are typically trained and tested on brain responses to each modality in isolation. Recent advances in multimodal pretraining have produced transformers that can extract aligned representations of concepts in language and vision. In this work, we used representations from multimodal transformers to train encoding models that can transfer across fMRI responses to stories and movies. We found that encoding models trained on brain responses to one modality can successfully predict brain responses to the other modality, particularly in cortical regions that represent conceptual meaning. Further analysis of these encoding models revealed shared semantic dimensions that underlie concept representations in language and vision. Comparing encoding models trained using representations from multimodal and unimodal transformers, we found that multimodal transformers learn more aligned representations of concepts in language and vision. Our results demonstrate how multimodal transformers can provide insights into the brain’s capacity for multimodal processing.

Sunday, October 29, 2023

Transmission Versus Truth, Imitation Versus Innovation: Children vs. LLMs

Yiu, E., Kosoy, E., & Gopnik, A. (2023). Transmission Versus Truth, Imitation Versus Innovation: What Children Can Do That Large Language and Language-and-Vision Models Cannot (Yet). Perspectives on Psychological Science, 0(0). https://doi.org/10.1177/17456916231201401

Abstract: Much discussion about large language models and language-and-vision models has focused on whether these models are intelligent agents. We present an alternative perspective. First, we argue that these artificial intelligence (AI) models are cultural technologies that enhance cultural transmission and are efficient and powerful imitation engines. Second, we explore what AI models can tell us about imitation and innovation by testing whether they can be used to discover new tools and novel causal structures and contrasting their responses with those of human children. Our work serves as a first step in determining which particular representations and competences, as well as which kinds of knowledge or skills, can be derived from particular learning techniques and data. In particular, we explore which kinds of cognitive capacities can be enabled by statistical analysis of large-scale linguistic data. Critically, our findings suggest that machines may need more than large-scale language and image data to allow the kinds of innovation that a small child can produce.

Wednesday, October 25, 2023

ConvNets Match Vision Transformers at Scale

Sunday, October 1, 2023

Internal feedback in the cortical perception-action loop enables fast and accurate behavior

Jing Shuang (Lisa) Lia, Anish A. Sarmaa, Terrence J. Sejnowskic, and John C. Doyle, Internal feedback in the cortical perception-action loop enables fast and accurate behavior, PNAS, September 22, 2023 120 (39) e2300445120 https://doi.org/10.1073/pnas.2300445120, asXiv: https://arxiv.org/abs/2211.05922

Significance

Internal feedback projections—signals flowing from motor areas or late sensory processing regions back to early sensory processing regions such as primary visual and auditory areas—are ubiquitous in the sensorimotor nervous system and are as or more numerous than feedforward projections. However, the function of internal feedback is poorly understood, particularly in the context of task performance. We leverage control theory and simple models to demonstrate that internal feedback facilitates good task performance when there are communication limitations such as internal time delays and speed–accuracy trade-offs, which motivate compensatory feedback signals to counter self-generated and predictable movements. Control theory explains why motor-related signals are found throughout the sensory cortex and why the motor cortex is dominated by internal dynamics.

Abstract

Animals move smoothly and reliably in unpredictable environments. Models of sensorimotor control, drawing on control theory, have assumed that sensory information from the environment leads to actions, which then act back on the environment, creating a single, unidirectional perception–action loop. However, the sensorimotor loop contains internal delays in sensory and motor pathways, which can lead to unstable control. We show here that these delays can be compensated by internal feedback signals that flow backward, from motor toward sensory areas. This internal feedback is ubiquitous in neural sensorimotor systems, and we show how internal feedback compensates internal delays. This is accomplished by filtering out self-generated and other predictable changes so that unpredicted, actionable information can be rapidly transmitted toward action by the fastest components, effectively compressing the sensory input to more efficiently use feedforward pathways: Tracts of fast, giant neurons necessarily convey less accurate signals than tracts with many smaller neurons, but they are crucial for fast and accurate behavior. We use a mathematically tractable control model to show that internal feedback has an indispensable role in achieving state estimation, localization of function (how different parts of the cortex control different parts of the body), and attention, all of which are crucial for effective sensorimotor control. This control model can explain anatomical, physiological, and behavioral observations, including motor signals in the visual cortex, heterogeneous kinetics of sensory receptors, and the presence of giant cells in the cortex of humans as well as internal feedback patterns and unexplained heterogeneity in neural systems.

Friday, March 3, 2023

High-resolution image reconstruction with latent diffusion models from human brain activity

Yu Takagi, Shinji Nishimoto, bioRxiv 2022.11.18.517004; doi: https://doi.org/10.1101/2022.11.18.517004

Abstract: Reconstructing visual experiences from human brain activity offers a unique way to understand how the brain represents the world, and to interpret the connection between computer vision models and our visual system. While deep generative models have recently been employed for this task, reconstructing realistic images with high semantic fidelity is still a challenging problem. Here, we propose a new method based on a diffusion model (DM) to reconstruct images from human brain activity obtained via functional magnetic resonance imaging (fMRI). More specifically, we rely on a latent diffusion model (LDM) termed Stable Diffusion. This model reduces the computational cost of DMs, while preserving their high generative performance. We also characterize the inner mechanisms of the LDM by studying how its different components (such as the latent vector of image Z, conditioning inputs C, and different elements of the denoising U-Net) relate to distinct brain functions. We show that our proposed method can reconstruct high-resolution images with high fidelity in straight-forward fashion, without the need for any additional training and fine-tuning of complex deep-learning models. We also provide a quantitative interpretation of different LDM components from a neuroscientific perspective. Overall, our study proposes a promising method for reconstructing images from human brain activity, and provides a new framework for understanding DMs. Please check out our webpage at https://sites.google.com/view/stablediffusion-with-brain/.

Tuesday, November 22, 2022

Bilingual readers possess two distinct visual word form areas (VWFA)

Abstract for the article linked above:

In expert readers, a brain region known as the visual word form area (VWFA) is highly sensitive to written words, exhibiting a posterior-to-anterior gradient of increasing sensitivity to orthographic stimuli whose statistics match those of real words. Using high-resolution 7T fMRI, we ask whether, in bilingual readers, distinct cortical patches specialize for different languages. In 21 English-French bilinguals, unsmoothed 1.2 mm fMRI revealed that the VWFA is actually composed of several small cortical patches highly selective for reading, with a posterior-to-anterior word similarity gradient, but with near-complete overlap between the two languages. In 10 English-Chinese bilinguals, however, while most word-specific patches exhibited similar reading specificity and word-similarity gradients for reading in Chinese and English, additional patches responded specifically to Chinese writing and, surprisingly, to faces. Our results show that the acquisition of multiple writing systems can indeed tune the visual cortex differently in bilinguals, sometimes leading to the emergence of cortical patches specialized for a single language.

Tuesday, October 4, 2022

Linearly Mapping from Image to Text Space

The paper's conclusion:

In this paper, we test the extent to which the representations of language models encode information about the non-linguistic world in terms of their ability to use image representations to perform vision-language tasks. We show through LiMBeR (Linearly Mapping Between Representation spaces) that training a linear (thus, distance-preserving) transformation to connect image features to an LM’s input space is competitive on image captioning and visual question answering benchmarks with similar models like MAGMA that tune both image and text networks. However, we also find that such transfer is highly dependant on the amount of linguistic supervision the image encoder backbone had during its pretraining phase. BEIT, which is a vision-only image encoder underperforms compared to CLIP, which was pretrained with natural language captions. We explore what conceptual information transfers successfully, and find through probing, clustering, and analysis of generated text that the representational similarity between LMs and vision-only image representations is mostly restricted to coarse-grained concepts of perceptual features. Our findings indicate that large LMs do appear to form models of the visual world along these perceptual concepts to some extent, but are biased to form categorical concepts of words that are not distinguished by vision-only models. We are excited by future work applying LiMBeR to other domains and modalities as a behavioral tool for understanding the representations of LMs and other deep neural networks.

Wednesday, May 18, 2022

A good tweet-stream on modeling vision in machines and humans

The whole thread is worthwhile, with links to other papers.

Friday, April 29, 2022

Input to the visual cortex

Tuesday, April 5, 2022

Remarks on the visual system and artificial neural nets

Saturday, December 18, 2021

Vision isn't "solved" [AI]

Abstract of linked article, Overinterpretation reveals image classification model pathologies:

Image classifiers are typically scored on their test set accuracy, but high accuracy can mask a subtle type of model failure. We find that high scoring convolutional neural networks (CNNs) on popular benchmarks exhibit troubling pathologies that allow them to display high accuracy even in the absence of semantically salient features. When a model provides a high-confidence decision without salient supporting input features, we say the classifier has overinterpreted its input, finding too much class-evidence in patterns that appear nonsensical to humans. Here, we demonstrate that neural networks trained on CIFAR-10 and ImageNet suffer from overinterpretation, and we find models on CIFAR-10 make confident predictions even when 95% of input images are masked and humans cannot discern salient features in the remaining pixel-subsets. We introduce Batched Gradient SIS, a new method for discovering sufficient input subsets for complex datasets, and use this method to show the sufficiency of border pixels in ImageNet for training and testing. Although these patterns portend potential model fragility in real-world deployment, they are in fact valid statistical patterns of the benchmark that alone suffice to attain high test accuracy. Unlike adversarial examples, overinterpretation relies upon unmodified image pixels. We find ensembling and input dropout can each help mitigate overinterpretation.

Wednesday, July 14, 2021

Deep Neural Networks are Surprisingly Reversible: A Baseline for Zero-Shot Inversion

This bears comparison with what William Powers had to say about imagination in Behavior: The Control of Perception (Aldine 1973).

Sunday, June 20, 2021

Heritable functional architecture in human visual cortex

Abstract of the linked article:

How much of the functional organization of our visual system is inherited? Here we tested the heritability of retinotopic maps in human visual cortex using functional magnetic resonance imaging. We demonstrate that retinotopic organization shows a closer correspondence in monozygotic (MZ) compared to dizygotic (DZ) twin pairs, suggesting a partial genetic determination. Using population receptive field (pRF) analysis to examine the preferred spatial location and selectivity of these neuronal populations, we estimate a heritability around 10-20% for polar angle preferences and spatial selectivity, as quantified by pRF size, in extrastriate areas V2 and V3. Our findings are consistent with heritability in both the macroscopic arrangement of visual regions and stimulus tuning properties of visual cortex. This could constitute a neural substrate for variations in a range of perceptual effects, which themselves have been found to be at least partially genetically determined. These findings also add convergent evidence for the hypothesis that functional map topology is linked with cortical morphology.

Thursday, April 22, 2021

Wiring diagram of the visual cortex

Abstract of the linked article:

The laminar location of the cell bodies and terminals of interareal connections determines the hierarchical structural organization of the cortex and has been intensively studied. However, we still have only a rudimentary understanding of the connectional principles of feedforward (FF) and feedback (FB) pathways. Quantitative analysis of retrograde tracers was used to extend the notion that the laminar distribution of neurons interconnecting visual areas provides an index of hierarchical distance (percentage of supragranular labeled neurons [SLN]). We show that: 1) SLN values constrain models of cortical hierarchy, revealing previously unsuspected areal relations; 2) SLN reflects the operation of a combinatorial distance rule acting differentially on sets of connections between areas; 3) Supragranular layers contain highly segregated bottom-up and top-down streams, both of which exhibit point-to-point connectivity. This contrasts with the infragranular layers, which contain diffuse bottom-up and top-down streams; 4) Cell filling of the parent neurons of FF and FB pathways provides further evidence of compartmentalization; 5) FF pathways have higher weights, cross fewer hierarchical levels, and are less numerous than FB pathways. Taken together, the present results suggest that cortical hierarchies are built from supra- and infragranular counterstreams. This compartmentalized dual counterstream organization allows point-to-point connectivity in both bottom-up and top-down directions.

Saturday, April 10, 2021

Convolutional neural nets and human vision processing

The abstract of the linked article:

Convolutional neural networks (CNNs) are increasingly used to model human vision due to their high object categorization capabilities and general correspondence with human brain responses. Here we evaluate the performance of 14 different CNNs compared with human fMRI responses to natural and artificial images using representational similarity analysis. Despite the presence of some CNN-brain correspondence and CNNs’ impressive ability to fully capture lower level visual representation of real-world objects, we show that CNNs do not fully capture higher level visual representations of real-world objects, nor those of artificial objects, either at lower or higher levels of visual representations. The latter is particularly critical, as the processing of both real-world and artificial visual stimuli engages the same neural circuits. We report similar results regardless of differences in CNN architecture, training, or the presence of recurrent processing. This indicates some fundamental differences exist in how the brain and CNNs represent visual information.

See the discussion of Levick's Law in my post, Showdown at the AI Corral, or: What kinds of mental structures are constructable by current ML/neural-net methods? [& Miriam Yevick 1975].