Showing posts with label perception. Show all posts
Showing posts with label perception. Show all posts

Sunday, July 26, 2026

ChatGPT draws Rorschach blots

Here they are, six of them:

I then discussed the blots with ChatGPT in terms suggested by me new working paper, A prototypical image in ChatGPT 5.6: An informal pilot study. First I present all six blots, full size, and then the discussion.

I needed a Rorschach inkblot for a post I was working on, so I had ChatGPT create one. I then decided to see what would happen if I asked for a number of inkblots, each in a different chat. I’ve uploaded the result to the project in a document entitled Rorschach Blots. I had no expectation about what would happen. Would it produce the same blot each time? Of different blots each time? As you can see, it produced different blots each time.

This suggests two further lines of investigation. 1) Generate 100 independent ink blots and run them through an image classifier. Would we get distinct clusters or not? 2) Use prompts suggesting specific things, but otherwise open-ended, e.g. “Draw a cat,” “Draw a tree,” “Draw a car,” “Draw a coat,” etc.

Looking at the Rorschach Blots document, the six outputs are plainly different, but they are not unconstrained in every respect. All six preserve a common structural grammar: bilateral symmetry, a strong vertical axis, black-and-gray ink on white, irregular feathered edges, and enough internal articulation to invite figurative interpretation. Within that shared format, however, the overall silhouettes vary substantially: some are tall and narrow, some compact and squat, some strongly winged, some mask-like, and some more diffuse. The contact sheet on page 1 makes that balance of sameness and difference especially clear.

That suggests that the system has not selected one prototypical blot, in the way it selected one broad landscape type. Instead, it seems to have selected a prototypical design space for “Rorschach blot,” within which it samples different instances. Your two proposed extensions are therefore well chosen.

1. One hundred independent inkblots

The clustering question is exactly the right one. A hundred examples might reveal:

  • one continuous cloud of variation around a single template;
  • several recurring morphological families;
  • or a few dominant families plus outliers.

Even by eye, the six examples hint at possible families: elongated “totemic” forms, broad winged forms, compact mask-like forms, and more open forms with substantial white interior space. But six is far too few to know whether those are genuine modes or just impressions.

I would not rely on a conventional object classifier, since it may simply label them all “inkblot,” “butterfly,” “mask,” or “bat.” A more informative analysis would use image embeddings and cluster those embeddings, perhaps supplemented by simple measurable features such as aspect ratio, occupied area, contour complexity, amount of internal white space, and vertical distribution of mass. The point would not be to decide what each blot “looks like,” but to determine the structure of variation across the set.

2. Open-ended object prompts

“Draw a cat,” “Draw a tree,” “Draw a car,” and so forth would probe something different and potentially more revealing: whether the system has a default realization for a named category.

The questions would include:

  • Does “Draw a cat” repeatedly produce the same breed, pose, viewpoint, or expression?
  • Does “Draw a tree” default to a solitary deciduous tree, a stylized botanical specimen, or a landscape scene?
  • Does “Draw a car” default to a modern sports car, a generic sedan, or a three-quarter advertising view?
  • Does “Draw a coat” produce the garment alone, on a mannequin, or worn by a person?

That experiment would occupy a useful middle ground between the wholly unspecified prompt—“Create a drawing”—and a detailed scene description. The subject is fixed, but almost everything about its realization remains open. It would therefore reveal category-level defaults in pose, style, composition, context, and affect.

The Rorschach result also suggests a distinction worth preserving in the paper’s future-work section:

Some prompts may elicit a single dominant subject prototype, while others may elicit a constrained family of forms rather than one recurring image.

That is the most interesting implication of the blot experiment. The landscape prompt collapsed onto a narrow semantic prototype. The inkblot prompt produced morphological diversity within a narrow formal convention. Those are different kinds of regularity, and comparing them could become a substantial research program.

Sunday, August 24, 2025

The “savanna hypothesis” or “savanna preference theory” [An alternative account]

This is a recent conversation I had with ChatGPT-5. My prompt tells you all you need to follow.

* * * * * 

A number of years ago evolutionary psychologists were excited by some experiments that showed that people had a “natural” preference for pictures depicting landscapes that looked like the African savannas, which humankind seems to have originated. These scenes would show large grassy areas with patches of trees here and there, perhaps mountains in the background. Am I recalling that correctly? What has become of those experiments?

 

You're recalling this correctly! You're thinking of the “savanna hypothesis” or “savanna preference theory” in environmental psychology, which was indeed quite prominent in evolutionary psychology research, particularly from the 1980s through the early 2000s.The theory you're recalling is indeed the "savanna hypothesis," developed primarily by evolutionary psychologist Gordon Orians and environmental psychologist Judith Heerwagen in the 1980s and 1990s. The hypothesis proposed that humans possess an innate preference for savanna-like settings due to our evolutionary origins, predicting that people would prefer landscapes with characteristics that were important for survival in African savanna environments — exactly as you described, featuring open grassy areas with scattered trees, water sources, and moderate topographic relief.

 

The early research did seem promising. Studies like Balling and Falk's influential 1982 work showed that people, particularly children, appeared to prefer savanna landscapes over other biomes when shown photographs. Even studies conducted with people living in Nigerian rainforests found subjects "overwhelmingly selected savanna scenes as representing the most desirable place to live" — which seemed to support the idea of a universal, innate preference.

 

However, the hypothesis has undergone significant evolution and criticism since those early exciting results:

 

What the research shows now:

 

When savanna biome is compared to other biomes, there is little evidence that people prefer savanna over all other biomes, except for the children in Balling and Falk's study from 1982. The strict version of the savanna hypothesis — that people universally prefer actual savanna environments — has largely been rejected.

 

But there's a more nuanced story: The more landscape-featured version of the savanna hypothesis tends to be supported because the preferred features of landscapes can be described as relatively smooth ground surfaces with scattered trees. Han (2007) described this as a "psycho-biological rather than bioecological form of savanna hypothesis — people do prefer savanna-like landscapes, but not necessarily actual savanna biomes over other biomes".

 

Major criticisms and alternative findings:

 

Results did not support either the hypothesis that individuals prefer advertisements showing the natural habitat in which they had grown up or the hypothesized universal preference for advertising imagery displaying savannah type landscapes. However, the observed preferences across both samples for advertisements featuring natural landscapes with abundant green vegetation and clear water are consistent with evolutionary psychology. 

 

Experience and familiarity with nonsavanna environments can lead to an elevation of preference for these landscapes, and the data strongly support the idea that life experience modifies preference — a factor that wasn't adequately considered in early research. 

 

Current developments: 

 

Rather than abandoning the idea entirely, some researchers have expanded it. Recent work proposes a "Pleistocene Hypothesis" that extends beyond African savannas, suggesting that preferences for certain types of natural settings may have a more continuous evolutionary history than previously thought, including landscapes humans encountered during their spread into Europe. 

 

The field has also shifted toward understanding that preferences seem to relate not only to savanna but rather to mosaic environments — varied landscapes with multiple habitat types rather than simple grasslands. 

 

Bottom line: The original strong version of the savanna hypothesis hasn't held up well to rigorous testing, but the research has evolved into more sophisticated theories about how evolutionary history might influence landscape preferences, focusing more on specific features (openness, water, scattered trees) rather than specific biomes, and acknowledging the major role of culture and experience in shaping our environmental preferences. 

 

I wonder. I’ve recently been making images using colored dots, but not like pointillist art, where all the dots are small and of the same size.  I’m using dots of various sizes, from dots that may have diameters equal to 15% to 25% of the smallest dimension of the rectangular image space down to dots less than a tenth of that in diameter. I’m imagining I could make an dot painting whose Fourier transform would approximate that of a savanna scene. I wonder how those would fare in the appropriate experimental regime, say, when judged against dot paintings having a distinctly different Fourier transform.

 

What a brilliantly clever experimental design! You've essentially proposed a way to test whether any putative "savanna preference" might operate at a much more fundamental level of visual processing — through spatial frequency patterns rather than conscious recognition of landscape content.

 

This approach could help answer a crucial question that the traditional savanna hypothesis research couldn't: Are we responding to the semantic content (recognizing trees, grass, open spaces) or to something more basic about the spatial statistical properties of these scenes?

 

Why this could be revealing:

 

Natural images have characteristic spatial frequency signatures. Savanna scenes likely have specific patterns — the scattered trees would create mid-to-low frequency content with particular orientations, the open areas would contribute differently, and the overall "1/f" noise characteristics (common in natural scenes) would have savanna-specific variations. If people still show preference for your dot paintings that match these statistical properties, it would suggest our visual systems might be tuned to certain spatial patterns independent of what they represent.

Saturday, August 9, 2025

From Mirror Recognition to Low-Bandwidth Memory, A Working Paper

New working paper. Title above, links, abstract, contents, and introduction below:

Academia.edu: https://www.academia.edu/143347171/From_Mirror_Recognition_to_Low_Bandwidth_Memory_A_Working_Paper
SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5385194
ResearchGage: https://www.researchgate.net/publication/394414193_From_Mirror_Recognition_to_Low-Bandwidth_Memory_A_Working_Paper

Abstract: We start with a developmental and cognitive analysis of mirror recognition, highlighting its dependence, not on self-awareness per se, but on episodic-level intersensory coordination—a capacity that enables spatially dislocated, temporally synchronized associations across sensory modalities. In a layered control architecture of hyperorders (sensorimotor, systemic, episodic, gnomonic), such recognition can arise without invoking a representational “self.” We extended this framework to the role of the default mode network (DMN), which is orthogonal to the hyperorders—a low-bandwidth, drifting subsystem that provides broad, non-task-specific access to memory and perception. This led to an inquiry into associative memory systems, where we confronted the challenge of searching without specific content-based probes. To address this, we proposed the design of an “associative drift engine”: a cognitive module capable of variable-bandwidth access, modulating the precision, noise, and scope of its memory probes. This system mirrors the DMN’s exploratory function and suggests a foundational mechanism for spontaneous recollection, creative association, and cognitive play—essential features of both natural and artificial minds.

Background Notes 1
1. Self and Mirror Recognition 1
2. ChatGPT’s assessment of the account mirror recognition 6
3. Default Mode Network 9
4. Toward an Associative Drift Engine 11
Summary of the discussion 15    

Background Notes

This first major section of this document consists of pages 74 to 84 from my 1978 dissertation, Cognitive Science and Literary Theory, Department of English, State University of New York at Buffalo. Yes, I was in the English Department and the dissertation uses examples from literature, the Oedipus story, the evolution of narrative form, and Shakespeare’s Sonnet 129. But I was also working closely with David Hays in the Linguistics Department. He was a first-generation researcher in machine translation, which transformed itself into computational linguistics in the mid-1960s.

During the period when I was in his research group – 1974 to 1978, when I finished my degree – we were sketching schemes for how to ground a symbolic cognitive system in the operations of a sensorimotor system organized as a stack of control systems in a scheme suggested by William Powers, Behavior: The Control of Perception (1973). Thus, when I talk about the sensorimotor hyperorders in the dissertation excerpt, I’m talking about a system modeled on Powers. By contrast, the systemic, episodic, and gnomonic hyperorders are symbolic systems, all directly linked to the sensorimotor system.

One thing else I want to emphasize is that, by this time, I had come to understand that the physical construction of the nervous system, both in its layout in the brain and its relationship with the external world, that structure carried information that did not have to be explicitly represented inside the system itself – I’ve used yellow highlighting to emphasize those sections. The account I offer of mirror recognition depends on this.

About this document

This document contains four things:

1. A passage from my 1978 dissertation in which I discuss mirror recognition,
2. ChatGPT’s (current) assessment of that passage,
3. A discussion of the Default Mode Network (DMN) in the brain, and
4. Some speculation from ChatGPT on how to construct, in effect, a DMN for an artificial associative memory, something it calls “an associative drift engine.”

Monday, June 16, 2025

Emergence of human-like object concept representations in multimodal LLMs

Du, C., Fu, K., Wen, B. et al. Human-like object concept representations emerge naturally in multimodal large language models. Nat Mach Intell (2025). https://doi.org/10.1038/s42256-025-01049-z

Abstract: Understanding how humans conceptualize and categorize natural objects offers critical insights into perception and cognition. With the advent of large language models (LLMs), a key question arises: can these models develop human-like object representations from linguistic and multimodal data? Here we combined behavioural and neuroimaging analyses to explore the relationship between object concept representations in LLMs and human cognition. We collected 4.7 million triplet judgements from LLMs and multimodal LLMs to derive low-dimensional embeddings that capture the similarity structure of 1,854 natural objects. The resulting 66-dimensional embeddings were stable, predictive and exhibited semantic clustering similar to human mental representations. Remarkably, the dimensions underlying these embeddings were interpretable, suggesting that LLMs and multimodal LLMs develop human-like conceptual representations of objects. Further analysis showed strong alignment between model embeddings and neural activity patterns in brain regions such as the extrastriate body area, parahippocampal place area, retrosplenial cortex and fusiform face area. This provides compelling evidence that the object representations in LLMs, although not identical to human ones, share fundamental similarities that reflect key aspects of human conceptual knowledge. Our findings advance the understanding of machine intelligence and inform the development of more human-like artificial cognitive systems.

Here's a preprint version at arXiv.

Friday, January 31, 2025

Football • {sports commentary} • [Media Notes 154]

I confess, football doesn’t interest me very much, never has. But I’m an America and I live in America. So I can’t escape it.

When I was in junior high school and high school I played in the marching band. That required me to attend every football game so we could provide half-time entertainment. We were so good, however, that I suspect some people came to the games more to hear us play than to see the game itself.

That’s the only time in my life that I ever watched football regularly. Of course, we played a bit of touch football in gym class, but that was it. I attended one football game in college. I was in the band. When one half of the band finished a tune eight bars ahead of the other half, that’s when I decided to blow this pop stand.

When I was in graduate school at SUNY Buffalo a roommate bequeathed me a small B&W portable. I watched a number of football games on it. This was during the O. J. Simpson years and I’d watch the Bills games to see him run. He was sensational. Football I didn’t care about, human excellence, that’s another matter.

After that, sure, every once in a while I’d catch a game. At least I assume I did. As I said, I’m living in America. Then, for some reason, a couple of weeks ago I decided to catch a play-off game on Netflix. Why? Why not? So I watched the Baltimore Ravens vs. the Pittsburgh Steelers. I’m from Pennsylvania, so that inclines me toward the Steelers. (The name “Franco Harris” sticks in my mind, so I must have watched some games when he was playing). I sent to school in Baltimore, which would tip me toward the Ravens, thought it was the Colts in Baltimore when I was there (Johnny Unitas as QB). Fact is, I could have cared less who won. Didn’t even watch the fourth quarter.

A week later it was the Buffalo Bills vs. Kansas City Chiefs. I made it the whole way through on that one. But I would hardly say I watched the game. Us, I did watch it, in fits and starts. But I also cruised the web doing this and that.

Which brings up a question: Let’s say the total elapsed time from the beginning of a game to the end is two to two-and-a-half hours. Only an hour of that is game time, which is interrupted for various reasons for varying amounts of time. During those interruptions we’re either getting some kind of commentary on the game, or we’re getting commercials. Add up the total time devoted to the game and commentary on the game. Add up the total time devoted to commercials during the broadcast. What’s the ratio between the two? My guess is that game time would be the larger number, but I’d guess the ratio is closer to 3/2 than to 2/1 in favor of game time.

So, I guess we could say that the commercials exist so that we can watch the game. But it could easily go the other way. Of course, if you’re a football fan, and so heavily invested in the game, that that’s certainly your priority. But if you’re not a fan, then it could almost go the other way.

Which brings me to the real reason for this note: the commentary. That fascinates me. I’m interested in it as a perceptual and cognitive activity. As I understand it, we generally have two commenters, one commenting on the action (play-by-play) and the other commenting on this and that. I believe the second is doing color commentary.

The play-by-play commenter is expected to comment on what’s happening as it happens. That requires them to have had a great deal of experience watching football games, more experience than I’ve had. You have to be able to instantly recognize hundreds of different patterns of activity and associate them with appropriate verbal comments. This is not a time for careful deductive reasoning. It’s an associative process. The commentary must be so fluid as to be of a piece with the perceptual act.

I wonder how long it takes to develop this capacity to the level we see in professional commentators? 10,000 hours? I don’t know. Let’s do a quick calculation. Ten thousand hours works out to something less than 5000 games, somewhere between 4000 and 4500. Let’s say it’s 100 games a year, two games a week. That’s forty to forty-five years. That’s possible, but I think the pros get in the game well before that. So it’s not 10,000 hours. It’s less than half that.

And then there’s the color commentary. That doesn’t have to track the action moment by moment, so it’s not constrained in that way. But still, it’s not an occasion for deductive reasoning. The commentator has to have access to a large range of relevant information about the players and the game, past and present, and come up with relevant bits and pieces in a matter of seconds. So it’s still pretty much an associative process.

And THAT, those last three paragraphs, that’s why I’m writing this not. As for the rest, why note? It’s context.

Tuesday, January 9, 2024

Musical notation, its varieties and its problems

This is a very interesting video on musical notation. If you aren't a musician, the topic probably seems, shall we say, empty. If you are a musician and you read music in the standard Western notation, then it's likely a topic on which you have some possibly quite strong opinions. Of course, you might be quite a skilled musician, but also unable to read music. In that case you might not care, or you may be ticked off, or something else.

I'm a musician and I read written music. But I can improvise as well and, on the whole, would prefer not to be bothered with notation. Alas, I can't avoid it. For one thing, I've often performed in situations where I had to read music. Beyond that, I've also written some music, and so had to struggle with the practical difficulty of figuring out just how to notate the music.

Beyond even that, I know quite about about cognition and perception, so I'm interested in the perceptual, cognitive, and neuro-psychology of the problem. This video gets into those issues when it discusses the practical problems presented by various notation conventions, but doesn't discuss them in the terminology of the relevant psychology disciplines. That's fine. I have no complaint on that score. But, as I listen to the video, I'm translating it into those terms, and that's what I find fascinating. One could write a book about that, and perhaps someone will do so one day.

From the YouTube page:

921,423 views Nov 3, 2023

Many people feel that western notation makes it unnecessarily hard to read music. If we want to sight read, learn music theory or just practice an instrument, surely there's a better way? Right? This has been a hot topic for almost 1000 years... AND I PUT IT TO BED RIGHT HERE!

00:00 - Setting the stage*
09:26 - Notation must die INTRO!
15:12 - Ancient Greek notation
19:40 - The history of western notation
31:57 - Chromatic staves
36:15 - The piano roll
42:54 - Clefs (and resistance to change)
44:42 - Muto method
46:45 - Notation & the aristocracy
48:25 - Tablature
51:40 - Guitar Hero
52:51 - Klavarskribo
54:08 - Other types of keyboard notation
55:47 - Musitude!
1:01:18 - Dodeca
1:03:01 - Accessibility
1:04:31 - Farbige Noten
1:06:49 - Jullian Carrillo's system
1:08:26 - The best of the rest
1:11:52 - Where can we go from here?

*Note: This opening section is about chess notation. It serves to introduce the problem of notation, though we aren't told that at the beginning. Unless chess notation actually interests you, you might want to skip over this.

Sunday, October 1, 2023

Internal feedback in the cortical perception-action loop enables fast and accurate behavior

Jing Shuang (Lisa) Lia, Anish A. Sarmaa, Terrence J. Sejnowskic, and John C. Doyle, Internal feedback in the cortical perception-action loop enables fast and accurate behavior, PNAS, September 22, 2023 120 (39) e2300445120 https://doi.org/10.1073/pnas.2300445120, asXiv: https://arxiv.org/abs/2211.05922

Significance

Internal feedback projections—signals flowing from motor areas or late sensory processing regions back to early sensory processing regions such as primary visual and auditory areas—are ubiquitous in the sensorimotor nervous system and are as or more numerous than feedforward projections. However, the function of internal feedback is poorly understood, particularly in the context of task performance. We leverage control theory and simple models to demonstrate that internal feedback facilitates good task performance when there are communication limitations such as internal time delays and speed–accuracy trade-offs, which motivate compensatory feedback signals to counter self-generated and predictable movements. Control theory explains why motor-related signals are found throughout the sensory cortex and why the motor cortex is dominated by internal dynamics.

Abstract

Animals move smoothly and reliably in unpredictable environments. Models of sensorimotor control, drawing on control theory, have assumed that sensory information from the environment leads to actions, which then act back on the environment, creating a single, unidirectional perception–action loop. However, the sensorimotor loop contains internal delays in sensory and motor pathways, which can lead to unstable control. We show here that these delays can be compensated by internal feedback signals that flow backward, from motor toward sensory areas. This internal feedback is ubiquitous in neural sensorimotor systems, and we show how internal feedback compensates internal delays. This is accomplished by filtering out self-generated and other predictable changes so that unpredicted, actionable information can be rapidly transmitted toward action by the fastest components, effectively compressing the sensory input to more efficiently use feedforward pathways: Tracts of fast, giant neurons necessarily convey less accurate signals than tracts with many smaller neurons, but they are crucial for fast and accurate behavior. We use a mathematically tractable control model to show that internal feedback has an indispensable role in achieving state estimation, localization of function (how different parts of the cortex control different parts of the body), and attention, all of which are crucial for effective sensorimotor control. This control model can explain anatomical, physiological, and behavioral observations, including motor signals in the visual cortex, heterogeneous kinetics of sensory receptors, and the presence of giant cells in the cortex of humans as well as internal feedback patterns and unexplained heterogeneity in neural systems.

Friday, September 15, 2023

A remark on how the geometry of photographs doesn't echo the eye's geometry

This photo exemplifies a phenomenon I've experienced many times. Look at that white statue, a bust that we're observing from the side, in the middle just to the left of center. It appears relatively small in relation to the size of the entire image. It appeared considerably larger than that to my eye when I took the photo, that is, larger in relation to the whole field of view. The eye/brain's optical system appears to magnify objects at the center of the field of view. I've observed this with countless photographs, and yet it still surprises me.

Sunday, July 9, 2023

The neural basis of color perception

Monday, July 25, 2022

Visual adaptation to an inverted visual field [time-course of neural change]

This article speaks to the issue raised in my recent post, Physical constraints on computing, process and memory, Part 1 [LeCun], under the somewhat strange notion of the brain has a hyperviscous mesh, which is about timescales of stability in patterns of connectivity, where it is understood that 'information' is registered in those patterns.

Timothy P. Lillicrap · Pablo Moreno‐Briseño · Rosalinda Diaz · Douglas B. Tweed · Nikolaus F. Troje · Juan Fernandez‐Ruiz, Adapting to inversion of the visual field: a new twist on an old problem, Experimental Brain Research 228(3), May 2013, https://doi.org/10.1007/s00221-013-3565-6

Abstract:

While sensorimotor adaptation to prisms that displace the visual field takes minutes, adapting to an inversion of the visual field takes weeks. In spite of a long history of the study, the basis of this profound difference remains poorly understood. Here, we describe the computational issue that underpins this phenomenon and presents experiments designed to explore the mechanisms involved. We show that displacements can be mastered without altering the updated rule used to adjust the motor commands. In contrast, inversions flip the sign of crucial variables called sensitivity derivatives-variables that capture how changes in motor commands affect task error and therefore require an update of the feedback learning rule itself. Models of sensorimotor learning that assume internal estimates of these variables are known and fixed predicted that when the sign of a sensitivity derivative is flipped, adaptations should become increasingly counterproductive. In contrast, models that relearn these derivatives predict that performance should initially worsen, but then improve smoothly and remain stable once the estimate of the new sensitivity derivative has been corrected. Here, we evaluated these predictions by looking at human performance on a set of pointing tasks with vision perturbed by displacing and inverting prisms. Our experimental data corroborate the classic observation that subjects reduce their motor errors under inverted vision. Subjects' accuracy initially worsened and then improved. However, improvement was jagged rather than smooth and performance remained unstable even after 8 days of continually inverted vision, suggesting that subjects improve via an unknown mechanism, perhaps a combination of cognitive and implicit strategies. These results offer a new perspective on classic work with inverted vision.

From the article's introduction:

Visuomotor adaptation to perturbations that displace the visual field, for example, from left to right, is widely studied and well characterized (Harris 1965; Kohler 1963; Kornheiser 1976; Redding and Wallace 1990). Pointing, throwing, and reaching tasks have been used to assess adaptation, and in these tasks, human subjects adapt quickly and smoothly to displacements, typically within minutes (Fernandez-Ruiz et al. 2006; Kitazawa et al. 1997; Martin et al. 1996; Redding et al. 2005; Redding and Wal- lace 1990). When prisms are worn, which displace targets and responses to the right and thus initially produce a left- ward error (Fig. 1b), subjects reduce their errors by correct- ing in the leftward direction on subsequent trials. Doing so, subjects make use of an implicit assumption about how error vectors ought to be used to update motor commands (Fig. 1a). the assumption, which holds in the case of dis- placed vision, is that the relationship between commands and errors (i.e., the sensitivity derivatives) has not been altered.

Comparatively, little is understood about visuomotor adaptation to inversions of the visual field—for example, a perturbation which flips the visual field from left to right about the midline (Fig. 1c). Studies have reported that, although subjects are initially severely impaired by inversions, they were eventually able to reacquire even complex sensorimotor skills, such as riding a bicycle (Harris 1965; Kohler 1963). However, most studies have been qualitative in nature (Rock 1966, 1973; Stratton 1896, 1897) or else have focused on perceptual rather than motor adaptations (Linden et al. 1999; Sekiyama et al. 2000). thus, the reason for the profound difference in the time course of visuomotor adaptation, the manner in which adaptation unfolds, and the mechanisms involved are not well studied.

Gradient vs. cognitive processes:

Superficially, our experimental data agree with the class of gradient-based models which update their feedback learning rule. However, closer examination of our results suggests that adaptation to inversions involves a complex mixture of implicit (i.e., gradient or reinforcement learning) and explicit or “cognitive” processes (e.g., Mazzoni and Krakauer 2006), which is not well modeled by the existing theory.

Tuesday, May 17, 2022

Neural Recognizers: Some [old] notes based on a TV tube metaphor [perceptual contact with the world]

Yet another bump can't hurt. Why? Because yesterday I saw this tweet by Kevin Mitchell: “A useful perspective shift is to think of a neuron (or brain area) as actively monitoring its inputs as opposed to being passively driven by them.” [5.17.22]

Another bump to the top can't hurt.  [Sept 2021]

I'm bumping this to the top of the queue because GPT-3. I'm reconfiguring and restructuring like crazy. More later.
Introduction: Raw Notes

A fair number of my posts here at New Savanna are edited from my personal intellectual notes. In this post the notes are unedited. This is an idea that dates back to my graduate school days in English at SUNY Buffalo. Since I keep my notes in Courier – a font that harks back to the days of manual typewriters – I’ve decided to retain that font for these posts and to drop justification.

Since these notes are “raw” you’re pretty much on your own. Sorry and good luck.


* * * * *

Diagrams from D'Arcy Thompson, On Growth and Form.



1.26.2002 – 1.27.2002

This is the latest version of an idea I first explored at Buffalo back in the late 1970s. It was jointly inspired by William Powers’ notion of a zero reference level at the top of his servo stack and by D’Arcy Thompson (see diagrams above). I’ve transcribed some of those notes into the next section. A version of this appeared in the paper DGH (David Hays) and I wrote on natural intelligence, where we talked in terms of Pribram’s neural holography and Spinelli’s OCCAM model for the cortical column:

  • W. Benzon and D. Hays. Principles and Development of Natural Intelligence. Journal of Social and Biological Structures 11, 293 - 322, 1988.
  • Powers, W.T. (1973). Behavior: The Control of Perception. Chicago: Aldine.
  • Pribram, K. H. (1971). Languages of the Brain. Englewood Cliffs, New Jersey: Prentice-Hall.
  • Spinelli, D. N. (1970). Occam, a content addressable memory model for the brain. In (K. H. Pribram & D. Broadbent, Eds): The Biology of Memory. New York: Academic Press, pp. 293-306.

“TV Tube Recognizer”

2.13.1979

Imagine a TV screen with a circle painted on it and with controls which allow you to operate on and manipulate the projection system in various useful ways. We’re going to use this to conduct an active analysis of the input to the screen.

Assume that the object to be analyzed is projected onto the screen in such a way that its largest dimension doesn’t extend beyond the circle painted on it. The analysis consists of twiddling the [control] dials until the area between the outer border of the object and the inner border of the circle is as small as possible. That is “minimize area between object and circle” is the reference signal for this servo-mechanical procedure, while “twiddle the dials” is the output function. Think of those knobs as operating on the coordinates in Thompson's illustrations above. (Notice that we are not operating on the input signal to the TV screen.)


TV-tube-active-analysis

Active Analysis

One thing we might do by dial twiddling is to operate on the coordinate system of the projection (I’m thinking here of D’Arcy Thompson’s grids whereby a bass on one coordinate grid becomes a flounder when projected onto another different grid.) Thus if the input is a vertical ellipse a horizontal stretch would lower the area between the ellipse and circle [painted on the TV screen]. One could bend the axes or distort them in various ways. Or, how about allowing the system to partition the screen in various ways and then make local alterations in the coordinate system within the partition.

Endless possibilities.

It doesn’t make much difference what [we do], the point is that the system have some way of operating on the image on the screen (without messing around with the input to the screen ...). The settings on the dials when the area between the projected object and the circle is at a minimum then constitutes the analysis of the object. To the extent that objects differ, the differences in the dial settings differentiate between objects (we are limited, of course, by the resolving power of the system). The settings which are best for a buzzard won’t be best for a flounder, nor a pine tree, nor a start, etc.


* * * * *


[1.26.2002]

The most obvious difficulty with this story is that it depends on someone observing the TV screen and twiddling the control knobs. We want to eliminate that someone so that the system can achieve the desired result itself.

The obvious way to do this is to call on the self-organizing capacity of cortical neural tissue. That tissue is itself the TV screen and control knobs while the appropriate subcortical thalamic nucleus is the source of input to the recognizer. The reference level is alpha oscillation, reflecting the observation that alpha energy is high when the stimulus is familiar and low when it is not. Unfamiliar input disturbs the oscillation and the recognizer seeks to restore oscillation by temporarily altering the properties (twiddling the dials) of the input array (thalamic nucleus).

Neural Recognizer

The neocortex is conceived as a patchwork of pattern recognizers; each is a sheet of cortical columns. Neighboring columns are mutually inhibitory, as in OCCAM (Spinelli 1970). A high level of output from one column will suppress output in its neighbors. The patterns are recognized in the primary input (input array) to a given recognizer. Let us assume a recognizer whose primary input is subcortical and let us set aside consideration of other inputs. The recognizer also generates primary output, which goes to the subcortical source of primary input. The computing capacity of a recognizer is far greater than that of its primary input.

The base state of such a recognizer occurs when the input is random (of a certain unspecified quality). In this base state the columns in the array oscillate – given a rather old notion that high alpha means low arousal, I’ve been thinking this would be at alpha; but, perhaps in view of Freeman’s work, I should revise this in favor of intrinsic chaos. The recognizer acts to maintain this base state under all conditions. When there is a non-random perceptual signal that signal will necessarily perturb the array so that it no longer oscillates smoothly. The array proceeds to form an impression of that input by sending (inhibitory) signals to the primary input. Some cortical columns will necessarily play a stronger role in this process than others. Eventually the recognizer will find some combination of outputs (modifying the properties of the input array) that restores randomness, and hence smooth oscillation. When this point is reached, the impression has been formed. This impression is of the input. In common parlance, we might want to say it represents that input.

Now some process must take place in the array so that the current perturbation can either be habituated into the background or an impression be taken, that is, can become part of the permanent repertoire of the recognizer. This latter, presumably, involves Hebbian learning and is triggered by reinforcement. In the manner of Spinelli’s OCCAM, the recognizer has many such impressions stored in its synaptic weights. A perceptual signal is presented across the entire array and, if it is of a kind that has already made an impression on the array, that impression will be evoked from the array and restore the recognizer to periodic oscillation. If it is of a kind that has not yet made an impression, then a new impression must be made.

Now, in fact, each recognizer has a variety of secondary inputs coming from other recognizers and it generates secondary outputs to them. All of them are attempting to account for their input simultaneously; through their secondary inputs and outputs the recognizers “help” one another out. Further it has inputs from subcortical nuclei which send neuromodulators to the array and it sends outputs to those nuclei which indicate its state of operation. The neuromodulators cause the recognizer to switch between its different operating modes.

I see these operating modes as follows:

Baseline: There is no perceptual load. The array is oscillating at alpha (chaos?).

Tracking: Perceptual input is accounted for. The array has recognized the input and is oscillating at alpha (chaos?).

Matching: The array is under a perceptual load and is attempting to match the input using its current set of impressions. EEG: “desynchronized,” gamma?

Forming (an impression): The array is under a perceptual load, but is unable to match it from its current impression repertoire. It is now forming a new impression. Obviously one critical aspect of the recognizer’s operation is switching from an unsuccessful matching operation to forming. EEG: “desynchronized,” gamma?

Habituating: The array is under a perceptual load, a new impression has been formed, and it has been assimilated into the background.

Fixing: A new impression has been formed. It must now become part of the permanent repertoire of impressions. This is the beginning of LTP. EEG: high alpha?

Group Expressive Behavior

We could apply this line of thought to group expressive behavior where the members of the group are regarded as oscillators coupled to one another through mutual perception and coordinated action. The simplest such behavior would be moving together, or clapping, to an isochronous pulse.

Assume a group moving to an isochronous pulse. Further assume that this activity is cortically controlled. Now, imagine that various members of the group are driven by subcortical impulses to inflect their movement in noticeable ways. These inflections will be transmitted to others through the coupling. Adjustments made to accommodate these inflections become, in effect, the group’s impression of those subcortical impulses.

This needs to be worked through rather more carefully, which will certainly change things a bit. But what I’m driving at is that these group impressions will become the stuff of culture. Here’s where we get memes and performance trajectories [as those are defined in Beethoven’s Anvil].

Thursday, January 27, 2022

Update on Color Terms: Nature or Nuture?

More on color terms [1.27.22]:
 
 
* * * * *
Thinking about this again so I'm bumping this to the top. It's from December 2012.

Color vision, from genetics through neuropsychology to color terms, is one of the most intensely investigated topics in cognitive science. The subject is interesting for two reasons:
  1. as psychological phenomena go, it's relatively simple, and accessible to investigation through a variety of methods, and
  2. in particular, it can be studied cross-culturally and thus shed light on the nature-nurture question.
Mark Changizi has recently argued the color vision has evolved for the purpose of allowing us to obtain clues about a person's health and state of mind from changes in skin color, blushing, bruising, and the like. He's stated this idea at some length in the first chapter of The Vision Revolution (BenBella Books, 2009, pp. 5-48). Folks have been discussing this idea at some length over at Crooked Timber.

In thinking through that discussion I formulated this question and sent it to Changizi: Question: Lots of languages have rather impoverished systems of color terms. Would folks speaking a language that lacked a term for green thereby have more fine-grained perception of greens? He didn't have an answer but indicated that people are working on that kind of issue. He sent me reprints of two papers that are indeed relevant. 

Color terminology seems subject to constraints that are universal but, at the same time, differences in color naming across cultures does seem to cause differences in color perception. How do you like them apples? Both universal and different at the same time.

* * * * *

Paul Kay and Terry Regier. Language, thought and color: recent developments. TRENDS in Cognitive Sciences Vol.10 No.2 February 2006, pp. 51-54

Here's how Kay and Regier state matters as they existed, say, a quarter of a century ago:
Color naming varies across languages; however, it has long been held that this variation is constrained. Berlin and Kay [1] found that color categories in 20 languages were organized around universal ‘focal colors’ – those colors corresponding principally to the best examples of English ‘black’, ‘white’, ‘red’, ‘yellow’, ‘green’ and ‘blue’. Moreover, a classic set of studies by Eleanor Rosch found that these focal colors were also remembered more accurately than other colors, across speakers of languages with different color naming systems (e.g. [2]). Focal colors seemed to constitute a universal cognitive basis for both color language and color memory.
Research conducted in the last decade or so has called those conclusions into question. Kay and Regier present and discuss this work and offer this summary:
The debate over color naming and cognition can be clarified by discarding the traditional ‘universals versus relativity’ framing, which collapses important distinctions. There are universal constraints on color naming, but at the same time, differences in color naming across languages cause differences in color cognition and/or perception. The source of the universal constraints is not firmly established. However, it appears that it can be said that nature proposes and nurture disposes. Finally, ‘categorical perception’ of color might well be perception sensu stricto, but the jury is still out.
The key proposition is that "differences in color naming across languages cause differences in color cognition and/or perception."

* * * * *

Paul Kay and Terry Regier, Resolving the question of color naming universals, PNAS, vol. 100, no. 15, July 22, 2003, pp. 9085-9089.

Abstract:
The existence of cross-linguistic universals in color naming is currently contested. Early empirical studies, based principally on languages of industrialized societies, suggested that all languages may draw on a universally shared repertoire of color categories. Recent work, in contrast, based on languages from nonindustrialized societies, has suggested that color categories may not be universal. No comprehensive objective tests have yet been conducted to resolve this issue. We conduct such tests on color naming data from languages of both industrialized and nonindustrialized societies and show that strong universal tendencies in color naming exist across both sorts of language.
From the methodology discussion:
The central empirical focus of our study was the color naming data of the Word Color Survey (WCS). The WCS was undertaken in response to the above-mentioned shortcomings of the BK [Berlin and Kay] data (1): it has collected color naming data in situ from 110 unwritten languages spoken in small-scale, nonindustrialized societies, from an average of 24 native speakers per language (mode: 25 speakers), insofar as possible monolinguals. Speakers were asked to name each of 330 color chips produced by the Munsell Color Company (New Windsor, NY), representing 40 gradations of hue at eight levels of value (lightness) and maximal available chroma (saturation), plus 10 neutral (black-gray-white) chips at 10 levels of value. Chips were presented in a fixed random order for naming. The array of all color chips is shown in Fig. 1. (The actual stimulus colors may not be faithfully represented there.) In addition, each speaker was asked to indicate the best example(s) of each of his or her basic color terms. The original BK study used a color array that was nearly identical to this, except that it lacked the lightest neutral chip. The languages investigated in the WCS and BK are listed in Tables 1 and 2.
The concluding paragraph:
The application of statistical tests to the color naming data of the WCS has established three points: (i) there are clear cross-linguistic statistical tendencies for named color categories to cluster at certain privileged points in perceptual color space; (ii) these privileged points are similar for the unwritten languages of nonindustrialized communities and the written languages of industrialized societies; and (iii) these privileged points tend to lie near, although not always at, those colors named red, yellow, green, blue, purple, brown, orange, pink, black, white, and gray in English.

Thursday, August 19, 2021

Rhythm and the perception of syntax

Tuesday, June 1, 2021

Some thoughts on why systems like GPT-3 will always have trouble with common sense knowledge

I develop an analogical argument about why natural language systems trained only on text will never be able to deal with common-sense reasoning. I begin by presenting Herbert Simon’s famous parable of the ant and follow it with some information about sensory deprivation. From those I conclude that our mental apparatus depends on access to the world to achieve stability. I then tap-dance my way to the assertion that common sense reasoning depends on those sensory motor systems which are, in turn, dependent on the world.

AI language engines are enmeshed in language, with no access to the physical world. Consequently common sense reasoning will forever be elusive. Common sense grounds us in the physical world.

Simon’s ant

In Chapter 3, “The Psychology of Thinking: Embedding Artifice in Nature,” of The Sciences of the Artificial (2nd Ed., 1981), Herbert Simon gives us a parable, a story to think with. Simon asks us to imagine an ant moving about on a beach:

We watch an ant make his laborious way across a wind- and wave-molded beach. he moves ahead, angles to the right to ease his climb up a steep dunelet, detours around a pebble, stops for a moment to exchange information with a compatriot. Thus he makes his weaving, halting way back to his home. So as not to anthropomorphize about his purposes, I sketch the path on a piece of paper. It is a sequence of irregular, angular segments – not quite a random walk, for it has an underlying sense of direction, of aiming toward a goal.

After introducing a friend, to whom he shows the sketch and to whom he addresses a series of unanswered questions about the sketched path, Simon goes on to observe:

Viewed as a geometric figure, the ant’s path is irregular, complex, hard to describe. But its complexity is really a complexity in the surface of the beach, not a complexity in the ant. On that same beach another small creature with a home at the same place as the ant might well follow a very similar path.

That is, because the beach has a complex surface, the ant is able to walk a complex path on that surface using rather simple mechanisms. In posing this parable Simon is, of course, asking us to think of the beach as the world in full and that ant is us. Relative to the world’s complexity, our conceptual apparatus is relatively simple.

I would like to propose that the nervous system requires environmental support if it is to maintain its physical stability and coherence. Note that Simon was not at all interested in the physical requirements of the nervous system. Rather, he was interested in suggesting that we can get complex behavior from relatively simple devices, and simplicity translates into design requirements for a nervous system. That’s fine, but I’m suggesting that the nervous system actively seeks out the world and so is dependent upon finding it, in a more or less orderly fashion.

Our sensory systems don’t ‘represent’ (if that’s the right word, many reject it) the world in great detail. Their apprehension of the world is rough and ready. One doesn't need to represent apples and oranges in full detail in order to distinguish them, nor cats and dogs, cars and bicycles, and so forth. Our systems need only ‘grab on’ to the things and events in the world. The world itself will ‘fill out’ our perceptions in real time.

Now, consider this variation on Simon’s story. What would happen if we put the ant on an absolutely featureless surface and let it walk about? What kind of paths would it trace then? As that surface lacks any of the normal cues in the ant’s environment I would imagine the ant would either not move at all or move in a genuinely random or perhaps a rigidly stereotypic way (e.g. around and around in a circle). Or perhaps the ant would hallucinate.

Sensory deprivation

That is what seems to happen to humans when we are deprived of sensory input. Early on in The Ghost Dance, a classic anthropological study of the origins of religion, Weston La Barre considers what happens under various conditions of deprivation. Consider this passage about Captain Joshua Slocum, who sailed around the world alone at the turn of the 20th Century:

Once in a South Atlantic gale, he double-reefed his mainsail and left a whole jib instead of laying-to, then set the vessel on course and went below, because of a severe illness. Looking out, he suddenly saw a tall bearded man, he thought at first a pirate, take over the wheel. this man gently refused Slocum’s request to take down the sails and instead reassured the sick man he would pilot the boat safely through the storm. Next day Slocum found his boat ninety-three miles further along on a true course. That night the same red-capped and bearded man, who said he was the pilot of Columbus’ Pinta, came again in a dream and told Slocum he would reappear whenever needed.

La Barre goes on to cite similar experiences happening to other explorers and to people living in isolation, whether by choice, as in the case of religious meditation, or force, as in the case of prisoners being brainwashed.

In the early 1950s Woodburn Heron, a psychologist in the laboratory of Donald Hebb, conducted some of the earliest research on the effects of sensorimotor deprivation [2]. The subjects were placed on a bed in a small cubicle. They wore translucent goggles that transmitted light, but no visual patterns. Sound was masked by the pillow on which they rested their heads and by the continuous hum of air-conditioning equipment. Their arms and hands were covered with cardboard cuffs and long cotton gloves to blunt tactile perception. They stayed in the cubicle as long as they could, 24 hours a day, with brief breaks for eating and going to the bathroom.

The results were simple and dramatic. Mental functioning as measured by simple tests administered after 12, 24, and 48 hours or isolation deteriorated. Subjects lost their ability to concentrate and to think coherently. Most dramatically, subjects began hallucinating. They would begin with simple forms and designs and evolve into whole scenes. One subject saw dogs, another saw eyeglasses, and they had little control over what they saw; no matter how hard they tried, they couldn’t change what they were seeing. A few subjects had auditory and tactile hallucinations. Upon emerging from isolation the visual world appeared distorted with some subjects reporting that the room appeared to be moving. Woodburn concluded, as have other investigators, that the waking brain requires a constant flux of sensory input in order to function properly.

Of course, one might object to this conclusion by pointing out that, in particular, these people were deprived interaction with other people and that is what causes the instability, not mere sensory deprivation. But, from our point of view, that is no objection at all. For other people are a major part of the environment in which human beings live. The rhythms of our intentional structures are stable only if they are supported by the rhythms of the external world. Similarly, one might object that, while these people were cut off from the external physical world, their brains, of course, were still operating in the interior milieu. Consequently the instabilities they experienced reflect “pressure” from the interior milieu that is not balanced by activity in the external world. This may well be true, I suspect that it is, but it is no objection to the idea that the waking brain requires constant input from the external world in order to remain stable. Rather, this is simply another aspect of that requirement.

Thus I suggest that detaching one’s attention from the immediate world to “think” may cause problems. And yet it is the capacity for such thought that is one aspect of the mental agility that distinguishes us from our more primitive ancestors. How do we keep the nervous system stable enough to think coherently? The answer to that question depends, of course, on just what is causing the instability. Part of the answer may well be that we periodically “tune” our cortical circuits through music and dance. As long as the brain gets such tuning on a regular basis it can maintain its stability well enough during episodes of extending thinking--whether the merest day dreaming, or concentrated intellectual activity of one sort or another. But, without regular tuning, the brain begins to lose its stability.

What does that have to do with GPT-3?

GPT-3, and many other contemporary systems, is trained on large bodies of text. That text, of course, consists of words. Well, they are words to us, who can say them, spell them, offer definitions, and use them correctly in utterances and writing. To the computer those are merely word forms, symbols that are not attached to meanings in the way that word forms are attached to meanings in the human mind. The object of these systems is to somehow approximate word meanings by calculating over the distribution of words in texts. The underlying assumption, which Warren Weaver articulated in his famous memo on machine translation back in 1949 [1], is that words that appear together share some aspect of meaning. Thus if we can ‘examine’ words in a sufficiently large number of contexts each, we should be able to approximate their meanings.

And indeed, GPT-3’s ability to generate coherent text of some length suggests it has manages a remarkable approximation. But it exhibits failings in commonsense reasoning [3], a failing that the symbolic systems of four decades ago exhibited as well. I believe that much commonsense reasoning takes place ‘close to the ground’ as it were [4]. Because GPT-3 only has access to word forms, not to the sensory-motor schemas that directly support, give meaning to, many of them it lacks the basis on which common sense reasoning functions. It is, in effect, lost.

It is in the situation Slocum found himself when isolated at sea. Lacking people to talk to, his mind began unraveling. And so it happens with people in sensory deprivation. Without sensory input to stabilize their perceptual system they begin hallucinating.

Let’s return to Simon’s ant. There is the world, and there is the ant with its mental apparatus. The human situation is more complex. There is the world, and there is our sensorimotor apparatus. And for the first two years of life, that’s pretty much it. But then language begins to develop, and language depends on both those sensorimotor systems and on interaction with others. GPT-3 lacks both sensorimotor apparatus and conversation with others. That is, during training GPT-3 has no contact with human interlocutors. Once trained GPT-3 is given prompts to which it responds. But, it is my understanding that it does not learn from those interactions.

How can GPT-3, and similar engines, possibly make sense of language that functions ‘close to the world’? I suppose one can hope that by considering a very very large number of texts an AI Engine can somehow ‘fill in’ the information it is missing because it lacks direct access to the world. That has not worked so far. What reason do we have think that some day it will?

Consider Simon’s ant once again. By examining the paths it traces we can approximate the beach’s micro-geography. How many paths must we examine and superimpose in order fix the location of every pebble and dunelet – forget about grains of sand?

Now, until AI language systems and rich and flexible access to the world and the capacity to develop a rich analog or quasi-analog account of the world, until that happens, common sense reasoning will remain elusive.

References

[1] Warren Weaver, “Translation”, Carlsbad, NM, July 15, 1949, 12. pp. Online: http://www.mt-archive.info/Weaver-1949.pdf.

[2] Heron, W. (1957). The pathology of boredom. Scientific American, 196, 52–56. https://doi.org/10.1038/scientificamerican0157-52.

[3] This has been much discussed in the literature. I have offered a modest example in a post where I had GPT-3 explain the punch line to a Jerry Seinfeld joke: Analyze This! Screaming on the flat part of the roller coaster ride [Does GPT-3 get the joke?], May 7, 2021, https://new-savanna.blogspot.com/2021/05/analyze-this-screaming-on-flat-part-of.html.

[4] I have argued this in a working paper, GPT-3: Waterloo or Rubicon? Here be Dragons, Version 2, August 20, 2020, 34 pp., https://www.academia.edu/43787279/GPT_3_Waterloo_or_Rubicon_Here_be_Dragons_Version_2. See pp. 21-25. See also my post, Computation, Mind, and the World [bounding AI], September 28, 2019, https://new-savanna.blogspot.com/2019/12/computation-mind-and-world-bounding-ai.html.