Showing posts with label Attractor-net. Show all posts
Showing posts with label Attractor-net. Show all posts

Thursday, December 7, 2023

UPDATE AT 76: Reflections on entering my eighth decade and why it portends to be the most productive one of my life

I posted this six years ago, a few days after my 70th birthday. It’s still a reasonable overview of things I’ve accomplished, though the future looks a bit different. GPT-3 and then ChatGPT came along unexpectedly. I’ve devoted a great deal of time on ChatGPT over the last year, over 100 blog posts and 11 working papers. I’d hoped to finish a 12th paper by today, but other things intervened so I’ve still got a day’s or so worth of work on it. Then I’ve got to look over that work, arrive at some conclusions, and set directions for future work. That’ll take several more days.

One of those working papers, Xanadu, GPT, and Beyond: An adventure of the mind, outlines the connections between this strand of current work and my earlier work going back to “Kubla Khan” and even a bit before. Back in 2022 I posted Relational Nets Over Attractors, A Primer: Part 1, Design for a Mind, which knits together the work I did with David Hays, first on cognitive networks in the 1970s, and then our two papers on the brain in 1988, Principles and Development of Natural Intelligence and Metaphor, Recognition, and Neural Process. That was unexpected.

It also appears that I’ve undertaken yet another overview of literary criticism, GOAT Literary Critics. I’ve posted two pieces already and have tentative plans for five more. I don’t know just when I’ll do them as the ChatGPT wrap-up has higher priority.

Finally, back in 2017 I wrote about how I made Jersey City my home: through photographs. Now Hoboken has become my home and, once more, it’s been through photographs. I post photos to a Facebook group for Hoboken photos and am planning a small exhibit of some of my photos. I’m working with one of my Jersey City friends, Greg Edgell, on that (& Greg’s gotten married and moved to the house near Morristown where he grew up).

Things are moving along. Like my friend Al used to say, “These are the good old days. The best days are yet to come.” 

* * * * *

Reflections on entering my eighth decade and why it portends to be the most productive one of my life


KK in Arches 70
In destinies sad or merry, true men can but try.
– Sir Gawain and the Green Knight 
In scientific prognostication we have a condition analogous to a fact of archery—the farther back you draw your longbow, the farther ahead you can shoot.
– Buckminster Fuller
Birthdays are generally just that, even “major” birthdays, like my most recent one, my 70th. They are an occasion for a celebration, perhaps a modest one, perhaps not quite so modest – our house was crowded with dinner guests for my father’s 50th birthday – perhaps even extravagant. I’ve never been to an extravagant birthday party. Birthdays may also be a time for reflection, but by no means necessarily so.

But in my experience birthdays rarely correspond to major life events. What’s a major life event? Getting married, whether at a small civil ceremony before a judge or an elaborate wedding with 100s of guests into the church and out to the reception where a great band – like me and my colleagues in The Out of Control Rhythm and Blues Band – performs for hours of dancing, that’s a major life event. Climbing a mountain you’ve trained for over a period of years, that’s a different kind of major life event. Graduating from school, or completing basic training in the military, passing the bar exam, all major.

Years ago, in my early 20s, I was in a rock band called “St. Matthew Passion.” It was our last gig, the sax player and I were jamming a whacked out intro to “She’s Not There” and suddenly it all disappeared, me, the musicians, the room, the world, all into a brilliant, but soft, white light. Only the light and the music. It lasted what, half a second, a second, two seconds? Whatever. Those few moments challenged me for years, changing my sense of myself and the world.

Major life events come in all forms and durations, but they rarely coincide with a birthday. Birthdays simply mark the passage of time.

And so it set out to be on Thursday, December 7, 2017, when I turned 70. I woke up, cruised the web, made four posts to New Savanna, had breakfast and then, and then I decided to go out and take some photos, including some of that green platform pump I’ve been having so much fun photographing.

20171207-_IGP1519

That was a bit unusual because I generally write in the morning, and perhaps I was motivated by my birthday to do something a bit different. But that’s all it was, a change in routine. It’s no big deal; I do it all the time.

But then I realized, sometime in the afternoon, that this birthday IS a big deal, and that I really can make it a big deal. How? By finishing my working paper, Calculating meaning in “Kubla Khan” – a rough cut. Why is that important, major milestone important? Because I’ve been working on it almost 50 years, all my adult life.

Calculating Kubla Khan 3
Teaser: Calculating meaning in “Kubla Khan”, 2017
This image likely makes little sense. Don’t worry about it. There are two more images that won’t make much sense. Don’t worry about them either. Just look at them as you would displays in a museum or gallery and move on.
Not that paper, no, not that. It’s the poem, “Kubla Khan”, by Samuel Taylor Coleridge, that I’ve been tracking all my adult life. I’ve been working on it since the spring of 1969 when I read it in Earl Wasserman’s class in my senior year at The Johns Hopkins University in Baltimore – where my father had gone to school. To have completed a project that framed one’s adult life, that is indeed a major event. Not completed, not in the sense that it’s all over and done with – for it isn’t, but in some deep and fundamental sense, things have changed. I’ve got a new understanding, and new obligations to go with it.

The first five lines:
In Xanadu did Kubla Khan
A stately pleasure-dome decree:
Where Alph, the sacred river, ran
Through caverns measureless to man
          Down to a sunless sea.
There are 49 more.

Let me explain. Perhaps then you will understand why I expect the next decade to be the most productive one of my life. And not just my intellectual life. There is the Bergen Arches Project as well. And who knows what else? I wonder if Rita Moreno is available for salsa lessons?

Monday, August 16, 2021

Ramble – Spacecraft, social nature of truth, Green Villain, the architecture of mind

Once again, I’m feeling a need to ramble around in my mind, see what’s going on.

Spacecraft

I’m still thinking about Tim Morton’s Spacecraft. I figure I’ve got two, perhaps three, more posts on it and then perhaps a formal review. One post will look at how he talks of spacecraft, and his basic distinction between spacecraft and spaceship. Even before that, though, I think I need to talk about the distinction between real spacecraft and fictional ones. By real spacecraft I mean the 100s if not 1000s of satellites we’ve launched since Sputnik back in 1957, the various probes and landers, manned orbital craft and, of course, the Apollo missions, etc. There’s been a lot of them.

But while Morton mentions one or three of them here and there, they aren’t of much interest to him. He’s only interested in fictional space craft, with the Millennium Falcon being his primary example. And I pretty much assumed that’s what he’d be talking about when I asked him to have a copy of the book sent to me. It’s a bit as though someone had chosen two write a book about unicorns or dragons and did so in much the same way as one would write a book about sheep or crows. Of course Morton knows these various spacecraft are fictional, as is hyperspace, which I’ve already written about. But he sees no need to remark on that fact, much less to examine it.

That in itself is interesting.

Then I want to write one about the future. For that’s what was on my mind when I asked Tim for a copy: I wonder if he’ll say something about the future? And he does, page 73. Not sure whether I can get a whole post out of that, we’ll see. But the future sort of goes along with science fiction, no? Nor, come to think of it, does he talk much, if any, about science fiction as a kind of fiction, but that’s what he’s writing about, no? That too is interesting, especially since he’s trained as a literary critic.

That is to say, what’s most interesting about Spacecraft is the framing that’s not there.

The social construction of truth

I was going to write a 3QD column about the social construction of truth, but bailed on it. It would have been an expansion and refinement of a recent blog post, What’s the difference between a conspiracy theory and a far-out, but reasonable, idea? When does signaling go haywire? (July 21). I’m not quite sure why I bailed on it. Though I will note that, for one thing, it’s been shifting beneath my feet.

Perhaps I’m still figuring out an approach to truth. One thing that’s clear in a general sort of way, lots of people in what I will call my society believe lots of different things and this collectivity of beliefs is by no means coherent, mutually consistent, and stable. Do I need an approach to that multiplicity, a way to move it into the foreground instead of keeping it in the background? I’m inclined to think of that as the nature of belief in society and truth is something we carve out of it.

OK, this is going somewhere. The deep problem with asserting the social nature of truth is the implied split between nature and culture (and this goes back to Morton and his OOO). Given that split, society (that is to say, us) is viewed as a source of distortion and error. Truth is something outside us, it’s stable. What is in fact going on, however, is that, once we reach substantial agreement, we project that agreement into nature and fool ourselves into thinking it has nothing to do with us.

And in a sense it doesn’t, because it is about the world. But that agreement is something we’ve created. Science, in this view, is a collection of techniques for framing propositions about the world such that observations can be made that compel agreement from a given scientific community. Which brings us to the case of fundamental physics, where we’ve had half a century of theorizing without observational confirmation. Where’s that going?

What if scientists in other fields decided that they didn’t need to worry about observational confirmation either. That would surely make things easier for psychologists with their replication crisis. Just toss out observation all together. The crisis disappears.

Do these physicists know that their recalcitrance is threatening the very nature of science?

So, what’s at state in my other examples: 1) GOP Trumpism and conspiracy theories (e.g. QAnon), 2) memetics, and 3) the tech singularity? The first is mostly about socio-political reality where current truth is determined by journalistic means. The second was a proposal about science that just fizzled out. And the third is about things that will happen in the future.

I’m thinking that doing all that much stretch the limits of a 3QD column. It’s the framework implied in the previous paragraph that’s tricky. That is, almost all thinking about epistemology takes place within a Cartesian framework, where mind and body are split and truth is a problematic negotiation between a mind and the world. Drop that framework and truth becomes something different, a matter of reading agreement.

Green Villain

Instead of writing about the social construction of truth I choose to do a photo essay about the Green Villain loft space 51 Pacific. That was fun, and didn’t require much thought, just arranging the photos and providing a bit of introductory prose. The trick was to choose the photos, as I’ve got 100s to choose from. The idea was to sample the space.

Here’s another one, of Serringe in process:

In going through the photos, it occurred to me that if this space existed today (the building was demolished in 2015) it would be possible to charge people, say, $5 to enter and tour the building. That wouldn’t have been possible when the space was alive and kicking, back in the early 2010s. Things have changed.

The architecture of mind

And then there’s this recent post: Attractor nets, from basic vertebrates to humans, August 13. That belongs to my general attractor nets project, where the goal is to produce a primer on attractor nets. Does it make sense to think in terms of an explicit construction for each of the five principles listed in the brain paper (Principles and Development of Natural Intelligence)? The basic goal of the attractor net project is to work out constructions for the fifth principle (indexing). That earlier post mentioned the first principle (feeling, on-blocks). What about the other three? I’m thinking I need to leave them alone for awhile.

Seinfeld working paper

Yes, I need to get one of those done, one based on my explication of various bits.

Friday, August 13, 2021

Attractor nets, from basic vertebrates to humans

I recently had a post where I tacked this on at the end [1]:

That’s what I had in mind when, over two decades ago, I played around with a notation system I called attractor nets [2]. There I was imagining a logical structure implemented over a large and various attractor landscape. The logical structure is represented by a network formalism developed by Sydney Lamb [3]. The network is quite different from those generally used in classical symbolic systems, where nodes represented objects of various kinds and arcs represented types of relationships among those objects. In Lamb’s nets the nodes are logical operators (OR, AND) while the arcs carry the content of the net. In an attractor net each arc corresponds to a basin of attraction. The net then represents a logical structure over basins of attraction.

For some reason I’d never expressed my intention in just that way, though there’s nothing new there.

And so I began thinking: What’s the simplest creature that requires a logical structure over the attractor landscape? I considered the possibility that language is what required an attractor net. That would imply that all animals from worms to great apes have an unadorned attractor landscape. I rejected that.

Instead I’ve decided that coordination between the sensory systems and the motor system is what necessitated an attractor net structure. In “Principles and Structure of Natural Intelligence”[4] David Hays and I made the following remarks about the basic vertebrate nervous system:

The primitive vertebrate nervous system is reticular (Best, 1972; Bowsher, 1973). The only principle active at this level is the modal principle. Within a given mode, behavior is governed by on-blocks, where the conditional elements are innate releasing mechanisms and the executed programs are fixed action patterns (Lorenz, 1969). These on-blocks are executed as they are triggered by the interaction of environmental stimuli and organismic modal shifts. There is little or no autonomous chaining of these on-blocks.

Those on-blocks would require the structure provided by an attractor net. Beyond that consider a creature navigating its way through the environment. Such paths are ‘littered’ with contingencies, and contingencies would likely require an attractor net to weave perception and action together.

Thus language is only the most sophisticated behavior requiring attractor nets. With language the attractor net weaves the language system into cognition thereby allowing the system to take ‘arbitrary’ walks through cognition. Those walks are what we call ‘thinking,’ in the common sense use of the term.

References

[1] Dual-system mentation in humans and machines [updated], New Savanna, August 11, 2021, https://new-savanna.blogspot.com/2021/08/dual-system-mentation-in-humans-and.html.

[2] I never produced an account that I thought was ready for others to read. But, in addition to piles of notes, I did produce two documents intended to summarize the work for my own purposes. The first of these two documents does explain Lamb’s notation while the second document is a collection of diagrams, that is, constructions.

William Benzon, Attractor Nets, Series I: Notes Toward a New Theory of Mind, Logic, and Dynamics in Relational Networks, Working Paper, 52 pp., https://www.academia.edu/9012847/Attractor_Nets_Series_I_Notes_Toward_a_New_Theory_of_Mind_Logic_and_Dynamics_in_Relational_Networks.

William Benzon, Attractor Nets 2011: Diagrams for a New Theory of Mind, Working Paper, 55 pp., https://www.academia.edu/9012810/Attractor_Nets_2011_Diagrams_for_a_New_Theory_of_Mind.

[3] Sydney Lamb, the computational linguist, believed this as well. He argues it in Pathways of the Brain, Amsterdam: John Benjamins (1998), pp. 181-182.

[4] William Benzon and David Hays, Principles and Development of Natural Intelligence, Journal of Social and Biological Structures, Vol. 11, No. 8, July 1988, 293-322, https://www.academia.edu/235116/Principles_and_Development_of_Natural_Intelligence.

Wednesday, August 11, 2021

Dual-system mentation in humans and machines [updated]

I’m interested in an article recently posted to arXiv:

Maxwell Nye, Michael Henry Tessler, Joshua B. Tenenbaum, Brenden M. Lake, Improving Coherence and Consistency in Neural Sequence Models with Dual-System, Neuro-Symbolic Reasoning, 6 July 2021, arXiv:2107.02794 [cs.AI]

We need to consider it in the context of current debates over whether or not machine learning approaches are full adequate for the creation of ‘intelligent’ systems – whatever they are – of whether we need ‘classical’ symbolic systems as well [1].

But first I want to make some remarks about language and its relation to cognition. Then we can take a look at the dual-system model.

Language and cognition

Language and cognition are often considered in conjunction with one another. Without cognition language is an empty formal system. While cognition can and is studied independently of language – animals, after all, have considerable cognitive capabilities while lacking language, though they have more limited means of communication – the particular range and power of human cognition is often, and I believe properly, attributed to the ways language allows us to “grab hold of” and “extend” our cognitive abilities.

We should note that language is itself a ‘full-depth’ system. Thus it has components in direct interaction with the external world, hearing and speaking, seeing and writing, and more abstract components, syntax and discourse. This system has a rich internal structure independent of its linkage to general cognition.

With that in mind, consider this passage from an old paper [2]:

Nonetheless, the linguist Wallace Chafe has quite a bit to say about what he calls an intonation unit, and that seems germane to any consideration of the poetic line. In Discourse, Consciousness, and Time Chafe asserts that the intonation unit is “a unit of mental and linguistic processing” (Chafe 1994, pp. 55 ff. 290 ff.). He begins developing the notion by discussing breathing and speech (p. 57): “Anyone who listens objectively to speech will quickly notice that it is not produced in a continuous, uninterrupted flow but in spurts. This quality of language is, among other things, a biological necessity.” He goes on to observe that “this physiological requirement operates in happy synchrony with some basic functional segmentations of discourse,” namely “that each intonation unit verbalizes the information active in the speaker’s mind at its onset” (p. 63).

While it is not obvious to me just what Chafe means here, I offer a crude analogy to indicate what I understand to be the case. Speaking is a bit like fishing; you toss the line in expectation of catching a fish. But you do not really know what you will hook. Sometimes you get a fish, but you may also get nothing, or an old rubber boot. In this analogy, syntax is like tossing the line while semantics is reeling in the fish, or the boot. The syntactic toss is made with respect to your current position in the discourse (i.e. the current state of the system). You are seeking a certain kind of meaning in relation to where you are now.

The important point is that, contrary to what you might think, you don’t necessarily know what you’re thinking until you have verbalized your thoughts. Thus conversation is often filled with disfluencies where you start and stop, hunt for the right word or phrase, and then continue. Writing can be like that as well, especially when dealing with new and challenging material.

Let’s amplify that point by considering the general outline of language acquisition offered by the great Russian developmental psychologist, Lev Semyonovich Vygotsky [3]. Vygotsky argued, in effect, that thinking – considered as a quasi-verbal voice in the mind – is just internalized speech [4]. The young child learns to speak to herself as others speak to her and thereby gains a mechanism affording some control over her mind. Think of the process Vygotsky describes in terms of the scaffolding metaphor that’s become popular. The adult’s speech scaffold’s the child's behavior, both in action and perception. An adult can direct the child’s perception (“see the dog”) and action (“come here”) through language. The child gradually learns to use her own speech to perform these functions. Finally there’s no need for external scaffolding, that is, no speech either from an adult or from the child herself. One just ‘thinks.’

Now let us think of Vygotsky’s story in relation to the one I previously told about speaking, where we don’t know what we’re going to say until we’ve actually said it, at which point what we hear is either satisfactory, and we keep going, or not, so we stop and hazard another guess. Vygotsky has the child speaking in the presence an adult who ‘scaffolds’ their activity. They see some object, say a ball. The child looks at it and offers a word. If the word is correct, the adult nods approval. If not, the adult so indicates and the child tries again. In time the adult can be dispensed with. By that time the child has moved beyond the simple act of naming things and is formulating assertions about them. At that point the child is in the situation Chafe described.

Dual-System, Neuro-Symbolic Reasoning

Now let’s consider the dual-system model developed by Maxwell Nye and his colleagues. Here’s the abstract from their article:

Human reasoning can often be understood as an interplay between two systems: the intuitive and associative (“System 1”) and the deliberative and logical (“System 2”). Neural sequence models—which have been increasingly successful at performing complex, structured tasks—exhibit the advantages and failure modes of System 1: they are fast and learn patterns from data, but are often inconsistent and incoherent. In this work, we seek a lightweight, training-free means of improving existing System 1-like sequence models by adding System 2-inspired logical reasoning. We explore several variations on this theme in which candidate generations from a neural sequence model are examined for logical consistency by a symbolic reasoning module, which can either accept or reject the generations. Our approach uses neural inference to mediate between the neural System 1 and the logical System 2. Results in robust story generation and grounded instruction-following show that this approach can increase the coherence and accuracy of neurally-based generations.

The terms “System 1” and “System 2” are from the well-known work of Daniel Kahneman (Thinking, fast and slow, 2013).

Consider the following diagram from the article. Unfortunately it’s illegibly small in this presentation, but you don’t really need to read the text [you can click on it to enlarge it]. You simply need to know what’s happening in the boxes.

The upper pair of boxes represent System 1, which is automatically generated through a machine leaning process. The lower pair of boxes represent System 2, which is an Old School symbolic system created though hand-coding. The process proceeds as follows:

  1. The upper left box represents a story as it exists at a certain point in time.
  2. GPT-3 then parses that story into the hand-coded world model (lower left).
  3. System 1 then generates several candidate continuations of the story (upper right).
  4. GPT-3 parses those candidates into propositions consistent with the world model (lower right).
  5. Those propositions are checked against the current state of the story (lower middle).
  6. A proposition consistent with the model is chosen and the corresponding sentence is attached to the ongoing story (far right).

What is going on here is roughly, and I do mean roughly, what I wrote about in discussing language. Something is proposed, checked, and then redone if necessary. In this case the checking is done by a symbolic model. I suggest that is the case with natural language in human speakers as well.

With that in mind let’s consider these remarks about future possibilities for the dual system model:

A promising direction for future work is to incorporate learning into the System 2 world model. Currently, the minimal world knowledge that exists in System 2 can be easily modified, but changes must be made by hand. Improvements would come from automatically learning and updating this structured knowledge, possibly by incorporating neuro-symbolic learning techniques (Ellis et al., 2020; Mao et al., 2019).

Learning could improve our dual-system approach in other ways, e.g., by training a neural module to mimic the actions of a symbolic System 2. The symbolic System 2 judgments could be used as a source of supervision; candidate utterances rejected by the symbolic System 2 model could be used as examples of contradictory sentences, and accepted utterances could be used as examples of non-contradictory statements. This oversight could help train a neural System 2 contradiction-detection model capable of more subtleties than its symbolic counterpart, especially in domains where labeled examples are otherwise unavailable. This approach may also help us understand aspects of human learning, where certain tasks that require slower, logical reasoning can be habitualized over time and tackled by faster, more intuitive reasoning.

I can’t help but think that the training of “a neural System 2 contradiction-detection model” is rather like the child’s internalization of the scaffolding functions provided by sympathetic adults. The net result would be that a neural system internalizes the structure of a classical symbolic model, thereby coming to implement that model on a neural foundation.

Roughly speaking, think of the child language learner as the neural model, as System 1. The scaffolding adult provides the symbolic logical model, System 2. Over time, the logical structure provided by System 2 becomes incorporated into the ‘texture’ of neural System 1.

* * * * *

In thinking about this I considered attempting to draw the interaction Vygotsky described – I do draw it out in my longer exposition[3] in such a way that it matched the system diagram for the dual-system model. But I decided against it. It would have been tricky at best, and likely something of a force-fit, and I’m not at all sure the effort would have been rewarded by commensurate insight. The rough and ready correspondence that I’ve suggest seems sufficient to my purpose. And my primary purpose has been to assure myself that this model represents a significant step toward the eventual goal of implementing symbolic systems on a neural foundation. [See some tweets I've appended to the end.]

That’s what I had in mind when, over two decades ago, I played around with a notation system I called attractor nets [5]. There I was imagining a logical structure implemented over a large and various attractor landscape. The logical structure is represented by a network formalism developed by Sydney Lamb [4]. The network is quite different from those generally used in classical symbolic systems, where nodes represented objects of various kinds and arcs represented types of relationships among those objects. In Lamb’s nets the nodes are logical operators (OR, AND) while the arcs carry the content of the net. In an attractor net each arc corresponds to a basin of attraction. The net then represents a logical structure over basins of attraction.

It was an interesting conceptual experiment. It needs to be taken farther.

References

[1] I believe that we do need symbolic systems. But I also believe that they can ultimately be implemented in some kind of neural net. See this recent post for further remarks, “Geoffrey Hinton says deep learning will do everything. I’m not sure what he means, but I offer some pointers. Version 2,” New Savanna, May 31, 2021, https://new-savanna.blogspot.com/2021/05/geoffrey-hinton-says-deep-learning-will_31.html.

[2] “Kubla Khan” and the Embodied Mind, PsyArt: A Hyperlink Journal for the Psychological Study of the Arts, Article 030915, November 29, 2003, https://www.academia.edu/8810242/_Kubla_Khan_and_the_Embodied_Mind.

[3] I’ve explained this in a bit more detail in a blog post, “Vygotsky Tutorial (for Connected Courses),” New Savanna, September 10, 2020, https://new-savanna.blogspot.com/2014/10/vygotsky-tutorial-for-connected-courses.html.

That’s excerpted from a long article, First Person: Neuro-Cognitive Notes on the Self in Life and in Fiction, PsyArt: A Hyperlink Journal for Psychological Study of the Arts, August 21, 2000. Downloadable version, https://www.academia.edu/8331456/First_Person_Neuro-Cognitive_Notes_on_the_Self_in_Life_and_in_Fiction.

[4] Sydney Lamb, the computational linguist, believed this as well. He argues it in Pathways of the Brain, Amsterdam: John Benjamins (1998), pp. 181-182.

[5] I never produced an account that I thought was ready for others to read. But, in addition to piles of notes, I did produce two documents intended to summarize the work for my own purposes. The first of these two documents does explain Lamb’s notation and how I adapted it to my purposes, though the exposition is a bit round-about and perhaps attempts too much. The second document is a collection of diagrams, that is, constructions. While the diagrams are heavily notated, they presume that you understand that basic conventions as set forth in the first document. I am currently working on a document that will be a more suitable introduction to the notation.

William Benzon, Attractor Nets, Series I: Notes Toward a New Theory of Mind, Logic, and Dynamics in Relational Networks, Working Paper, 52 pp., https://www.academia.edu/9012847/Attractor_Nets_Series_I_Notes_Toward_a_New_Theory_of_Mind_Logic_and_Dynamics_in_Relational_Networks.

William Benzon, Attractor Nets 2011: Diagrams for a New Theory of Mind, Working Paper, 55 pp., https://www.academia.edu/9012810/Attractor_Nets_2011_Diagrams_for_a_New_Theory_of_Mind

* * * * *

Saturday, June 26, 2021

Summer Ramble 6.26.21: Literary Studies, Mind Models, Seinfeld, Joyful Nature

First I list the projects currently at the top of my to-do list. Then I discuss them by subject area.

Specific projects on deck

Three working papers in late stages of development:

The road to Xanadu led to Mars: I went one way, you all went another: I figure out how it is that structuralism led me through cognitive science and computation to a new approach to literary studies.

Crisis Among HUMANITIES DISCIPLINES in the Twenty-First Century: Built around the book, Permanent Crisis: The Humanities in a Disenchanted Age, in which Paul Reitter and Chad Wellmon argue that the institutional culture of the humanities is grounded in discord and discontent that follows from 19th debates about the nature of the university.

To Model the Mind: Speculative Engineering as Philosophy: Built around the book, Models of the Mind: How physics, engineering and mathematics have changed out understanding of the brain, by Grace Lindsey. I’m interested in an engineering approach to understanding both the human mind and AIs.

Big intellectual project on deck: To Model the Mind: A Primer on Attractor-Nets: It’s time to dust off the notes on Attractor-Nets which I made in the mid-2000s and get something in a form that other people can understand. Why? Because those notes speak to current debates in artificial intelligence about artificial neural-nets and symbolic computing.

Book project: Playing for Peace: Reclaiming our Human Nature, which is Volume 3 in the peace series I’ve been working on with Charlie Keil: Local Paths to Peace Today. This consists of various papers, by Charlie, me, and others, on the importance of music and play in life.

Seinfeld working paper: A working paper based on my current series of analyses of bits from his book, Is This Anything?

Literary criticism and the academy

While Crisis Among HUMANITIES DISCIPLINES is about the humanities in general, I concentrate on literary criticism because that’s the discipline I know the best. It turns out that much of the discussion centers on the difficulties I have had defining my relationship to a discipline that has, in effect, rejected me. The road to Xanadu led to Mars is about ideas, how mine evolved from structuralism, while the profession ultimately was unable to accommodate structuralist insights and was forced to deconstruction and post-structuralism (which doesn’t go beyond structuralism, but rejects it). Here’s a short chronology that encompasses events mentioned in both working paper:

1972: Finish my master’s thesis on “Kubla Khan” – The thesis was structuralist in conception and spirit, but moved to structuralism’s farthest edge and sent me toward computation and cognitive science. Central to Xanadu/Mars.

1976: Publish my article “Cognitive Science and Literary Semantics” in MLN’s Centennial Issue in 1976. I’ve now gone beyond the boundaries of literary criticism as those boundaries would evolve into the 1980s, but that wasn’t obvious at the time. At the time those boundaries were fluid. Central to Xanadu/Mars.

1985: I was denied tenure at RPI and unable to secure another academic post. Sometime in the wake of that I negotiate (with myself) a Socratic bargain with the profession. That bargain is modeled on The Crito. I discuss this extensively in Crisis.

1995: I begin reconsidering the nature of my intellectual enterprise in relation to literary criticism as critics become interest in cognitive science, but not the cognitive science that informed my work. I decide that I had all along been interested in form. I discuss this extensively in Xanadu/Mars

2001: I begin corresponding with Mary Douglas. She gets me interested in ring-form composition. In 2003 I send her a long post about ring-form in “The Nutcracker Suite” episode of Fantasia. Relevant to Xanadu/Mars

2010: I renegotiate the Socratic bargain. I am no longer loyal to the profession, but simply to the free and unfettered pursuit of truth in an international community of scholars. Again, central to Crisis.

2021: Make sense of it all: Crisis and Xanadu/Mars.

Modeling the mind

Finishing the working paper, To Model the Mind: Speculative Engineering as Philosophy, should be relatively straightforward, though tricky. The main task to be accomplished is to talk in a meaningful way about all the kinds of mind-like entities, artificial, real, and hybrid, that we are currently dealing with. I’m not talking about a thorough discussion, just an indication of the range along with some sense of the kinds of issues involved.

The attractor net primer is perhaps the most important item on my current agenda. I’m guessing the primer is going to run 40 to 50 pages or more and will have lots of diagrams. I have to get it into the current arena of discussion about AI and the mind.

Joyous nature book

I’m currently editing Playing for Peace: Reclaiming our Human Nature. I need to write an introduction, finish the guide-to-contents, and tell Charlie (Keil) what help I need from him to finish it up.

Working Paper on Jerry Seinfeld

Provisional title of the working paper: Seinfeld’s Comedy: Jokes are Intricately Crafted Machines. I need to do analyses of three more bits, two of which I’ve already done, though in another form, and write an introduction.

Monday, May 31, 2021

Geoffrey Hinton says deep learning will do everything. I’m not sure what he means, but I offer some pointers. Version 2.

This is updated from a previous version to include a passage by Sydney Lamb.

* * * * *

Late last year Geoffrey Hinton had an interview with Karen Hao [1] in which he said “I do believe deep learning is going to be able to do everything,” with the qualification that “there’s going to have to be quite a few conceptual breakthroughs.” I’m trying to figure out whether or not, to what extent, in what way I (might) agree with him.

Neural Vectors, Symbols, Reasoning, and Understanding

Hinton believes that “What’s inside the brain is these big vectors of neural activity” and that one of the breakthroughs we need is “how you get big vectors of neural activity to implement things like reason.” That will certainly require a massive increase in scale. Thus while GPT-3 has 175 billion parameters, the brain has trillion, where Hinton treats each synapse as a parameter. 

Correspondingly Hinton rejects the idea that symbolic reasoning is primitive to the nervous system (my formulation), rather “we do internal operations on big vectors.” What about language? He doesn’t address the issue directly but he does say that “symbols just exist out there in the external world.” I do think that covers language, speech sounds, written words, gestural signs, those things are out there in the external world. But the brain uses “big vectors of neural activity” to process those. 

Hinton’s remark about symbols bears comparison with a remark by Sydney Lamb: “the linguistic system is a relational network and as such does not contain lexemes or any objects at all. Rather it is a system that can produce and receive such objects. Those objects are external to the system, not within it”[2]. Lamb has come to think of his approach as neurocognitive linguistics and, while his sense of the nervous system is somewhat different from Hinton’s, they agree on this issue and, in the current intellectual climate, that agreement is of some significance. For Lamb is a first generation researcher in machine translation  and so was working when most AI research was committed to symbolic systems. We’ll return to Lamb later as I think the notation he developed is a way to bring symbolic reasoning within range of Hinton’s “big vectors of neural activity”.

But now let’s return to Hinton with a passage from an article he co-authored with Yann LeCun and Yoshua Bengio [3]:

In the logic-inspired paradigm, an instance of a symbol is something for which the only property is that it is either identical or non-identical to other symbol instances. It has no internal structure that is relevant to its use; and to reason with symbols, they must be bound to the variables in judiciously chosen rules of inference. By contrast, neural networks just use big activity vectors, big weight matrices and scalar non-linearities to perform the type of fast ‘intuitive’ inference that underpins effortless commonsense reasoning.

I note, moreover, commonsense reasoning seems to be problematic for everyone.[4]

Let’s look at one more passage from the interview:

For things like GPT-3, which generates this wonderful text, it’s clear it must understand a lot to generate that text, but it’s not quite clear how much it understands.

I’m not sure that it is at all useful to say that GPT-3 understands anything. I think that, in using that term, Hinton is displaying what I’ve come to think of as the word illusion.[5] Briefly, GPT-3’s language model is constructed over a corpus consisting entirely of word forms, of signifiers without signifieds, to use an old terminology. Hinton knows that, of course, but, after all, he understands texts from seeing or hearing word forms alone, as do we all, and so, in effect, credits GPT-3 with somehow having induced meaning from a statistical distribution. The text it generates looks pretty good, no? Yes. And that is something we do need to understand, just what is GPT-3 doing and how does it do it? But this is not the place to enter into that.[6]

I think that GPT-3’s remarkable performance based on such ‘shallow’ material should prompt us into reconsidering just what humans are doing when we produce everyday ‘boilerplate’ text. Consider this passage from LeCun, Bengio, and Hinton, where they are referring to the use of an RNN:

This rather naive way of performing machine translation has quickly become competitive with the state-of-the-art, and this raises serious doubts about whether understanding a sentence requires anything like the internal symbolic expressions that are manipulated by using inference rules. It is more compatible with the view that everyday reasoning involves many simultaneous analogies that each contribute plausibility to a conclusion
.

In dealing with these utterly remarkable devices, we would be rein in our narcissistic investment in the routine use of our ‘higher’ cognitive and linguistic capacities as opposed to our mere sensory-motor competence. It’s all neural vectors. 

Note, however, that it is one thing to say that “we do internal operations on big vectors.” I agree with that. That’s not quite the same as saying we can do everything with deep learning. Deep learning is a collection of architectures, but I’m not sure such architectures are adequate for internalizing the vectors needed to effectively mimic human perceptual and cognitive behavior. The necessary conceptual breakthroughs will likely take us considerably beyond deep learning engines. With that qualification, let’s continue.

How the brain might be doing it

I find that, with the caveats I’ve mentioned, this is rather congenial. Which is to say that I can make sense of it in terms of issues I’ve thought through in my own work.

Some years ago David Hays and I wanted to come to terms with neuroscience and ended up reviewing a wide range of work and writing a paper entitled, “Principles and Development of Natural Intelligence.”[7] The principles are ordered such that principle N assumed N-1. We called the fifth and last principle indexing:

The indexing principle is about computational geometry, by which we mean the geometry, that is, the architecture (Pylyshyn, 1980) of computation rather than computing geometrical structures. While the other four principles can be construed as being principles of computation, only the indexing principle deals with computing in the sense it has had since the advent of the stored program digital computer. Indexed computation requires (1) an alphabet of symbols and (2) relations over places, where tokens of the alphabet exist at the various places in the system. The alphabet of symbols encodes the contents of the calculation while the relations over places, i.e. addresses, provide the means of manipulating alphabet tokens in carrying out the computation. [...] Within the context of natural intelligence, indexing is embodied in language. Linguists talk of duality of patterning (Hockett, 1960), the fact that language patterns both sounds and sense. The system which patterns sound is used to index the system which patterns sense.

In short, “indexing gives computational geometry, and language enables the system to operate on its own geometry.” This is where we get symbols and complex reasoning.

I should note that, while we talked of “an alphabet of symbols” and “relations over places” we were not asserting that that’s what was going on in the brain. That’s what’s actually going on in computers, but it applies only figuratively to the brain. The system that is using sound patterns to index patterns of sense is using one set of neural vectors (though we didn’t use that term) to index a different set of neural vectors.

How do we get deep learning to figure that out? I note that automatic image annotation is a step in that direction [8], but have nothing to say about that here.

Instead I want to mention some informal work I did some years ago on something I call attractor nets.[9] The general idea was to use Sydney Lamb’s relational networks, in which nodes are logical operators, as a tertium quid between the symbol-based semantic networks Hays and I had worked on in the 1970s and the attractor landscapes of Walter Freeman’s neurodynamics. I showed – informally, using diagrams – how using logical operators (AND, OR) over attractor basins in different neurofunctional areas could reconstruct symbolic systems represented as directed graphs. Each node in a symbolic graph corresponds to a basin of attraction, that is, an attractor. In the present context we can think of each neurofunctional area as corresponding to a collection of neural vectors and the attractors as objects represented by those vectors. An attractor net would then become a way of thinking about how complex reasoning could be accomplished with neural vectors.

In the attractor net notation word forms, or signifiers, are distinct from word meanings, of signifieds. Is that distinction important for complex reasoning? I believe it is, though I’m not interested in constructing an argument at this point. That, I believe, puts a limit on what one can expect of engines like GPT-3. That too requires an argument.

So, what about natural vs. artificial intelligence?

The notion of intelligence is somewhat problematic. As a practical matter I believe that a formulation by Robin Hanson is adequate: “’Intelligence’ just means an ability to do mental/calculation tasks, averaged over many tasks.”[10] As for the difference between artificial and natural, that comes down to four things:

1) a living system vs. an inanimate system,
2) a carbon-based organic electro-chemical substrate vs. a silicon-based electronic substrate,
3) real neurons (having on average 10K connections with others) vs. considerably simpler artificial neurons realized in program code, and
4) the neurofunctional architecture and innate capacities of a real brain vs. the system architecture of a digital computing system.

Make no mistake, those differences are considerable. But I think we now have in hand a body of concepts and models that is rich enough to support ever more sophisticated interaction between students of neuroscience and students of artificial intelligence. To the extent that our research and teaching institutions can support that interaction I expect to see progress accelerate in the future. I offer no predictions about what will come of this interaction.

Some related posts

William Benzon, Showdown at the AI Corral, or: What kinds of mental structures are constructible by current ML/neural-net methods? [& Miriam Yevick 1975], New Savanna, June 3, 2020, https://new-savanna.blogspot.com/2020/06/showdown-at-ai-corral-or-what-kinds-of.html.

William Benzon, What’s AI? – Part 2, on the contrasting natures of symbolic and statistical semantics [can GPT-3 do this?], New Savanna, July 17, 2020, https://new-savanna.blogspot.com/2019/11/whats-ai-part-2-on-contrasting-natures.html.

William Benzon, A quick note on the ‘neural code’ [AI meets neuroscience], New Savanna, April 20, 2021, https://new-savanna.blogspot.com/2021/04/a-quick-note-on-neural-code-ai-meets.html.

References

[1] Interview with Karen Hao, AI pioneer Geoff Hinton: “Deep learning is going to be able to do everything”, MIT Technology Review, Nov. 3, 2020. https://www.technologyreview.com/2020/11/03/1011616/ai-godfather-geoffrey-hinton-deep-learning-will-do-everything/

[2] Sydney Lamb, Linguistic Structure: A Plausible Theory, Language Under Discussion, 4(1) 2016, 1–37, https://doi.org/10.31885/lud.4.1.229.

[3] From Yann LeCun, Yoshua Bengio, and Geoffrey Hinton, Deep Learning, Nature, 521 28 May 2015, 436-444, https://doi.org/10.1038/nature14539.

[4] As an example I offer a recent post in which I quiz GPT-3 about a Jerry Seinfeld bit: Analyze This! Screaming on the flat part of the roller coaster ride [Does GPT-3 get the joke?], May 7, 2021, https://new-savanna.blogspot.com/2021/05/analyze-this-screaming-on-flat-part-of.html.

[5] See my post, The Word Illusion, May 12, 2021, https://new-savanna.blogspot.com/2021/05/the-word-illusion.html.

[6] For some extended remarks, see my working paper, GPT-3: Waterloo or Rubicon? Here be Dragons, Working Paper, Version 3, August 20, 2020, 34 pp., https://www.academia.edu/43787279/GPT_3_Waterloo_or_Rubicon_Here_be_Dragons_Version_3.

[7] William Benzon and David Hays, Principles and Development of Natural Intelligence, Journal of Social and Biological Structures, Vol. 11, No. 8, July 1988, 293-322, https://www.academia.edu/235116/Principles_and_Development_of_Natural_Intelligence.

[8] Wikipedia, Automatic image annotation, https://en.wikipedia.org/wiki/Automatic_image_annotation.

[9] William Benzon, Attractor Nets, Series I: Notes Toward a New Theory of Mind, Logic, and Dynamics in Relational Networks, Working Paper, 52 pp., https://www.academia.edu/9012847/Attractor_Nets_Series_I_Notes_Toward_a_New_Theory_of_Mind_Logic_and_Dynamics_in_Relational_Networks.

William Benzon, Attractor Nets 2011: Diagrams for a New Theory of Mind, Working Paper, 55 pp., https://www.academia.edu/9012810/Attractor_Nets_2011_Diagrams_for_a_New_Theory_of_Mind.

William Benzon, From Associative Nets to the Fluid Mind, Working Paper. October 2013, 16 pp. https://www.academia.edu/9508938/From_Associative_Nets_to_the_Fluid_Mind.

[10] Robin Hanson, I Still Don’t Get Foom, Overcoming Bias, July 24, 2014, https://www.overcomingbias.com/2014/07/30855.html.

Friday, May 14, 2021

Geoffrey Hinton says deep learning will do everything. I’m not sure what he means, but I offer some pointers.

Superseded by Version 2 with an additional paragraph about Sydney Lamb.

* * * * *

Late last year Geoffrey Hinton had an interview with Karen Hao [1] in which he said “I do believe deep learning is going to be able to do everything,” with the qualification that “there’s going to have to be quite a few conceptual breakthroughs.” I’m trying to figure out whether or not, to what extent, in what way I (might) agree with him.

Neural Vectors, Symbols, Reasoning, and Understanding

Hinton believes that “What’s inside the brain is these big vectors of neural activity” and that one of the breakthroughs we need is “how you get big vectors of neural activity to implement things like reason.” That will certainly require a massive increase in scale. Thus while GPT-3 has 175 billion parameters, the brain has trillion, where Hinton treats each synapse as a parameter.

Correspondingly Hinton rejects the idea that symbolic reasoning is primitive to the nervous system (my formulation), rather “we do internal operations on big vectors.” What about language? He doesn’t address the issue directly but he does say that “symbols just exist out there in the external world.” I do think that covers language, speech sounds, written words, gestural signs, those things are out there in the external world. But the brain uses “big vectors of neural activity” to process those.

Let’s look at a passage from an article Hinton co-authored with Yann LeCun and Yoshua Bengio [2]:

In the logic-inspired paradigm, an instance of a symbol is something for which the only property is that it is either identical or non-identical to other symbol instances. It has no internal structure that is relevant to its use; and to reason with symbols, they must be bound to the variables in judiciously chosen rules of inference. By contrast, neural networks just use big activity vectors, big weight matrices and scalar non-linearities to perform the type of fast ‘intuitive’ inference that underpins effortless commonsense reasoning.

I note, however, commonsense reasoning seems to be problematic for everyone.[3]

Let’s look at one more passage from the interview:

For things like GPT-3, which generates this wonderful text, it’s clear it must understand a lot to generate that text, but it’s not quite clear how much it understands.

I’m not sure that it is at all useful to say that GPT-3 understands anything. I think that, in using that term, Hinton is displaying what I’ve come to think of as the word illusion.[4] Briefly, GPT-3’s language model is constructed over a corpus consisting entirely of word forms, of signifiers without signifieds, to use an old terminology. But Hinton knows that, of course, but, after all, he understands texts on the basis of word forms alone, as do we all, and so, in effect, credits GPT-3 with somehow having induced meaning from a statistical distribution. The text it generates looks pretty good, no? Yes. And that is something we do need to understand, just what is GPT-3 doing and how does it do it? But this is not the place to enter into that.[5]

I think that GPT-3’s remarkable performance based on such ‘shallow’ material should prompt us into reconsidering just what humans are doing when we produce everyday ‘boilerplate’ text. Consider this passage from LeCun, Bengio, and Hinton, where they are referring to the use of an RNN:

This rather naive way of performing machine translation has quickly become competitive with the state-of-the-art, and this raises serious doubts about whether understanding a sentence requires anything like the internal symbolic expressions that are manipulated by using inference rules. It is more compatible with the view that everyday reasoning involves many simultaneous analogies that each contribute plausibility to a conclusion
.

In dealing with these utterly remarkable devices, we would be rein in our narcissistic investment in the routine use of our ‘higher’ cognitive and linguistic capacities as opposed to our mere sensory-motor competence. It’s all neural vectors. 

Note, however, that it is one thing to say that “we do internal operations on big vectors.” I agree with that. That’s not quite the same as saying we can do everything with deep learning. Deep learning is a collection of architectures, but I’m not sure such architectures are adequate for internalizing the vectors needed to effectively mimic human perceptual and cognitive behavior. The necessary conceptual breakthroughs will likely take us considerably beyond deep learning engines. With that qualification, let’s continue.

How the brain might be doing it

I find that, with the caveats I’ve mentioned, this is rather congenial. Which is to say that I can make sense of it in terms of issues I’ve thought through in my own work.

Some years ago David Hays and I wanted to come to terms with neuroscience and ended up reviewing a wide range of work and writing a paper entitled, “Principles and Development of Natural Intelligence.”[6] The principles are ordered such that principle N assumed N-1. We called the fifth and last principle indexing:

The indexing principle is about computational geometry, by which we mean the geometry, that is, the architecture (Pylyshyn, 1980) of computation rather than computing geometrical structures. While the other four principles can be construed as being principles of computation, only the indexing principle deals with computing in the sense it has had since the advent of the stored program digital computer. Indexed computation requires (1) an alphabet of symbols and (2) relations over places, where tokens of the alphabet exist at the various places in the system. The alphabet of symbols encodes the contents of the calculation while the relations over places, i.e. addresses, provide the means of manipulating alphabet tokens in carrying out the computation. [...] Within the context of natural intelligence, indexing is embodied in language. Linguists talk of duality of patterning (Hockett, 1960), the fact that language patterns both sounds and sense. The system which patterns sound is used to index the system which patterns sense.

In short, “indexing gives computational geometry, and language enables the system to operate on its own geometry.” This is where we get symbols and complex reasoning.

I should note that, while we talked of “an alphabet of symbols” and “relations over places” we were not asserting that that’s what was going on in the brain. That’s what’s actually going on in computers, but it applies only figuratively to the brain. The system that is using sound patterns to index patterns of sense is using one set of neural vectors (though we didn’t use that term) to index a different set of neural vectors.

How do we get deep learning to figure that out? I note that automatic image annotation is a step in that direction [7], but have nothing to say about that here.

Instead I want to mention some informal work I did some years ago on something I call attractor nets.[8] The general idea was to use Sydney Lamb’s relational networks, in which nodes are logical operators, as a tertium quid between the symbol-based semantic networks Hays and I had worked on in the 1970s and the attractor landscapes of Walter Freeman’s neurodynamics. I showed – informally, using diagrams – how using logical operators (AND, OR) over attractor basins in different neurofunctional areas could reconstruct symbolic systems represented as directed graphs. Each node in a symbolic graph corresponds to a basin of attraction, that is, an attractor. In the present context we can think of each neurofunctional area as corresponding to a collection of neural vectors and the attractors as objects represented by those vectors. An attractor net would then become a way of thinking about how complex reasoning could be accomplished with neural vectors.

In the attractor net notation word forms, or signifiers, are distinct from word meanings, of signifieds. Is that distinction important for complex reasoning? I believe it is, though I’m not interested in constructing an argument at this point. That, I believe, puts a limit on what one can expect of engines like GPT-3. That too requires an argument.

So, what about natural vs. artificial intelligence?

The notion of intelligence is somewhat problematic. As a practical matter I believe that a formulation by Robin Hanson is adequate: “’Intelligence’ just means an ability to do mental/calculation tasks, averaged over many tasks.”[9] As for the difference between artificial and natural, that comes down to four things:

1) a living system vs. an inanimate system,
2) a carbon-based organic electro-chemical substrate vs. a silicon-based electronic substrate,
3) real neurons (having on average 10K connections with others) vs. considerably simpler artificial neurons realized in program code, and
4) the neurofunctional architecture and innate capacities of a real brain vs. the system architecture of a digital computing system.

Make no mistake, those differences are considerable. But I think we now have in hand a body of concepts and models that is rich enough to support ever more sophisticated interaction between students of neuroscience and students of artificial intelligence. To the extent that our research and teaching institutions can support that interaction I expect to see progress accelerate in the future. I offer no predictions about what will come of this interaction.

Some related posts

William Benzon, Showdown at the AI Corral, or: What kinds of mental structures are constructible by current ML/neural-net methods? [& Miriam Yevick 1975], New Savanna, June 3, 2020, https://new-savanna.blogspot.com/2020/06/showdown-at-ai-corral-or-what-kinds-of.html.

William Benzon, What’s AI? – Part 2, on the contrasting natures of symbolic and statistical semantics [can GPT-3 do this?], New Savanna, July 17, 2020, https://new-savanna.blogspot.com/2019/11/whats-ai-part-2-on-contrasting-natures.html.

William Benzon, A quick note on the ‘neural code’ [AI meets neuroscience], New Savanna, April 20, 2021, https://new-savanna.blogspot.com/2021/04/a-quick-note-on-neural-code-ai-meets.html.

References

[1] Interview with Karen Hao, AI pioneer Geoff Hinton: “Deep learning is going to be able to do everything”, MIT Technology Review, Nov. 3, 2020. https://www.technologyreview.com/2020/11/03/1011616/ai-godfather-geoffrey-hinton-deep-learning-will-do-everything/.

[2] From Yann LeCun, Yoshua Bengio, and Geoffrey Hinton, Deep Learning, Nature, 521 28 May 2015, 436-444, https://doi.org/10.1038/nature14539.

[3] As an example I offer a recent post in which I quiz GPT-3 about a Jerry Seinfeld bit: Analyze This! Screaming on the flat part of the roller coaster ride [Does GPT-3 get the joke?], May 7, 2021, https://new-savanna.blogspot.com/2021/05/analyze-this-screaming-on-flat-part-of.html.

[4] See my post, The Word Illusion, May 12, 2021, https://new-savanna.blogspot.com/2021/05/the-word-illusion.html.

[5] For some extended remarks, see my working paper, GPT-3: Waterloo or Rubicon? Here be Dragons, Working Paper, Version 2, August 20, 2020, 34 pp., https://www.academia.edu/43787279/GPT_3_Waterloo_or_Rubicon_Here_be_Dragons_Version_2.

[6] William Benzon and David Hays, Principles and Development of Natural Intelligence, Journal of Social and Biological Structures, Vol. 11, No. 8, July 1988, 293-322, https://www.academia.edu/235116/Principles_and_Development_of_Natural_Intelligence.

[7] Wikipedia, Automatic image annotation, https://en.wikipedia.org/wiki/Automatic_image_annotation.

[8] William Benzon, Attractor Nets, Series I: Notes Toward a New Theory of Mind, Logic, and Dynamics in Relational Networks, Working Paper, 52 pp., https://www.academia.edu/9012847/Attractor_Nets_Series_I_Notes_Toward_a_New_Theory_of_Mind_Logic_and_Dynamics_in_Relational_Networks.

William Benzon, Attractor Nets 2011: Diagrams for a New Theory of Mind, Working Paper, 55 pp., https://www.academia.edu/9012810/Attractor_Nets_2011_Diagrams_for_a_New_Theory_of_Mind.

William Benzon, From Associative Nets to the Fluid Mind, Working Paper. October 2013, 16 pp. https://www.academia.edu/9508938/From_Associative_Nets_to_the_Fluid_Mind.

[9] Robin Hanson, I Still Don’t Get Foom, Overcoming Bias, July 24, 2014, https://www.overcomingbias.com/2014/07/30855.html.

Friday, July 17, 2020

What’s AI? – Part 2, on the contrasting natures of symbolic and statistical semantics [can GPT-3 do this?]

In view of current excitement over GPT-3 I'm bumping this to the top of the queue. This is about the implementation of Old School semantic or cognitive nets in neural nets where the nodes and edges of the semantic net each represent activity in regions of the neural net. 

* * * * *

Comparison of Semantic Nets (left) with Attractor Nets (right). See below.
Color me pleased with the AI essay I just published in 3 Quarks Daily, Some Notes On Computers, AI, And The Mind, and my follow-up here, What’s AI? – @3QD [& the contrasting natures of symbolic and statistical semantics]. These notes are, in turn, a follow-up to my follow-up. In particular, I want to elaborate on these concluding paragraphs:
Now, let’s compare these two systems. In the symbolic computation semantics we have a relational graph where both the nodes and the arcs are labeled. All that matters is the topology of the graph, that is, the connectivity of the nodes and arcs, and the labels on the nodes and arcs. The nature of those labels is very important.

In the statistical system words are in fixed geometric positions in a high-dimensional space; exact positions, that is, distances, are critical. This system, however, doesn’t need node and arc labels. The vectors do all the work.
These statements are, I feel, at the right level of generalization and abstraction. The purpose of these notes is to lay out some of the things that would have to be taken into consideration in order to further develop those ideas.

HOWEVER, it’s sketchy & full of holes. First the sketchy stuff. Then I’ve got links to material that’s somewhat more worked out. It’ll help full in the holes and flesh out the details.

Stream of consciousness ramble-through
Caveat: These are informal notes, mostly for myself. I list them here as place-holders for work that needs to be done.
How do you make inference in the system? Symbolic model allows for a logic based on node and arc types. Such a model treats the knowledge structure as a database and makes inferences over it. Humans are not like that. In statistical systems there is no capacity for ‘externally’ guided inference. All inference is ‘black-boxed’ and internal to the system.

How do you add “meaning” to a statistical model? What is “meaning” anyhow? Does meaning ultimately require/imply interactive coupling with [living in] a world? And doesn’t it imply/require an intending subject?

Is this the ultimate import of critiques such as those of Dreyfus and Searle? But aren’t those ‘cheap’ critiques, arrived at without knowledge of how these systems work? And what of it?

How do we get rid of arc and node labels? Doesn’t that amount to coupling with the external world? See David Hays in Cognitive Structures (HRAF Press 1981) and Sydney Lamb’s notation for his stratificational grammar. And then we have my attractor nets, where the nodes are logical operators over attractor landscapes and the arcs are attractors within those landscapes.

If symbolic systems are more flexible and ‘deeper’ how come the statistical systems are more powerful in actual applications? The symbolic systems get their knowledge directly from the humans that design them. The makers decide on node and arc types and the way to make inferences over them; and the makers encode knowledge directly into the system. Statistical systems learn. NLP systems in effect back in to a simulacrum of the knowledge humans have ‘baked-in’ to the texts they generate. But what about visual systems based on unsupervised learning? Those systems are directly ‘in touch with’ an ‘external world’, no?

What about the phenomenal power of chess and Go programs? These programs are now more powerful than the best human players. I note as well that neither chess nor Go require contact with a rich physical world. Both take place in a 2D world containing a very limited universe of objects and in which there are very limited opportunities for action. For all practical purposes, there is no world external to the computer, unlike what we have for translation systems, text generation systems, or car-driving systems. Does this imply that perhaps the best and most powerful use of these systems is in the internal configuration of and maintenance of computational systems themselves? But won’t that make these systems even more ‘black-boxy’ than they already are?

And speech-to-text, text-to-speech?

Note that learning systems of various sorts do require exposure to mountains of data, whether externally generated and presented to the system or created by the system itself (as in game systems that compete against themselves). Humans do not require exposure to so many cases for effective learning. What’s this about? Well, for one thing it’s about having a brain and body that evolved to fit the world: no blank slate. Does is all cascade back on to that, or is there more?

Natural intelligence and metaphor

While much of my work with David Hays was anchored in a symbolic systems approach to the mind, we did venture elsewhere. Thus when the symbolic systems approach collapsed in the mid-1980s we were, if not exactly prepared, able to keep moving on. We had other conceptual irons in the forge.

In Cognitive Structures, which I linked above, Hays came up with a scheme to ground a digital cognitive system in an analog sensorimotor system. A few years after that he and I worked on a paper where we grounded our whole system in neural systems:
William L. Benzon and David G. Hays, Principles and development of natural intelligence, Journal of Social and Biological Systems 11, 293-322, 1988, https://www.academia.edu/235116/Principles_and_Development_of_Natural_Intelligence.

Abstract: The phenomena of natural intelligence can be grouped into five classes, and a specific principle of information processing, implemented in neural tissue, produces each class of phenomena. (1) The modal principle subserves feeling and is implemented in the reticular formation. (2) The diagonalization principle subserves coherence and is the basic principle, implemented in neocortex. (3) Action is subserved by the decision principle, which involves interlinked positive and negative feedback loops, and resides in modally differentiated cortex. (4) The problem of finitization resolves into a figural principle, implemented in secondary cortical areas; figurality resolves the conflict between pro-positional and Gestalt accounts of mental representations. (5) Finally, the phenomena of analysis reflect the action of the indexing principle, which is implemented through the neural mechanisms of language.

These principles have an intrinsic ordering (as given above) such that implementation of each principle presupposes the prior implementation of its predecessor. This ordering is preserved in phylogeny: (1) mode, vertebrates; (2) diagonalization, reptiles; (3) decision, mammals; (4) figural, primates; (5) indexing. Homo sapiens sapiens. The same ordering appears in human ontogeny and corresponds to Piaget's stages of intellectual development, and to stages of language acquisition.
I note that we conceived of the second principle, diagonalization, as involving holograph-like processing, which involves convolution. Convolution is important in some forms of contemporary neural net systems. Subsequently we published a paper in which we used convolution to account for metaphor:
William L. Benzon and David G. Hays, Metaphor, Recognition, and Neural Process, The American Journal of Semiotics, Vol. 5. No. 1 (1987), 59-80, https://www.academia.edu/238608/Metaphor_Recognition_and_Neural_Process.

Abstract: Karl Pribram's concept of neural holography suggests a neurological basis for metaphor: the brain creates a new concept by the metaphoric process of using one concept as a filter — better, as an extractor — for another. For example, the concept "Achilles" is "filtered" through the concept "lion" to foreground the pattern of fighting fury the two hold in common. In this model the linguistic capacity of the left cortical hemisphere is augmented by the capacity of the right hemisphere for analysis of images. Left-hemisphere syntax holds the tenor and vehicle in place while right-hemisphere imaging process extracts the metaphor ground. Metaphors can be concatenated one after the other so that the ground of one metaphor can enter into another one as tenor or vehicle. Thus conceived metaphor is a mechanism through which thought can be extended into new conceptual territory.

Wednesday, June 3, 2020

Showdown at the AI Corral, or: What kinds of mental structures are constructible by current ML/neural-net methods? [& Miriam Yevick 1975]

There’s a big conversation going on in the AI world these days about appropriate architectures: Can new-style machine-learning/neural-net methods take us all the way to the Holy Land or do we need to incorporate old-style structured symbolic models? My sense the that most current practitioners learn toward the former position; I favor the latter. (As for the Holy Land, it’s a mirage.)

In the first section I present an important, if somewhat neglected, paper from 1975 and what David Hays and I made from it. Then some tweets culled from the current stream. I conclude with some crazy stuff, my notes on attractor networks and the fluid mind. That crazy stuff is about blending the two kinds of logic, intermixing structured symbolic systems with data-driven machine/deep learning systems. Something like that.

Yevick’s Law: Two Kinds of Logic

As Louis Armstrong used to say, it’s one of those old time good ones:
Yevick, Miriam Lipschutz (1975) Holographic or Fourier logic. Pattern Recognition 7: 197-213.
https://doi.org/10.1016/0031-3203(75)90005-9

Abstract: A tentative model of a system whose objects are patterns on transparencies and whose primitive operations are those of holography is presented. A formalism is developed in which a variety of operations is expressed in terms of two primitives: recording the hologram and filtering. Some elements of a holographic algebra of sets are given. Some distinctive concepts of a holographic logic are examined, such as holographic identity, equality, containment and “association”. It is argued that a logic in which objects are defined by their “associations” is more akin to visual apprehension than description in terms of sequential strings of symbols.
Yes, it was published in 1975, which is ancient times in the world of artificial intelligence. It was inspired by a body of theorizing and evidence – promulgated by Karl Pribram, among others – that the neocortical processing was based on holographic principles rather than those of propositional/symbolic logic. It seems to me that what Yevick called holographic logic is similar in spirit, and even in mathematics in some respects, to current work on neural networks, while, in contrast, ordinary logic is as the abstract has it, "description in terms of sequential strings of symbols."

A decade later David Hays and I called on that paper in a highly speculative synthesis of a variety of work in cognitive, neural, perceptual, and comparative psychology with computational orientation:
William Benzon and David Hays, Principles and Development of Natural Intelligence, Journal of Social and Biological Structures, Vol. 11, No. 8, July 1988, 293-322.
https://www.academia.edu/235116/Principles_and_Development_of_Natural_Intelligence
We sketched out five principles. The fourth principle, which we called the figural principle, was based on Yevick’s work. Here’s how we opened the discussion:
The figural principle concerns the relationship between Gestalt or analogue process in neural schemas and propositional or digital processes. In our view, both are necessary; the figural principle concerns the relationship between the two types of process. The best way to begin is to consider Miriam Yevick's work (1975, 1978) on the relationship between 'descriptive and holistic' (analogue) and 'recursive and ostensive' (digital) processes in representation.
Fig. 10. Yevick's law. The curves indicate the level of representational complexity required for a good identification
The critical relationship is that between the complexity of the object and the complexity of the representation needed to ensure specific identification. If the object is simple, e.g. a square, a circle, a cross, a simple propositional schema will yield a sharp identification, while a relatively complex Gestalt schema will be required for an equivalently good identification (see Fig. 10). Conversely, if the object is complex, e.g. a Chinese ideogram, a face, a relatively simple Gestalt (Yevick used Fourier transforms) will yield a sharp identification, while an equivalently precise propositional schema will be more complex than the object it represents. Finally, we have those objects which fall in the middle region of Figure 10, objects that have no particularly simple description by either Gestalt or propositional methods and instead require an interweaving of both. That interweaving is the figural principle.

Definition. The figural mechanism brings environments of moderate complexity within the limits of computability by putting a propositional assemblage of local narrow band-width Gestalts into a framework provided by global wide band-width analysis to achieve cross-validation of the two analyses.
At various points later in the essay of the propositional reconstruction Gestalt processes. In our view the mind evolves through an interweaving of these two kinds of logic. Both are necessary.
Subsequently we used this line of thought in a paper about metaphor:
William Benzon and David Hays, Metaphor, Recognition, and Neural Process, The American Journal of Semiotics, Vol. 5, No. 1 (1987), 59-80.
https://www.academia.edu/238608/Metaphor_Recognition_and_Neural_Process

Karl Pribram's concept of neural holography suggests a neurological basis for metaphor: the brain creates a new concept by the metaphoric process of using one concept as a filter — better, as an extractor — for another. For example, the concept “Achilles” is “filtered” through the concept “lion” to foreground the pattern of fighting fury the two hold in common. In this model the linguistic capacity of the left cortical hemisphere is augmented by the capacity of the right hemisphere for analysis of images. Left-hemisphere syntax holds the tenor and vehicle in place while right-hemisphere imaging process extracts the metaphor ground. Metaphors can be concatenated one after the other so that the ground of one metaphor can enter into another one as tenor or vehicle. Thus conceived metaphor is a mechanism through which thought can be extended into new conceptual territory.
Caveat: Students of cognitive linguistics should think of this as an account of blending rather than (cognitive) metaphor.

Some current work, from the Twitterverse