Showing posts with label artificial minds. Show all posts
Showing posts with label artificial minds. Show all posts

Wednesday, April 2, 2025

AI & humans, then and now

In “The Evolution of Cognition” (1990) David Hays and I argued that the long-term evolution of human culture flows from the architectural foundations of thought and communication: first speech, then writing, followed by systematized calculation, and most recently, computation. In discussing the importance of the computer, we remark:

One of the problems we have with the computer is deciding what kind of thing it is, and therefore what sorts of tasks are suitable to it. The computer is ontologically ambiguous. Can it think, or only calculate? Is it a brain or only a machine?

The steam locomotive, the so-called iron horse, posed a similar problem for people at Rank 3. It is obviously a mechanism and it is inherently inanimate. Yet it is capable of autonomous motion, something heretofore only within the capacity of animals and humans. So, is it animate or not? Perhaps the key to acceptance of the iron horse was the adoption of a system of thought that permits separation of autonomous motion from autonomous decision. The iron horse is fearsome only if it may, at any time, choose to leave the tracks and come after you like a charging rhinoceros. Once the system of thought had shaken down in such a way that autonomous motion did not imply the capacity for decision, people made peace with the locomotive.

The computer is similarly ambiguous. It is clearly an inanimate machine. Yet we interact with it through language; a medium heretofore restricted to communication with other people. To be sure, computer languages are very restricted, but they are languages. They have words, punctuation marks, and syntactic rules. To learn to program computers we must extend our mechanisms for natural language.

Back then the question was mostly an academic one. That is to say, it had little purchase on the daily lives of most people. Consequently, however intently a relatively small cadre of academics debated the question, it was of relatively little interest to ordinary people.

That changed quite dramatically when, late in November 2022, OpenAI released ChatGPT on the web where anyone with an internet account and a web browser to access it and play with it. Overnight millions did so. The question of whether or not this thing was dead or alive, that is inanimate or animate, mindless or conscious, impressed itself on millions of users. It was no longer an academic question. It was a live question, and to some it was even existential: How long before this, this, this THING, goes rogue and destroys us?

Fortunately, that has not happened. We are all alive to debate the issue. And we do so, using terms that existed long prior to the release of ChatGPT. That’s a problem.

Back in the days when the questions of computer intelligence, of computational minds, and of artificial consciousness were academic, we had no examples of devices whose behavior was phenomenologically problematic. Computers played an inferior game of chess, though that ended in 1997 when IBM’s Deep Blue defeated Gary Kasparov, and were at best halting, clumsy, and relentlessly stupid with language. You could take whatever position you wished about the possibility of artificial intelligence (AI), artificial general intelligence (AGI), a term coined early in the millennium, or even superintelligence, a term popularized by Nick Bostrom’s 2013 book of that title, when it came to actual devices, it was clear that they were not intelligent or conscious.

ChatGPT could “talk,” just like a human being, or so much so that one had to work hard to find a meaningful difference. Many users proceeded as though there were no meaningful difference. Now the question of AI, AGI, or even superintelligence has taken on a different valence. Any chatbot “knows” a wider range of subjects than even the most brilliant and learned of humans. In that specific and limited sense, these things are superintelligent. While no one, so far as I know, has claimed that these chatbots are superintelligent in the fullest sense (as in Bostrom’s book, Superintelligence) you see the problem. Don’t you?

Just as we don’t know how the human mind, the human brain, works. So we don’t know how these chatbots, these large language models (LLMs) work. Do they work like we do or not? At some level obviously not. Computer hardware is quite different from biological “wetware” (brains). But when we consider function, that we don’t know. As long as we stick to symbolic behavior, the ability to write natural language, and increasingly to speak it, to write computer code, and to worth with mathematics, our ability to distinguish the real, that is, humans, from the artificial, that is, computers, is problematic. Thus the terms, the concepts, we have inherited from the pre-GPT-3 era are no longer adequate to problems we now face.

That is what makes the question of computer intelligence both so urgent and so deeply problematic. For the moment, I’m fond of a formulation Steven Harnad expressed somewhere on the web: The behavior chatbots exhibit is astonishing when you consider the fact that they don’t understand another. Those words are mine, but the thought is Harnad’s. This formulation, however, is no more than a stop gap.

We need new concepts, and new conceptual framework. That’s easier called for than accomplished. The accomplishment will take an intellectual generation.

* * * * *

I’ve made this general point, but at greater length and in different terms in a paper I finished in January of 2023: ChatGPT intimates a tantalizing future; its core LLM is organized on multiple levels; and it has broken the idea of thinking.

Wednesday, August 16, 2023

GPT4 still challenged by artithmetic

Abstract for the linked paper, Faith and Fate: Limits of Transformers on Compositionality:

Transformer large language models (LLMs) have sparked admiration for their exceptional performance on tasks that demand intricate multi-step reasoning. Yet, these models simultaneously show failures on surprisingly trivial problems. This begs the question: Are these errors incidental, or do they signal more substantial limitations? In an attempt to demystify Transformers, we investigate the limits of these models across three representative compositional tasks—multi-digit multiplication, logic grid puzzles, and a classic dynamic programming problem. These tasks require breaking problems down into sub-steps and synthesizing these steps into a precise answer. We formulate compositional tasks as computation graphs to systematically quantify the level of complexity, and break down reasoning steps into intermediate sub-procedures. Our empirical findings suggest that Transformers solve compositional tasks by reducing multi-step compositional reasoning into linearized subgraph matching, without necessarily developing systematic problem- solving skills. To round off our empirical study, we provide theoretical arguments on abstract multi-step reasoning problems that highlight how Transformers’ performance will rapidly decay with increased task complexity.

Monday, February 13, 2023

Why LLMs are good at generating code

By comparison, see my old working paper, PowerPoint Assistant: Augmenting End-User Software through Natural Language Interaction (2013). Here's the abstract:

This document sketches a natural language interface for end user software, such as PowerPoint. Such programs are basically worlds that exist entirely within a computer. Thus the interface is dealing with a world constructed with a finite number of primitive elements. You hand-code a basic language capability into the system, then give it the ability to ‘learn’ from its interactions with the user, and you have your basic PPA.

Sunday, August 7, 2022

AGI as shibboleth, symbols [reacting to Jack Clark]

Jack Clark has a LONG tweet stream on AI policy. Though I don’t agree with every tweet – would anyone? – it’s worth at least a quick look. I want to comment on two of the tweets.

AGI as shibboleth, and beyond

Has AGI ever been anything other than a shibboleth? I believe the term was coined in the 1990s because some researchers felt that AI had become stale and focused on specialized domains, so-called “narrow” AI. The phrase “artificial general intelligence” (AGI) was as a banner under which to revive the founding goal of AI, to construct the artificial equivalent of human intelligence.

What researchers actually construct are mechanisms. But no one knows how to specify a mechanism or set of mechanisms for AGI. Oh, sure, there’s the Universal Turing machine which can, in point of abstract theory, compute any computable function. It may be a mechanism, but the idea so abstract that it provides little to no guidance in the construction of computer systems.

AGI, like AI before it, is an abstract goal, a beacon, without a procedure that will lead to it. No matter how vigorously you chase over the surface of the earth for the North Star, you’re never going to get there. And so AGI simply functions as a shibboleth. If you want into the club, you have to pledge allegiance to AGI.

But you don’t need to pledge allegiance in order to construct interesting and even useful systems. So why invent this unreachable goal? Is it just to define a club?

Meanwhile I’ve written a paper in which I define the idea of an artificial mind. I begin by defining mind:

A MIND is a relational network of logic gates over the attractor landscape of a partitioned neural network. A partitioned network is one loosely divided into regions where the interaction within a region is (much) stronger than the interactions between regions. Each of these regions will have many basins of attraction. The relational network specifies relations between basins in different regions.

Note that the definition takes the form of specifying a mechanism involving logic gates and a neural network. Given that:

A NATURAL MIND is one where the substrate is the nervous system of a living animal.

And:

An ARTIFICIAL MIND is one where the substrate is inanimate matter engineered by humans to be a mind.

There are other definitions as well as some caveats and qualifications.

However, those definitions come after 50 pages of text and diagrams in which I lay out the mechanisms that support those definitions. The paper is primarily about the human brain, but one can imagine constructing artificial devices that meet those specifications. Now, whether those specifications are the right specifications, that’s open for discussion. However that discussion turns out, it is a discussion about mechanisms, not myths and magic.

The paper:

Relational Nets Over Attractors, A Primer: Part 1, Design for a Mind, Version 2, Working Paper, July 13, 2022, pp. 76, https://www.academia.edu/81911617/Relational_Nets_Over_Attractors_A_Primer_Part_1_Design_for_a_Mind

Ah, symbols

Here’s a twofer:

It's the first tweet that interests me, but let’s look Richard Sutton’s bitter lesson. Here’s his opening paragraph:

The biggest lesson that can be read from 70 years of AI research is that general methods that leverage computation are ultimately the most effective, and by a large margin. The ultimate reason for this is Moore's law, or rather its generalization of continued exponentially falling cost per unit of computation. Most AI research has been conducted as if the computation available to the agent were constant (in which case leveraging human knowledge would be one of the only ways to improve performance) but, over a slightly longer time than a typical research project, massively more computation inevitably becomes available. Seeking an improvement that makes a difference in the shorter term, researchers seek to leverage their human knowledge of the domain, but the only thing that matters in the long run is the leveraging of computation. These two need not run counter to each other, but in practice they tend to. Time spent on one is time not spent on the other. There are psychological commitments to investment in one approach or the other. And the human-knowledge approach tends to complicate methods in ways that make them less suited to taking advantage of general methods leveraging computation. There were many examples of AI researchers' belated learning of this bitter lesson, and it is instructive to review some of the most prominent.

Sutton then goes on to list domains where there has proven so: chess, Go, speech recognition, and computer vision. He then draws some conclusions, which I want to bracket.

Note, however, that Sutton talks of researchers seeking “to leverage their human knowledge of the domain.” Is that what’s going on symbolic AI? Perhaps in expert systems, which may have been the most pervasive practical result of GOFAI. But I don’t think that’s an accurate general characterization. That’s not what was going on in computational linguistics, for example, or in much of the work on knowledge representation. That research was based on the belief that much of human knowledge is inherently symbolic in character and therefore that we must create models that capture that symbolic character.

Why did those models collapse? I think there are several factors involved:

1. Combinatorial explosion: Symbolic systems tend to generate large numbers of alternative with little or no way of choosing among them.

2. Hand coding: Symbolic systems have to be painstakingly hand-coded, which takes time.

3. Too many models, difficult to choose among them: This exacerbates the hand-coding problem.

4. Common sense has proven elusive: But then it has proven elusive for deep learning as well.

Perhaps the first problem can be solved through more computing power, though exponential search can easily outstrip the addition of CPU cycles and memory. The third problem is one for science, and is, I believe, entangled with the fourth one. The second problem is inconvenient, but, alas, if hand-coding is necessary, then it’s necessary. But perhaps if we’re clever....

On the fourth one, here’s what I said in my GPT-3 paper:

A lot of common-sense reasoning takes place “close” to the physical world. I have come to believe, but will not here argue, that much of our basic (‘common sense’) knowledge of the physical world is grounded in analogue and quasi-analogue representations. This gives us the power to generate language about such matters on the fly. Old school symbolic machines did not have this capacity nor do current statistical models, such as GPT-3.

Thus the problem is not specific to symbolic systems. It is quite general. It’s not at all clear that we can deal with this problem without having robots out and about in the world. I note that the working paper I mentioned in the previous section, Relational Nets Over Attractors, is about constructing symbolic structures over quasi-analog representations, which, following the terminology of Saty Chary, I characterize as structured physical systems.

Let’s return to Sutton’s paper. Here’s his final paragraph:

The second general point to be learned from the bitter lesson is that the actual contents of minds are tremendously, irredeemably complex; we should stop trying to find simple ways to think about the contents of minds, such as simple ways to think about space, objects, multiple agents, or symmetries. All these are part of the arbitrary, intrinsically-complex, outside world. They are not what should be built in, as their complexity is endless; instead we should build in only the meta-methods that can find and capture this arbitrary complexity. Essential to these methods is that they can find good approximations, but the search for them should be by our methods, not by us. We want AI agents that can discover like we can, not which contain what we have discovered. Building in our discoveries only makes it harder to see how the discovering process can be done.

I’m hesitant to think of symbol systems as being “simple ways to think about the contents of minds.” That strikes me as rhetorical overkill. But Sutton is right about “the arbitrary, intrinsically-complex, outside world.” He says that “we should build in only the meta-methods that can find and capture this arbitrary complexity.” Well, sure, why not?

But are we doing that now? That’s not at all obvious to me. it seems likely to me that the DL community is hoping that they’ve discovered the metamethods, or will do so in the near future, and so we don’t have to think about what’s going on inside either human minds or the machines we’re building. Well, if human minds use symbols, and it seems all but self-evident that we do – if language isn’t a symbol system, what is? – then the current repertoire of DL methods is not up to the task.

What meta-methods are needed to detect patterns of symbolic meaning and construct those quasi-analog representations?

My GPT-3 paper:

GPT-3: Waterloo or Rubicon? Here be Dragons, Version 4.1, Working Paper, May 7, 2022, 38 pp., https://www.academia.edu/43787279/GPT_3_Waterloo_or_Rubicon_Here_be_Dragons_Version_4_1

Thursday, June 30, 2022

Steven Pinker and Scott Aaronson debate scaling

Scott Aaronson has hosted Steven Pinker to a discussion at Shtetl-Optimized.

Pinker on AGI:

Regarding the second, engineering question of whether scaling up deep-learning models will “get us to Artificial General Intelligence”: I think the question is probably ill-conceived, because I think the concept of “general intelligence” is meaningless. (I’m not referring to the psychometric variable g, also called “general intelligence,” namely the principal component of correlated variation across IQ subtests. This is a variable that aggregates many contributors to the brain’s efficiency such as cortical thickness and neural transmission speed, but it is not a mechanism (just as “horsepower” is a meaningful variable, but it doesn’t explain how cars move.) I find most characterizations of AGI to be either circular (such as “smarter than humans in every way,” begging the question of what “smarter” means) or mystical—a kind of omniscient, omnipotent, and clairvoyant power to solve any problem. No logician has ever outlined a normative model of what general intelligence would consist of, and even Turing swapped it out for the problem of fooling an observer, which spawned 70 years of unhelpful reminders of how easy it is to fool an observer.

If we do try to define “intelligence” in terms of mechanism rather than magic, it seems to me it would be something like “the ability to use information to attain a goal in an environment.” (“Use information” is shorthand for performing computations that embody laws that govern the world, namely logic, cause and effect, and statistical regularities. “Attain a goal” is shorthand for optimizing the attainment of multiple goals, since different goals trade off.) Specifying the goal is critical to any definition of intelligence: a given strategy in basketball will be intelligent if you’re trying to win a game and stupid if you’re trying to throw it. So is the environment: a given strategy can be smart under NBA rules and stupid under college rules.

Since a goal itself is neither intelligent or unintelligent (Hume and all that), but must be exogenously built into a system, and since no physical system has clairvoyance for all the laws of the world it inhabits down to the last butterfly wing-flap, this implies that there are as many intelligences as there are goals and environments. There will be no omnipotent superintelligence or wonder algorithm (or singularity or AGI or existential threat or foom), just better and better gadgets.

Aaronson responds:

Basically, one side says that, while GPT-3 is of course mind-bogglingly impressive, and while it refuted confident predictions that no such thing would work, in the end it’s just a text-prediction engine that will run with any absurd premise it’s given, and it fails to model the world the way humans do. The other side says that, while GPT-3 is of course just a text-prediction engine that will run with any absurd premise it’s given, and while it fails to model the world the way humans do, in the end it’s mind-bogglingly impressive, and it refuted confident predictions that no such thing would work.

Though I’m with Pinker on the definition of AGI, I also take the second of the two positions Aaronson set forth, which is, I take it, Aaronson’s position while the first is Pinker’s position. That’s why I wrote GPT-3: Waterloo or Rubicon? Here be Dragons (Version 4.1).

Aaronson continues:

I freely admit that I have no principled definition of “general intelligence,” let alone of “superintelligence.” To my mind, though, there’s a simple proof-of-principle that there’s something an AI could do that pretty much any of us would call “superintelligent.” Namely, it could say whatever Albert Einstein would say in a given situation, while thinking a thousand times faster. Feed the AI all the information about physics that the historical Einstein had in 1904, for example, and it would discover special relativity in a few hours, followed by general relativity a few days later. Give the AI a year, and it would think … well, whatever thoughts Einstein would’ve thought, if he’d had a millennium in peak mental condition to think them.

If nothing else, this AI could work by simulating Einstein’s brain neuron-by-neuron—provided we believe in the computational theory of mind, as I’m assuming we do. It’s true that we don’t know the detailed structure of Einstein’s brain in order to simulate it [...]. But that’s irrelevant to the argument. It’s also true that the AI won’t experience the same environment that Einstein would have—so, alright, imagine putting it in a very comfortable simulated study, and letting it interact with the world’s flesh-based physicists. A-Einstein can even propose experiments for the human physicists to do—he’ll just have to wait an excruciatingly long subjective time for their answers. But that’s OK: as an AI, he never gets old.

Next let’s throw into the mix AI Von Neumann, AI Ramanujan, AI Jane Austen, even AI Steven Pinker—all, of course, sped up 1,000x compared to their meat versions, even able to interact with thousands of sped-up copies of themselves and other scientists and artists. Do we agree that these entities quickly become the predominant intellectual force on earth—to the point where there’s little for the original humans left to do but understand and implement the AIs’ outputs (and, of course, eat, drink, and enjoy their lives, assuming the AIs can’t or don’t want to prevent that)?

Eh. Now that I have an explicit definition of artificial minds, I have no need for a definition of artificial intelligence. While my primer (Relational Nets Over Attractors, A Primer: Part 1, Design for a Mind) is mostly about the human mind and the human brain, the fact that I was able to propose a substrate-neutral definition of “mind” has the side-effect that I can talk about artificial minds as mechanisms, not magic, to use Pinker’s formulation.

Aaronson also notes:

I should clarify that, in practice, I don’t expect AGI to work by slavishly emulating humans—and not only because of the practical difficulties of scanning brains, especially deceased ones. Like with airplanes, like with existing deep learning, I expect future AIs to take some inspiration from the natural world but also to depart from it whenever convenient. The point is that, since there’s something that would plainly count as “superintelligence,” the question of whether it can be achieved is therefore “merely” an engineering question, not a philosophical one.

That is consistent with the view I have articulated in the primer.

Aaronson has more to say, as does Pinker. As of this moment, the dialog has attracted 100 comments (including two from me). It’s worth exploring.

Thursday, May 19, 2022

Meaning and Semantics, Relationality and Adhesion

For some time now I’ve been making a distinction between meaning and semantics, where I use meaning as a function of intention, in the more-or-less standard philosophical sense of intention as “aboutness.” When I talk of semantics I am talking about the elements in the language system. I have now decided that semantics has two aspects: relationality and adhesion.

I suppose we can think of meaning as inhering in the intentional relationship between the person – and we are talking about human beings here, but we could be talking about animals or maybe, just maybe, artificial minds – and the world. Walter Freeman regarded meaning as inherent in the total trajectory of the brain’s state during some experience, whether conversation, reading a text, or out and about in the world. There is more to the brain’s state than the operations of the language and cognitive systems. Thus meaning is necessarily different from semantics.

The standard philosophical arguments (such as Searle’s Chinese room) about artificial intelligence (which I’m now calling artificial minds), focus on meaning and intention to the utter neglect of semantics, as though it doesn’t exist. It may well be the case that all these artificial systems fail on the grounds of intentionality. It seems to me that the success of this line of argument is also a pyrrhic victory, for it leaves the philosopher powerless to reason about what these systems can do. It leads to a false binary where either the system is a human or it is a worthless artifact.

But that’s an aside. I’m much interested in that philosophical argument at the moment. I’m intersted in semantics, with its aspects of adhesion and relationality. Roughly speaking, adhesion is what ‘connects’ a concept to the world through perception. If we use a standard semantic network diagram (below), adhesion is carried on the REP (represent) arcs.

Relationality is carried on the arcs linking concepts with one another. Thus VAR in the diagram is for variety; beagles and collies are varieties of dog. We can also think of adhesion as being about compression (data reduction) and categorization – Gärdenfors’ dimensions in concept spaces. Relationality is about relations between objects in different concept spaces. But that’s only a rough characterization.

Large language models, such as GTP-3, are exploiting semantic relationality – the argument I made in my GPT-3 working paper, but have no access to adhesion. Vision systems are gounded in adhesion and may also exploit aspects of relationality.

[If we use Powers’s notion of intensities, where perception and cognition have to account for incoming intensities, then adhesion is about compression of intensities while relationality is about distribution of them over different concepts.]

More later.

Thursday, May 5, 2022

Artificial Minds, Day 3: Do I Still Believe It?

Two days ago I posted about the need for a Gestalt Switch: From Artificial Intelligence to Artificial Minds. I argued that intelligence is best conceived as a performance measure while a mind “is something that is implemented in some kind of device, to use a fairly generic term.” I certainly believe just that much. But that alone doesn’t justify speaking of these artificial devices, these computer systems, as minds, does it?

I believe it does, providing these devices meet the definition I have given for a mind. That’s what this post is about.

* * * * *

I have always maintained that we will never “build a machine to equal the human brain,” as I recounted in a recent post (Apr. 5, 2022). I still believe that. I have never, or at least not that I can remember, believed that human minds are the only minds that exist. Why can’t we say that animals have minds? I see no reason why not, though I understand that others do. The same goes for consciousness. Are animals conscious? Yes.

Let’s look at the definition of mind I gave in that post:

A MIND is a relational network of logic gates over the attractor landscape of a partitioned neural network. A partitioned network is one loosely divided into regions where the interaction within a region is (much) stronger than the interactions between regions. Each region compresses and categorizes its inputs, with each category having its own basin of attraction, and sends the results to other regions. Each region will have many basins of attraction. The relational network specifies relations between basins in different regions.

There is nothing in there that wouldn’t apply to many, most, all(?), animals. It’s possible that the brain of C. elegans, with its 302 neurons, doesn’t meet that definition. There might well some minimum number of neurons required to cross the threshold to mind. I’m not prepared to speculate about what that number, and in what architecture, would be.

Here's a passage from a note I sent to an old friend and colleague the other day:

The complexity literature is vast and I certainly don’t know it well, but in what I’ve seen, I’ve never seen anyone talking about a system consisting of multiple interconnected attractor landscapes. As far as I know, Walter Freeman never did so, and he spent a career investigating the complex dynamics of the nervous system. I’m talking about I don’t know how many landscapes – 150 to 1000, who knows – each with tens or 100s of thousands of attractor basins. Each thing you can identify visually has its own basin. Each word has a half dozen basins associated with it. And so forth. It’s crazy, but then how could an account of the brain not be crazy?

That’s what the numbers game is about. But if a machine, a computer, can meet the terms of the definition, then, yes, it has a mind.

I don’t know whether any of the current deep learning systems meet the terms of the definition. They are very different from organic wetware in many respects. The one that comes most readily to mind is that they require two phases of operation: learning and inference. The system that does the learning has to be much larger than the one that does the inference. I don’t quite know how to handle that, but perhaps what matters for the definition is the inference system.

I can easily imagine that the current inference engines do not meet the terms of the definition. If so, however, that is as likely to be a matter of architecture as of size. But then it may also be the case that this two-phase existence is an automatic fail. But it’s not obvious to me that the two-phase existence is inherent in digital technology. It is recognized as a problem and work is being done to move beyond it.

So, for the moment my position is that, yes, artificial minds are possible, but it is an open question whether or not any current systems meet the terms of the definition.

What about consciousness? I see no reason to deny consciousness to animals. Am I willing to extend it to machines that meet the terms of the definition I’ve offered for MIND? That is a very interesting question, one that I haven’t thought about until just now. But I do not see that coming up with an answer is an urgent matter. It can wait.

Tuesday, May 3, 2022

Gestalt Switch: From Artificial Intelligence to Artificial Minds

For many, both pro AI and con, the difference between the two terms – artificial intelligence and artificial minds – is of relatively little consequence, a matter of emphasis, perhaps. I take a different view. I believe that we are ready to switch from  thinking about the duck of AI to thinking about the rabbit of artificial minds.

Intelligence as a Measure of Performance

Intelligence is best conceived as a measure of performance. In the case of IQ it is the performance of a human mind. But one can also measure the performance of artificial devices of all kinds. For example, acceleration from 0 to 60 mph is a standard way of measuring the performance of automobiles. We also meaure the performance of computational systems of all kinds. Both artificial intelligence and computational linguistics have developed many measures of performance.

Talk of building an AI device is thus a category error, and talk of human-level AI, or AGI, simply compounds the error. We can imagine constructing many different kinds of devices intended to perform tasks heretofore performed only by the human mind. Indeed, we have done so, and have, for the most part, done it in the name of artificial intelligence. But we have yet to formulate a coherent account of just what this intelligence-thing is, of how it is constructed and operates, much less a human-level intelligence-thing.

It’s time we give it up. Let’s figure out what a mind is and then figure out how to build one of those.

For many, both pro AI and con, I suppose that suggestion is rather like leaping from the frying pan into the fire. Are you crazy? No, I’m not. But as you might expect, I do take a different view. I offer an explicit definition of mind and proceed from there. Are you crazy? No, I’m not. But I am a fearless speculator.

Minds, Both Natural and Artificial, a Speculative Definition

An artificial mind is something that is implemented in some kind of device, to use a fairly generic term. The artificial mind emerges from the operation of the device. Can we build such a device?

I’m not sure. I don’t know whether or not any of the devices we’ve built in the name of artificial intelligence qualify as minds in the sense I am about to propose, but some of the more recent machine learning devices might. I’ve certainly been thinking a lot about them recently, e.g. GPT-3: Waterloo or Rubicon? Here be Dragons, Version 4, Working Paper, April 26, 2024, https://www.academia.edu/43787279/GPT_3_Waterloo_or_Rubicon_Here_be_Dragons_Version_4.

I propose to define mind as follows:

A MIND is a relational network of logic gates over the attractor landscape of a partitioned neural network. A partitioned network is one loosely divided into regions where the interaction within a region is (much) stronger than the interactions between regions. Each region compresses and categorizes its inputs, with each category having its own basin of attraction, and sends the results to other regions. Each region will have many basins of attraction. The relational network specifies relations between basins in different regions.

Yes, I know, there’s a fair amoung of quasi-technical jargon there. I don’t intend to unpack it here and now. I’m working on thant off-line. A bit later I will, however, say a few words about where that came from.

Right now I offer some CAVEATS:

The underlying neural network, whether real or artificial, must meet certain minimal conditions. Those conditions, I assume, would have to do with:

  1. the SIZE of the network,
  2. its internal DIFFERENTIATION, and
  3. the nature of its ACCESS TO THE WORLD external to it.

I assume that a proper theory would address each of those issues, and others as well. Such a theory would of course be subject to empirical verification. I also assume that, to some extent, empirical investigation would precede such a theory and thus would contribute to its development.

Some further definitions:

A NATURAL MIND is one where the substrate is the nervous system of a living animal.

An ARTIFICIAL MIND is one where the substrate is inanimate matter engineered by humans to be a mind.

A SOCIAL MIND is capable of fluid interaction with natural or artificial minds, or both. That interaction could be mediated by language, but perhaps by music as well, or other means.

For the moment I take those as self-evident, but also as provisional, as is the definition of mind.

My point is that these definitions specify DEVICES of some kind. They tell us something about how such devices are constructed.

Where’d This Come From?

I have been working on such things for a long time, my whole career. This is not the place to recount that story. You can find the most recent version in this post, On the Differences between Artificial and Natural Minds: Another version of my intellectual biography, with specific comments on recent weeks in this post, What I've been up to in the last two weeks, across the Continental Divide and on to the Pacific [how the mind works]. What I would like to do here is outline the intellectual genealogy behind the terms in my basic definition of mind.

Here’s the first sentence, which contains the basic definition – the other sentences comment on it: A MIND is a relational network of logic gates over the attractor landscape of a partitioned neural network. I first learned about relational networks from David Hays in the linguistics department at SUNY Buffalo, where I went to graduate school. They were in common use in the cognitive sciences for characterizing mental structures and can been seen as deriving, at least in part, from associationist psychology.

Such networks commonly used the nodes of the network to represent conceptual objects and the arcs or edges of the network to represent links between those objects. Sydney Lamb developed networks in which the nodes were logical operators and the content of the network was carried on the arcs. Lamb was specifically inspired by the nervous system and that’s why I chose his notation, though I interpret it differently than he did.

That brings us to the second clause of the definition: ... over the attractor landscape of a partitioned neural network. The notion of an attractor landscape comes from complexity theory of various kinds which I picked up in various places over the years. The idea of a partitioned neural network is simply my way of talking about the fact that the neocortex is divided in various regions that seem to be functionally distinct, though the borders between these regions are not distinct. Let us assume that interaction within a region is (much) stronger than the interactions between regions. If that weren’t the case it’s hard to see the point of differentiating between the regions.

That is, the neocortex is partitioned into functionally specialized regions: Each region will have many basins of attraction. The notion of basins of attraction comes from complexity theory. I was particularly influenced, however, by the work of the late Walter Freeman on the complex dynamics of the cortex. He did a lot of work on the olfactory cortex (mostly of rats and rabbits I believe), where showed that each odor corresponds to a specific basin of attraction. When a new ordor is learned, a new basin is added to the landscape, which is reconfigured as a result.

And that brings us back to those logic gates. When a rat recognizes an ordor, its olfactory cortext enters the basic of attraction associated with that odor. It can enter only one basin at a time. The basins thus stand in an competitive relationship with one another, logical OR. Now, think about some cortical patch that receives inputs from, say, three other cortical regions. Its attractor landscape thus reflects the interation between outputs from those regions. That corresponds to logical AND. Those are the two basic building blocks of Lamb’s network: The relational network specifies relations between basins in different regions.

There’s one final bit: Each region compresses and categorizes its inputs, with each category having its own basin of attraction, and sends the results to other regions. Back in the 1970s, when I was working with him, David Hays wrote a book, Cognitive Structures (1981). He talked of the parameters of perception as mediating between sensorimotor activity and cognition proper. Though he didn’t put it in these terms, those parameters are an aspect of the compression and categorizing process. More recently Peter Gärdenfors has argued for Conceptual Spaces (2000) and The Geometry of Meaning (2014). Each conceptual space characterizes objects along several dimensions. Those dimensions also characterize the compression and categorizing process.

That’s it. Well, the core of it anyhow.

Do I Believe It?

Not at this time, no. But I don’t disbelieve it either. It’s too soon. I hold the idea in epistemic suspension, if you will.

Rather, I offer it as speculation to guide further investigation. The scientist can take these speculations as pointers for experiment and observation. The engineer can use them as guides for the design and fabrication of devices.

This domain is a complicated one. The scientist needs to think like an engineer, to reverse engineer the mind and brain, as a way of coming up with predictions for experimental verification. And the engineer needs to conjure up their inner scientist in order to investigate and observe the operations of the devices they construct.

I offer these ideas, this speculative engineering, in the hope and belief that they will prove useful in these endeavors.