Jimmy Ba, Geoffrey Hinton, Volodymyt Mnih, Joel Z. Leibo, Catalin Ionescu, Using Fast Weights to Attend to the Recent Past, arXiv:1610.06258v3 [stat.ML] 5 Dec 2016
Abstract: Until recently, research on artificial neural networks was largely restricted to systems with only two types of variable: Neural activities that represent the current or recent input and weights that learn to capture regularities among inputs, outputs and payoffs. There is no good reason for this restriction. Synapses have dynamics at many different time-scales and this suggests that artificial neural networks might benefit from variables that change slower than activities but much faster than the standard weights. These “fast weights” can be used to store temporary memories of the recent past and they provide a neurally plausible way of implementing the type of attention to the past that has recently proved very helpful in sequence-to-sequence models. By using fast weights we can avoid the need to store copies of neural activity patterns.
Imanol Schlag, Kazuki Irie, Jürgen Schmidhuber, Linear Transformers Are Secretly Fast Weight Programmers, arXiv:2102.11174v3 [cs.LG] 9 Jun 2021
Abstract: We show the formal equivalence of linearised self-attention mechanisms and fast weight controllers from the early ’90s, where a “slow” neural net learns by gradient descent to program the “fast weights” of another net through sequences of elementary programming instructions which are additive outer products of self-invented activation patterns (today called keys and values). Such Fast Weight Programmers (FWPs) learn to manipulate the contents of a finite memory and dynamically interact with it. We infer a memory capacity limitation of recent linearised softmax attention variants, and replace the purely additive outer prod- ucts by a delta rule-like programming instruction, such that the FWP can more easily learn to correct the current mapping from keys to values. The FWP also learns to compute dynamically changing learning rates. We also propose a new kernel function to linearise attention which balances simplicity and effectiveness. We conduct experiments on synthetic retrieval problems as well as standard machine translation and language modelling tasks which demonstrate the benefits of our methods.
Elias Najarro, Shyam Sudhakaran, Sebastian Risi, Towards Self-Assembling Artificial Neural Networks through Neural Developmental Programs, arXiv:2307.08197v1 [cs.NE]
Abstract: Biological nervous systems are created in a fundamentally different way than current artificial neural networks. Despite its impressive results in a variety of different domains, deep learning often requires considerable engineering effort to design high-performing neural architectures. By contrast, biological nervous systems are grown through a dynamic self-organizing process. In this paper, we take initial steps toward neural networks that grow through a developmental process that mirrors key properties of embryonic development in biological organisms. The growth process is guided by another neural network, which we call a Neural Developmental Program (NDP) and which operates through local communication alone. We investigate the role of neural growth on different machine learning benchmarks and different optimization methods (evolutionary training, online RL, offline RL, and supervised learning). Additionally, we highlight future research directions and opportunities enabled by having self-organization driving the growth of neural networks.
From the introduction:
The study of neural networks has been a topic of great interest in the field of artificial intelligence due to their ability to perform complex computations with remarkable efficiency. However, despite significant advancements in the development of neural networks, the majority of them lack the ability to self-organize, grow, and adapt to new situations in the same way that biological neurons do. Instead, their structure is often hand-designed, and learning in these systems is restricted to the optimization of connection weights.
Biological networks on the other hand, self-assemble and grow from an initial single cell. Additionally, the amount of information it takes to specify the wiring of a sophisticated biological brain directly is far greater than the information stored in the genome (Breedlove and Watson, 2013). Instead of storing a specific configuration of synapses, the genome encodes a much smaller number of rules that govern how to grow a brain through a local and self-organizing process (Zador, 2019). For example, the 100 trillion neural connections in the human brain are encoded by only around 30 thousand active genes. This outstanding compression has also been called the “genomic bottleneck” (Zador, 2019), and neuroscience suggests that this limited capacity has a regularizing effect that results in wiring and plasticity rules that generalize well.
In this paper, we take first steps in investigating the role of developmental and self-organizing algorithms in growing neural networks instead of manually designing them, which is an underrepresented research area (Gruau, 1992; Nolfi et al., 1994; Kow Aliw et al., 2014; Miller, 2014). Even simple models of development such as cellular automata demonstrate that growth (i.e. unfolding of information over time) can be crucial to determining the final state of a system, which can not directly be calculated (Wolfram, 1984). The grand vision is to create a system in which neurons self-assemble, grow, and adapt, based on the task at hand.
Towards this goal, we present a graph neural network type of encoding, in which the growth of a policy network (i.e. the neural network controlling the actions of an agent) is con- trolled by another network running in each neuron, which we call a Neural Developmental Program (NDP). The NDP takes as input information from the connected neurons in the policy network and decides if a neuron should replicate and how each connection in the network should set its weight. Starting from a single neuron, the approach grows a functional policy network, solely based on the local communication of neurons. Our approach is different from methods like NEAT (Stanley and Miikkulainen, 2002) that grow neural networks during evolution, by growing networks during the lifetime of the agent. While not implemented in the current NDP version, this will ultimately allow the neural network of the agents to be shaped based on their experience and environment.
I speculate that such a regime seems to be a way of eliminating the problem that current models cannot be readily modified once they have been trained. Rather, they must be completely retrained if one wishes to add new information to them. This represents a step toward polyviscosity.
That's (the stuff outside the brackets) is the title of a most interesting paper:
Mundt, M.; Pliushch, I.; Majumder, S.; Hong, Y.; Ramesh, V. Unified Probabilistic Deep Continual Learning through Generative Replay and Open Set Recognition. J. Imaging 2022,8,93. https://doi.org/10.3390/jimaging8040093
Abstract: Modern deep neural networks are well known to be brittle in the face of unknown data instances and recognition of the latter remains a challenge. Although it is inevitable for continual-learning systems to encounter such unseen concepts, the corresponding literature appears to nonetheless focus primarily on alleviating catastrophic interference with learned representations. In this work, we introduce a probabilistic approach that connects these perspectives based on variational inference in a single deep autoencoder model. Specifically, we propose to bound the approximate posterior by fitting regions of high density on the basis of correctly classified data points. These bounds are shown to serve a dual purpose: unseen unknown out-of-distribution data can be distinguished from already trained known tasks towards robust application. Simultaneously, to retain already acquired knowledge, a generative replay process can be narrowed to strictly in-distribution samples, in order to significantly alleviate catastrophic interference.
From the paper:
In particular, should we wish to apply and extend the system to an open world, where several other animals (and non animals) exist, there are two critical questions: (a) How can we prevent obvious mispredictions if the system encounters a new class? (b) How can we continue to incorporate this new concept into our present system without full retraining? With respect to the former question, it is well known that neural networks yield overconfident mispredictions in the face of unseen unknown concepts [3], a realization that has recently resurfaced in the context of various deep neural networks [4–6]. With respect to the latter question, it is similarly well known that neural networks, which are trained exclusively on newly arriving data, will overwrite their representations and thus forget encoded knowledge—a phenomenon referred to as catastrophic interference or catastrophic forgetting [7,8]. Although we have worded the above questions in a way that naturally exposes their connection: to identify what is new and think about how new concepts can be incorporated, they are largely subject to separate treatment in the respective literature. While open-set recognition [1,9,10] aims to explicitly identify novel inputs that deviate with respect to already observed instances, the existing continual learning literature predominantly concentrates its efforts on finding mechanisms to alleviate catastrophic interference (see [11] for an algorithmic survey).
‘A structured physical system has the necessary and sufficient means for specific intelligent response’. By structured physical system, I mean, an analog design, e.g. a Rube Goldberg apparatus, or a Braitenberg (!) vehicle, etc. This is in contrast to this: PSSH - Physical Symbol System Hypothesis - 'A physical symbol system has the necessary and sufficient means for general intelligent action'.
I would add the slide rule as another example. From Wikipedia:
The slide rule is a mechanical analog computer, which is used primarily for multiplication and division, and for functions such as exponents, roots, logarithms, and trigonometry. It is not typically designed for addition or subtraction, which is usually performed using other methods. Maximum accuracy for standard linear slide rules is about three decimal significant digits, while scientific notation is used to keep track of the order of magnitude of results. [...]
At its simplest, each number to be multiplied is represented by a length on a pair of parallel rulers that can slide past each other. As the rulers each have a logarithmic scale, it is possible to align them to read the sum of the numbers' logarithms, and hence calculate the product of the two numbers.
My father used a slide rule for his entire career as an engineer. I learned to use one in my teens – everyone did back that – but never had any use for one. They’ve been replaced by cheap electronic calculators and PCs.
But that’s a digression. It’s the more general Structured Physical System Hypothesis that interests me, as Saty has been argued that the brain is such a system I agree (see this recent post, Once more around the merry-go-round: Is the brain a computer?). Here’s my reply to Saty:
Hi, Saty. I like your hypothesis – Structured Physical System Hypothesis (SPSH) – a lot. I think that a lot about the brain is consistent with it. For example, we know that mappings from one area to another as we move from the sense organs (or muscles), to the subcortex, and into the cortex tend to preserve topological relations between neurons. That’s of obvious value in the visual and motor systems. But there are subtleties. In the cortical visual system we have a so-called What-system and a so-called Where-system beyond the primary cortex. The What-system tracks location in space while the Where-system identifies objects. I assume that the Where-system has links to the hippocampus and I’d expect it can deal with both ego-centric and geocentric coordinates (this must be in the literature). But I’d think the What-system deals in object-centered coordinates. And so forth and so on.
In thinking about Freeman’s results (HERE and HERE), and others, I’ve coined a phrase: polyviscous connectivity. Thus I say that the cortical network as a whole exhibits polyviscous connectivity. What do I mean? Some connections are highly resistant to change, and thus have high viscosity. Others change quite readily, and have low viscosity. There is a literature on long-term (LTP) and short-term potentiation (STP) of neural connectivity that is certainly relevant here, but I’ve not looked at it in quite a while.
Consider Freeman’s results. He’s measuring neural activity with an 8 by 8 array of electrodes mounted on the cortical surface. They’re going to detect activity of neurons at varying levels of viscosity. Let’s a assume that the patterns of connectivity encoding odorants that rat already recognizes have a relatively high viscosity. Let’s further assume that the neurons most susceptible to learning new odorants have a relatively low viscosity.
Once they’ve formed a stable response to the new odorant, that will result in a new pattern of neural activity for the ensemble. But it is also going to change the patterns exhibited by already learned odorants even though the high-viscosity connections haven’t changed. The high viscosity connections maintain the overall integrity of the ensemble. In time, if the new odorant continues to be encountered, the connections registering it will increase in viscosity. So, polyviscous connectivity allows a structured connectionist physical system to maintain its overall integrity while adding new items to its repertoire.
More later.
* * * * *
Note: Looking around on my hard-drive I found an article which is about what I have called polyviscosity, though it doesn’t use that term:
Poonam Mishra and Rishikesh Narayanan, Stable continual learning through structured multiscale plasticity manifolds, Current Opinion in Neurobiology 2021, 70:51–63, https://doi.org/10.1016/j.conb.2021.07.009
Abstract: Biological plasticity is ubiquitous. How does the brain navigate this complex plasticity space, where any component can seemingly change, in adapting to an ever-changing environment? We build a systematic case that stable continuous learning is achieved by structured rules that enforce multiple, but not all, components to change together in specific directions. This rule-based low-dimensional plasticity manifold of permitted plasticity combinations emerges from cell type–specific molecular signaling and triggers cascading impacts that span multiple scales. These multiscale plasticity manifolds form the basis for behavioral learning and are dynamic entities that are altered by neuromodulation, metaplasticity, and pathology. We explore the strong links between heterogeneities, degeneracy, and plasticity manifolds and emphasize the need to incorporate plasticity manifolds into learning-theoretical frameworks and experimental designs.
* * * * *
Note: I’d previously been using the terms “hyperciscosity” or “hyperviscous”, and you’ll find them in my posts and notes on this topic going back to 2013 (when I wrote about From Associative Nets to the Fluid Mind). But I have reluctantly decided to coin a new term since “hyperviscosity” is already being used.
* * * * *
Addendum, 8.13.22: On the Structured Physical System Hypothesis, see this post where I feature remarks by Rodney Brooks, Has the computer metaphor for the mind run out of steam? As the title suggests, Brooks is wondering whether or not it makes sense to think about nervous systems in terms of computation. Thus he wonders:
Is information processing the right metaphor there? Or are control
theory and resonance and synchronization the right metaphor? We need
different metaphors at different times, rather than just computation.
Physical intuition that we probably have as we think about computation
has served physicists well, until you get to the quantum world. When you
get to the quantum world, that physical intuition about stuff and place
gets in the way.
So, viscosity. Honey is more viscous than, say, water. It flows more slowly, much more. What happens if you drop a lump of honey into a tumbler of water? It sinks to the bottom in a continuous lump and flattens out along the bottom. It will begin to diffuse into the water along the boundary, but I don’t know how long, if ever, it will take to mix completely. Now, put a stick down into the tumbler until it extends into the honey. Give is a stir or three, but no more. Now you’ll have gobs and threads of honey mixed in with water in a complex and somewhat irregular and ragged way. That’s a simple polyviscous fluid. It has regions of relatively high viscosity and other regions of relatively low viscosity. Now imagine a fluid with 5, 10, 27, 48, and more different levels of viscosity, from all but solid like cold tar through the wispiest whatever. Polyviscosity.
As the title of the year-old post indicates, I was thinking in terms of connectivity:
Thus I say that the cortical network as a whole exhibits polyviscous connectivity. What do I mean? Some connections are highly resistant to change, and thus have high viscosity. Others change quite readily, and have low viscosity.
OK. Now let’s shift our thinking just a bit and think of the mind as a polyviscous fluid. The mind, as the saying goes, is what the brain does. And that is very complex.
Imagine that you’re watching a movie, make it a Hong Kong martial arts movie. Your mind is entrained to the images on the screen. During a fight scene the level of mental viscosity is relatively how. The fight is over and the hero rests, contemplating the sunset, let’s say. The viscosity is somewhat higher.
Yet, while you’re entrained by the film, you’re not completely absorbed into. While the hero contemplates the sunset, you take a bit of popcorn. And maybe you were munching furiously during the fight. So, even as you were watching the film you slipped in some mental popcorn “frames” among the film frames. Very slippery, low viscosity.
When I wrote my book on music, Beethoven’s Anvil, I talked of the mind as neural weather. Thus (p. 72):
If the functional proclivities of a patch of neural tissue are not relevant for a current activity, those neurons will not be firing very often, but they will still generate some output. The only neuron that does not generate any output is a dead one. Neurons that are firing at low intensity one moment may well be recruited to more intense activity the next. As Walter Freeman has said, a low level of activity is still a means of participating in the evolving mental state.
The mind, in this view, is thus like the weather. The same environment can have very different kinds of weather. And while we find it natural to talk of weather systems as configurations of geography, temperature, humidity, air pressure etc., no overall mechanism regulates the weather. The weather is the result of many processes operating on different temporal and spatial scales.
At the global level and on a scale of millennia we have the long-term patterns governing the ebb and flow of glaciers which, in one commonly accepted theory, is a function of wobble and tilt in the earth’s spin axis and the shape of the earth’s orbit. At the global level and operating annually we have the succession of seasons, which is caused by the orientation of the earth with respect to the sun as it moves through the year. We can continue on, considering smaller and smaller scales until we consider the wind ripping through the twin towers of the World Trade Center or even the breeze coming in through your open window and blowing the papers off your desk.
Weather is regular enough that one can predict general patterns at scales of hours, days, and months, but not so regular that making such predictions is easy and routinely reliable. Above all, there is no central mechanism governing the weather. It just happens.
So, neural weather, polyviscous fluid. Perhaps we’re getting somewhere. The mind IS what the brain does, and what the brain does is complex and varies along a wide range of time scales. The brain’s overall physical structure is relatively constant throughout life, barring injury and disease. But the connectivity changes over a variety of time scales from seconds through hours and days and even longer (think of cortical plasticity). There is much, perhaps most, millisecond to millisecond, activity that produces no synaptic change at all. A very fluid phenomenon, over multimer time scales.
— Bill Benzon, BAM! Bootstrapping Artificial Minds (@bbenzon) December 27, 2022
From the article:
Future computer systems, said Hinton, will be take a different approach: they will be "neuromorphic," and they will be "mortal," meaning that every computer will be a close bond of the software that represents neural nets with hardware that is messy, in the sense of having analog rather than digital elements, which can incorporate elements of uncertainty and can develop over time.
"Now, the alternative to that, which computer scientists really don't like because it's attacking one of their foundational principles, is to say we're going to give up on the separation of hardware and software," explained Hinton.
"We're going to do what I call mortal computation, where the knowledge that the system has learned and the hardware, are inseparable."
These mortal computers could be "grown," he said, getting rid of expensive chip fabrication plants.
"If we do that, we can use very low power analog computation, you can have trillion way parallelism using things like memristors for the weights," he said, referring to a decades-old kind of experimental chip that is based on non-linear circuit elements.
"And also you could grow hardware without knowing the precise quality of the exact behavior of different bits of the hardware."
The new mortal computers won't replace traditional digital computers, Hilton told the NeurIPS crowd. "It won't be the computer that is in charge of your bank account and knows exactly how much money you've got," said Hinton.
"It'll be used for putting something else: It'll be used for putting something like GPT-3 in your toaster for one dollar, so running on a few watts, you can have a conversation with your toaster."*
Early on in my reading and studying neuroscience I read about the glia, brain cells between the neurons. Not much was known about them at the time and they seem not to have been much studied. With my recent interest in polyviscosity I decided to check up on the glia.
Things have changed. Quite a bit is now known about them. They seem to be crucial. One recent article:
Robertson JM. The Gliocentric Brain. Int J Mol Sci. 2018 Oct 5;19(10):3033. doi: 10.3390/ijms19103033. PMID: 30301132; PMCID: PMC6212929.
Abstract: The Neuron Doctrine, the cornerstone of research on normal and abnormal brain functions for over a century, has failed to discern the basis of complex cognitive functions. The location and mechanisms of memory storage and recall, consciousness, and learning, remain enigmatic. The purpose of this article is to critically review the Neuron Doctrine in light of empirical data over the past three decades. Similarly, the central role of the synapse and associated neural networks, as well as ancillary hypotheses, such as gamma synchrony and cortical minicolumns, are critically examined. It is concluded that each is fundamentally flawed and that, over the past three decades, the study of non-neuronal cells, particularly astrocytes, has shown that virtually all functions ascribed to neurons are largely the result of direct or indirect actions of glia continuously interacting with neurons and neural networks. Recognition of non-neural cells in higher brain functions is extremely important. The strict adherence of purely neurocentric ideas, deeply ingrained in the great majority of neuroscientists, remains a detriment to understanding normal and abnormal brain functions. By broadening brain information processing beyond neurons, progress in understanding higher level brain functions, as well as neurodegenerative and neurodevelopmental disorders, will progress beyond the impasse that has been evident for decades.
I take it, then, that the glia are central to consciousness, reorganization, and polyviscosity. And that’s only one article.
It seems to me that one effect of the computational view of neural function has been implicitly to encourage treating networks of neurons as passive switching networks that just happen to be constituted of living cells. But the fact that cells are living has been treated as contingent and not essential to their switching functions. A great deal of neuroscience and cognitive reads like this and AI even more so. For that matter, that’s more or less how I thought about matters for years. That had begun to change by September of 2014 when I wrote a post, What’s it mean, minds are built from the inside? Here’s three paragraphs:
If we want a computer to hold vast intellectual resources at its command, it’s going to have to learn them, and learn them from the inside, just like we do. And we’re not going to know, in detail, how it does it, any more than we know, in detail, what goes on in one another’s minds.
How do we do it? It starts in utero. When neurons first differentiate they are, of course, living cells and further differentiation is determined in part by the neurons themselves. Each neuron “seeks” nutrients and generates outputs to that end. When we analyze neural activity we tend to treat it, and its activities, as components of a complicated circuit in service of the whole organism. But that’s not how neurons “see” the world. Each neuron is just trying to survive.
Think of ants in a colony or bees in a swarm. There may be some mysterious coherence to the whole, but that coherence is the result of each individual pursuing its own purposes, however limited those purposes may be. So it is with brains and neurons.
So, I’ve been moving away from the passive-switching-network view for a while.
But it took that 1988 paper by Fodor and Pylyshyn (pp. 34-45):
Classical theories are able to accommodate these sorts of considerations because they assume architectures in which there is a functional distinction between memory and program. In a system such as a Turing machine, where the length of the tape is not fixed in advance, changes in the amount of available memory can be affected without changing the computational structure of the machine; viz by making more tape available. By contrast, in a finite state automaton or a Connectionist machine, adding to the memory (e.g. by adding units to a network) alters the connectivity relations among nodes and thus does affect the machine’s computational structure. Connectionist cognitive architectures cannot, by their very nature, support an expandable memory, so they cannot support productive cognitive capacities. The long and short is that if productivity arguments are sound, then they show that the architecture of the mind can’t be Connectionist. Connectionists have, by and large, acknowledged this; so they are forced to reject productivity arguments.
Jerry A. Fodor; Zenon W. Pylyshyn (1988). Connectionism and cognitive architecture: A critical analysis. Cognition, 28(1-2), 0–71. doi:10.1016/0010-0277(88)90031-5.
THAT focused my attention on the problem of memory as a physical process. And that, in turn, led me back to Walter Freeman, which I discuss in this post, Physical constraints on computing, process and memory, Part 1 [LeCun]. And that in turn let me to my thoughts about polyviscosity – though I’d initially used the term “hyperviscosity,” but have abandoned it because it is already in use.
So, memory presents a physical problem (Fodor and Pylyshyn). That problem thus requires a physical solution: polyviscosity. And polyviscosity, it would seem, requires living tissue. Perhaps we’ll figure out how to implement it in inanimate materials, but at the moment living tissue is what we’ve got. And glial cells are central to the mechanisms of polyviscosity.
I’m tempted to say something like – and here I’m rambling again – that the glia implement consciousness. And the function of consciousness is to ‘operate’ the neuromolecular mechanisms of reorganization. And if that seems a bit circular, well, that’s the best I can do at the moment. The important point is that consciousness, reorganization, polyviscosity, and the glia and involved in the same phenomena.
In Part 1 I connected polyviscosity with the account of consciousness and reorganization offered by William Powers in Behavior: The Control of Perception (1973). In Part 2 I rambled on about how neural ‘fluidity’ required 1) that neural circuity be ready for reorganization at any time, but that 2) commitment to reorganization was retrospective, and 3) that entire system had to undergo a global transformation in an instant in times of emergency (the so-called startle response). I also speculated about the active nature of the inter-cellular space in and around synapses. Now I want to talk about substrate independence.
Substrate independence is the idea that the mind (considered as an information processing system) can be implemented on any substrate in which the necessary functions can be realized. As I have previously suggested – see posts listed below – this implies a distinction between hardware and software that doesn’t (seem to) apply to natural nervous systems. That distinction also implies a symbolic computing regime in which computational process is physically separate from memory. That is not realistic either.
A mind, at least of a natural kind, would seem to imply a polyviscous substrate. Does that require living tissue? I don’t know but my default belief would be that it does. I keep saying that minds are organized from the inside (see post listed below). I’m thinking that requires that the constituent elements be living cells, not only neurons, but glial cells as well.
Of course, given a sufficiently detailed description of how neural circuitry works, it might be possible to simulate these processes. But that would be computationally expensive.
I note, as well, that those definitions did not take into account the requirements of polyviscosity and memory as I am coming to understanding them. That thinking is subsequent to finishing that paper. It might well have to be revised with those considerations in mind.
As I explained previously – Consciousness, reorganization and polyviscosity, Part 1 – William Powers identified consciousness with the process of reorganization, which is responsible for learning and the formation of memories in his model. Ever since I read that (in his 1973 Behavior: The Control of Perception) I have adopted that as my preferred account of consciousness, though I have read other accounts on principle. I like it because it follows naturally from his overall model, which David Hays and I adopted though with modifications, and seems to match the phenomenology of consciousness.
Yet I had one problem: It implied that reorganization is constant as long as we are awake. But Hays was comfortable with that, though I forget just how he phrased it, and so I accepted it. Now that I’m thinking of the nervous system as simultaneously operating on several levels of viscosity, that makes sense. (I know, a lousy formulation. I’m just making it up.)
Conscious is, in effect, “negotiating” between these different levels. We must be ready to change, to commit something to memory, at any moment. But we can’t always anticipate these changes. So connectivity is always a bit fluid. Some synapses are always in “change” mode, but the potential changes are not necessarily committed. Commitment represents a retrospective judgment.
This is consistent with a query I put to Walter Freeman some years ago: Does consciousness allow for ‘instantaneous’ global change, as in the startle reflex? Freeman thought that was a sensible suggestion.
Consciousness is what allows for mental “fluidity.” Neural circuits are not passive switches like ordinary electrical switches. They are active.
Experience can “flow” through us without setting any change in motion. The moment, however, that the experiential flow meets resistance, the flow becomes turbulent and that turbulence initiates reorganization.
My guess is – and that’s all it is, a guess – is that we are now dealing with molecular mechanisms operating at the synaptic level. These mechanisms maintain interneural space. Synapses are not just empty gaps. They are structured gates. The molecular mechanisms maintaining that structure give rise to polyviscosity.
And when we do take time off, we struggle to relax: A 2022 survey of over 20,000 professionals found that 54 percent of people said they weren’t sure they could fully “unplug from work” while taking paid time off.
This vacationing failure has consequences: Some research suggests that being a “work martyr” who doesn’t take time off or works through a vacation isn’t good for work performance. Of course, it’s not so great for people personally, either, increasing stress and the risk of burnout. Some companies are going so far as to mandate that employees take time off — a measure that is perhaps only necessary in a culture where people feel guilty for not being at work, and then feel ashamed of working when they should be relaxing.
Why is it difficult for professionals to relax from work? That’s a good question, a deep one. The next paragraph suggests an answer:
“Maybe we can blame the Puritans,” Emma Goldberg wrote recently in The Times. “Those settling in America in the 17th century thought idleness was sinful, and a six-day workweek sensible.”
People need a vacation. They always have. But especially when the office is closed, and work is what happens when you’re near your phone, which is to say every waking hour, employees need to recharge. Some are quietly asking permission to rest. Others know that their break is overdue, and now they’re getting nudges from the boss: log off.
I assume that “recharge” is a metaphor; we’re not talking about plugging into some kind of device that charges our, our what? Circuits? What circuits? Neural circuits perhaps?
I think there’s something there. It’s about the long-term maintenance of the brain, about behavioral-mode and long-term plasticity. Work requires a certain general configuration of neural ‘readiness’ if you will; over time that can ‘harden into brittleness’ and that’s not good. You need to get away from that configuration into a different one, one with different requirements and options. Just why, we don’t know. But, you know, I’m inclined to think it has something to do with the requirements of living in the early human environment, whatever and wherever that was. Here’s what Wikipedia says about life during the Paleolithic era:
Nearly all of our knowledge of Paleolithic human culture and way of life comes from archaeology and ethnographic comparisons to modern hunter-gatherer cultures such as the !Kung San who live similarly to their Paleolithic predecessors. The economy of a typical Paleolithic society was a hunter-gatherer economy.[24] Humans hunted wild animals for meat and gathered food, firewood, and materials for their tools, clothes, or shelters.[24]
Human population density was very low, around only 0.4 inhabitants per square kilometre (1/sq mi). This was most likely due to low body fat, infanticide, women regularly engaging in intense endurance exercise, late weaning of infants, and a nomadic lifestyle. Like contemporary hunter-gatherers, Paleolithic humans enjoyed an abundance of leisure time unparalleled in both Neolithic farming societies and modern industrial societies. At the end of the Paleolithic, specifically the Middle or Upper Paleolithic, humans began to produce works of art such as cave paintings, rock art and jewelry and began to engage in religious behavior such as burials and rituals.
That’s very different from the more sedentary lives of office workers. Does “recharge” mean something like “return to a more primitive way of life,” if only for a week or two? Why does the brain need that, if that’s what’s going on?
Let’s get back to Vanderkam’s article:
But is enforcing the binary between work and time off, sharpening those blurred boundaries, the only solution? I don’t think so. Particularly for those of us who enjoy our work, if you can’t — or, let’s face it, won’t — disconnect from work on vacation, let me assure you: It is probably OK. Work is no worse a way to spend vacation downtime than watching TV or perusing Instagram — and creative work can sometimes even be a welcome break from the chaos of a family vacation.
Ah, but watching TV or cruising Instagram aren’t a very good alternative to work mode. They’re more like work mode without doing any work.
Continuing on:
It is also OK, however, to take little vacations during working hours. An hour outside reading a novel, an afternoon bike ride, lunch with a friend, leaving the office (or desk at home) a little early to shop for and cook a special dinner: If you’re thoughtful and intentional about it, dispensing with strict boundaries between work and the rest of life can make a fuller, less burned-out life possible.
OK, but that’s presupposing the kind of behavioral flexibility that’s missing in work-conditioned lives. Notice that phrase, “thoughtful and intentional.” That implies the possibility of deliberate choice. A bit later Vanderkam notes:
When I did a time diary study in 2013 and 2014 of women who had professional jobs and kids at home, about half said they worked what I call a “split shift” — leaving work on the early side to spend time with their children, then doing work at home at night after the kids went to bed. Moving work around in terms of where and when it is done made it more possible for these women to have a big career and a meaningful family life. Men don’t talk about this time shifting as much, but some do it too.
Still later:
I hope we can begin to understand that, for many, work is a collection of tasks, not a collection of hours in a certain place. And time is a finite resource, but one that cannot always be neatly divided into “work time” and “free time.” Taking time for yourself during the work day doesn’t make you lazy, and working a bit on vacation doesn’t make you a workaholic. Dispensing with strict time boundaries should also mean ditching the guilt you might feel for either.
What’s the appropriate mixture of tasks? How do we engender the flexibility to move back and forth between them? What’s the proper childhood foundation? Of course I’m going to suggest that music is important –requiring, as it does, both focused technical practice to acquire and sharpen physical skill and wild-and-crazy jamming with friends – but that’s more than I can squeeze into this post.
Another observation from my own experience seems germane. Back in my days as a university faculty member, I noticed that I was not in a really good research frame of mind until three or four weeks after the Spring semester had ended—my brain had to have one set of modes to handle the academic routine of teaching and committee work and another set for intense thinking. Transition from one set of modes to the other took time [14].
Extended vacations may well afford a similar change in modal organization. One takes a month off from work and spends two weeks on safari in Africa; then boards a small sailing boat and island-hops in the Caribbean for a week, and concludes with a climb up El Capitan. With all that time away from work, the mind changes and we enter different modes of experience. Reading travel books, or novels, even the best, is quite different from going there. Physically restructuring the mind requires time and a steady regime of different sensations, desires, and acts. That “willing suspension of disbelief for the moment” (Coleridge, 1817, p. 6) through which the Rank 3 reader transports him/herself to another world is but a transition between currently available modes. What happens after days and weeks of exploration has a different quality.
The nature of consciousness is one of the big mysteries of contemporary thought. The best account of consciousness I know of is that offered by William Powers in Behavior: The Control of Perception (1973). That’s what this post is about. My objective is simple, to link Powers’s account of consciousness to the concept of polyviscosity that I offered about a week ago, The structured physical system hypothesis (SPSH), Polyviscous connectivity [The brain as a physical system]. Unfortunately, Powers’s concept, while basically simple, is simple only in the context of his overall model of mind, and that is not something that can readily be conveyed in a single blog post. Thus this post is mostly for my own benefit.
Powers’ model consists of two components: 1) a stack of servomechanisms – see the post In Memory of Bill Powers – regulating both perception and movement, and 2) a reorganizing system. The reorganizing system is external to the stack, but operates on it to achieve adaptive control, an idea he took from Norbert Wiener. Powers devoted “Chapter 14, Learning” to the subject (pp. 177-204). Reorganization is the mechanism through which Powers achieves learning.
Here’s an extensive passage that gets at the heart of the present matter (pp. 199-201):
To the reorganizing system, under these new hypotheses, the hierarchy of perceptual signals is itself the object of perception, and the recipient of arbitrary actions. This new arrangement, originally intended only as a means of keeping reorganization closer to the point, gives the model as a whole two completely different types of perceptions: one which is a representation of the external world, and the other which is a perception of perceiving. And we have given the system as a whole the ability to produce spontaneous acts apparently unrelated to external events or control considerations: truly arbitrary but still organized acts.
As nearly as I can tell short of satori, we are now talking about awareness and volition.
Awareness seems to have the same character whether one is being aware of his finger or of his faults, his present automobile or the one he wishes Detroit would build, the automobile’s hubcap or its environmental impact. Perception changes like a kaleidoscope, while that sense of being aware remains quite unchanged. Similarly, crooking a finger requires the same act of will as varying one’s bowling delivery “to see what will happen.” Volition has the arbitrary nature required of a test stimulus (or seems to) and seems the same whatever is being willed. But awareness is more interesting, somehow.
The mobility of awareness is striking. While one is carrying out a complex behavior like driving a car through to work, one’s awareness can focus on efforts or sensations or configurations of all sorts, the ones being controlled or the ones passing by in short skirts, or even turn to some system idling in the background, working over some other problem or musing over some past event or future plan. It seems that the behavioral hierarchy can proceed quite automatically, controlling its own perceptual signals at many orders, while awareness moves here and there inspecting the machinery but making no comments of its own. It merely experiences in a mute and contentless way, judging everything with respect to intrinsic reference levels, not learned goals.
This leads to a working definition of consciousness. Consciousness consists of perception (presence of neural currents in a perceptual pathway) and awareness (reception by the reorganizing system of duplicates of those signals, which are all alike wherever they come from). In effect, conscious experience always has a point of view which is determined partly by the nature of the learned perceptual functions involved, and partly by built-in, experience-independent criteria. Those systems whose perceptual signals are being monitored by the reorganizing system are operating in the conscious mode. Those which are operating without their perceptual signals being monitored are in the unconscious mode (or preconscious, a fine distinction of Freud’s which I think unnecessary).
This speculative picture has, I believe, some logical implications that are borne out by experience. One implication is that only systems in the conscious mode are subject either to volitional disturbance or reorganization. The first condition seems experientially self-evident: can you imagine willing an arbitrary act unconsciously? The second is less self-evident, but still intuitively right. Learning seems to require consciousness (at least learning anything of much consequence). Therapy almost certainly does. If there is anything on which most psychotherapists would agree, I think it would be the principle that change demands consciousness from the point of view that needs changing. Furthermore, I think that anyone who has acquired a skill to the point of automaticity would agree that being conscious of the details tends to disrupt (that, is, begin reorganization of) the behavior. In how many applications have we heard that the way to interrupt a habit like a typing error is to execute the behavior “on purpose”—that is, consciously identifying with the behaving system instead of sitting off in another system worrying about the terrible effects of having the habit? And does not “on purpose” mean in this case arbitrarily not for some higher goals but just to inspect the act, itself?
That, then, is consciousness as Powers conceives it. It is correlated with reorganization. If we are to reorganize a perception or action, we must be aware of it. The fact that we spend most of our lives in some state of consciousness implies that we are always learning or, perhaps, maintaining ourselves in a state of readiness to learn.
What has this to do with polyviscosity? Here I am thinking of neural connectivity. It is polyviscous in that some connections are highly resistant to change while others change readily. Reorganization, that is to say learning, requires that neural connectivity change. Connections of various levels of viscosity are likely to be intermingled in any given volume of cortical tissue.
Classical theories are able to accommodate these sorts of considerations because they assume architectures in which there is a functional distinction between memory and program. In a system such as a Turing machine, where the length of the tape is not fixed in advance, changes in the amount of available memory can be affected without changing the computational structure of the machine; viz by making more tape available. By contrast, in a finite state automaton or a Connectionist machine, adding to the memory (e.g. by adding units to a network) alters the connectivity relations among nodes and thus does affect the machine’s computational structure. Connectionist cognitive architectures cannot, by their very nature, support an expandable memory, so they cannot support productive cognitive capacities. The long and short is that if productivity arguments are sound, then they show that the architecture of the mind can’t be Connectionist. Connectionists have, by and large, acknowledged this; so they are forced to reject productivity arguments.
Physically, the nervous system appears to be connectionist in character. And so adding new items to the system is physically problematic. That’s the problem that is solved by polyviscous connectivity – see my post, Physical constraints on computing, process and memory, Part 1 [LeCun] (Note: in that post I use the term “hyperviscous” rather than “polyviscous”). Some connections must remain stable while others change. The stable connections maintain the overall structural integrity of the network while the changing connections introduce new items into that structure.
Here's a recent article that’s relevant, though it doesn’t use the term “polyviscious”: Poonam Mishra and Rishikesh Narayanan, Stable continual learning through structured multiscale plasticity manifolds, Current Opinion in Neurobiology 2021, 70:51–63, https://doi.org/10.1016/j.conb.2021.07.009
Abstract: Biological plasticity is ubiquitous. How does the brain navigate this complex plasticity space, where any component can seemingly change, in adapting to an ever-changing environment? We build a systematic case that stable continuous learning is achieved by structured rules that enforce multiple, but not all, components to change together in specific directions. This rule-based low-dimensional plasticity manifold of permitted plasticity combinations emerges from cell type–specific molecular signaling and triggers cascading impacts that span multiple scales. These multiscale plasticity manifolds form the basis for behavioral learning and are dynamic entities that are altered by neuromodulation, metaplasticity, and pathology. We explore the strong links between heterogeneities, degeneracy, and plasticity manifolds and emphasize the need to incorporate plasticity manifolds into learning-theoretical frameworks and experimental designs.
Kotseruba, I., Tsotsos, J.K. 40 years of cognitive architectures: core cognitive abilities and practical applications. Artif Intell Rev 53, 17–94 (2020). https://doi.org/10.1007/s10462-018-9646-y
Abstract: In this paper we present a broad overview of the last 40 years of research on cognitive architectures. To date, the number of existing architectures has reached several hundred, but most of the existing surveys do not reflect this growth and instead focus on a handful of well-established architectures. In this survey we aim to provide a more inclusive and high-level overview of the research on cognitive architectures. Our final set of 84 architectures includes 49 that are still actively developed, and borrow from a diverse set of disciplines, spanning areas from psychoanalysis to neuroscience. To keep the length of this paper within reasonable limits we discuss only the core cognitive abilities, such as perception, attention mechanisms, action selection, memory, learning, reasoning and metareasoning. In order to assess the breadth of practical applications of cognitive architectures we present information on over 900 practical projects implemented using the cognitive architectures in our list. We use various visualization techniques to highlight the overall trends in the development of the field. In addition to summarizing the current state-of-the-art in the cognitive architecture research, this survey describes a variety of methods and ideas that have been tried and their relative success in modeling human cognitive abilities, as well as which aspects of cognitive behavior need more research with respect to their mechanistic counterparts and thus can further inform how cognitive science might progress.
Laird, J. E., Lebiere, C., & Rosenbloom, P. S. (2017). A Standard Model of the Mind: Toward a Common Computational Framework across Artificial Intelligence, Cognitive Science, Neuroscience, and Robotics. AI Magazine, 38(4), 13-26. https://doi.org/10.1609/aimag.v38i4.2744
A standard model captures a community consensus over a coherent region of science, serving as a cumulative reference point for the field that can provide guidance for both research and applications, while also focusing efforts to extend or revise it. Here we propose developing such a model for humanlike minds, computational entities whose structures and processes are substantially similar to those found in human cognition. Our hypothesis is that cognitive architectures provide the appropriate computational abstraction for defining a standard model, although the standard model is not itself such an architecture. The proposed standard model began as an initial consensus at the 2013 AAAI Fall Symposium on Integrated Cognition, but is extended here through a synthesis across three existing cognitive architectures: ACT-R, Sigma, and Soar. The resulting standard model spans key aspects of structure and processing, memory and content, learning, and perception and motor, and highlights loci of architectural agreement as well as disagreement with the consensus while identifying potential areas of remaining incompleteness. The hope is that this work will provide an important step toward engaging the broader community in further development of the standard model of the mind.
This Common Model of Cognition divides humanlike thought into multiple modules, with a short-term memory module at the center of the model. The other modules – perception, action, skills and knowledge – interact through it.
Learning, rather than occurring intentionally, happens automatically as a side effect of processing. In other words, you don’t decide what is stored in long-term memory. Instead, the architecture determines what is learned based on whatever you do think about. This can yield learning of new facts you are exposed to or new skills that you attempt. It can also yield refinements to existing facts and skills.
The modules themselves operate in parallel; for example, allowing you to remember something while listening and looking around your environment. Each module’s computations are massively parallel, meaning many small computational steps happening at the same time. For example, in retrieving a relevant fact from a vast trove of prior experiences, the long-term memory module can determine the relevance of all known facts simultaneously, in a single step.
It’s that middle paragraph that caught my attention. Why? Because it is and isn’t true. Sure, a lot of learning is a side effect, as they say. Speaking is perhaps the classic example. But it is also the case that we do devote enormous effort to deliberate learning. That’s what happens in school. Just why they gloss over it is a mystery. However...
Automatic vs. deliberate learning
This speaks to the issues I raised in my recent post, Physical constraints on computing, process and memory, Part 1 [LeCun], where I was concerned with the distinction that Jerry Fodor and Zenon Pylyshyn made between “classical” theories of cognition where there is an explicit distinction between memory and program and connectionist accounts where memory and program are interwoven in one structure. Classical systems can easily acquire new knowledge by adding more memory; the structure of the program is unaffected. Connectionist systems are not like that.
To a first approximation the human nervous system seems to be a connectionist system. Each neuron seems to be both an active unit and a memory unit. There is no obvious division between a central processor, where all the programming resides, and a passive memory store. And yet, we learn, all the time we learn. How is that possible?
In that post I cited research by Walter Freeman on the sense of smell. It seems that when a new odorant is learned, the entire ‘landscape’ of odorant memory is changed. That is, not only is a new item added to the landscape, but the response patterns of existing items are change. That’s what we would expect in a connectionist model. Just how the brain does this is obscure, though I offered an off-the-cuff speculation.
Anyhow, let’s say that what Freeman was observing was the automatic memory that happens in the course of ordinary processing. Let us say that automatic memory is consonant with those ordinary processes. Deliberate memory is necessary to learn things that a dissonant with those processes. Let’s leave those two terms, consonant and dissonant, undefined beyond their contrastive use. We – me or someone else – can worry about a more thorough characterization later.
Deliberate learning: arithmetic, the method of loci
As an example of deliberate learning, consider arithmetic. It begins with learning the meaning of number names by enumerating collections of objects and then by learning the tables for addition, subtraction, multiplication, and division. This process requires considerable drill. Let’s hypothesize that that is necessary to overcome the inertia, the viscosity – to use a term I introduced in that earlier post – of the automatic process.
As a result of this drill, a foundation is laid on which one can then learn how to do more complex calculations. Considerable drill is required to become fluent in that process. But we’ve got three kinds of drill going on.
1. Meaning of number words: this is an episodic procedure that establishes the meaning of a small number of words. To determine whether any of the words applies to a collection of object, execute the procedure.
2. Learning arithmetic table: this is straight memorization of system items, each having the form: numeral, operation, numeral, equals, numeral.
3. Learning multiple-digit calculation: this is an episodic level set of procedures in which one calls up the items in the arithmetic tables and applies them in succession to pairs and n-tuples of multiple digit numbers.
The episodic procedures, 1 and 3, are dissonant with respect to ordinary episodic processes, such as moving about the physical world, while the system procedures, 2, are dissonant with respect to the ordinary processes of learning the meanings of words.
As another example, consider the method of loci, sometimes known as the memory palace. Here’s the account I gave in my working paper on Visual Thinking:
The locus classicus for any discussion of visual thinking is the method of loci, a technique for aiding memory invented by Greek rhetoricians and which, over a course of centuries, served as the starting point for a great deal of speculation and practical elaboration — an intellectual tradition which has been admirably examined by Frances Yates. The idea is simple. Choose some fairly elaborate building, a temple was usually suggested, and walk through it several times along a set path, memorizing what you see at various fixed points on the path. These points are the loci which are the key to the method. Once you have this path firmly in mind so that you can call it up at will, you are ready to use it as a memory aid. If, for example, you want to deliver a speech from memory, you conduct an imaginary walk through your temple. At the first locus you create a vivid image which is related to the first point in your speech and then you “store” that image at the locus. You repeat the process for each successive point in the speech until all of the points have been stored away in the loci on the path through the temple. Then, when you give your speech you simply start off on the imaginary path, retrieving your ideas from each locus in turn. The technique could also be used for memorizing a speech word-for-word. In this case, instead of storing ideas at loci, one stored individual words.
The process starts with choosing a suitable building and memorizing it. That’s deliberate learning. Think of it as analogous to the three kinds of drill involved in learning arithmetic calculation.
Actually using the memorized building for a specific task, that is deliberate learning as well. Here the deliberation is confined to associating items to be learning with positions in the palace. One learning a collection of system links. The idea that one is to use vivid images no doubt reflects the inherent nature of the nervous systems, it is an exhortation to consonance.
This post is a response to a long video posted by Dr. Tim Scarfe which raised a number of important issues. One of them is about physical constraints in the implementation of computing procedures and memory. I’m thinking this may well be THE fundamental issue in computing, and hence in human psychology and AI.
I note in passing that John von Neumann’s The Computer and the Brain (1958) was about the same issue and discussed two implementation strategies, analog and digital. He also suggested that the brain perhaps employed both. He also noted that, unlike digital computers, where you have and active computational unit linked to passive memory through fetch-execute cycles, that each unit of the brain, i.e. neuron, appears to be an active unit.
Physical constraints on computing
Here’s the video I was talking about. It is from the series Machine Learning Street Talk, #78 - Prof. NOAM CHOMSKY (Special Edition), and is hosted by Dr. Tim Scarfe along with Dr. Keith Duggar and Dr. Walid Saba.
As I’m sure you know, Chomsky has nothing good to say about machine learning. Scarfe is not so dismissive, but he does seem to be a hard-core symbolist. I’m interested in a specific bit of the conversation, starting about about 2:17:14. One of Scarfe’s colleagues, Dr. Keith Duggar, mentions a 1988 paper by Fodor and Pylyshyn, Connectionism and Cognitive Architecture: A Critical Analysis (PDF). I looked it up and found this paragraph (pp. 22-23):
Classical theories are able to accommodate these sorts of considerations because they assume architectures in which there is a functional distinction between memory and program. In a system such as a Turing machine, where the length of the tape is not fixed in advance, changes in the amount of available memory can be affected without changing the computational structure of the machine; viz by making more tape available. By contrast, in a finite state automaton or a Connectionist machine, adding to the memory (e.g. by adding units to a network) alters the connectivity relations among nodes and thus does affect the machine’s computational structure. Connectionist cognitive architectures cannot, by their very nature, support an expandable memory, so they cannot support productive cognitive capacities. The long and short is that if productivity arguments are sound, then they show that the architecture of the mind can’t be Connectionist. Connectionists have, by and large, acknowledged this; so they are forced to reject productivity arguments.
That’s what they were talking about. Duggar and Scarfe agree that this is a deep and fundamental issue. A certain kind of very useful abstraction seems to depend on separating the computational procedure from the memory on which it depends. Scarfe (2:18:40): “LeCun would say, well if you have to handcraft the abstractions then learning's gone out the window.” Duggar: “Once you take the algorithm and abstract it from memory, that's when you run into all these training problems.”
OK, fine.
But, as they are talking about a fundamental issue in physical implementation, it must apply to the nervous system as well. Fodor and Pylyshyn are talking about the nervous system too, but they don’t really address the problem except to assert that (p. 45), “the point is that the structure of ‘higher levels' of a system are rarely isomorphic, or even similar, to the structure of ‘lower levels' of a system,” and therefore the fact that the nervous system appears to be a connectionist network need not be taken as indicative about the nature of the processes it undertakes. That is true, but no one has, to my knowledge, provided strong evidence that this complex network of 86 billion neurons is, in fact, running a CPU and passive memory type of system.
Given, that, how has the nervous system solved the problem of adding new content to the system, which it certainly does? Note that here is their specific phrasing, from the paragraph I’ve quoted: “adding to the memory (e.g. by adding units to a network) alters the connectivity relations among nodes and thus does affect the machine’s computational structure.” The nervous system seems to be able to add new items to memory without, however, having to add new physical units, that is neurons, to the network. That is worth thinking about.
Human cortical plasticity: Freeman
The late Walter Freeman has left us a clue. In an article from 1991 in Scientific American (which was more technical in those days), entitled “The Physiology of Perception,” he discusses his work on the olfactory cortex. He’s using an array of electrodes mounted on the cortical surface (of a rat) to register electrical activity. Note that he’s NOT making recordings of the activity of individual neurons. Rather, he’s recording activity in a population of neurons. He then made 2-D images of that activity.
The shapes we found represent chaotic attractors. Each attractor is the behavior the system settles into when it is held under the influence of a particular input, such as a familiar odorant. The images suggest that an act of perception consists of an explosive leap of the dynamic system from the " basin" of one chaotic attractor to another; the basin of an attractor is the set of initial conditions from which the system goes into a particular behavior. The bottom of a bowl would be a basin of attraction for a ball placed anywhere along the sides of the bowl. In our experiments, the basin for each attractor would be defined by the receptor neurons that were activated during training to form the nerve cell assembly.
We think the olfactory bulb and cortex maintain many chaotic attractors, one for each odorant an animal or human being can discriminate. Whenever an odorant becomes meaningful in some way, another attractor is added, and all the others undergo slight modification.
Let me repeat that last line: “Whenever an odorant becomes meaningful in some way, another attractor is added, and all the others undergo slight modification.” That the addition of a new item to memory should change the other items in memory is what we would expect of such a system. But how does the brain manage it? It would seem that specific memory items are encoded in whole populations, not in one or a small number of neurons (so-called ‘grandmother’ cells). Somehow the nervous system is able to make adjustments to some subset of synapses in the population without having to rewrite everything.
In this connection it’s worth mentioning my favorite metaphors for the brain, as a hyperviscous fluid. What do I mean by that? A fluid having many components, of varying viscosity (some very high, some very low, and everything in between), which are intermingled in a complex way, perhaps fractally. Of course, the brain, like most of the body’s soft tissue, is mostly water, but that’s not what I’m talking about. I’m talking about connectivity.
Perhaps I should instead talk about a hyperviscous network, or mesh, or maybe just a hyperviscous pattern of connectivity. Some synaptic networks have extremely high viscosity and so change very slowly over time while others have extremely low viscosity, and change rapidly. The networks involved in tracking and moving in the world in real time must have extremely low viscosity while those holding our general knowledge of the world and our own personal history will have a very high viscosity.
In the phenomenon that Freeman reports, we can think of the overall integrity of the odorant network as being maintained at, say, level 2, where moment-to-moment sensations are at level 0. The new odorant is initially registered at level 1 and so will affect the level 1 networks across all odorants. That’s the change registered in Freeman’s data. But the differences between odorants are still preserved in level 2 networks. Over time the change induced by the new odorant will percolate from level 1 to level 2 synaptic networks. Thus a new item enters the network without disrupting the overall pattern of connectivity and activation.
That is something I just made up. I have no idea whether or not, for example, it makes sense in terms of the literature on long-term potentiation (LTP) and short-term potentiation (STP), which I do not know. I do note, however, that the term “viscosity” has a use in programming that is similar to my use here.
Addendum: Are we talking about computation in the Freeman example?
I take it as self-evident that an atomic explosion and a digital simulation of an atomic explosion are different kinds of things. Real atomic explosions are enormously destructive. If you want to test an atom bomb, you do so in a remote location. But you can do a digital simulation in any appropriate computer. You don’t have to encase the computer in lead and concrete to shield you from the blast and radiation, etc. And so it is with all simulations. The digital simulation is one thing, and real phenomenon, another.
That’s true of neurons and nervous systems too. [...] However, back in 1943 Warren McCulloch and Walter Pitts published a very influential paper (A Logical Calculus of the Ideas Immanent in Nervous Activity) in which they argued that neurons could be thought of as implementing circuits of logic gates. Consequently many have, perhaps too conveniently, assumed that nervous systems are (in effect) evaluating logical expressions and therefore that the nervous system is evaluating symbolic expressions.
I think that’s a mistake. Nervous systems are complex electro-chemical systems and need to be understood as such. What happens at synapses is mediated by 100+ chemicals, some more important than others. It seems that some of these processes have a digital character while others have an analog character. [...] I have come to the view that language is the simplest phenomenon that can be considered symbolic, thought we may simulate those processes through computation if we wish. That implies that there is no symbolic processing in animals and none in humans before, say, 18 months or so. Just how language is realized in a neural architecture that seems so contrary to the requirements of symbolic computing, that is a deep question, though I’ve offered some thoughts about that in the working paper I mentioned in my original comment.
If the phenomenon Freeman describes is not about computation, and it is NOT according to my current beliefs, then how does the problem brought up by Fodor & Pylyshyn apply?
And yet there IS a problem, isn’t there. There is a physical network with connections between the items in the network. Those connections must be altered in order to accommodate a new phenomenon. We can’t just add a new item to the end of the tape. That is, it IS a physical problem of the same form. So perhaps this technicality doesn’t matter.
This was originally published 11.30.12, but I'm bumping it to the top of the queue because I'm thinking about the general idea of the fluid mind. Here's a working paper, From Associative Nets to the Fluid Mind,
A couple of mornings ago I had a good idea while lounging in the tub. A good idea, but not a great or surprising idea. Good’s enough.
As I explained a year or so ago, I do a lot of thinking while lounging in the tub in the morning—or, for that matter, at other times of the day, on occasion. So this was not at all unusual. Still, I thought I’d post a brief note to the blog, you know, just to note the cognitive utility of hanging out in the tub. ‘Cause tending to one’s mind is important, but a rather obscure and tricky business.
Alas, I forgot to do so. And I forget just now which good idea I had that day, though I have a sense that I did act on it. Anyhow, it happened again today. So I made a point of writing a note this time.
Today’s insight is simple, that I could, and perhaps even should and will, talk about manga and anime at the end of my next, and I hope penultimate post, on pluralism. Working title for the post: Facing up to Relativism: Pluralist Axiology.
A mouthful, that: “axiology.” It’s basically a cover term for ethics and aesthetics. What’s a pluralist got to say about the fact that different peoples have different Life Ways?
Negotiation, that’s what. Latour talks about it in “Exploring Common Worlds” in Politics of Nature. While I had no specific itinerary in mind when I set out on this venture into OOO-land, I certainly didn’t expect to find myself with an occasion to talk about cartoons. But now...
* * * * *
Enough. This is about the bathtub, not about a post I’m going to write sometime in the next several days (I hope).
What is it about bathtub lounging that’s so congenial to meandering thought, to reverie? And why’s such thought useful?
It’s certainly not about working out details, about dotting i’s and crossing t’s. That requires sustained attention. Bathtub thinking seems to allow things to float to the surface and there wander around and mingle together.
I like to think of the mind as fluid, and has having many different viscosities at once. Call it hyperviscosity. Some things move very slowly, like chilled molasses, only slower. Other things move rapidly, like gas in a flame. But, in the mind, these things are all going on at once and consciousness, well, it attends sometimes to the fast things, sometimes to the slow things, and sometimes to the glacial things.
It’s not just that, in the tub, you really don’t have anything else to do; it’s not just the open time. It’s also the warm water. That’s important. It doesn’t have any effect on the brain’s temperature, of course, as that’s regulated to be body temperature, but the relaxation does have an effect on the overall state of the mind.
Tub thinking seems to affect the dynamics of hyperviscosity. The different viscosities tend to segment into different layers. Tub thinking gets the layers to interact with one another.
I’m thinking that some very slow things move just a bit faster and creep up a layer or three. Some of the fastest things dissipate and get out of the way. Other slow down, even way down, and sing. Thus there’s a gentle turnover in the depths and surfaces of the mind.
In view of current excitement over GPT-3 I'm bumping this to the top of the queue. This is about the implementation of Old School semantic or cognitive nets in neural nets where the nodes and edges of the semantic net each represent activity in regions of the neural net.
* * * * *
Comparison of Semantic Nets (left) with Attractor Nets (right). See below.
Now, let’s compare these two systems. In the symbolic computation semantics we have a relational graph where both the nodes and the arcs are labeled. All that matters is the topology of the graph, that is, the connectivity of the nodes and arcs, and the labels on the nodes and arcs. The nature of those labels is very important.
In the statistical system words are in fixed geometric positions in a high-dimensional space; exact positions, that is, distances, are critical. This system, however, doesn’t need node and arc labels. The vectors do all the work.
These statements are, I feel, at the right level of generalization and abstraction. The purpose of these notes is to lay out some of the things that would have to be taken into consideration in order to further develop those ideas.
HOWEVER, it’s sketchy & full of holes. First the sketchy stuff. Then I’ve got links to material that’s somewhat more worked out. It’ll help full in the holes and flesh out the details.
Stream of consciousness ramble-through
Caveat: These are informal notes, mostly for myself. I list them here as place-holders for work that needs to be done.
How do you make inference in the system? Symbolic model allows for a logic based on node and arc types. Such a model treats the knowledge structure as a database and makes inferences over it. Humans are not like that. In statistical systems there is no capacity for ‘externally’ guided inference. All inference is ‘black-boxed’ and internal to the system.
How do you add “meaning” to a statistical model? What is “meaning” anyhow? Does meaning ultimately require/imply interactive coupling with [living in] a world? And doesn’t it imply/require an intending subject?
Is this the ultimate import of critiques such as those of Dreyfus and Searle? But aren’t those ‘cheap’ critiques, arrived at without knowledge of how these systems work? And what of it?
If symbolic systems are more flexible and ‘deeper’ how come the statistical systems are more powerful in actual applications? The symbolic systems get their knowledge directly from the humans that design them. The makers decide on node and arc types and the way to make inferences over them; and the makers encode knowledge directly into the system. Statistical systems learn. NLP systems in effect back in to a simulacrum of the knowledge humans have ‘baked-in’ to the texts they generate. But what about visual systems based on unsupervised learning? Those systems are directly ‘in touch with’ an ‘external world’, no?
What about the phenomenal power of chess and Go programs? These programs are now more powerful than the best human players. I note as well that neither chess nor Go require contact with a rich physical world. Both take place in a 2D world containing a very limited universe of objects and in which there are very limited opportunities for action. For all practical purposes, there is no world external to the computer, unlike what we have for translation systems, text generation systems, or car-driving systems. Does this imply that perhaps the best and most powerful use of these systems is in the internal configuration of and maintenance of computational systems themselves? But won’t that make these systems even more ‘black-boxy’ than they already are?
And speech-to-text, text-to-speech?
Note that learning systems of various sorts do require exposure to mountains of data, whether externally generated and presented to the system or created by the system itself (as in game systems that compete against themselves). Humans do not require exposure to so many cases for effective learning. What’s this about? Well, for one thing it’s about having a brain and body that evolved to fit the world: no blank slate. Does is all cascade back on to that, or is there more?
Natural intelligence and metaphor
While much of my work with David Hays was anchored in a symbolic systems approach to the mind, we did venture elsewhere. Thus when the symbolic systems approach collapsed in the mid-1980s we were, if not exactly prepared, able to keep moving on. We had other conceptual irons in the forge.
In Cognitive Structures, which I linked above, Hays came up with a scheme to ground a digital cognitive system in an analog sensorimotor system. A few years after that he and I worked on a paper where we grounded our whole system in neural systems:
Abstract: The phenomena of natural intelligence can be grouped into five classes, and a specific principle of information processing, implemented in neural tissue, produces each class of phenomena. (1) The modal principle subserves feeling and is implemented in the reticular formation. (2) The diagonalization principle subserves coherence and is the basic principle, implemented in neocortex. (3) Action is subserved by the decision principle, which involves interlinked positive and negative feedback loops, and resides in modally differentiated cortex. (4) The problem of finitization resolves into a figural principle, implemented in secondary cortical areas; figurality resolves the conflict between pro-positional and Gestalt accounts of mental representations. (5) Finally, the phenomena of analysis reflect the action of the indexing principle, which is implemented through the neural mechanisms of language.
These principles have an intrinsic ordering (as given above) such that implementation of each principle presupposes the prior implementation of its predecessor. This ordering is preserved in phylogeny: (1) mode, vertebrates; (2) diagonalization, reptiles; (3) decision, mammals; (4) figural, primates; (5) indexing. Homo sapiens sapiens. The same ordering appears in human ontogeny and corresponds to Piaget's stages of intellectual development, and to stages of language acquisition.
I note that we conceived of the second principle, diagonalization, as involving holograph-like processing, which involves convolution. Convolution is important in some forms of contemporary neural net systems. Subsequently we published a paper in which we used convolution to account for metaphor:
Abstract: Karl Pribram's concept of neural holography suggests a neurological basis for metaphor: the brain creates a new concept by the metaphoric process of using one concept as a filter — better, as an extractor — for another. For example, the concept "Achilles" is "filtered" through the concept "lion" to foreground the pattern of fighting fury the two hold in common. In this model the linguistic capacity of the left cortical hemisphere is augmented by the capacity of the right hemisphere for analysis of images. Left-hemisphere syntax holds the tenor and vehicle in place while right-hemisphere imaging process extracts the metaphor ground. Metaphors can be concatenated one after the other so that the ground of one metaphor can enter into another one as tenor or vehicle. Thus conceived metaphor is a mechanism through which thought can be extended into new conceptual territory.