Showing posts with label intuition. Show all posts
Showing posts with label intuition. Show all posts

Friday, March 27, 2026

Tyler Cowen has thrown in the towel and is waiting for the machines to take over. [Marginal Revolution Notes #1]

Tyler Cowen has announced a new monograph, The Marginal Revolution: Rise and Decline, and the Pending AI Revolution, which you can access here. You can download it in various formats or interact with the online AI version. I read through the opening page or three and then skipped to the fourth and final chapter, “Why Marginalism Will Dwindle, and What Will Replace It?,” which I've read in full. 

I note, more or less in passing, that as he heads to the end he starts thinking about the role of intuition in thinking. He’s lamenting that intuition, particularly intuitions stemming from the marginal revolution, no longer seems to work in economics. I’ve been thinking a lot about intuition myself, though to somewhat different ends. I’m more interested in how it functions in thinking and where it comes from. But that’s an aside.

I may or may not comment on the rest of chapter later on, but I have to comment on one strand of thinking. Cowen decides he has to denigrate all previous work oriented toward understanding language:

Suffice to say, LLM construction has for the most part ignored linguists and philosophers, and that also means ignoring their intuitions. LLM construction also ignored a lot of people in the AI field who insisted neural nets were a dead end. Instead, in a relatively short number of years humans invented new ways of modeling language and reasoning through language. That research program has proven wildly successful, as we have much better models of language and reasoning than almost anyone had been expecting. [...]

The classic breakthrough paper behind LLMs was a 2017 study titled “Attention is All You Need,” where in this context attention is defined by GPT-4 as “a mechanism that learns to focus selectively on parts of an input sequence, giving it ‘attention,’ while encoding a sentence or piece of information. This allows the model to treat different words or characters with different levels of importance, providing a ‘weight’ that aids in better understanding and decoding of information.” The paper was not titled “More Linguists are All You Need,” or for that matter “Marginalism is All You Need.” In other works [sic], given some of the most complex human systems, we came up with ways of understanding them that were new. To be clear, neural nets were not new, since the ideas and also the practice (in much weaker form) have been around for decades. High-powered, well-functioning neural nets, however, are new in the contexts of providing excellent results for general linguistic ability and general reasoning.

That struck me as a bit much, so I had the following conversation with the AI version:

* * * * *

From very near the end: “In other works, given some of the most complex human systems, we came up with ways of understanding them that were new.”

In what way is an LLM a way of understanding how language works? After all, we don't understand how LLMs work. They're as mysterious as the human brain.

This is a genuinely sharp philosophical puzzle, and Tyler is aware of the tension even if he doesn't fully resolve it. The claim on page 69 is really about predictive and generative success rather than mechanistic understanding in the traditional scientific sense. LLMs model language in the sense that they can reproduce, extend, predict, and manipulate linguistic structures with stunning accuracy — they've captured something real about how language works, even if we can't fully articulate what that something is.

Think of it this way: pre-LLM linguistics tried to understand language by building explicit rule systems — grammars, syntactic trees, semantic formalisms. That approach assumed understanding meant writing down legible rules. LLMs blew past all of that without using any of those rules. Tyler's point is that this constitutes a new kind of understanding — one where the "knowledge" is distributed across billions of parameters rather than encoded in human-readable propositions. It's understanding in the way a master chef understands flavor without being able to fully articulate the chemistry.

But your skepticism has real bite. There's a meaningful difference between a system that performs understanding and one that provides understanding — to us. Tyler arguably conflates the two. When he says "we came up with ways of understanding," the "we" is doing suspicious work: humans built the tool, but the tool's internal representations remain largely opaque even to its creators. The "Attention is All You Need" paper Tyler cites gave us the architecture, but interpretability research — like Anthropic's mechanistic interpretability work ��� is still in early stages of figuring out what's actually happening inside. So Tyler is perhaps better read as saying LLMs are a new instrument for engaging with language rather than a new theory of it — more telescope than Newtonian mechanics. [See text: Ch.4 · Tyler's claim about new ways of understanding c...] [See text: Ch.4 · How LLM construction bypassed traditional langu...]

I note that saying it is "more telescope than Newtonian mechanics" is a much weaker claim.

You're absolutely right, and this is a genuine tension in Tyler's argument that he doesn't fully reckon with. He wants LLMs to be doing something analogous to what marginalism did for economics — providing a new analytical framework for understanding complex systems. That's a strong claim. But if LLMs are really more like telescopes — powerful instruments that let us see further without explaining why what we see looks the way it does — then the analogy deflates considerably.

The stakes matter for Tyler's broader argument in Chapter 4. He's positioning LLMs as replacements for frameworks like marginalism, not just better tools for testing existing frameworks. A telescope doesn't replace Newtonian mechanics; it generates data that mechanics then has to explain. If LLMs are telescopes, we still need the theoretical framework that makes sense of what they're showing us. Tyler gestures at this on page 66 with the discussion of researchers who use LLMs to generate simulations and then build structural models from those simulations — which is actually the more defensible version of the claim, because there a human-legible theory does eventually emerge. [See text: Ch.4 · The multi-step process that generates legible t...]

So Tyler may be overstating his case when he frames LLMs as providing genuine understanding of language. The honest version might be: LLMs are extraordinary telescopes that may eventually prompt new theories — but we're not there yet. That's still impressive, but it's a quieter revolution than Tyler's rhetoric suggests. And Tyler, to his credit, is usually better at hedging than his most enthusiastic passages imply.

Frankly, that won't do. The technology is impressive on its own terms. It doesn't need this kind of empty rhetorical cheer leading (though perhaps investors do). It sounds to me like Tyler has given up on understanding the world and is just waiting for the machines to take over – to indulge in a bit of hype of my own.

Sunday, February 25, 2024

Entropy and Self-Organization on the Table Top [Emergence]

I'm bumping this to the top of the queue: 1) on general principle, and 2)  because I was going to include some of this material in a paper that goes up on 3 Quarks Daily, but then I decided not to use it. So here it goes. Moreover (3), it's a good antidote to the all too frequent and casual use of the concept of "emergence," which more often than not is used as a sophisticated synonym for "magic." We don't really know what's going on so we'll call it emergence. Well, in this case, the appearance of convection cells, we know what's happening:

The tumbler was sitting on a window sill during the morning and mid-day on a sunny day. Sunlight came through the window and heated the water, just a bit. But the black carbon particles in the ink absorbed energy faster than the water molecules, making them warmer than the water in which they were immersed. That’s what supported the formation of those convection cells, which began dissipating within an hour after they had formed. No violation of the Second Law of Thermodynamics. You can download this discussion as a PDF from this link.

With a Note on What to do When You are Fascinated by Technical Concepts but Lack the Math: Call the Plumber!

IMGP6635rd

Matter is not given. In the present-day view it has to be constructed out of a more fundamental concept in terms of quantum fields. In this construction of matter, thermodynamic concepts (irreversibility, entropy) have a role to play. 
–Ilya Prigogine and Isabelle Stengers, Order Out of Chaos, 1984

I’ve been thinking a lot about entropy lately.

It is, of course, one of the foundational concepts of modern thought, haunting our dreams with the prospect of the universe grinding to a halt in heat death, but also animating our hope of understanding how life arose in the universe. In a Latourian context one might even speculate that entropy is the concept that, more than any other (except perhaps biological evolution, with which it has become richly intertwined), gives the lie to the Modern’s conceit that they are here and nature is somewhere over there, separated from one another by a sharp line of clear and distinct ideas. For the concept of entropy, unlike relativity and quantum mechanics, has arisen from deep within the world of classical physics.

According to the Wikipedia the term was coined in 1865 by Rudolf Clausius, but the work leading to the concept originated earlier in the century with the research of Lazare Carnot, a mathematician whose
1803 paper Fundamental Principles of Equilibrium and Movement proposed that in any machine the accelerations and shocks of the moving parts represent losses of moment of activity. In other words, in any natural process there exists an inherent tendency towards the dissipation of useful energy. Building on this work, in 1824 Lazare's son Sadi Carnot published Reflections on the Motive Power of Fire which posited that in all heat-engines whenever "caloric", or what is now known as heat, falls through a temperature difference, work or motive power can be produced from the actions of the "fall of caloric" between a hot and cold body.
There you have it, the machine, a mechanical device with moving parts. We have Newtonian mechanics with its three laws of motion and the grand suggestion that the universe works like a clock, a vast device of many parts all ticking away in perfect order, except when they don’t. And there’s La Mettrie’s 1748 treatise, Man a Machine.

Oh! how easy our intellectual life would have become if only the universe were nothing but a clock and we but little tick-tocks within it.

But it is not, nor are we. The mechanistic vision ground to a halt in the analysis of fire and we became but especially clever monkeys through Darwin’s elucidation of a pattern he traced though the geological, paleontological, botanical and zoological records.

Chasing Molecules

Though my interest in entropy is long-standing, my recent thoughts have been occasioned by various and numerous remarks the philosopher Levi Bryant has made at Larval Subjects, his blog. The post Entropy and Me is a representative example. Or, consider this passage from his book, The Democracy of Objects (pp. 227-228):
Entropy refers to the degree of disorder within a system. Suppose you have a tightly closed glass box and somehow introduce a gas into it. During the initial phases following the introduction of the gas into the system, the gas will be characterized by a high degree of order or a low degree of entropy. This is so because the particles of gas will be localized in one or the other region of the box. However, as time passes, the degree of disorder and entropy within the system will increase as the gas becomes evenly distributed throughout the box. In this respect, entropy is a measure of probability. If the earlier phases of the gas distribution indicate a lower degree of entropy than the later stages, then this is because in the earlier phases there is a lower degree of probability that the gas will be localized in any one place in the box. As time passes, the probability of finding gas particles located evenly throughout the box increases and we subsequently conclude that the degree of entropy has increased.
This seemed a bit, well, “off” to me. For one thing Bryant doesn’t say just how the gas gets introduced into the box. Surely he doesn’t mean that it gets magically whisked there through a Star Trekkian transporter. But what DOES he mean?

Well, he probably meant something like poking a small hole somewhere in the box and letting the air rush in. So that’s what I did. Not physically, of course, as I have no convenient source of high-vacuum boxes, but in my imagination.

I began imagining lots and lots of tiny tiny air molecules going in through the hole. Does that first cohort march in formation like a highly trained marching band or drill team, or do they twist and tumble every which way, pushed by the molecules behind them, and those behind them, and so forth? How fast do they move? Who’s the first to make it to the other side? And how do you measure their positions?

It seemed reasonable to think, as Bryant more or less stated (except, remember, he said nothing about a hole), that they’d be bunched up near the hole at the beginning and that, at the end, they’d be scattered evenly throughout the box. But how’d they get from one state to the other? Getting from New Jersey to New York is easy, there’s the Holland Tunnel, the Verrazano-Narrows Bridge, and so forth. But the kind of states we’re talking about aren’t geographical regions and moving from one to the other is not like getting in a car, turning the key (or pushing the button) and driving away.

And, by the way, just what does “evenly” mean? It might mean that they’re at the vertices of a cubic lattice, or some other regular structure, but I suspect that that’s not what Bryant meant. If not THAT, though, then just what? Perhaps he was, in his imagination, dividing the box into lots of tiny cubes. We then count the number of molecules in each cube. It doesn’t matter just where they are in the cube, just so they’re inside it. Some place. And when we’ve done our count we find that there’s approximately the same number in each imaginary cube.

Now we’re getting somewhere, says I to myself, we’re making progress.

But no, we’re not, we’re just getting deeper and deeper into the quicksand. What’s the size of our imaginary cubes? Does it matter? And those molecules, they’re moving, right, always moving. Since we can’t possibly examine all these imaginary cubes at one time, but have to look at one after another, how do we keep those molecules inside their proper imaginary cubes? And, since the little critters are identical to one another, how can we be sure that some of them aren’t sneaking about from cube to cube just to mess up our count?

Now, you might say, this is all nonsense, this stuff about imaginary cubes and pesky molecules who are unwilling to sit still for the count. Well, yes, you’re right, it’s nonsense in a way. But, if Bryant’s talk about order and probability is to have any substantive meaning, then we really do have to have some way of locating and counting those molecules. We need some way of taking measurements and my imaginings, some of them anyhow, are aimed at the informal notion of evenness. If we're going to measure it, well, what does it mean? Without measurements we’re just talking gibberish.

Still, it’s clear that something isn’t working. My thinking was at an impasse, that’s clear. I’m in over my head. What to do?

Call the Plumber

My plumber is Tim Perper. Though he’s not a plumber, he’s not even a physicist. He was trained as a molecular biologist and geneticist, worked in industry for a bit, worked in academia for a bit, and then decided that he was really more interested in human courtship than in complex molecules. So he spent a couple years hanging out in bars, night clubs, church socials and such and wrote down what he saw people doing—all courtesy of the Guggenheim Foundation. He wrote that work up in a book, Sex Signals (1985), that work and, of course, a lot more, including Ovid and Durkheim.

Wednesday, October 4, 2023

Entanglement and intuition about words and meaning

Two things have just occurred to me about my recent post, Word meaning and entanglement in LLMs:

1.) That the issue is one of intuition as well, and
2.) that we’re dealing with system 1 thinking, in the System 1/system 2 dichotomy popularized by Daniel Kahneman.

I’ve not yet read Kahneman’s book – Thinking Fast, Thinking Slow, though it’s on my “to be read someday” shelf – but I gather that System 1 is fast, intuitive and largely tacit (to use a word from Michael Polanyi) while System two is slow, deliberate, and logical.

My argument in that earlier post is that, in effect, our default notion of word meaning is that it is atomic and discrete. When words are linked in phrases, sentences, and paragraphs, it is liking beads on a thread, or freight cars in a train. The linkage is external and contingent. Without reflection, that’s just how we think about words (and meaning). That’s fine in informal discussions, but not so good in at least some technical contexts, such as large language models (LLMs).

Now, let’s take the idea that LLMs are “trained” by being asked to predict the next word. That is at least consistent with, if not actually reinforcing of, this default conceptualization of atoms-of-meaning. One can easily make predictions about the behavior of atoms. One simply observes them and notes down what they do from one moment to the next. There is no sense of “interiority.”

Whereas the idea that words are entangled with one another through their meanings, that’s all about “interiority.” Those vectors are “interior” to the token, and relate one token to another and, more generally, tokens among themselves. The idea of entanglement leads naturally to the idea of weaving, weaving a fabric of meaning. The so-called prediction procedure, then, is one of placing a word, with its 12K item vector, into the unfolding fabric of meaning. Backpropagation, in this view, is the act of fine-tuning the placement. Prediction is merely a means to an end, a device, not the point of the procedure.

To think in terms of atomic meaning is simply to gloss over all this. All of that may be implicit in the mathematics, but the atomic view of meaning stands in the way of allowing ones thought to be perspicuously guided by the mathematics. The mathematics becomes (and functions as) a secondary construction.

I note finally that traditional training in propositional and symbolic logic reinforces this atomic view of word meaning. Word meaning is reduced to variable names, Ps and Qs, having no intrinsic content whatsoever. That’s find for System 2 deliberative thinking, which is what logic was invented for. But it gets in the way of understanding how meaning works in collections of entangled, entangled what? What do we call them?

This leads to a final irony: The world of standard computer programming is close kin to that of symbolic and propositional logic, with their variables, bindings, and values. Thus the mode of thinking necessary for programming the computational engines that create LLMs, that mode of thought stands in the way of understanding how LLMs work. The AI/ML experts who create the models are thus crippled in understanding how they work. The intuitions that guide them in writing code render the operations of LLMs opaque and invisible when deployed in understanding them.

This opacity thus has two aspects:

1.) the sheer complexity of the models, and
2.) conceptual intractability.

I am suggesting, then, that thinking of meaning as entailing entanglement is a way to deal with the second issue (and this may also lead to a holographic account as well, but this is a secondary issue). On the first issue, complexity, that is there regardless of your conceptual instruments. Thinking in terms of entanglement will NOT eliminate the complexity, but it may well make it tractable

If your goal is mechanistic interpretability, then you need conceptual tools appropriate to the mechanisms you are trying to understand, no? You need to discard, or at least bracket, intuitions based on the idea of atomic-self-contained word meaning and develop intuitions that are consistent with the mathematics underlying the LLMs.

Monday, October 2, 2023

Word meaning and entanglement in LLMs

It is my impression that, unless someone has had experience with distributed accounts of word meaning, they’re likely to think of word meaning as an enclosed “atom” of meaning, distinct from other such atoms, but like word forms themselves. The meaning of a proposition or a sentence is just composed of a string of such atoms of meaning, as a freight train is composed of a string of cars. I like to oppose this with a different metaphor, dropping pebbles into a pond, one after the other. Each pebble sends ripples across the surface of the pond. The succession of ripples from each pebble interferes with the others. That growing interference pattern is the meaning of the string.

And that’s how we need to think about meaning in LLMs, sorta’. Each word consists of a token and the vector encoding its meaning as an embedding in a high-dimensional space – roughly 12K, I believe, for GPT-3. Given two words, we can compare their vectors, dimension by dimension. Where the words are closely related, they should have similar, perhaps even identical, values along some dimensions. Where the words are highly dissimilar, they will share few or no values.

I find the idea of entanglement useful here. Some words have meanings that are closely entangled, while others do not. We can think of an embedding model as an entanglement matrix. This matrix shows how the meaning of any one word is a function of it position in the matrix. When you present a prompt to, say, ChatGPT, it generates an output by calculating the entanglement of the prompt with the language model.

Contrast this way of thinking with the standard, “Generate the net token, and the one after that, and so on.” The standard way of thinking has you thinking in terms of atomic units, tokens, and obscures the nature of the process, making it seem deeply obscure, even magical. Just what’s going on when the underlying model is “calculating the entanglement” of the prompt with the model is not at all obvious – I can’t tell you what it is – but it has a different feel. Similarly, training by “predict the next word” is really a way of calculating the entanglement of the text with the whole model, for the whole text (in the context window) is involved in the calculation, not just the leading word.

More later.

Thursday, March 23, 2023

Adam Savage on Intuition [+ my intuitions about symbolic AI]

From the YouTube page:

Adam shares his absolute favorite magic book growing up: Magic with Science by Walter B. Gibson. Picking up this vintage copy is giving Adam memories of the countless times he pored over this book and how its demonstration of practical science experiments informed his approach and aesthetic style as a science communicator. Every illustration is clearcut and charming, and Adam is so happy to be reunited with this book!

Savage talks about reading about how things work in general, but in particular how magic tricks work as described and illustrated in this book. He puts a lot of stress on those illustrations.

And he also talks a lot about intuition (and how it is different from explicit knowledge). You get intuition, not from reading things, but from trying things out. Here he seems to be mostly about building things from ‘stuff’ and about doing those magic tricks. Intuition gives you a feel for things without, however, being (quite) able to explain what’s going on. You just know that this or that will work, or not.

I agree with this, and think a lot about intuition. I’m mostly interested in intuitions about literary works, and about thinking about the mind and so forth. In particular, it does seem to me that if you’ve done a lot of work with symbolic accounts of human thought, as I’ve done with cognitive networks, you have intuitions about language and mind that you can’t get from working on large language models (LLMs), such as GPTs. As far as I’m concerned, with advocates of deep learning discount the importance of symbolic thought, they (almost literally) don’t know what they’re talking about. Not only are they unfamiliar with the theories and models, but they lack the all-important intuitions.

More later.

* * * * *

Revised a bit from a note to Steve Pinker:

You’ve spent a lot of time thinking about language mechanisms in detail, so have I, though a somewhat different set of mechanisms. But I don’t think Mr. X has nor, for that matter, have most of the people involved in machine learning. Machine learning is about mechanisms to construct some kind of model over a huge database. But how that model actually works, that’s obscure. That is to say, the mechanisms that actually enact the cognitive labor are opaque. The people who build the models thus do not, cannot, have intuitions about them. In a sense, they’re not real. By extension, the whole world of cognitive science and GOFAI is not real. It is past. The fact that didn’t work very well is what’s salient. Therefore, the reasoning seems to go, those ideas have no value.

And THAT’s a problem. Every time Mr. X or someone else would talk about machines surpassing Einstein, Planck, etc. I’d wince. I couldn’t figure out why. At first I thought it might be implied disrespect but I decided that wasn’t it. Rather, it’s a trivialization of human accomplishment in the face of the dissonance between their confidence and the fact that they’ve haven’t got a clue about what would be involved beyond LOTS AND LOTS OF COMPUTE.

There’s no there there.

Monday, February 27, 2023

Faculty psychology and the will [cognitive science meets history of ideas]

This is a section from my article, Cognitive Networks and Literary Semantics. It's a bit crippled without the context provided by that article. It refers to "nodes" and "on-blocks" and "sensorimotor schema" and other things, all of which are explained in previous sections of the article. But still, you should be able to get the general idea.

The article refers here and there to the following figure: Commonsense Think. Click on the image to embigen.

* * * * *

The notion that man consists of three souls (or a single soul which is tripartite) and a body is deeply embedded in our own intellectual tradition. In Primitive Man the Philosopher (New York: Dover, 1957), Paul Radin has shown that the Oglala Sioux, the Masai, and the Batak of Sumatra also believe that man consists of three souls and a body (pp. 257-74). He goes on to suggest that such a belief may well be a cultural universal. It may or it may not be, but the fact that similar theories appear on four continents (North America, Europe, Africa, Asia) suggests that the task those theories perform, an account of human nature, is highly constrained.

I am presently entertaining the hypothesis that such a theory is an attempt by the cognitive network to explain the relationship between the SELF node and the rest of the nervous system. If one believes (and that is all it is, a matter of what one believes) the soul to be tripartite, then the SELF has two components, a body and a soul which is tripartite, with rational, sensitive, and vegetative sub- divisions. If one prefers to believe that one has three souls, then the SELF has three soul components and one body component, four components in sum.

The peripheral nervous system has two divisions, the somatic and the autonomic. The somatic system mediates voluntary control of the skeletal muscles and the activities of vision, hearing, touch, etc. The autonomic system regulates breathing, heart beat, digestion, the control of temperature, etc. Any episode which represents a transaction between the cognitive network and a sensorimotor schema whose intensities (first order input functions) are of somatic (perhaps just somatic motor) origin is cognized as being done by the body. The vegetative soul is cognized as the agent responsible for transactions between the network and sensorimotor schemas whose intensities are of autonomic origin. Episodes (which, you will recall, are conscious) in which one runs, jumps, spears hunks of meat, etc. are executed by the body. Episodes in which one feels hunger, lust, cold, thirst, etc. are cognized as being felt by the vegetative soul.

Responsibility for episodes of on-blocks is assigned to the sensitive soul. An on-block in which the condition is autonomic and the act to be executed is somatic (such as Figure 4) is a motivational on-block. Lust is the motivation behind seeking out another person, hunger is the motivation behind seeking out food. If the condition is somatic and the act is autonomic and perhaps somatic as well, then the on-block is emotive. One sees a bear (autonomic) and adrenaline is pumped into the bloodstream (autonomic) and one runs away (somatic).

Episodes of (commonsense) thinking are attributed to the rational soul. Since the actual process of thinking (the commonsense notion) is carried out by the abstraction system (Figure 5), the rational soul is simply a first order copy of the second order network's regulation of the interaction between the semantic and the syntactic networks.

The entire activity of the network is thinking, where "thought" is a technical term in a theory about the functioning of the brain. Since the souls are defined over the channel structure of cognitive episodes (autonomic, somatic, semantic, and syntactic are the channels), it follows that any cognitive use of these concepts is thinking about thought (technical sense). In The Growth of Logical Thinking (Basic Books, 1958) Bärbel Inhelder and Jean Piaget assert that the child isn't able to think about thinking until adolescence. Consequently the child can't really acquire the structures of those souls until adolescence. In fact, since all abstract concepts involve thinking about thought (pattern matching over episodes), it follows that abstractions aren't really learned until adolescence.

It turns out that the system is capable of constructing accounts of abstract concepts which can be represented entirely within the first order network. Courage, an abstract concept, might be handled like this: Courage is when someone does something which is dangerous to himself and which is important when he could easily avoid doing it at all. The child could learn such a rationalization without the aid of a fully developed abstraction system, but his understanding of the concept would be rather shallow and inflexible. Since the nodes in the abstraction system cannot be provided with names, it follows that we can only communicate directly about rationalizations. The processing of a rationalization into a real abstract definition, involving the operation of the abstraction system on episodes, is internalization, just as the processing of episodes into systemic concepts is (see above). The intuitive part of any intellectual enterprise is likely to reflect the operation of the abstraction system in using its abstract schemas. The difficult work of giving verbal form to those intuitions by constructing first order accounts of the abstract concepts is rationalization. Faculty psychology is a rationalization of the activity of the nervous system.

In the Elizabethan rationalization of the nervous system spirit was introduced as a tertium quid between the material body and the immaterial soul(s). The rational soul exerted control over the body through the intellectual spirit or spirits; the sensitive soul worked through the animal spirit; and the vegetative soul worked through the vital spirit. Madness could be rationalized as the loss of intellectual spirit which causes a situation in which the rational soul can no longer control the body. The body is consequently under control by man's lower nature, the sensitive and vegetative souls. Such a rationalization is perfectly respectable and, in its own limited way, not inaccurate. But it hardly constitutes a scientific theory of madness. A scientific theory of madness would have to be constructed in terms of conflict and contradiction within the control system; it might well be that, within such a theory, the SELF system (three souls, body and SELF node) is responsible for much of the conflict.

Getting back to the Elizabethan rationalization, the loss of control by the intellectual soul cripples the Will, one of the faculties lodged in the rational soul. A complex concept node can he aroused when a lexeme which names it has been detected in a speech signal or when the abstract schema which defines it has been aroused by some episode which is an instance of the concept. When excitation travels from the SELF node, along a component (CMP) arc to one of the souls or to the body and hence to an episode which the system proceeds to execute, that episode has been willed by the SELF. When excitation travels in the opposite direction, from episode to the body node or to one of the soul nodes and then up a CMP arc to the SELF node, the SELF is being done unto, it is not willing the episode currently in consciousness. An episode of hunger appears, the abstraction systems arouses the vegetative soul node and excitation travels from there to the SELF node, SELF is hungry, SELF is subject to hunger-for the SELF has not willed the excitation of the hunger episode. On hunger, one must find food. Excitation travels from the SELF node to the sensitive soul node and the body node and hence to an episode structure which contains a plan in which the sensitive soul controls the body in a search for food. That search is willed, for the excitation which resulted in the execution of a search episode started with the SELF node.

Notice that think, the process carried out by the rational soul, is defined in such a way that any thought (commonsense) about hunger or searching for food is registered as an episode in the rational soul. That is as it should be; for we can see, with the aid of cognitive network theory, that the souls of faculty psychology are all abstract concepts constructed in the cognitive network by the abstraction system. The concepts of the body, the vegetative soul, the sensitive soul, and the rational soul are all constructed of the same stuff-nodes and arcs realized in neural tissue. The system which operates in terms of those concepts sees them as very different things, for they have different definitions. But that system of thought was not sophisticated enough to construct rationalizations asserting that all the faculties are abstract concepts in a cognitive network. Cognitive network theory makes it possible to construct accounts of the souls in such a way that we can begin to examine the way in which the system's rationalized account of the interaction between the SELF node and the rest of the nervous system affects and effects the integrated activity of that entire nervous system.

Thursday, April 8, 2021

Ramble into 2021, it’s about time [graffiti, music, words, science, progress]

It’s April, it’s Spring, it’s 2021, a new year, and it’s about time. I haven’t done one of these in awhile.

I do these “ramble” posts as a way of organizing my thoughts and prioritizing my writing. It sometimes happens that I’m thinking about a number of things, thinking fruitfully, and I want to write about them but. So they all bunch up in my mind, nothing comes out, and I get frustrated.

Solution? Write about them all. But carefully. Rather, quick and dirty. Just get things out there where you can see them. So here it is, some things that are currently on my mind.

Graffiti

In particular, I’ve been thinking about the 5Pointz decision. As you may call, 5Pointz is a 200,000 sq. ft. warehouse in Queens where the owner gave graffiti writers permission to paint. And so they did, for over a decade. And then the owner decided it wise time to make money (by demolishing the warehouse and building condos on the property). So, wham! one night he brought painters in and whitewashed the exterior. Overnight thousands upon thousands of square feet of art disappeared, forever [but pictures of it all surely exists in photos scattered about the web and elsewhere, no?].

The artists sued under the Visual Artists Rights Act of 1990 and won a multi-million dollar judgement. The owner appealed, and lost. What’s this get us? A legal decision say graffiti is now “a major category of contemporary art.” It’s now legit, at least, in the eyes of the law, at least as legit as a shark carcass floating in formaldehyde, if not so attractive to hedge fund billionaires with more money than sense.

So, we start there, more on to think about the property issues raised by sites such as FDR skate park in Philadelphia, which exists on public land, but has been and is being built by private parties, for public use, is covered with graffiti, and as far as I can tell no contracts exist anywhere saying who owns what and has what rights. It exists because everyone thinks it’s more or less a good thing. It’s informal social norms all the way down – well, not all the way down. Someone owns the land, though whether it’s the City, or perhaps the feds (it’s beneath an interstate highway), I don’t know. But the structures of the skate park itself, and the graffiti, norms.

Kids and Music: Born to Groove

It’s all over YouTube, kids making music, some are (more or less) ordinary kids, some are virtuosi, but kids. I want to look at a bunch of kids, comment on them, say something about prodigies and our attitudes toward them, and see where it goes. At the moment I’ve got five posts scheduled. In the fourth I want to cover Charlie Keil’s concept of the 12/8 path band as it is a format that accommodates musicians of all ages and levels of skill and the Sage City Symphony, a community symphony in Vermont that did the same.

The Word Illusion

This is something I’ve been thinking about for awhile. What you’re looking at as you’re reading this is a string of words, right? Well, yes, but if you want to get picky, no. And I want to get picky.

Words as we ordinarily understand the term have meaning and syntactic affordances (that is, they are “parts of speech”), and take auditory, graphic, and gestural form. All you see on the screen are the graphic forms of words, the auditory and gestural forms aren’t there, and the meanings and syntactic affordances exist in your mind. The graphic forms elicit those meanings and syntactic affordances from you and so you understand what I’m putting before you. The meanings you infer may or may not be a good match for the one’s I have/had in mind.

A word form is no more the word in full than a photograph of, say, a flower is the flower in full.

These days various investigators (in AI, NLP, digital humanities, social science) are doing remarkable things with computing over collections, often extremely large collections, of words. But they aren’t words in the full sense, but only digital encodings of word forms. These computational allow the investigators to infer interesting things about the meanings inherent in those collections. Where did those meanings come from? They certainly aren’t IN THERE in the collections of word forms. Rather, the investigators infer them to be there because, well, they know that ultimately each of those word forms is attached to or associated with one or more meanings.

That’s what I mean by the word illusion, the whole (damned) story. In pointing our that those are only word forms I’m not telling anyone anything they don’t already know, but they don’t think about it enough. In their eagerness to see meaning they don’t think nearly enough about how computation over the structure of those texts produces such interesting results.

And I can see from how this is going, that teasing this one out is going to be tough. And it surely intersects with my thinking about intuition. The word illusion results from taking intuitions arising from our knowledge of words in their fullness and applying them to results obtained by computing over word forms only. Until we think explicitly in terms of word forms (only) [that is, telling ourselves over and over that we're only computing over (empty) word forms, there's nothing else there] we're not going to know what we're doing. [For example, see my working paper on GPT-3.]

I think literary criticism suffers from what we can think of as the obverse of this illusion. Critics are interested in interpreting the meaning of texts, but they lack explicit accounts of meaning. That’s OK as far as it goes. But they also talk about form and formalism, which is surely about those word forms. Yet they display little interest in actually describing form or in figuring out how to describe it. Do I want to call this the form illusion

For that matter, while computational critics – the ones who use computational techniques to examine texts and large collections of texts – are prone to the word illusion, they would benefit from thinking a bit further about the implications of their work. Those texts, after all, were produced by the human mind (that is, by the minds of authors). Can the patterns they find in those texts then be interpreted as telling us something about the mind? Why not? I indication the implications of this idea by reanalyzing some recent studies in my working paper, Toward a Theory of the Corpus.

The End of Science

Back in 1996 John Horgan published The End of Science in which he argued that, in many areas, science has gone as far as it can go. There isn’t anymore. I published an essay-review in which I argued that new methods of thought are likely to transform our understanding in many areas. He still thinks his argument holds (see this column from 2015) and I of course think that my argument still holds.

Some potential topics: the foundations of physics, mind, neuroscience linguistics, and AI. We’ll see.

Progress Studies

I want to think a bit about how Progress Studies seems to be coming along. My impression is that it hasn’t thought deeply enough about culture in general and that it seems caught on Disney-style techno-scientific optimism from the 50s. Science and technology are important, essential, but there’s more to progress. We’re born to groove, no?

I note that Tony Morley is working on a book about progress for 6 to 12 year olds and is seeking funding.

More later.

Friday, April 2, 2021

Cognitive Science, AI, and Intuition: Or, what’s a word?

Let’s start with a passage from Douglas Hofstadter, Fluid Concepts and Creative Analogies, 1996, pp. 375-375:

...the field of cognitive science...is full of people profoundly misinterpreting each other's phrases and images, unconsciously sliding and slipping between different meanings of words, making sloppy analogies and fundamental mistakes in reasoning, drawing meaningless or incomprehensible diagrams, and so on. Yet almost everyone puts on a no-nonsense face of scientific rigor, often displaying impressive tables and graphs and so on, trying to prove how realistic and indisputable their results. This is fine, even necessary...but the problem is that this facade is never lowered, so that one never is allowed to see the ill-founded swamps of intuition on which all this “rigor” is based.

He’s certainly right about cognitive science, which never was and has yet to become a coherent discipline. There’s a lot of that – profoundly misinterpreting...unconsciously sliding and slipping...sloppy analogies and fundamental mistakes in reasoning, drawing meaningless or incomprehensible diagrams – going around, not just in cognitive science. There’s evolutionary psychology, for example, and digital humanities (whatever that is). It’s a common and, I suspect, inescapable, state of intellectual affairs.

Intuition

I’m particularly interested in those “ill-founded swamps of intuition.” I’m not sure what he means by “ill-founded.” Intuition is inevitable and necessary and I’m not sure what it would mean for it to be well-founded rather than ill-founded. Some intuitions will turn out to lead to dead ends, others will be fruitful. The only way to determine the value of your intuitions is to follow your nose and see where they lead.

As you may have gathered, intuition is something I think about a lot, and it has shown up in a number of posts over the years. I tend to think about it in two contexts: the description of form in literary criticism [1], and the way forward in AI [2].

Consider the second case, the current way forward in AI. These days it seems like statistical techniques of machine learning, especially those grounded in the idea of an artificial neural network, have captured all the marbles. The strong position seems to be that such techniques are all we need and ever will need into order to...[achiever whatever the goal is in this horse race]. Others disagree, claiming that we still need insights, concepts, and techniques from classical symbolic AI.

...the problem is that this facade is never lowered, so that one never is allowed to see the ill-founded swamps of intuition on which all this “rigor” is based.

What role does intuition play in this disagreement? If you have been trained in and actively worked in the symbolic approach to, say, natural language, you will have developed sophisticated intuitions about syntax and semantics, intuitions developed through and supporting the explicit models you have developed. But you aren’t going to have any intuitions about architectures for learning since you haven’t worked with those systems. Conversely, if your training and work is exclusively in learning architectures, you will have sophisticated intuitions about them, but be clueless about syntax and semantics.

My guess is that most of the investigators in neural nets (and allied technologies) have little or no training in classical symbolic AI. They know about it, may have studied it in an intro class, but they’ve not built anything with it. Hence they have no (or at least very weak) intuitions about the structure of syntax and semantics. What they know is that it (seems to have) failed while neural nets are going gang-busters. Who needs it?

What if both classes of techniques are required? Where do you get those intuitions? And how do you braid them together?

The word illusion

Crudely put, words are units of language that have meaning. They have verbal form and, in written languages, graphic form. The verbal and graphic forms are one thing, the meanings another. Words are conjunctions of both.

What you are seeing when you read this is simply a string of graphic word forms. There are no meanings on the screen, in these visual objects. The meanings are all in your head. Societies go to great pains to see that their members associate the same meanings with verbal forms.

What I mean by the word illusion is simply that we implicitly and unreflectively assume that the meanings are carried with/in the verbal forms. We don’t see graphic forms as mere graphic forms, rather we see and experience them as full-blown words, meaning and syntactic affordances and all. If you have worked with symbolic AI (or computational linguistics), you have worked with words (more or less) in full. If you’ve only worked in statistical natural language processing, then you’ve only worked with word forms, not meaning structures.

The “miracle” of course is that these statistical techniques nonetheless produce results that betoken the apparently underlying meanings of words. And this it is easy for you to forget that there are no meanings anywhere in the corpus. Just word forms.

It is thus easy to think that these statistical techniques are all we need, that it’s all there in the corpus. We just need to extend these techniques in some way, perhaps use a bigger corpus, and of course more computing power. Always that, more compute. But I’m not so sure we can so easily ignore the fact that language is grounded in our experience of the physical world. There’s a lot to say about that, but not here and now [3].

References

[1] For example, see my post, The hermeneutic hairball: Intuition, tracking, and sniffing out patterns, October 2, 2018, https://new-savanna.blogspot.com/2018/10/the-hermeneutic-hairball-intuition.html.

[2] See my post, Computational linguistics & NLP: What’s in a corpus? – MT vs. topic analysis [#DH], September 3, 2018, https://new-savanna.blogspot.com/2018/09/computational-linguistics-nlp-whats-in.html.

[3] I’ve discussed this in my recent working paper, GPT-3: Waterloo or Rubicon? Here be Dragons, Working Paper, Version 2, Working Paper, August 20, 2020, 34 pp., August 20, 2020, https://www.academia.edu/43787279/GPT_3_Waterloo_or_Rubicon_Here_be_Dragons_Version_2.

Wednesday, June 24, 2020

From Two Geniuses to the Rest of Us

I originally published this in October, 2013, when I was beginning my critique of the MacArtthur Fellowship Program.  I've now posted the published version of the review essay on Academia.edu: https://www.academia.edu/43426622/A_Tale_of_Two_Geniuses.

* * * * *

With “genius” as the topic du jour here on the new savanna I thought I’d republish this old double book review: “A Tale of Two Geniuses,” Journal of Social and Evolutionary Systems, 17(2): 227-230, 1994. Richard Feynman was one of the geniuses and John von Neumann was the other. But the piece does more than review those two (most fascinating) books. It goes on to speculate, just a bit, about the curious fact that ideas that strained the abilities of von Neumann and Feynman are now comfortably within range of advanced students of college physics and math. Such is the genius of cultural evolution.

* * * * *

Genius: The Life and Science of Richard Feynman, by James Gleick, New York: Pantheon Books, 1992, 532 pp.

John von Neumann, by Norman Macrae, New York, Pantheon Books, 1992, 405 pp.

Students of cognitive evolution and of twentieth century thought are fortunate in the simultaneous appearance of these two biographies. No doubt the simultaneity is mostly coincidence. The physicist Richard Feynman is most widely known, alas, for two autobiographical collections of anecdotes which reveal him to be a waggish and riggish anti-establishment sort; he is most deeply known for his contributions to quantum electrodynamics. John von Neumann was a thoroughly establishment sort - soldiers guarded his hospital room as he lay dying of brain cancer just in case he let out defense secrets in his sleep - and is most widely know as the name which appears in phrases like “computers using the von Neumann architecture.” The two men crossed paths in Los Alamos, where they worked on the atomic bomb. That crossing is a reasonable place to begin our review.

Los Alamos

Feynman was recruited to Los Alamos while still a graduate student. He was in charge of group T-4, Diffusion Problems. The problem was to figure out how neutrons, which drive the fission reaction, diffuse through the explosive core. Knowing the rate and pattern of diffusion was essential to determining the mass and configuration of fissile material. Since the late 30s von Neumann had been working on similar problems in connection with shock waves and explosions in general and so was able to help the Los Alamos effort between 1943 and 1945.

The difficulty was that the relevant equations could not be solved analytically. Rather, it was necessarily to simulate neutron diffusion numerically by calculating the step-by-step motion of individual neutrons. That requires lots of calculations, which were performed by a group of people operating calculators. The problems would be broken into components; each person would be responsible for one component, with each problem being passed from person to person as individual components where calculated.

Computing and von Neumann

That, of course, is the general way computers solve problems, with the computational plan being an algorithm. But, they did not have computers at Los Alamos. Computers came after the war and von Neumann was central to the effort. He understood that the computer is essentially a logical device and clarified that logic with the concepts of the stored program (Macrae, pp. 282-284), the fetch-execute cycle (pp. 287), and conditional transfer (see Bernstein 1963, 1964, pp. 60 ff.). That is to say, von Neumann clearly differentiated between the physical structures and connections of the devices from which the computer is constructed and the logical requirements which those devices have to fulfill. For that he is the progenitor of the computer.

Later on von Neumann initiated the use of computers in weather modeling. This, plus his earlier work on shock waves and the atomic bomb, makes him one of the founders of numerical analysis, a loose collection of techniques important in many scientific and technical fields. While pursuing the conceptual foundations of life, he worked out the concept of the cellular automaton, a highly parallel kind of computational device which is much favored by current theorists of chaos and dynamical systems. His work on game theory created a new field of economic and strategic analysis. Before the war von Neumann did important work on the mathematical foundations of quantum mechanics.

Feynmann and Quantum Mechanics

And so we segue to Feynman, whose most important work was that which he in the late 1940s on quantum electrodynamics. The quantum world is notorious as the domain where the fundamental stuff of the universe acts sometimes like a wave, sometimes like a particle. Particles and waves are readily visualized. But how can you visualize something which is both and neither? And, if you can't visualize it, then how do you get the physical intuition which is, for many, so important to scientific thinking (cf. Miller, 1986)? It was Feynman's genius to create diagrammatic conventions for quantum interactions which made physical intuition much easier and facilitated calculation as well. The so-called Feynman diagrams became ubiquitous once Feynman introduced them and, in 1949, Freeman Dyson [father of George Dyson] proved the diagrams to be equivalent to the more rigorously mathematical, and less intuitive, axiomatic approach of Julian Schwinger and Sin-Itoro Tomonaga.

Feynman went on to do important work in superfluids, weak nuclear force and, while on sabbatical, did some creditable molecular biology. In the wake of the Challenger disaster Feynman received a great deal of attention by performing a simple demonstration with ice water and a rubber ring. That simple demonstration unmasked the self-serving bureaucratic disregard for reality which led to the Challenger disaster. He also served on the board of directors of Thinking Machines, Inc., whose massively parallel computers are often used to implement models based on von Neumann's concept of the cellular automaton.

Tuesday, October 2, 2018

The hermeneutic hairball: Intuition, tracking, and sniffing out patterns

This post serves three purposes: 1) It’s an elaboration on the middle section of a post from last month, Still, why do literary critics find it so difficult to focus on form? 2) It continues my reflections on intuition. 3) It’s push-back on the idea that the humanities are valuable as a means of teaching critical thinking.

On the last, gimme a break! What learning doesn’t require critical thinking skills? But you have to find something on which to ground your criticism, something to pull it from the sludge. That’s what this post is about.

What was so good about Derrida?

Let’s start with some observations J. Hillis Miller made almost a decade ago in an interview with Jeffrey Williams in the minnesota review [1]:
I learned a lot from myth criticism [referring to Northrup Frye], especially the way little details in a Shakespeare play can link up to indicate an “underthought” of reference to some myth or other. It was something I had learned in a different way from Burke. Burke came to Harvard when I was a graduate student and gave a lecture about indexing. What he was talking about was how you read. I had never heard anybody talk about this. He said what you do is notice things that recur in the text, though perhaps in some unostentatious way. If something appears four or five times in the same text, you think it's probably important. That leads you on a kind of hermeneutical circle: you ask questions, you come back to the text and get some answers, and you go around, and pretty soon you may have a reading.
That’s it, noticing patterns in texts, unostentatious patterns. That’s where the good stuff is.

Derrida was good at it as well. Referring to a passage in Remembrance of Things Past:
What Derrida did that I never would have thought of was to notice that the whole passage is based on words in pris: apprendre, comprendre, prendre. Those words are perfectly translatable, but lose their play on pris—I understood: j'ai compris. Derrida noticed these words and their recurrence in a way that helps you to understand the way the passage is put together and the meaning it has. Derrida was a genius in doing that sort of reading. That's why Derrida for me is even more important for his way of reading than for his invention of big concepts like différance.
There’s that notion again, “recurrence”. And why would Miller think Derrida’s ability to spot such patterns is more important that his ability to invent “big concepts”? Could it be that (at least some of) those big concepts are about what you uncover as you come to terms those unobtrusive patterns. If you’re out to deconstruct something, if that’s your critical game, those patterns tell you where to start pulling on the threads to unravel the cloth.

But how do you LEARN to do that, notice interesting things, diagnostic signs, in texts? Surely you learn it by doing it. A teacher works though it in lecture, and prods and nudges during class discussion. You read examples. You write your own examples, getting feedback from teachers and more experienced students. Getting good at it takes years. There are no shortcuts.

Markings and patterns

And that’s what I was doing, largely for myself, during those months and years working on “Kubla Khan” decades ago at Johns Hopkins. By that time I had had four years of undergraduate instruction in various subjects, including literature. I’d written, say, a dozen or more critical papers, and gotten them critiqued. I had the example of what Lévi-Strauss had done with myth, in The Raw and the Cooked, and in various essays, including one on some Winnebago myths. And there was some essays about poems as well, by Lévi-Strauss, Roman Jakobson, and Michael Riffaterre. But it was mostly Lévi-Strauss on myth.

That got me started. But it wasn’t enough to bring me home. I made worksheets, worksheet after worksheet. I’d type out the text of “Kubla Khan” in double-space or triple-space and then mark it up. I must have done that half a dozen times or so. Here’s a fragment from one such sheet:

KK-working-text-old-2cr72.jpg

Notice that I’ve numbered the lines (down the left edge) and that I’ve made various kinds of notations in three, maybe four colors of in, black, blue, red, and I believe green (running vertically in the left margin). I’ve underlined and circles various words and phrases, used various lines and brackets to connect things across lines – in this I was inspired by an illustration in the notes to “Kubla Khan” in the edition I was working from, edited I believe by Kathleen Coburn. And I’ve written various kinds of comments all over the place. Some comments seem descriptive: “spatial limitation” (line 6), “enclosing fertility” (7), “water” (8), “radiating” (9), and “earth” (10). Others are more interpretive: “music takes one down & away” (to the right of ll. 3 & 4), “Nature” (on the rightmost red bracket), and “Man as object” (line 1). Notice, though, that the interpretation is rather “shallow”. I’m not trying to decode symbols. I’m classifying and organizing. I’m looking for patterns.

Monday, September 24, 2018

Still, why do literary critics find it so difficult to focus on form?

A topic I’ve been thinking about off and on for some time, most recently: Once more, and thinking of ring composition: Why aren’t literary critics interested in describing literary form?

A meaning-focused discipline

It’s surely that literary criticism has focused on meaning. But why should that distract from an interest in form, especially since so-called formalism looms large in methodological discussions and in practical criticism? That is, why does formalism have so little to do with actual form? That’s the question.

Of course, form in the sense that I mean, isn’t completely invisible. It’s quite visible in formal verse, where patterns of rhyme, meter, and so forth have been extensively catalogued. This is, however, a relatively peripheral matter for literary critics, though it may be very important to some working poets.

Verse forms are visible because they are forms of sound, and sound is readily objectified. What of ring forms? Well, in the small, they’re known as chiasmus, and are well known and often remarked. But chiasmus often involves symmetrical arrangements of sound. Large-scale ring forms do not. They’re more difficult to identify, and certainly more problematic.

That they are large scale, encompassing the whole work, is one source of difficulty. You have compare features across the whole text, that’s much more difficult than comparing features within a line or a few lines of poetry. Moreover it’s not obvious what features you should be comparing. It’s not sound. And objectifying subject matter is trickier, no?

The road to form

In my own experience, it takes a fair amount of tedious work to identify ring-form structures. I can think of one case where I suspected a ring-form at the outset, Tezuka’s Metropolis, and another where I started with a clue, David Bordwell’s remark about symmetry in King Kong. In both of these cases it took me some hours of work over a day or three to conduct the analysis. And then there’s Gojira, where I’d worked on the film off and on over a couple of years before I suspected it might be a ring composition. And then, again, it was hours of work to verify my suspicion.

I note moreover that I did this work after, long after, I had adopted a computational view of literary texts. Indeed, I did most of this work after my 2006 paper on literary morphology [1]. By that time I had, of course, done the work on “Kubla Khan” and on Tezuka’s Metropolis. But it wasn’t until several years later I ran up a post about “The Nutcracker Suite” and “sorcerer’s Apprentice” episodes of Fantasia [2], though I must have worked on “The Nutcracker Suite” in 2006 or 2007 as I’d written a long email to Mary Douglas about it and she died in May of 2007.

My point is simply that I hadn’t begun this work in a systematic way until I’d found a example or two by the by and until I had a theoretical reason – computational form – to look for them. I’d been driven to computation by “Kubla Khan” years ago, after I found those nested structures that must “smelled” like computation.

KK-triple-72.jpg
Nested structure in “Kubla Khan”, ll. 1-36.
And it was easy and natural enough to extend computation to Lévi-Strauss’s notion of the armature. And finally, having adopted a computational view of language, it followed that literature must have a computational aspect as well. The remaining issue is whether or not literary texts display large-scale computational structures rather than simply being a concatenation of sentence-level structures. The existence of ring-form texts suggests that they do.

Monday, September 3, 2018

Computational linguistics & NLP: What’s in a corpus? – MT vs. topic analysis [#DH]

What’s in a corpus? Words, words organized into texts. Of course.

But that obvious answer not quite what I’m after. I’m interested in how we think about corpora, their role in our work. “We”, who’s that? I’m not sure it much matters, not exactly. It will emerge.

How are these corpora connected to the world? What do we hope to understand about the world by analyzing these corpora? What are our intuitions in these matters. To some extent I’m trying to track down something I don’t know how to conceptualize. This post from October 20, 2017 is a good example of that:
Borges redux: Computing Babel – Is that what’s going on with these abstract spaces of high dimensionality? [#DH], http://new-savanna.blogspot.com/2017/10/borges-redux-computing-babel-is-that.html
But when stalking such an abstract beast, it helps to have specific examples in mind. So I’m thinking about the role of corpora in statistical machine translation (MT) vs. their role in topic analysis. Roughly speaking, in MT statistical analysis of corpora are is a means to an end. In topic analysis statistical analysis of a corpus is the end. That difference entails a somewhat different way of thinking about corpora.

Martin Kay, an “ignorance model”

I’m basing this post on some observations by Martin Kay, one of the grand old men of MT. Kay apprenticed with Margaret Masterman at the Cambridge Language Research Unit in the 1950s. In 1951 David Hays hired him to work with the RAND group in MT. He went on to a distinguished career in computational linguistics at the University of California, Irvine, the Xerox Palo Alto Research Center, and Stanford.

Research in MT was undertaken to achieve a practical end, the translation of texts from one language to another. The United States government was particularly interested in obtaining translation of Russian texts. The researchers who undertook this work had various motivations, but some of them were interested in linguistic science and were happy enough to have their work funded by a government agency, the Department of Defense, with a practical goal.

In 2005 the Association for Computational Linguistics gave Kay a Lifetime Achievement Award. On that occasion he looked back over his career and made some observations about the relative merits of statistical and symbolic approaches to MT [1]. He speaks as a man fundamentally interested in basic knowledge who has, however, at times undertaken work with practical ends.

At the beginning of the following passage Kay distinguishes between computational linguistics and natural language processing (NLP). The distinction is a common one, albeit a bit problematic as well [2]. But the distinction Kay makes is clear enough (p. 5):
Computational linguistics is not natural language processing. Computational linguistics is trying to do what linguists do in a computational manner, not trying to process texts, by whatever methods, for practical purposes. Natural Language Processing, on the other hand, is motivated by engineering concerns. I suspect that nobody would care about building probabilistic models of language unless it was thought that they would serve some practical end. There is nothing unworthy in such an enterprise. But ALPAC’s conclusions are as true today as they were in the 1960’s—good engineering requires good science. If one’s view of language is that it is a probability distribution over strings of letter or sounds, one turns one’s back on the scientific achievements of the ages and foreswears the opportunity that computers offer to carry that enterprise forward.
I agree with Kay’s fundamental point, though I note that humanists using NLP techniques are often pursuing basic knowledge rather than a practical end.

Kay wrote that passage in 2005. I don’t know just when literary critics first began exploring NLP techniques, but I first became aware of digital humanities work in topic modeling sometime in 2012 [3]. That’s well after Kay wrote those words.

Let’s return to his remarks. Nearing the end of his talk, Kay remarks (p. 12):
My professional life almost encompasses the history of computational linguistics. But I was only fourteen when Warren Weaver wrote his celebrated memorandum drawing a parallel between machine translation and code breaking. He said that, when he saw a Russian article, he imagined it to be basically in English, but encrypted in some way. To translate it, what we would have to do is break the code and the statistical techniques that he and others had developed during the second world war would be a major step in that direction. However, neither the computer power nor large bilingual corpora were at hand, and so the suggestions were not taken up vigorously at the time. But the wheel has turned, and now statistical approaches are pursued with great confidence and disdain for what went before. In a recent meeting, I heard a well known researcher claim that the field had finally come to realize that quantity was more important than quality.

The young Turks blame their predecessors, the advocates of so-called symbolic systems, for many things. Here are just four of them. First, symbolic systems are not robust in the sense that there are many inputs for which they are not able to produce any out- put at all. Second, each new language is a new challenge and the work that is done on it can profit little, if at all, from what was done previously on other languages. Third, symbolic systems are driven by the highly idiosyncratic concerns of linguists rather than real needs of the technology. Fourth, linguists delight in uncovering ambiguities but do nothing to resolve them. This is actually a variant of the third point.
Kay mounts a quick defense on the first three points, but says a bit more about the fourth, ambiguity (pp. 12-13):
This, I take it, is where statistics really come into their own. Symbolic language processing is highly nondeterministic and often delivers large numbers of alternative results because it has no means of resolving the ambiguities that characterize ordinary language. This is for the clear and obvious reason that the resolution of ambiguities is not a linguistic matter. After a responsible job has been done of linguistic analysis, what remain are questions about the world. They are questions of what would be a reasonable thing to say under the given circumstances, what it would be reasonable to believe, suspect, fear or desire in the given situation. If these questions are in the purview of any academic discipline, it is presumably artificial intelligence. But artificial intelligence has a lot on its plate and to attempt to fill the void that it leaves open, in whatever way comes to hand, is entirely reasonable and proper. But it is important to understand what we are doing when we do this and to calibrate our expectations accordingly. What we are doing is to allow statistics over words that occur very close to one another in a string to stand in for the world construed widely, so as to include myths, and beliefs, and cultures, and truths and lies and so forth. As a stop-gap for the time being, this may be as good as we can do, but we should clearly have only the most limited expectations of it because, for the purpose it is intended to serve, it is clearly pathetically inadequate. The statistics are standing in for a vast number of things for which we have no computer model. They are therefore what I call an “ignorance model”.
That last point is very important.

By the mid-to-late 1960s computational linguists began to realize that they would have to tackle semantics, which covers the relationship between language and the world, in contrast to syntax, morphology, and phonology, which are all internal to the language system. And so computational linguists did that, along with cognitive psychologists and researchers in artificial intelligence. The work then, and now, was interesting and fruitful, but not terribly useful for practical tasks, such as MT. By the 1990s the conjunction of large amounts of cheap computing power and large bodies of digital texts gave statistical approaches a definitive edge in practical applications, an edge which remains.

Friday, July 13, 2018

Training your mind, Michael Nielsen on Anki and human augmentation

Michael Nielsen, an AI researcher at Y Combinator Research, has written a long essay, Augmenting Long-term Memory, which is about Anki, a computer-based tool for training long-term memory.
In this essay we investigate personal memory systems, that is, systems designed to improve the long-term memory of a single person. In the first part of the essay I describe my personal experience using such a system, named Anki. As we'll see, Anki can be used to remember almost anything. That is, Anki makes memory a choice, rather than a haphazard event, to be left to chance. I'll discuss how to use Anki to understand research papers, books, and much else. And I'll describe numerous patterns and anti-patterns for Anki use. While Anki is an extremely simple program, it's possible to develop virtuoso skill using Anki, a skill aimed at understanding complex material in depth, not just memorizing simple facts.

The second part of the essay discusses personal memory systems in general. Many people treat memory ambivalently or even disparagingly as a cognitive skill: for instance, people often talk of “rote memory” as though it's inferior to more advanced kinds of understanding. I'll argue against this point of view, and make a case that memory is central to problem solving and creativity. Also in this second part, we'll discuss the role of cognitive science in building personal memory systems and, more generally, in building systems to augment human cognition. In a future essay, Toward a Young Lady's Illustrated Primer, I will describe more ideas for personal memory systems.

The essay is unusual in style. It's not a conventional cognitive science paper, i.e., a study of human memory and how it works. Nor is it a computer systems design paper, though prototyping systems is my own main interest. Rather, the essay is a distillation of informal, ad hoc observations and rules of thumb about how personal memory systems work. I wanted to understand those as preparation for building systems of my own. As I collected these observations it seemed they may be of interest to others. You can reasonably think of the essay as a how-to guide aimed at helping develop virtuoso skills with personal memory systems. But since writing such a guide wasn't my primary purpose, it may come across as a more-than-you-ever-wanted-to-know guide.

To conclude this introduction, a few words on what the essay won't cover. I will only briefly discuss visualization techniques such as memory palaces and the method of loci. And the essay won't describe the use of pharmaceuticals to improve memory, nor possible future brain-computer interfaces to augment memory. Those all need a separate treatment. But, as we shall see, there are already powerful ideas about personal memory systems based solely on the structuring and presentation of information.
The method of loci is well-known, and I'm sure you can come up with a lot of information just by googling the term. You might even come up with my encyclopedia article, Visual Thinking, where I treat it as one form of visual thinking among others.

Before returning to Nielson and Anki, I want to digress to a different form of mental training. When I was young people didn't have personal computers, nor even small hand-held electronic calculators. If you had to make a lot of calculations, you might have used a desktop mechanical calculator, a slide rule–my father had become so fluent with his that he didn't even have to look at it while doing complex multi-step calculations, or you might have mastered mental calculation.

Some years ago I reviewed biographies of John von Neumann and Richard Feynman; both books mentioned that their subjects were wizards of mental calculation. I observe:
Feynman and von Neumann worked in fields were calculational facility was widespread and both were among the very best at mental mathematics. In itself such skill has no deep intellectual significance. Doing it depends on knowing a vast collection of unremarkable calculational facts and techniques and knowing one's way around in this vast collection. Before the proliferation of electronic calculators the lore of mental math used to be collected into books on mental calculation. Virtuosity here may have gotten you mentioned in "Ripley's Believe it or Not" or a spot on a TV show, but it wasn't a vehicle for profound insight into the workings of the universe.

Yet, this kind of skill was so widespread in the scientific and engineering world that one has to wonder whether there is some connection between mental calculation, which has largely been replaced by electronic calculators and computers, and the conceptual style, which isn't going to be replaced by computers anytime soon. Perhaps the domain of mental calculations served as a matrix in which the conceptual style of Feynman, von Neumann, (and their peers and colleagues) was nurtured.
Then, citing the work of Jean Piaget, I suggest every so briefly why that might be so. However, once powerful handheld calculators became widely available, skill in mental calculation was no longer necessary. These days one may hear of savants who have such skills, but that's pretty much it.

Returning to Nielsen and Anki, as his essay evolves, he suggests that more than mere memory is at stake. After explaining Anki basics he describes how he used Anki to help him learning enough about AlphaGo–the first computer system that beat the best human experts at Go–to write an article for Quanta Magazine. Alas
I knew nothing about the game of Go, or about many of the ideas used by AlphaGo, based on a field known as reinforcement learning. I was going to need to learn this material from scratch, and to write a good article I was going to need to really understand the underlying technical material.
He then explains what he did. The upshot:
This entire process took a few days of my time, spread over a few weeks. That's lot of work. However, the payoff was that I got a pretty good basic grounding in modern deep reinforcement learning. This is an immensely important field, of great use in robotics, and many researchers believe it will play an important role in achieving general artificial intelligence. With a few days work I'd gone from knowing nothing about deep reinforcement learning to a durable understanding of a key paper in the field, a paper that made use of many techniques that were used across the entire field. Of course, I was still a long way from being an expert. There were many important details about AlphaGo I hadn't understood, and I would have had to do far more work to build my own system in the area. But this foundational kind of understanding is a good basis on which to build deeper expertise.
He then explains how he he used Anki to do shallow reads of papers. I'm not going excerpt or summarize that material, but I'll point out that doing shallow reads is a very useful skill. When I was in graduate school I prepared abstracts of the current literature for The Journal of Computational Linguistics. While some articles and tech reports had good abstracts, many did not. In those cases I'd have to read the article and write an abstract; I gave myself an hour, perhaps a bit more to write a 250-word abstract. I gave those articles a shallow read. How'd I do it? Hmmmm... I'll get back to you on that. It's quite possible that Nielsen's Anki process is better than the one I used.

Yet:
Really good resources are worth investing time in. But most papers don't fit this pattern, and you quickly saturate. If you feel you could easily find something more rewarding to read, switch over. It's worth deliberately practicing such switches, to avoid building a counter-productive habit of completionism in your reading. It's nearly always possible to read deeper into a paper, but that doesn't mean you can't easily be getting more value elsewhere. It's a failure mode to spend too long reading unimportant papers.
My process was certainly good enough to make that go-nogo decision.
Nielson then goes on to discuss this and that use of Anki, suggesting:
Anki isn't just a tool for memorizing simple facts. It's a tool for understanding almost anything. It's a common misconception that Anki is just for memorizing simple raw facts, things like vocabulary items and basic definitions. But as we've seen, it's possible to use Anki for much more advanced types of understanding. My questions about AlphaGo began with simple questions such as “How large is a Go board?”, and ended with high-level conceptual questions about the design of the AlphaGo systems – on subjects such as how AlphaGo avoided over-generalizing from training data, the limitations of convolutional neural networks, and so on.

Part of developing Anki as a virtuoso skill is cultivating the ability to use it for types of understanding beyond basic facts. Indeed, many of the observations I've made (and will make, below) about how to use Anki are really about what it means to understand something.
That's the good stuff.

Where's he going? Human augmentation:
The human-computer interaction (HCI) community has tried to achieve it in the systems they build, not just for memory, but for augmenting human cognition in general. But I don't think it's worked so well. It seems to me that they've given up a lot of boldness and imagination and aspiration in their design** As an outsider, I'm aware this comment won't make me any friends within the HCI community. On the other hand, I don't think it does any good to be silent, either. When I look at major events within the community, such as the CHI conference, the overwhelming majority of papers seem timid when compared to early work on augmentation. It's telling that publishing conventional static papers (pdf, not even interactive JavaScript and HTML) is still so central to the field. . At the same time, they're not doing full-fledged cognitive science either – they're not developing a detailed understanding of the mind. Finding the right relationship between imaginative design and cognitive science is a core problem for work on augmentation, and it's not trivial.

In a similar vein, it's tempting to imagine cognitive scientists starting to build systems. While this may sometimes work, I think it's unlikely to yield good results in most cases. Building effective systems, even prototypes, is difficult. Cognitive scientists for the most part lack the skills and the design imagination to do it well.

This suggests to me the need for a separate field of human augmentation. That field will take input from cognitive science. But it will fundamentally be a design science, oriented toward bold, imaginative design, and building systems from prototype to large-scale deployment.
* * * * *

See also my post, Beyond "AI" – toward a new engineering discipline, in which I excerpt Michael Jordan, "Artificial Intelligence — The Revolution Hasn’t Happened Yet". Jordan discusses human augmentation under the twin rubrics of "Intelligence Augmentation" and "Intelligence Infrastructure".