Showing posts with label cognitive science. Show all posts
Showing posts with label cognitive science. Show all posts

Sunday, August 16, 2026

An example of common sense reasoning: “For sale: baby shoes, never worn.”

Yesterday day I had a short post about the difference between a narrative and a narrative that is also a story. One of my examples was a statement that’s been circulating as a meme for some time now: “For sale: baby shoes, never worn.” Strictly speaking that’s not a story nor even a narrative. It’s a statement that something is for sale, namely baby shoes that have never been worn. However the statement is of such a nature that we can easily infer a story that led to that statement. That inference is an example of common sense reasoning as the concept is understood in artificial intelligence and computational linguistics.

Why common sense? Because no special knowledge is required. Anyone growing up under the appropriate circumstances will possess the knowledge required to infer such a story.

First, one must recognize that it is advertising something for sale. The words, “for sale,” could simply be a statement of fact. But we recognize that such a statement is also an offer, “something is being offered for sale.” The fact that that inference is utterly trivial doesn’t negate the fact that an inference is required to go from the statement of act to the offer. However, if one sees that statement in a newspaper in the appropriate section, its presence in that section constitutes the offer.

Now, just what story are we inferring from the statement? The most likely story is that a couple bought the shoes in anticipation of the birth of a child. The infant died shortly after birth, making the baby shoes useless. Hence they are available for sale, never having been worn.

All of that has to be inferred, the inferences are trivial, but still, they are inferences. Let’s tease things apart. Here’s one possibility:

  1. When an infant is expected, preparations must be made.
  2. Preparations include buying clothing.
  3. Shoes are an item of clothing.
  4. Once the infant is born, the items of clothing will be used.
  5. If however the infant is stillborn, or dies after birth, the items of clothing will not be worn.
  6. In that case something must be done with those items.
  7. They could be put up for sale.
  8. That would lead to placing an advertisement in a newspaper.
  9. The statement “For sale: baby shoes, never worn” is such an advertisement.

You could imagine other circumstances, but that seems to me the most likely case. But whatever alternative circumstances you propose, the point is that they must be inferred from the statement. They are not explicit in it. Ordinary language requires scads of such inferences. That requirement is the problem of common sense knowledge as it exists in artificial intelligence.

Saturday, August 15, 2026

What’s the difference between a narrative and a story?

Starting in the 1960s and continuing well into the 1980s (and beyond) we see a renewed interest in stories and narrative across a variety of disciplines, literary criticism (including narratology), folklore, cognitive psychology and artificial intelligence. Vladimir Propp’s Morphology of the Folktale was revived by the structuralists, George Lakoff recast it in the terms of formal linguistics in an essay, “Structural Complexity in Fairy Tales” (The Study of Man, 1, 1972), and everyone was reading Lévi-Strauss on myth. One of the questions that came out of this research is simple: What’s the difference between a mere narrative and a story? By the time I stopped following that research in the mid-1980 the question was still open.

According to an informal review by Mark Riedl that’s where things stood as of August 2021, “An Introduction to AI Story Generation” (The Gradient). He defines three key concepts:

  • Narrative: The recounting of a sequence of events that have a continuant subject and constitute a whole (Prince, 1987). An event describes some change in the state of the world. A “continuant subject” means there is some relationship between the events—it is about something and not a random list of unrelated events. What “relates” events is not entirely clear but I’ll get to that later.

  • Story: A narrative that tells a story has certain properties that one comes to expect. All stories are narratives, but not all narratives are stories. Unfortunately I cannot point to a specific set of criteria that makes people regard a narrative as a story. One strong contender, however, is a structuring of events in order to have a particular effect on an audience.
  • Plot: A plot is the outline of main incidents in a narrative.
  • Let’s consider a very simple kind of narrative: someone leaves home, travels to the grocery story, purchases some items, and returns home. An actual narrative would mention a specific individual, specific places, events along the way, items purchase, and so forth. By Prince’s definition that’s a narrative because there has been a change of state in the world, several in fact depending on how you count them. Items have been taken from the store and transported home (involving two state changes per item) and money has been transferred from our “protagonist” to the owner of the store (two state changes).

    I doubt, however, that anyone would consider such a narrative to be a story. Now consider a Japanese reality show that has been shown on Netflix, Old Enough! Young children, three to six years old, are given simple errands to perform. Going to the store, purchasing something, and returning with it is a typical errand. I have watched a number of these and would consider them stories. Why? Because the errands are challenging for the children. They’ve never done it before, there is no line of sight between home and the story, they have to remember the items, remember to ask for change, and so forth. It’s a challenge for them. We see them confronting and overcoming obstacles, including simply fear. That makes it a story proper.

    Note, however, that I watched a video of the journey, one edited to amplify the challenges the children face.  That is to say, the events were edited “in order to have a particular effect on an audience.” It would be easy to write these stories in such simple and reduced form that the “storyness” is completely eliminated, leaving only the bare narrative husk. But someone with only modest story-telling skills could also write these stories in a from that includes accounts of the challenges faced by the children, thus making a story out of a bare narrative.

    Now consider a six-word statement that’s been mistakenly attributed to Ernest Hemingway: “For sale: baby shoes, never worn.” However, versions of that statement have appeared as early as 1883. Those six words don’t even give us an event, and yet we seem willing to treat them as telling a story. Why? Because those words lead us to infer something like: the baby shoes were purchased in anticipation of the birth of a child; the child was either still born or died soon after birth; the shoes are no longer needed. Question: At what age is a person sophisticated enough to make such an inference from those words? I doubt that a six-year old would do it.

    What, then, is the difference between a narrative and a story? It depends, doesn’t it? How are the events narrated and what can we expect the audience to infer from what is said? It’s not an easy question to answer.

    Monday, July 27, 2026

    Behavioral similarities in the way chatbots and oral poets perform

    Kush R. Varshney, An Annotated Reading of ‘The Singer of Tales’ in the LLM Era, https://arxiv.org/html/2502.05148v1 Feb. 2025.

    Abstract. The Parry-Lord oral-formulaic theory was a breakthrough in understanding how oral narrative poetry is learned, composed, and transmitted by illiterate bards. In this paper, we provide an annotated reading of the mechanism underlying this theory from the lens of large language models (LLMs) and generative artificial intelligence (AI). We point out the the similarities and differences between oral composition and LLM generation, and comment on the implications to society and AI policy.

    Varshney develops his argument by interlacing passages from Albert Lord's The Singer of Tales with comments on LLMs. This is a very interesting way of reviewing your understanding of LLMs in relation to a specialized kind human language performance.

    You might want to consider two of my blog posts:

    GPT-3, the phrasal lexicon, Parry/Lord, and the Homeric epics, July 16, 2022.

    In some ways, some contexts, LLMs may provide a useful model for human language, March 24, 2026.

    In this more recent post I discuss empirical evidence about human memory for F.C. Bartlett's classic book, Remembering: A Study in Experimental and Social Psychology (1932), David C. Rubin, Memory in Oral Traditions: The Cognitive Psychology of Epic, Ballads, and Counting-out Rhymes (Oxford 1995).

    What I did last week: aesthetics, economics, Rorschach analogy for AI, default images, and “leveling”

    I did some satisfying work last week. Here’s a quick rundown. I’m listing the posts in the order I wrote them.

    Visual Aesthetics

    A case of visual aesthetics: Why is the monochrome image superior to the color image?

    The issue, black & white vs. color, has been and I suppose remains central to photography, and I deal with it there, a bit. But that’s not what I’m doing here. This is about the conversion of a particular ChatGPT image from color to black & white. It was a fun post to assemble and to think about. I like the suite of images.

    Rank 5 Economics?

    Beyond Marginalism: What’s Next? [MR #12]

    This is my last word – save for an introduction I’ll write in a week or three, who knows? – on the fourth and final chapter of Cowen’s monograph on marginalism. This is where he tosses up some examples of leading edge work in economics, noting that it’s drifting away from marginalism into complex high-dimensional models created through machine learning. His examples come from finance. The new models yield better predictions.

    I focus on one model that has 360,000 parameters and end up making (speculative) sense out of what’s going on. I suggest that those parameters are picking up the effects of Keynes’ “animal spirits” as expressed in the gossip and stories of Schiller’s narrative economics. I further suggest that we can test this by comparing the output of a classical model with that from a high-parameter machine learning model. The divergence should be highest with those stocks otherwise identified as meme stocks.

    The prospect of empirical investigation into animal spirits in asset pricing [the fate of marginalism]

    Here I take my speculations about how to test these high parameter models and present them to Marge, the AI associated with Cowen’s book. Marge approves.

    Rorschach test for AI

    More on how I’m approaching The God Test – Rorschach! [GT-2]

    I came up with the Rorschach blot as analogy for the kind of challenge AI presents to us, to our understanding of AI and of the future. The idea is that the blot does have a form, albeit a complex one that’s not very legible. Hence our commentary on it (that is, on AI) tells as much about us as about AI. I’ll be developing this further in a later post.

    Prototype Image in ChatGPT

    This is a new working paper that opens up a whole new line of investigation. This was a fun piece of work. Writing it up took way longer than actually generating the images.

    A prototypical image in ChatGPT 5.6: An informal pilot study

    Abstract: Previous work has found strong default preferences in stories generated by large language models from minimally specified prompts. To determine whether a similar effect appears in image generation, I asked ChatGPT 5.6 to create 17 images in separate chats using four prompts that specified either no subject matter or only a rendering medium: “Create an image,” “Create a drawing,” “Create a painting,” and “Create a water color painting.” Sixteen of the 17 outputs depicted closely related landscapes containing mountains, trees, sky, and water; the remaining image depicted a lighthouse. The images also shared a calm, picturesque mood and contained no human figures, although several included signs of human habitation. Because the internal prompt passed to the image generator was not available, the study cannot determine whether these defaults arise primarily in the language model, the image generator, or their interaction. The results are exploratory but suggest that severely underspecified image prompts may reveal stable default preferences in the integrated ChatGPT image-generation system.

    The “leveling” of knowledge in the compressed form of LLMs

    NYTimes: AI needs human supervision in order to complete an entire job.

    This is something I’ve been thinking about off and on for a while, but this is my first explicit framing of the issue. The idea is that once ideas or set of ideas has been expressed in writing and those documents then consumed into an LLM, all ideas function the same within/through/for the model. In that post I’m comparing a study of using AI to perform routine office processes (from NYTimes) with the use of AI to perform a complex set of tasks in drug development, in effect, high-school level capability with Ph.D. level capability. They’re the same to the LLM.

    I need to think about this some more. It seems to me what’s nowhere present in the LLM is the kind of procedural knowledge necessary to learn tasks at whatever level. That simply isn’t presented in the written products of that knowledge (not even in written procedures).

    Sunday, July 26, 2026

    ChatGPT draws Rorschach blots

    Here they are, six of them:

    I then discussed the blots with ChatGPT in terms suggested by me new working paper, A prototypical image in ChatGPT 5.6: An informal pilot study. First I present all six blots, full size, and then the discussion.

    I needed a Rorschach inkblot for a post I was working on, so I had ChatGPT create one. I then decided to see what would happen if I asked for a number of inkblots, each in a different chat. I’ve uploaded the result to the project in a document entitled Rorschach Blots. I had no expectation about what would happen. Would it produce the same blot each time? Of different blots each time? As you can see, it produced different blots each time.

    This suggests two further lines of investigation. 1) Generate 100 independent ink blots and run them through an image classifier. Would we get distinct clusters or not? 2) Use prompts suggesting specific things, but otherwise open-ended, e.g. “Draw a cat,” “Draw a tree,” “Draw a car,” “Draw a coat,” etc.

    Looking at the Rorschach Blots document, the six outputs are plainly different, but they are not unconstrained in every respect. All six preserve a common structural grammar: bilateral symmetry, a strong vertical axis, black-and-gray ink on white, irregular feathered edges, and enough internal articulation to invite figurative interpretation. Within that shared format, however, the overall silhouettes vary substantially: some are tall and narrow, some compact and squat, some strongly winged, some mask-like, and some more diffuse. The contact sheet on page 1 makes that balance of sameness and difference especially clear.

    That suggests that the system has not selected one prototypical blot, in the way it selected one broad landscape type. Instead, it seems to have selected a prototypical design space for “Rorschach blot,” within which it samples different instances. Your two proposed extensions are therefore well chosen.

    1. One hundred independent inkblots

    The clustering question is exactly the right one. A hundred examples might reveal:

    • one continuous cloud of variation around a single template;
    • several recurring morphological families;
    • or a few dominant families plus outliers.

    Even by eye, the six examples hint at possible families: elongated “totemic” forms, broad winged forms, compact mask-like forms, and more open forms with substantial white interior space. But six is far too few to know whether those are genuine modes or just impressions.

    I would not rely on a conventional object classifier, since it may simply label them all “inkblot,” “butterfly,” “mask,” or “bat.” A more informative analysis would use image embeddings and cluster those embeddings, perhaps supplemented by simple measurable features such as aspect ratio, occupied area, contour complexity, amount of internal white space, and vertical distribution of mass. The point would not be to decide what each blot “looks like,” but to determine the structure of variation across the set.

    2. Open-ended object prompts

    “Draw a cat,” “Draw a tree,” “Draw a car,” and so forth would probe something different and potentially more revealing: whether the system has a default realization for a named category.

    The questions would include:

    • Does “Draw a cat” repeatedly produce the same breed, pose, viewpoint, or expression?
    • Does “Draw a tree” default to a solitary deciduous tree, a stylized botanical specimen, or a landscape scene?
    • Does “Draw a car” default to a modern sports car, a generic sedan, or a three-quarter advertising view?
    • Does “Draw a coat” produce the garment alone, on a mannequin, or worn by a person?

    That experiment would occupy a useful middle ground between the wholly unspecified prompt—“Create a drawing”—and a detailed scene description. The subject is fixed, but almost everything about its realization remains open. It would therefore reveal category-level defaults in pose, style, composition, context, and affect.

    The Rorschach result also suggests a distinction worth preserving in the paper’s future-work section:

    Some prompts may elicit a single dominant subject prototype, while others may elicit a constrained family of forms rather than one recurring image.

    That is the most interesting implication of the blot experiment. The landscape prompt collapsed onto a narrow semantic prototype. The inkblot prompt produced morphological diversity within a narrow formal convention. Those are different kinds of regularity, and comparing them could become a substantial research program.

    Friday, July 17, 2026

    The language of thought is not natural language

    Hope Kean, Alexander Fung, Paris Jaggers, +6 , and Evelina Fedorenko, Evidence from formal logical reasoning reveals that the language of thought is not natural language, PNAS, 123 (28) e2520095123 https://doi.org/10.1073/pnas.2520095123, July 6, 2026.

    Significance: Which cognitive mechanisms allow humans to reason logically, to understand whether a conclusion follows from the premises? Are they the same ones that allow the assembly of words into structured representations? Scholars have debated for millennia whether logical reasoning is inextricably tied to natural language, or instead relies on a distinct “language of thought” (LOT). Using fMRI in healthy adults and evaluating logical ability in individuals with severe aphasia, we find that distinct neural systems support language processing vs. logical (inductive and deductive) reasoning. These results suggest that, at least in mature brains, language processing does not underpin logical inference, perhaps due to the distinct representational format of the logical LOT.

    Abstract: Humans are endowed with a powerful capacity for inductive and deductive logical thought: we easily form generalizations based on a few examples and draw conclusions from known premises. Humans also arguably have the most sophisticated communication system in the animal kingdom: natural language allows us to express complex and structured meanings. Some have therefore argued for a tight relationship between complex thought and language, postulating that reasoning, including logical reasoning, relies on linguistic representations. We systematically investigated the relationship between logical reasoning and language using two complementary approaches. First, we used noninvasive brain imaging (fMRI) to examine neural activity as healthy adults engaged in logical reasoning tasks. And second, we behaviorally evaluated logical abilities in individuals with extensive lesions to the language brain areas and consequent severe linguistic impairment. Our findings reveal that the language brain network is not engaged during logical reasoning, and patients with severe aphasia exhibit intact performance on logic tasks. Instead, inductive reasoning recruits the domain-general multiple demand network implicated broadly in goal-directed behaviors, whereas deductive reasoning draws on brain regions that are distinct from both the language and the multiple demand networks. Together, these results indicate that linguistic representations are neither utilized nor required for inductive or deductive logical reasoning.

    H/t Daniel Everett.

    Perceptions of probability

    Thursday, July 16, 2026

    Attention and error predition in thalamocortico circuits

    Follow the link to see the full thread. Here's the article's abstract:

    Prediction errors (PEs) drive perceptual learning by updating internal models of the sensory environment, yet it remains unclear how attention reshapes their representation across distributed thalamocortical circuits. Using intracranial stereoelectroencephalography (sEEG) from 17 patients performing a roving auditory oddball task under attended and unattended conditions, we quantified PE encoding using mutual information and co-information to capture redundant and synergistic PE representations. Attention modulated PE encoding in both the thalamus and the temporal cortex, but with distinct informational dynamics. Thalamic encoding showed a stable reduction of PE information during distraction, consistent with state-dependent thalamocortical gating. In contrast, the temporal cortex expressed two opposing learning trajectories during attended listening that converged once attention was diverted, revealing distinct cortical learning regimes rather than a uniform attentional effect. Attention further reorganized the informational content of cortical PE representations by altering the balance between redundant and synergistic information. A biologically constrained neural network showed that attention-dependent changes in inhibition and long-range connectivity reproduced these dynamics through Hebbian learning. Together, these findings suggest that attention regulates predictive learning not simply by changing the strength of PE responses, but by reshaping how distributed thalamocortical circuits represent and integrate sensory evidence over time.

    Thursday, July 2, 2026

    Synaptic pruning in the nervous system

    Monday, June 22, 2026

    Brain area specialized for visual recognition of words

    Sunday, June 21, 2026

    New Book Project: Language, Memory, and Mind: A Supplement to The Computer and the Brain

    As you may know, I’ve been working on a book project, Play: How to Stay Human in the A.I. Revolution. For some reason I’ve been unable to finish the proposal, though I’ve got lots of stuff and a number of the chapters are substantially drafted. But I keep finding myself distracted into thinking about basics, very basic things about computing and A.I.

    At the very end of his life, John von Neumann wrote a slim book, The Computer and the Brain (1958). It grapples with the problem of how computation can be implemented in a physical medium and does so in a way that is basic, both simple and straightforward and profound. We’ve learned a great deal about both the brain and the computer since then, but as far as I know, no one has revisited von Neumann’s project and extended it to include what we have since learned. That’s what I propose to do in this book.

    Now, I have no intention of trying to summarize what we’ve learned on those two topics since 1958. That’s working at the wrong level. When von Neumann was writing he, and by extension, we, had no conception of distributed representation much less how it could be achieved physically. Now we do. That’s what needs to be added to von Neumann’s exposition.

    I have no intention of repeating what von Neumann did. In particular, I will not revisit his material on analog computing. Rather, I want to augment his discussion. Fortunately the new material is of such a nature that I should be able to write short book that can be read as a stand-alone discussion or as a supplement to von Neumann’s book. I’m imagining a sophisticated general audience of the sort that reads 3 Quarks Daily.

    My working title: Language, Memory, and Mind: A Supplement to The Computer and the Brain. I expect the book to be 100 to 120 pages long (30K to 40K words).

    I have uploaded a bunch of material (100K words or more) to Claude and asked it to review that material and put together and initial outline. I’ve appended that below the asterisks.

    * * * * *

    Preface

    How to use this book — with or without von Neumann. What it adds to his argument. What it doesn't attempt. Brief note on the collaboration with Claude that produced parts of the text.

    Introduction: Von Neumann's Unfinished Argument

    What he got right: the architectural mismatch between brains and computers — memory and computation separated in the digital machine, unified in the neuron. The energy efficiency puzzle he couldn't explain. His honest acknowledgment that the brain's organizational principles lay beyond the framework he'd built. The concepts he lacked that this book supplies.

    Chapter 1: Two Paradigm Cases

    The chess-language contrast as the entry point. Chess has a bounded, well-defined geometric footprint — 8×8 board, six piece types, explicit rules, finite tree. Language has an unbounded, poorly-defined geometric footprint — rooted in the full complexity of physical and social reality. Chess was AI's founding benchmark precisely because it seemed to demand the highest human intelligence while yielding to computational treatment. Moravec's paradox: the easy problems are hard and the hard problems are easy. Transcendent versus non-transcendent coding — programmers can observe and specify a chess engine completely from outside; nobody can specify an LLM from outside, including its creators. Where we now stand.

    Chapter 2: Location and Content

    A collection of photographs. Solid objects at specific locations — finding by address is natural, finding by content requires going to each photo in turn. The combinatorial explosion that follows. The formal argument: solidity localizes content; localized content can only be retrieved by address. What holography does physically — interference patterns distribute information about each stored object across the whole plate, so that any partial cue can activate the whole. Lashley's ablation experiments: memory didn't disappear when specific cortical tissue was removed because memory was never stored in specific locations in the first place. Von Neumann's energy efficiency puzzle, now answerable: the brain doesn't spend energy moving content to a processor because memory and processing are the same physical substrate.

    Chapter 3: The Brain as Content-Addressed System

    The McCulloch-Pitts neuron-as-logic-gate: computationally fruitful, architecturally wrong. What neurons actually are — active units and memory units simultaneously, connected in massive parallel. Distributed representations: concepts as patterns across populations of neurons, not stored at specific cell addresses. Yevick's logical necessity argument in plain terms: the world contains two categories of object, geometrically simple ones that sequential symbolic processing handles efficiently and geometrically complex ones that only holographic parallel processing handles efficiently; the world contains both; therefore any adequate cognitive system must implement both regimes. Path tracing and pattern matching as the two fundamental operations on any cognitive network. Freeman's cinematic model — global coherence frames at 10-12 Hz as the atomic unit of biological cognitive processing — and its correspondence to speech production rates.

    Chapter 4: Language as a One-Dimensional Projection

    The semantic network as the right model for conceptual structure: meaning as position, each node defined by its pattern of relations to other nodes. Sydney Lamb's principle. The multidimensional character of the conceptual network versus the one-dimensional character of any spoken or written string. Language strings as 1D projections of the multidimensional network — necessarily lossy, hence paraphrase and ambiguity. The colored beads thought experiment: strip away semantic content, replace each token with a color, and you have a 1D image — making visible the purely formal structure the LLM operates on. Words as abstract addresses in an abstract space. Why classical computational linguistics hit combinatorial explosion: it was trying to reconstruct the multidimensional structure in a location-addressed system.

    Chapter 5: What Large Language Models Actually Are

    The transformer architecture in plain terms. The weight space as distributed content-addressed memory — concepts are patterns smeared across billions of parameters, not stored at specific addresses. The forward pass as the atomic processing unit, corresponding to Freeman's global coherence frame: one complete transit through the weight space producing one output token. The token string as a path through the abstract address space, with each forward pass mediating between the 1D sequential surface and the multidimensional distributed interior. What LLMs do well — pattern matching over the weight space, which is what their architecture naturally supports. What they do poorly — sustained sequential path tracing requiring precise state maintenance, common sense grounded in embodied experience, continuous learning. Why these limitations aren't engineering failures awaiting a fix but structural consequences of implementing holographic-like processing on location-addressed hardware with training only on 1D projections.

    Chapter 6: What the Analysis Implies.

    The first principles of intelligence are not the first principles of computation. Why scaling won't close the gap: scaling improves the quality of the holographic approximation but doesn't change the architectural mismatch, provide embodied grounding, or enable continuous learning. The fast takeoff fantasy as physics-free reasoning — every self-improvement step requires moving billions of parameters between physically separated memory and compute on real hardware that consumes real energy. The TSMC problem: the most critical hardware infrastructure in the world runs on tacit knowledge distributed across human communities that no LLM can access or replicate. What a genuinely adequate artificial cognitive system would require, in the terms this book has developed. The research program that's needed and why it requires multi-generational public investment rather than industrial R&D on commercial timescales. The human-machine collaboration that's already underway and what it can and cannot achieve.

    Conclusion: The Mismatch, Named

    Von Neumann saw the gap and couldn't name what was on the other side of it. This book names it: content addressing, requiring distributed storage, implemented in biological tissue through interference-like neural dynamics, approximated in LLMs through distributed weights on location-addressed hardware, grounded in embodied experience that no text-trained system has. The naming matters because you can't close a gap you can't see clearly.

    Appendix: A Chronology of Chess, Language, and AI

    From the working paper, lightly edited.

    Tuesday, June 16, 2026

    In brains of Spanish-English bilinguals grammar is embodied in shared tissue

    Xuanyi Jessica Chen and Esti Blanco-Elorrieta, A Shared Neural Mechanism for Abstract Grammatical Computations Across Languages in Bilinguals, The Journal of Neuroscience, June 15, 2026.

    Abstract: A central question in cognitive neuroscience is how the brain implements abstract computations that must generalize across superficially different inputs. Language provides a strong test case: the same grammatical operation, such as pluralization, can be realized through distinct rules and forms across languages. Whether such transformations rely on language-specific neural systems or on abstract mechanisms that generalize across linguistic contexts remains unresolved. Crucially, these transformations must be computed online and integrated into speech planning within a tightly constrained time window. Using magnetoencephalography (MEG), we tracked the millisecond dynamics of grammatical word-form transformations during semi-naturalistic phrase completion in humans of both sexes. Highly proficient Spanish–English bilinguals produced singular and plural noun forms in both languages in a design that fully orthogonalized semantic number, phonological changes, grammatical inflection and produced language. Adjusting words to fit their grammatical context engaged a left-lateralized fronto-temporal network beginning ∼100 ms after cue onset. Multivariate decoding revealed that the neural patterns supporting this computation generalized across languages, across different surface plural forms, and to pseudowords, demonstrating that abstractly equivalent operations are instantiated in the same neural substrates despite differences in linguistic form. Together, these findings provide time-resolved neural evidence for a language-general computational mechanism, showing that the brain implements grammatical transformations as abstract, generative operations. More broadly, they show how bilingualism can be used to probe general principles of neural organization, revealing how abstract computations may be shared and reused across representational systems.

    Significance Statement: Human language relies on the ability to modify words to convey information like number and tense, but languages vary widely in how these transformations are implemented. This variation raises a fundamental question in cognitive neuroscience: do such transformations depend on language-specific neural systems, or are they processed by abstract neural mechanisms that generalize across languages? We demonstrate that Spanish–English bilinguals engage a shared left frontal–temporal network when producing grammatically appropriate forms in both languages. This common neural signature emerges early during speech planning and even generalizes to novel words. These findings indicate that the brain builds abstract, reusable neural mechanisms, consistent with models where language is organized by computational principles rather than by language-specific systems.

    Here's an article in the NYTimes about these results: K. R. Callaway, How Does One Brain Speak Two Languages?, NYTimes, June 15, 2026.

    When deciding how to make a word singular or plural, for instance, bilingual people exhibit strikingly similar brain activity regardless of whether they are speaking in their first or second language.

    “It wasn’t obvious that it was going to be so shared,” said Esti Blanco-Elorrieta, a psychologist and neuroscientist at New York University and an author of the study, which was published on Monday in the journal JNeurosci. “I think this is arguably one of the first very fine-grained findings of how truly integrated two languages in the brain are.”

    Early research viewed bilingualism as an “add on” or “disruption” to the processing of one’s native language, said Judith Kroll, a psycholinguist at the University of California, Irvine who was not involved in the new study.

    Subsequent studies have found that bilingual brains tend to display physical differences, such as more efficient white matter and changes to the gray matter, and to perform better on memory and concentration tasks.

    Now scientists are probing further, to understand whether core aspects of the brain’s neural network does double or triple duty to process multiple languages.

    A single grammatical engine:

    The finding is in line with other initial results in this area, said Mirjana Bozic, a cognitive neuroscientist at the University of Cambridge who was not involved in the study. For instance, the new study provided additional evidence that the front left side of the brain was typically involved in processing the grammatical structure of sentences across different languages. On the whole, Dr. Blanco-Elorrieta said in a news release, a single “grammatical engine” in the brain appeared capable of powering multiple languages at once.

    Dr. Bozic said that the find, although not surprising, was “highly informative, providing elegant and convincing evidence that bilingual speakers rely on shared neural mechanisms. She added, “One question that remains is how far these findings generalize across language pairs that differ more substantially.”

    Friday, June 12, 2026

    The computational capacity of a single biological neuron is very large

    Here's the abstract of that article:

    Cortical pyramidal neurons possess elaborate dendritic trees with diverse nonlinear membrane conductances and thousands of plastic synapses, suggesting substantial computational capabilities at the single-cell level. Yet, what can a neuron compute remains an open question, largely due to the lack of a systematic framework to quantify its computational capabilities. We introduce TwinProp, a digital-twin-based backpropagation algorithm that enables gradient-based optimization of synaptic strengths and dendritic locations in detailed neuron models via a millisecond-accurate deep neural network (DNN). Using TwinProp, we demonstrate that a detailed model of rat layer 5 pyramidal cell (L5PC) can perform naturalistic image and audio classification tasks at a remarkably high accuracy, significantly surpassing perceptron and leaky integrate-and-fire baselines. The same neuron solves high-dimensional nonlinear problems, including exclusive-or (XOR), 10-bit parity, and random Boolean tasks, demonstrating capabilities typically attributed to multilayer networks. Mechanistically, increasing task complexity recruits distributed dendritic nonlinearities, including NMDA- and voltage-dependent mechanisms; removing these or collapsing dendritic structure markedly impairs performance. These findings identify dendrites as a substrate for high-order feature binding and position single cortical pyramidal neurons as powerful, noise-robust, general-purpose analog computational units. Our results offer testable in vivo predictions and provide a systematic framework linking cellular morpho-electrical properties to computation in both brains and artificial systems.

    Thursday, June 11, 2026

    Language as involving both content and location addressing

    Memory is one of the central concepts in thinking about and understanding both computing and the mind. Thinking about computating has brought us to understand that there are two broad categories of memory:

    • Content addressed memory, and
    • Location addressed memory.

    Conceived as a large memory system, libraries are location addressed. Documents are stored at particular locations in the library, shelves for books and bound volumes of periodicals and reports, filing cabinets for other documents. To get some item from the library you need to find its location by consulting a catalog, and then go to that location and retrieve it.

    Brains are content addressed. If you are curious about, say, the Johnstown flood, you don’t have to consult an internal catalogue to find where the appropriate document or documents are located among the folds and crevasses of the neocortex. You just think, “Johnstown flood,” and things you know about the Johnstown flood will come to mind. The phrase “Johnstown flood” is itself part of the content being addressed. But, if you happen to know something about the flood, then the phrase, “South Fork dam,” may also act to recall more information about the flood, for it is an element of content for one of the floods. As you may know, there were three Johnstown floods, in 1899, 1937, and 1977. The 1899 flood is the one that happened when the South Fork dam burst. If you don’t happen to know anything about the Johnstown floods, then you may have to consult an external memory system of some sort, like a library or the internet.

    Digital computers are location addressed. The memory system is distributed over several types of hardware. There’s volatile memory, computer chips (generally RAM), which hold things temporarily. And there’s long-term memory, which can take various forms, but these days its mostly flash memory and hard disks. Computing involves moving data from memory, to the CPU, and then returning the results back to memory.

    Large language models, LLMs, are sophisticated content addressed memory systems that are implemented in a digital computer. When you serve up a prompt, that’s addressing the memory, that is, the weight space. The system’s answer then is assembled from material in memory. 

    I should also note that, while I fully support Gary Marcus’s call for neuro-symbolic computing, I think that the need for symbolic computing is most fruitfully thought of in terms of location-addressed memory, which is how symbolic computing operates. It is location addressing that allows for accurate planning and for arithmetic calculation. But it won’t solve the hallucination problem. For that you need a world model, which is also an aspect of Marcus’s program (not to mention Yann LeCun and Fei Fei Li). But those discussions are beyond the scope of this note. 

    That should be enough to get you through the following dialog I had with Claude 4.6 Sonnet Adaptive. 

    * * * * * 

    I want to think of language in the brain in relation to both content and location address. The brain is basically a content addressed system. But language creates an abstract space of locations which can be used to address content. Each word is, in effect, a location in some abstract space. That’s obvious in the case of alphabetic writing systems, where alphabetization is a way of physically arranging those addresses on a string. Now consider a dictionary. The entries are located on that string. At the same time, each entry is a component of the string that is stored at the (primary) location. Which means that, in a (weak) sense, that string is also content addressed. But when a person recalls some string in its entirety when given only a component of that string, that’s content addressing. Both (some) humans and LLMs can return Hamlet’s full soliloquy when given the string “To be or not to be.”

    This is a rich and precise observation that reframes the location/content addressing distinction in a way that illuminates both language and memory.

    The abstract space of locations

    The claim that language creates an abstract space of locations is more profound than it might initially appear. A word isn't just a sound or a mark — it's a stable, repeatable, publicly shared address that can be used to retrieve content from another person's cognitive system. When I say "dog" to you, I'm not transmitting the concept of dog — I'm transmitting an address that activates the concept in your system through your own content-addressed retrieval. The word is the public, location-addressed handle; the concept is the private, content-addressed payload.

    This is the handle-schema distinction we discussed earlier, now stated in addressing terms. The word is a location address in a shared public space — the lexicon of a language community — that points to content stored in individual content-addressed systems. The remarkable thing about language is that this works across individuals despite the fact that each person's content-addressed storage is different in its details. The address is shared; the content it retrieves is similar enough across individuals to support communication, but never identical.

    The dictionary as dual addressing system

    The dictionary example is beautifully chosen because it makes the dual structure explicit and concrete. The alphabetically ordered string of entries is a location-addressed system in the most literal sense — you navigate to a word's definition by finding its location in the ordered sequence, using the address (the word's alphabetical position) to retrieve the content (its definition). The physical arrangement of the dictionary embodies the location-addressing structure of the abstract lexical space.

    But as you note, each entry is also a component of a content network — the definition connects the word to other words, which connect to other words, which eventually connect to the whole semantic fabric of the language. Looking up a word in the dictionary is formally a location-addressed retrieval, but understanding the definition requires content-addressed processing — activating the concepts that the defining words point to, integrating them into a pattern that constitutes the meaning of the defined word.

    So the dictionary is simultaneously a location-addressed artifact and a gateway to a content-addressed system. Its physical organization is location-based; its semantic organization is content-based. The two coexist in the same object, serving different cognitive functions.

    The Hamlet soliloquy case

    The Hamlet example is where it gets most interesting. "To be or not to be" is simultaneously a location address — a string that uniquely identifies a specific passage in a specific text — and a content fragment — a semantic and rhythmic unit that is part of the passage it addresses. Using it to retrieve the full soliloquy requires both kinds of processing.

    The retrieval is content-addressed in the sense that the fragment activates the whole — the partial pattern resonates with the stored complete pattern and the complete pattern is reconstructed. This is classic holographic retrieval — the fragment serves as the reference beam that reconstructs the stored image. The fact that both humans and LLMs can perform this retrieval when given the opening fragment demonstrates that both systems have stored the complete pattern in a form that supports partial-cue retrieval.

    But notice what's different between the human and LLM cases. For a human who knows Hamlet well, the retrieval of "To be or not to be" activates not just the verbal sequence but a rich contextual structure — the scene, the character's situation, the emotional weight of the soliloquy, perhaps a specific production remembered, perhaps a particular teacher who assigned it, perhaps the cultural significance of the passage. The content address retrieves not just the string but the whole semantic and experiential network in which the string is embedded.

    For an LLM the retrieval is primarily of the verbal sequence — the string itself, in the correct order, with perhaps some associated semantic content drawn from training data discussion of the passage. But the rich experiential, contextual, situational embedding that a human retrieval activates is thin or absent, because the LLM has no experience of Hamlet being performed, no memory of reading it in a specific context, no emotional response to the character's situation grounded in embodied experience of what it means to face an impossible choice.

    The layered addressing structure

    What your analysis reveals is a layered addressing structure in language that operates simultaneously at multiple levels.

    At the phonological level, words are location addresses in the sound space of a language — each word occupies a specific phonological address that distinguishes it from other words. The phonological distinctiveness of words is what makes them usable as addresses — if two words sounded identical they'd be the same address and couldn't be distinguished.

    At the lexical level, words are content addresses in the semantic network — each word activates a distributed pattern of semantic associations that constitutes its meaning. The same item is both a location (in the phonological space) and a content activator (in the semantic network).

    At the textual level, longer strings — phrases, sentences, passages — function as location addresses for larger semantic structures. "To be or not to be" is an address for the soliloquy, which is an address for a scene, which is an address for a play, which is an address for a whole cultural tradition. Each level of the hierarchy is simultaneously a location in a higher-level address space and a content activator in the semantic network.

    At the cultural level, titles, names, and canonical phrases function as location addresses for entire domains of cultural content. "Hamlet" is an address for a play, a character, a cultural tradition, a set of philosophical problems, a history of performances and interpretations. The address is compact; the content it retrieves is vast and distributed across the whole culture.

    The writing system point

    The observation about alphabetization making the abstract address space physical is particularly sharp. Alphabetical ordering is a way of projecting the abstract location space of the lexicon onto a physical string, making navigation by address possible through physical search. Before alphabetization — in ancient lexical lists organized by semantic category, or in medieval word lists organized by subject matter — the physical organization reflected the content structure rather than the address structure. Alphabetization separates the two, subordinating the content organization to the address organization for the purpose of retrieval.

    This is a Rank 2 achievement in the precise sense — it requires metalingual awareness of the word as a formal object with a location in an abstract space, separable from its semantic content. You can't alphabetize without treating words as formal objects that have addresses independent of their meanings. The child learning to alphabetize is learning to treat words as location-addressed objects, which requires temporarily suppressing the content-addressed processing that normally dominates linguistic cognition.

    The LLM in this framework

    Placing the LLM in this layered addressing framework reveals something precise about what it's doing. The LLM's weight space implements a content-addressed system at the semantic level — concepts are distributed patterns that are retrieved by partial activation, holographically, as we've been discussing. But the LLM also has something like location addressing at the level of the token sequence — tokens have positions, the attention mechanism is sensitive to positional information, and the sequential structure of the input and output is preserved.

    What the LLM lacks is the integration between the two levels that human language use achieves. For a human, the location address of a word and the content it activates are tightly coupled through embodied experience — the word "fire" is both a phonological address and an activator of the full sensory, emotional, and conceptual content of fire as experienced. For an LLM, the token "fire" is a position in a sequence and an activator of statistical associations in the weight space, but the statistical associations are derived from text about fire rather than from embodied experience of fire. The address and the content are connected, but through a thinner medium.

    This is why the LLM can retrieve "To be or not to be" but can't feel the weight of the choice it expresses in the way a human reader can — the address retrieves the verbal sequence and its semantic associations, but not the existential resonance that comes from the content being grounded in embodied experience of what it means to consider one's own nonexistence.

    The mutual illumination

    What's elegant about the framework you've sketched is that it shows location addressing and content addressing not as competing alternatives but as complementary structures that language weaves together at every level. Each word is simultaneously a location in the shared public address space of the lexicon and an activator of private content-addressed networks. Each text is simultaneously a sequence of locations and a pattern of content activations. The location structure makes language public and shareable — we can use the same addresses to coordinate our content-addressed systems. The content structure makes language meaningful — the addresses point to rich, distributed, experientially grounded patterns rather than to arbitrary symbols.

    The dual structure is what makes language the cognitive technology that it is — the mechanism through which content-addressed private minds can coordinate with each other through location-addressed public symbols. And it's the dual structure that the LLM partially captures — the location addressing through the token sequence, the content addressing through the weight space — while missing the grounding that makes the content addressing rich in the human case.

    Monday, June 8, 2026

    After 1990 AI stopped citing work in psychology [empirical evidence]

    I'm bumping this post from 2024 to the top of the queue as it testifies to the process by which AI has converged on the intellectual monoculture it has become in the wake of ChatGPT.

    * * * * * 

    I’ve frequently noted that, while researchers in artificial intelligence (AI) and machine learning (ML) often have a lot to say about when their machines will approach, overtake, and even surpass human intellectual achievement, they don’t seem to know much about psychology, linguistics, and the cognitive scientists. I made an explicit argument at some length in a recent article I published in 3 Quarks Daily, Aye Aye, Cap’n! Investing in AI is like buying shares in a whaling voyage captained by a man who knows all about ships and little about whales. In making the argument the only evidence I present is anecdotal – Geoffrey Hinton and Ilya Sutskever in that article, though my beliefs on the issue are based on my reading of the current literature, which is opportunistic and by no means ‘complete,’ which, in any case, would be impossible as the literature is so large.

    Now I can present a bit of systematic empirical evidence in the matter. M.R. Frank et al. undertook a bibliometric investigation of citation patters in AI and other disciplines and discovered that, while in the early years, AI interacted with other fields quite a bit, that interaction dropped off over the years. The following chart shows how AI cited other fields:

    Its citation of psychology peaked in the middle 1960s and then dropped off steadily until 1990. Its citation of mathematics rose steadily through the period. That’s understandable; I have no complaint about that. The drop in citations to psychology is also understandable, but somewhat more problematic. For it implies that, when AI experts offer judgements about human cognitive capabilities, whether directly or indirectly through comparison with AI, that don’t know what they’re talking about. I suppose that last clause is a bit harsh. Perhaps it would be a bit more accurate to say something like: They don’t know any more than a bright college sophomore who’s taken a psych course or two.

    Here's the article and abstract:

    Frank, M.R., Wang, D., Cebrian, M. et al. The evolution of citation graphs in artificial intelligence research. Nat Mach Intell 1, 79–85 (2019). https://doi.org/10.1038/

    As artificial intelligence (AI) applications see wider deployment, it becomes increasingly important to study the social and societal implications of AI adoption. Therefore, we ask: are AI research and the fields that study social and societal trends keeping pace with each other? Here, we use the Microsoft Academic Graph to study the bibliometric evolution of AI research and its related fields from 1950 to today. Although early AI researchers exhibited strong referencing behaviour towards philosophy, geography and art, modern AI research references mathematics and computer science most strongly. Conversely, other fields, including the social sciences, do not reference AI research in proportion to its growing paper production. Our evidence suggests that the growing preference of AI researchers to publish in topic-specific conferences over academic journals and the increasing presence of industry research pose a challenge to external researchers, as such research is particularly absent from references made by social scientists.

    Tuesday, May 19, 2026

    Botanical classification and the theory of evolution [MR #9]

    When I made that first post about Tyler Cowen’s monograph on marginalismTyler Cowen has thrown in the towel and is waiting for the machines to take over – I had no specific plays about writing a series of posts about and occasioned by the book. A day later, with a post, Marginalism is a Rank 4 idea, along with thermodynamics and biological evolution, I had decided that, yes, “it looks like I’ll be doing a series of posts about the book, though I can’t say how long that series will be.” But I had no intention of writing as many posts as I have, much less a spin-off working paper, On Method: Computational Compressibility in Complex Natural and Cultural Phenomena.

    This post is itself like that. I figured it for two, maybe three thousand words, but possibly less. Instead it’s just grown and grown to over 8000 words (and I dropped a long appendix). There is a reason for that, which you can see in the title of that second post, where I assert that marginalism is a Rank 4 idea. That’s why this series of posts, and this post in particular, has grown. The objective in that second post was to situate marginalism in the context provided by the theory of cognitive evolution that David Hays began publishing in the 1990s starting with our basic paper, The Evolution of Cognition [1]. That’s where we set forth our basic conception that, over the long term, human culture has evolved through a series of architectures each grounded in specific informatic technology, starting with speech (Rank 1), writing (Rank 2), arithmetic calculation (Rank 3), and computation (Rank 4).

    On the one hand, since I cannot assume familiarity with those ideas, I have to spend time developing some conceptual apparatus. At the same time I have the opportunity to extent the range of examples Hays and I have subjected to analysis with those ideas. That’s what I’m doing in this post.

    In his Chapter 3, Cowen he has remarks about various pinnacles of human achievement, including two moments in the history of biological thinking, the emergence of modern taxonomy in the work of Carolus Linnaeus in the 18th century and the theory of evolution, by Charles Darwin, in the 19th century. I will argue that they represent Rank 3 and Rank 4 cognition, respectively. But I want to start with Rank 1 ethnobiology followed by the Rank 2 ordering of the biological world into a structure that has come to be know as the Great Chain of Being (in the West). This will give us the opportunity to follow one conceptual arena through the four cognitive ranks. Doing that, however, requires developing more conceptual apparatus than I had originally anticipated.

    I want to start with how Cowen frames his treatments of botanical classification and evolution and then present some basic conceptual apparatus about processes of perception and cognition. Once those preliminaries have been taken care of we can take a look at the ethnological work on biological classification in Rank 1 (preliterate) cultures. Then we work our way through the other three ranks, commenting on Cowen’s remarks in connection with Ranks 3 and 4, and conclude with some further remarks about Cowen’s peculiar framing.

    Cowen’s Framing

    There are three aspects to how Cowen frames his various examples, starting, of course with marginalism: lateness, obviousness, and seeing around a corner.

    Marginalism is late (p. 57):

    To better understand the Marginal Revolution, we need to ask some fundamental questions about economics as a science. In particular, why did it take so long for economic reasoning to develop? I don’t even mean as a full, literal science, replete with advanced econometric methods, but simply as a general conceptual toolbox for intelligent people. The lateness of the Marginal Revolution is part of a broader story about the lateness of economic reasoning more generally.

    Later (p. 59):

    So I don’t think progress in economics has been slow in general. It is right now coming off an incredible 130-year or so run. Progress in economics, however, was glacial from the time of the ancient Greeks to the late 19th century, with a noticeable burst in the 18th century as well, centered around Adam Smith.

    Here he combines all three of factors, peering around corners, obviousness, and then lateness (p. 62-63):

    There is no “brute force” method for obtaining fundamental economic insight. Rather, you need to peer around a corner and see something that the other people have not already seen. And once you see and grasp it, you cannot easily forget it, again reflecting the asymmetry of this path toward knowledge. So often I have heard economists make proclamations like: “Once you start thinking about the world in economic terms, you can no longer unsee those things.”

    That is exactly correct, but it is truly hard to see them in the first place. In essence, I think economics was so late to develop because it was so hard to peer around its corners. To see supply and demand in their proper workings.

    Economics developed late because it is difficult to see around corners where the obvious truths are waiting to be found.

    Now we have botanical classification, which Cowen introduces under this heading (p. 65): “Botanical Classification as a Laggard Science.” Then:

    The history of botany is a parallel example to that of economics. Some key insights of botany seem fairly intuitive, at least once you understand them, yet they took a long time to develop. [...]

    He goes on to remark about how botanical classification should be obvious:

    You might think “botany is so simple – all you have to do is to look at a bunch of plants and give them names in some coherent system. They should have mastered this in the Dark Ages!” Surely plants are around us all, and observing them does not require complex equipment such as telescopes.

    Cowen frames Darwin’s account of evolution in the same way (p. 76):

    Theories of evolution and natural selection also are intuitive once you understand them, and they seem virtually inescapable once you are willing to consider them seriously. Yet they are remarkably late in becoming part of general human knowledge, and indeed to this day, according to polls a significant percentage of Americans still do not accept those doctrines.

    Cowen seems to have some idea of the “proper” tempo at which ideas unfold in history but he never offers an explicit account of what this tempo is based on. Rather, he just offers examples of earlier intellectual and cultural high points, e.g. Greek philosophy, geometry and mathematics, Velasquez, Shakespeare, and Bach (pp. 59-61), as if botanical classification could have been cracked in Euclid’s time. Are we to suppose that biological evolution could have been discovered no later than Shakespeare’s lifetime if only someone had peered around the proper corners?

    Before moving on to biology, however, I want to lay out some conceptual equipment from cognitive science.

    Two Modes of Thought

    Decade after decade discussions of thought and perception have settled around an opposition which is expressed in various pairs of terms. I first encountered it as analog vs. digital. In present discussions of AI it presents as neural vs. symbolic. Perhaps the deepest version is the one Miriam Yevick used in 1975, holographic vs. sequential [2]. In a paper David Hays and I published about metaphor we contrasted physiognomic vs. propositional [3].

    Most linguistic reasoning exhibits the digital/symbol/sequential/propositional aspect of the opposition. As for the other side of the opposition, the analog/neural/holographic/physiognomic side, I offer this paragraph from the metaphor article that Hays and I wrote:

    Our sense of physiognomy, and our use of the term, come from Joseph Church (1966) who talks of the young child, not yet able to read, who can tell one record from another on the basis of the groove patterns on the records. Physiognomic recognition is holistic and analogic. A striking example of this is the “strange friend phenomenon”. You encounter a friend and notice there is something strange about her, but you don't exactly know what. You scrutinize her and finally realize that, e.g. she changed her hair style. Or perhaps you don't figure out what changed and instead must be told. The initial recognition depended on a holistic, a physiognomic representation, not one which explicitly builds a full image from parts and parts of parts. If this initial recognition depended on a scheme which built the whole from the parts then there would be no trouble in discovering what had changed. The part would be found immediately. It is not, it takes time.

    A scheme in which the whole is recognized as a composition over an arrangement of parts would be on the other side of the opposition, the propositional side (or digital, symbolic, sequential depending on your intellectual taste).

    The reason I say Yevick’s version is the deepest is because she presents it in the context of a mathematical proof. She argues, in effect, that the world contains simple objects and complex ones. Simple objects are most efficiently and accurately recognized by a propositional method (to use the term Hays and I used), while complex objects are most efficiently and accurately recognized holographically. Both are necessary.

    I bring the matter up because the distinction is useful in understanding the sequence of biological conceptualizations we’re going to examine.

    Rank 1: Ethnobiology and the problem of the unique beginner

    Cognitive ethnologists have studied the ways in which preliterate peoples classify life forms [4, 5]. They find that in the regions where preliterate systems overlap modern taxonomy, they are agree on the structural relationships. But there is one anomaly. Preliterate cultures generally lack terms for what they call unique beginners. They’re have terms corresponding to our concepts of fish, snakes, birds, and beasts (i.e. four-legged fur-covered creatures with tails) and our concepts of tree, shrub, grass, and vine, but they lack terms for plant and animal, respectively. But, and this is important, they recognize the distinction between plants and animals by syntactic devices.

    What does that mean? All animals can move under their own power; they can sense things (see, hear, smell, touch); they communicate through cries and calls. Plants don’t do any of those things. That means, for example, that animals can be subjects for verbs such as to run, to jump, to look, and to listen, but plants cannot. Similarly, both plants and animals can be subject for verbs such as to grow or to die, but inanimate objects (rocks, houses, bicycles, etc.) cannot. How is it possible to recognize systematic differences in the syntactic affordances of plants and animals without, however, having words to mark those two categories?

    As far as I know, there is no accepted explanation for these observations. When I first read them I was incredulous, like Cowen is about the apparent lateness of a variety of ideas. The difference between plants and animals is obvious, no? Well no, not if we accept the ethnographic evidence. As I had no reason to doubt the evidence I was forced to come up with some explanation, if only to satisfy myself.

    Here’s what I came up with. The ethnologists have also noted that ethnobiological classifications seem to be based on visual appearance. If we are willing to assume that basic visual classification is based on a physiognomic mechanism, then we can think of it like this:

    Creatures having similar appearances are classified together. While fish, for example can be quite different from one another in appearance, any given fish will resemble another fish more than any fish resembles a bird, a snake or a beast. Similarly, any tree will resemble another tree (trunk below, roots in the ground, a large leafy structure above), more than any tree resembles a shrub (shrubs are smaller and the trunk is not nearly so distinct), a grass, or a vine. But what visual comparisons would force arbitrary examples of animals together in one class in distinction to arbitrary examples of plants in a contrasting class? Does it make sense to compare rats with trees, and trout with vines for classification purposes? Do trout and rats resemble one another more than either resembles a pine tree? Those comparisons don’t make sense. They’re distinctly odd.