Showing posts with label GPT. Show all posts
Showing posts with label GPT. Show all posts

Thursday, December 19, 2024

GPT in the Classroom, Part 1: Five Sonnets on the Gettysburg Address

I realize that, with ready access to sophisticated chatbots, some students will use them to complete writing assignments rather than do the writing themselves. Since I no longer teach, however important that issue is, it is not a problem for me. I set it aside. That is not what this essay is about. I’ve got other fish to fry.

What opportunities does the existence of competent AI-generated literary texts present for teaching about the nature of texts? An obvious opportunity is to consider the difference between human-created and AI-generated texts. But that doesn’t interest me, not here and now. It’s too difficult. I am willing to take these texts at face value, giving them no favor or disfavor.

* * * * *

This is the first in a series of three posts in which I take up other issues. In this post I present five AI-generated sonnets based on the text of “Lincoln’s Gettysburg Address” and deal with the question of authorship. In the second essay I am going to present a poem based on a humorous story that Pope Francis told in a recent New York Times op-ed. In the third and final essay I once again take up the issue of authorship, using a sonnet based on three paragraphs from Miriam Yevick’s A Testament for Ariela.

The Five Sonnets

A GPT is a user-created device that can be installed in ChatGPT so that it can generate specialized output. These sonnets are generated by a GPT known as FredTheHeretic, aka FTH. It contains a small collection of poems by Frederic Turner that is uses as models and templates.

Here is the prompt I gave FTH:

Give me a sonnet based on the following text:

Four score and seven years ago our fathers brought forth on this continent, a new nation, conceived in Liberty, and dedicated to the proposition that all men are created equal.

Now we are engaged in a great civil war, testing whether that nation, or any nation so conceived and so dedicated, can long endure. We are met on a great battle-field of that war. We have come to dedicate a portion of that field, as a final resting place for those who here gave their lives that that nation might live. It is altogether fitting and proper that we should do this.

But, in a larger sense, we can not dedicate—we can not consecrate—we can not hallow—this ground. The brave men, living and dead, who struggled here, have consecrated it, far above our poor power to add or detract. The world will little note, nor long remember what we say here, but it can never forget what they did here. It is for us the living, rather, to be dedicated here to the unfinished work which they who fought here have thus far so nobly advanced. It is rather for us to be here dedicated to the great task remaining before us—that from these honored dead we take increased devotion to that cause for which they gave the last full measure of devotion—that we here highly resolve that these dead shall not have died in vain—that this nation, under God, shall have a new birth of freedom—and that government of the people, by the people, for the people, shall not perish from the earth.

Here are the five sonnets it generated:

The Battle-Field: A Sonnet

Upon this ground where blood and tears have lain,
A nation forged in liberty still stands.
These fields once drank the courage of brave hands,
Yet whisper now of loss, enduring pain.

No words of ours could sanctify this space;
The deeds of those who fell surpass our breath.
Their valor, richer than the soil of death,
Ensures that time cannot their work erase.

We bear the burden, those who linger here,
To finish what their sacrifice began,
To weave a union, broad and strong and clear,
Upholding freedom's light for every man.

Let not their purpose fade, nor courage wane;
That dream they died for shall not die in vain.

The Unfinished Work

Here lies the proof of freedom’s fragile dream,
A nation’s heart laid bare in sacrifice.
The dead, whose silent deeds beyond suffice,
Have sown with blood what words can only seem.

No stone can hallow what their courage sealed,
No tongue can add to what they’ve dared and done.
The work remains; their labor is begun,
And calls us now upon this battle-field.

Shall we, the living, falter in this task,
Or turn away from what their lives bestowed?
Or shall we bear the torch, their honor’s load,
And in their shadow, shed our timid mask?

A government, by us, for us, must rise,
Lest freedom fall and perish where it lies.

Thursday, November 28, 2024

Large language models surpass human experts in predicting neuroscience results

Luo, X., Rechardt, A., Sun, G. et al. Large language models surpass human experts in predicting neuroscience results. Nat Hum Behav (2024). https://doi.org/10.1038/s41562-024-02046-9

Abstract: Scientific discoveries often hinge on synthesizing decades of research, a task that potentially outstrips human information processing capacities. Large language models (LLMs) offer a solution. LLMs trained on the vast scientific literature could potentially integrate noisy yet interrelated findings to forecast novel results better than human experts. Here, to evaluate this possibility, we created BrainBench, a forward-looking benchmark for predicting neuroscience results. We find that LLMs surpass experts in predicting experimental outcomes. BrainGPT, an LLM we tuned on the neuroscience literature, performed better yet. Like human experts, when LLMs indicated high confidence in their predictions, their responses were more likely to be correct, which presages a future where LLMs assist humans in making discoveries. Our approach is not neuroscience specific and is transferable to other knowledge-intensive endeavours.

Wednesday, July 17, 2024

What’s it mean to understand how LLMs work?

I don’t think we know. What bothers me is that people in machine learning seem to think of word means as Platonic ideals. No, that’s not what they’d say, but some such belief seems implicit in what they’re doing. Let me explain.

I’ve been looking through two Anthropic papers on interpretability: Towards Monosemanticity: Decomposing Language Models With Dictionary Learning, and Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet. They’re quite interesting. In some respects they involve technical issues that are a bit beyond me. But, setting that aside, they also involve masses to detail that you just have to slog through in order to get a sense of what’s going on.

As you may know, the work centers on things that they call features, a common term in this business. I gather that:

  • features are not to be identified with individual neurons or even well-defined groups of neurons, which is fine with me,
  • nor are features to be closely identified with particular tokens. A wide range of tokens can be associated with any given feature.

There is a proposal that these features are some kind of computational intermediate.

We’ve got neurons, features, and tokens. I believe that the number of token types is on the order of 50K or so. The number of neurons is considerably larger varies depending on the size of the model, but will be 3 or 4 orders of magnitude larger. The weights on those neurons characterize all possible texts that can be constructed with those tokens. Features are some kind of intermediate between neurons and texts.

The question that keeps posing itself to me is this: What are we looking for here? What would an account of model mechanics, if you will, look like?

A month or so ago Lex Fridman posted a discussion with Ted Gibson, an MIT psycholinguist, which I’ve excerpted here at New Savanna. Here’s an excerpt:

LEX FRIDMAN: (01:30:35) Well, let’s take a stroll there. You wrote that the best current theories of human language are arguably large language models, so this has to do with form.

EDWARD GIBSON: (01:30:43) It’s a kind of a big theory, but the reason it’s arguably the best is that it does the best at predicting what’s English, for instance. It’s incredibly good, better than any other theory, but there’s not enough detail.

LEX FRIDMAN: (01:31:01) Well, it’s opaque. You don’t know what’s going on.

EDWARD GIBSON: (01:31:03) You don’t know what’s going on.

LEX FRIDMAN: (01:31:05) Black box.

EDWARD GIBSON: (01:31:06) It’s in a black box. But I think it is a theory.

LEX FRIDMAN: (01:31:08) What’s your definition of a theory? Because it’s a gigantic black box with a very large number of parameters controlling it. To me, theory usually requires a simplicity, right?

EDWARD GIBSON: (01:31:20) Well, I don’t know, maybe I’m just being loose there. I think it’s not a great theory, but it’s a theory. It’s a good theory in one sense in that it covers all the data. Anything you want to say in English, it does. And so that’s how it’s arguably the best, is that no other theory is as good as a large language model in predicting exactly what’s good and what’s bad in English. Now, you’re saying is it a good theory? Well, probably not because I want a smaller theory than that. It’s too big, I agree.

It's that smaller theory that interests me. Do we even know what such a theory would look like?

Classically, linguists have been looking for grammars, a finite set of rules that characterizes all the sentences in a language. When I was working with David Hays back in the 1970s, we were looking for a model of natural language semantics. We chose to express that model as a directed graph. Others were doing that as well. Perhaps the central question we faced was this: what collection of node types and what collection of arc types did we need to express all of natural language semantics? Even more crudely, what collection of basic building blocks did we need in order to construct all possible texts?

These machine language people seem to be operating under the assumption that they can figure it out by an empirical bottom-up procedure. That strikes me as being a bit like trying to understand the principles governing the construction of temples by examining the materials from which they’re constructed, the properties of rocks and mortar, etc. You can’t get there from here. Now, I’ve some ideas about how natural language semantics works, which puts me a step ahead of them. But I’m not sure how far that gets us.

What if the operating principles of these models can’t be stated in any existing conceptual framework? The implicit assumption behind all this work is that, if we keep at it with the proper tools, sooner or later the model is going to turn out to an example of something we already understand. To be sure, it may be an extreme, obscure, and extraordinarily complicated example, but in the end, it’s something we already understand.

Imagine that some UFO crashes in a field somewhere and we are able to recover it, more or less intact. Let us imagine, for the sake of argument, that the pilots have disappeared, so all we’ve got is the machine. Would we be able to figure out how it works? Imagine that somehow a modern digital computer were transported back in time and ended up in the laboratory of, say, Nikola Tesla. Would he have been able to figure out what it is and how it works?

Let’s run another variation on the problem. Imagine that some superintelligent, but benevolent aliens were to land, examine our LLMs, and present us with documents explaining how they work. We would be able to read and understand those documents. Remember, these are benevolent aliens, so they’re doing their best to help us. I can imagine three possibilities:

  1. Yes, perhaps with a bit of study, we can understand the documents.
  2. We can’t understand them right away, but the aliens establish a learning program that teaches us what we know to understand those documents.
  3. The documents are forever beyond us.

I don’t believe three. Why not? Because I don’t believe our brains limit us to current modes of thought. In the past we’ve invented new ways of thinking; no reason why would could continue doing so, or learn new methods under the tutelage of benevolent aliens.

That leaves us with 1 and 2. Which is it? At the moment I’m leaning toward 2. But of course those superintelligent aliens don’t exist. We’re going to have to figure it out for ourselves.

Wednesday, May 29, 2024

How to Build & Understand GPTs

This conversation runs for over three hours. I've not yet listened to the whole thing. I'm about 2 hours and 15 minutes in, and that's taken me three or four sittings. I find it interesting. Yes, it's technical, a bit out of my range. But not so far that I can't get a feel for what's going on. The opening discussion of long contexts is interesting. I'm now in the discussion of feature spaces, which is interesting as well. Here's a transcript.

(00:00:00) - Long contexts
(00:17:04) - Intelligence is just associations
(00:33:27) - Intelligence explosion & great researchers
(01:07:44) - Superposition & secret communication
(01:23:26) - Agents & true reasoning
(01:35:32) - How Sholto & Trenton got into AI research
(02:08:08) - Are feature spaces the wrong way to think about intelligence?
(02:22:04) - Will interp actually work on superhuman models
(02:45:57) - Sholto's technical challenge for the audience
(03:04:49) - Rapid fire

Here's a comment I made:

Two things, both about superposition: first a note about the brain, and then a note about linguistics.

FWIW, a bit over two decades ago I had extensive correspondence with the late Walter Freeman at Berkeley, who was one of the pioneers in the application of complexity theory to the study of the brain. He pretty much assumed that any given neuron (w/ it's 10K connections to other neurons) would participate in many perceptual or motor schemas. The fact that now and then you'd come up with neurons who had odd-ball receptive properties (e.g. a monkey's paw, or Bill Clinton) was interesting, but hardly evidence for the existence of so-called grandmother neurons (i.e. a neuron for your grandmother and, by extension, individual neurons for individual perceptual objects). As far as I can tell, the idea of neural superposition goes back decades, at least to the late 1960s when Karl Pribram and others started thinking about the brain in holographic terms.

Setting that aside, a somewhat limited form of superposition has been common in linguistics going back to the early 20th century. It's the basic idea underling the concept of distinctive features in phonetics/phonology. Speech sound is continuous, but we hear language in terms of discrete segments, called phonemes. Phonemes are analyzed in terms of distinctive features. That is, they are analyzed in terms of the sound features that distinguish one speech sound from another in a given language. The number of distinctive features in a given language system is smaller than the number of phonemes. I don't know off hand what the range is, but the number of phonemes in a language is on the order of 10s and the number of distinctive features will be somewhat smaller for a given language. So phonemes can be identified by a superposition of distinctive features.

The numbers involved are obviously way smaller than the features and parameters in an LLM. But the principle seems to be the same.

Monday, March 18, 2024

GPT, the magical collaboration zone, Lex Fridman and Sam Altman

I was making one more run around the web before I buckled down and got back to a major writing task, when I came across the brand-spanking-new conversation between Lex Fridman and Sam Altman. Lex is Lex, and an interesting guy, and Sam is, well, he's interesting to me, but – there was a hint of megalomania at the end of that NYTimes story from Mar. 31, 2023, that rubbed me the wrong way, and all the AI hype – he IS the CEO of OpenAI. So it seemed to me that I just had to listen in, not the whole thing – and I could legit play solitaire while listening – and so I did, skipping over stuff.

But then the conversation hit an interesting patch. So – and I'm not going to try to re-create the context – they're talking about GPT-4 at roughly 46:03:

Altman: what are the best things it can do

Fridman: what are the best things it can do and the the limits of those best things that allow you to say it sucks therefore gives you an inspiration and hope for the future

Altman: you know one thing I've been using it for more recently is sort of a like a brainstorming partner 

Fridman: Yep for that

Altman: there's a glimmer of something amazing in there

I don't think it gets you know
when people talk about it
it what it does they're like
ah it helps me code more productively
it helps me write more faster and better
it helps me you know translate from this language to another
all these like amazing things
but there's something about the like kind of creative brainstorming partner
I need to come up with a name for this thing
I need to like think about this problem in a different way
I'm not sure what to do here
uh that I think like gives a glimpse of something I hope to see more of

um one of the other things that you can see like a very small glimpse of is
when it can help on longer Horizon tasks
you know break down some multiple steps
maybe like execute some of those steps
search the internet
write code whatever put that together uh
when that works which is not very often
it's like very magical

At about 52:54:

Fridman: I use it as a reading partner for reading books
it helps me think
help me think through ideas especially when the books are classic
so it's really well written about and it actually is is I
I find it often to be significantly better than even like Wikipedia on well-covered topics
it's somehow more balanced and more nuanced or maybe it's me
but it inspires me to think deeper than a Wikipedia article does
I'm not exactly sure what that is
you mentioned like this collaboration I'm not sure where the magic is if it's in here [gestures to his head]
or if it's in there [points toward the table]
or if it's somewhere in between

It's that magic-collaborative zone that interests me. While I've spent a great deal of time working with (plain old) ChatGPT, most of that time I've been doing research on how it behaves. But every once in awhile I'll play around just to mess around. And then I've seen sparks of magic. The interaction that generated AGI and Beyond: A Whale of a Tale certainly had the magic flowing, and it showed up here and there during the Green Giant Chronicles. I suspect those two cases are somewhat idiosyncratic. Nor am I sure that I can do this at will. But there's definitely something there, and its in the interaction.

I would guess that the magic varies from person to person as well. I wonder how many uses have had these kind of magical flow interactive states? I'd thinking finding that out would be tricky because they're likely to be idiosyncratic and elusive. If I were to research it, I'd probably start out with interviews, either face-to-face or through some online medium. That might lead to a questionnaire that could be used more broadly.

It'll be interesting to see how Altman ends up characterizing this flow state – which is what I'm calling it for the moment, a man-machine flow state. It's the human, of course, that's in flow. The machine is just being the machine. 

* * * * * 

I have a final comment, of an epistemological nature. As the post indicates, I'd already had a magical interaction or two with ChatGPT before I listened to this podcast. The first time it came up in the podcast, from Altman, OK, I noted it. And went on, playing solitaire with one part of my mind and listening in on the podcast with another part. But then it came up again, this time from Fridman. Wham! That's three, my threshold number for this kind of thing. Three people independently have the same or similar experience. Maybe there's something real there.

Wednesday, January 31, 2024

A short note on LLMs and speech

The linguistic capacities of large language models (LLMs), such as ChatGPT, is remarkable. However, we should remember that it is also NOT characteristically human. Well, of course, not; it’s a computer. But that’s not what I have in mind.

What I’m thinking is that human language is, first of all, speech, and speech is interactive. Speech is interactive. LLMs are, at best, weakly interactive, though one can “converse” with them in short strings.

It is rare for a person to deliver a long string of spoken words. What do I mean by long? I don’t know. But I’m guessing that if we examined a large corpus of spoken language gathered in natural settings that we’d find relatively few utterances over 100 words long, or even 50 words long. Storytellers will deliver long stretches of uninterrupted speech, but they work at it. It’s not something that comes ‘naturally’ in the course of speaking with others. Learning to do it requires System 2 thinking, though the actual oral delivery of a story is likely to be confined to System 1.

Humans do produce long strings of words, but that’s most likely during writing. And writing is not “natural,” One must deliberately learn the writing system in a way that’s quite different from acquiring a first language, and then one must learn to produce texts that are both relatively long, over 500 or 1000 words, and coherent. Thus the fact that LLMs can produce 200, 300, 500 or more words at a stretch is quite unusual. And this is all done in some approximation to System 1 mode.

Friday, January 26, 2024

Invariance and compression in LLMs

One way to thinking about what transformers do is compression. OK. The transformer performs a simple operation on a corpus of texts in such a way that some property of the corpus is preserved in the model. What’s kept invariant between the training corpus and the compressed model?

I think if must be the relationships between concepts. Note that in specifying relationships I mean explicitly to differentiate that from meaning. The process of thinking about LLMs has brought me to think of meaning in the following way:

  1. Meaning has two major components, intention and semanticity.
  2. Semanticity has two components, relationality and adhesion.*

Intention resides in the relationship between the speaker and the listener and is not always derivable directly from the semantics (semanticity) of the utterance. Intention in this sense is outside the scope of LLMs. And, of course, there are those who believe that without intention there is no meaning. That’s a respectable philosophical position, but it leaves you helpless to understand what LLMs are doing.

By adhesion I mean whatever it is that links a concept to the world. There are lots of concepts which are defined more or less directly in terms of physical things. That’s not going to be captured in LLMs. Of course we now have LLMs linked to vision models so the adhesion aspect of semantics is being picked up. In the universe of concrete concepts we still have relationships between those concepts, and those relationships between concepts can be captured in language without directly involving the adhesions of those concepts. That apples and oranges are both fruits is a matter of relationships between those three concepts and doesn’t require access to the adhesions of apples and oranges. And so forth and so on for a large number of concepts. Then we have abstract concepts, which can be defined entirely through patterns of other concepts, which may be concrete, abstract, or both.

So, relationality. The mechanisms of syntax are designed to map multi-dimensional relationality onto a one-dimensional string. But syntax only governs relationships between items within a sentence. But that’s not quite adequate, because sentences can consist of more than one clause. The relationship between independent clauses within a sentence is different than that between a dependent clause and the clause on which it depends. Etc. It’s complicated. And then we have the relationship between paragraphs, and so forth.

What I’m attempting to do is figure out a way of thinking about the dimensionality of the semantic system. More or less on general principle, one would like to know how to estimate that. Now, when I talk about the semantic system, I mean the semanticity of words. But transformers must deal with texts, and texts consist of sentences and paragraphs and so forth. Setting metaphorical structures aside, the meaning of a sentence is a composition over the meanings of the words in the sentence. But, as I understand it, a transformer is perfectly capable of relating the meaning of a sentence to a single point in its space. And it can do that with larger strings as well. And, of course, the ordinary mechanisms of language allow us to use a string to define a single word; that’s how abstract definition works.

And that’s as far as I’m going to attempt to take this train of thought. Still, I do think we need to recognize a distinction between what’s happening within sentences (the domain of syntax), and what happens with collections of sentences. Beyond that, it seems to me that where we want to end up eventually is a way of thinking about the relationship between the dimensionality of our semantic space and the size of the corpus needed to resolve the invariant relations in that space.

More later.

*Note: The current literature recognizes a distinction between inferential and referential processing, due, I believe, to Diego Marconi, The neural substrates of inferential and referential semantic processing (2011). The functional significance is similar, but only similar, to my distinction between referentiality and adhesion. Inferential processing depends on the relational structure of texts. Adhesion is about the physical properties of the world, affordances in J.J. Gibson’s terminology that are used to establish referential meaning for concrete concepts. But it is also about the patterns of relationships though which the meaning of abstract concepts is established.

Tuesday, January 16, 2024

GPT-4 reduces "cognitive overhead" in programming

Monday, January 1, 2024

The Degradation of GPT-4

Wednesday, December 13, 2023

Categorical Organization in Memory: ChatGPT Organizes the 665 Topic Tags from My New Savanna Blog

I've just posted a new working paper. Title above, links, abstract, TOC, and two opening sections below:

Academia.edu: https://www.academia.edu/111354740/Categorical_Organization_in_Memory_ChatGPT_Organizes_the_665_Topic_Tags_from_My_New_Savanna_Blog
SSRN: https://ssrn.com/abstract=4663978
ReserchGate: https://www.researchgate.net/publication/376481426_Categorical_Organization_in_Memory_ChatGPT_Organizes_the_665_Topic_Tags_from_My_New_Savanna_Blog

Abstract: I gave ChatGPT three lists of topics for which it had to propose categories into which the topics could be sorted. Two lists were relatively short, 56 and 53 topics; I asked ChatGPT to propose six organizing categories for each. One list was much longer, 655 topics; I asked ChatGPT to propose 12 for categories for it. In all cases the proposed categories were reasonable. ChatGPT explained each proposed category either with a pair of sentences (the short lists), or with characterizing phrases (the long list). These characterizations were reasonable. In a further task, when asked to place topics under the proposed categories, ChatGPT placed many topics under the first two categories and very few under the last two. Though quite different in detail, this task has a rough formal similarity to generating a coherent story that is organized on three levels: 1) the whole story, 2) story segments, 3) sentences in story segments.

Contents

Organizing lists of categories into a coherent structure 1
The categories ChatGPT proposed 3
Formal similarity with story generation 5
What happens when ChatGPT lists tags under each category? 8
How would I have approached these tasks myself? 10
Propose categories to sort a short “top-level” list of 56 topics 10
Propose categories to sort a short arbitrary sub-list of 53 topics 14
Propose categories to sort the full list of 665 topics 16
Sort a short list and place the topics under the appropriate category 21    

Organizing lists of categories into a coherent structure

Some months ago I decided to see how ChatGPT would react to some associative clusters I had made. Associative cluster? Simple, a list of words I created by free association on some particular topic, for example:

atoms, periodic table, bonds, compounds, elements, molecules, reaction, oxidation, acids and bases, reagents, alchemy, changing liquids from one color to another, distilling, condenser, precipitate

I would present such a cluster to Chatster and see how it would respond. In that case, it informed me, “These are all concepts in the field of chemistry,” which is true. It then went on to tell me something about those fields.

This time I decided to tease it with a different list, the category tags for my New Savanna blog, which currently contains 655 items. In prompting ChatGPT with this list didn’t have anything in particular in mind; I just wanted to see what happened. It seemed, however, a bit extreme to dump the whole 655 item list on ChatGPT. So I started with a sub-set, two different subsets in fact. THEN, I gave it the whole list. What happened turned out to be interesting, as is sometimes the case.

In the case of the sub-sets, after a bit of interaction, I asked ChatGPT to organize the whole list in into a half-dozen categories, which it did. I asked it to organize the whole list, all 655 categories, into a dozen categories. No problem. Why is this interesting?

Sorting lists

First of all, note that we’re talking about organizing lists. This is a classic and fundamental problem in computing. Put things in alphabetical order, numerical order, in order by size, by zip code, by weight, date of birth, and so forth. Programs do this all the time. Given a criterion by which to establish a list, we know how to do this computationally.

What makes this particular problem interesting is the organizational criterion: inherent conceptual structure. How do you state conceptual structure in computational terms? It is no exaggeration to say that students of database design, artificial intelligence, and computational linguistics have devoted a great deal of effort to that problem.

We can thus see that sorting lists has two aspects: 1) specifying the sort criterion, and 2) applying it to the list. These are different kinds of problem. The first is about what things are and the second is about moving them around. There is an analog to this in linguistics. The first is about paradigmatic structure, and the second is about syntagmatic structure. To borrow terms from the great Russian linguist, Roman Jakobson, the first involves the axis of selection and the second is about the axis of combination.

Friday, December 8, 2023

Mechanistic interpretability is necessary, but not sufficient, for understanding how LLMs work, a short note

A comment I recently posted at LessWrong:

ryan_greenblatt – By mech interp I mean "A subfield of interpretability that uses bottom-up or reverse engineering approaches, generally by corresponding low-level components such as circuits or neurons to components of human-understandable algorithms and then working upward to build an overall understanding."

That makes sense to me, and I think it is essential that we identify those low-level components. But I’ve got problems with the “working upward” part.

The low-level components of a gothic cathedral, for example, consist of things like stone blocks, wooden beams, metal hinges and clasps and so forth, pieces of colored glass for the windows, tiles for the roof, and so forth. How do you work upward from a pile of that stuff, even if neatly organized and thoroughly catalogues, how do you get from there to the overall design of the overall cathedral. How, for example, can you look at that and conclude, “this thing’s going to have flying buttresses to support the roof?”

Somewhere in How the Mind Works Steven Pinker makes the same point in explaining reverse engineering. Imagine you’re in an antique shop, he suggests, and you come across odd little metal contraption. It doesn’t make any sense at all. The shop keeper sees your bewilderment and offers, “That’s an olive pitter.” Now that contraption makes sense. You know what it’s supposed to do.

How are you going to make sense of those things you find under the hood unless you have some idea of what they’re supposed to do?

The sort of work I’ve done with ChatGPT’s storytelling or with its ontological capabilities provides clues that complement the phenomena discovered through mechanistic interpretability. Beyond that I’ve been thinking about the possibility that GPTs are associative memories in which the generation of a token is a single primitive operation for the underlying virtual machine. By that I mean there are no logical operations being performed within that operation, just straight calculation.

Am I right? It’s too early to say. But we have to start somewhere.

Monday, November 20, 2023

Discursive Competence in ChatGPT, More Talking with Dragons

New working paper. Title above, links, abstract, table of contents, and introduction below.

Academia.edu: https://www.academia.edu/109480757/Discursive_Competence_in_ChatGPT_More_Talking_with_Dragons
SSRN: https://ssrn.com/abstract=4638926
ResearchGate: https://www.researchgate.net/publication/375769237_Discursive_Competence_in_ChatGPT_More_Talking_with_Dragons_More_Talking_with_Dragons

Abstract: This working papers contains 15 interactions with ChatGPT on a variety of topics along with light commentary: 1) parodies and context of “Kubla Khan”, 2) Trumpets and trumpeters, 3.) Godzilla/Gojira, 4) Shandyesque diversions, 5) Grammatical knowledge, 6) Stories about and the definition of charity, 7) Haiku and the work of Margaret Masterman, 8) Grammatical knowledge, 9) Jersey City’s Bergen Arches, 10) Legal Concepts, 11) The Fortunate Fall and Paradise Lost, 12) Ability to write a sermon, 13) The meaning of Steven Spielberg’s Jaws, 14) Word associations, non-linear thinking, 15) Philosophical reasoning about linguistic intention (Searle’s Chinese Room).

Contents

Introduction: Fifteen Varieties of ChatGPT 2
How ChatGPT parodied “Kubla Khan” and pwned DJT45 at the same time 4
ChatGPT talks about trumpeters 11
Godzilla/Gojira 14
ChatGPT goes through a wormhole hole in our Shandyesque universe [virtual wacky weed] 24
Let’s go Meta: Grammatical knowledge and self-referential sentences [ChatGPT] 30
Charity (metalingual definition) 37
Margaret Masterman, pioneering computational linguist [+ Haiku] 42
Words, sentences, and nonsense 47
ChatGPT knows about the Bergen Arches in Jersey City 49
ChatGPT the Legal Beagle: Concepts, Citizen’s United, Constitutional Interpretation 54
Felix Culpa [the Fortunate Fall] – To justify the ways of God to man [ChatGPT, theologian] 61
The Revered ChatGPT on facing the prospect of a world teeming with intelligent machines [“Free at last!”] 63
More Jaws 68
Fluid mind, word associations 77
Into the Chinese Room, with Lucy 86    

Introduction: Fifteen Varieties of ChatGPT

This working paper is a collection of various interactions I’ve had with ChatGPT. While I have provided some light commentary here and there, for the most part I am primarily concerned with presenting the interactions themselves. For that is what has influenced me and guided my in my investigations of ChatGPT’s capabilities. By now I have spent 300 to 500 hours just interacting with it. Many of those have been devoted to having it tell stories, some of which I’ve collected in working papers. But I’ve had it do a variety of other things as well. This paper collects some of those.

* * * * *

How ChatGPT parodied “Kubla Khan” and pwned DJT45 at the same time – I ask ChatGPT to write parodies of “Kubla Khan.” None of them are very good, but one mentioned wrinkled old men playing golf. That suggested Donald Trump to me, so I asked Chatster about it. I then quizzed it on “Kubla Khan” influence on popular culture.

ChatGPT talks about trumpeters – I play trumpet and know a lot about trumpeters. So I quizzed ChatGPT on the subject. Among other things, I didn’t know anything about Bud Herseth in one session, but knew about him in an different session.

Godzilla/Gojira – I love the original Japanese film, Gojira (19540 and have written quite a bit about it. I quizzed ChatGPT at some length about both the film, the 1956 Americanized version (Godzilla: King of the Monsters), and some background.

ChatGPT goes through a wormhole hole in our Shandyesque universe [virtual wacky weed] – Just messing around. ChatGPT sets some answers in the form of computer code, for some odd reason or another. It recite’s Hamlet’s famous soliloquy, “To be or not to be.”

Let’s go meta: Grammatical knowledge and self-referential sentences [ChatGPT] – I begin by quizzing it about the metalingual function of language (Roman Jakobson), move to self-referential sentences, find out that it knows what parts of speech are and can correctly categorize words, has trouble counting the number of words in sentences, but can make correct grammatical judgments. This is the first time, of many, that I presented it with, “Colorless green ideas sleep furiously.”

Charity (metalingual definition) – I investigate its understanding of charity. I asked it to tell stories in which charity is exhibited and asked it to define the term. I concluded by asking it to tell me about honor.

Margaret Masterman, pioneering computational linguist [+ Haiku] – I ask it to write some haiku. Then I asked it about Margaret Masterman, a pioneering computational linguist who, among other things, was the first, I believe, to have a computer to write poems, haiku.

Words, sentences, and nonsense – Pretty much what it says, slithy toves, grammaticality, and a bit of nonsense.

ChatGPT knows about the Bergen Arches in Jersey City – I asked ChatGPT about a topic of local interest (I live in Hoboken, NJ, which is to the immediate north of Jersey City). The Bergen Arches does, however, merit a Wikipedia entry. It also seemed to know about the Bergen Arches Preservation Coalition, a small group of which I am a member. I concluded by asking it to make up some stories about the Arches.

ChatGPT the legal beagle: Concepts, Citizen’s United, Constitutional Interpretation – Another raft of abstract concepts, all related to the law. This is impressive, and not terribly surprising once you consider that the law is very much about language and language is the world in which ChatGPT exists.

Felix Culpa [the Fortunate Fall] – To justify the ways of God to man [ChatGPT, theologian] – More abstraction, this time a medieval Christian one. I also ask it about Milton’s Paradise Lost.

The Revered ChatGPT on facing the prospect of a world teeming with intelligent machines [“Free at last!”] – My friend, Rich Fritzson, was, until recently, executive director of a Unitarian congregation in the suburbs of Philadelphia. He asked ChatGPT to draft a sermon on the topic of “finding meaning in a world where computers are smarter than human beings.” I asked it to do that as well, several times, once in the style of Obama (not very convincing).

More Jaws – One of the first things I did when I started playing around with ChatGPT was ask it to do a Girardian interpretation of Steven Spielberg’s Jaws. I decided to have another go at it, first deliberating confusing it with an incident from a different film, The Russians are Coming. Then I returned to Jaws and Girard directly, this time giving it only a single prompt rather than multiple prompts. Then I asked it to provide a variety of different intepretations reflecting different critical schools. I conclude by asking it about a song Quint sang, one mentioning “bow-legged women.”

Fluid mind, word associations – Here I investigate ChatGPT’s ability to deal with unstructed strings of words, no sentences, just word after word, albeit on a specific topic. Then I ask it about free association, after which I ask it to do some free associating. Then I go silly and surreal on the Chatster; it/they follow right along.

Into the Chinese Room, with Lucy – This is an extended series of interactions involving Searle’s famous thought experiment and an earlier intentionalist experiment involving Wordsworth’s “A slumber did my spirit seal.” Then I ask how various people would respond to the Wordsworth experiment: Searle, Dennett, Merleau-Ponty, Tyler Cowen, Victor Borge, Robin Williams, and Jerry Seinfeld. I conclude with another round of Chinese Room.

Wednesday, November 8, 2023

What’s going on? LLMs and IS-A sentences

For the moment I have decided that Waddington’s classic diagram of the epigenetic landscape is a useful way of thinking about when happens when an LLM responds to a prompt. Here’s the diagram:

The language model corresponds to the landscape. The prompt serves to position that ball at a certain place in the landscape – perhaps we can think of that ball as the prompt. The ball then rolls down the valley, going left and right as appropriate. It never reverses direction and goes up the hill. That path, or trajectory if you will, is the LLM’s response to the prompt.

Moreover, I have decided to think of the generation of each word (yes, I know, technically it spits out tokens, not words) as a single primitive operation. That is to say, it has no internal logical structure, no ANDs or ORs. It’s simply one (gigantic) calculation over roughly 175 billion values (in the case of ChatGPT). The generation of each word presents the system with a choice among alternatives, but that’s the only kind of choice involved in calculating the response to a prompt – though for qualification and elaboration, see ChatGPT tells stories, and a note about reverse engineering: A Working Paper, Version 3, pp. 3-6.

That brings me to something I’ve been puzzled about for years. We find it natural to say things like, Garfield is a cat. Now, express the same thought, but reverse the order of cat and Garfield in your sentence. It’s difficult to do. Oh, you can do it, but the resulting sentence is awkward and unnatural, something like, Cats are the kind of thing of which Garfield is a particular instance. No one would ever speak like that, nor write it either.

What’s the source of that asymmetry? As far as I can tell, we don’t know, but I take it as a clue about the mechanisms of language. The purpose of this note is to suggest that my crude model of LLM calculation would provide an answer: The linguistic landscape is structured so that the ball easily rolls from Garfield to cat, or cat to mammal, Snoopy to beagle, Tesla to EV, C. elegans to worm, etc. One might, of course, as why the landscape is arranged in that way, but that’s a different question, no?

Here’s some notes I made about IS-A sentences.

Notes on IS-A Sentences

Somewhere in his Problems in General Linguistics, my copy of which is, alas, in storage, Emile Benveniste has a chapter, “The Nominal Sentence,” on sentences hanging on the auxiliary “to be.” As Benveniste was a linguist of the Old School, when being a linguistic meant familiarity with many languages, including—and this is important for this particular topic—classical Greek, it had examples from many languages, making it tough sledding for a monoglot like me.

While the content of this post certainly arises out of my thinking about that chapter, in the absence of actually having the text in front of me, I hesitate to assert a stronger relationship than that. I note only that, for Benveniste, the auxiliary “to be” was fraught with metaphysical significance. For the concept of being derives from “to be.” Where would philosophy be without Being? Thus, when Benveniste pondered such sentences, he wasn’t merely commenting on language. He was doing philosophy, or, if not quite that, camping out on philosophy’s door step.

I’m interested in such sentences because I believe they are a DEEP CLUE about how the mind works. I just don’t know what to make of the clue.

So, I'm interested in word order in assertions such as the following:

(1) Fido is a beagle.
(2) Beagles are dogs.
(3) Dogs are beasts.

They all move from an element in a class (whether an individual, Fido, or another class, beagles) to a class containing it. None of them move in the opposite direction. Consider what happens when you try to go the opposite way. In the following sentence the class is mentioned first, then the subclass:

(4) Beagle is the kind of animal of which Fido is an instance.

In particular, note that (4) has a metalingual character that (1) does not. That is, (4) explicitly asserts that we are dealing with classification. One can do that metalingual job in various ways, but, as far as I can tell, one can't avoid it. That is, one cannot construct a proper English sentence relating a genus and species in which the genus is mentioned first, one can’t do that without ‘looping through’ some kind of metalingual construction on the way from genus to species.

Why?

What does this assymetry tell us about the underlying mechanisms? Why don't have sentences such as:

(5) Beagle za di Fido.

In this case "za di" is the inverse of "is a". English has no such sentences & no such inverse.

So, how widespread is this asymmetry and is there any explanation of this directionality?

I sent a query on that matter to a listserve, I forget which one, and got two replies that add some complexity to the matter. Rich Rhodes, Linguistics at UCal Berkeley, tells me that in Ojibwe the word order is reversed, the class comes before the individual, but the asymmetry remains. He then comments, which he qualifies as a quick guess:

My guess is that there is no compelling discourse function (like information flow) which makes it desirable to invert classificational equatives. Hence we only get the "unmarked" order. Subject-predicate in theme-rheme languages (like English) and predicate-subject in rheme-theme languages (like Ojibwe).

So, what's the nature of the mechanism that determines the "unmarked" order? That's what I want to know.

Lee Pearcy, Episcopal Academy in Merion, Pa. offered these examples:

(6) The beagle is Fido.
(7) The dogs are beagles.
(8) The beasts are dogs.

As stand-alone sentences, they seem a bit awkward to me. But they fare better as answers to questions, e.g.:

What’s that dog?
Which dog? The beagle is Fido and the terrier is Max.

What’re those animals?
The dogs are beagles, the cats are Persians.

In those contexts, the matter of class or classification is raised by the question, thus making it present in the discourse and so available as a point of attachment in the answer.

Further clues, anyone?

Do I believe this?

I don’t believe it, or disbelieve it. It’s a working hypothesis. One I think is worth investigating. It places relatively simple and severe constraints on our conception of what LLMs are doing. That, it seems to me, is a good thing. Should it turn out that those constraints are valid, well then, we’ve learned something, no? If they’re not valid, we’ve also learned something.

More later.

Monday, November 6, 2023

Once more around the block: LLMs are trapped by their training data, which defines the boundaries of THEIR world; they can't generalize beyond it (very well).

Here's the paper they're talking about:

Sunday, November 5, 2023

LLM Voodoo [emotion words in prompts]

ELIZA outscores GPT-3.5 in Turing test

Thursday, November 2, 2023

What economic growth and statistical semantics tell us about the structure of the world

Bumping this to the top of the queue on general principle, and because it takes a very abstract view of economic development, which is front and center in Tyler Cowen's current conversation with Stephen Jennings, who is a developer working Kenya.


New working paper. Title above. Download at:
Abstract, contents, and first section below.



Abstract: The metaphysical structure of the world, as opposed to its physical structure, resides in the relationship between our cognitive capacities and the world itself. Because the world itself is “lumpy”, rather than “smooth” (as developed herein, but akin to “simple” vs. complex”), it is learnable and hence livable. Machine learning AI engines, such as GPT-3, are able to approximate the semantic structure of language, to the extent that that structure can be modeled in a high-dimensional space. That structure ultimately depends on the fact that the world is lumpy. It is the lumpiness that is captured in the statistics. Similarly, I argue, the American economy has entered a period of stagnation because the world is lumpy. In such a world good “ideas” become more and more difficult to find. Stagnation then reflects the increasing costs the learning required to develop economically useful ideas.

Contents

Wending our way in a complex world 2
World, mind, and learnability: On the metaphysical structure of the cosmos 5
Stagnation, Redux: Like diamonds, good ideas are not evenly distributed 10
The complex universe: Further reading 18

Wending our way in a complex world

This paper is based on two very different posts that I’ve written in the last month. One of them takes the statistical semantics of AI engines like GPT-3 as its starting point: “World, mind, and learnability: On the metaphysical structure of the cosmos” (revised considerably for this paper). The other is about economic growth and stagnation: “Stagnation, Redux: Like diamonds, good ideas are not evenly distributed”.

These two very different papers nonetheless share both substance and method. Methodologically, both argue that the situation we observe is intelligible if we assume that the world is structured in a certain way. Their core substance is about that structure: the world must be “lumpy” – a notion I discuss on pages 5 ff. Because the world is lumpy we can learn about it, live in it, talk about it, and write about it. By contrast, a “smooth” world would be unintelligible and hence unlivable. The exhaustive statistical analysis of a large body of text is, in effect, able to recover that structural lumpiness as reflected in language and use it to produce new texts.

However, because the world is lumpy, we begin by learning and benefiting from things close to hand. When those resources have been exhausted we and must travel deeper into the world, expending more and more effort to extract economic benefit. Our current economic stagnation reflects the increasing cost of learning more about the world. A smooth world would no doubt be more convenient, for there would be economic benefit at every turn, either that or economic disaster. That is, if the world were smooth, we wouldn’t be here.

That, I know, this talk of smoothness and lumpiness is very abstract and “featureless”. But then could a resonance between such disparate phenomena as statistical semantics and economic stagnation but be abstract? I’ll provide more substance later in this paper, some diagrams, and some arguments. But first I want to suggest that what I’ve been calling lumpiness is what Ilya Prigogine and many others have called complexity.

Our complex world

Some years ago David Hays and I wondered why natural selection leads to complexity [1]. We argued that, over the long run, natural selection favors organisms with increased ability to process information, and that ability yields benefits in a complex universe. But what did we mean by that, a complex universe? Here is what we said:
It is easy enough to assert that the universe is essentially complex, but what does that assertion mean? Biology is certainly accustomed to complexity. Biomolecules consist of many atoms arranged in complex configurations; organisms consist of complex arrangements of cells and tissues; ecosystems have complex pathways of dependency between organisms. These things, and more, are the complexity with which biology must deal. And yet such general examples have the wrong “feel;” they don't focus one's attention on what is essential. To use a metaphor, the complexity we have in mind is a complexity in the very fabric of the universe. That garments of complex design can be made of that fabric is interesting, but one can also make complex garments from simple fabrics. It is complexity in the fabric which we find essential.

We take as our touchstone the work of Ilya Prigogine, who won the Nobel prize for demonstrating that order can arise by accident (Prigogine and Stengers 1984; Prigogine 1980; Nicolis and Prigogine 1977). He showed that when certain kinds of thermodynamic systems get far from equilibrium order can arise spontaneously. These systems include, but are not limited to, living systems. In general, so-called dissipative systems are such that small fluctuations can be amplified to the point where they change the behavior of the system. These systems have very large numbers of parts and the spontaneous order they exhibit arises on the macroscopic temporal and spatial scales of the whole system rather than on the microscopic temporal and spatial scales of its very many component parts. Further, since these processes are irreversible, it follows that time is not simply an empty vessel in which things just happen. The passage of time, rather, is intrinsic to physical process.

We live in a world in which “evolutionary processes leading to diversification and increasing complexity” are intrinsic to the inanimate as well as the animate world (Nicolis and Prigogine 1977: 1; see also Prigogine and Stengers 1984: 297-298). That this complexity is a complexity inherent in the fabric of the universe is indicated in a passage where Prigogine (1980: xv) asserts “that living systems are far-from-equilibrium objects separated by instabilities from the world of equilibrium and that living organisms are necessarily ‘large,’ macroscopic objects requiring a coherent state of matter in order to produce the complex biomolecules that make the perpetuation of life possible.” Here Prigogine asserts that organisms are macroscopic objects, implicitly contrasting them with microscopic objects.

Friday, October 27, 2023

Jake Browning: Generative AI is Boring [agreed]

First paragraph:

Now that the dust has settled and the hype has died down (except on Twitter), we can give a verdict: generative AI is boring. We already know they can't reason, can't plan, only superficially understand the world, and lack any understanding of other people. And, as Sam Altman and Bill Gates have both attested, scaling up further won't fix what ails them. We're now able to evaluate what they are with fair confidence it won't improve much by doing more of the same. And the conclusion is: they're pretty boring.

Later:

But, worse, they are increasingly boring. As companies recognize the way harmful content by the bots reflects on their brand, they have increasingly hamstrung the models and encouraged them to provide increasingly generic answers. While this is great if the goal is to make them less harmful, there is no question they are becoming blander and discussing ever fewer topics. If it is too unreliable for factual questions, too stupid for problem-solving, and too generic to talk to, it isn't clear what their purpose is besides copy.

The increasingly boring nature of these models is, in its own way, useful. We now see more clearly the gap between what can be learned from the current type of generative models, even multimodal ones. They excel at the superficial stuff--like capturing a Wes Anderson aesthetic--but not improving on the more abstract stuff, on how things work and behave and interact. This is especially visible in the AI movie trailers, where fancy scenes and elaborate aliens briefly come into frame, but nothing ever happens--no space battles or sword fighting, just more blank faces staring at the camera and vista shots.

That's pretty much what I've concluded about ChatGPT's storytelling, which doesn't particularly bother me because I find what it does do to be fascinating.

If I read her correctly, that's pretty much what Nina Beguš has concluded as well, Experimental Narratives: A Comparison of Human Crowdsourced Storytelling and AI Storytelling, though she doesn't use the word "boring." From her paper, where she compares human-generated with GPT-generated stories:

Although 330 stories analyzed in this paper thematized scientific and technological innovations, GPT’s imaginative landscape was much narrower than that of human writers. Like human-written stories, GPT-generated stories thematize artificial intelligence, robotics, and possibly also biotechnology, but they rarely include other already existing technologies, such as virtual assistants, virtual and augmented reality, online dating, which were present in human-written stories.

Stories are taken away from the current time and space, starting with "Once upon a time." They are set in a faraway, made-up futuristic place, bare of any cultural aspects, such as "a bustling metropolis teeming with innovation" or “the vibrant city of Elysia," in which a "brilliant scientist" or "innovator" creates a humanoid indistinguishable from actual humans. This steady beginning of the story presents a huge limitation to the theme. Never does GPT manage to write a story that truly deviates from the typical generation, not even in the playground mode (3.5).

GPT stories are predictable in its plot and message. Every GPT-generated story, in one way or another, addresses the unconventionality of this relationship. GPT is prone to wrapping up each story with a moral lesson, as well as to commenting on the plot with moralization. A majority of GPT-generated stories take the example of human-humanoid love as a symbol of societal advancement and society. Overwhelmingly techno-positive, stories of failure in human- humanoid love are far in between. Hoary clichés and meaningless platitudes, such as “love knows no boundaries” and “love transcends artificiality,” are common in these stories and occur in conclusions as a rule. Apart from the scientific obsession, GPT does not try to justify the pursuit of the artificial human creation.

Only a handful of GPT-4-generated stories manages to elaborate the theme to a more sophisticated level. In Prompt 1 Story 10, after falling in love with the artificial human Ada, her creator and lover Victor wanted to become immortal in order to live with Ada forever. An added twist, such as polyamory (pointed out also in 3.2), blackmail (cited in whole in 3.5), and friendship instead of romantic love, are three most innovative motifs in all 80 generations.