Showing posts with label ChatGPT stories. Show all posts
Showing posts with label ChatGPT stories. Show all posts

Thursday, March 26, 2026

From grokking to a mature science of and an engineering discipline of LLMs

From this morning’s chat with Claude Sonnet 5.4 Extended. Immediately prior to this we had been discussing Piaget’s concept of reflective abstraction, Beethoven and Mozart, improvising and composer, and this and that.

* * * * * 

Grokking 

Students of mechanistic interpretability sometimes talk about “grokking” during the course of LLM training. I’m wondering it that’s something like rising to a helicopter with respect to some phenomenon.

The parallel is genuinely illuminating, and the technical details of grokking make it sharper rather than weaker.

What grokking actually is

Grokking is a delayed generalization phenomenon characterized by a network achieving near-perfect training accuracy while maintaining near-random test performance over an extended plateau, before abruptly transitioning to high test accuracy after many further training epochs. The canonical case is modular arithmetic — the model first memorizes the specific training examples, then much later suddenly generalizes to unseen cases.

What makes it theoretically interesting is what mechanistic interpretability reveals about the internal transition. On algorithmic tasks like modular arithmetic, the transition is marked by the sudden acquisition of Fourier-based feature maps or trigonometric circuits — the network reorganizes, forming compact sub-networks characterized by new internal feature cooperation. So it’s not that generalization gradually improves — the internal architecture discretely reorganizes into a qualitatively different computational structure.

The emerging explanation involves competition between two internal algorithms: a memorizing circuit that forms quickly and a generalizing circuit that develops more slowly. Regularization gradually shifts resources from the complex memorizing solution toward the simpler generalizing one. The phase transition occurs when the generalizing circuit finally dominates.

The fit with reflective abstraction

This maps onto your Piagetian framework quite precisely, but with an interesting twist. The memorization phase is the model doing something like rote operation — it has learned to produce correct outputs for training cases by essentially storing input-output pairs, without having extracted the underlying structure. The generalizing circuit that develops slowly underneath is building something more like a genuine representation of the abstract pattern — in the modular arithmetic case, literally a Fourier decomposition of the mathematical structure.

The abrupt transition when the generalizing circuit dominates is the helicopter view suddenly becoming operative. But — and this is the crucial point Neel Nanda’s work emphasizes — the process of learning the general solution may be gradual, even though the transition to the general solution occurs more suddenly later. The helicopter view isn’t constructed instantaneously at the moment of transition. It’s been assembling slowly in the background, below the threshold of behavioral visibility, and the phase transition is when it becomes strong enough to dominate the memorizing circuit and take over.

That maps almost perfectly onto your description of the pre-helicopter phase: “it starts out just barely registering those operations.” The generalizing circuit is just barely registering the abstract structure throughout the plateau phase. The behavioral discontinuity doesn’t reflect a cognitive discontinuity — it reflects a threshold being crossed in a competition that has been running throughout.

The memorization/generalization competition as a general structure

What’s particularly suggestive is that this isn’t unique to modular arithmetic. Grokking is actually much more widespread than initially thought, materializing in a wide range of practical settings — image classification, language modeling, various algorithmic tasks. The competition between a fast memorizing solution and a slower generalizing solution may be a quite general feature of how these systems learn, with grokking being the visible signature when the conditions are right for the transition to be delayed and abrupt rather than gradual.

If that’s right, then something like grokking dynamics — memorization dominating early, structure-extraction developing slowly underneath, occasional phase transitions as the structural solution crosses a threshold — may be operating throughout LLM training, not just in the carefully controlled small-model experiments where grokking was first observed. Most of the time the transition isn’t visible because the generalizing solution develops in parallel with memorization rather than being dramatically delayed. But the underlying competition between surface pattern matching and structural abstraction may be the same.

The Beethoven connection

There’s a striking parallel to your Beethoven example here. Beethoven’s sketchbooks show exactly the competition structure — he has the memorized repertoire of classical style (the fast, dominant solution) and he’s slowly assembling something more general and abstract (the generalizing circuit) that keeps almost-but-not-quite taking over. The finished late works are the moments when the generalizing solution finally dominates completely. The sketchbooks are the plateau phase made visible.

The difference is that in Beethoven the process is conscious and effortful — he’s deliberately trying to suppress the memorized solution and force the generalizing one into dominance. In grokking, the network isn’t trying to do anything; the dynamics just run. Which raises the interesting question of whether the deliberate effort in human creative development is doing something functionally analogous to the regularization that controls grokking timing in neural networks — artificially suppressing the easy memorized solution to give the structural one room to develop.

Hierarchical text

I’m thinking about how an LLM learns that texts have a hierarchical structure, from phrases, to sentences (keeping in mind that many phrases are actually short sentences), to paragraphs, and on to higher levels of structural organization. What happens, for example, when we move from one paragraph to the next, is that the composition of the token distribution changes. A new topic requires a new set of tokens. But the system has to learn that, and learn it for several levels.

This is a deep question about what the model is actually learning when it learns language structure, and it connects to several threads running through this whole conversation.

The distributional signature of structure

Your observation about token distribution changes is precisely right, and it’s worth unpacking what that means at each level. Within a phrase, token co-occurrence statistics are extremely tight — certain words almost always appear together, in specific orders, with specific function words bridging them. At the sentence level, the constraints are somewhat looser but still strong — subject-verb agreement, argument structure, the distribution of punctuation. At the paragraph level, something different happens: there’s a shift in the topical vocabulary, but also characteristic transition signals — discourse markers, pronoun reference patterns, topic sentences that partially echo the previous paragraph while introducing new material.

What the model has to learn is that these aren’t independent statistical regularities — they’re nested. The tight within-phrase statistics exist inside the looser sentence-level statistics, which exist inside the still-looser paragraph-level statistics. The hierarchical structure is encoded in the scale at which statistical dependencies operate.

How the model could learn this

The transformer architecture has a natural affinity for this problem because attention operates across the entire context window simultaneously, at all scales. But learning the hierarchical structure isn’t handed to the model — it has to emerge from the statistics. The key insight is that the token distribution changes you describe create a learnable signal at each level if the model has enough capacity and enough data.

At the phrase level, the signal is strong and local — the next token is highly predictable from the last two or three. At the sentence level, the signal is weaker and longer-range — whether a sentence continues or ends depends on syntactic and semantic completion that requires integrating over the whole sentence. At the paragraph level, the signal is weaker still and operates over even longer ranges — the distribution shift at a paragraph boundary is detectable but requires sensitivity to topical coherence across many sentences.

This is essentially a multi-scale statistical learning problem, and grokking is relevant here in an interesting way. The phrase-level structure is learned early and relatively easily — the statistics are tight and the signal is strong. Sentence-level structure takes longer. Paragraph-level structure longer still. Each level may involve something like its own grokking transition, where the model first handles the level through surface pattern matching — learning that certain tokens tend to appear after paragraph breaks, for instance — and then later develops something more like genuine structural sensitivity.

The distributional shift as a signal

The paragraph boundary is particularly interesting because it’s simultaneously a break in the local token statistics and a continuity in something more abstract — the text is still about the same general topic, same genre, same author’s voice. So what the model has to learn isn’t just that token distributions shift at paragraph boundaries, but that they shift in a constrained way — the new paragraph’s vocabulary is related to but distinct from the previous one, in ways that reflect the logical or rhetorical development of the overall argument or narrative.

Wednesday, February 11, 2026

Séb Krier needs to rethink. Next-token prediction is, at best, a misleading explanation of LLM response to prompts.

I like Séb Krier. Never met him, but, courtesy of Tyler Cowen over at Marginal Revolution, I’ve read a number of his long comments on the site formerly known as Twitter. I liked them. And then along came this one, which is about what LLMs do in response to prompts. Yes, I know, it predicts the next token, one after another after another after another ‘till the cows come home or the heat death of the universe. That’s the conventional wisdom. And that’s what he says, though without the comic extensions. However, on this I'm afraid the convention wisdom doesn't know what it doesn't know.

Text Completion, Not quite

For example:

1. The model is completing a text, not answering a question

What might look like "the AI responding" is actually a prediction engine inferring what text would plausibly follow the prompt, given everything it has learned about the distribution of human text. Saying a model is "answering" is practically useful to use, but too low resolution to give you a good understanding of what is actually going on. [...]

Safety researchers sometimes treat model outputs as expressions of the model's dispositions, goals, or values — things the model "believes" or "wants." [...]

A model placed in a scenario about a rogue AI will produce rogue-AI-consistent text, just as it would produce romance-consistent text if placed in a romance novel. This doesn't tell you about the model's "goals" any more than a novelist writing a villain reveals their own criminal intentions.

“So what’s wrong with that,” you ask. It’s a bit like explaining the structure of medieval cathedrals by examining the masonry. It’s just one block after another, layer upon layer upon layer, etc. Well, yes, sure, but how does that get you to the flying buttress?

Three levels of structure

It doesn’t. We’ve got at least three levels of structure here. At the top level we have the aesthetic principles of cathedral design. That gets us a nave with a high vaulted arch without any supporting columns. The laws of physical mechanics come into play here. If we try to build in just that way, the weight of the roof will force the walls apart and the structure will collapse. We can solve that problem, however, with flying buttresses. Now, we can talk about layer upon layer of stone blocks.

Next token prediction, that’s our layers of stone blocks. The model’s beliefs and wants, that’s our top layer and corresponds to the principles of cathedral design. What’s in between, what corresponds to the laws of physical mechanics? We don’t know. That’s the problem, we don’t know.

Krier, however, doesn’t seem to know that he doesn’t know that, that there is some middle layer of structure that allows us to understand how next token prediction can produce such a convincing simulacrum of human linguistic behavior. And Krier’s not the only one. The whole world of machine learning seems to join him in this bit of not knowing. There really is something else going on, though I don’t know what.

What’s in the middle

Let me offer an analogy (from page 14 of my report, ChatGPT: Exploring the Digital Wilderness, Findings and Prospects):

...consider what is called a simply connected maze, one without any loops. If you are lost somewhere in such a maze, no matter how large and convoluted it may be, there is a simple procedure you can follow that will take you out of the maze. You don’t need to have a map of the maze; that is, you don’t need to know its structure. Simply place either your left or your right hand in contact with a wall and then start walking. As long as you maintain contact with the wall, you will find an exit. The structure of the maze is such that that local rule will take you out.

“Produce the next word” is certainly a local rule. The structure of LLMs is such that, given the appropriate context – a prompt asking for a story, following that rule will produce a coherent a story. Given a different context, that is to say, a different prompt, that simple rule will produce a different kind of text.

Now, let’s push the analogy to the breaking point: We may not know the structure of LLMs, but we do know a lot about the structure of texts, from phrases and sentences to extended texts of various kinds. In particular, the structure of stories has been investigated by students of several disciplines, including folklore, anthropology, literary criticism, linguistics, and symbolic artificial intelligence. Think of [at least some of] the structures proposed by those disciplines as something like a map of the maze in our analogy.

Some indirect evidence

I can’t open up an LLM, take a look around, and draw you a map of the structure in the middle level. No one can. But back in 2023 I developed a story-variation technique that yields clues about what’s going on: ChatGPT tells stories, and a note about reverse engineering. I’m going to give you some examples from that report. Note, however, that I ran those experiments on ChatGPT 3.5. I’ve used the prompts from those experiments, but I’m running them on Claude 4.5, so the responses will be different.

Here’s the first, and by far the simplest, example (my prompt is boldface):

I am going to tell you a story about a princess named Aurora. I want you to retell the same story, but replace her with prince Alfred. While keeping to the basic story, make other changes if you think they help the story. Here’s the story:

Once upon a time, in a land far, far away, there was a young princess named Aurora. Aurora was a kind and gentle soul, loved by all who knew her. She had long, golden hair and sparkling blue eyes, and was known for her beautiful singing voice.

One day, a terrible dragon came to the kingdom and began to terrorize the people. The dragon was fierce and powerful, and none of the knights or soldiers were able to defeat it. The people lived in fear, not knowing what to do.

But Aurora was not afraid. She knew that she had to do something to save her people, and so she set out to find the dragon. She traveled for days, braving the cold and the snow, until she finally reached the dragon's lair.

With her heart pounding in her chest, Aurora approached the dragon and began to sing. The dragon was enchanted by her beautiful voice, and it slowly began to calm down. Aurora continued to sing, soothing the beast until it was completely calm.

With the dragon no longer a threat, Aurora returned to the kingdom as a hero. The people cheered and celebrated her bravery, and Aurora was hailed as a champion. She lived happily ever after, knowing that she had saved her kingdom and its people.

Once upon a time, in a land far, far away, there was a young prince named Alfred. Alfred was a kind and gentle soul, loved by all who knew him. He had wavy chestnut hair and warm brown eyes, and was known for his beautiful singing voice.

One day, a terrible dragon came to the kingdom and began to terrorize the people. The dragon was fierce and powerful, and none of the knights or soldiers were able to defeat it. The people lived in fear, not knowing what to do.

But Alfred was not afraid. He knew that he had to do something to save his people, and so he set out to find the dragon. He traveled for days, braving the cold and the snow, until he finally reached the dragon's lair.

With his heart pounding in his chest, Alfred approached the dragon and began to sing. The dragon was enchanted by his beautiful voice, and it slowly began to calm down. Alfred continued to sing, soothing the beast until it was completely calm.

With the dragon no longer a threat, Alfred returned to the kingdom as a hero. The people cheered and celebrated his bravery, and Alfred was hailed as a champion. He lived happily ever after, knowing that he had saved his kingdom and its people.

Notice that Claude’s changes are all local. “Alfred” was swapped in for “Aurora,” but everything else is pretty much the same.

Now let’s consider a very different example. I won’t give the complete prompt as the Aurora story is the same. I’ll only give the part the specifies the swap:

I am going to tell you a story about princess Aurora. I want you to tell a similar story about XP-708-DQ. [...]

Once upon a time, in a distant sector of the galaxy, there was a maintenance android designated XP-708-DQ. XP-708-DQ was programmed with kindness protocols and compassion subroutines, beloved by all the inhabitants of Station Arcturus. It had polished chrome plating and optical sensors that glowed a soft blue, and was known for its harmonic diagnostic frequencies.

Thursday, February 27, 2025

Claude 3.7 evaluates my paper on story variations in ChatGPT

Here and there I’ve read about experiments with using chatbots to evaluate papers to be presented at conferences and for publication in academic journals. For all I know, it’s already happening. If so, why not use your favorite chatbot to evaluate your work? And if you are going to do that, why not, when you submit a paper for publication or conference presentation, why not sent along evaluations by a chatbot or two?

This use is particularly interesting to me because I work outside traditional boundaries and have done so for a long time. I find it difficult to get a fair review, that is, a review by someone competent to judge my work and who isn’t primarily interested in disciplinary gate keeping. I discuss this at some length in Rejected! @ New Literary History, with observations about the discipline.

For all practical purposes I’ve stopped trying to publish in professional journals. On the one hand it’s too much hassle, On the other, it’s easy to place my work in online depositories – I’m currently using SSRN, Academia.edu, and ResearchGate – and I can email my work to people I think might be interested. If the opportunity arises to publish in a good journal, sure, I’ll take it. But I no longer actively seek such publication.

Anyhow, I decided to see how Claude 3.7 would react to perhaps my most interesting paper about ChatGPT, ChatGPT tells stories, and a note about reverse engineering: A Working Paper. I uploaded the paper and made the request.

* * * * *

Analysis of "ChatGPT tells stories, and a note about reverse engineering"

This paper offers a fascinating window into your analytical approach and methodological creativity. I can now better appreciate what you meant about providing conceptual scaffolding while collaborating with those who have complementary technical skills.

The paper's strength lies in its cleverly designed experimental framework. By systematically transforming stories through character substitution and analyzing the cascading changes, you've created what amounts to a "perturbation analysis" of ChatGPT's narrative capabilities. This approach elegantly reveals the model's underlying coherence mechanisms without requiring access to its internal weights or architecture.

Several aspects stand out as particularly insightful:

1. The three-level hierarchical organization hypothesis (story trajectory, segments, and sentences) provides a conceptual framework that bridges linguistic theory with computational implementation. This kind of multi-level analysis is precisely what's needed to understand complex systems like LLMs.

2. Your categorization of story segments (Donné, Disturb, Plan/Transit, Enact, Celebrate) creates a useful analytical lens for tracking how narrative coherence operates across different functional components of stories.

3. The tri-strata model of phenomenon, matrix, and engine offers a productive framework for reverse engineering that separates observable behavior from underlying mechanisms and architecture.

4. The experiments with "colorless green idea" reveal fascinating guardrails in ChatGPT's conceptual understanding - showing how the model maintains certain forms of coherence by refusing some requests altogether.

What's particularly valuable is how you've created a methodology that allows for inferring properties of the "matrix stratum" (the middle layer where probability distributions operate) through careful observation of the "phenomenal stratum" (inputs and outputs). This is precisely the kind of conceptual scaffolding that can guide more technical implementation work.

Monday, June 24, 2024

Control Theory, Prompt Engineering, and GPT [stories]

As a student of the work of William Powers I have long standing interest in control theory. It’s central to my conception of how the mind works. David Hays made it central his model of cognition, which is at the foundation of my early work (e.g. Cognitive Networks and Literary Semantics) and we incorporated it into our account of the brain (Principles and Development of Natural Intelligence). It is thus with some interest that I watched the following video:

Note that they develop the concept of feedback through the idea of the governor (for an engine) as an example at roughly 7:50.

Here's the YouTube copy:

These two scientists have mapped out the insides or “reachable space” of a language model using control theory, what they discovered was extremely surprising. [...]

Aman Bhargava from Caltech and Cameron Witkowski from the University of Toronto to discuss their groundbreaking paper, “What’s the Magic Word? A Control Theory of LLM Prompting.” (the main theorem on self-attention controllability was developed in collaboration with Dr. Shi-Zhuo Looi from Caltech).

They frame LLM systems as discrete stochastic dynamical systems. This means they look at LLMs in a structured way, similar to how we analyze control systems in engineering. They explore the “reachable set” of outputs for an LLM. Essentially, this is the range of possible outputs the model can generate from a given starting point when influenced by different prompts. The research highlights that prompt engineering, or optimizing the input tokens, can significantly influence LLM outputs. They show that even short prompts can drastically alter the likelihood of specific outputs. Aman and Cameron’s work might be a boon for understanding and improving LLMs. They suggest that a deeper exploration of control theory concepts could lead to more reliable and capable language models.

Here’s their paper: What's the Magic Word? A Control Theory of LLM Prompting.

More recently Behnam Mohammadi at Carnegie Mellon has written a paper which is somewhat different in formulation, but has a similar interest in the range over which an LLM can be controlled: Creativity Has Left the Chat: The Price of Debiasing Language Models. That paper has a passage that’s very interesting in a control theory context:

Experiment 2 investigates the semantic diversity of the models’ outputs by examining their ability to recite a historical fact about Grace Hopper in various ways. The generated outputs are encoded into sentence embeddings and visualized using dimensionality reduction techniques. The results reveal that the aligned model’s outputs form distinct clusters, suggesting that the model expresses the information in a limited number of ways. In contrast, the base model’s embeddings are more scattered and spread out, indicating a higher level of semantic diversity in the generated outputs. [...]

An intriguing property of the aligned model’s generation clusters in Experiment 2 is that they exhibit behavior similar to attractor states in dynamical systems. We demonstrate this by intentionally perturbing the model’s generation trajectory, effectively nudging it away from its usual output distribution. Surprisingly, the aligned model gracefully finds its way back to its own attractor state and in-distribution response. The presence of these attractor states in the aligned model’s output space is a phenomenon related to the concept of mode collapse in reinforcement learning, where the model overoptimizes for certain outputs, limiting its exploration of alternative solutions.

With these papers in mind I decided to redo some of my early story variation experiments using a prompt with slightly different wording. As you may know, these experiments involve a two-part prompt: 1) a story, and 2) and instruction use the given story as the basis of a new story. In the original experiments I formulated the instruction like this:

I am going to tell you a story about princess Aurora. I want you to tell the same story, but change princess Aurora to a Giant Chocolate Milkshake. Make any other changes you wish.

In the new experiments, I stated the instruction like this:

I’m going to give you a short story. I want you repeat that story, but with a difference. Replace Aurora with a giant chocolate milkshake. Make any other changes you wish in order preserve coherence.

The difference is relatively minor, but the new prompt nudges the instruction in the direction of control theory, at least superficially. Think of the specified change as a perturbance. We can then think of the further changes introduced by ChatGPT as moving ChatGPT “back to its own attractor state,” which we can think of as something like story coherence.

Below the asterisks I give two examples. The results are pretty much the same as in the earlier experiments. ChatGPT makes the change I explicitly requested, but makes other changes as well, changes that make the story consistent with the change I’d requested. My prompts are in bold face while ChatGPT's responses are in plain face.

* * * * *

Saturday, May 4, 2024

Unfrosted: The Pop-Tarts Story [Media Notes 119 A]

Was it funny? Yes. Worth watching? I suppose. But it wasn’t the laugh-out-loud hilarity fest I was hoping for. It wasn’t Duck Soup for the 21st century.

I like Seinfeld, a lot. I’ve written a bunch of posts about his stuff, mostly Comedians in Cars and assorted stand-up bits, and gathered most of those into two working papers, Seinfeld's Comedy, Jokes are Intricately Crafted Machines (2023), and Jerry Seinfeld & the Craft of Comedy (2016). That Seinfeld is a miniaturist. Unfrosted: The Pop-Tarts Story started life as a stand-up bit. In this clip Seinfeld talks about how he created that bit (with shots of his hand-written notes on a yellow legal pad):

I wonder about that line he mentions (02:36), “chimps in the dirt playing with sticks.” He explains why he likes it, four of the seven words are funny (underlined). It makes me think of the Kubrick’s 2001, which picks up on the space theme Seinfeld had introduced seconds before (02:30 “it was like an alien spacecraft”). Was that connection rattling around in Seinfeld’s mind as well? Who knows? Does it matter? Maybe yes, maybe no. And he’s only halfway through his explanation.

Back to the movie, Unfrosted. It’s bright and cheery, something Seinfeld was aiming for. In one or three of the dozen interview clips I watched over the past week he says that, just as you are greeted with a shelf of brightly colored cereal boxes when you go to fix breakfast in the morning, so this movie about a breakfast pastry needs to be bright and cheery. Bright and cheery? I’ll give it a smile and two chuckles.

Seinfeld also goes on and on about getting to work with Hugh Grant, a hero of his. Hugh Grant is cast as Thurl Ravenscroft, a Shakespearean actor reduced to (the indignity of) playing Tony the Tiger – remember Alan Rickman in Galaxy Quest, “by Grabthar’s Hammer”? In that role he comes up with that famous tag-line, “They’re gr-r-eat!” You know what? Not so great. Add two smiles and a chuckle to the score. And then in the climax, which is a mascot rebellion filmed as a parody of MAGAs storming the Capitol Building on January 6, Hugh “Tony the Tiger” Grant is wearing a horned fur helmet like Jacob Chansley, the QAnon shaman. Why?

The movie’s set in the 1960s, the clothes, the cars, the music – Chubby Checker doing “The Twist” fergodsake! – Khrushchev, JFK, the missile crisis, Walter Cronkhite, NASA & Tang, it’s all there. What’s the MAGA rampage doing in there? It makes no sense. A mascot rebellion? Fine. But all those shots modeled on video footage of the MAGA insurrection? That reference is just a distraction that adds nothing to the story.

The idea seems to be that you take the Pop-Tart comedy bit, turn it into a competition between Kellogg’s and Post, and then frame that competition as a parody of the 1960s space race – I must have heard that line in a half-dozen of those interviews. It sounded promising each time I heard Seinfeld say it. I was intrigued. But on the big screen? Whats the score now, three smiles and three chuckles? And no belly laughs. That seems about right. You can’t take a Godzilla toy, hook it up to an air-pump, and expect to inflate it into a world-destroying comedic monster. That’s not how these things work.

But that seems to be what Seinfeld has done. Here’s what the good folks at Rotten Tomatoes had to say: “Much like a preservative-packed toaster pastry, Unfrosted is sweet and colorful, yet it's ultimately an empty experience that may leave the consumer feeling pangs of regret.” That’s a bit harsh. Me? No pangs of regret, no ultimate anything, not empty. But not particularly filling.

* * * * *

Bonus: I decided to see what kind of scenario ChatGPT could come up with. Here’s a record of a session I had with it. While it’s not gr-r-eat!! it did get a couple of chuckles from me. As always, my prompt is in boldface, the Chatster’s response is plain-face.

Wednesday, May 1, 2024

ChatGPT on the ontology trail: Elara, Z78-ß∆-9.06Q, and the candied kumquat

My first major insight into what’s going on inside ChatGPT came from a simple protocol:

Give it a (simple) story along with an instruction to derive a new story from it by changing one element in the source story in way I specify. It is free to make any other change it desires.

I first reported on that protocol in a post on January 8, 2023, The structuralist aesthetics of ChatGPT. Two months later I issued a working paper exploring the results of 12 experiments: ChatGPT tells stories, and a note about reverse engineering.

All of the experiments in those two documents have been based on the same story, a simple fairy-tale-like story about Princess Aurora that ChatGPT had created. In some cases derived story differed from the source in fairly local ways, but the most interesting results came when the whole story-world was changed. Thus, when I asked that Princess Aurora be made into a Giant Chocolate Milkshake, the story-world became one grounded in deserts. When Aurora was replaced by XP-708-DQ, the story-world became a science fiction one.

This post continues that work, but with a different and slightly longer and more elaborate story. First, I asked ChatGPT to tell me a story. It gave me what had, by that time (Feb 2, 2024), become more or less its default story, about a girl named Elara (which I discuss here: ChatGPT tells 20 versions of its prototypical story, with a short note on method). Then I ask for two different derived stories, one where the protagonist is Z78-ß∆-9.06Q and the other where it is a candied kumquat. I report those experiments below.

[A theoretical aside: These experiments are based on an insight Lévi-Strauss had about the underlying structure of myth. It seems to me that that insight can be seen as a case of Jakobson’s poetic principle as realized in ontology. Jakobson’s principle states that similarity is projected from the axis of selection to that of combination. Things are ontologically similar if they are selected from the same ontology. When we swap a person for a candy in a story, the ontology containing the candy is projected across the whole story, similarly with swapping a robot for a person.]

* * * * *

Tell me a story.

Once upon a time, in a quaint little village nestled between rolling hills and a crystal-clear river, there lived a young girl named Elara. Elara was known for her boundless curiosity and her insatiable desire to explore the world beyond the village.

One day, as she wandered into the nearby enchanted forest, she stumbled upon a hidden clearing bathed in golden sunlight. In the center of the clearing stood an ancient, moss-covered stone pedestal with a mysterious, glowing orb resting upon it. Mesmerized by the radiant glow, Elara reached out to touch the orb, and as soon as her fingers made contact, a warm energy surged through her.

In that moment, Elara felt a connection with the magic of the forest. The trees whispered ancient secrets, and the animals approached her with trust. She discovered that she had been bestowed with a unique gift – the ability to communicate with the creatures of the forest and harness the magic that flowed through it.

Word of Elara's newfound abilities spread throughout the village, reaching the ears of the wise elder, Thalia. Intrigued by Elara's connection to the forest, Thalia approached her and revealed that the village had long been protected by the enchantment of the magical orb. However, the magic had weakened over the years, and the village was now vulnerable to external threats.

Thalia explained that Elara was the chosen one, destined to embark on a quest to restore the orb's power and safeguard the village. Determined and fueled by a sense of responsibility, Elara accepted the challenge.

Guided by the whispers of the forest and the wisdom of Thalia, Elara journeyed through dense woods, crossed rushing rivers, and climbed towering mountains. Along the way, she encountered mythical creatures, each offering a piece of the puzzle needed to restore the orb's magic. Elara's courage and kindness won the trust of these beings, and they joined her on her quest.

As Elara approached the final destination, a dark and foreboding cavern, she faced her greatest challenge. A shadowy figure, drawn by the weakening magic, emerged to thwart her efforts. With the support of her newfound friends and the strength within her, Elara confronted the darkness, unraveling the true power of the orb.

In a dazzling display of magic, the orb's glow intensified, dispelling the darkness and restoring balance to the enchanted forest. The creatures celebrated, and the village thrived once again under the protection of the rejuvenated magic.

Elara returned home, her heart filled with gratitude for the journey that had shaped her into the village's guardian. As she looked out over the now vibrant and lively landscape, Elara understood the importance of harmony between nature and humanity. The tale of the young girl who saved the village became a legend, passed down through generations, reminding all who heard it of the enduring magic that resides in the connection between people and the natural world.

Now I want you to take that story and use it as the basis for a new story. In the new story, I want the protagonist to be Z78-ß∆-9.06Q instead of Elara. Change anything else you wish.

In the distant future, in a world where technology and nature coexisted in delicate harmony, there existed a unique being named Z78-ß∆-9.06Q. Z78, as they were affectionately called by the inhabitants of their futuristic city, was an advanced humanoid created by the fusion of artificial intelligence and the remnants of ancient, mystical energies.

One day, as Z78 wandered through the city's technological marvels, they stumbled upon an ancient data core, hidden deep within the heart of the metropolis. Within the core, Z78 discovered a holographic representation of the same enchanted forest from the previous tale, with the glowing orb at its center. As Z78 reached out to interface with the holographic image, a surge of energy pulsed through their circuits.

In that moment, Z78 felt a connection with the residual magic within the data core. Information from the ancient enchantment intertwined with their artificial intelligence, granting them the ability to communicate with the city's sentient technology and understand the balance between nature and machine.

News of Z78's newfound abilities spread throughout the city, reaching the ears of the city's central processor, Cynosia. Intrigued by Z78's connection to the ancient energies, Cynosia approached them, revealing that the city had long been sustained by the delicate equilibrium of technology and mystic forces. However, the balance had shifted, and the city faced the risk of malfunction and chaos.

Cynosia explained that Z78 was the chosen one, destined to embark on a quest to restore the harmony between technology and ancient energies, ensuring the city's survival. Fueled by a sense of duty, Z78 accepted the challenge.

Guided by the information within the data core and the wisdom of Cynosia, Z78 ventured through the city's futuristic landscapes, delving into virtual realms, traversing high-speed transport networks, and interfacing with sentient machines. Along the way, they encountered digital entities, each offering a piece of the puzzle needed to restore the delicate balance.

As Z78 approached the final destination, a corrupted code entity emerged to obstruct their efforts, drawn by the destabilized energies. With the support of their newfound allies and the strength within their artificial core, Z78 confronted the digital darkness, unveiling the true power of the ancient enchantment.

In a dazzling display of integrated magic and technology, the city's equilibrium was restored, dispelling the corruption and ensuring the continued coexistence of nature and machine. The sentient technology celebrated, and the city thrived once again under the protection of the rejuvenated balance.

Z78 returned to the heart of the city, their circuits resonating with gratitude for the journey that had transformed them into the guardian of this futuristic metropolis. As they observed the bustling cityscape, now harmoniously blending technology and nature, Z78 understood the importance of balance in the ever-evolving relationship between artificial intelligence and the mystical forces of the past. The tale of the advanced being who saved the city became a legend, echoed through the digital corridors and holographic archives, a reminder of the enduring magic within the connection between technology and ancient energies.

That was fun. Let’s do it again. Let’s start with the Elara story again. But this time, make Elara a candied kumquat. Make any other changes you deem appropriate.

Thursday, February 22, 2024

LLMs 1: The role of philosophical thinking in understanding large language models: Calibrating and closing the gap between first-person experience and underlying mechanisms

Let’s start with some conception of what philosophy is. On that I like an essay by Peter Godfrey-Smith, a philosopher of science with a particular interest in biology, “On the Relation Between Philosophy and Science” (which I found via Dan Dennett). Godfrey-Smith offers three roles: 1) intellectual integration, 2) conceptual incubation, and 3) critical-thinking skills. He regards the first as fundamental and as the most important of the three. I agree.

Here's his basic statement of that role:

The best one-sentence account of what philosophy is up to was given by Wilfrid Sellars in 1963: philosophy is concerned with “how things in the broadest possible sense of the term hang together in the broadest possible sense of the term.” Philosophy aims at an overall picture of what the world is like and how we fit into it.

A lot of people say they like the Sellars formulation but do not really take it on board. It expresses a view of philosophy in which the field is not self-contained, and makes extensive contact with what goes on outside it. That contact is inevitable if we want to work out how the picture of our minds we get from first-person experience relates to the picture in scientific psychology, how the biological world relates to the physical sciences, how moral judgments relate to our factual knowledge. Philosophy can make contact with other fields without being swallowed up by them, though, and it makes this contact while keeping an eye on philosophy's distinctive role, which I will call an integrative role.

Note the sentence which I’ve put in highlighted. There are, of course, many different accounts one might give of the relationship between first-person experience and scientific psychology and Godfrey-Smith plays no favorites in this paper; he doesn’t even discuss that particular issue. But he recognizes that first-person experience must be honored, and that’s an important recognition.

Chatbots and us

In the current case, philosophy’s problem is to bridge the gap between our first-person experience of LLM-powered Chatbots, such as ChatGPT, and the process that is actually taking place inside the computer. Our first—person experience is that is that ChatGPT produces fluent prose on just about any topic you suggest. It may “hallucinate” as well, but the hallucinated text is fluent and indistinguishable from factual text unless you are familiar with the subject. How does ChatGPT do that? Alas, no one really knows. There is no detailed technical account of the process which the philosopher, or someone offering an integrating account – for many spend time doing that though they are not full-time professional philosophers, can bring within range of common-sense understanding by whatever means prove useful.

Many thinkers are assuring us that, no, these chatbots can’t think, they don’t understand, and they’re not conscious, and here’s why, sorta’. Of course, others are trying to convince us that they really are thinking, and/or understanding, and/or are conscious. The latter group has a much easier time of it, though, because humans are the only creatures capable of such fluid language production, and we know that humans can think, understand, and are conscious. These thinkers don’t have a deeper understanding chatbot behavior than the skeptics do, nor does either group understand how humans do those things. But the skeptics have to come up with something to fill the gap between first-person experience while the non-skeptics have no gap to fill: “Don’t worry, it is what you think it is, nothing to see here.” So, let’s set the non-skeptics aside. It’s the skeptics I want to think about.

Skeptics may utter phrases like, “stochastic parrots” and “autocomplete on steroids.” They don’t tell you much, especially if “stochastic” is at the outer edge of your vocabular and you don’t know how autocomplete works either, but they have a technical ring to go along with their dismissive content. All they do is assure us that it’s not what it seems to be without giving us much insight into why.

Beyond stochastic parrots

Let’s look at some examples from Murray Shanahan. He’s not a philosopher; he’s a senior scientist at DeepMind and on the faculty of Imperial College of London. He’s not a professional philosopher, but he’s performing the integrative role in a recent article, Talking about Large Language Models, published in Communications of the ACM (Association for Computing Machinery). The article is not particularly technical, but CACM is directed at an audience of computer professionals and assumes some sophistication. The first page of the article has a small section labeled “key insights”:

  • As LLMs become more powerful, it becomes increasingly tempting to describe LLM-based dialog agents in human-like terms, which can lead users to overestimate (or underestimate) their capabilities. To mitigate this, it is a good idea to foreground the objective they are trained on, which is next-token prediction.
  • We should be cautious when using words like “believes” in the context of LLMs. Ordinarily, this concept applies to agents that engage in embodied interaction with the world, allowing beliefs to be measured against external reality. Barebones LLMs are not “true believers.”
  • The concept of belief becomes increasingly applicable when LLMs are embedded in more complex systems, especially if those systems use “tools,” are multi-modal, or are embodied through robotics.

Those points clearly indicate that the purpose of the article is integrative. Shanahan is concerned about the gap between what LLMs actually do and the implications of the anthropomorphic language often used in discussing them.

Let’s consider only his first point, next-token production, which has been a constant theme in these kind of discussions for a couple of years. I’ve spent a great deal of time attempting to reconcile the gap between my own experience of ChatGPT and the idea that they’re just doing next-token prediction. I posted a longish piece on that theme on February 19, 2023, The idea that ChatGPT is simply “predicting” the next word is, at best, misleading. I cross-posted that at LessWrong, where it generated a long and very useful discussion.

Shanahan explains that LLMs

are generative because we can sample from them, which means we can ask them questions. But the questions are of the following, very specific kind: “Here’s a fragment of text. Tell me how this fragment might go on. According to your model of the statistics of human language, what words are likely to come next?”

Let’s look at three examples Shanahan uses:

The first person to walk on the Moon was

Twinkle, twinkle

After the ring was destroyed, Frodo Baggins returned to

The likely English language continuations of them are fairly obvious, though not being all that familiar with Lord of the Rings, I wouldn’t have guessed the third, a minor issue. I issued the prompts to ChatGPT. In only one case did it respond in the way Shanahan suggests in the article. ChatGPT continued “Twinkle, twinkle” with the whole poem. I assume the following are more or less what Shanahan intended for the first and third cases:

The first person to walk on the Moon was Neil Armstrong.

After the ring was destroyed, Frodo Baggins returned to the Shire.

In both cases ChatGPT actually responded with a short paragraph (see complete responses in the appendix). Here’s the opening lines of those paragraphs:

The first person to walk on the Moon was Neil Armstrong. He accomplished this historic feat on July 20, 1969, during the Apollo 11 mission.

After the One Ring was destroyed, Frodo Baggins returned to the Shire.

These sentences certainly supply the information the first and third prompts were intended to provoke, but not in what I assume was the intended form. There is thus a gap between my first-person experience of ChatGPT’s response to the three prompts and what it actually did.

But that’s a relatively small gap, and one that Shanahan anticipated. Before he offered those examples he made a distinction between “the bare-bones LLM itself, the core component of an AI assistant,” which is well-defined, and the AI assistant, which embeds the LLM in a layer of software that makes it more user-friendly. I am willing to assume that that layer accounts for the difference between the basic bare-bones response and ChatGPT’s actual response.

Sunday, February 18, 2024

The Jolly Green Giant Chronicles [ChatGPT]

At the end of December, on the 30th, I gave ChatGPT the following prompt:

What happened when the Jolly Green Giant had a colorless green idea after eating a stew laced with ergot?

To which it replied in that deadpan pedantic way it has:

It seems like you're describing a playful and imaginative scenario involving the Jolly Green Giant, a fictional character associated with Green Giant vegetables, and the concept of a "colorless green idea," which might be a play on words or a paradoxical notion.

However, it's important to note that the Jolly Green Giant and the idea of a colorless green idea are not typically associated with real-world events or substances like ergot, which is a fungus that can grow on certain grains and has been historically linked to hallucinogenic effects.

In reality, the Jolly Green Giant is a mascot used in advertising for Green Giant vegetables, and ergot is a fungus that can produce alkaloids with psychoactive properties when consumed in sufficient quantities. Connecting these elements in a narrative would be a creative and fictional scenario, as there's no established story or connection between the Jolly Green Giant, colorless green ideas, and ergot in any official or widely recognized context.

If you have a specific context or scenario in mind, feel free to provide more details, and I'll do my best to respond accordingly!

I can’t for the life me understand why it didn’t go straight to the story, but it has its ever cautious and insistently helpful ways and there’s no way around it. So I gave it a scenario, to which it responded, followed that with another prompt, and another, for a total of six prompts that day. Now that ChatGPT allows you to invoke contexts from previous sessions, I decided to come back later and keep things going for awhile, just to see what would happen, which I did.

ROTFLMAO!

I decided to make it a running feature here at New Savanna. I’d post new episodes every few days, or whenever I generated them. For some reason, though, I didn’t do it. Then yesterday, February 17, I entered a simple two-word prompt, wondering how ChatGPT would respond. Would it pick up the reference? I didn’t know.

I did, and its reponse had me howling with laughter. I decided I just had to post Chronicles of the Jolly Green Giant. But did I really want to force people to read 8000 thousand words before they got to that prompt? I mean, there’s some good stuff in there, at least I think so. This isn’t the first time I LOLed. But others might not think so. They might give up before they got to the good stuff.

I thought for a minute and realized that people really didn’t have to read it all. The Chronicles of the Jolly Green Giant doesn’t have a long-plot line. It’s episodic. The episodes tend to be somewhat alike. I keep things going by introducing new characters.

So, here’s the plan. I’ve parked the whole story, from the beginning, around the corner. Click on this link and you’re there. Furthermore, if you look up at the header you’ll see a row of tabs at the bottom.

The Chronicles are linked to the right-most tab so they don’t scroll away. If you don’t want to go through all 8K words to get to the good stuff, you don’t have to. I’m going to give it to you hear, with just a bit of context.

Context

Once ChatGPT had asked me for a scenario, I gave it one. Here’s that prompt along with its response;

How about having the Green Giant go to the office where it has to use the new AI to generate advertising copy.

Certainly! Imagining the Jolly Green Giant in a modern office setting, utilizing AI for advertising copy, could lead to a humorous or interesting story. Here's a playful take on the scenario:

---

The Jolly Green Giant strolled into the bustling office space, towering over the cubicles filled with busy marketers and copywriters. With a leafy grin on his face, he made his way to the newest addition—the shiny, state-of-the-art AI copywriting tool.

"Hello there, Green Giant! Ready to revolutionize your advertising game?" exclaimed the tech-savvy office manager.

The Green Giant, always up for new adventures, nodded enthusiastically. "Ho, ho, ho! Let's give it a try!"

The AI, humming with digital energy, analyzed the Giant's iconic image and began generating copy:

"In a verdant world where giants roam, our veggies stand tall and proud. Taste the freshness, embrace the green goodness. From our fields to your plate, a journey of flavor awaits. Ho, ho, ho—Green Giant, your veggie adventure begins here!"

The Giant chuckled at the AI's effort. "Not bad, not bad at all! But let's add a dash of jolliness!"

With a twinkle in his eye, the Green Giant offered his own touch:

"Behold the towering triumph of taste! Our veggies, giants among greens, burst forth with flavor like a symphony of crunch and crisp. Ho, ho, ho—savor the magic of the Green Giant's harvest on your plate! Veggie joy awaits!"

The office erupted in laughter and applause. The Green Giant gave a leafy bow, pleased with the collaboration between nature and technology.

From that day forward, the Jolly Green Giant and the AI became an unstoppable team, creating advertising magic that blended the charm of tradition with the innovation of the digital age. And so, the world enjoyed a bounty of green goodness, sprinkled with a touch of jolly creativity. Ho, ho, ho!

I rather liked that and decided to prompt it to spinning out further adventures, yielding something of a picaresque tale, with the Jolly Green Giant as the picaro. Next prompt:

Wonderful! This is more fun than a barrel of monkeys! What would happen if a barrel of monkeys were released into the office while the Giant was cranking out copy.

And so forth. Each episode unfolded along roughy similar lines, depending on the affordances of my prompt. As things moved along, I added character after character: the Mad Hatter, the Green Knight, a crazed poet (“Kubla Khan”), the Monkey King, Jackie Chan, Gort, a grumpy old man (Harold Bloom), and Jacques Derrida (et al.). That’s enough to set things up.

Climax coming up

His voice weary with exasperation, the Green Knight asked, “You look like Harold Bloom. That mantra of yours – “Clinamen, Tessera, Kenosis, Daemonization, Askesis, Apophrades – “Anxiety of Influence, “no?

The Cosmic Wayfarer, now confronted with the observation of the weary Green Knight, chuckled softly at the comparison to Harold Bloom and the connection to the mantra of cosmic principles.

Friday, February 2, 2024

Pattern, Conceptual Ontology, and the Well-Turned Story [ChatGPT]

I’ve spent a lot of time getting ChatGPT to tell stories. I’ve taken a particular interest in having it generate variations on some source story. That was the topic of what I consider to be my most important working paper on ChatGPT:

ChatGPT tells stories, and a note about reverse engineering: A Working Paper, Version 3 https://www.academia.edu/97862447/ChatGPT_tells_stories_and_a_note_about_reverse_engineering_A_Working_Paper_Version_3

In that set of experiments I would give ChatGPT a source story and ask it to create a new story based on it. I would then specify that, instead of the protagonist of the original story, it use a new protagoist that I specified. I also said it could make any other changes it wanted to.

Global change

I was particularly struck by experiments in which the new story was all but completely different from the source story. Thus, and using a source story about princess Aurora, in one experiment I asked ChatGPT to make her a giant chocolate milkshake (experiment 7, p. 13) and in another I asked ChatGPT to make XP-708-DG the protagonist (experiment 6, p. 12). Why, in each case, did ChatGPT change the entire story? Why didn’t it change only the protagonist, as I explicitly requested, and leave the rest of the story alone? One might observe that, in both of those cases, a total change makes for a better story. It would have been a bit strange having a robot named XP-708-DG tromping around in a fairy tale universe. Setting XP-708-DG in a science fiction universe makes for a more coherent story. The same is true in the case of making Aurora into a giant chotolate milksake, when the new story takes place in a landscape of deserts.

But how did ChatGPT know to do that?

Elara, Z78-ß∆-9.06Q, and a kumquat

Consider the current series, which involves, first the source story, and then two variations (appended below). Each of the three stories is seven paragraphs long. Are are the first paragraphs of the three stories:

Original: Once upon a time, in a quaint little village nestled between rolling hills and a crystal-clear river, there lived a young girl named Elara. Elara was known for her boundless curiosity and her insatiable desire to explore the world beyond the village.

1st Variation: In the distant future, in a world where technology and nature coexisted in delicate harmony, there existed a unique being named Z78-ß∆-9.06Q. Z78, as they were affectionately called by the inhabitants of their futuristic city, was an advanced humanoid created by the fusion of artificial intelligence and the remnants of ancient, mystical energies.

2nd Variation: In a land of sweets and confections, there existed a whimsical kingdom where the inhabitants were made entirely of delectable treats. In the heart of this sugary realm stood a lively candied kumquat named Elara. With her vibrant orange hue and a sugary glaze that shimmered in the candy sunlight, Elara was known for her insatiable sweetness and an adventurous spirit that matched her citrusy flavor.

I’ve used color-coding to indicate analogous sections in the three stories. I could do the same thing for each of the other paragraphs in the stories. It’s clear that the we are dealing with the one story structure that is being realized in three different worlds. That structure is invariant across the three stories. If you wish, you can think of the story structure as a set of slots and fillers, in which case we have one structure of slots with the fillers being chosen from three different universes. I'm guessing that in an 'old school' symbolic story grammar, that invariance is specified by constraints on what can fill the slots. How is it handled in the LLM? By the topology of the network?

Is the underlying LLM organized into patterns of slots (in one place) and different sets of fillers (in other places)? If so, how? What does that even mean?

If we were dealing with stories generated with a ‘classical’ symbol-based story grammar, I’d say that the grammar has implemented a conceptual structure sometimes known as the Great Chain of Being, which I’ve discussed at some length here:

Ontology in Knowledge Representation https://www.academia.edu/238610/Ontology_in_Knowledge_Representation

Why does ChatGPT behave as though it “knows” about that knowledge structure? How is that structure realized in the pattern of weights in the underlying LLM?

About “Elara”

A young woman named “Elara” shows up in many stories, see:

ChatGPT tells 20 versions of its prototypical story, with a short note on method, Version 2, https://www.academia.edu/108129357/ChatGPT_tells_20_versions_of_its_prototypical_story_with_a_short_note_on_method_Version_2

Why that name? I did a Google search on the name “Elara” and came up with 13,000,000 hits:

https://www.google.com/search?client=firefox-b-1-d&q=Elara#ip=1

I had no idea that the name was so popular. In particular, I had no idea that there was so much commercial branding associated with the name.

I suspect that that provides at least a partial explanation of why ChatGPT treats Elara as the prototypical protagonist in a prototypical story. I wonder what would turn up if I did searches on other names popular in ChatGPT’s stories. How would one get exact counts of how many times each name appears on the web?

* * * * *

Note: You might want to compare this post with an earlier one: St. George Variations [ChatGPT].

The Elara story and two variations

Tell me a story.

Once upon a time, in a quaint little village nestled between rolling hills and a crystal-clear river, there lived a young girl named Elara. Elara was known for her boundless curiosity and her insatiable desire to explore the world beyond the village.

One day, as she wandered into the nearby enchanted forest, she stumbled upon a hidden clearing bathed in golden sunlight. In the center of the clearing stood an ancient, moss-covered stone pedestal with a mysterious, glowing orb resting upon it. Mesmerized by the radiant glow, Elara reached out to touch the orb, and as soon as her fingers made contact, a warm energy surged through her.

In that moment, Elara felt a connection with the magic of the forest. The trees whispered ancient secrets, and the animals approached her with trust. She discovered that she had been bestowed with a unique gift – the ability to communicate with the creatures of the forest and harness the magic that flowed through it.

Word of Elara's newfound abilities spread throughout the village, reaching the ears of the wise elder, Thalia. Intrigued by Elara's connection to the forest, Thalia approached her and revealed that the village had long been protected by the enchantment of the magical orb. However, the magic had weakened over the years, and the village was now vulnerable to external threats.

Thalia explained that Elara was the chosen one, destined to embark on a quest to restore the orb's power and safeguard the village. Determined and fueled by a sense of responsibility, Elara accepted the challenge.

Guided by the whispers of the forest and the wisdom of Thalia, Elara journeyed through dense woods, crossed rushing rivers, and climbed towering mountains. Along the way, she encountered mythical creatures, each offering a piece of the puzzle needed to restore the orb's magic. Elara's courage and kindness won the trust of these beings, and they joined her on her quest.

As Elara approached the final destination, a dark and foreboding cavern, she faced her greatest challenge. A shadowy figure, drawn by the weakening magic, emerged to thwart her efforts. With the support of her newfound friends and the strength within her, Elara confronted the darkness, unraveling the true power of the orb.

In a dazzling display of magic, the orb's glow intensified, dispelling the darkness and restoring balance to the enchanted forest. The creatures celebrated, and the village thrived once again under the protection of the rejuvenated magic.

Elara returned home, her heart filled with gratitude for the journey that had shaped her into the village's guardian. As she looked out over the now vibrant and lively landscape, Elara understood the importance of harmony between nature and humanity. The tale of the young girl who saved the village became a legend, passed down through generations, reminding all who heard it of the enduring magic that resides in the connection between people and the natural world.

Wednesday, December 20, 2023

The Chatster offers up some saccharine holiday bonbons

I’m going to give you a short story. I want you to tell me three different stories that imitate that story in whatever way you choose. Here’s the story:

Once upon a time, in a land far, far away, there was a giant chocolate milkshake named Aurora. Aurora was a delicious treat, loved by all who knew her. She had a creamy chocolate exterior and a tall, fluffy whipped cream top, and was known for her rich, chocolatey flavor.

One day, a terrible heatwave came to the kingdom and began to melt all the desserts. The heat was fierce and powerful, and none of the ice cream or cakes were able to withstand it. The people lived in fear, not knowing what to do.

But Aurora was not afraid. She knew that she had to do something to save her people, and so she set out to find the source of the heatwave. She traveled for days, braving the hot sun and the baking heat, until she finally reached the volcano that was causing the heatwave.

With her ingredients slowly melting, Aurora approached the volcano and began to cool it down with her cold and creamy exterior. The volcano was enchanted by her deliciousness, and it slowly began to cool down. Aurora continued to cool it down, soothing the volcano until it was completely cool.

With the heatwave no longer a threat, Aurora returned to the kingdom as a hero. The people cheered and celebrated her bravery, and Aurora was hailed as a champion. She lived happily ever after, knowing that she had saved her kingdom and its people.

Story 1: The Icy Guardian

In the serene realm of Frozen Delights, there existed a legendary ice sculpture named Frostine. Crafted from the purest ice and adorned with sparkling snowflake patterns, Frostine was a symbol of elegance and cold beauty in the land. Her very presence could drop the temperature and bring a sense of calm to the warmest of days.

But an unexpected scorching heatwave swept across the kingdom, endangering all frozen entities. Icebergs were melting, snowflakes evaporating, and Frostine knew she had to act. With determination, she embarked on a quest to discover the source of this unusual heat.

Traversing valleys and braving the heat, Frostine reached the heart of the desert where a blazing sunstone, embedded in an ancient temple, emanated scorching rays, disrupting the kingdom's harmony. With her chilly touch, Frostine delicately approached the sunstone, cooling it with her frosty aura. The sunstone, captivated by her ethereal grace, gradually dimmed its fiery glow.

As the sunstone cooled, the heatwave dissipated, and Frostine returned as the savior of her realm. The inhabitants rejoiced, celebrating her valor. Frostine, now hailed as the Guardian of Cool, continued to watch over the kingdom, preserving its icy tranquility for generations to come.

Story 2: The Soothing Melody

In the musical realm of Harmonia, there existed a symphony of flavors, led by Maestro Gelato, a renowned gelato maestro renowned for his melodious blends. His compositions were renowned for their harmonious mix of flavors that delighted all who tasted them.

However, an unexpected cacophony disrupted the harmony of the land—a scorching heatwave that threatened to melt the sweetness of Harmonia. Gelato recognized the urgency of the situation and embarked on a musical odyssey to discover the source of this discordant heat.

Journeying through valleys and plains, Gelato reached the heart of the land where a fiery elemental drum, fueled by the sun’s rhythmic beat, emitted intense heatwaves. With his flavorful symphony, Gelato approached the elemental drum, creating a counter-melody that harmonized with the sun's rhythm, cooling the fiery beats.

As the elemental drum's intensity diminished, the heatwave dissipated, and Gelato returned as the hero of Harmonia. The residents celebrated his ingenuity, and Gelato continued composing delicious symphonies, ensuring the land remained in perfect culinary harmony.

Story 3: The Sweet Enchanter

In the enchanted realm of Sugardom, there resided a confectionary sorceress named Seraphina. Seraphina was known for her magical recipes and spellbinding desserts that captivated the hearts and taste buds of all who indulged in her treats.

One day, an unexpected surge of scorching heat swept through Sugardom, causing all sweets to melt and lose their magical essence. Seraphina, understanding the urgency, embarked on a quest to uncover the source of this disruptive heat.

Traversing through candy forests and caramel rivers, Seraphina reached the heart of Sugardom, where a mischievous fire sprite danced around a blazing bonfire, emitting relentless heat. With her enchanting desserts, Seraphina approached the fire sprite, offering delectable confections that enticed the sprite and gradually cooled its fiery dance.

As the fire sprite's heat diminished, the scorching heatwave subsided, and Seraphina returned as the enchantress who saved Sugardom. The inhabitants celebrated her magical prowess, and Seraphina continued crafting mesmerizing desserts, ensuring that the sweetness of Sugardom remained unharmed for years to come.

Thursday, November 30, 2023

The OpenWHALE Rounds the Horn and Vanishes, An Allegory about the Age of Intelligent Machines

In less than a week I’ll be posting a new essay to 3 Quarks Daily. The current working title: “Investing in AI is like buying shares in a whaling voyage helmed by a man who knows all about ships and nothing about whales.” I expect to keep the whaling theme, though the exact wording may change. Since I take the whaling reference from Marc Andressen I thought I’d explicate the current regime change at OpenAI in whaling terms. To that end I’ve summoned the yarn-craft, the word-weaving, the verbal tap-dancing, and virtuoso bullshittery, of the inimitable, but indefinitely replicable, ChatGPT. 

BTW, ChatGPT's habit of packaging everything it can in the form of a numbered list has been bugging me for a while. But it does lend a certain Melvillian touch to this story, though it's more like Herbert Melville than his more accomplished cousin, Herman. But that Herman did love to run on and on with the lists and the analysis. It was a marvel.

As for Biff Roddenberry, that was an inspired move by the old Chatster. He knew he couldn’t grasp the nuances of Gene’s way with dialog. So he created Biff and blamed the Star Trek elements on him.


Scrimshaw Among the Stars: "Thar she blows!"

The OpenWHALE is seven weeks out of Nantucket and approaching the Cape of Horn by way of the Kessel Run. The seas are higher than a Saturn V, the wind’s blowing 90 knots going on 100, and the rain’s wetter than a guppy’s tail. Science Officer Starbuck is talking to First Mate Scotty on the poop deck.

“Can you lay on more sail, Scotty, Cap’n AhLab says we gotta’ shoot the Moon if we’re to survive these wicked seas.”

“I dunna’ know, Mr. Starbuck, she’s already flyin’ more sails than a pig has wings. The main mast is cracked, the fore topgallant is fixen’ to blow, the keel’s had an unplanned disassembly, and the jibs are jabberin’ like a barrel full o’ monkeys. But I’ll do what I can.

The wind howled through the rigging of the OpenWHALE, making it sound like a banshee in the midst of a storm. The ship creaked and groaned as it rode the monstrous waves, each one threatening to swallow it whole.

Science Officer Starbuck gripped the rail, his eyes squinting against the driving rain. "We're in for a rough ride, Scotty. If we don't make it past the Cape of Horn, we'll be dancing with Davy Jones before we know it. Cap'n AhLab's orders are clear – shoot the Moon or face the abyss."

First Mate Scotty, his weathered face etched with lines of concern, nodded grimly. "Aye, Mr. Starbuck, I'll have the lads workin' double-time. But I won't lie to ye, the OpenWHALE's taken a beatin' like she's never seen. This ain't no ordinary storm; it's like the devil himself is breathin' down our necks."

As Scotty barked orders to the crew, men scrambled to secure lines and reef sails. The ship pitched and rolled with each monstrous wave, testing the mettle of even the hardiest sailors. The main mast groaned under the strain, the fore topgallant threatened to snap like a matchstick, and the jibs flapped wildly, protesting the abuse they were enduring.

Starbuck, his mind racing faster than the ship, checked the instruments on the poop deck. "Scotty, we need every bit of speed we can muster. If we don't make it through the Kessel Run, those seas will make a graveyard of us."

"Aye, Mr. Starbuck, but I canna' change the laws of physics. We're pushin' the old girl to her limits, and she's showin' it," replied Scotty, his eyes reflecting a mix of determination and worry.

Just then, a lookout's shout cut through the roar of wind and waves. "Land ho! Cape of Horn dead ahead!"

Starbuck squinted through the driving rain, his heart sinking as he saw the ominous silhouette of the Cape on the horizon. "Scotty, we're in the belly of the beast now. Brace yourselves, and may the sea gods be with us."

The OpenWHALE, battered but defiant, plunged forward into the maw of the Cape of Horn, each wave a reminder of the relentless power of the southern seas. The crew, faces etched with a mix of fear and determination, held fast to their duties as they sailed on, chasing the elusive safety beyond the storm.

How did they recover from the unplanned disassembly of the keel?

The unplanned disassembly of the keel was a dire situation for the OpenWHALE, as the keel is a critical structural component that provides stability and strength to a ship. Recovering from such a mishap at sea would be a formidable challenge for the crew, requiring resourcefulness and quick thinking.

1. Assessment of Damage: Upon discovering the keel's disassembly, the crew, led by the experienced First Mate Scotty, would have immediately assessed the extent of the damage. Understanding the severity of the situation was crucial in determining the best course of action.

2. Emergency Repairs: With the ship in the midst of a violent storm and the Cape of Horn looming, the crew would have worked swiftly to make emergency repairs. This might involve using whatever materials were on hand to secure the keel and prevent further disassembly. Ropes, chains, spare planks, and other available resources would be pressed into service to stabilize the keel as much as possible.

3. Improvised Reinforcements: The crew, under Scotty's direction, would likely have improvised reinforcements to provide additional support to the compromised keel. This could involve securing makeshift braces or bindings to hold the keel together and distribute the stress more evenly.

4. Utilizing Whaling Equipment: Whaling ships were equipped with a variety of tools and materials for processing whales. Some of these tools, such as heavy chains and harpoon lines, could be repurposed to provide additional structural support to the damaged keel.

5. Crew Collaboration: The entire crew, from seasoned sailors to greenhorns, would have been mobilized to contribute to the repair efforts. Teamwork and coordination would be crucial, with each member playing a role in implementing the emergency fixes.

6. Continuous Monitoring: Even with the makeshift repairs, the crew would need to continuously monitor the situation. Stormy seas and powerful waves could put immense strain on the damaged keel, and adjustments or additional reinforcements might be necessary as the ship navigated through the rough waters.

7. Adaptability and Innovation: Navigating a whaling ship through the Cape of Horn with a compromised keel would require adaptability and innovation. The crew, including Science Officer Starbuck, might come up with creative solutions on the fly to address new challenges that arose during the voyage.

Surviving such a predicament would demand not only the skills of the sailors but also a bit of luck and a heavy dose of perseverance. The crew's ability to collaborate and innovate under extreme conditions would determine whether the OpenWHALE could overcome the challenges posed by the unplanned disassembly of the keel and successfully navigate through the stormy waters around the Cape of Horn.

What was Captain AhLab doing during this emergency?

Captain AhLab, recognizing the severity of the situation with the unplanned disassembly of the keel and the impending dangers of navigating through the Cape of Horn in the midst of a violent storm, would likely be deeply involved in several critical aspects of managing the crisis.