Showing posts with label AI-whaling. Show all posts
Showing posts with label AI-whaling. Show all posts

Tuesday, July 28, 2026

Framing my discussion of The God Test, Part 1: Rorschach, reason, and whaling – [GT-3]

I’ve got to bite the bullet: I’m just going to have to go through a bunch of (preliminary) stuff before I can really engage with The God Test. My current target is to be in a position to publish a proper review of the book in 3 Quarks Daily for the week of August 9.

Rorschach Recap

I want start by recapping the Rorschach metaphor I introduced in the previous post, More on how I’m approaching The God Test – Rorschach! [GT-2]. What I like about it is that has a shape, there’s something there, but it’s not clear what. So we have little choice but to project onto it in order to (begin to) make sense of it.

First: It is a new kind of thing, an artifact we can converse with in an open-ended and natural way. The steam engine was the same kind of thing. It was an inanimate object that moved over the surface of the earth under its own power. Previously only animals (& humans as animals) had that power. So it becomes an iron horse. Just what are AIs? What’s their nature? That’s one thing.

Second: How it works is opaque. We know how to create large language models (LLMs), but we don’t know how they work. That’s new. We may not have understood the deep physics of the steam engine, but we certainly knew how they worked.

Third: We don’t know what they portend for the future. To some extent this is a function of the first two: How can we, how should we, interact. But it is also a function of the future, which is undetermined. We just don’t know.

Rhetorical force over reason

This is an argument I made in the first working paper I published after the release of ChatGPT in November of 2022: ChatGPT intimates a tantalizing future; its core LLM is organized on multiple levels; and it has broken the idea of thinking (February 6, 2023).

What do I mean by that, has broken the idea of thinking? Prior to ChatGPT it was obvious that humans could think and computers could not. [Yeah, I know, there’s Deep Blue defeating Kasparov in chess. That just changes the dates, not the argument.] The difference in performance was so obvious that the fact that we don’t really know how humans think wasn’t much of an issue. Now it is. Sure, we can still say that we can think and the AI’s can’t, but that’s just a line and without good explanations on both sides of the line, it seems a bit arbitrary, if not desperate.

I made a particular argument about Searles’ (in)famous Chinese Room thought experiment. I read it when it was first published in Brain and Behavioral Science in 1980. I wasn’t impressed. Why not? He didn’t say anything about any of the techniques used in AI or computational linguistics (CL). How could anyone possibly take that seriously?

He talked about intention, that’s how. Meaning requires intention and only living things can have intention, a remark he made at the end of the article. Without intention the most you get is syntax, but no meaning. Searle could get away with that because, in the first place, the concept of intention has a long history within philosophy – it has a subtle meaning, but that can wait for a later post – and so philosophers, his main audience, were comfortable with it. That’s one thing.

But there’s something more important, something that we can see only in retrospect, and that’s the simple fact computers very obviously could not translate from Chinese into English or into any other language. That difference carried tremendous weight. We don’t have a subtle behavioral difference between computers and humans that requires a subtle and sophisticated argument. To a first approximation, almost any argument would do. As far as I was concerned, “intention” was just a fancy word for something we don’t understand. But that’s not an argument anyone needs to take seriously. The behavioral distance is quite sufficient to carry the argument for those who insist that computers can’t and will never be able to think like humans.

Now the behavioral evidence has changed. Sure, differences remain, but the evidence is shifting. The old arguments remain and those who believed them still do so, but it’s getting harder. The need for explicit arguments grounded in explicit accounts of computers, and also brains, is growing.

Whaling and expertise

What’s an expert in machine learning and LLMs actually expert in? For some time now I’ve been arguing that investing in AI is like investing in a whaling venture where the captain and crew of the ship know all there is to know about the ship and how to handle it but know little or nothing about whales and their behavior and about navigating around the Cape Horn and in the South Pacific, where the whales live. What are the chances of that voyage being successful? Not very good.

The people who have created the current AI technology are like that captain and crew. The know how to sail the ship. But they don’t know much about language or cognition. They don’t actually know much about the human mind. Here my point is not about the fact that the models are opaque, but that human language and cognition are highly structured and they don’t believe that one needs to know (much of) anything about that not only to build AI but to make confident prediction about the future of AI.

Gary Marcus, Subbarao Kambhampati, and others have been consistently arguing that, yes, the current technology is remarkable, but we are going to have to adopt classical symbolic techniques if we are to fully develop the technology so that we have accurate and safe systems. Marcus is arguing from his knowledge of human language and cognition. As far as I can tell, Wright doesn’t take that seriously. I know that he had Marcus on his NonZero podcast, and that he lists Marcus in his acknowledgements, but that he doesn’t discuss Marcus’s ideas. I conclude that he doesn’t take that line of argument seriously.

That’s a mistake, but this is not the place to make my own arguments on this issue. My point is simply that expertise in AI is no generally construed to encompass knowledge of, expertise in, human cognition and language. I can’t see how that is going to work out well in the future.

[Note: If you’re curious about my views, on this subject, read the article linked in the first paragraph of this section. My views all over the place here at New Savanna, particularly around the work of the mathematician Miriam Yevick. Also, check out the experimental work I’ve done with LLMs.]

Wednesday, April 29, 2026

AI spending is out of control [Ahab in pursuit of Moby Dick]

Karen Weise, A.I. Spending Sets a Record, With No End in Sight, NYTimes, April 29, 2026.

For the past two years, Amazon, Google, Microsoft and Meta have repeatedly set records for how much they are spending on artificial intelligence.

On Wednesday, the four giants did it again.

In the first three months of the year, the four companies reported in their financial results, they plowed a total of $130.65 billion into capital expenditures, largely spending on data centers that power A.I. That figure — which was another record — was more than three times what the Manhattan Project cost to develop nuclear bombs and 71 percent higher than what the tech giants spent in the same quarter a year earlier.

All of the companies said they would be spending even more, totaling roughly $700 billion this year. Meta, for one, raised its spending forecast for 2026 to between $125 billion and $145 billion, up from its previous prediction of $115 billion to $135 billion. Google also boosted its projection, to at least $180 billion, and said its spending would be “significantly” higher next year.

The big four – Google, Microsoft, Amazon, and Meta can afford it because they “continue to dominate in core businesses that spew cash, such as serving ads on YouTube or Instagram, delivering items in a few hours or tallying cells in Excel.” They've entered into circular relationships and Anthropic and OpenAI which have, in turn, “committed to spending hundreds of billions on computing power that the tech giants provide.” Yada yada so forth and so on et cetera et cetera:

Some of the tech companies have justified their building binge by saying they cannot meet all the demand. But analysts said there were risks if the companies became too dependent on two young customers: OpenAI and Anthropic.

More than 40 percent of Microsoft’s $625 billion in outstanding cloud contracts, for example, come from OpenAI, the company said in January. This week, Microsoft and OpenAI announced new terms that loosened their ties.

Betting so much on OpenAI and Anthropic is a gamble. But even if the start-ups flop, the tech giants are likely to weather the losses because of their size, scale and other businesses, said Matt Stucky, who manages tech investments for Northwestern Mutual.

“The core business,” he said, “is good.”

I think they're taking the economy for a Nantucket sleigh ride.

Monday, April 27, 2026

Those new AI metrics: From AGI to bragawatts

Erin Griffith, How Do You Measure A.I. Firms’ Gargantuan Energy Plans? In ‘Bragawatts.’ NYTimes, April 26, 2026

The term started more than a decade ago in the energy industry, used to describe power from a solar or wind project that had no chance of being built. Last year, A.I. executives began boasting with increasing boldness about their plans. A.I. watchers, including Waldemar Szlezak, the head of infrastructure at the private equity firm KKR, repurposed the term in a Financial Times column imploring investors to look past the A.I. hype and focus on the reality of today’s power grid. Since then, the term has popped up in media headlines, analyst reports and on social media, typically with a healthy dose of skepticism about how quickly such projects can realistically be built.

The numbers being announced are staggering. Nvidia estimated that as much as $4 trillion would be spent on A.I. infrastructure this decade. OpenAI said it had committed to spend $1.4 trillion to build data centers around the world. (It later lowered that target to a mere $600 billion.)

Brad Gastwirth, global head of research and market intelligence at Circular Technology, a supply chain services firm, said that projects highlighting a gigawatt or more of energy are the most likely to be bragawatts.

“That’s where you can have some scratching of the heads,” he said.

Likewise for any infrastructure projects announced by companies that haven’t already secured the land to build the project, he noted. “That’s definitely the braganomics.”

There's more at the link.

Monday, February 3, 2025

TNSTASFL, that goes for knowledge too. OR: Why there’s so much AI hype. [once more with the whales]

Why is there so much hype about AI? Sure, it’s new, it’s interesting, and certainly has transformative potential. That’s one thing. But all this talk about AGI in five years, possibly followed by ASI, and then, who knows, perhaps DOOM! The machines will take over and humans will either be reduced to slavery or be eliminated entirely. Where’d that come from?

[AGI=artificial general intelligence. ASI=artificial superintelligence.]

Well, yeah, there’s fantasy. But I think something else is going on as well.

While there are other things going on, the excitement is centered on LLMs (large language models), the things that power chatbots such as ChatGPT, Gemini, Claude, and others. You don’t have to know much of anything about language, cognition, the imagination, or the human mind in order to create an LLM. You need to know something about programming computers, and you need to know a lot about engineering large-scale computer systems. If you have those skills, you can create an LLM. That’s where all your intellectual effort goes, into creating the LLM.

The LLM then goes on to crank out language, and really impressive language at that. That doesn’t require any intellectual effort from you. As far as you’re concerned, it’s free. It took some genuine insight to come up with the transformer architecture. That’s what all these LLMs are built on. That was created by engineers at Google.

OpenAI got ahold of the idea and built their GPT series. That’s all engineering. I saw some output from GPT-2. Not very impressive. GPT-3 was much more impressive. It was built on the same design as GPT-2, but just bigger. GPT-2 had 1.5 billion parameters; GPT-3 had 175 billion. I assume that the size difference required some very skillful engineering; but the underlying concept was the same.

From an intellectual point of view the dramatically increased performance from one model to the next was free. The same goes for the difference between GPT-3 and GPT-3.5 (which powered the original ChatGPT). And so it goes for GPT-4. (We’re still waiting for GPT-5.)

In that situation, where increased performance, even radically increased performance, imposes no similar increase in intellectual insight, in scientific understand if you will, in that situation it’s easy to give-in to one’s fantasies and generate hype by the bucket load. Forget the buckets. Let’s go for swimming pools, giant Olympic-sized swimming pools filled with hype.

And so, once again, I trot out my whaling analogy. Nineteenth-century whaling ships were three-masted square-rigged vessels, just like the merchant ships used between ports in Europe and America for trading purposes. The skills needed to sail them are quite different. Now, take an expert captain and crew from a merchant vessel, put them on a whaler, and what happens? For one thing the whaler has a try-works midships. It’s used to render whale oil from blubber. The merchant seamen have never seen that. But that’s a skill easily learned.

But sailing the treacherous seas around Cape Horn, that’s another matter. Once you’re through, now you’ve got to hunt whales in the Pacific Ocean. If you’ve never done it before, how do you know where to look? And if you’ve spotted a whale, what then? How do they behave? How do you after them and kill them? No, I’m afraid the skills of a merchant seaman aren’t adequate to the task.

That’s what we’ve got in the case of deep leaning, LLMs, and language. The people who’ve created the technology don’t know anything about language and cognition. They get the performance for free and don’t have intellectual tools for thinking about what’s going on. So they throw hype into the void and hope it’ll make things right.

It won’t. They’re lost, and don’t know it.

Drinking Silicon Valley Joy Juice is not a formula for long-term success. 

* * * * *

NOTE: I used ChatGPT to create the image. If you look closely at the sign in the upper left you'll see that it elaborated TM (for trademark) into TMI (too much information), which is interesting, but not appropriate.

Tuesday, January 28, 2025

The DeepSeek breakthrough – What’s it mean? [On the difference between engineering and science]

Frankly, at the moment I’m inclined to think it means that Silicon Valley just got handed its lunch. Strutting around about AGI this, $500 billion that... Voilà! We’re Masters of the Universe. They may or may not be “masters of their domain,” (likely not) in the Seinfeldian sense, but Masters of the Universe they are NOT.

Engineering automobiles, or rockets – that’s one thing. Engineering artificial minds. What’s going on is unprincipled hacking. Throw enough person-hours, compute, and money at it and, sure, you’ll do something....

Mutt: “Slow down, son, you’re ranting!”

Jeff: “OK, OK, I’ll slow down.”

Full disclosure: My priors

I was trained in computational semantics back in the mid-1970s by David Hays, who had been a first-generation researcher in what was originally called machine translation (MT) but got rebranded as computational linguistics (CL) in the mid-1960s when it got defunded for failing to deliver practical benefit to the US Military. At the time I worked with him Hays had shifted his attention to semantics. In keeping with that focus he read a great deal about cognitive and perceptual psychology and neuroscience. He wanted his models to have some grounding in scientific fact.

He tended to think of AI researcher as unprincipled hackers. If the program worked, that’s all that mattered. They didn’t seem very interested in possible psychological reality.

And that’s what the current regime of work on deep learning looks like to me: hacking. To be sure, it’s brilliant, perhaps inspired hacking. If I didn’t believe that I wouldn’t have spent a great deal of my time over the last two years working with ChatGPT and now Claude 3.5, working to tease out clues about what’s going on under the hood.

That’s a problem, no one actually knows how these models work. Oh, there’s interesting research on that problem, some of it under the rubric of mechanistic interpretability. But that doesn’t seem to be a priority. Instead, the emphasis is on scaling up, more data, more compute, more parameters, more more more! (Slow down son!)

The impact of DeepSeek

Given that scaling up has worked in the past, and in the absence of any deep insight into how these things work, the scaling hypothesis, as it is sometimes called, had a certain superficial validity. The Chinese have just blown a big hole in the scaling hypothesis. Here’s Kevin Roose in The New York Times:

The first is the assumption that in order to build cutting-edge A.I. models, you need to spend huge amounts of money on powerful chips and data centers.

It’s hard to overstate how foundational this dogma has become. Companies like Microsoft, Meta and Google have already spent tens of billions of dollars building out the infrastructure they thought was needed to build and run next-generation A.I. models. They plan to spend tens of billions more — or, in the case of OpenAI, as much as $500 billion through a joint venture with Oracle and SoftBank that was announced last week.

DeepSeek appears to have spent a small fraction of that building R1. [...] But even if R1 cost 10 times more to train than DeepSeek claims, and even if you factor in other costs they may have excluded, like engineer salaries or the costs of doing basic research, it would still be orders of magnitude less than what American A.I. companies are spending to develop their most capable models. [...]

But DeepSeek’s breakthrough on cost challenges the “bigger is better” narrative that has driven the A.I. arms race in recent years by showing that relatively small models, when trained properly, can match or exceed the performance of much bigger models.

What that means is that the industry’s intuitive understanding of what the late Dan Dennett liked to call the design space for AI, that understanding is wrong. Yes, bigger does sometimes/often get you more performance. But if smaller can yield comparable performance, than something else is going on, something we can’t identify.

Were the Chinese just lucky? Or do they know something, something deep, that we don’t? In the absence of any further information, I’d guess that it’s both.

Can we figure out what they did and do it ourselves? Sure, no problem. DeepSeek is an open-source model and the researchers have released good documentation. We’ll replicate what they’ve done and perhaps improve on it, as they will also do. That’s not the issue.

The issue is understanding. We already knew that these models are black boxes. And we guessed/hoped that buiding a bigger box would make it better. We now know that that’s not necessarily true. What else don’t we know? How are we going to find out?

Thursday, January 23, 2025

ChatGPT critiques benchmarks as a measure of LLM performance and then elaborates on my whaling analogy for what’s wrong with the AI business.

I just read this article: Kevin Roose, When A.I. Passes This Test, Look Out, NYTimes, Jan. 23, 2025.

“Fiddlesticks! Fiddlesticks and moonbeams!” said I to myself. “Not a benchmark-to-end-all-benchmarks. Will these people never learn? Why do they keep looking for the house keys under the streetlight when they lost them at the front door? Are they drunk or something?” Oh, they don’t know what they’re doing, do they?

So I put the article aside, intending to write up a post about it and to register my complaint about the inadequacy of benchmarks. Then I had an idea: I’ve not used ChatGPT in a while, so why don’t I have it summarize the article for me.” Once the Chatster had done that, I had an idea: Why not ask it what it thinks about bench-marking?

I didn’t have any particular expectations about what kind of answer it would give. To my surprise and delight it came up with a substantial critique. I decided to press on and told the story about how I came to write about Spielberg’s Jaws. Again, the Chatster rose to the occasion and wrote a mini-essay, “The Role of Discovery in Intellectual Work.” It was on point, so I gave it my whaling analogy. Chatster broke it down very nicely.

“Damn!” thought I to myself. “Chatster’s on FIRE!”

And away we go. 

NOTE: See my post, Competence and Performance in Bench-marking LLMs.

* * * * *

Summarize this article, please. Assume I am sophisticated and well-educated.

The article, "A Test So Hard No AI System Can Pass It — Yet," examines the development and significance of a new benchmark called "Humanity’s Last Exam," designed to evaluate the limits of artificial intelligence (AI) capabilities. Created by Dan Hendrycks and his team at the Center for AI Safety in collaboration with Scale AI, the test comprises 3,000 highly challenging questions spanning disciplines like philosophy, physics, and engineering, contributed by experts. Its goal is to assess AI's ability to perform complex intellectual tasks, approximating a general intelligence score.

The impetus for this exam arises from the rapid advancements in AI, which have rendered traditional benchmarks obsolete as models like those from OpenAI and Google surpass graduate-level challenges. Despite their advanced abilities, current AI models failed Humanity’s Last Exam, with the best-performing model achieving only 8.3%. Hendrycks anticipates significant improvement, potentially marking a point where AI systems surpass human experts in answering complex questions.

The article highlights concerns over AI's uneven progress—astounding in specialized tasks but deficient in others—and discusses the broader implications for fields requiring nuanced problem-solving beyond exams, such as research. It concludes with a call for innovative methods to measure AI’s impact beyond standardized testing, particularly as these systems continue to evolve.

What do you think of this whole “benchmark” business? How useful are they?

Benchmarks like Humanity's Last Exam have value, but their utility is nuanced. Here’s my take:

Strengths of Benchmarks:

  • Tracking Progress: Benchmarks serve as yardsticks, offering a quantitative way to measure the trajectory of AI systems over time. They allow researchers to compare different models and identify strengths and weaknesses systematically.  
  • Challenge Motivation: By creating hard problems, benchmarks inspire innovation. They push AI researchers to design models that can tackle tasks requiring deeper reasoning, creativity, and generalization.  
  • Defining Boundaries: Benchmarks like this can help delineate areas where AI falls short, highlighting tasks that still require human expertise, such as novel discovery or conceptual synthesis.

Limitations and Challenges:

  • Artificial Metrics: Benchmarks may overfit AI development to specific tasks, leading to “training to the test” rather than true generalization. Models can excel at benchmarks without demonstrating real-world problem-solving abilities.  
  • Narrow Framing: Many benchmarks, including even sophisticated ones, focus on clearly-defined problems with deterministic answers. But much of human intellectual labor—creative writing, ethical decision-making, or forming new scientific hypotheses—doesn’t fit neatly into these molds. 
  • Temporary Relevance: As seen with earlier AI challenges, benchmarks can become obsolete quickly. Once AI models surpass them, they lose their diagnostic power, necessitating constant updates or more difficult tests.  
  • Underestimating Complexity: Benchmarks often fail to account for the broader, messier contexts in which humans operate. AI might pass a theoretical test on physics but struggle to apply those principles in practical, experimental settings.

Broader Perspective:

Benchmarks are useful for building a snapshot of what AI can do at any moment, but they must evolve to remain meaningful. To go beyond mere scoring, we might need:

  • Dynamic Benchmarks: Tasks that adapt to the model’s performance, such as questions whose answers are not yet known or which require multi-modal reasoning (e.g., combining visual and textual information).  
  • Real-World Applications: Shifting focus from exams to practical contributions, like assessing AI’s impact on medicine, climate science, or engineering breakthroughs. 
  • Human-Centric Evaluation: Considering AI’s utility in collaborative settings, where it aids but does not replace human expertise.

Benchmarks like Humanity’s Last Exam have an important place in measuring AI progress. However, we must also look beyond them to gauge AI’s real-world significance and its potential to address humanity’s most pressing challenges. They’re a piece of the puzzle, not the whole picture.

* * * * *

I agree. The problem I have with benchmarks is that it asks for something specific. But that’s not always how life is lived. Often we have to figure out whether there’s a problem or something of interest. Once we’ve done that, we have to figure out an approach.

I’m trained as a literary and cultural critic. I read a lot of books, see a lot of films and TV shows. I write about many of these texts, broadly considered. It’s up to me to notice that there’s something interesting about a text. For example, a couple of years ago I was watching Steven Spielberg’s movie “Jaws.” After I had watched the movie I read the Wikipedia entry, which mentioned various interpretations. It also mentioned that the sequels were not as good. So I looked at them (couldn’t finish 3 and 4). All the while I was thinking.

I started comparing Jaws 2 with the original. The original was much tighter. I started thinking about the role of Quint. “Why did he have to die?” I asked myself. Sure, he wasn’t a nice man, but he hadn’t hurt anyone. All that had to happen to give the movie a “happy” ending was that the shark gets killed. It wasn’t necessary to have anyone else die. So, what’s the point of having Quint die? How would the film be different if he hadn’t been killed by the shark?

It was at that point that it hit me: Girard, sacrifice, mimetic desire. NOW I had a specific question: How do Girard’s ideas apply to Jaws? It took me awhile to come up with an answer that I liked. But that’s how it goes. The really important thing was coming up with the idea to apply Girard’s ideas to the movie. That’s what took the hard work.

These benchmarks don’t test that kind of work at all. But that’s where the really hard and fruitful problems lie.

You’ve pinpointed a profound challenge in evaluating AI: life—and much of human intellectual labor—is far messier and more open-ended than what benchmarks can capture. Your example of watching Jaws illustrates exactly the kind of complex, emergent thinking that is central to creative, critical, and interpretive work. Let’s unpack your critique a bit further.

Sunday, December 10, 2023

State of AI @3QD, plus AI culture and Captain Ahab

Earlier this week I posted another article to 3 Quarks Daily:

The title says it all, sort of. The article centers on the fact that, as far as I can tell, many of our AI experts aren’t so expert, that is, they often don’t know what they’re talking about. People will criticize their arguments, but won’t call them out for not having the expertise they claim they have.

Actually, I don’t quite get that far in the article. That’s one topic for this post. The other is an extension of the whaling metaphor in the title. It seems to me that the mad dash for AGI is a bit like Ahab’s quest for Moby Dick, which did not end well for either of them.

What’s an AI expert expert about? NOT human language & cognition.

I chose that whaling analogy to emphasize the peculiar nature of expertise in machine learning: You don’t have to know much about human language and cognition in order to build an AI engine that mimics human cognitive behavior astonishingly well, much better than anyone would have predicted as recently as 2019, the year before GPT-3 was unveiled. Hence the analogy, and AI expert is like a whaling captain who knows all about his ship, but little about whales.

How did this come about and, more to the point, why do we let them get away with it? I didn’t actually pose the latter question, but my essay did suggest an answer to it: This feature of AI-culture has become quasi-institutionalized so that responsibility for pronouncements made by individual researchers must be apportioned between the culture and the individuals. If we question their expertise directly, rather than simply criticizing their arguments, that critique threatens our quasi-institutional understandings about the scope of that culture. Once we start down that road, what other institutional understandings will start to unravel? Let’s not go there.

But that’s a digression. I’m more interested in saying a bit more about how this situation came about.

As I pointed out in the paper, the issue can be traced back to Turing’s famous paper, “Computing Machinery and Intelligence” (1950). “That’s the paper in which he proposed the so-called Turing Test for evaluating machine accomplishment. The test explicitly rejects comparisons based on internal mechanisms, regarding them as intractably opaque and resistant to explanation, and instead focuses on external behavior.” That was (perhaps) a reasonable thing to do at the time. That test, however, was rendered useless in the late 1960s by Joseph Weizenbaum’s ELIZA, a simple computer program simulated human conversation if a very compelling way. At that time, the actual accomplishments of the discipline were not very compelling, at least to researchers outside the discipline, and over-reaching proclamations (about when computers will surpass humans) had few implications for practical action. That is no longer the case. Billions of dollars are being wagered on predictions offered by AI experts.

If we look at what actually happened – and here I’m sketching out a history I haven’t researched thoroughly, I’m just making this up out of what’s already in my mind, so beware of LLM-like confabulation – we see that early researchers did attend to human cognition. That’s most obvious in the case of chess, where research into human chess playing was undertaken, and the ubiquitous expert systems, where human experts were interviewed about their thought processes as preparation for designing the system. Around the corner, researchers in the sibling discipline of computational linguistics (originally machine translation) called on linguistics and cognitive psychology for insight into the design of these systems.

These two lines research of research came together in a large project sponsored by the Defense Department, the ARPA Speech Understanding Project. It extended over five years in the mid-1970s and involved three separate projects involving perhaps a half-dozen research organizations in universities and other research organizations. They undertook to create systems in which a computer would take spoken language questions and provide answers in written English. This was a massive research effort that recruited expertise both in computer science and engineering and in human perception and cognition. No one researcher was expert in all the disciplines involved, but the enterprise encompassed them all, and I assume that at least some researchers read all the reports produced by the project in which they took part, if not all the reports from all the projects. (As bibliographer for Computational Linguistics at the time, I scanned and prepared abstracts for them all.)

Things began to change in the 1980s, when research in connectionist neural networks re-commenced and statistical machine learning techniques began emerging. These techniques effectively separated computational expertise from domain knowledge. It was no longer so necessary to bring deep domain expertise to bear on the design of computer systems. That divergence widened into a yawning chasm with the use of GPUs in the second decade of this century. But the implicit quasi-institutional understandings that had developed back in the 1950s and 1960s remained in place.

The people who develop the computer systems are THE experts. And they certainly are experts. But as far as I can tell, the level of expertise these machine experts have in linguistics and human cognition is nothing to write home about. In that sense they are like the hapless captain of a whaling vessel who knows about his ship, but not about whales.

This is not a healthy situation.

Ahab, Moby Dick, and AI Doom

When I first came up with my title it was but a device to point out the divergence between AI researcher’s knowledge of their programs and their ignorance about language. I had no intention of elaborating it into a conceit that I’d employ at various points throughout the essay. Nor did I have in mind that I’d actually reference Melville’s Moby Dick. That just happened (near the end).

And now that it has, I wonder. Is the pursuit of (the mythical) AGI like Ahab’s pursuit of Moby Dick? Is their fear that the AGI will turn on them, is that like Ahab’s fear of and animosity toward Moby Dick? Is the underlying psychology pretty much the same despite all the obvious differences between the two passions? I don’t know. But as Edward Mendelson argued back in 1976, Moby Dick is an encyclopedic narrative, one that set out to encompass the whole of mid-19th century America. The danse macabre between Ahab and the whale is not thus a private affair between a man and an animal; it is a figure for something at the heart of America. What? Or was Melville just imagining things?

Whatever.

The high-tech industry’s dash to AI supremacy has a similar sweep. We're on a Nantucket sleigh-ride. Instead of 45 ton Moby Dick pulling a 35 ft. whaleboat, a bunch of techbros and Silicon Valley billionaires have hitched the earth to a rampaging Jupiter (318 times the mass of the earth). Look at the numbers in this tweet:

I recognize every one of those names, though I know more about some than others. I assume that the mythical everyone knows who Elon Musk is, but not many are likely to know who Zvi Mowshowitz is. I don’t either, not really. That is, I can’t recite the story of how he came to be regarded as an expert, but I do read his long posts at LessWrong. Musk thinks there’s a 20% to 30% chance of AI killing us all; Mowshowitz puts the number at 60%. The whole range is from 10% to 90%.

Those numbers are insane, Ahab-like insane. If you were running in a marathon and someone came up to you during the race and told you that there was a mere 10% chance you will die unless you exit the race, NOW, what would you do? You’d stop running, just as you would if the chances were 20%, 50%, or 90%. In what way do these people believe those numbers? To what reality are they tethered? To what extent are the tech industry’s actions guided by perceptions no more grounded in reality than the so-called hallucinations of a large language model? 

* * * * *

Addendum 12.11.23: In terms of the informal game theory argument offered in the section, “Deconstructing AI Doom,” of my article, A New Counter Culture, those numbers are a Schelling point, a rallying point for a new (counter) culture. As such, what’s important about those numbers IS NOT their plausibility, though they do come draped in epistemic theatre intended to create the appearance plausibility, but their distinctiveness. Those numbers are not tethered to the reality of mainstream media, whatever that is. They’re a clear demarcation of a different way of looking at the world, one regarded as superior by its proponents.



  

* * * * *

Addendum: See this most interesting article by Andrew Schenker, After Melville, in The Baffler.

Thursday, November 30, 2023

The OpenWHALE Rounds the Horn and Vanishes, An Allegory about the Age of Intelligent Machines

In less than a week I’ll be posting a new essay to 3 Quarks Daily. The current working title: “Investing in AI is like buying shares in a whaling voyage helmed by a man who knows all about ships and nothing about whales.” I expect to keep the whaling theme, though the exact wording may change. Since I take the whaling reference from Marc Andressen I thought I’d explicate the current regime change at OpenAI in whaling terms. To that end I’ve summoned the yarn-craft, the word-weaving, the verbal tap-dancing, and virtuoso bullshittery, of the inimitable, but indefinitely replicable, ChatGPT. 

BTW, ChatGPT's habit of packaging everything it can in the form of a numbered list has been bugging me for a while. But it does lend a certain Melvillian touch to this story, though it's more like Herbert Melville than his more accomplished cousin, Herman. But that Herman did love to run on and on with the lists and the analysis. It was a marvel.

As for Biff Roddenberry, that was an inspired move by the old Chatster. He knew he couldn’t grasp the nuances of Gene’s way with dialog. So he created Biff and blamed the Star Trek elements on him.


Scrimshaw Among the Stars: "Thar she blows!"

The OpenWHALE is seven weeks out of Nantucket and approaching the Cape of Horn by way of the Kessel Run. The seas are higher than a Saturn V, the wind’s blowing 90 knots going on 100, and the rain’s wetter than a guppy’s tail. Science Officer Starbuck is talking to First Mate Scotty on the poop deck.

“Can you lay on more sail, Scotty, Cap’n AhLab says we gotta’ shoot the Moon if we’re to survive these wicked seas.”

“I dunna’ know, Mr. Starbuck, she’s already flyin’ more sails than a pig has wings. The main mast is cracked, the fore topgallant is fixen’ to blow, the keel’s had an unplanned disassembly, and the jibs are jabberin’ like a barrel full o’ monkeys. But I’ll do what I can.

The wind howled through the rigging of the OpenWHALE, making it sound like a banshee in the midst of a storm. The ship creaked and groaned as it rode the monstrous waves, each one threatening to swallow it whole.

Science Officer Starbuck gripped the rail, his eyes squinting against the driving rain. "We're in for a rough ride, Scotty. If we don't make it past the Cape of Horn, we'll be dancing with Davy Jones before we know it. Cap'n AhLab's orders are clear – shoot the Moon or face the abyss."

First Mate Scotty, his weathered face etched with lines of concern, nodded grimly. "Aye, Mr. Starbuck, I'll have the lads workin' double-time. But I won't lie to ye, the OpenWHALE's taken a beatin' like she's never seen. This ain't no ordinary storm; it's like the devil himself is breathin' down our necks."

As Scotty barked orders to the crew, men scrambled to secure lines and reef sails. The ship pitched and rolled with each monstrous wave, testing the mettle of even the hardiest sailors. The main mast groaned under the strain, the fore topgallant threatened to snap like a matchstick, and the jibs flapped wildly, protesting the abuse they were enduring.

Starbuck, his mind racing faster than the ship, checked the instruments on the poop deck. "Scotty, we need every bit of speed we can muster. If we don't make it through the Kessel Run, those seas will make a graveyard of us."

"Aye, Mr. Starbuck, but I canna' change the laws of physics. We're pushin' the old girl to her limits, and she's showin' it," replied Scotty, his eyes reflecting a mix of determination and worry.

Just then, a lookout's shout cut through the roar of wind and waves. "Land ho! Cape of Horn dead ahead!"

Starbuck squinted through the driving rain, his heart sinking as he saw the ominous silhouette of the Cape on the horizon. "Scotty, we're in the belly of the beast now. Brace yourselves, and may the sea gods be with us."

The OpenWHALE, battered but defiant, plunged forward into the maw of the Cape of Horn, each wave a reminder of the relentless power of the southern seas. The crew, faces etched with a mix of fear and determination, held fast to their duties as they sailed on, chasing the elusive safety beyond the storm.

How did they recover from the unplanned disassembly of the keel?

The unplanned disassembly of the keel was a dire situation for the OpenWHALE, as the keel is a critical structural component that provides stability and strength to a ship. Recovering from such a mishap at sea would be a formidable challenge for the crew, requiring resourcefulness and quick thinking.

1. Assessment of Damage: Upon discovering the keel's disassembly, the crew, led by the experienced First Mate Scotty, would have immediately assessed the extent of the damage. Understanding the severity of the situation was crucial in determining the best course of action.

2. Emergency Repairs: With the ship in the midst of a violent storm and the Cape of Horn looming, the crew would have worked swiftly to make emergency repairs. This might involve using whatever materials were on hand to secure the keel and prevent further disassembly. Ropes, chains, spare planks, and other available resources would be pressed into service to stabilize the keel as much as possible.

3. Improvised Reinforcements: The crew, under Scotty's direction, would likely have improvised reinforcements to provide additional support to the compromised keel. This could involve securing makeshift braces or bindings to hold the keel together and distribute the stress more evenly.

4. Utilizing Whaling Equipment: Whaling ships were equipped with a variety of tools and materials for processing whales. Some of these tools, such as heavy chains and harpoon lines, could be repurposed to provide additional structural support to the damaged keel.

5. Crew Collaboration: The entire crew, from seasoned sailors to greenhorns, would have been mobilized to contribute to the repair efforts. Teamwork and coordination would be crucial, with each member playing a role in implementing the emergency fixes.

6. Continuous Monitoring: Even with the makeshift repairs, the crew would need to continuously monitor the situation. Stormy seas and powerful waves could put immense strain on the damaged keel, and adjustments or additional reinforcements might be necessary as the ship navigated through the rough waters.

7. Adaptability and Innovation: Navigating a whaling ship through the Cape of Horn with a compromised keel would require adaptability and innovation. The crew, including Science Officer Starbuck, might come up with creative solutions on the fly to address new challenges that arose during the voyage.

Surviving such a predicament would demand not only the skills of the sailors but also a bit of luck and a heavy dose of perseverance. The crew's ability to collaborate and innovate under extreme conditions would determine whether the OpenWHALE could overcome the challenges posed by the unplanned disassembly of the keel and successfully navigate through the stormy waters around the Cape of Horn.

What was Captain AhLab doing during this emergency?

Captain AhLab, recognizing the severity of the situation with the unplanned disassembly of the keel and the impending dangers of navigating through the Cape of Horn in the midst of a violent storm, would likely be deeply involved in several critical aspects of managing the crisis.

Friday, November 10, 2023

Whaling, Andreesen, Investment, and the Risk Premium

Whale's tooth scrimshawed by my uncle, Rune Ronnberg.

In yesterday’s ramble I noted that my 3 Quarks Daily column for December would be about the current state of AI and that I would open with the following conceit:

Would you invest in a whaling voyage captained by someone who knows all there is to know about his ship, and is able to helm it in a day sail to and from home port, but whose knowledge of sailing on the open ocean, of the weather, of navigation and, above all else, of whales and whaling, is no greater than that of the typical landlubber?

I then went on to explain how that conceit applied to the AI business and a bit of why I thought it appropriate.

But that’s not what this post is about. Why’d I pick whaling? When the conceit first occurred to me it wasn’t about a whaling vessel; it was just about a sea-going ship. It was only later that I decided it should be a whaler.

Why? Because, Marc Andreesen, the well-known venture capitalist, billionaire, and blatherskite tech-optimist, likes to talk about whaling, specifically, how the financing of whaling ventures is like venture capital financing. That’s next part of this post. I conclude by looking at the overall value of whaling as an investment medium.

Whaling and Venture Capital

Here’s what Andreesen said in a recent interview with Noah Smith:

There’s something very old about what venture capital is -- Tyler Cowen uses the term “project evaluation”, the process of sorting through many possible configurations of people and ideas and then picking a few to back with money and effort to try to create something new and important in the world. In venture capital, this idea traces back to the whaling industry of centuries past, where independent financiers would fund captains and ships to hunt whales -- legend has it this is the origin of the term “carried interest”, which originally meant the share of the whale carried by the ship and kept by the captain and crew. The same “project evaluation” pattern has played out repeatedly for centuries, for many kinds of large-scale, risky projects, from colonial settlements like the Plymouth colony, to music/film/television projects, to the world’s largest private equity transactions.

But of course there’s also something very new about what venture capital is -- we fund the most cutting edge ideas and projects in the world, brand new conceptions of what technology makes possible. The founders we fund routinely break rules and create new models that people think are impossible until they happen. This includes many of the leading edge crypto ideas I discussed above, many of which assume a very different form of industrial organization than a classic joint stock company.

So we sit at the vortex of this combination of the very old and the very new. It’s certainly possible that venture capital itself gets pulled into this vortex and comes out the other side radically transformed, and in fact this is what some of the smartest crypto experts are predicting. And yet...there is, at least so far, no substitute to someone doing the work of sorting and filtering all of the potential projects and making big and risky bets.

Just set aside those last two paragraphs as they play no role in the rest of this post. I just included them for resonance, if you will.

So, whaling was a risky business and so is venture capital. It is thus no accident that they share a similar financing structure. That’s the sort of thing that would have appealed to one of my undergraduate professors, Arthur Stinchcombe, who devoted a book to the relationship between the structure of organizations and the uncertainties they faced, Information and Organizations (1990).

The Overall Value of Whaling as an Investment

What about profitability? What about ROI? A recent article by Barbara L. Coffee in the International Journal of Maritime History sheds some light on that. Her article contains this interesting paragraph:

Historically, whaling was presented as a prosperous industry and one that the country (Americans) should support. For instance, Daniel Webster, on 3 May 1828, gave a Senate speech supporting the construction of a breakwater at Nantucket: ‘There is a population of eight or nine thousand persons living here on the sea, adding largely every year to the amount of national wealth by the boldest and most preserving industry.’ Whaling was also presented positively in the business magazines of the time:

The importance of this traffic, not only in its profits, which have, perhaps, been greater than those of any other single object of our national enterprise, the capital which is invested in its expeditions, embracing nearly one tenth part of the tonnage of the country, the importance of the moral interests which it involves, comprising the conditions of that large and valuable class of seamen who are its active agents, and the circumstances bordering on the sublime which attend its hazardous expeditions, all render it an interesting subject to our commercial and mercantile population.

In its importance as augmenting the wealth of the country, it is not equaled by any other species of traffic, and presents a marked example of productive labor.

... and the luxurious edifices which adorn many of the cities, attest the enterprise of those who are engaged in the traffic and the success of their labors.

In his speech to Congress on 1 May 1844, Grinnell stated: ‘Commercial History furnishes no account of any parallel; our ships now outnumber those of all other nations combined, and the proceeds of its enterprise are in proportion, and diffused to every part of our country.’

Coffee’s work, achieved by benefit of comprehensive hindsight not available to investors in the 19th century, shows quite a different picture. She concludes:

The nineteenth-century American whaling industry is steeped in tales of great wealth, tragic losses and considerable fortitude. While rich in stories of voyages good and bad, the whaling lore is poor in the accounting of the financial stakes and returns of the thousands of voyages that took place. The goal of this article was to establish the financial facts behind the legends. To accomplish this, the financial returns of the owners of nineteenth-century whaling vessels were reviewed using materials that have been born digital, sources that have been digitized and those still only in print. [...]

The current article takes the financial appraisals of Davis et al a step further by expanding the number of voyages analysed, and by reviewing more years and more harbours. Its main finding relates to risk premiums, which may be defined as the additional return to compensate investors for tolerating risk greater than a risk-free asset. During the nineteenth century, US government bonds, a risk-free asset, returned an average of 4.6%; whaling, a risky asset, returned a mean of 4.7%. This shows 0.1% as the risk premium for whaling over US government bonds.

Think about that. Coffee examined 11,257 voyages undertaken from 1800 to 1899 and found that, for all the risk it entailed, investing in whaling ventures was only slightly better, a 0.1% risk premium, than investing in US government bonds. I wonder what the risk premium is for the venture capital business taken as a whole?

People in the nineteenth century would not have had all that data, along with the analysis, available to them when considering where to invest their money. Note that we’re not just talking about wealthy people, people who can afford to lose some money. Coffee quotes a remark in Moby Dick (Chapter 16: The Ship), “People in Nantucket invest their money in whaling vessels, the same way that you do yours in approved state stocks bringing in good interest.” Their decisions may well have been biased by anecdotes about successful voyages, while forgetting those voyages that lost money or worse, were lost at sea. These days investors in tech ventures are no doubt attracted by stories of unicorns, but at least the law limits investment in venture funds to those who can prove they have the means to do so. You and I are safe from losing our money in that particular way.