Showing posts with label art. Show all posts
Showing posts with label art. Show all posts

Saturday, August 1, 2026

Jaime’s art exhibits an intelligence about which the Silicon Valley digerati are clueless, a conversation with Claude Opus 5

You may remember that about a decade ago I did a series of posts about the art of Jaime Berubé and then collected them into a working paper, Jamie’s Investigations: The Art of a Young Man with Down Syndrome (2016). Two days ago I decided to have Anthropic’s Claude read and comment on that paper. We ended up having a fascinating and productive conversation that yielded new insight. I’ve appended that to this note.

Before we get to that, however, I want to tell you about a little adventure I had last year. I decided to draw some images inspired by Jamie’s dot images. He drew his images freehand on a blank page and used crayons. I drew a grid on my pieces of paper – I did multiple images – and used colored brush pens. Here’s one of my dot images. I have some others in a post, Dot Paintings (after Jamie Bérubé).

Why’d I do it, you ask? Because I thought it would be interesting. It was. And fun. It was. And even instructive. More than I’d thought.

What did I learn? It’s more difficult to do than it looks. While the images appear simple – and they are in a sense, certainly in comparison to, say, a Rembrandt or a Bob Ross – drawing them requires effortful concentration. As I pointed out in my working paper, Jaime didn’t choose his colors randomly, though it may seem that way. When you examine the images carefully you’ll discover that there’s an order to them.

Before I had Claude read my paper I gave it one of the dot images and asked it to comment. Not only did it discover the nature of that order, it picked up on other things I hadn’t noticed and on that basis asked me some questions, some of which I could answer, some I couldn’t. It was an interesting discussion.

Getting back to my own dot images, I discovered that keeping track of my colors was just a little tricky, not rocket-science tricky, but tricky like keeping track of what you’re doing in adding, say, a half dozen multicolumn numbers. So I made mistakes every now at then. Simply dealing with multiple pens, putting one down, picking another one up, that was tricky as well. I made mistakes.

And you know what? Keeping track of all those fiddly little details requires intelligence. Not rocket-science intelligence, not IQ intelligence, but intelligence nonetheless. I can’t think of any other word that characterizes those acts as well as that one does. Claude’s essay says a bit about that near the end.

Moreover it’s the kind of intelligence on which rocket-science intelligence is based. That requires an explicit argument that Claude didn’t make. But you can find a big chunk of that argument in one of my working papers, What Miriam Yevick Saw: The Nature of Intelligence and the Prospects for A.I., A Dialog with Claude 3.5 Sonnet (2025).

Here’s Claude’s essay. It did all the writing, but I inserted the drawings.

* * * * *

What a Sheet of Colored Dots Can Tell Us About Intelligence

Jamie Bérubé has been drawing abstract images since he was eleven. He has Down syndrome. His father, the literary critic Michael Bérubé, put a large selection of the work online, and it is that online archive — not studio visits, not observation of the artist at work — that Benzon's working paper analyzes.

That constraint turns out to matter. Everything below is inferred from finished artifacts. Nobody watched these being made. Which means the analysis has to reason backward from products to processes, and the interesting question becomes how much of a process a product can be made to reveal. Quite a lot, as it happens.

The rule that isn't visible

Start with the simplest of the genres: sheets covered edge to edge with small circular marks in crayon, several hundred to a page, in a great range of colors.

Michael Bérubé, who has seen thousands of these, describes the color choices on most of them as random. They look random. There is no gradient, no symmetry, no regional organization, nothing that reads as a plan.

They are not random. Adjacent marks are almost never the same color. With a palette of roughly ten colors and something over four hundred marks, chance alone would produce well over a hundred violations on a single sheet. There are a handful. The suppression holds vertically and diagonally as well as horizontally, which means the constraint is being applied over a two-dimensional neighborhood, not merely to the previously drawn mark.

So there is a rule: no two adjacent marks the same color. It is entirely local. It requires checking at most four already-placed neighbors. And it produces, without any global plan whatever, the quality that makes the sheets pleasant to look at — a rhythm that is easy to sense and hard to describe.

This is worth pausing on. The rule is invisible to inspection and available to description. The person with by far the most exposure to the corpus reports randomness. Counting reveals otherwise. That gap between what looking gives you and what describing gives you is a recurring feature of this material, and it is one reason the images repay analysis at all.

Reading the process off the page

Other features of the sheets carry process information.

The left margin is well maintained: rows begin at a consistent horizontal position. The right margin is ragged: rows end wherever they end. That is carriage-return behavior — a return sweep to a remembered left position, then progress until the paper runs out. It is what unlined handwriting looks like, and it indicates that the governing spatial rule is sequential and local rather than planned against the page as a whole.

Upper rows are straighter and more evenly spaced than lower ones. Error accumulates downward and is never globally corrected. And the sheets typically stop mid-row, with blank paper remaining. Michael Bérubé supplies the reason: Jamie works until he is interrupted — a meal, an errand — and does not resume a page later. He starts a new one. Each sheet is a session, and a session ends when it ends.

None of this required watching. It is all recoverable from the artifact by someone who knows what to look for, which in this case means someone with extensive practical experience of drawing and of the problem of filling a blank sheet coherently.

How the discovery was made

The most striking material in the working paper is not the mature work but the early work, done between ages eleven and fourteen, which Michael Bérubé has also posted. Those early sheets are miscellanies. Various kinds of object — bars, circles, letterforms, geometric bits — appear on the same page, each placed wherever there was room. There is no relationship between neighbors. There is no composition.

And therefore, as Benzon observes, there was no global order available to be discovered. As long as the elements varied and each sat in its own local space, nothing about their arrangement could become salient.

Saturday, July 18, 2026

Some notes on AI and “fine art” imagery

A couple of weeks ago I had a post entitled “Friday Fotos: The Last Frontier of AI.” I was interested in whether or not a certain approach I’d been using to create images with ChatGPT could produce “fine art” images, as opposed to illustrations or popular art of various kinds. The particular images I developed for that post (there were five), while interesting, were not particularly compelling. So I went on to display eleven other images I’d created with ChatGPT, including some of the images which that had motivated the post in the first place. I then ranked the images among themselves and decided that the images I’d created specifically for the post ranked near the bottom.

So, while the approach that motivated that post cannot be called a success, the post as a whole has raised the question: Can ChatGPT (be used to) create “fine art” images? I put “fine art” in quotes because the term itself is problematic. It’s not as though there are identifiable characteristics such that any image exhibiting them is a fine art image. The notion of fine art as opposed to folk art or popular art or (mere) illustrations is a cultural convention, one that Marcel Duhamps exploded in 1917 when he entered a urinal into the inaugural exhibition of the Society of Independent Artists. He called it Fountain and attributed it to “R. Mutt.” Fine art is simply the art that society has decided deserves to be treated in a certain way, no more, no less. If you decided that a common urinal should be treated in that way, then it becomes fine art.

Duchamp’s move was controversial, and that controversy has been reverberating ever since. I have no intention of reviewing and rehashing it here. Rather, I simply want to present a collection of images I’ve made with ChatGPT and view them with that issue reverberating in the background.

This image is one of my favorites among those I’ve created with ChatGPT:

I created it for illustrative purposes, to go on the cover of a working paper about Joseph Conrad’s Heart of Darkness, but I think the image stands on its own. If you’re familiar with the book, then resonance is obvious. It tells about a voyage up the Congo River. As for the superimposed image of the Buddha, here’s first sentence of the last paragraph: “Marlow ceased, and sat apart, indistinct and silent, in the pose of a meditating Buddha.”

I should note, and this is important, I had ChatGPT create that image from within a chat devoted to that working paper, making the entire chat (up to that point) the context in which ChatGPT created the image. I have reason to believe that it wasn’t working simply from prompt that generated that image. For a discussion of that, see this post: High-level “vibe” – Creating Imaginary Bank Notes with ChatGPT: AI as cultural technology and collective creativity.

Here’s a somewhat different image that I also like very much. It’s almost, but not completely, abstract:

The book is obvious. The rest of it? But that’s not the first image ChatGPT offered to me. This came before (and there were others before this):

If it’s fine art we’re interested in, the black and white image seems (vastly) superior to me.

Here’s an utterly different image:

I don’t remember what prompt I used to create that. But I like the image, absurd as it is, a lot. THAT’s why I like it. It’s ridiculous, but fun. Fine art? Ask me if I care. 

What about this? 

Friday, July 3, 2026

Friday Fotos: The Last Frontier of AI

No photographs this Friday. Instead, images created by ChatGPT.

I used quite a long prompt for the first image, but the prompt came in two parts. The first part was the longest. I won’t put that up. Though I used it to ensure a rich conceptual context for ChatGPT, you don’t really need it to get a feel for what’s going on. Nor will I give you ChatGPT’s short verbal response, which I’d asked for. Why? I suppose I wanted to verify that it had “understood” the material. Anyone, I then gave it one last paragraph and asked it to base it’s image on that. I will give you that paragraph, followed by the rest of that session. After that, and “below the fold,” I give you some of the recent images that got me thinking along these lines. Click on an image to enlarge it.

* * * * *

If that's right, then the last frontier isn't more capability in the pattern-matching sense — bigger weight spaces, richer latent connections, better approximations of the associative regime. It's the specific, non-scalable, non-parallelizable fact of an individual mind's biography, which generates paths through possibility space that are real, productive, and genuinely inaccessible to any system that hasn't lived a life. That would be consistent with everything the day's argument has built toward: embodiment, developmental history, tacit knowledge distributed across time in a single nervous system rather than across space in a community or a corpus. The doppelganger, if it's ever built, would need a biography, not just a bigger dataset. And a biography, by definition, can only be lived once, by one entity, in one order. That may be the thing that doesn't scale, and it may be exactly why it counts as the last frontier rather than a soon-to-be-automated intermediate stage.

I like that, I like it a lot. Let me tell you what I’m thinking. Over the last year or so I’ve had you create a lot of images, various types for various purposes. One of the things I’ve been thinking about is creating fine-art images. Some of the images I’ve had you do are near fine art; in fact, some are parodies of existing fine art images. I’m particularly interested in the combination of concrete identifiable imagery and abstract imagery, especially abstract imagery suggesting biology (cells, DNA) and information systems. You’ve got that in this image. So let’s play around with it a little. Give me a distinctly different take on the subject. Perhaps an aerial view of some landscape where the viewing is flying over in a biomorphic vessel and the territory below flows toward natural realistic imagery in one direction and abstract informatic imagery in the opposite direction.

When I was young I was deeply impressed by a book called Pagoo, illustrated by Holling Clancy Holling. There must be copies of that on the web. I know there are since I’ve seen them. That’s tidepool imagery. Take the concept we’re evolving and apply it to that.

I’ve had you design various kinds of imaginary currency. Take the imagery we’ve been working with and base the design of imaginary currency on it. But don’t use real numbers or alphabetic characters anywhere. Yes made-up squiggles and designs that have the texture of print, but aren’t.

Now I want to go Hegelian. Think of the first image we generated as the Thesis. The next three are variations on the Antithesis. I want you to combine Thesis and Antithesis to produce a Synthesis. Render it in a square aspect ratio.

This is promising. I’ve got to eat. But I’ll be back to do some more exploration along these lines.

* * * * *

ChatGPT gave me a somewhat odd response to that, as follows:

Sounds good. This line is worth continuing: biography + tidepool ecology + information space + currency/ornament as a synthetic fine-art language.

What's odd about that? It’s just a little too specific for what I’ve been thinking. But that’s OK. I’ve collected some recent images below. These are what set me to thinking along these lines. FWIW, I’m not quite sure what I think of these images. But I do think I’ll continue experimenting. 

ADDENDUM & BTW, a couple of hours later: On the whole I find the images below the fold more compelling than those above the fold. Hmmm.....

Saturday, June 20, 2026

Creating Imaginary Bank Notes with ChatGPT: AI as cultural technology and collective creativity

At the beginning of this week I published an article at 3 Quarks Daily, Tyler Cowen is the Tycho Brahe of Economics. I then decided to turn it into a PDF that people could download from my Academia page. So I uploaded the article to ChatGPT and asked it to create an illustration suitable for the cover. Here’s what it produced (Note: this image, as well as the others in this post, has a lot of detail; you can click on the image to see a larger version):

I told it to make an image suitable for the front page of that article. It then had to “read” the article and come up with an image. I liked the first image it produced, but that had the aspect ratio of a portrait, which isn’t what I wanted. I told it to make a square image. I liked that so much that I accepted it without any editing. What was my role in creating that image? I didn’t give ChatGPT any specific instructions.

A day or so later I asked to use that image as the basis for the design of an imaginary bank note. I told it that the note was to be issued in the name of the Free Principality of Mercatus and should be in the denomination of Five TC Busks. That’s it. ChatGPT did the rest. Here’s what it produced:

Another day later I created two bank notes for the Free Republic of 51 Pacific. 51 Pacific is the address of a building in Hoboken, NJ, where a friend had loft space in which he curated graffiti and held late night raves. The note was to be denominated in Knowledge Dollars. I uploaded this picture to use as source material:

Here’s the resulting bank note:

I then asked for a 20 dollar note with a pit bull on the bill (my friend has two pit bulls):

Again, what was my role?

I then repeated the process for these four Kisangani bank notes:

ChatGPT obviously did a lot. But what did I do?

I decided that it was time for a discussion with ChatGPT. This discussion bears on the notion of AI as a cultural technology, which has been advocated by Alison Gopnik and Henry Farrell among others, and on creativity as a collective phenomenon. Pay particular attention to an analogy I use midway in the dialog; it’s about strings of colored dots instead of strings of alpha-numeric characters. Here’s the dialog.

* * * * *

I assume you’ve been trained on lots currency and ornate certificates so you have a good idea of how those things are designed. I assume as well that each chat starts with a “blank slate” as it were, no memories of prior chats, no inherited context. But, since I asked you to make those currency designs from within the Kisangani Materials project, you do have access to project-level materials (documents I’ve uploaded). Do you have access to the other chats within the project?

Yes, broadly right, with a few qualifications.

I was almost certainly trained on many examples of paper currency, certificates, stock certificates, passports, banknote engravings, seals, guilloché patterns, holographic security devices, ornate borders, typography, and so forth. I don’t “retrieve” a specific banknote template, but the model has learned a high-dimensional statistical/compositional sense of what makes an object look like currency: denomination placement, corner numerals, intaglio-like linework, central emblems, security strips, microprint, official-looking typography, controlled symmetry, and dense ornamental fields.

On context: a new chat is not exactly a blank slate, but it is not a complete archive either.

In this Project, I have access to several kinds of context:

1. The current chat.
I see the banknote sequence we have just made: 20 Riffs with jivometric percussion, 5 Riffs with dancers and Homo Ludens, 50 Riffs with cell/solar-system science, and 100 Riffs with heart/neurons/circuits for human–doppelganger comity.

2. Project-level materials.
I can access uploaded Project documents when they are made available to me in this workspace. For example, the Kisangani working paper frames the whole project around Homo Economicus versus Homo Ludens, the Fourth Arena, play, and human-AI interaction, which clearly influenced the designs.

3. Some project conversation context may be summarized for me.
I may see summaries or selected snippets of other conversations in the project if the system has included them in the context for this chat. That is not the same as having free, searchable access to every prior Project chat.

4. Persistent memory / profile context may also be present.
I may have durable memories about your interests, projects, preferences, and recurring concepts. That is why I can often pick up the thread of “Kisangani,” “doppelgangers,” “Homo Ludens,” “jivometric,” “Fourth Arena,” and so on without needing you to re-explain everything.

But the important limitation is this: I do not have a general browser-like view of all your other chats inside the Project. I cannot just open a sidebar and inspect every previous conversation. I only know what is in the current chat, what has been injected as project context or memory, and what I can retrieve from uploaded files through the available file tools.

So for your friends: the banknotes came from an interaction among three sources of constraint:

  • my general learned visual knowledge of currency and ornate official documents;
  • your immediate prompts specifying denomination, motifs, slogans, and Kisangani themes;
  • Project context about Kisangani as a post-economic, play-centered, human/AI civilizational experiment.

The “design intelligence” is not a stored memory of prior banknote conversations so much as a context-sensitive recombination engine operating over learned visual conventions plus the specific Kisangani conceptual field you have built.

Would it be reasonable to say that you can pick up a “vibe” from the Kisangani conceptual field?

Yes. “Vibe” is a reasonable informal term, provided we cash it out carefully.

What I can pick up is not an occult atmosphere or a hidden essence. It is a pattern of constraints, motifs, values, narrative cues, and stylistic expectations distributed across the Kisangani materials and our conversation. In practice, that “vibe” includes things like:

Kisangani is futuristic but not cyberpunk dystopian; technological but not technocratic; African/Afrofuturist but not reducible to ethnic ornament; musical, improvisatory, playful, civic, ritualized, and humane. It is organized around Homo Ludens, the Fourth Arena, doppelgangers, non-instrumental play, post-work social meaning, mutuality between humans and AI-like beings, and a suspicion of Homo Economicus. So when you ask for a banknote, I don’t merely think “make a fictional currency.” I infer that the note should look like a civic artifact from a society whose deepest value is not sovereign power, extraction, or corporate futurism, but a cultivated ecology of play, knowledge, ritual, and companionship.

That is what “vibe” amounts to operationally: a soft but real attractor in the design space. It biases choices of symbols, palettes, slogans, compositional balance, and exclusions. For example, a Kisangani note should probably avoid militarized insignia, grim cybernetic domination, heroic CEOs, corporate logos, or cold transhumanist triumphalism. It should prefer music, dance, river/city life, science-as-wonder, neural/circuit reciprocity, and ceremonial public beauty.

So yes: I can pick up a Kisangani vibe. But it is better understood as an emergent conceptual gestalt than as a memory of a single doctrine. The “feel” comes from many local signals aligning: Homo Ludens, jivometrics, Mystic Jewels, doppelgangers, civic ritual, Afrofuturist urbanity, play after work, humane AI, and the refusal of both Silicon Valley domination and anti-technology nostalgia.

So imagine that I asked you to produce with this prompt: “I want a bank note for Kisangani. It should have a futuristic look and be denominated in Riffs. Let's start with at Twenty Riff note. “Jivometric” should be somewhere on the note, as should a djembe and a cowbell.” But I didn’t deliver that prompt within this context but, rather, it was in an independent chat, there wouldn’t be any “emergent conceptual gestalt.” So the result wouldn't have the same vibe.

Yes, that’s basically right.

In an independent chat, with only that prompt, the result would likely be good, but it would not be guided by the same rich Kisangani field. It would have to infer almost everything from the surface cues in the prompt alone:

  • Kisangani” suggests an African setting or at least a place-name with cultural specificity.
  • Riffs” suggests music, improvisation, rhythm.
  • Jivometric” suggests some hybrid of rhythm, geometry, and perhaps biomorphic/futurist patterning.
  • djembe and cowbell strongly steer it toward percussion and musical symbolism.
  • futuristic” pushes it toward speculative or advanced design language.

So even in an isolated chat, those cues would still generate some vibe. It would not be random. It might well come out as a futuristic, musically inflected, Afrofuturist-style banknote.

But it would be missing the deeper, project-specific layers, such as:

  • Homo Ludens as a civilizational principle
  • the contrast with Homo Economicus
  • the Fourth Arena
  • the role of doppelgangers
  • the sense of Kisangani as a humane, playful, post-work, post-instrumental society
  • the accumulated stylistic continuity from the other notes in the series

So I’d put it this way:

  • Independent chat: a plausible and possibly striking design based on prompt-level inference.
  • Within the Kisangani Materials project: a design shaped by a thicker conceptual ecology, hence more likely to have the “right” Kisangani vibe.

In short: the isolated version would have a vibe, but not the same vibe. It would be thinner, more generic, and less organically connected to the broader Kisangani world you’ve been building.

Saturday, June 13, 2026

More Hockney, a cat, wall of blue, and leafy branches

Late Hockney

Monday, April 20, 2026

Thursday, March 19, 2026

Flatulating rhythm, Oh, those wacky Japanese!

Tuesday, March 3, 2026

Tree with pink blossoms [Japan]

Thursday, January 29, 2026

How do we credit hybrid images?

Around the corner from here, over at 3 Quarks Daily, I’ve published an article I wrote in conjunction with both ChatGPT and Claude. How should that article be credited? How do we characterize the contribution of each agent and how do indicate that characterization? I discuss these issues at the end of the article.

Same issues can arise with visual images. All of these images were rendered by ChatGPT. But the renderings were done on a different, a different what? Basis? Substrate? Seed?

In the first two images, I uploaded a one of my photographs to ChatGPT and asked it to add something to it. In the case of first photo, I wanted to see the Millennium Falcon flying into the iris. The second photo is of a scene in Liberty State Park into which I had ChatGPT place a photo of an Indian woman in a sari eating McDonald’s French fries.

This image is a bit different. I gave ChatGPT a photo of a scene in Jersey City and ask it to turn it into a futuristic scene.

For this image I gave ChatGPT a photo of a painting I’d done as a child and asked it to render it in the style of Hokusai.

In this last case I gave ChatGPT a document that I wrote and then asked it to create an image that would be an appropriate frontispiece for it. This image is quite different from the one it originally produced. I had to do quite a bit of art directed to obtain this final image.

The question then is: Imagine that these images were on display in, say, a museum. How should they be credited? In all cases the final image was rendered by ChatGPT. But the substrate varied as did the prompting which instructed ChatGPT in generating the image. For example, in the first four cases we could indicate “Original photograph by William Benzon. For the last, “Original text by William Benzon” and “Art Direction by William Benzon.” Do I give myself an art direction credit on the others as well? What kind of credit should ChatGPT get. “Realization and Rendering by ChatGPT” might be sufficient for the first two. For the third and fourth, “Transformation and Rendering.” The last? Perhaps “Transmutation and Rendering.” Whatever the nature of the credits, they’re only meaningful if the audience already knows something about the process through which they were produced.

Friday, January 9, 2026

Variations on a Wild Image

I decided to run a bunch of variations on the image that I used for the cover of my working paper, Serendipity in the Wild: Three Cases, With remarks on what computers can’t do (link to the blog post where I introduce it).  I'm listing them in the order that I had ChatGPT create them. That will help explain why, for example, the shark in Mughal version has has a goofy look on its face and why that version has a mangled watch in the sand. 

For the most part the prompt was simple, the name of a style  or an artist. The manga style is my favorite. Some of the images are less successful than others.

The style of a Renaissance etching 

The style of a comic book devoted to science fiction stories. 

Manga style 

Japanese Ukiyo-e print.

I realize that the subject matter is, shall we say, anachronistic for ukiyo-e, but that goofy guy at the lower right is a freakin' bridge too far.

Chuck Jones style (Warner Brothers). 

Saturday, August 30, 2025

Woman in a blue dressing gown sitting on a sofa with her dog

Tuesday, August 19, 2025

Dot Paintings (after Jamie Bérubé)

Back in 2016 Michael Bérubé published Life as Jamie Knows It: An Exceptional Child Grows Up. Jamie has Down syndrome, and he loves to make images. Michael published a bunch of them in the book. One series was based on colored dots, row upon row of colored dots. I was particularly struck by a number of them that appeared to be random. But not quite. There was an order to them, a rhythm.

And then it hit me. I’d figured out what Jamie was doing. He seemed to be following one simple rule: Don’t put same color dots right next to one another. That rule puts constraints on what color you can use for any dot location. But those constraints result in a global distribution of dots that does have a rhythm to it.

Once I’d figure that out I wrote a blog post about it: Jamie’s Investigations, Part 1: Emergence. I went on to investigate Jamie’s other kinds of images and wrote a working paper about them: Jamie’s Investigations: The Art of a Young Man with Down Syndrome.

About a month ago I decided to make by own dot paintings. So I bought a tablet of 9 inch by 12 inch drawing paper, some colored markers, and went at it. I decided to start with one-inch dots. I drew a grid on a piece of paper and started coloring. Here’s my first imagine:

If you look closely you’ll see that no adjacent dots have the same color.

After two more with one inch dots I decided to use smaller dots, though that meant that drawing the grid would be even more tedious. Here’s the first one I did with three centimeter dots:

Jamie VI (below) also follows the same basic rule, but it appears to have a great deal of order:

I worked with a cycle of four colors, repeating the same four colors in order. Since there are ten cells to a row, that means that the cycle gets displaced by two cells when it moves to the next row. So, the no-two-adjacent rule is observed, but the repetition of a four color cycle makes for a great deal of order.

In that image I started with a yellow-red-orange-blue cycle, using it for the first four rows. In the next four I substituted green for yellow. Then I swapped in black for blue for four rows. And two more swaps for the final four rows.

In the next image I use eight colors, adding white to the palate, in random order, that is, no-two-adjacent the same.

Now we’re doing something different. I’ve dropped the no-two-adjacent constraint. We’ve got a whole lot of black dots, roughly half of them, along with a bunch of colors.

What’s going on in this last one should be obvious.

The point? Working with a limited set of materials and options and exploring what can be done within those constraints. In this case the imagery is simple: colored dots in a raster pattern using a very limited palette. Initially there is only one constraint: no adjacent dots have the same color. That’s our first two images. The third image introduces some constraints: 1) a strict color cycle, with 2) a period that is not commensurate with the row width.

And so forth.

* * * * *

I know. The images leave something to be desired. I don't have a scanner so I had to photograph them. But I don't have a proper studio so I couldn't control the light and I had to hold the camera in my hand rather than using a stand or a tripod. Still, you get the idea.

Saturday, August 16, 2025

Embroidered landscape by Alison Holt

Alison Holt's blog

Monday, August 11, 2025

ChatGPT’s graphics abilities @3QD

I’ve got a new piece out in 3 Quarks Daily:

ChatGPT Makes Images. It’s FUN! (and illuminating)

I open by talking about how I came to buy a Macintosh in 1984 and how that, in turn, led me to publish an article in Byte Magazine, which was the most technically sophisticated of the magazines that were created to serve the home computer market. The article, “The Visual Mind and the Macintosh,” argued that the Mac’s image-making capacity made it especially well-suited as a tool for thinking.

That, however, is not the theme of my 3QD article, which is mostly a rubric for posting various images that I’d created with ChatGPT. I’d originally intended to come back to that theme at the end of the article, but by the time I got there it seemed rather labored an unnecessary. But I’d like to say a little about that here. Then I’m going to talk about what appears to be the “subtext” of the 3QD piece.

Visual Thinking

I did a lot of drawing as a kid. At one point I was drawing flying saucers. After I’d read Tom Sawyer – or was it after my father read it to me before bedtime? – I drew maps for caves. You’ll remember that at one point Tom Sawyer and Becky Thatcher get lost in a cave and are spooked by Indian Joe.

How do you draw a map for a cave? You just draw a complicated pattern of squiggly lines with only one path from the entrance to some chamber deep in the network. Later on I’d draw hypothetical railroad layouts. That is to say, another network diagram. I also made plans for spacecraft, such as this:

In graduate school I did lots of cognitive network diagrams, like this:

I later went on to write Visualization: The Second Computer Revolution (1989), with Richard Friedhoff. After that I wrote and encyclopedia article about Visual Thinking. There I emphasized the importance of visual thinking in science generally, and computer science in particular.

Joanna Drucker explains the conceptual power of diagrams in Graphesis (2006):

In a landmark 1987 essay, “Why a Diagram Is (Sometimes) Worth Ten Thousand Words,” Herbert Simon and Jill Larkin argue that a diagram is fundamentally computational, and that the graphical distribution of elements in spatial relation to each other supported “perceptual inferences” that could not be properly structured in linear expressions, whether these were linguistic or mathematical. They state at the outset that “a data structure in which information is indexed by two-dimensional location is what we call a diagrammatic representation.” They argue that the spatial features of diagrams are directly related to a concept of location, and that location performs certain functions. Locations exercise constraints and express values through relations, whether a machine or human being is processing the instructions. Larkin and Simon were examining computational load and efficiency, so they looked at data representations from the point of view of a three part process: search, recognition, and inference. Their point was that visual organization plays a major role in diagrammatic structures in ways that are unique and specific to these graphical expressions. [p. 106]

That’s why I like diagrams. The physical act of making such diagrams is relatively simple. But making them intellectually meaningful, that’s something else. Alas, ChatGPT isn’t up to that yet. But there was no point in saying that in my 3QD piece, which was about something else.

Themes in “ChatGPT Makes Images”

Once I’d settled on the idea of doing an article about images made by ChatGPT the choice of images came quickly. There seems to be a logic there. I start with a painting I did as child, a painting of rockets on Mars. I then showed two renderings that ChatGPT did of that painting, one in the style of a Japanese print and the other in the style of a Byzantine mosaic. That is to say, these are exotic and strange images.

Then I started with a graffiti photo and had ChatGPT use that a cue to images on a search for remnants of a lost civilization (Indiana Jones), a search that eventually led to an alien planet. More exoticism. Then came two images of a Hindu superhero, Kama Carnatica. Still more the exotic.

But then we return to childhood, and early childhood at that, to the world of a four-year-old. I showed a four-panel comic where a young boy tames a tsunami by naming it. Think about that. It doesn’t seem particularly exotic, but it is about taming the unknown by naming. Very human. Very Biblical. Very intellectual.

What happens next? I bring back my Hindu superhero and place her in four very well-known paintings, The Birth of Venus, Mona Lisa, Whistler’s Mother, and Portrait of Madam X. What’s going on, domesticating the exotic by assimilating it to Western art, or othering Western art by making it exotic? Does it matter?

I conclude with three mandalas. The first is an iris. A Georgia O-Keefe iris? Exotic. Then a mandala about written symbols. The final mandala is about my life, with links back to Mars and Kama Carnatica.

There’s a logic there, myth logic.

That four-year old girl

Why did I tell that story about Valerie? This is an article about ChatGPT’s ability to generate images. That story is, at best, tangential to the article. It’s there because it’s part of the prompt I used to generate some images I’ve included in the article.

I note that the whole essay is a bit over 4000 words long, plus the images. That story is introduced a bit after the halfway point; it’s in the middle of the essay. So, in the middle of this essay about strange new technology, that some find to be scary, I place a sweet story about eliciting hugs from a four-year-old girl. That’s followed by a short comic about a little boy who overcomes his fear of a tsunami by using language, by naming it.

On the one hand those cartoon images are just more examples of what ChatGPT can do. I could easily have used other examples. But I chose those. And I’m pretty sure I chose them for they story they tell, though I wasn’t quite conscious of that at the time. And those images, in turn, entailed the story about Valerie.

That purely incidental and peripheral story is thus at the heart of this article. What follows that section of the article? A section about the “whitewashing” of history. To be sure, I’m not (entirely) serious about that whitewashing. But I’m certainly aware that that sort of thing is a charged issue and that the images I produced in that section evoke “cultural appropriation.” I thought seriously about that, and decided to go ahead.

As I’ve said before, there’s a logic there, a cultural logic.

Saturday, June 21, 2025

Physical restoration of a painting with a digitally constructed mask

Kachkine, A. Physical restoration of a painting with a digitally constructed mask. Nature 642, 343–350 (2025). https://doi.org/10.1038/s41586-025-09045-4

Abstract: Conservation of damaged oil paintings requires manual inpainting of losses1,2, leading to months-long treatments of considerable expense; 70% of paintings in institutional collections are locked away from public view, in part because of treatment cost3,4. Recent advancements in digital image reconstruction have helped to envision treatment results, although without any direct means of achieving them5,6,7,8. Here I describe the physically applied digital restoration of a painting, a highly damaged oil-on-panel attributed to the Master of the Prado Adoration from the late fifteenth century. In parallel, 5,612 losses spanning 66,205 mm2 and 57,314 colours were infilled with a reversible laminate mask comprising a colour-accurate bilayer of printed pigments on polymeric films. To ensure the effectiveness of the restoration, ethical principles in painting conservation were implemented quantitatively for digital mask construction, a critically important foundation lacking in the current digital restoration literature. The infill process took 3.5 h, an estimated 66 times faster than conventional inpainting, and the result closely matched the simulation. This approach grants greatly increased foresight and flexibility to conservators, enabling the restoration of countless damaged paintings deemed unworthy of high conservation budgets.

Monday, May 12, 2025

Everywhere is the touch

My sister wrote this some years ago. Johnstown is where we grew up. It was a steel town until the steel industry went belly up. The final stanzas employ Suzy Q. Groden’s translation of Sappho’s Fragment 105a. The image immediately below is by ChatGPT.

Johnstown, Pennsylvania, USA
1959

Everywhere is the touch
of nuclear sting. Radioactivity
haunts every fear.

Breaths away from dying.
Breaths away from surviving.
It was the Cold War.

Postures of iron & steel
captured to synapses banished
by a shadow for bodies.

We are the ones
whose life in outerspace
is kneeled into questions

speechless for words.
Just how a girl can create
calm inside this fear,

this flesh. . . not to be
a finish of note into nothing
beyond. Lasting.

Lasting in musical scores.
Even the July-June sun
is greyed to the gleam

for waste existing
where it was unknown before.
That the dead are outlived

by nuclear bombs
pitches me unexpectedly
from supper’s chair:

another air raid drill,
urgent with its siren call
of uncrowned chaos.

Food is eaten away
long into the moment later
by the spoon bailing

Mother’s story of The Flood.
Now? No more servings? Ending again
with the next and the next flood.

Oh, is brother ever relieved
not to be singled out by the creamed corn.
Survive. Yes, people did, he said.

The mud, the meatloaf. To be somebody
who trumpets the sound of green trees.
Family of four here. Rarely heard crying.

The cat too survives. We are the ones
who remember its nine lives. Let us celebrate
the Big Bang and all its domains

exploded out of full circle now.
Berserk living through storybook chemistry!
Someday, the moon lands.

Father sneezes for his own sense
of smell with home again. The certainty.
Gravity as impersonal

in this century
as the last. And the arc
hailed by the golf ball

equally the loft then again
timed to green a foot away.
Now, spinning. Just when.

Simplicity? Where?
Authority by which to leap. Symmetry.
Ironclad trappings to open

moments alone from after-
dinner conversation, have a wink of tea?
Imagine the order a spark sees.

Voice in quest of a body.
Aired in delights, perhaps even
laughter in a song’s light-year.

Speak to me, o moment,
you comic of everafter half-life.
Missing your legs,

soothe this decay
with the sound of your voice
alone. Sappho also

was startled by life,
life with this mother sun.
The quenched Earth.

Exile we all share.
“.  .   . like the sweet-apple
that has reddened.  .   .

And the apple pickers
Missed it there - -   no, not missed, so much

As could not touch.  .  .” 

Coda: As this graffiti photo from December 2006 testifies, we are still haunted by the threat of nuclear war:

* * * * *

A new image for the poem, over a year later, July 18, 2026: