Showing posts with label aesthetics. Show all posts
Showing posts with label aesthetics. Show all posts

Monday, July 20, 2026

A case of visual aesthetics: Why is the monochrome image superior to the color image?

Last Saturday (July 18, 2026) I posted, Some notes on AI and “fine art” imagery, which included these two images, remarking that the first seemed superior to the second:

But I didn’t ask why it is superior. Now I’m asking: Why?

FWIW, ChatGPT produced the first image based on my prompt. I forget the exact working of the prompt, but it was simple. After uploaded a long document about virtual reading I directed it to create an image for the cover of that document. That color image is what it gave me. I liked the form, but found the color distracting, so I asked it to render the image as a Renaissance engraving, resulting in the first image.

To investigate the question I used Photoshop to alter the color version. For this one I reduced the level of saturation:

Here I rendered the image in monochrome blue:

Finally, monochrome grayscale:

All three strike me as being superior to the original color image. The de-saturated version still has color in it, but the color is not so prominent. The monochrome blue has color information, obviously, but only one hue, so there is no contrast between hues. Finally, the grayscale image has no color information at all. Of the two monochrome images, is one better than the other? At the moment I like the blue better than the grayscale, but I’m not sure it matters much.

But I like the engraving rendition better than either of those. Why? It has a much richer texture. Without bothering to has things out, let’s posit that as part of the reason for the superiority of that version over the original color version.

But that doesn’t explain why the other three are superior to the original. What they have in common is less color information. There’s no color information in the grayscale image and the range of color information in the other two is reduced, drastically so in the blue image.

Why should that matter? Let me speculate. The vision system handles brightness differently than it does hue. Luminance is picked up by so-called rods in the retina while color is picked up by cones. To my eye the color information is just there in the color image. It forms no interesting pattern one its own and it doesn’t seem to be interacting with brightness in any way. It adds nothing but color itself.

Let’s speculate a bit farther. Some years ago Mark Changizi published The Vision Revolution (2009). He argued that the biological purpose of color vision is to allow us to read one another’s emotions more accurately based on gradations in skin color. I find his arguments persuasive. Now, while that may be the primary function of color vision, color vision obviously is not confined to how we see other humans. We use it generally in the world.

Now, let’s look at that image. The only recognizable object in it is that book, and it has very little color in it. The rest of it is just spots, lines and swirls that doesn’t look like anything. But maybe it looks a little like veins beneath the skin, not much, but perhaps enough to “invite” the color system to look for emotional resonance, and fail. It’s that invitation-and-failure that makes the color information, not simply irrelevant, but actively distracting.

Do I believe this? Yes and no. It’s speculation. I just made it up. To move from there to belief I’ve have to come up with a way of actually investigating the question, which is more than I’m in a position to do at the moment.

Saturday, July 18, 2026

Some notes on AI and “fine art” imagery

A couple of weeks ago I had a post entitled “Friday Fotos: The Last Frontier of AI.” I was interested in whether or not a certain approach I’d been using to create images with ChatGPT could produce “fine art” images, as opposed to illustrations or popular art of various kinds. The particular images I developed for that post (there were five), while interesting, were not particularly compelling. So I went on to display eleven other images I’d created with ChatGPT, including some of the images which that had motivated the post in the first place. I then ranked the images among themselves and decided that the images I’d created specifically for the post ranked near the bottom.

So, while the approach that motivated that post cannot be called a success, the post as a whole has raised the question: Can ChatGPT (be used to) create “fine art” images? I put “fine art” in quotes because the term itself is problematic. It’s not as though there are identifiable characteristics such that any image exhibiting them is a fine art image. The notion of fine art as opposed to folk art or popular art or (mere) illustrations is a cultural convention, one that Marcel Duhamps exploded in 1917 when he entered a urinal into the inaugural exhibition of the Society of Independent Artists. He called it Fountain and attributed it to “R. Mutt.” Fine art is simply the art that society has decided deserves to be treated in a certain way, no more, no less. If you decided that a common urinal should be treated in that way, then it becomes fine art.

Duchamp’s move was controversial, and that controversy has been reverberating ever since. I have no intention of reviewing and rehashing it here. Rather, I simply want to present a collection of images I’ve made with ChatGPT and view them with that issue reverberating in the background.

This image is one of my favorites among those I’ve created with ChatGPT:

I created it for illustrative purposes, to go on the cover of a working paper about Joseph Conrad’s Heart of Darkness, but I think the image stands on its own. If you’re familiar with the book, then resonance is obvious. It tells about a voyage up the Congo River. As for the superimposed image of the Buddha, here’s first sentence of the last paragraph: “Marlow ceased, and sat apart, indistinct and silent, in the pose of a meditating Buddha.”

I should note, and this is important, I had ChatGPT create that image from within a chat devoted to that working paper, making the entire chat (up to that point) the context in which ChatGPT created the image. I have reason to believe that it wasn’t working simply from prompt that generated that image. For a discussion of that, see this post: High-level “vibe” – Creating Imaginary Bank Notes with ChatGPT: AI as cultural technology and collective creativity.

Here’s a somewhat different image that I also like very much. It’s almost, but not completely, abstract:

The book is obvious. The rest of it? But that’s not the first image ChatGPT offered to me. This came before (and there were others before this):

If it’s fine art we’re interested in, the black and white image seems (vastly) superior to me.

Here’s an utterly different image:

I don’t remember what prompt I used to create that. But I like the image, absurd as it is, a lot. THAT’s why I like it. It’s ridiculous, but fun. Fine art? Ask me if I care. 

What about this? 

Friday, July 3, 2026

Friday Fotos: The Last Frontier of AI

No photographs this Friday. Instead, images created by ChatGPT.

I used quite a long prompt for the first image, but the prompt came in two parts. The first part was the longest. I won’t put that up. Though I used it to ensure a rich conceptual context for ChatGPT, you don’t really need it to get a feel for what’s going on. Nor will I give you ChatGPT’s short verbal response, which I’d asked for. Why? I suppose I wanted to verify that it had “understood” the material. Anyone, I then gave it one last paragraph and asked it to base it’s image on that. I will give you that paragraph, followed by the rest of that session. After that, and “below the fold,” I give you some of the recent images that got me thinking along these lines. Click on an image to enlarge it.

* * * * *

If that's right, then the last frontier isn't more capability in the pattern-matching sense — bigger weight spaces, richer latent connections, better approximations of the associative regime. It's the specific, non-scalable, non-parallelizable fact of an individual mind's biography, which generates paths through possibility space that are real, productive, and genuinely inaccessible to any system that hasn't lived a life. That would be consistent with everything the day's argument has built toward: embodiment, developmental history, tacit knowledge distributed across time in a single nervous system rather than across space in a community or a corpus. The doppelganger, if it's ever built, would need a biography, not just a bigger dataset. And a biography, by definition, can only be lived once, by one entity, in one order. That may be the thing that doesn't scale, and it may be exactly why it counts as the last frontier rather than a soon-to-be-automated intermediate stage.

I like that, I like it a lot. Let me tell you what I’m thinking. Over the last year or so I’ve had you create a lot of images, various types for various purposes. One of the things I’ve been thinking about is creating fine-art images. Some of the images I’ve had you do are near fine art; in fact, some are parodies of existing fine art images. I’m particularly interested in the combination of concrete identifiable imagery and abstract imagery, especially abstract imagery suggesting biology (cells, DNA) and information systems. You’ve got that in this image. So let’s play around with it a little. Give me a distinctly different take on the subject. Perhaps an aerial view of some landscape where the viewing is flying over in a biomorphic vessel and the territory below flows toward natural realistic imagery in one direction and abstract informatic imagery in the opposite direction.

When I was young I was deeply impressed by a book called Pagoo, illustrated by Holling Clancy Holling. There must be copies of that on the web. I know there are since I’ve seen them. That’s tidepool imagery. Take the concept we’re evolving and apply it to that.

I’ve had you design various kinds of imaginary currency. Take the imagery we’ve been working with and base the design of imaginary currency on it. But don’t use real numbers or alphabetic characters anywhere. Yes made-up squiggles and designs that have the texture of print, but aren’t.

Now I want to go Hegelian. Think of the first image we generated as the Thesis. The next three are variations on the Antithesis. I want you to combine Thesis and Antithesis to produce a Synthesis. Render it in a square aspect ratio.

This is promising. I’ve got to eat. But I’ll be back to do some more exploration along these lines.

* * * * *

ChatGPT gave me a somewhat odd response to that, as follows:

Sounds good. This line is worth continuing: biography + tidepool ecology + information space + currency/ornament as a synthetic fine-art language.

What's odd about that? It’s just a little too specific for what I’ve been thinking. But that’s OK. I’ve collected some recent images below. These are what set me to thinking along these lines. FWIW, I’m not quite sure what I think of these images. But I do think I’ll continue experimenting. 

ADDENDUM & BTW, a couple of hours later: On the whole I find the images below the fold more compelling than those above the fold. Hmmm.....

Saturday, June 20, 2026

Creating Imaginary Bank Notes with ChatGPT: AI as cultural technology and collective creativity

At the beginning of this week I published an article at 3 Quarks Daily, Tyler Cowen is the Tycho Brahe of Economics. I then decided to turn it into a PDF that people could download from my Academia page. So I uploaded the article to ChatGPT and asked it to create an illustration suitable for the cover. Here’s what it produced (Note: this image, as well as the others in this post, has a lot of detail; you can click on the image to see a larger version):

I told it to make an image suitable for the front page of that article. It then had to “read” the article and come up with an image. I liked the first image it produced, but that had the aspect ratio of a portrait, which isn’t what I wanted. I told it to make a square image. I liked that so much that I accepted it without any editing. What was my role in creating that image? I didn’t give ChatGPT any specific instructions.

A day or so later I asked to use that image as the basis for the design of an imaginary bank note. I told it that the note was to be issued in the name of the Free Principality of Mercatus and should be in the denomination of Five TC Busks. That’s it. ChatGPT did the rest. Here’s what it produced:

Another day later I created two bank notes for the Free Republic of 51 Pacific. 51 Pacific is the address of a building in Hoboken, NJ, where a friend had loft space in which he curated graffiti and held late night raves. The note was to be denominated in Knowledge Dollars. I uploaded this picture to use as source material:

Here’s the resulting bank note:

I then asked for a 20 dollar note with a pit bull on the bill (my friend has two pit bulls):

Again, what was my role?

I then repeated the process for these four Kisangani bank notes:

ChatGPT obviously did a lot. But what did I do?

I decided that it was time for a discussion with ChatGPT. This discussion bears on the notion of AI as a cultural technology, which has been advocated by Alison Gopnik and Henry Farrell among others, and on creativity as a collective phenomenon. Pay particular attention to an analogy I use midway in the dialog; it’s about strings of colored dots instead of strings of alpha-numeric characters. Here’s the dialog.

* * * * *

I assume you’ve been trained on lots currency and ornate certificates so you have a good idea of how those things are designed. I assume as well that each chat starts with a “blank slate” as it were, no memories of prior chats, no inherited context. But, since I asked you to make those currency designs from within the Kisangani Materials project, you do have access to project-level materials (documents I’ve uploaded). Do you have access to the other chats within the project?

Yes, broadly right, with a few qualifications.

I was almost certainly trained on many examples of paper currency, certificates, stock certificates, passports, banknote engravings, seals, guilloché patterns, holographic security devices, ornate borders, typography, and so forth. I don’t “retrieve” a specific banknote template, but the model has learned a high-dimensional statistical/compositional sense of what makes an object look like currency: denomination placement, corner numerals, intaglio-like linework, central emblems, security strips, microprint, official-looking typography, controlled symmetry, and dense ornamental fields.

On context: a new chat is not exactly a blank slate, but it is not a complete archive either.

In this Project, I have access to several kinds of context:

1. The current chat.
I see the banknote sequence we have just made: 20 Riffs with jivometric percussion, 5 Riffs with dancers and Homo Ludens, 50 Riffs with cell/solar-system science, and 100 Riffs with heart/neurons/circuits for human–doppelganger comity.

2. Project-level materials.
I can access uploaded Project documents when they are made available to me in this workspace. For example, the Kisangani working paper frames the whole project around Homo Economicus versus Homo Ludens, the Fourth Arena, play, and human-AI interaction, which clearly influenced the designs.

3. Some project conversation context may be summarized for me.
I may see summaries or selected snippets of other conversations in the project if the system has included them in the context for this chat. That is not the same as having free, searchable access to every prior Project chat.

4. Persistent memory / profile context may also be present.
I may have durable memories about your interests, projects, preferences, and recurring concepts. That is why I can often pick up the thread of “Kisangani,” “doppelgangers,” “Homo Ludens,” “jivometric,” “Fourth Arena,” and so on without needing you to re-explain everything.

But the important limitation is this: I do not have a general browser-like view of all your other chats inside the Project. I cannot just open a sidebar and inspect every previous conversation. I only know what is in the current chat, what has been injected as project context or memory, and what I can retrieve from uploaded files through the available file tools.

So for your friends: the banknotes came from an interaction among three sources of constraint:

  • my general learned visual knowledge of currency and ornate official documents;
  • your immediate prompts specifying denomination, motifs, slogans, and Kisangani themes;
  • Project context about Kisangani as a post-economic, play-centered, human/AI civilizational experiment.

The “design intelligence” is not a stored memory of prior banknote conversations so much as a context-sensitive recombination engine operating over learned visual conventions plus the specific Kisangani conceptual field you have built.

Would it be reasonable to say that you can pick up a “vibe” from the Kisangani conceptual field?

Yes. “Vibe” is a reasonable informal term, provided we cash it out carefully.

What I can pick up is not an occult atmosphere or a hidden essence. It is a pattern of constraints, motifs, values, narrative cues, and stylistic expectations distributed across the Kisangani materials and our conversation. In practice, that “vibe” includes things like:

Kisangani is futuristic but not cyberpunk dystopian; technological but not technocratic; African/Afrofuturist but not reducible to ethnic ornament; musical, improvisatory, playful, civic, ritualized, and humane. It is organized around Homo Ludens, the Fourth Arena, doppelgangers, non-instrumental play, post-work social meaning, mutuality between humans and AI-like beings, and a suspicion of Homo Economicus. So when you ask for a banknote, I don’t merely think “make a fictional currency.” I infer that the note should look like a civic artifact from a society whose deepest value is not sovereign power, extraction, or corporate futurism, but a cultivated ecology of play, knowledge, ritual, and companionship.

That is what “vibe” amounts to operationally: a soft but real attractor in the design space. It biases choices of symbols, palettes, slogans, compositional balance, and exclusions. For example, a Kisangani note should probably avoid militarized insignia, grim cybernetic domination, heroic CEOs, corporate logos, or cold transhumanist triumphalism. It should prefer music, dance, river/city life, science-as-wonder, neural/circuit reciprocity, and ceremonial public beauty.

So yes: I can pick up a Kisangani vibe. But it is better understood as an emergent conceptual gestalt than as a memory of a single doctrine. The “feel” comes from many local signals aligning: Homo Ludens, jivometrics, Mystic Jewels, doppelgangers, civic ritual, Afrofuturist urbanity, play after work, humane AI, and the refusal of both Silicon Valley domination and anti-technology nostalgia.

So imagine that I asked you to produce with this prompt: “I want a bank note for Kisangani. It should have a futuristic look and be denominated in Riffs. Let's start with at Twenty Riff note. “Jivometric” should be somewhere on the note, as should a djembe and a cowbell.” But I didn’t deliver that prompt within this context but, rather, it was in an independent chat, there wouldn’t be any “emergent conceptual gestalt.” So the result wouldn't have the same vibe.

Yes, that’s basically right.

In an independent chat, with only that prompt, the result would likely be good, but it would not be guided by the same rich Kisangani field. It would have to infer almost everything from the surface cues in the prompt alone:

  • Kisangani” suggests an African setting or at least a place-name with cultural specificity.
  • Riffs” suggests music, improvisation, rhythm.
  • Jivometric” suggests some hybrid of rhythm, geometry, and perhaps biomorphic/futurist patterning.
  • djembe and cowbell strongly steer it toward percussion and musical symbolism.
  • futuristic” pushes it toward speculative or advanced design language.

So even in an isolated chat, those cues would still generate some vibe. It would not be random. It might well come out as a futuristic, musically inflected, Afrofuturist-style banknote.

But it would be missing the deeper, project-specific layers, such as:

  • Homo Ludens as a civilizational principle
  • the contrast with Homo Economicus
  • the Fourth Arena
  • the role of doppelgangers
  • the sense of Kisangani as a humane, playful, post-work, post-instrumental society
  • the accumulated stylistic continuity from the other notes in the series

So I’d put it this way:

  • Independent chat: a plausible and possibly striking design based on prompt-level inference.
  • Within the Kisangani Materials project: a design shaped by a thicker conceptual ecology, hence more likely to have the “right” Kisangani vibe.

In short: the isolated version would have a vibe, but not the same vibe. It would be thinner, more generic, and less organically connected to the broader Kisangani world you’ve been building.

Saturday, May 23, 2026

Making it strange on the tabletop @3QD

I’ve got a new piece at 3 Quarks Daily:

What’s a Tablescape? I’m Glad You Asked.

It’s about photos I’ve taken atop tables where I’m eating a meal, mostly breakfast in recent years. The basic rule is that the bottom of the camera must be resting on the top of the table. As a practical matter, I cannot compose these shots with any care. Why not? Because I can’t get my eye down to the camera where I can look through the view-finder (DSLR) or look at the viewscreen (point-and-shoot).

That is to say, not only is there an element of chance in the process – I don’t quite know what image is being captured – but that the resulting image is from a POV that’s not available to me. It’s a world as observed by a creature whose eyes and between an inch and three inches above the surface on which it is standing. Hence the title of this post, “Making it strange on the tabletop.”

“Making is strange” is an old slogan of the modernist avant-garde, with “making it new” as a variant. It’s a doctrine of aesthetic alienation, though alienation is a positive rather than a negative sense. The idea is to bring you closer to the world, to make you more observant, but giving you perspectives you’ve never had, and hence cannot have become habituated too.

Tuesday, July 8, 2025

What’s a “good” photograph & the camera as maker of patterns

Here’s a good photo, not wonderful by any means, just a decent photo:

This is pretty much the same shot, but it is a bad photo:

Why is it a bad photo? Because the camera moved during the relatively long exposure needed to capture the image in the dim evening light. I was using a hand-held camera (I don’t even own a tripod), so that’s a problem. But, you know what, I actually like that photo, perhaps even as much if not more than the previous photo, the “good” one. Why do I like it? Because of the colors and composition. It’s a pleasing image.

It's bad (only) in the sense that the blur interferes with the photo’s representational function. The photo is supposed to represent something, buildings in the Hudson Yards development in Manhattan. What if you don’t care about that function? Now, if you want to sell photos to magazines, then yes, you have to care about how well the photos represent their subjects. But if you’re not in that business – and I’m not – then that source of badness just disappears. What’s left is a pleasing image.

Let’s play around with that. In the following three photos I use Photoshop’s pixilate filter to alter the image. As the pixels get larger the photos representation function recedes further into the background:

Here’s another series, but I use a different version of the pixilate function.

Notice that this time the representation function is pretty much obliterated in the last image, and very badly degraded in the one before.

Lesson? 

* * * * *

BTW, this is how I discovered ICM (intentional camera movement). I had this image that was blurred from camera movement, but I liked it. So I said to myself, Why don't I do that intentionally. That's when I started experimenting with deliberate camera movement. It's become a regular feature of my repertoire.

Saturday, March 15, 2025

2025 New York Times photo review

I recently entered the 2025 New York Times Portfolio Review, which was free. It’s simple: You submit 10 photos, if you make the cut those photos will be critiqued by four or five judges. I didn’t make the cut, alas.

Why not? I don’t know, they didn’t say. Which is par for the course for these things.

My guess is that it went something like this: They received 2500 applicants for 160 slots. I don’t know the actual procedure, but whatever it was, I suspect that, in effect, they divided the entries into two piles, possibles, and rejects. They then had to choose among the possibles. I’m sure they did that as carefully as they could, but they might as well have done it randomly. The future will tell whether or not they made the right selection. And here’s the tricky part, the selection they made was an initial step in determining that future.

I have no idea who entered, but anyone could do so. I would think that most unknowns would take a deep breath before entering, but surely there were some entrants who had no business entering their photos. Mostly likely all of them ended up in the reject pile. But there were others in the reject pile as well. How large was the reject pile? I don’t know. Maybe 1000, less than half, maybe 2000, considerably more than half. Who knows. However large the pile of possibles, I’m guessing that they were all of roughly the same quality. THAT’s why I said choosing the actual entrants from that group was essentially arbitrary. Reasons were no doubt given, in the minds of the judges, perhaps on score sheets of some kind, and they were “real,” but also arbitrary, ex post facto rationalizations.

Why am I leaving it up to the future to sort things out? I’m not saying that there were no worthwhile differences among the possibles. Not at all. What I’m saying is that we don’t know, at this time, what those differences are.

I have much the same problem with my own photos. Beyond a certain point, I don’t know which ones are better than the others. I can’t tell. I’ve not yet trained myself to make such a discrimination.

Saturday, January 25, 2025

Photoshop’s Generative AI Rocks!

BEFORE:

AFTER:

For my money, the difference between the first image and the second is (just about) as remarkable as any of the verbal wizardry I’ve seen from Claude or ChatGPT. The visual difference between the two images is easy to see. It’s the sort of thing, where, if you didn’t actually have to figure out how to transform the first into the second, you’d think it could be done in the blink of an eye, with the flick of the wrist.

And now you can, almost. With generative AI it really is almost that easy. Almost, but not quite. We’re not yet to the point where one can say, “computer, clean that up for me,” and it’ll know what you want and be able to do it. While I’m tempted to say we’re not far from it now, I don’t really know. The phrase “clean that up” is doing a lot of work in that command. I’m not sure we’re to the point where a so-called agent powered by a ginormous so-called Foundation Model can do that. That agent might have to be personally trained by the use to know what the scope of that phrase is.

The thing is, I’m not sure we need such an agent. What I actually had to do was not difficult and did not take much time, say two, three, five minutes, and this was my first time through. I had to:

  1. pick a tool,
  2. set a parameter on the tool, which took an adjustment to get it right,
  3. manually trace over the area I wanted altered, which took a little skill, but nothing beyond the reach of anyone capable of using Photoshop at all, and finally,
  4. direct Photoshop to execute the operation.

That last took a half-minute to a minute on my machine, which is maybe three years old. It would go much quicker with a new machine. Now, if I had fifty of those to do, in that case, yes, it would be nice to be able to get the job done with a simple command,

Now, if I were a painter with the requisite draftsmanship, I could paint a picture like either of those photos. Neither of them would present a technical problem to such a craftsman. But that craftsman is unlikely to want to execute an image like the first. Why junk up the scene with that intrusive table and what looks like a piece of a bike rack in the background.

Things are a bit different for the photographer. Sure, taking a shot like the second would as easy as the first shot, if those blasted things weren’t in the way. But they were, and that’s a big problem. I would think the problem would have been almost impossible to handle with traditional analog photography. You would have to manually paint over the intrusions. That’s not at all practical. With digital photography things are different. You can easily go in change any pixel you want to. You could, in theory, manually edit the first image so that it comes out like the second. But it would be hellish and time-consuming. There are probably people who can and have done that sort of thing. I hope they get paid and arm and a leg for doing it. But the need for that kind of skill is now all-but-over.

Now, notice that green smudge at the lower right. If I were shooting this for a magazine, I’d probably have to get rid of it. That would be difficult to do and I’m not sure how well the AI would be able to paint the girl’s feet. Perhaps I’ll give it a try some time.

But I’m in no hurry. “Why not?” you ask. “Because I like it there.” “Why, pray tell,” you ask. Because it gives a sense of distance, of space. Whatever it is, probably some kind of plant, it’s between the photographer and the kids. The photographer, that’s me, likes such things. He’ll even deliberately introduce such “defects” into his photos. They’re part of his aesthetic.

In this case, however, I doubt that there was any deliberation. I saw the kids move out of the corner of my eye. So I turned and took the shot. The green blob intruded, which is fine. But so did that table, not so good.

Now, back to the underlying AI tech. As I said up top, the difference between those two images is as remarkable to me as the verbal skills of ChatGPT or Claude. But, and this is very important, you need to understand that, to a first approximation, LLM tech treats language the same way this visual tech treats visual information. LLMs treat language as though it consisted of strings of colored beads. You and I know that those colored beads are in fact letters that spell out words and the spaces and punctuation between words. The AI tech doesn’t “know” that. You and I know that those strings are words, symbols; the AI tech doesn’t know that. As far as it is concerned, you might as well feed it (images of) strings of colored beads.

Remember step 3 above, where I trace over the part of the image I want removed? That’s a prompt. Or rather, that plus the source image constitute a prompt to the AI. It then produces a new image with the changes specified in the prompt. Considered at the appropriate level of abstraction, it’s the same as prompting ChatGPT or Claude to take a sad story as input and return a happy one.

[As Sean Connery said in that movie, “Here endeth the lesson.”]

Monday, January 20, 2025

Claude 3.5 Sonata describes two different photos and approaches the concept of a minimal photo.

I’ve been asking Claude 3.5 to describe photos since December 6, 2024. It’s an exercise. I’m not looking for anything in particular. I just what to see how Claude describes a photo and where the conversations goes. Describing images can be easy, but it can also be tricky difficult. It depends on the nature of the image and what one attempts in the description. Often Claude will end their statement with a question or two. Sometimes I’ll take them up, but sometimes I’ll go in a different direction. It depends.

On the whole, it’s an interesting task, without being terribly demanding. Some conversations are more interesting than others. Some turn out to be very interesting, e.g. Claude 3.5 describes an iris on the street, and then we go off into the urban pastoral, or Claude 3.5 Describes a Big Mac (and then gets lost in the weeds, does it ever! Sheesh!). Until now, I’ve started all conversations with a single photo, but in a few cases I’ve added a photo into the conversations as things moved along, such as that urban pastoral conversation or, more recently, Claude 3.5 Sonata describes a sunrise photo of some weeds, and then we talk. This time I decided to start with two photos.

The idea is to present two very different photos along with a question: In what way are these two photos the same? I had a specific answer in mind: They are what I call minimal photographs, a concept which I define in the course of the conversation. Would Claude figure it out without my help?

It turns out that Claude didn’t do that, which is neither here nor there. I just wanted to find out. I wouldn’t necessarily expected a human interlocutor to come up with that answer either. Some might. There was, after all, an art movement known as “Minimalism.” But I wouldn’t expect everyone to know that.

In any event, the conversation turned out to be quite interesting. Claude told me a thing or three about those photos that I hadn’t noticed. They’re a pretty good art critic.

* * * * *

I want you to describe each of these photos and tell me what they have in common. Why would someone think they are, in some way, the same kind of photo?

These two images are similar in that they both capture an obscured or diffused view of the sun, creating moody, atmospheric compositions.

In Image 1, we see ships silhouetted on the horizon against what appears to be a hazy, golden-brown sky. The sun's light is severely muted by what could be fog, smoke, or atmospheric pollution, creating an eerie, monochromatic seascape.

In Image 2, we see bare tree branches silhouetted against a bright, diffused circle of light - presumably the sun - behind a gray, overcast sky. The intricate pattern of the branches creates a stark contrast against the glowing orb behind them.

What these photos have in common is their artistic approach to capturing an obscured sun. In both cases, the sun is present but veiled - in the first by atmospheric conditions creating a sepia tone, and in the second by cloud cover creating a gray palette. Both photos use silhouettes (ships in one, tree branches in the other) to create contrast against the muted light source.

Both images also evoke a similar mood - there's something slightly melancholic or mysterious about them. They demonstrate how atmospheric conditions can transform the usually bright, clear sun into something more subdued and ethereal. Both photographers have chosen to capture moments where nature has created a natural filter over the sun, allowing us to see its presence in an unusual way.

Interesting. I took the “sepia” one at sunrise on a misty day. I was looking north on the Hudson River. The sun would have been to my right and so not in the field of view, if it could be seen at all through the fog (I don’t remember). I took the other one at mid-day out in the suburbs, and of course I was looking directly at the sun, as I sometimes like to do. And you’re right about the mood.

Ah, I apologize for my misinterpretation of the first image! That's fascinating - what I interpreted as a sun diffused by haze was actually just the natural coloring of dense river fog illuminated by the morning light. Thank you for the correction and context.

This actually makes the images even more interesting as a pair, because they show two very different ways that atmospheric conditions can create similar visual effects. In the Hudson River shot, the fog itself becomes luminous, creating that golden-brown color field that fills the frame, with just enough visibility to make out the ship silhouettes. In the suburban shot, it's the cloud cover that transforms the harsh midday sun into that softer glowing orb.

I'm curious - as someone who likes to photograph the sun directly (which can be quite challenging), what draws you to that kind of shot? There's something particularly striking about how the bare branches frame and segment the diffused light in that second image.

Also, I imagine the actual experience of taking these two photos felt very different - one in the early morning quiet by the river, and the other looking up at a bright midday sky. Yet they both captured similarly contemplative moods in the final images.

On the second issue, shooting into the sun, what draws me to the shot is, at least in part, the challenge. The bright light floods the sensor so that I have to do a bit if finagling to draw out an image in processing. And, yes, in that particular kind of shot I like the way the clouds diffuse the light.

Yikes! 31 thousand photos over 20 years, how do you make sense of it all?

With this photo I defined my style. What's in the center of the image, the focal point? Nothing really. Just space. My style's about space, and light.

That’s what I started doing over the weekend. I started taking photos in the summer of 2004 when I went to Chicago to give a keynote at LACUS (Linguistic Association of Canada and the United States). I saw that the new Millennium Park was opening and decided I needed to photograph it. So a bought a Canon Powershot A75 and off I went.

I had no further plans for the camera, but I had it, so why not? I started walking the neighborhood, taking photos, and got interested in the graffiti. “Why don’t I photograph that,” sez I to myself, “I bet there’s more of that around here.” Was there ever.

I started with graffiti, then moved to the city around it, and before you knew it I was photographing everything. When I moved from Hamilton Park to Layafette, new neighborhood, more photos. Then out to the suburbs, Maplewood, Peter building his bookshelf, then to Hoboken. Yay! I’m three blocks from the Hudson River. The New York skyline, boats on the river, and back to Jersey City and Liberty State Park for the wildflowers. Then the irises and tiger lillies in the 11th St. flower beds in Hoboken. Halloween! The Arts and Music Festival. The Demolition Exhibition in Jersey City. Selfies and shots out the window.

And that’s how I ended up with 31,474 photos on Flickr sorted into 39 albums. [And at least as many more – maybe 3 or 4 times as many – on my computer that I’ve never uploaded.] Some contain a half-dozen or fewer photos, but others contain a thousand or more (4485 for the Hoboken album). Some albums are organized by geographical area, others by subject matter. And many photos are in two, three, or four, or maybe five different albums.

How do I make sense of it all?

Click on the image. You'll get a new window containing the album. Click on "Back to album" at the upper left.

I decided to create a new album, called Sample Photos. I’d go through all my other albums and pick out one or three to put into the Sample Photos album. I figured it would be a 100 or so photos. HAH! It’s up to 389 photos and I’m not done. It’ll probably be 500 or so photos by the time I’m done.

And you know, that’s too many. Too many for me to grasp. Now, you understand, it’s not like I somehow held all those photos in my memory. I’d forgotten most of them. But then, when I see them, I remember that I took it. Sometimes. I did it all, but don’t know I did it all.

And I want to boil it all down so I can somehow grasp it, make sense of it. Not so much in terms of being able to explain it. But simply to feel that I grasp it. I figure maybe I can grasp 50 photos.

But how can I go from 500 photos to 50 photos? Can I do it in one step? Or do I take it down to, say, 200 photos, and then down to 50 from there? I could just go through the samples album and pick every fifth photo and put it into a Samples 2 album. I know, that seems rather arbitrary, and it is. Still, it might be worth a try.

We’ll see.

[All the while I’m thinking, what kind of a mental process is this? That’s half the reason I’m doing it.]

Tuesday, December 31, 2024

GPT in the Classroom, Part 2: Escape to America

Time to return to the classroom. I want to return to the question of authorship. First, I present the interaction that resulted in a poem, “Escape to America.” Once we’ve been through that we can ask: Who’s the author? 

* * * * *

The poem, “Escape to America”

I want you to create a sonnet based on three short texts by Miriam Yevick. Use the first text for the first quatrain, the second text for the second quatrain, and the third text for the third quatrain. Create the final couplet as you will. Call the sonnet: Escape to America. Here are the three texts:

The tanks exuded power. The soldiers cast their vigor into the rows of women. The women’s hot breath floated back to them on the cheers. Girls showered flowers and kisses into the warrior’s path. This was the scent of war! The orderly formations loosened, females leaped from the sidewalk and ran into the mêlée. Soldiers broke rank and grabbed at random. The caravan driven on by the motorized units undulated forward amidst frenzied embraces.

Luckily the trains to Paris were still running normally. We were greatly surprised to encounter a city still believing in the protection offered by the Maginot Line (circumvented by the Nazis the previous week) and oblivious to what had happened and was happening some hundred miles to the North.

The next day we took a walk on Broadway. I saw a striking dress in a shop window: sky blue with bold purple stripes descending diagonally from both shoulders to meet along the centerline of the body. I went into the store and bought the dress, which accentuated my youthful shape, for $ 1. Ready-made!

Here is a sonnet, as requested, based on Miriam Yevick’s three texts:

Escape to America

The tanks exuded power, soldiers’ cheers
Roared as their vigor met the women's cries.
Bouquets of kisses, flowers, flung through tears,
Unordered ranks dissolved in frenzy’s guise.

The train to Paris held us; still we sped
Toward streets naïve, in thrall to the Maginot,
Where brittle hopes ignored the northern dread—
The unseen storm that forced us here to go.

And Broadway’s throng—a different world it seemed.
In sky-blue dress, bold purple stripes askew,
I saw myself transformed, as if I dreamed;
Ready-made youth, America’s own hue.

And now, the past fades dim beyond my sight—
A foreign land, forsaken for this light.

I like what you did a lot, but I want two changes: In the first quatrain, somehow indicate that the soldiers are Nazis. Maybe do it by mentioning the swastika. In the third quatrain, I’m not sure about “America’s own hue.” The second line mentions purple stripes, but purple is not particularly identified with America.

Here is the revised sonnet with your requested changes:

Escape to America

The tanks exuded power, swastikas spread,
Their vigor met the women’s frantic cries.
Bouquets of kisses, flowers, flung through dread,
Unordered ranks dissolved in frenzy’s guise.

The train to Paris held us; still we sped
Toward streets naïve, in thrall to the Maginot,
Where brittle hopes ignored the northern dread—
The unseen storm that forced us here to go.

And Broadway’s throng—a different world it seemed.
In sky-blue dress, bold stripes in proud display,
I saw myself transformed, as if I dreamed;
Ready-made youth, new freedoms on my way.

And now, the past fades dim beyond my sight—
A foreign land, forsaken for this light.

* * * * *

But who is the author?

There’s no doubt that ChatGPT, with the aid of FredTheHeretic (aka FTH), generated the words. In what sense is ChatGPT the author? It couldn’t have done it without FTH, which is a GPT, a user-created device that can be installed in ChatGPT so that it can generate specialized output. FTH contains a small collection of poems by Frederick Turner. Does that make Frederick Turner the author? He didn’t actually write the words, but they’re (more or less) in his style.

What about Miriam Yevick? The poem is about her life, her escape from a Europe in the process of being swallowed up by Nazi Germany. Not only is it about her life, but it’s based on her words, words from her memoire, A Testament for Ariela. Surely Yevick deserves some authorship credit.

But how did FTH obtain those words? They didn’t just leap out of the book and into the computer. I selected the passages from the book – they are not contiguous in the text – and presented them to FTH. When FTH’s first try was a little wonky, I suggested changes. Surely I deserve some credit.

So far we’ve got ChatGPT, FTH, Miriam Yevick, and me. ChatGPT didn’t spontaneously emerge into existence one day when a computer had some CPU cycles to spare. It was created by a team of programmers, engineers, data scientists, and technicians at OpenAI. And FTH was created by Paul Fishwick and his graduate students at the University of Texas at Dallas. How do we credit these people?

This is not a new problem. The motion picture industry has been up against it for years and has evolved a rather elaborate set of conventions for doling out credit, credits negotiated with the various parties, both individuals, and organizations, involved. I don’t intend to propose a solution in this case. But the problem is now here and we’re going to have to deal with it.

One final point: It seems to me that denying “Escape to America” is meaningful because the words were actually produced by a computer, using that as an excuse to assert that it’s not a poem, that’s bone-headed, stupid, and short-sighted. We’ve got intellectual work to do.

Thursday, June 6, 2024

That Shakespeare Thing

I'm bumping this to the top of the queue for several reasons, general principle among them. The opinion I offer in the last half or so requires a proper argument and, while I've written various things in that direction, I haven't gotten around to such an argument. Nor can I see putting it high on the agenda now. Consider it a stake in the ground, or a promissory note.
 
* * * * *
 
It is a truth fervently believed, at least among those who have beliefs about such things, that Shakespeare is the greatest writer the world has ever seen. Without question. Period. End of story. So help me god. Cross my heart and hope to die.

Bollocks!

It’s not that I doubt Shakespeare’s excellence. Of course he’s good. But not that good. No mortal human is, or could possibly be, that good. For THAT good is not about history, it’s about mythology.

And that mythology has got to stop. We can’t treat our literary culture as though it were but an appendage to Shakespeare’s large, various, and excellent output. Debts are owed, certainly. But appendages to, certainly not.

It’s simple: We can’t enter into the 21st century as long as we keep swearing fealty to The Bard, even if we cross our fingers behind our backs while so swearing. The world’s changing, it’s been changing since Shakespeare’s time. The old guy can no longer keep up. It’s time to put him on a raft, and cut the raft free. Let him float out to sea.

Who’s this WE you’re talking about?

Good question. Tricky question. I suppose I could say Harold Bloom and the Bloomistas and be done with it. In fact, that’s what I will say: Bloom and the Bloomistas!

Call it a figure.

Though Harold Bloom is real enough. His admiration for Shakespeare is well known – didn’t he write a fat book explaining how we’re all Shakespeare’s children? And he’s set himself up as the Defender of the Western Literary Canon, the Finger in the Dike that Protects Western Civ from the Sea.

Give the finger a rest. Let the water flow. Life goes on.

What brought this on, you ask?

It’s been a long time coming. At least since late 1989 when The New Republic published a special 75th anniversary issue in which their long-term film critic, Stanley Kauffmann, reflected on film’s history and accomplishments. Toward the middle he had a Big Paragraph:
After we have at the beginnings [of film], what can we say today about the results? What has film accomplished since then? Once, after a meeting in which films were glowingly discussed, a well-known poet challenged me to name one film that was the equal of the greatest work in other arts, the work of Sophocles or Dante or Michelangelo or Bach or Tolstoy. The answer was, is, double. First, there is no such film. There may never be such a film. Second, who is the Sophocles or Bach of the 20th century? Great artists there have certainly been in our time, and I would not blithely equate even the best films I know with the work of Joyce or Picasso. But few would rank Joyce or Picasso or the other masters of our time with the greatest artists in history; and the overwhelming fact is that, arriving in this century that has been stormy even for the oldest arts, film has created hundreds of works that are now part of the cultural legacy of every civilized human being.
He’s got that last clause right: Film HAS created hundreds of works that are now part of the cultural legacy of every civilized human being. The rest of it is pious nonsense, spun from the same thread as Bloomistavision.

That’s what set me on this line of thought. What set me off is Apocalypse Now. I didn’t see it when it first came out. Don’t know exactly why. Perhaps because it was a Hollywood film, so how could it be Really Good? Saw the Redux version; don’t remember what I thought. But then the DVD set came out, Apocalypse Now: The Complete Dossier.

I was stunned. Watching it on my small computer monitor with the add-on speakers – decent sound, but not as good as my stereo system, and certainly not theatrical – I was stunned. Thought I to myself: Shakespeare couldn’t do this.

Shakespeare couldn’t do this.

It’s not simply that he didn’t have the technology, though he didn’t, and that’s not irrelevant. Whatever Coppola was saying, Shakespeare couldn’t say it, because he didn’t know it. He didn’t know it.

Shakespeare didn’t know it. It’s not his fault. Our world isn’t his world. Our world makes its own demands, and Shakespeare’s knowledge cannot be equal to them, no more.

Let him rest.