Showing posts with label information. Show all posts
Showing posts with label information. Show all posts

Tuesday, February 25, 2025

Search as an interface between informatics and economics: How is the distribution of good ideas like the distribution of gold deposits?

The Answer: Both distributions are highly irregular. What that means in the case of gold deposits is pretty obvious: Gold ore is a physical substance that is found in the earth, a huge mass of physical substance. But there is no obvious order to just where you can find deposits of ore. So you have to go looking for them, which is called prospecting.

Ideas, though, are not things. What does it mean to talk about their distribution? Is there some kind of abstract space where ideas exist? If so, how do we map and describe that space.

So, first I’m going to talk a bit about locating ideas in space. Then I’m going to present a conversation I had with Claude 3.5. First I talk about locating ideas in some abstract space, then I present the conversation I had with Claude, which starts with gold and ends with AI.

Ideas in space

Well, think of a library. Libraries contain books and books contain ideas. Books are physical objects and so have locations in physical space, library shelves. So, how are books placed on those shelves? Off hand, there seems to be two principles: 1.) alphabetically by author name, and 2.) according to subject matter. Fiction tends to be organized according to the first principle while non-fiction is organized by the second. This means that novels placed on the same shelf might are likely to be very different in character. (Take a look at this alphabetized list in Wikipedia.) Non-fiction is arranged by subject matter, so books that occupy the same self will be generally about the same thing. But there is a limit to how far that principle takes us. For one thing, any given book can be about many different things. It can’t be on selves with all of them. Still, you get the idea. We have ideas arranged in space, albeit a space limited to three dimensions.

Things get more interesting when we have a collection of electronic documents. In this case, actual physical space is irrelevant. What matters is, well, idea space. Back in the 1960s and 1970s Gerard Salton came up with the idea of representing documents using vectors containing terms in the documents. Thus, each document will occupy a position in some high-dimensional space. Documents located close together in that space will be about similar topics. Salton’s Vector Space Model is one of the foundations of machine learning – something I discuss with Claude 3.5 in this post, Memory in machines and minds, a conversation with Claude 3.5 Sonnet.

OK, so that gets us a way of thinking about ideas as existing in an abstract space. But how do we motivate the idea that some ideas are closer to us than other ideas? Well, imagine that you occupy some position in idea space, a high dimensional space. There are a bunch of ideas you understand very well. We can say they are close to you. Other ideas are utterly unintelligible. By convention we will say that they are far from you. In order to reach one of those distance ideas you are going to have to learn the ideas that are on a path between where you are now and where that idea is.

Yeah, I know, it’s fuzzy and indistinct. But you get the idea. I explore that notion in these working papers: Stagnation and Beyond: Economic growth and the cost of knowledge in a complex world, and, What economic growth and statistical semantics tell us about the structure of the world. The second one is shorter and more directly to the point.

It seems pretty clear to me that the further development of machine learning and A.I. is going to teach us quite a bit about the structure of idea space. Every LLM is a high-dimensional space jam-packed with billions and billions of ideas. We’ve just barely begun to map that space, something I talk about in my recent report, ChatGPT: Exploring the Digital Wilderness, Findings and Prospects, see especially pp. 34-38, 42-44.

* * * * *

That’s enough of that. The rest of this post consists of a dialog I had with Claude 3.5 about the relationship between technology and the search for irregularly distributed things. We start with gold, move on to oil and diamonds, then documents, drug discovery, movies and finally new AI technology. As always, my prompts are bold face while Claude’s replies are plain face.

Gold

How have gold prospecting methods changed from the 19th century through the 20th century?

Let me walk you through the key changes in gold prospecting methods across these two centuries.

In the 19th century, prospecting was largely manual and relied on simple tools and techniques:

Panning was the most iconic method - prospectors would use flat pans to swirl river sediment, letting water wash away lighter materials while heavier gold settled at the bottom. This was supplemented by sluice boxes, which were long wooden channels that used flowing water and riffles to separate gold from sediment on a larger scale.

Placer mining dominated early operations, focusing on loose gold deposits in streams and riverbanks. As surface deposits became depleted, hydraulic mining emerged - using high-pressure water cannons to blast away hillsides, though this was later restricted due to environmental damage.

The 20th century brought significant technological advances:

Early in the century, dredging became widespread - floating mechanical dredges could process huge volumes of river sediment. Hardrock mining expanded with better drilling and crushing equipment, allowing miners to extract gold from solid rock formations more efficiently.

Later developments included:

- Geochemical prospecting using sophisticated analysis of soil and rock samples
- Aerial surveys and remote sensing to identify promising deposits
- More precise drilling techniques for sampling
- Advanced recovery methods like heap leaching and carbon-in-pulp processing
- Mercury-free extraction techniques addressing environmental concerns

Modern prospecting now combines traditional methods with high-tech approaches like:

- GPS mapping and GIS software for precise location data
- Ground-penetrating radar to detect deposits
- Portable XRF analyzers for rapid mineral identification
- Environmental monitoring systems
- Computer modeling of deposits

This progression shows how gold prospecting evolved from relatively simple manual techniques to a sophisticated industry incorporating advanced technology while becoming more environmentally conscious.

OK, so we have roughly three technological eras: 1) manual, 2) chemical sample analysis and aerial and remote sensing, and 3) high-tech computer-based methods for detection and mapping.

Yes, that's a good way to break it down! Let me build on that framework:

Era 1 (Manual) - roughly 1800s to early 1900s:

Primary methods were direct physical observation and mechanical separation. Prospectors relied on their understanding of geology, visible gold indicators, and mechanical processes like panning, sluicing, and early dredging. Skills were often passed down through experience and apprenticeship.

Era 2 (Chemical/Remote) - roughly 1920s to 1970s:

Brought systematic scientific analysis into prospecting. Geochemical sampling could detect trace elements indicating gold deposits. Aerial photography and later satellite imagery allowed prospectors to identify promising geological formations from above. This era marked the shift from prospecting as an art to a more systematic science.

Era 3 (Digital) - 1980s onward:

Computerization transformed the field. GPS, digital mapping, and data analysis tools allowed prospectors to integrate multiple data sources and create detailed geological models. Ground-penetrating radar and portable analyzers provided immediate feedback in the field. Environmental monitoring became increasingly sophisticated and integrated into prospecting practices.

What's particularly interesting is how these eras overlap and build on each other - modern prospectors often combine techniques from all three eras, using traditional knowledge alongside the latest technology.

Carry-over skills

What kinds of skills carried over from one era to the next? For example, what did highly skilled 19th century prospectors know that remained valuable in the second era? And then from second to third.

This is a fascinating question about the evolution of expertise! Let me trace these skill transfers:

From Era 1 to Era 2:

- Understanding of geological formations and how they relate to gold deposits remained crucial. Early prospectors' knowledge of where gold typically accumulated (like placer deposits in river bends or specific rock formations) helped inform where to conduct chemical sampling and aerial surveys
- Reading the landscape for signs of mineralization - knowledge of indicator minerals, rock types, and geological structures remained valuable even when supplemented by chemical analysis
- Water flow dynamics understanding - experience with how gold moved and concentrated in waterways helped inform sampling strategies
- Assessment of ore quality - the ability to visually evaluate mineral-bearing rock remained important even as chemical assays became available

From Era 2 to Era 3:

- Interpretation of indirect evidence - the skill of making conclusions from multiple indicators (developed during the chemical/remote sensing era) became even more valuable when working with computer models and multiple data streams
- Sampling methodology - understanding of how to collect representative samples remained crucial for computer analysis - Pattern recognition in geological features - this skill trans
ferred well to interpreting digital mapping and radar data
- Knowledge of seasonal and weather effects on prospecting - this remained important for planning surveys and interpreting remote sensing data

What's particularly interesting is how these generations of prospectors had to learn to integrate their traditional knowledge with new technologies - for example, using their understanding of geology to know where to deploy expensive new sensing equipment most effectively.

The core skill that seems to have remained valuable across all eras is the ability to think in three dimensions about underground structures and visualize how geological processes might have created and moved gold deposits. This spatial thinking was valuable whether using a gold pan, analyzing aerial photographs, or interpreting 3D computer models.

Tuesday, October 13, 2020

Cultural Evolution and Oral Tradition: ‘Information transfer’ at the micro scale

Bumping this to the top of the queue because cultural evolution is on my mind these days. This books speaks to a number of issues, including what makes a story culturally "fit": what features make a story easy to memorize/remember?
 
It’s clear that one problem I have with Dennett’s memetics is this his conception face-to-face mechanisms of cultural evolution – like the transfer of information from one computer to another – seems rather thin, unrealistically so. I tend to think that meaning is something arrived at through negotiation whereas Dennett writes as though one-shot one-way ‘information transfer’ is sufficient to the process.

I want to present some passages from David Rubin, Memory in Oral Tradition: The Cognitive Psychology of Epic, Ballads, and Counting-out Rhymes (Oxford UP 1995) that I think merit close consideration. These are passages about oral epic and so are relevant to thinking about folktales, myth and such, stories that are held in memory and delivered to an audience without benefit of written prompt. One thing we need to keep in mind is that, in oral culture, the notion of faithful repetition is not the same as it is in literate culture. In the literate world repetition means word-for-word. In oral cultures it does not. A faithful recounting of a story is one where the same characters are involved in the same (major) incidents in (pretty much) the same order. Word-for-word recounting is not required; in fact, such a notion is all but meaningless. With no written (or otherwise recorded) verification, how do you tell?

This passages illustrates that nicely (pp. 137-138):
Avdo Medjedovic was the best singer recorded by Lord and Parry. An example of his learning a new song provides insights into what it is that the poetic-language learner must learn about his genre (Lord, 1960; Lord & Bynum, 1974). A singer sang a song of 2,294 lines that Avdo Medjedovic had never heard before. When the song was finished, Avdo Medjedovic was asked if he could sing the same song. He did, only now the song was 6,313 lines long. The basic story line remained the same, but, to use Lord's description, "the song lengthened, the ornamentation and richness accumulated, and the human touches of character, touches that distinguish Avdo Medjedovic from other singers, imparted a depth of feeling that had been missing" (p. 78). Avdo Medjedovic's song retold the same story in his own words, much as subjects in a psychology experiment would retell a story from a genre with which they were familiar, but Avdo Medjedovic's own words were poetic language and his story was a song of high artistic quality. Although the particular words changed, the words added were all traditional; and so the stability of the tradition, if not the stability of the words of a particular telling of a story, was ensured.

Several aspects of this feat are of interest. First, the song was composed without preparation and sung at great speed. There was no time for preparation before the 6,313 lines were sung, and once the song began, the rhythm allowed little time for Avdo Medjedovic to stop and collect his thoughts. Such a feat implies a well-organized memory and the equivalent of an efficient set of rules for production. Second, the song expanded yet remained traditional in style, demonstrating that more than a particular song was being recalled. Rather, rules or parts drawn from other songs were being used. Third, although Avdo Medjedovic was creative by any standards, he was not trying to create a novel song; he believed that he was telling a true story just the way he had heard it, though perhaps a little better. To do otherwise would be to distort history.
So, an expert listens to a story than runs to 2,294 lines and then immediately repeats it back, but embellished to 6,313. Would he be able to do the same thing the next day or ten days or a year later? Probably.

Saturday, September 7, 2019

Different languages and similar encoding efficiency during speech despite differences in Shannon information or speech rate

Christophe Coupé1, Yoon Oh, Dan Dediu1, and François Pellegrino, Different languages, similar encoding efficiency: Comparable information rates across the human communicative niche, Science Advances 04 Sep 2019: Vol. 5, no. 9, eaaw2594
DOI: 10.1126/sciadv.aaw2594
Abstract: Language is universal, but it has few indisputably universal characteristics, with cross-linguistic variation being the norm. For example, languages differ greatly in the number of syllables they allow, resulting in large variation in the Shannon information per syllable. Nevertheless, all natural languages allow their speakers to efficiently encode and transmit information. We show here, using quantitative methods on a large cross-linguistic corpus of 17 languages, that the coupling between language-level (information per syllable) and speaker-level (speech rate) properties results in languages encoding similar information rates (~39 bits/s) despite wide differences in each property individually: Languages are more similar in information rates than in Shannon information or speech rate. These findings highlight the intimate feedback loops between languages’ structural properties and their speakers’ neurocognition and biology under communicative pressures. Thus, language is the product of a multiscale communicative niche construction process at the intersection of biology, environment, and culture.

INTRODUCTION

Language is universally used by all human groups, but it hardly displays undisputable universal characteristics, with a few possible exceptions related to pragmatic and communicative constraints (1, 2). This ubiquity comes with very high levels of variation across the 7000 or so languages (3). For example, linguistic differences between Japanese and English lead to a ratio of 1:11 in their number of distinct syllables. These differences in repertoire size result in large variation in the amount of information they encode per syllable according to Shannon’s theory of communication. Despite those differences, Japanese and English endow their respective speakers with linguistic systems that fulfill equally well one of the most important roles of spoken communication, namely, information transmission. We show here that the interplay between language-specific structural properties (as reflected by the amount of information per syllable) and speaker-level language processing and production [as reflected by speech rate (SR)] leads languages to gravitate around an information rate (IR) of about 39 bits/s. This finding, based on quantitative methods applied to a large cross-linguistic corpus of 17 languages, highlights the intimate feedback loops between languages and their speakers due to communicative pressures. We suggest that this phenomenon is rooted in the human neurocognitive capacity, probably present in our lineage for a long time (4), and that human language can be analyzed as the product of a multiscale communicative and cultural niche construction process involving biology, environment, and culture (5).

Each human language provides its speakers with a communication system that fulfills their needs for transmitting information to their peers. The Uniform Information Density hypothesis (6) and similar approaches [e.g., (7) and (8)] suggested that speakers distribute information along the speech signal following a smooth distribution rather than high-amplitude fluctuations. Compatible with Shannon’s theory, this optimization process guarantees the robust information transmission at a rate close to the channel capacity. We adopt here a quite different perspective, where we compare, across very different languages, the average rates at which information is emitted. This approach enables us to estimate the channel capacity and to assess whether the large differences observed among languages in terms of encoding result in analog differences in channel capacity or, conversely, whether there exist compensating strategies that go beyond the local adaptation operating during speech production. Therefore, we investigate the interaction between information encoding and average SR and, more specifically, whether the variation among languages in IR is regulated by communicative constraints. Thus, does too low an IR hinder communicative efficiency? And, at the other extreme, does pushing it too high incur too heavy physiological and cognitive costs? While a negative correlation between average SR and the informativeness of linguistic constituents has been demonstrated in a small multilanguage corpus (9), the distribution of IRs across human languages is almost totally unknown despite its crucial importance for understanding human spoken communication. While our data here come only from speech production (information encoding), our results, nevertheless, implicitly address also speech perception (information retrieval) and processing, as they are all intimately coupled and coevolve during language acquisition, use, and change (10).
Victor Mair comments at Language Log.  There's an article in The Atlantic as well.

Tuesday, December 18, 2018

Stagnation 1.2: Energy efficiency and the cost of deepening our understanding of the world

Yesterday (12.17.18) Alex Tabarrock had a post about the energy efficiency of the ever cheaper logic engines in our many digital devices. That, of course, is in the territory covered by one of the three case studies in Bloom, Jones, Van Reenen, and Webb [1]. In the next section of this post I explicate that connection and then offer some rough and informal remarks about the energy costs of understanding the world.

Silicon productivity and energy efficiency

Over the past 60 years, the energy efficiency of ever-less expensive logic engines has improved by over one billion fold. No other machine of any kind has come remotely close to matching that throughout history.

Consider the implications even from 1980, the Apple II era. A single iPhone at 1980 energy-efficiency would require as much power as a Manhattan office building. Similarly, a single data center at circa 1980 efficiency would require as much power as the entire U.S. grid. But because of efficiency gains, the world today has billions of smartphones and thousands of datacenters.
Here is Mills’ statement of Jevon’s Paradox:
Put differently: the purpose of improved efficiency in the real world, as opposed to the policy world, is to capture the benefits from an engine. So long as people and businesses want more of those benefits, the declining cost of their use increases demand, which in turn outstrips efficiency gains. Jevons understood (and logic dictates) that efficiency gains must come at the same capital cost; but magic really happens when hardware costs decline.
Bloom, Jones, Van Reenen, and Webb took a look at research productivity in the semiconductor industry and discovered that, while the increase in circuit density captured in Moore’s Law, has continued through the late 20th century and into the early decades of this one, the R&D effort goes up steadily (Figure 4, p. 17). Later in the paper they offer an observation that incorporates Jevons’ Paradox (p. 44):
Research productivity for semiconductors falls so rapidly, not because that sector has the sharpest diminishing returns — the opposite is true. It is instead because research in that sector is growing more rapidly than in any other part of the economy, pushing research productivity down. A plausible explanation for the rapid research growth in this sector is the “general purpose” nature of information technology. Demand for better computer chips is growing so fast that it is worth suffering the declines in research productivity there in order to achieve the gains associated with Moore’s Law.
As information processing technology gets cheaper its use spreads further and deeper into social and cultural processes. The productivity of the semiconductor industry may be dropping in economic terms, but that productivity loss enables an enormous increase in the energy efficiency of information processing. The increasing costs of semiconductor R&D are easily covered.

On the whole I’m inclined to reconceptualize that productivity loss as the increasing cost of learning more about the world: what’s out there and how do we build things?

Energy, information, and evolution

Some years ago David Hays and I published an article entitled, “A Note on Why Natural Selection Leads to Complexity” [2]. Here’s the abstract:
While science has accepted biological evolution through natural selection, there is no generally agreed explanation for why evolution leads to ever more complex organisms. Evolution yields organismic complexity because the universe is, in its very fabric, inherently complex, as suggested by Ilya Prigogine's work on dissipative structures. Because the universe is complex, increments in organismic complexity yield survival benefits: (1) more efficient extraction of energy and matter, (2) more flexible response to vicissitudes, (3) more effective search. J.J. Gibson's ecological psychology provides a clue to the advantages of sophisticated information processing while the lore of computational theory suggests that a complex computer is needed efficiently to perform complex computations (i.e. sophisticated information processing).
What about information processing in the evolution of human culture?

In the second quarter of the previous century the anthropologist Leslie White was interested in socio-cultural complexity and put energy consumption at the center of this thinking [3]. More recently David Hays examined a large literature and devoted a chapter to the subject in his history of technology [4]. Hays was interested in the energy usage of preliterate, literate but pre-industrial, industrial, and emerging informatic forms of social and economic organization. To that end he provided estimates in each category of per capita energy flux, energy densities per square mile of inhabited land, “energy taken from the environment, energy delivered to useful purposes, human labor required, [and] material welfare produced”. And he broke down energy usage by category of use: food, domestic, agriculture, transportation, and industry.

But he didn’t segregate informatics as a specific category of energy usage. Nor do I intend to do so here – if anyone knows of someone who has done so, please let me know. But I will offer some observations.

We might begin by asking at what point in socio-cultural evolution we find individuals who are full-time specialists in ‘information processing’ – which is not a particularly good term. Of course we should specify just what that means, but, as I’m only after a quick sketch, I’ll forgo that and suggest that we’re looking for a religious specialist. It is my impression – I recalling a literature I haven’t read in a number of years – that they appear well before the emergence of literacy, but not in the simplest societies.

With literacy would have various information specialists, priests, scribes, philosophers of various kinds, lawyers, engineers, architects and, for that matter, various kinds of artists. We also see the emergence of schools. And some of these activities would require energy beyond that used by their brains, e.g. the production of paper and writing implements.

This brings us back to Mark Mills, Energy and the Information Infrastructure Part 1: Bitcoins & Behemoth Datacenters (RealClear Energy 11.19.2018):
Society has not seen a new “energy service” vector arrive for two centuries, until now. Fouquet’s analysis doesn’t include energy in service of information. Fouquet could, in theory, have calculated energy for information services across those same five centuries. The energy cost to make paper in a single book, while far less now than centuries ago, is still equivalent to the fuel consumed driving a Prius 10 miles. There were also energy costs associated with building the libraries that housed the books, etc. But to be fair to Fouquet, those numbers were so tiny compared to energy for heating that they’d disappear from visibility.

History’s ignition point with regard to the energy cost of information becoming visible can be traced to 1946. The world’s first, and then only, datacenter was ENIAC’s room full of 20,000 burning hot vacuum tubes, which demanded 150 kW. But the proliferation of the new data infrastructures didn’t begin to explode until the Internet’s expansion started at the end of the 20th century — i.e., it began when Fouquet’s history ends. Now the power level of a single ENIAC is found in every dozen square feet inside the billions of square feet of datacenters.

There is no dispute that a new “energy service” has arrived. The core question is whether in fact the trajectory will look like all others in history. [...] The odds are the “information service” trajectory will look the same as for other services Fouquet mapped. By 2050, society will likely use more energy for data than was used for illumination in 1950.
And after that?

* * * * *

What happens to those productivity calculations in that world? The question, no doubt, is ill-posed. But I can live with that. As you may suspect, I think this business of stagnation is somehow ill-posed, as though it harbors a secret wish for the proverbial free lunch, where the lunch is increased knowledge of the world. It’s not free, nor is there any reason why it should be.

In any event, we seem headed for a world in which more and more energy, human and otherwise, is ‘information processing’, with more and more of that being embedded in artificial devices. While I’m skeptical about fantasies about super-intelligent computers – in part because I don’t see much evidence of a conception of intelligence that’s useful for engineering purposes – I’m quite sure that, if we don’t drastically degrade our world or even destroy ourselves, there are surprises over that horizon.

There is a major transformation coming. It involves computers and computation. But also our understanding of them, and, more than likely, of ourselves as well.

More later.

References

[1] Nicholas Bloom, Charles I. Jones, John Van Reenen, and Michael Webb, Are Ideas Getting Harder to Find? March 5, 2018, https://web.stanford.edu/~chadj/IdeaPF.pdf.

[2] William Benzon and David G. Hays, A Note on Why Natural Selection Leads to Complexity, Journal of Social and Biological Structures 13: 33-40, 1990, Academia: https://www.academia.edu/8488872/A_Note_on_Why_Natural_Selection_Leads_to_Complexity; SSRN: https://ssrn.com/abstract=1591788.

[3] Leslie White, Energy and the Evolution of Culture, American Anthropologist, 1943 45:335-356, Download at https://deepblue.lib.umich.edu/bitstream/handle/2027.42/99636/aa.1943.45.3.02a00010.pdf?sequence=1

[4] David Hays, “Energy”, The Evolution of Technology through Four Cognitive Ranks, 1995, Metagram Press. Online (the book’s only form), http://asweknowit.ca/evcult/Tech/CHAPTER3.shtml.

Tuesday, October 23, 2018

Semantic information, agency, and statistical physics

Kolchinsky A, Wolpert DH. 2018 Semantic information, autonomous agency and non-equilibrium statistical physics. Interface Focus 8: 20180041. http://dx.doi.org/10.1098/rsfs.2018.0041
Abstract: Shannon information theory provides various measures of so-called syntactic information, which reflect the amount of statistical correlation between systems. By contrast, the concept of ‘semantic information’ refers to those correlations which carry significance or ‘meaning’ for a given system. Semantic information plays an important role in many fields, including biology, cognitive science and philosophy, and there has been a long-standing interest in formulating a broadly applicable and formal theory of semantic information. In this paper, we introduce such a theory. We define semantic information as the syntactic information that a physical system has about its environment which is causally necessary for the system to maintain its own existence. ‘Causal necessity’ is defined in terms of counter-factual interventions which scramble correlations between the system and its environment, while ‘maintaining existence’ is defined in terms of the system's ability to keep itself in a low entropy state. We also use recent results in non-equilibrium statistical physics to analyse semantic information from a thermodynamic point of view. Our framework is grounded in the intrinsic dynamics of a system coupled to an environment, and is applicable to any physical system, living or otherwise. It leads to formal definitions of several concepts that have been intuitively understood to be related to semantic information, including ‘value of information’, ‘semantic content’ and ‘agency’.

1. Introduction

The concept of semantic information refers to information which is in some sense meaningful for a system, rather than merely correlational. It plays an important role in many fields, including biology [1–9], cognitive science [10–14], artificial intelligence [15–17], information theory [18–21] and philosophy [22–24].1 Given the ubiquity of this concept, an important question is whether it can be defined in a formal and broadly applicable manner. Such a definition could be used to analyse and clarify issues concerning semantic information in a variety of fields, and possibly to uncover novel connections between those fields. A second, related question is whether one can construct a formal definition of semantic information that applies not only to living beings but also any physical system—whether a rock, a hurricane or a cell. A formal definition which can be applied to the full range of physical systems may provide novel insights into how living and non-living systems are related.

The main contribution of this paper is a definition of semantic information that positively answers both of these questions, following ideas publicly presented at the FQXi's 5th International Conference [31] and explored by Carlo Rovelli [32]. In a nutshell, we define semantic information as ‘the information that a physical system has about its environment that is causally necessary for the system to maintain its own existence over time’. Our definition is grounded in the intrinsic dynamics of a system and its environment, and, as we will show, it formalizes existing intuitions while leveraging ideas from information theory and non-equilibrium statistical physics [33,34]. It also leads to a non-negative decomposition of information measures into ‘meaningful bits’ and ‘meaningless bits’, and provides a coherent quantitative framework for expressing a constellation of concepts related to ‘semantic information’, such as ‘value of information’, ‘semantic content’ and ‘agency’.
From a special issue, ‘Computation by natural systems’ organised by Dominique Chu, Christian Ray and Mikhail Prokopenko .

Tuesday, July 31, 2018

Sabine Hossenfelder clarifies entropy

1. Entropy doesn’t measure disorder, it measures likelihood.

Really the idea that entropy measures disorder is totally not helpful. Suppose I make a dough and I break an egg and dump it on the flour. I add sugar and butter and mix it until the dough is smooth. Which state is more orderly, the broken egg on flour with butter over it, or the final dough?

I’d go for the dough. But that’s the state with higher entropy. And if you opted for the egg on flour, how about oil and water? Is the entropy higher when they’re separated, or when you shake them vigorously so that they’re mixed? In this case the better sorted case has the higher entropy.

Entropy is defined as the number of “microstates” that give the same “macrostate”. Microstates contain all details about a system’s individual constituents. The macrostate on the other hand is characterized only by general information, like “separated in two layers” or “smooth on average”. There are a lot of states for the dough ingredients that will turn to dough when mixed, but very few states that will separate into eggs and flour when mixed. Hence, the dough has the higher entropy. Similar story for oil and water: Easy to unmix, hard to mix, hence the unmixed state has the higher entropy.

Tuesday, July 10, 2018

Artificial General Intelligence (AGI): Curiouser and Curiouser [#AI #Mars]

A  few years ago physicist David Deutsch mused on the possibility of fully general artificial intelligence. The bottom line: We don't know what we're doing. Though he takes a round about way to get there.
In 1950, Turing expected that by the year 2000, ‘one will be able to speak of machines thinking without expecting to be contradicted.’ In 1968, Arthur C. Clarke expected it by 2001. Yet today in 2012 no one is any better at programming an AGI than Turing himself would have been.

This does not surprise people in the first camp, the dwindling band of opponents of the very possibility of AGI. But for the people in the other camp (the AGI-is-imminent one) such a history of failure cries out to be explained — or, at least, to be rationalised away. And indeed, unfazed by the fact that they could never induce such rationalisations from experience as they expect their AGIs to do, they have thought of many.
It certainly seems that way, that we're no closer than we were 50 years ago. At least we've managed to toss out a lot of ideas that won't work. Or have we?

Deutsch thinks the problem is philosophical: "I am convinced that the whole problem of developing AGIs is a matter of philosophy, not computer science or neurophysiology, and that the philosophical progress that is essential to their future integration is also a prerequisite for developing them in the first place."

He places great emphasis on the ideas of Karl Popper, whom I admire, but I don't quite see what Deutsch sees in Popper. Still, he manages to marshall an interesting quote:
As Popper wrote (in the context of scientific discovery, but it applies equally to the programming of AGIs and the education of children): ‘there is no such thing as instruction from without … We do not discover new facts or new effects by copying them, or by inferring them inductively from observation, or by any other method of instruction by the environment. We use, rather, the method of trial and the elimination of error.’ That is to say, conjecture and criticism. Learning must be something that newly created intelligences do, and control, for themselves.
That is to say, minds are built from within. And the building is done by fundamental units that are themselves alive and so trying achieve something, if only the minimal something of remaining alive. Those fundamental units are, of course, cells.

Sunday, October 11, 2015

The Diary of a Man and His Machines, Part 2: How’s this Stuff Organized?

As I indicated in the previous post in this series, I’m in the process of transferring my “stuff” to a new computer, my third for this century. So, I thought I’d talk a little about how I organize my stuff, not in detail, though. I’m more or less interested in the simple fact that, after all, you have to organize things some how.

Organization wasn’t an issue for my oldest machines, the NorthStar and the “toaster” Macs, because they didn’t have hard drives. There wasn't anything on them for ME to organize. I kept all my data (and programs for the Mac) on floppy drives. Of course, I generally had more than one document on a floppy, and I’d keep the same kind of stuff on a floppy. Organizing floppies is the same kind of problem as organizing books or files. You put books on shelves and keep similar books together on the shelves. Of course, a given book might be like several others. For example, John Bowlby’s Attachment is psychology, infant behavior, primatology, and psychoanalysis. Just where it goes on the shelves depends on whatever else I’m putting on shelves. Similarly, you put files in boxes, with similar files together, or perhaps alphabetically, but alphabetically by what?

Organizing lots of stuff isn’t easy. There’s a reason library science is called a science. Organizing a library is tricky.

Well, I’ve got 30 years of work accumulated on my various machines. That’s a library. And organizing it is a b*tch.

My Performa, mid-90s, had a hard-drive. So now I could do something other than put disks in boxes. But it was a small hard-drive by today’s standards and I don’t remember anything about it. But I still had lots of floppies hanging around. And I used them.

When I got the G3 Mac I’m pretty sure I moved everything off floppies I could. Now keeping track of my hard-drive became a problem. But there’s always search. And I used it. Still do.

The thing is, when you do as much work as I do, it’s hard to keep things in order. Heck, beyond a certain point it’s hard to even know what order is. And files have accumulated at a fierce clip in the last seven or eight years, spanning my two previous machines. For one thing, I started taking photos. I’ve uploaded almost 18,000 photos to Flickr since 2006. Given that those are just photos that I’ve processed from RAW files, and that I don’t process all my RAWs, that implies that I’ve got 60,000 or more photos floating around on my machine. I’d had to think what a full-time professional photographer has to deal with.

I started blogging at The Valve in December of 2005 and wrote I don’t know how many posts until I logged off in March of 2012, but 100s. I started New Savanna in April of 2010 and have published 3450 posts so far (not counting this one); but 1083 were mostly photos, though some of those would have had a bit commentary. That’s a lot of writing, on lots of topics – literature, cognitive science, neuroscience, film and animation, music jazz jamming, Jersey City, graffiti, my life here and there, and so forth. And I’ve got notes all over the place on all those topics, drafts of papers, and other documents.

And then there’s all the material I’ve downloaded, probably thousands of papers on all those topics and more. I’ve got at least half a dozen folders filled the miscellaneous collections of downloaded stuff and each of those folders has 30 to 100 items in it. And I’ve got a nice little pile of music files, but they’re mostly tucked away in iTunes, where they have some kind of quasi-order. And I’ve even put together a few modest videos, stitching together photos to go along with sound, or even shooting a dozen or so videos of me playing trumpet.

So one of the things I’m doing while moving to this new machine is cleaning up things a bit. But I could easily devote several days, if not a week or more, to doing nothing but looking around and re-organizing. Is it worth the effort?

Sunday, July 19, 2015

Is capitalism dissolving around us?

Capitalism, it turns out, will not be abolished by forced-march techniques. It will be abolished by creating something more dynamic that exists, at first, almost unseen within the old system, but which will break through, reshaping the economy around new values and behaviours. I call this postcapitalism.

As with the end of feudalism 500 years ago, capitalism’s replacement by postcapitalism will be accelerated by external shocks and shaped by the emergence of a new kind of human being. And it has started.

Postcapitalism is possible because of three major changes information technology has brought about in the past 25 years. First, it has reduced the need for work, blurred the edges between work and free time and loosened the relationship between work and wages. The coming wave of automation, currently stalled because our social infrastructure cannot bear the consequences, will hugely diminish the amount of work needed – not just to subsist but to provide a decent life for all.

Second, information is corroding the market’s ability to form prices correctly. That is because markets are based on scarcity while information is abundant. The system’s defence mechanism is to form monopolies – the giant tech companies – on a scale not seen in the past 200 years, yet they cannot last. By building business models and share valuations based on the capture and privatisation of all socially produced information, such firms are constructing a fragile corporate edifice at odds with the most basic need of humanity, which is to use ideas freely.

Third, we’re seeing the spontaneous rise of collaborative production: goods, services and organisations are appearing that no longer respond to the dictates of the market and the managerial hierarchy. The biggest information product in the world – Wikipedia – is made by volunteers for free, abolishing the encyclopedia business and depriving the advertising industry of an estimated $3bn a year in revenue.

Almost unnoticed, in the niches and hollows of the market system, whole swaths of economic life are beginning to move to a different rhythm. Parallel currencies, time banks, cooperatives and self-managed spaces have proliferated, barely noticed by the economics profession, and often as a direct result of the shattering of the old structures in the post-2008 crisis.
Beyond the left:

Friday, March 20, 2015

Where are the polymaths of days gone by? Who's going to make sense of it all?

Jonathan Haber dove whole-hog into online courses put up by major universities and worked his way through bachelor's worth of courses in a year, documenting his work online (@ degreeoffreedom.org). As J. Peder Zane says in the Good Old Gray Lady (aka the NYTimes):
Mr. Haber’s project embodies a modern miracle: the ease with which anyone can learn almost anything. Our ancient ancestors built the towering Library of Alexandria to gather all of the world’s knowledge, but today, smartphones turn every palm into a knowledge palace.

And yet, even as the highbrow holy grail — the acquisition of complete knowledge — seems tantalizingly close, almost nobody speaks about the rebirth of the Renaissance man or woman. The genius label may be applied with reckless abandon, even to chefs, basketball players and hair stylists, but the true polymaths such as Leonardo da Vinci and Benjamin Franklin seem like mythic figures of a bygone age.
Well, of course not. The Renaissance is, you know, so quattrocento.

Monday, February 2, 2015

Some Quick Thoughts on Cultural Evolution

Here are some thoughts I’ve been having on cultural evolution. All of them need fuller exposition, but I don’t have time for that now.

1. Cultural Evolution, so What?

I’ve got a fairly sophisticated narrative account of some varieties of popular music in 20th century America. The account centers on the interaction between African- and European-American populations and musical forms. When I was doing that work–in the late 1990s–I kept thinking that this is the kind of phenomenon that begs for an account in terms of cultural evolution. But I didn’t once use the term “memes” (my term of choice at the time) in the paper nor did I talk of cultural evolution (except in passing at the very end). It’s not at all obvious to me how any existing account of cultural evolution would lead to a deeper understanding of that history. It would just be an exercise in terminology mongering.

This is a very big deal for me, and I don’t yet know what to do about it.


2. Information and Apple Pie

On information, sure, at the highest level of generality and abstraction cultural evolution involves information. But that doesn’t get us much in itself. Until we understand how that information is embodied in the brain and in various media and can measure it, the idea isn’t a very useful or deep one. It’s one of mere terminology, verbal packaging. But see comment six, below.

3. Information, DNA, the Brain and the Rest

In biology we know that genetic information is embodied in DNA and we know a great deal about how that works. And we are rapidly increasing our ability to manipulate DNA.

We don’t know how information is encoded in the nervous system. But that’s not so much my point here. It belongs in #2 above.

My point here is that, in long-held view, the genetic information for culture isn’t in the brain. It’s in publicly accessible traits of physical objects and processes. That means that, whereas the genetic information of life forms is encoded in the same medium, DNA, that is not the case for cultural genetic information. The genetic information for culture has various embodiments.

From the standpoint of theoretical elegance/parsimony, this is not so good. But the fact is, no matter what kind of model you choose, if you want to deal with information, you are going to have to deal with multiple physical embodiments and the transformations between them.

4. Mutual Information for Culture

We’re interested in the ‘conservation’ of information in the social group. That requires shared access to common reference points. The meaning of those reference points is, in effect, negotiated through interaction. Those agreed reference points are the ‘genes’ of the cultural process.

Compare this with DNA replication. The double strand divides and then each half constructs the necessary complement. The result is two identical DNA molecules where there had been one.

Shared access to a common public reference plays the role in cultural informatics that DNA replication plays in biological informatics.

On mutual information in culture, see my Open Letter to Steven Pinker.

5. The mind as computer program

Programs have variables and variables require values. Where do mental programs get values for their variables? Some of them are generated internally.

And some of them come from the external world. Among those we have the cultural genetic material. They function as values for variables in mental programs.

See this post for further specification of this thought: Memes as Data: Targets, Couplers, and Designators. That is collected in my working paper, Cultural Evolution, Memes, and the Trouble with Dan Dennett.

Wednesday, January 28, 2015

Cultural Beings & Intertextuality: Information

I’ve got two quickish thoughts on cultural evolution, once concerning the concept of cultural beings and the other in my ongoing ‘war’ against the concept of cultural information.

Cultural Beings and Intertextuality

I’ve recently introduced the term “cultural being” to indicate a package or envelope of coordinators (aka a ‘text’) plus the ‘trajectories’ of that package in the minds of all who encountered it. Such misgivings as I still have stem from the fact that it is in fact an odd notion, though it’s not so different from how the concept of ‘the text’ is in fact used in literary criticism. Literary critics also talk about intertextuality, how texts are related to one another.

The fact that Greene’s Pandosto was reworked into Shakespeare’s The Winter’s Tale is thus a fact about the intertextuality of both texts. Shakespeare may have written his text, but he did so after having read Greene’s text and using is as a model for his. The Winter’s Tale is thus, in a sense, a daughter of Pandosto. And most of Shakespeare’s plays are daughters of multiple sources.

The cultural being that is associated with Pandosto would thus extend into Shakespeare’s mind (and into mine as well). That is the ‘route’ through which it ‘influenced’ Shakespeare’s play. More generally, later cultural beings are the result of the intermixing of earlier cultural beings in the minds of authors.

Information in Cultural Evolution

I have two objections to standard memetic talk of memes as cultural information. One is simply that the concept is not very clear (see the addendum). The other is that it is used to paper over the very tricky and interesting question of how one mind influences another. “Oh, one mind just transfers information to the other mind.”

Not only is this a mistaken account of what physically happens during communication (which Michael Reddy critiqued with is account of the conduit metaphor) but it also glosses over the fact that communication is often imperfect. That very imperfection is one potential source of cultural variation. The mind that reads a signal, any signal, is not identical to the mind that sends it, and so it may misread the signal. Where the two minds are in face-to-face interaction it may be possible to negotiate a satisfactory understanding. Where negotiation is impossible, misunderstanding may be inevitable and hard to eradicate.

One case that I find particularly interesting is that of the musical interaction between European-Americans and African-Americans in 20th century musical styles. Instrumental styles pass from black music to white music more easily than do vocal styles, but even there we can see differences. The basic harmonic and melodic practices transfer easily enough, but the rhythmic nuances do not. The neuromuscular systems that reconstruct the music are different from those that produce it. Thus Pat Boone’s versions of Little Richard’s tunes are anemic in comparison to the originals. The “information” didn’t quite “transfer.”

Addendum: John Wilkins on Information

This is from John Wilkins’ entry, Replication and Reproduction, in the Stanford Encyclopedia of Philosophy:
The literature dealing with information is both extensive and factious. Several different formal analyses of information can be found and very little agreement about which analysis is best for which subjects. On one point these scholars tend to agree—cybernetic information and communication-theoretic information will not do for replication in biological contexts. The best bet is semantic information (Sterelny 2000a; Godfrey-Smith 2000; Sarkar 2000). The trouble is that no widely accepted version of semantic information exists. Winnie (2000) distinguishes between Classical and Algorithmic Information Theory and opts for a revised version of the Algorithmic Theory. But once again, the problem is that no such formal analysis currently exists. In the face of all this disagreement and unfinished business, biologists such as Maynard Smith (2000) maintain either that informal analyses of “information” are good enough or that some future formal version of information theory will justify the sorts of inferences that they make. The sense of “information” as used in the Central Dogma of molecular biology, which states that information cannot flow from protein to DNA, is more like a fit of template, or the primary structure of the protein sequence compared to the sequence of the DNA base pairs. Attempts have been made in what is now known as bioinformatics to use Classical Information Theory (Shannon's theory of communication) to extract functional and phylogenetic information (Gatlin 1972; Maclaurin 1998; Wallace and Wallace 1998; Brooks and Wiley 1988), but it appears to have been unsuccessful in the main. While the most likely conclusion is that no version of information theory as currently formulated can handle “information” as it functions in biology (see Griffiths 2001 for further discussion), attempts have been made to formulate just such a version (Sternberg 2008; Bergstrom and Rosvall 2011). However, this undercuts the motivation for appealing to information theory to elucidate genes in the first instance.

Saturday, April 26, 2014

What is computing? It's more than doing sums!

Look at some of the presentations for the 2nd Workshop on Mind, Mechanism, and Mathematics (Mat 12-13, NYC):
  • Mark Braverman (Princeton University) - Protecting a Conversation Against Adversarial Interference
  • Rebecca Schulman (Johns Hopkins University) - Software for Matter: Programming the Morphogenesis, Replication and Metamorphosis of Everyday Things
  • Martin Davis (New York University and UC Berkeley) - Gödel, Mechanism, and Consciousness
  • Benjamin Koo (Tsinghua University) - CELL: A Cognitive Extreme Learning Lab
  • Paul Grant (University of Cambridge) - Synthetic Spatial Patterning Using Two-Channel Quorum-Sensing Signaling
Check out the Big Questions:
1. The Mathematics of Emergence: The Mysteries of Morphogenesis
2. Possibility of Building a Brain: Intelligent Machines, Practice and Theory
3. Nature of Information: Complexity, Randomness, Hiddenness of Information
4. How should we compute? New Models of Logic and Computation

Friday, June 14, 2013

Culture Memes Information WTF!

I’ve been thinking a lot about information recently, mostly as a consequence of reading Dan Dennett on memetics. I’m uncomfortable with his usage, and similar ones, and I can’t quite figure out why. Let me offer two passages, and then some comments by way of thinking out loud.

The first passage is from George Williams, a biologist. It’s in a chapter from a book edited by John Brockman, The Third Culture: Beyond the Scientific Revolution:
Evolutionary biologists have failed to realize that they work with two more or less incommensurable domains: that of information and that of matter. I address this problem in my 1992 book, Natural Selection: Domains, Levels, and Challenges. These two domains will never be brought together in any kind of the sense usually implied by the term "reductionism." You can speak of galaxies and particles of dust in the same terms, because they both have mass and charge and length and width. You can't do that with information and matter. Information doesn't have mass or charge or length in millimeters. Likewise, matter doesn't have bytes. You can't measure so much gold in so many bytes. It doesn't have redundancy, or fidelity, or any of the other descriptors we apply to information. This dearth of shared descriptors makes matter and information two separate domains of existence, which have to be discussed separately, in their own terms.

The gene is a package of information, not an object. The pattern of base pairs in a DNA molecule specifies the gene. But the DNA molecule is the medium, it's not the message. Maintaining this distinction between the medium and the message is absolutely indispensable to clarity of thought about evolution.

Just the fact that fifteen years ago I started using a computer may have had something to do with my ideas here. The constant process of transferring information from one physical medium to another and then being able to recover that same information in the original medium brings home the separability of information and matter. In biology, when you're talking about things like genes and genotypes and gene pools, you're talking about information, not physical objective reality. They're patterns.

I was also influenced by Dawkins' "meme" concept, which refers to cultural information that influences people's behavior. Memes, unlike genes, don't have a single, archival kind of medium. Consider the book Don Quixote: a stack of paper with ink marks on the pages, but you could put it on a CD or a tape and turn it into sound waves for blind people. No matter what medium it's in, it's always the same book, the same information. This is true of everything else in the cultural realm. It can be recorded in many different media, but it's the same meme no matter what medium it's recorded in.
It seems to me that that is more or less how the concept of information is used in many discussions. It’s certainly how Dennett tends to use it. Here’s a typical passage (it’s the fifth and last footnote in From Typo to Thinko: When Evolution Graduated to Semantic Norms):
There is considerable debate among memeticists about whether memes should be defined as brain-structures, or as behaviors, or some other presumably well-anchored concreta, but I think the case is still overwhelming for defining memes abstractly, in terms of information worth copying (however embodied) since it is the information that determines how much design work or R and D doesn’t have to be re-done. That is why a wagon with spoked wheels carries the idea of a wagon with spoked wheels as well as any mind or brain could carry it.
Here I can’t help but think that Dennett’s pulling a fast one. Information has somehow become reified in a way that has the happy effect of relieving Dennett of the task of thinking about the actual mechanisms of cultural evolution. That in turn has the unhappy effect of draining his assertion of meaning. In what way does a wagon with spoked wheels carry any idea whatsoever, much less the idea of itself?