On July 21 I posted some remarks about the final chapter in Tyler Cowen’s monograph on marginalism: Beyond Marginalism: What’s Next? [MR #12]. In that Chapter Cowen discussed an asset pricing model that used machine learning to create a 360,000 parameter model that gave better predictions than classic models using only four or five factors. The factors in those classic models are predefined on an intuitive basis. Cowen despairs of making intuitive sense of those 360,000 factors in the machine learning model.
In my post I argued that those 360,000 factors might be capturing the effects of Keynes’s animal spirits as expressed in the stories and gossip Schiller writes about as narrative economics. Two days later I discussed the idea with Marge, the AI attached to the online version of Cowen’s text. In that post I suggested a method for going on a “fishing expedition” to determine whether or not my suggestion had merit. Marge’s response: “The fishing expedition you're proposing is methodologically clean, and the prediction is specific enough to be falsifiable — which is more than can be said for most conjectures at this level of abstraction.”
That brings us to today, where Cowen has posted the abstract of an article about “a novel measure of firm-level cyber risk exposure based on the quarterly earnings calls of listed firms.” I read that as being complementary to my speculation. The authors of that article are looking for the effects of a specific line of narrative, and they found them. So I asked Marge to clarify the relationship between my speculation and their finding. Here’s that conversation.
* * * * *
Tyler just posted the abstract of this paper to Marginal Revolution: Jamilov, Rustam and Tahoun, Ahmed and Rey, Helene, The Anatomy of Cyber Risk (May 10, 2023). The Journal of Finance, Forthcoming. Here’s the abstract:
This paper uses computational linguistics to introduce a novel measure of firm-level cyber risk exposure based on the quarterly earnings calls of listed firms. Our measure covers more than 14,000 firms from over 90 countries between 2003 and 2025. The measure is validated using human auditors and a large language model. We show that cyber risk exposure affects stock returns and profits, is priced in the options market, predicts actual cyberattacks, and propagates from firm to sector level. Back-of-the-envelope estimates suggest that the global cost of cyber risk exposure is around $1 trillion per year.
So I took a quick look at the paper to find that novel measure. Here’s what I found. First this:
Our measurement approach consists of two phases. First, we follow Baker et al. (2016) and Hassan et al. (2019) to construct a comprehensive pre-defined dictionary of terms related to cybersecurity risk. We assemble this dictionary from three reputable institutional sources that act as information aggregators on the practical cyber-risk issues firms face on a daily basis. These libraries contain most of the keywords commonly used in cyber-related discussions among private market participants across industries.
Then this:
The second step of our measurement approach builds on the insights from Sautner et al. (2023) and involves a keyword discovery algorithm. This method is based on Meta AI Research lab’s word embedding tool FastText. The algorithm has been trained on a billion phrases in corporate earning calls to suggest related keywords.
And this:
We now construct our baseline quarterly measure of firm-level cyber risk exposure. Let the set of all terms in our final dictionary be C. Our algorithm counts the number of sentences in earnings calls that contain at least one term in C.
They then explain their measure. What’s important for my purposes is where they went looking for those keywords, in earnings calls. They’re measuring the effects of one kind of animal spirits/gossip/narrative on PRICING IN THE OPTIONS MARKET, among other things. I seems to me that they’re looking at one feature of the information that I’ve speculated is smeared of those 360,000 parameters.
That's a sharp observation, and it reframes what Jamilov, Tahoun, and Rey are doing in a way they probably wouldn't frame it themselves.
Their methodology is essentially a named-narrative extraction — they're pulling a specific, bounded narrative thread (cyber risk discourse) out of earnings calls and showing that it has measurable pricing effects. The earnings call is a fascinating data source for your purposes because it's a ritualized social performance in which management constructs and transmits narratives to analysts and investors. It's not raw gossip — it's institutionalized storytelling with real money riding on how the story lands. The fact that cyber risk language in these calls predicts actual cyberattacks and options pricing suggests the narrative isn't just epiphenomenal. It's carrying genuine information, but in a form that only becomes legible when you treat language itself as data rather than just a vehicle for explicitly stated facts. [See text: Ch.4 · ChatGPT digesting news predicts returns especia...]
Your conjecture about the 360,000-factor model is that it's picking up undifferentiated narrative residue — the full spectrum of animal spirits distributed across parameters with no single factor carrying an interpretable label. What Jamilov et al. are doing is the named version of the same operation: isolating one specific narrative thread and demonstrating it has pricing consequences. These are complementary rather than competing approaches. The named-narrative approach gives you interpretability and causal traction on specific mechanisms — you can say "cyber risk discourse in Q3 earnings calls predicts options pricing in Q4" and tell a coherent story about why. The high-parameter ML approach captures everything the named approach misses, which is presumably most of the narrative space. [See text: Ch.4 · Machine learning finds nonlinearities and inter...]
The research design implication is interesting. If you could run both models on the same dataset — the 360,000-factor ML model and the Jamilov-style named-narrative extraction — and then ask how much of the ML model's advantage over classical factors is absorbed by adding the named narrative measures, you'd get a rough estimate of how much of the distributed signal is actually identifiable narrative content versus genuinely opaque pattern. If the named narratives absorb a substantial fraction of the ML advantage, that's evidence for your animal spirits hypothesis in a fairly direct form. If they absorb very little, it suggests the ML model is picking up something even more diffuse than nameable narratives — perhaps the interaction structure between narratives, the way stories about cyber risk combine with stories about AI capability and geopolitical tension to produce pricing effects that none of the individual narratives would generate alone. [See text: Ch.4 · Machine learning models have tens of millions o...]
The options market finding is particularly telling for your framework. Options pricing is where sophisticated investors put explicit probability estimates on future states of the world. If cyber risk language — not cyber risk events, but the narrative about cyber risk — is priced into options, that's direct evidence that markets are treating narrative content as probabilistically informative about future fundamentals. Which is exactly what your conjecture requires: animal spirits aren't just noise distorting rational pricing, they're carrying genuine forward-looking information in a form that classical factor models can't extract but language-sensitive tools can. Keynes was right that animal spirits move markets, but the mechanism is more epistemically respectable than the term implies.
No comments:
Post a Comment