Wednesday, September 2, 2026

CBR: "Can AI Map Wall Street’s Genome?"

From the University of Chicago, Booth School of Business, Chicago Booth Review, August 31:

The same architectures used to process language could shed new light on financial markets. 

he AI boom began unassumingly, with a 2017 research paper titled “Attention Is All You Need” about how to translate text between languages. At the time, the findings did not look like the start of a trillion-dollar revolution. They looked more like a technical advance toward solving an engineering problem in natural language processing, then just one of several developing branches of artificial intelligence. It seems evident that the authors, all Google employees, had little inkling of the paper’s future fame in the way they simply went with a randomized order for their names rather than jockeying for first position.

Much has stemmed from that paper. It largely created the AI architecture that underpins large language models, which in turn run chatbots. It also gave it a name: the transformer. And although transformers grew out of natural language processing, they haven’t stayed in that domain. Netflix has trained its own version on viewing histories to improve recommendations; Stripe trained one on payment data to spot fraud; Google DeepMind trained one on weather patterns to improve forecasts; and the list goes on.

Now transformer models are being trained on the “language” of financial data, bringing their prediction-making capabilities to the world’s capital markets, whose worth is approaching $300 trillion. Sophisticated trading firms are already using them in the latest evolution of market sparring.

Markets may be the ultimate test for transformers, which are being used to illuminate unseen, dynamic, ever-changing relationships in stock returns, high-frequency trading data, and more. One group of researchers that includes Chicago Booth’s Ralph S. J. Koijen, collaborating on what they call the Market Genome Project, has turned a transformer model on investor portfolios to find out what moves prices. A stock’s momentum, a sector’s rotation, the aftershocks of an earnings surprise—these often aren’t random but reflect patterns that unfolded months earlier and expectations about future events. They’re traces of how investors think and act. The question is whether transformers can learn to read those traces.

How we got here 
Let’s start with a simple lesson about transformers. “Paris → France as Berlin → __________.” If you were presented with this analogy, you would know that the missing word is Germany. The solution does not involve step-by-step reasoning but rather recognizing a relationship and applying it to find the answer. Paris is the capital of France, a country. Berlin is also a capital, and therefore the missing information must be the country for which Berlin is the capital.

Before the 2017 paper came out, machine learning generally processed language word by word and struggled to capture contextual relationships spread out across text. The models did not have a good mechanism for saying, essentially, “This thing here depends on something that appeared earlier.”

Transformers changed that. They process whole sentences at once, solving the problem that earlier models had with understanding these relationships. More precisely, they learn which pieces of information pertain to which other pieces, regardless of position. That single shift turned context from a liability into an asset. It made it possible to pretrain models on vast, messy data and extract useful structures. The architecture proved able to capture long-range dependencies—that is, connect distant words that matter to each other.

Markets have their own sequences of long-range data, which would seem ready-made for transformers. Much of investing is front-loaded with information. Investors examine cash flows, compare business models, and think through how the economy might evolve over the next year or two—all before investing capital.

These analyses resolve into decisions about what to own, how much of that asset to hold, and what to pair it with. The choices compress all of the prior work into a single transaction. This bottom-up approach, focused on individual companies or industries, has traditionally framed how asset pricing is understood. Fundamental analysis drives investor decisions, and prices and portfolio allocations are the outcomes.

But this framing misses that investors’ actual portfolio choices contain information of their own. In 2015, Koijen and Princeton’s Motohiro Yogo began developing what they called a demand system approach to asset pricing. Since asset prices ultimately reflect a supply of assets and investor demand for them, the best place to look for what moves prices is in investor portfolios, the researchers argued.

Starting with institutional portfolios, Koijen and Yogo reviewed quarterly US equity holdings from Form 13F filings to the Securities and Exchange Commission between 1980 and 2017. These filings reported the stock positions of institutional investors managing more than $100 million. The data included in the sample covered roughly 68 percent of the US stock market during that time period.

Their findings upended conventional wisdom about prices. Supply dynamics such as changes in the institutional investors’ share counts (the number of shares outstanding) explained about 2 percent of the variation in returns. Changes in firm characteristics explained about 10 percent. Dividend yield explained less than 1 percent. Put together, all supply-side factors accounted for only about 12 percent of the variation.

Koijen and Yogo find that demand mattered more, but not in the obvious ways. Growth or shrinkage in investors’ assets under management explained just over 2 percent of variation, while changes in how investors weight firm characteristics explained about 5 percent.

The big driver was something harder to see: shifts in latent demand not captured by traditional characteristics such as market capitalization, profitability, and value. Changes in which stocks institutions chose to own explained about 23 percent of the variation. Changes in how much they owned explained nearly 58 percent. These demand shifts accounted for more than 80 percent of the differences in stock returns....

....MUCH MORE 

If one is so inclined see also November 2024's "While Google Slept: ChatGPT and Transformers" and December 2025's AI Architecture: "An AI Startup Looks Toward the Post-Transformer Era".