Essay

What the S&P 500 admits about AI (when the SEC is reading)

MIT FutureTech read ten years of SEC 10-K filings, the annual report a CEO personally certifies with securities law watching, and scored 510 S&P 500 firms on how deeply AI actually runs the business. By 2025, 11 percent had got there. Two out of three firms clearing the advanced bar are in tech. And on the number that was supposed to justify all of it, the authors report no productivity gains at any level of adoption.

Ask a CEO about AI on an earnings call and you get theater. Yu, Fleming, Hampton, Combemale, and Thompson, a team out of MIT FutureTech with a coauthor at Carnegie Mellon, posted an arXiv preprint last month that moves the question to the venue where theater gets expensive: the 10-K, which the CEO and CFO certify personally, and where securities law prohibits materially false or misleading statements. The SEC’s first AI-washing order against a public company shows the mechanism working. Presto Automation spent five registration statements claiming its drive-thru voice AI “eliminat[es] human order taking,” while agents in the Philippines and India entered every order its original system took. Once the Commission came asking, the unwind ran through the filings: the October 2023 10-K disclosed the humans, and by December the company had conceded that intervention ran at 100 percent of orders at most locations. The hype lived in the sales documents. The correction got filed. That price on hype is the closest thing to a “truth serum” the disclosure system has.

So what does it actually test? Less than the word promises. The law suppresses the checkable lie, and Presto’s was checkable: a claimed non-intervention rate against a countable reality. The paper’s rubric grades something softer, how centrally a firm narrates AI in its filing. Meta scores a 5 for boilerplate about “new generative AI experiences”; Best Buy scores a 2 for a sentence about training employees. Nothing on that gradient is a number anyone can audit. And the law runs one direction: it polices overstatement and says nothing about silence, so real use below the materiality line never has to reach the document at all. The grader, meanwhile, is GPT-5-mini, spot-checked by hand and pronounced “largely consistent with common expectations”, no gold set, no agreement statistics, with the Census Bureau’s adoption survey (correlation 0.76) and the fintech Ramp’s transaction data (0.87) as sector-level outside checks the authors themselves flag as not apples-to-apples. Call the instrument what it is: a certified-narration index.

Two numbers carry the finding. By 2025, 11 percent of S&P 500 firms scored a 5, deep integration. A further 10 percent were running AI at production level with explicit financial expectations attached. Together, 21 percent of the index clears a genuinely advanced bar, more than four times the 5 percent that cleared it in 2022. Two of every three of those firms sit in tech. And 18 percent of the S&P 500 in 2025 never mention AI at all, the cell the one-directional law lets stay quiet: read it as undisclosed, not necessarily unadopted.

All of the movement is three years old. The scores run essentially flat from 2016 through 2022, then break upward. The authors date the discontinuity to ChatGPT’s release. Hold that against the clocks of earlier general-purpose technologies: factory electrification needed about forty years to reach most American plants, and the computer needed eight from Solow’s productivity-paradox quip to the mid-nineties acceleration. Advanced adoption quadrupling in three years, inside the biggest and most process-bound companies in the American economy, is by those clocks a sprint. Which sharpens the real question: what has all that speed bought?

On margins, the paper gives two answers. Its conclusion says firms early in integration run 1 to 3 points below non-adopters and the deep embedders come in around 6 above. Its preferred regression says more: 12.6 points for the non-technology firms that go all the way in, 15 with sector-year controls, about five for tech. The summary quotes the blend; the table is louder. Before reading either as pay-first-collect-later, look at the cell the big number stands on: nineteen non-technology firms occupy the deepest level, a group the authors say carries too wide a confidence interval for the descriptive comparison to mean much, and on the neighboring measure, Tobin’s Q, the deep-cell anomaly is largely driven by two outliers, Align Technology and Paycom. Two firms can move these coefficients. What the design does earn: firm fixed effects measure each company against its own baseline as it climbs the rubric, which is as close to a J-curve reading as a descriptive panel gets, and the authors still refuse the causal word, “we are not making statements about causality.” With nineteen firms at the top, they are right to.

Run the same design on revenue per employee, the paper’s productivity proxy, and the J loses its upstroke. The authors’ own sentence is the flattest in the study: “we don’t find meaningful productivity gains due to AI adoption.” The tables are harsher: exploration- and pilot-stage firms run about 4 percent below non-adopters (the pilot estimate significant at the one percent level), and the deepest adopters are statistically indistinguishable from firms that never started. Two things keep that null honest. The proxy is a ratio, and deep adopters are the group hiring fastest, around 20 percent more headcount for non-tech and 36 for tech, so the denominator swells exactly where the gains should show. And it is still a sturdier null than the viral “MIT says 95 percent of AI pilots fail” headline of August 2025, which grew out of an interview-and-survey deck from a different MIT group; this one reads the legal record of the whole index, ten years deep.

The obvious defense wins. A J-curve’s clock starts at adoption, and the top of this index is even younger than the ChatGPT line suggests: deep integration stood at 5.4 percent in 2023 and 11.4 in 2025, so roughly half the firms at level 5 arrived within the last two years. One to three years into the pay phase, the Brynjolfsson model says the books should look exactly like this whether the payoff is en route or never coming. So the honest reading is narrower and stranger than either camp wants: the disclosure record of the biggest companies in the economy, read end to end, shows them deep in the pay phase of an enormous reorganization, nothing collected yet, and no way to show which future they are in. Neither side of the trade gets to mark it to market.

Worth naming where I sit before I use it. I left Opendoor for Closin on close to this exact bet, that the founder who earns a domain himself, fast, beats the founder waiting on his employer’s AI rollout to clear committee review. A paper showing institutions mid-dip is a convenient paper for that bet. Discount me accordingly, and go read it yourself before you take my read on it.

The paper’s half of the claim ends at the org chart. The authors scope individuals out in one sentence, a stand-alone AI tool “may increase worker productivity,” and decline to measure it. Fair: their unit is the enterprise. But their own reference list points at the missing half. Among the explanations for the null they cite Chan and Shedania, whose title is a question: why do micro-economic productivity gains from AI disappear at scale? Read from the operator’s side, that is no paradox. AI has lowered the cost of domain knowledge: expertise that used to cost a graduate degree and a hiring committee now costs a weekend of reading. A company deep-integrates by reorganizing tens of thousands of people around that fact, and the reorganization is the dip on its books. A founder reorganizes one person. The gains are real in a person’s hands and thin out on the climb up the org chart, and the index only reads the top, where the theater is performed and certified. My half runs on the shorter clock: wherever a few people now ship work that used to take a department, the test is already running, and it settles in quarters, with no 10-K required to report the result. The spread between those two clocks is the arbitrage.


Sources: