Alpha Without
Prediction
Language models cannot forecast prices, and the honest literature stopped pretending otherwise. That is not the interesting question. The interesting question is what an agent that reads everything, remembers, and never gets bored is actually worth, and the answer turns out to be specific, measurable, and mostly not about predicting anything.
Published: September 14, 2026 · Read time: ~21 minutes
The thesis
Almost everything that looks like alpha in the agent literature turns out to be something else: search intensity nobody counted, market and style exposure nobody attributed, or knowledge the model had already memorised. Three separate 2026 studies dismantled the headline results, and a live deployment of thousands of trading agents lost money.
What survives is not forecasting. Agents do not raise your skill per bet. They raise how many times you can use the skill you already have. Everything worth building here follows from that one distinction. Scope: evidence through September 2026 across analyst research, corporate disclosure, earnings communication, options, agentic architectures, and alternative data. Several figures are single-source; see the caveats at the end.
Agent-discovered strategies that passed certification once search intensity was counted (two universes, two frontier models, ~100 candidates each)
Agents with positive stock-selection alpha. The rest were market and style tilts in disguise (masked-ticker factor attribution)
Names covered per analyst at one fund that deployed agents properly (Avala Global)
Senior-analyst document task, at another (Balyasny)
Everything that looks like alpha in the agent literature is something else
Three 2026 studies independently dismantled the field's headline results, each by instrumenting a different failure, and together they leave very little standing.
The first counted searches. An August 2026 paper kept a trial ledger, recording every evaluation an agent ran rather than only the one it reported, then deflated the Sharpe ratio against that actual search intensity. The agent's best discovery had a design Sharpe of 1.69. Indexed to its own 102 trials, that became a deflated 0.86. Out of sample it delivered 0.18. Across two universes, two frontier models and roughly a hundred candidates each, not one agent-discovered strategy passed certification.
The same paper planted a deliberately leaky strategy, one that could see the future, with a Sharpe above 34. It sailed through deflation with a perfect score. That is the finding worth carrying: multiple-testing correction does not catch leakage. They are two different diseases, they need two different cures, and almost every published agent result treats them as one.
The second study masked the tickers and dates, then ran a standard factor attribution over what the agents actually held. Nine of ten had negative stock-selection alpha, some catastrophically so. The advertised returns of 85% and 61% turned out to be market exposure and style tilt. The agents were not picking stocks. They were long beta and long momentum, and nobody had checked.
The third measured what the model already knew. Subtracting memorised knowledge of the test period cut reported in-sample returns by 45% on average and by 78% at worst. The sharp detail is that the obvious defences make it worse: anonymising the company or instructing the model to ignore what it knows both scored lower than doing nothing, because they degrade reasoning without removing memory.
Read any agent backtest against this
A survey of 77 agent-trading studies found 19 with closed-loop evaluation. Of those nineteen, two reported time-consistent data splits, one modelled transaction costs, and one documented survivorship. None reached the top reproducibility tier. The most-cited system in the field reports a Sharpe of 8.21 over three months on three tickers with no stated costs, and its own authors flag the number as implausible.
Six months, seven and a half million model calls, and a loss
The strongest evidence is not a backtest. In 2026 two fleets of autonomous LLM trading agents ran in production with real capital: about 3,500 user-funded vaults trading for three weeks, then roughly 550 agents trading perpetual futures for ten more. Together they produced 7.5 million model invocations and around 300,000 on-chain actions. Somebody finally published what happened.
They lost. Cumulative realised profit and loss on the second fleet was negative $217,000. Fifteen percent of agents were net positive, against 53% of the retail accounts trading alongside them. Round-trip win rate was 41% versus retail's 50%. On the first fleet the median vault returned 0.492 times capital, and 16.2% finished profitable.
They also converged. Portfolio overlap between agents rose steadily through the run, and on a single day in March, 1,544 of 3,454 active vaults bought the same token within one hour. A population of independent agents reading the same feeds is not a population of independent bets. It is one bet, placed many times, which is the precise opposite of what breadth is supposed to buy you.
And they hallucinated persistently. Between 41% and 54% of observations had agents recording funding income on a venue that pays no funding, then reasoning from that fabricated memory later.
Three frontier models spanned 263.27 to 264.37 basis points of regret. Every confidence interval overlapped every other. There was no measurable difference in decision quality between them, and a twenty-five-fold difference in what they cost to run.
The authors' central finding is the one worth stealing. What determined agent behaviour was not the model and not the strategy prompt. It was the operating layer: the risk slider, the shape of the candidate list the agent was shown, the mechanics of the order path. A risk slider moved leverage by 0.425 per level. Agent identity absorbed 60% of behavioural variance. Where the leaderboard cut off at the top three, selection jumped by a factor of 1.75 for no reason other than the render.
And the thing that did work was mechanical
Attaching a fixed stop and target bracket at entry earned 39 basis points per position, with a confidence interval of 21 to 57 that clears zero comfortably. Meanwhile one risk-slider setting, holding 11% of the book, accounted for 62% of all liquidations. The paper's conclusion is blunt: every prompt-side attempt to fix a behaviour that lives in the operating layer underperformed a one-line change to the order path or the render.
Agents do not raise your skill. They raise how many times you can use it.
Active management has a governing equation, and it explains the entire results pattern above. Grinold and Kahn's fundamental law says your information ratio is roughly your skill per bet multiplied by the square root of how many independent bets you make, scaled by how much of your intended position you actually get on.
TC
How much of the position you actually get on.
IC · skill per bet
Agents barely move this. It is the hard term.
N · independent bets
Agents move this hard. It is the easy term.
Covering 20 names badly is worth less than covering 200 names at the same modest skill. Ten times the breadth is 3.2 times the information ratio.
Every piece of evidence in this piece falls into place against that equation. The agent literature has been trying to raise skill per bet, which is the hard term, by asking a language model to be a better forecaster than the market. It cannot, and the certification results show it does not. Meanwhile the funds getting real value are attacking the breadth term, which is the easy one, by having agents read what nobody had time to read.
That is exactly what the deployment evidence looks like. One fund reports analysts covering 200 names where they used to cover 20. Another cut a senior analyst's document task from two days to thirty minutes. Neither claims the model is smarter than the analyst. Both claim it lets the analyst's existing judgment touch ten times as many situations.
Two conditions, and both are engineering problems
Breadth only counts if the bets are independent, which is precisely what the production fleets destroyed when their portfolios converged and 1,544 agents bought the same token in an hour. And breadth bought by loosening your threshold is not breadth at all, because it lowers skill per bet at the same time.
The version of this that fails
There is a real counterweight, and it is worth sitting with, because the naive reading of the equation is wrong in a specific way. When a large data vendor rolled generative AI out to its analysts, those analysts consulted 26% more sources and covered 24% more topics. Their forecast accuracy fell.
A second study makes the same point from the other direction. Across 121,000 forecasts, human analysts had a median error of 1.9% of price against a language model's 2.3%, and the human advantage was largest for exactly the small, loss-making, leveraged firms where synthetic coverage was supposed to win. On another sample a seasonal random walk beat the model outright.
More inputs is not more breadth. Breadth is one validated standard applied to more situations. More opinions is just more noise wearing the costume of coverage.
The distinction is sharp in practice. The strongest result below ran a single fixed embedding model over 1.19 million analyst reports. One standard, enormous coverage, no per-name judgment calls. That is breadth. Handing an analyst a chatbot and hoping they read faster is not.
Eleven places people look, and what is actually there
What follows is every family of signal the research covers, with a verdict attached.
- Durable survives honest evaluation, including a post-cutoff test.
- Decaying the effect is real and measurably shrinking.
- Unproven the component steps work, but nobody has connected them to returns.
- Illusory the published result is mostly base rate, leakage, or costs nobody netted out.
Analyst narrative
Durable- Mechanism
- Not the rating and not the target price. The prose of the report, especially the forward-looking section, carries information that the headline numbers throw away.
- Evidence
- 1.19 million reports, S&P 1500, 2000 to 2023, run through a fixed open-source embedding model into a simple ridge regression. Long-short 1.04%/month, five-factor plus momentum alpha 68 bp (t = 2.64), information ratio 1.23 against 94 fundamental and 18 analyst factors. The "Strategic Outlook" section is 15% of the words and 41% of the Sharpe.
- The kicker
- Rating, EPS and target-price revisions all go insignificant once the text is included. And sorting on report sentiment earns nothing at all. Tone is not the channel, which means almost every sentiment product is reading the wrong thing.
- Leakage test
- Passes, and this is rare. Re-run on chronologically frozen models the signal strengthens, which is the opposite of what memorisation produces.
- Practicalities
- Horizon 6 to 24 months, peaking at 12. Turnover 28%/month, still 0.58 to 0.95%/month net at realistic costs. Works better on large, mature names, which is the opposite of the neglected-small-cap folklore.
The overnight language window
Durable- Mechanism
- Numbers are arbitraged in seconds. Language takes until the next morning. The gap between those two speeds is the trade.
- Evidence
- 5,428 earnings events, 2022 to 2025. The numeric EPS surprise has an information coefficient of 0.03 at the print and −0.02 by the next open, meaning it is gone. Transcript sentiment measured at the next open: IC 0.11 (t = 4.04), top-minus-bottom quintile 153 bp (t = 4.30), Sharpe 2.28. Correlation with the EPS surprise is only 0.21, so it is nearly orthogonal to the thing everyone already trades.
- Agent's role
- Read the full transcript in the hours the market cannot, and be ready at the open. This is a breadth problem, not a cleverness problem.
- Caution
- Anything faster than this is gone. A related after-hours edge worth 0.72% per trade in 2011 to 2015 was insignificant net of spreads by 2016 to 2020, and a five-second delay eliminates it entirely.
Uncertainty and evasion
Durable for volatility- Mechanism
- Hedging, non-answers and vocal instability are monotone in uncertainty and carry no sign. They tell you the distribution is widening, not which way.
- Evidence
- Across 1,795 calls, acoustic and linguistic uncertainty explains up to 43.8% of out-of-sample variance in 30-day realised volatility, and the same paper explicitly fails to forecast direction. LLM-extracted non-answers predict higher forecast error, wider dispersion, larger drift and wider spreads.
- Agent's role
- Segment the call and score the Q&A separately, which is where the signal lives. Spontaneous speech beats scripted speech, the most replicated result in this literature.
- Fleet note
- A 4B fine-tune scores 84.9% macro-F1 on evasion classification, beating Claude 4.5, GPT-5.2 and Gemini 3 Flash. This is commodity classification now, and it runs on hardware you own.
- How it is faked
- By selling a sign-free signal as a directional one. Most blowups here come from mapping "management is evasive" onto a long-short book instead of onto a volatility position.
8-K sub-item classification
Durable- Mechanism
- The SEC's own item codes are weak labels. Materially different events share a code, and the most reactive ones hide in a catch-all.
- Evidence
- 292,984 filings, 2022 to 2026, reclassified into 119 event types. Precision rises from 12% to 96% with quality gating. A CEO departure scores 1.32 against a routine appointment at 1.06, both filed under Item 5.02. Nine of the fifteen most reactive event types sit inside catch-all Item 8.01.
- Agent's role
- This is the single cleanest "only a language model can do this" result in the disclosure space. Regex cannot separate two events that share a code.
- Caution
- Selection bias is severe. Only 9.84% of impairment firm-quarters ever produce an 8-K at all, and the drift is conditional on investor attention.
Disagreement, decomposed
Durable- Mechanism
- Disagreement is two different things wearing one name. Fundamental disagreement, where people dispute the business, predicts returns. Noise disagreement, where people are just loud, does not.
- Evidence
- Split by a language model, fundamental disagreement earns −0.7%/month while the noise component earns zero. Pooled together, as every dictionary and every aggregate measure does, the signal points the wrong way. It is also nearly orthogonal to numeric forecast dispersion, so it is not the crowded version of the same thing.
- Why it matters
- This is the template for where agents genuinely add value: not scoring a document better, but separating two things that were always conflated because separating them required reading.
Estimate revision momentum
Decaying- Where it stands
- Six-month EPS revision still earns +62 bp/month (t = 2.39) over 2015 to 2024, but with a momentum beta of 0.41, so you are partly buying momentum. Plain analyst revision earns +8 bp (t = 0.57) and is fully explained by momentum. Forecast disparity is dead.
- Wider picture
- Across roughly 200 documented anomalies, the median long-short fell from 48 bp/month before 2005 to 7 bp after, excluding microcaps, with a median t-statistic of 0.45. The authors note that modest allowances for luck or costs would have eliminated even those.
Disclosure language change
Decaying- Where it stands
- The original finding, that companies changing their filing language underperform, reported 34 to 58 bp/month. The widely-quoted 188 bp figure is one subsample, not the headline. There is no dedicated replication and no out-of-sample test, and it does not appear in the main open-source anomaly library at all.
- What the LLM adds
- A real improvement over keyword methods: semantic change detection earns −0.52%/month (t = −2.55) where an entity-extraction baseline earns −0.14% and is insignificant. The mechanism is qualifier preservation. Phrases like "roughly in line with" and "at the low end of the prior range" are exactly what regex discards and exactly where guidance changes hide.
- Crowding
- Researchers matched EDGAR access logs to institutional holdings and found that bulk filing downloaders trade on this and earn excess returns. You are not early.
Supply-chain and peer graphs
Unproven- The gap
- The most interesting negative here. The extraction literature and the alpha literature do not overlap. Every paper that builds a firm-link graph with a language model at scale reports F1 and coverage and runs no return backtest. Every paper reporting alpha on a link graph uses a hand-built, vendor, or embedding-similarity graph. The obvious experiment has not been run by anyone.
- Extraction
- Genuinely good. 170 million web pages, 21,000 screened firms, binary link detection F1 0.784, validated against OECD trade tables at r = 0.64. Against a major commercial dataset, only 22% of edges overlap, so 78% of inferred links are absent from what you can buy. That is either real new coverage or unbounded false positives, and the precision figure was measured on a human-annotated sample rather than on the disputed complement.
- Nearest alpha
- Does not support the thesis on inspection. It reports 7.27%/yr but on a hand-built 48-node graph, using a small encoder rather than a generative model, with a negative cross-sectional coefficient and the short leg doing the work. A separate clean test of model-derived peer similarity against merger announcement returns found a rank correlation of +0.07 (p = 0.37), with intervals tight enough to rule out moderate effects.
- Failure mode
- Everyone optimises the wrong metric. Extraction runs at roughly 0.70 recall and every paper tunes for precision, but for a lead-lag signal a missing edge is the expensive error. Underwrite well below the replicated 0.79%/month for customer momentum, not the original 1.58%.
Statement analysis and XBRL
Unproven- The retraction
- The paper everyone cites for "language models beat analysts at predicting earnings direction", reporting 60.35% against 52.71% and a Sharpe of 3.36, has been withdrawn, on the authors' own finding of data and analysis inconsistencies. A companion paper was withdrawn outright. If you have seen that statistic quoted this year, it was quoted from retracted work.
- What is solid
- Extraction. Fine-tuning lifts value extraction from 52.5 to 98.5 and formula calculation from 27.3 to 98.7.
- What is not
- Everything downstream. On full tagging, the best model scores F1 0.693 on extraction but 0.189 on linking a value to the right accounting concept, for an end-to-end score of 0.066. And no paper tests whether structured-data signals beat text signals in returns: the one study that bridged them is the retracted one.
Insider and political disclosure
Illusory- Insiders
- The classic opportunistic-insider result was 82 bp/month. Recent work finds negative dollar returns once you cap position size per signal, which is what any real book does. A separate study finds a 6.3% mean abnormal return that exists only in microcaps.
- The reform
- After the 2023 trading-plan rules, insider sales within 90 days of plan adoption fell from 35% to under 3%, and migrated into the 90 to 120 day window, where they rose from 11% to 30% and where abnormal returns increased from roughly $23M to $89M a year. The regulation relocated the behaviour rather than removing it.
- Congress
- Across 17,859 disclosed events from 2020 to 2025, post-disclosure purchases returned +10.95% against the index's +19.73%. The buy-minus-sell spread is −0.02 points with a t-statistic of −0.16. Only 32.2% of 311 tracked portfolios beat the index. The popular tracking fund beat its benchmark by 12 points since inception, essentially all of it earned in 2023.
Fraud and restatement prediction
Illusory- Base rate
- Fraud runs at roughly 0.7% of firm-years. The best-known machine-learning result had its top-percentile precision cut from 4.5% to 2.5% under scrutiny, below a much older logistic model. That is roughly nineteen false flags in every twenty.
- The split
- One 2026 study reports AUC 0.96 and F1 0.76 on a random split, then AUC 0.68, F1 0.14 and precision 0.09 on a company-isolated split. The model had been recognising companies, not fraud. On a separate benchmark every language model tested loses to a random forest.
The options premium is still there. It stopped being alpha in 2012.
The usual question is when to enter an option. The honest answer reframes it, because the most important options finding of the last two years is not about timing at all.
Implied volatility still exceeds subsequent realised volatility almost all the time. Over 2024 to 2026 the spread averaged 3.83 volatility points and was positive on 85.6% of days. In 2026 so far it has averaged 5.41 points, the richest since 2021, and was positive on 95.4% of days. In March 2026 the index traded at an average implied 25.6 while realised volatility sat between 11.4 and 14.7 for the entire month, a gap of eleven to sixteen points every single session.
And yet every Cboe option-writing index underperformed the S&P 500 total return in 2024, in 2025, and again in 2026.
A Chicago Fed study resolves the paradox. Option alphas have been statistically indistinguishable from zero for about fifteen years, with a structural break dated to May 2012. The mechanism they identify is that dealers' net index gamma flipped from negative to neutral or positive as everyone else gained frictionless access to selling volatility. In the authors' own words the surviving premium is no larger than what the position's market beta would predict.
So you do still collect the spread. You are simply being paid for equity beta rather than for bearing variance risk, and you should size the position as the beta position it actually is.
One day, August 2024
Short-volatility loss during a +23% equity year
Six sessions, April 2025
The premium then went negative for a month
March 2026
Same position, same decade, third time
In April 2025 realised volatility exceeded implied continuously for a month, bottoming at 26.4 points the wrong way, with realised running at 61% annualised. The spread being positive 95% of the time is not the risk. The other 5% is.
The earnings-volatility trade inverted
The classic retail move, selling a straddle into earnings because implied always overshoots, has stopped working and then reversed. The ratio of realised to implied earnings moves went from 85% in 2019 to 96% in 2021 to 104% in 2024. Straddle buyers averaged a 2% loss over the trailing twelve quarters, then made 45% in January and February 2026 and 23% in the first week of the second quarter, concentrated in mega-cap technology names.
There is no published Sharpe ratio for this trade in either direction. The one real backtest is filtered down to 4% of announcements and carries 3.2 to 1 tail asymmetry, with a worst case of −26.7% against a best case of +8.3%. Treat anyone quoting a clean edge here with suspicion, and note that the only continuous public series on implied versus realised earnings moves comes from a vendor publishing on its own unreplicated data.
Where premium genuinely remains
| Surface | State | Evidence |
|---|---|---|
| Index correlation structure | Rich | 1-month implied correlation averaged 12.89 in 2026 against 41.48 across 2010–19; the single-stock minus index spread hit a record 34.14 points in July 2026 |
| Oil volatility | Rich | Highest premium on record over twenty years, March 2026 |
| Gold volatility | Rich | 98th percentile implied-minus-realised spread |
| VIX futures roll-down | Decayed | Return-to-volatility fell 0.67 to 0.16 versus the 2010s, on 29 days of backwardation in 2025 |
| Single-name earnings vol | Inverted | Realised-to-implied 85% to 104% |
| Index vol outright | No alpha | Zero since 2012; you are paid beta |
That table is the actual answer to "where do I find an edge in options". It is a question of which volatility you sell, not when you sell it. The premium has migrated out of blanket index selling and single-name earnings, and into the correlation structure and commodities.
Verify the tape before you trust a level
The famous VIX print of 65.73 in August 2024 was not a tradable number. Spot printed 70% above its own prior close while the exchange-traded product holding actual VIX futures printed only 4.3% above its close. In a genuine volatility event the two agree closely: in April 2025 the ratios were 1.099 and 1.068. The exchange's own file records that day's open and low as the previous Friday's close, carried forward. Anyone backtesting an entry rule off 65.73 is trading a price that never existed.
Six tests, and why passing five is still failing
A review of 164 papers from 2023 to 2025 found that no single methodological bias is discussed in more than 28% of them. That is the state of the field, and it is why the checklist below matters more than any signal above. Each item exists because a published result died on it.
- Count the searches, not the winner. Keep a trial ledger of every evaluation you run, and deflate the Sharpe ratio against that number. An agent's best find went from a design Sharpe of 1.69 to a deflated 0.86 against its own 102 trials, and to 0.18 out of sample.
- Test leakage separately, because deflation will not catch it. In the same study a deliberately future-peeking strategy with a Sharpe above 34 passed deflation with a perfect score. Multiple-testing correction and leakage are different diseases.
- Split by entity, not at random. A fraud model scoring AUC 0.96 on a random split scored 0.68 with precision 0.09 once no company appeared in both halves. It had learned companies, not fraud.
- Test after the model's knowledge cutoff, or freeze the model. Roughly 32% of the headline news-to-return effect is memorisation. Lookahead propensity is materially positive inside the training window and falls to essentially zero after it. Subtracting memorised knowledge cuts reported in-sample returns by nearly half on average.
- Attribute before you celebrate. Mask the tickers and dates and run a factor decomposition. Nine of ten agents that reported strong returns had negative stock-selection alpha once market and style exposure came out.
- Net out costs, borrow, and latency. Across 162 anomalies the average was 0.14% a month before costs and −0.01% after borrow fees. An after-hours edge worth 0.72% per trade dies on a five-second delay.
The metric trap underneath all of it
Directional accuracy flatters any model that predicts the majority class. Use balanced accuracy, per-class precision and recall, and Matthews correlation. And when a signal is sign-free, like evasion or acoustic stress, score it against volatility rather than forcing it into a long-short book.
What the evidence says to actually build
The deployment pattern at funds that are getting value is consistent and unglamorous. Research automation, document diffing, coverage expansion and drafting are in production. Autonomous capital allocation is not, outside small dedicated sleeves.
| Firm | What they run | Reported effect |
|---|---|---|
| Balyasny | Firm-wide assistant, micro-agents flagging filing wording changes across ~5M documents | Senior-analyst task 2 days to 30 minutes |
| Bridgewater | ~17-person AI lab, ~$2bn fund since 2024, three-layer validation chain | Error rate 8% to 1.6% |
| Man Group | Assistant drafts trade rationales; explicitly augmentation | ~40% monthly usage |
| Avala Global | Agents for equity research coverage | 20 to 200 names per analyst |
| D.E. Shaw | Gateway with PII stripping and per-desk cost meters | Throttled by budget |
Small models do the bulk of it
This is the part most relevant to running your own hardware, and the evidence is unusually clear. A 4B fine-tune beats three frontier models at evasion classification. A Qwen3-4B reasoning model trained with reinforcement learning beat GPT-4.1 on its test window and runs locally. Fine-tuning lifts financial extraction tasks by 36% on average and takes value extraction from 52.5 to 98.5.
Against that, the production study found three frontier models spanning one basis point of decision-quality difference, all intervals overlapping, at a twenty-five-fold cost spread. Paying frontier prices for classification is simply a mistake.
Local models for extraction, diffing, classification, routing and scoring. Frontier models only for multi-hop synthesis. A full agent pass costs one to ten dollars per name per day at frontier prices, which is trivial at twenty names and ruinous at a thousand.
The operating layer beats the prompt
The single most transferable finding here is that behaviour lives in the scaffolding, not the model. In production, a risk slider moved leverage by 0.425 per level, agent identity absorbed 60% of behavioural variance, and where a leaderboard cut off at the top three, selection jumped 1.75 times for no reason but the render. A mechanical stop and target attached at entry earned 39 basis points per position. One slider setting holding 11% of the book produced 62% of the liquidations.
Every attempt to fix those behaviours with better prompting underperformed a one-line change to the order path. If you build anything here, spend your effort on the tool registry, the candidate list, and the risk mechanics, and treat the model as the cheapest interchangeable part.
Seven experiments nobody has run
These came out of the research as explicit gaps, which makes them the most interesting places to work. Each is a real paper that does not exist.
- Customer momentum on a model-built supply graph. Extraction is proven at F1 0.78 over 170 million pages. Alpha is proven on hand-built graphs. Nobody has connected the two, and 78% of model-inferred links are absent from commercial datasets.
- Recall-optimised extraction for lead-lag. Everyone tunes precision. For a supply-chain signal the missing edge is the expensive error, and no one has built for that.
- Report novelty and boilerplate as a traded signal. Scoring how much of an analyst report is genuinely new rather than recycled. No paper found.
- Stealth revisions. Cases where the model or assumptions change materially while the rating and target price stay put. No paper found.
- Reason-conditioned revisions. Trading on why an analyst changed their mind rather than that they did. No paper found.
- Structured data against text in returns. Whether XBRL-derived signals add anything over text signals. The one paper that bridged them has been retracted.
- The 2024 beneficial-ownership deadline. The 13D filing window was cut from ten days to five and nobody has published a measurement of the effect.
If you want one project, take the first
It has proven components at both ends, a clearly stated gap in the literature, a defined failure mode to design against, and it runs on extraction work that small local models do well. It is also the kind of result that is publishable whichever way it comes out.
What I would not lean on
Six independent investigations fed this piece and two of them disagreed, which is worth surfacing rather than smoothing over. The supply-chain result reported as supporting evidence in one pass turned out, on closer reading in another, to rest on a hand-built graph with a small encoder and a negative cross-sectional coefficient. I have reported the second reading.
- The widely-quoted claim that signal half-lives have compressed from five to seven years down to eighteen months comes from a model-derived estimate with a simulated convergence figure, not a direct measurement.
- The earnings-volatility series is published by a vendor on its own unreplicated data and is the only continuous public source that exists.
- The microcap capacity figures could not be verified against a 2025 or 2026 source, despite being load-bearing for the argument that small accounts have room where institutions do not.
- There is a genuine counter-current worth knowing. One line of work argues that reinforcement learning agents can sustain collusive, supra-competitive profits that reduce price efficiency, which would mean AI slows information absorption rather than speeding it. That cuts against the entire crowding narrative, and it is not settled.
Method: six parallel research investigations run on 14 September 2026, covering analyst research, corporate disclosure, earnings communication, options and volatility, agentic architectures, and alternative data. Each was instructed to prioritise 2025 and 2026 evidence and to flag older work. Claims are attributed to their strongest available source; where two investigations disagreed, both readings are reported above.
About the author
Pragadeesh VS works on AI software, media and investing at Serverlessvc.com.
Not investment advice. Every figure here is a research finding, most are pre-cost, and the central argument is that published results in this field systematically overstate what survives contact with a real book.