In July 2026, host Corey Hoffstein asked Ben Wellington, Head of Complex Feature Engines at Two Sigma, a direct question: if large language models make feature construction and data exploration dramatically easier, will the technical advantages that firms spent years building be competed away?
Wellington’s answer can be summarised simply: a dataset is not a moat. The hard part is what happens next—knowing which question to ask, deciding whether a result is real or accidental, understanding whether it adds anything to an existing portfolio, and connecting it to a research system that works in live markets.
AI has not made alpha easy to find. It has made plausible-looking alpha exceptionally easy to produce. In doing so, it is accelerating two forces at once: alpha discovery and alpha decay.
01 / AI finds alpha—and accelerates its decay
The most immediate effect of AI on quantitative research is the falling cost of an experiment. Work that once took days to formulate, implement, and backtest can now be divided among specialised agents and completed within hours.
But a shorter experiment cycle does not create market inefficiencies at the same rate. Once a pattern becomes widely traded, capital pushes prices toward a new equilibrium and the excess return fades.
This dynamic predates AI. Earlier research found that some market anomalies weakened after publication, plausibly because investors began trading on them. AI simply turns the same machine faster.
The first opportunities to be competed away are likely to be those that are easy to describe, easy to copy, and easy to find in public data using standard workflows. AI erodes the rent from information-processing speed first. What becomes more valuable next is question selection, system design, and governance.
A July 2026 report from the CFA Institute Research and Policy Center reaches a similar conclusion. AI can accelerate information processing and price discovery, potentially shortening traditional arbitrage windows. Yet institutions that depend on similar models, data, and decision frameworks may generate similar signals and act at similar times, reducing the diversity of market views and amplifying correlations under stress.
The subtler risk is that every firm produces more signals while the industry collectively explores fewer genuinely distinct ideas.
Wellington describes this as a reduction in the entropy of research output. Consistency is progress in manufacturing; in portfolio construction, it can be a risk. A thousand models with similar data, assumptions, and exposures are not a thousand independent sources of alpha. A portfolio needs models that fail in different ways under different conditions.
This is why orthogonality is the lifeblood of alpha modelling. Researchers with different training, experience, and instincts can take the same dataset in very different directions. A one-click AI research tool may flatten those differences before they have a chance to compound.
Man AHL’s 2026 AlphaTrend experiment makes the point concrete. Given the same research task, different foundation models displayed different research personalities: some clustered around similar explanations, while others produced a broader range of signal families.
Diversity in AI-native research does not come from asking the same model the same question a hundred times. It must be designed into the combination of models, agent roles, context, research prompts, and evaluation criteria.
For industrial production, consistency is efficiency. For alpha production, consistency can be risk. A good AI tool should make each researcher more themselves—not make every researcher converge on the same answer.
02 / Hypotheses can scale. Evidence cannot.
If the old bottleneck in quantitative research was a shortage of candidate ideas, AI may create the opposite problem: an excess of plausible results.
Financial markets have an exceptionally low signal-to-noise ratio. Test enough variables, parameters, and sample windows, and eventually something will produce an impressive historical curve. Expanding the search space also expands the probability of finding noise by accident. Ten times more research output does not mean ten times more valid alpha.
The deeper issue is an asymmetry. Hypotheses can be generated in parallel; evidence cannot. There is only one realised market history, and genuinely out-of-sample time still arrives one day at a time. Ten thousand candidates can appear overnight, but compute cannot create ten thousand independent futures against which to test them.

Compute scales the supply of answers. It does not scale the supply of truth.
This quantity trap long predates generative AI. A large-scale replication of hundreds of market anomalies found that many lost statistical significance under stricter samples and standards. Research on the factor zoo likewise shows that a new factor can appear strong in isolation while adding little explanatory power beyond existing factors.
By 2026, the question is no longer whether an LLM can generate a factor. It is whether generation can become a research capability that deserves trust. AlphaBench, presented at ICLR 2026, shows meaningful potential in factor generation, evaluation, and search while also highlighting persistent problems in robustness, efficiency, and practical usability. The foundation model is raw material; the research system determines whether its output can be believed.
Man AHL’s AlphaTrend experiment offers another useful example. Researchers gave the system not only promising ideas but also hypotheses expected to fail and cases with uncertain outcomes. Its value was not limited to finding better-performing signals. It also learned to rule out dead ends and identify effects confined to particular market regimes.
In quantitative research, a credible no can be as valuable as a new factor.
Scarcity therefore shifts from candidate generation to selection and validation. Humans spend less time completing every experiment and more time defining worthwhile questions, setting risk boundaries, and deciding what evidence remains insufficient for real capital.
The moat moves with that shift. It is no longer only a secret factor or a more advanced model. It becomes a research constitution: which questions deserve resources, which results must be rejected, what evidence is strong enough for deployment, how failures re-enter the next research cycle, and where human accountability must remain.
Public models determine how many possibilities machines can propose. Proprietary evaluation and feedback systems determine which possibilities are allowed to become assets.
03 / More agents, more alpha—or more of the same?
In 2025, the industry was still asking whether AI could become a quantitative researcher. By 2026, the more practical question is what happens when dozens or hundreds of AI researchers operate simultaneously. How do institutions stop them from proposing the same hypothesis, believing the same mistake, and making the same trade at the same time?
This is not only a model problem. It is an organisational problem. If every agent reads similar material, calls similar tools, and is scored by the same backtest metric, multi-agent research may be little more than parallel computation. The system moves faster while exploring the same small part of the map.
Parallelism is not plurality. Convergence often begins before factor generation. If every institution starts with the same public reports, factor libraries, and prompt templates, even a stronger model may only move faster across the same terrain.
Models are good at expanding a question once it has been defined. They are less reliable at recognising which neglected market behaviour deserves to become a question in the first place. An anomaly observed in trading, a shift in an industry supply chain, or an intuition imported from another discipline may be where differentiation begins.
When answers become cheap, good questions become expensive. Non-consensus alpha starts not with a more complicated formula, but with a question the crowd has not learned to ask.
Different questions are not enough; disagreement must survive the automated workflow. Physicists, computer scientists, and traders can notice different relationships in the same data. If every idea is rewritten by one process and filtered through one objective function, those differences can disappear before they matter.
A strong system creates tension between roles: one agent constructs a hypothesis, another tries to falsify it, and others verify that the code actually implements the intended financial logic. One convenient score cannot become the sole arbiter. Once every agent learns to optimise for the same metric, the system can produce superficially different results with nearly identical underlying exposures.
Even preserved disagreement gets research only halfway. Once candidates arrive at scale, institutional quality depends on what gets rejected. When ideas were scarce, teams worried about missing a good one. When ideas are abundant, the greater risk is putting false alpha into production.
A beautiful backtest only earns a candidate an interview. The real interview happens out of sample, in paper trading, under realistic execution rules, and eventually with limited live capital absorbing actual slippage and market impact. A candidate must also explain why it should work, when it should stop working, and what it adds to the existing portfolio.
Models can expand the candidate set. Statistical tests can eliminate some illusions. The market ultimately reveals the constraints the researcher failed to imagine.
But market lessons do not automatically remain inside the institution. Imagine one agent exploring a dead end in March and another walking down the same path in June after changing a few parameters. Both experiments complete. Compute is fully utilised. Yet the institution has learned nothing.
Automation without memory industrialises repeated mistakes.
Research memory is more than an archive of reports and code. It allows the past to change what the system does next. Why did a path fail? Was the problem in the data, implementation, market regime, or original hypothesis? Failure becomes an asset only when those answers alter the next round of search, evaluation, and risk control.
XALPHA frames this as the transition from isolated automation to a closed research loop. It combines external knowledge with feedback from previous cycles, allowing what the system has read, tested, and learned to change the starting point of its next investigation.
CogAlpha approaches the challenge from another direction. Its code-level alpha representation, hierarchical agents, and evolutionary search are designed to create structural diversity during candidate generation. But any research framework becomes an asset-management capability only when it connects to rigorous portfolio construction, execution, risk management, and live-market feedback.
The new moat: questions, disagreement, rejection, and memory
These four gates form the new moat for AI-native quantitative research. Evaluating an institution should not stop at which model it uses, how many agents it deploys, or how many factors it generates. More revealing questions are: where do its research questions originate; can genuinely different paths remain distinct; why are results rejected; will a failed idea return in another form; and how quickly can the system detect, explain, and replace a decaying strategy?
This points to a measure more important than the half-life of alpha: the half-life of error.
Every alpha can decay as market structure and participant behaviour change. A durable institution does not promise that one answer will remain correct forever. It detects failure sooner, understands why its assumptions stopped working, and turns that failure into the starting point for the next research cycle.
AI is shortening the half-life of alpha. It may also shorten the half-life of research error. Once every institution has access to generation, competition will no longer be about whose AI appears smartest. It will be about who can preserve disagreement, reject attractive but false answers, and build a system in which people and machines remember how the market proved them wrong.
References
[1] Flirting with Models podcast, S7E32: Ben Wellington (Two Sigma), “The Future of Features Research,” July 2026.
[2] CFA Institute Research and Policy Center, Artificial Intelligence and the Future of Finance: A Framework for Structural Change, July 2026.
[3] Schwert, G. W., “Anomalies and Market Efficiency,” NBER Working Paper No. 9277.
[4] Ben Wellington, “Anything Can Be Language Now: My Thoughts on the Future of Features Research,” Two Sigma, July 2026.
[5] Man AHL, “AlphaTrend and Agentic Research Workflows,” February 2026.
[6] Hou, K., Xue, C., and Zhang, L., “Replicating Anomalies,” NBER Working Paper No. 23394.
[7] Feng, G., Giglio, S., and Xiu, D., “Taming the Factor Zoo,” NBER Working Paper No. 25481.
[8] Luo, H., Ko, H. T., Chen, J., Sun, D., Zhang, Y., and Liu, C., “AlphaBench: Benchmarking Large Language Models in Formulaic Alpha Factor Mining,” ICLR 2026.
[9] Liu, F., Fu, Y., Wang, Y., and Liu, Q., “XALPHA: A Memory-Driven AI Quant Researcher for Hypothesis-to-Code Alpha Discovery,” arXiv:2607.08332, July 2026.
[10] Liu, F. et al., “Cognitive Alpha Mining via LLM-Driven Code-Based Evolution,” ACL 2026.
