In a recent Bloomberg interview, Man Group’s Head of AI, Tushara Fernando, said the firm’s token usage had increased roughly 86-fold since January. The figure points to a broader change: AI is moving into workflows across the investment firm, from quantitative research and engineering to finance, operations, and human resources.

For investment teams, the first-order effect is a much wider range of usable alternative data. Card transactions, hiring activity, website traffic, vessel movements, and satellite imagery can now be analysed alongside podcasts, expert interviews, PDFs, complex contracts, and fragmented web content.

Models can transcribe, extract, classify, and connect these sources at scale. But the ability to process more information is not the same as the ability to make better investment decisions. The difficult work is still turning information into evidence—and evidence into a differentiated, testable, and tradable view.

From scattered information to investment evidence

Traditional data platforms depend heavily on predefined fields, relationships, and queries. Every new source or entity type can require another layer of cleaning rules, mapping, and engineering. Large language models reduce this upfront cost by allowing podcasts, filings, spreadsheets, and web pages to be handled within a shared natural-language task.

Longer context windows and agent workflows also let researchers follow a question across sources, collect supporting evidence, fill information gaps, and call tools without breaking the chain of inquiry.

Fernando described an investor studying the AI supply chain. An agent transcribed a technical podcast featuring an engineering leader at a major cloud provider. The discussion indicated that GPUs remained scarce, while data-centre capacity—and eventually connectivity between facilities—could become the next constraint.

That observation is not yet an investment thesis. It is evidence that creates better questions: is the bottleneck in chips, power, cooling, physical capacity, or networking? How might it affect orders, revenue, and margins? Which companies could benefit, and how much is already reflected in consensus expectations and prices?

AI did not make the investment decision. It brought evidence that might otherwise have been missed onto the same research surface.

This capability can also change which assets are economically worth researching. Dense prospectuses, non-standard contracts, indicative quotes, and incomplete reference data historically required substantial fixed investment in people and infrastructure. AI can reduce that cost by extracting terms, structuring reference data, and connecting complex instruments to existing research systems.

More data does not automatically mean more alpha

Extraction is only the beginning. A model can read every row in a card-transaction dataset without knowing whether a record represents a completed purchase, a pre-authorisation, or a refund; whether the sample represents the wider market; or whether the vendor later revised the historical data.

If those questions remain unresolved, AI may simply help a team reach the wrong conclusion faster. Before a dataset enters production, four areas deserve close scrutiny.

1. The investment question and economic mechanism

Start with the investment question, not the novelty of the dataset. Is the goal to estimate same-store sales, track market share, monitor pricing, forecast power demand, or assess the pace of data-centre construction? A credible research case should explain what the data measures and how that change could ultimately affect an asset’s value. Correlation without an economic mechanism is rarely enough.

2. Provenance, bias, and point-in-time integrity

Teams need to understand how the data was collected, how the sample was constructed, how coverage changed, and whether a vendor rewrites history or updates entity mappings. Point-in-time integrity is essential: a historical record used in a backtest must reflect what an investor could actually have known on that date.

3. Fitness for the investment horizon

Faster data is not always better data. Minute-level updates can add noise and cost to a strategy with a six-month horizon. Frequency, delivery lag, historical depth, identifiers, mappings, and maintenance requirements all need to match the intended use.

4. Incremental value and decay

A single backtest or Sharpe ratio cannot establish value. Teams need out-of-sample tests, regime analysis, ablation studies, and realistic transaction costs. Exclusive data is not necessarily useful, and widely available data is not necessarily worthless. The relevant question is what it adds to the institution’s existing information—and how quickly that advantage may be copied or arbitraged away.

The evidence-to-decision loop

Alternative data does not move directly from ingestion to signal. A practical research loop contains six connected layers:

1. Evidence: What happened, and can the claim be traced to its original source, timestamp, and data version?

2. State: What do multiple pieces of evidence imply about demand, inventory, capacity, customer validation, or the industry cycle?

3. Mechanism: Why is the change happening, and how might it flow through to prices, orders, revenue, and profit?

4. Expectations: What could happen under different scenarios, and what evidence would invalidate the thesis?

5. Pricing and positioning: Is the information already reflected in consensus, valuations, flows, holdings, and crowdedness?

6. Validation and learning: Did the thesis play out? If not, was the failure in the research, execution, sizing, or risk controls?

AI can support search, extraction, linking, testing, monitoring, and documentation throughout this loop. Humans still define the problem, judge the economic mechanism, set risk, and remain accountable for the final decision.

Three disciplines for an AI-native data process

First, a summary is not a fact. Important claims must remain traceable to the original passage, timestamp, and data version. Models can compress podcasts, reports, and websites, but they can also omit qualifications, confuse speakers, or turn speculation into assertion.

Second, cheaper experimentation makes research discipline more important. AI can generate a large number of features quickly, while also magnifying multiple-testing errors and overfitting. Automation expands experimental capacity; it does not guarantee better conclusions.

Third, AI costs should be assessed against decision value without requiring every exploratory call to show an immediate return. Better measures include broader research coverage, faster data onboarding and validation, fewer errors, and measurable improvements in mature workflows.

The real moat is institutional learning

Models can be purchased, and competitors may buy the same datasets. What is harder to replicate is the institutional memory built around them: historical versions, failed experiments, vendor records, entity mappings, validation standards, and post-investment reviews.

An AI-native investment firm is therefore not defined by how many models or agents it deploys. It is defined by how effectively each research project improves the next one. The strongest organisations will explore broadly, standardise what works, retire what does not, and preserve a rigorous link from evidence to decision.

AI helps an institution see more. Its research system determines whether that information becomes insight—or noise.

Sources

Bloomberg Odd Lots, “One of the World’s Largest Hedge Funds on Its 86x Growth in Token Spending.”

The Alternative Data Podcast, “The Natalya Dmitriyeva Episode.”

Neudata New York Summer Data Summit 2026, Natalya Dmitriyeva.

Read the original article on LinkedIn ↗︎