Nyquist Nyquist
CLOSED BETA · 2026 Start a 60-day pilot
Nyquist Research · No. 01 · Field Note

AI agents in trading: where the real boundary of applicability lies.

A practitioner-grade calibration of where LLM-based signal generation actually adds edge, where it destructively competes with classical factor methods, and what the infrastructure layer looks like when it all breaks at 9:31 AM on a CPI print day. The boundary is real, empirically measurable, and not where the vendor demos suggest it is.

Published
11 May 2026
Reading time
22 min
Author
Nyquist Research
Topic
Quant · LLM Signals · Factor Models · Infrastructure

Everyone is deploying AI agents in trading. Almost nobody is rigorously measuring whether they actually work — beyond the backtest, beyond the demo environment, beyond the carefully curated in-sample period that happened to include exactly the regime the model was trained on. The gap between "our LLM agent showed 25% annualized returns in simulation" and "our LLM agent survived live deployment for 18 consecutive months without cascading infrastructure failures, adverse selection, or silent signal degradation" is not a model problem. It is a systems problem — and it is precisely where most teams quietly fail, without ever publishing a post-mortem.

This is not a defense of classical factor models, nor an endorsement of the AI-agent-replaces-quant narrative that has saturated fintech discourse since late 2023. It is an attempt at a practitioner-grade calibration: where does LLM-based signal generation actually add edge, where does it destructively compete with better-established methods, and what does the infrastructure layer look like when it all breaks at 9:31 AM on a CPI print day.

01 — Two paradigms

Two paradigms, one goal: alpha generation.

The systematic equity space has operated for decades on the foundation of classical factor models. The Fama–French framework — three-factor, extended to five in 2015 with profitability (RMW) and investment (CMA) — remains the dominant risk decomposition and return-prediction paradigm. Layered on top: momentum (12-1 month return), mean-reversion overlays at shorter horizons, and statistical arbitrage pairs from cointegration or common-factor residuals. The virtues are well-documented and empirically robust across decades, asset classes, and geographies: interpretability, stable signal turnover, manageable transaction costs, and a research lineage dense with peer-reviewed scrutiny.

The framework rests on a deceptively simple foundation. The cross-sectional expected return of asset i is modelled as a residual idiosyncratic alpha αi plus the sum of systematic factor exposures (value, size, profitability, investment, momentum) times their loadings. The entire art of quant portfolio construction lies in estimating that residual alpha reliably enough that, aggregated across a diversified universe, it generates a positive and persistent information ratio.

LLM/AI-agent signal generation is built on a fundamentally different premise: that the information content of unstructured text — earnings call transcripts, 10-K and 10-Q filings, central bank communications, analyst reports, macro press releases — is systematically underpriced, and that large language models are the appropriate tool to extract that discrepancy at scale. The practical architecture varies: fine-tuned BERT-class models (FinBERT) generating cross-sectional sentiment scores; GPT-4-class models with retrieval-augmented generation producing multi-step trade ideas; or fully autonomous execution loops where an LLM orchestrates ingestion, signal, construction, and routing with minimal human oversight.

The appeal is genuine. FinBERT-class models can process upwards of 1,000 earnings call transcripts in the time a senior analyst team reads three; the marginal cost of expanding coverage from 500 to 5,000 names is near-zero once the pipeline is built. The risk is equally genuine: signal-to-noise in NLP-derived sentiment is structurally lower than factor signals, the turnover profile is incompatible with execution costs in most institutional contexts, and the models introduce a new failure mode — look-ahead bias embedded in pre-training corpora — that standard backtesting frameworks are poorly equipped to detect.

02 — The setup

Inputs & assumptions.

Before any empirical comparison, the experimental conditions must be stated precisely. Without this, any IC or Sharpe figure is uninterpretable.

ConditionSpecification
DataUS equities, daily OHLCV plus standardized fundamentals — EBITDA margins, book-to-market, operating profitability, asset growth / investment rate.
PeriodJan 2019 – Dec 2024. Two distinct volatility regimes (COVID 2020–21, rate normalization 2022) plus the 2023–24 AI rally.
UniverseTop 500 names by trailing 63-day average daily dollar volume, rebalanced quarterly.
BenchmarkEqual-weighted, rebalanced monthly — chosen over cap-weighted to avoid mega-cap concentration distorting evaluation.
LLM proxyFinBERT sentiment from quarterly earnings calls, normalized cross-sectionally by GICS sector, refreshed at transcript release.
Factor modelFama–French 5-factor (Mkt-RF, SMB, HML, RMW, CMA) + 12-1 momentum (UMD), rolling 36-month OLS, rebalanced monthly.
Transaction costsNot applied in base results. Sensitivity discussed separately.
The honest caveat

All headline metrics below are gross — pre-transaction-cost. This is the single most important adjustment academic benchmarks fail to make. A factor model with 1.3 gross Sharpe and 4× turnover likely retains 1.0–1.1 net; an LLM signal with 0.8 gross Sharpe and 12× turnover likely delivers 0.2–0.4 net, if anything.

03 — The evidence

Empirical comparison: the numbers that matter.

Signal propertyA · Factor modelB · LLM agentDelta
Information Coefficient0.05 – 0.070.02 – 0.04danger zone
Information Ratio (ICIR)0.6 – 1.10.2 – 0.5fails ICIR > 0.5
Hit rate52 – 55%49 – 53%marginal vs coin-flip
Annual turnover3 – 5×8 – 15×cost-drag eats edge
Max drawdown (gross)−8% to −12%−14% to −22%regime fragility
Sharpe (gross)1.1 – 1.50.6 – 1.0net → 1.0–1.1 vs 0.2–0.4

For the Fama–French + momentum composite, an IC of 0.05–0.07 is consistent with the literature: it aggregates multiple individually weak but low-correlation sub-signals — value ≈ 0.02–0.03, momentum ≈ 0.04–0.05, profitability ≈ 0.02–0.03 — whose combination benefits from diversification across signal types and holding periods. The LLM proxy IC of 0.02–0.04 is statistically non-trivial but sits in a danger zone: at this level, even a modest increase in transaction costs flips the strategy from modestly profitable to net-negative. The practitioner threshold — "IC below 0.03 with annual turnover above 8× is almost certainly net-negative after costs" — applies directly.

There is a deeper issue: the IC of LLM sentiment is highly non-stationary. During earnings seasons, realized IC spikes to 0.06–0.08 in individual months; between reporting windows it collapses toward zero or turns negative as the model operates on stale transcripts in a market that has moved on. This episodic structure creates a false impression of strength in aggregate statistics while masking the operational reality that the signal is silent for 8–10 weeks of every 13-week quarter.

The Information Ratio — mean IC over its standard deviation — tells the same story. The factor IR of 0.6–1.1 reflects both magnitude and stability; a standard quant PM threshold of ICIR > 0.5 for meaningful allocation is satisfied by classical factors and failed by the LLM proxy in most parameterizations. Hit rates of 52–55% for factors are modest but sufficient when combined with favourable skew; the LLM hit rate of 49–53% is at best marginally better than coin-flip, and during the 2022 rate shock it fell below 50% — sentiment captures market narrative efficiently, but narrative and subsequent returns are only loosely correlated at earnings-to-earnings horizons.

The turnover differential — 3–5× versus 8–15× — is the most consequential metric in the table, and the one most frequently glossed over. LLM turnover is high for a structural reason: the signal is document-driven, so every new transcript or revised guidance can flip the cross-sectional ranking, whereas a value or profitability signal moves only on quarterly restatement.

At 10× annual turnover on a 500-name universe, a realistic 5–10 bps one-way execution cost translates to 50–100 bps of annual cost drag — 0.3–0.6 Sharpe units — which eliminates the gross edge entirely in the median parameterization.
Page-one item in any internal review of an LLM signal strategy

The wider drawdown range for LLM agents (−14% to −22% versus −8% to −12%) reflects two dynamics. First, regime sensitivity: sentiment models trained predominantly on post-2009 bull-market transcripts are underweight the language of genuine distress — in March 2020, FinBERT-derived sentiment gave almost no advance warning, because management teams are coached to maintain confident language on calls. Second, narrative crowding: when a public-model signal becomes widespread, the more crowded the factor, the more violent the de-crowding event, because every participant faces the same exit signal simultaneously.

04 — The other side

Where AI agents actually win.

Taken in isolation, the comparison above would suggest LLM approaches are uniformly inferior. That conclusion is too strong and would be intellectually dishonest. There are specific, well-defined domains where LLM capabilities provide genuine and durable edge — and they are not the cross-sectional equity signal generation where the vendor demos concentrate.

  1. Alternative-data parsing at scale. A human team of ten cannot systematically track linguistic patterns across 2,000 earnings calls a quarter; FinBERT-class models do it in hours, extracting features consistently predictive of subsequent price behaviour — hedging density in forward guidance, sentiment divergence between prepared remarks and off-script Q&A, vague-language frequency. Strongest as an overlay on fundamental factors, not a replacement: a top-quintile value/profitability name with a positive shift in management tone is a more compelling long than fundamentals alone. The IMF and Bank of England both confirm structured NLP of corporate and central-bank text adds predictive content beyond standard economic variables.
  2. Regime detection and narrative shift. Arguably the highest-conviction application. Classical factors rank stocks cross-sectionally but are agnostic about the macro regime in which that ranking is evaluated — a value factor that worked 2000–06 and 2016–22 failed categorically 2010–15; the factor did not change, the regime did. The BIS 2025 paper Parsing the Pulse shows LLMs can decompose macro sentiment across growth, inflation, financial stability, and monetary policy — distinguishing "inflation from supply shock" from "inflation from demand overheating", two narratives with very different factor implications, with 4–6 weeks of lead time over hard data.
  3. Multi-asset macro synthesis. The connections between asset classes are mediated through narrative — Fed language shifts the curve, which shifts spreads, which shifts equity sectors, which shifts commodity demand. An LLM ingesting FOMC minutes, Treasury auction commentary, OPEC+/IEA supply reports, prime-broker color, and FX positioning constructs a cross-asset signal no pure factor model can replicate. Crucially, these signals have longer natural holding periods, which makes the turnover and cost problem far less severe.
  4. Autonomous research orchestration. The most durable application may not involve direct signal generation at all: hypothesis generation, pipeline construction, backtest synthesis, walk-forward validation, and systematic flagging of look-ahead and data-mining bias. The primary failure mode in quant research is not lack of ideas — it is lack of rigorous validation bandwidth. That is the bottleneck LLMs actually address.
05 — A failure unique to LLMs

The look-ahead bias problem.

One additional failure mode deserves its own section because it is specific to LLM-based approaches and has no analogue in classical factor research: look-ahead bias embedded in pre-training corpora. Standard backtesting assumes inputs are point-in-time correct — that the data available at each historical decision point reflects only information actually available then. Most data vendors provide point-in-time databases precisely for this.

LLMs trained on internet-scale corpora break that assumption. A model trained on data through December 2024 has, embedded in its weights, information about which companies performed well throughout 2019–2024. When used to generate sentiment scores in a backtest covering that same period, the resulting figures are biased in ways extremely difficult to detect and quantify.

The sobering number

Work at NeurIPS 2025 and on SSRN formalizing look-ahead benchmarks finds that models evaluated without correcting for pre-training contamination show 15–30% higher gross Sharpe ratios than on a strict holdout post-dating the training cutoff. Any backtest within the model's pre-training window is structurally compromised. The only reliable evaluation is live out-of-sample.

06 — Decay

The alpha decay acceleration problem.

Classical factor alpha decay — the shortening of factor holding periods as more capital pursues the same signals — has been structural since at least 2010, when factor ETFs began compressing premia. The half-life of a documented anomaly, from publication to near-complete arbitrage, has shortened from about 10 years in the 1990s to 4–5 years by the 2020s.

01
1990s factor anomaly

~10 years to arbitrage. The original pace.

A documented anomaly could persist roughly a decade before the published premium was substantially competed away.

02
2010s factor — ETF era

4–5 years. Capital arrives faster.

Factor ETFs compressed the diffusion-to-arbitrage cycle to roughly half a decade across value, size, and momentum.

03
LLM signal — 2020s

1–2 years. And accelerating.

FinBERT is open-source, the methodology is in dozens of papers and replicated by hundreds of teams. Models themselves accelerate diffusion of any insight they are trained on.

LLM-based signals inherit and accelerate this dynamic — a 5–7× compression over three decades. As IBKR Research noted in 2026, LLMs materially shorten the shelf-life of copyable alpha, not just because more people use the same models, but because the models, by synthesizing and redistributing research, accelerate the diffusion of any insight they ingest. The implication is direct: LLM-derived signals should be treated as high-decay alpha sources, allocated accordingly, and retrained at higher frequency than classical factor exposures.

07 — Ground truth

The infrastructure problem nobody talks about.

Here is the failure mode that does not appear in academic papers, vendor presentations, or most conference talks: the majority of AI agent failures in live trading are not model failures. They are infrastructure failures. Industry data on production AI agent deployments shows failure rates of 70–95%, with data quality, freshness, and integration issues cited as the primary cause rather than model error. In trading specifically: stale prices are deterministic inputs to position sizing; missing ticks serially miscalculate volatility; API inconsistencies between vendors cause agents to act on a data state that never existed.

We encountered all three of the following when stress-testing Nyquist's agent loops against real tick data, ranked by detectability — the top is where teams lose most before noticing:

  1. Silent degradation — most dangerous. Decision quality erodes gradually with no system-level alarm, because the degradation is in signal quality, not availability — the orchestration layer reports green. By the time it surfaces in P&L, the drawdown is already significant. Drivers: NLP corpus drift, uncalibrated regime change, stale cost models, data-format breakage.
  2. Race conditions — hard to reproduce. Signal, risk, and portfolio-construction agents share a common data layer without strict synchronization. Inconsistent reads produce recommendations that violate risk limits at the moment of generation — even though each agent behaved correctly given its view. Drivers: shared portfolio state under concurrent write, missing transactional boundaries.
  3. Temporal drift — most common. The agent operates on a data state that was accurate hours ago. Yesterday's closing volatility on a day VIX has moved 15 points intraday is not a malfunction — the agent reasons correctly on wrong data and routes orders with high confidence. Drivers: batch SLA misses, alt-data ingestion lag, microstructure gaps, CPI / NFP / FOMC stress windows.

The most systematically damaging events sat at the intersection of high message traffic and structural data complexity: opening auctions, economic release windows, and quarterly index rebalances — events where the agent's need for data accuracy is highest precisely when data delivery is most stressed.

Infrastructure is not a commodity layer beneath the model. For autonomous trading systems, it is the model's ground truth. A model that reasons correctly on wrong data produces wrong answers with high confidence — worse, from a risk perspective, than a model that produces no answers at all.
The case for treating the data layer as first-class
Why this matters more for agents than analytics: data-freshness requirements for autonomous execution are categorically stricter than for reporting workloads. Point-in-time guarantees, tick completeness, and latency budgets under concurrent load should be contractually defined — at the agent level, not the analytics level — before any agent goes live.
08 — Honesty & close

Caveats, and the relevant question for 2026.

The evidence here carries real limits. Survivorship bias is endemic to a fixed universe defined by current liquidity — the top-500 cut as of 2024 excludes everything delisted or impaired during the period, inflating both approaches, and likely more so for sentiment signals biased toward large, heavily covered names. FinBERT is a proxy — a sophisticated agent with tool use, RAG, and multi-step reasoning can produce higher-quality signals; read this as representative of commodity implementations, not a ceiling. All metrics are gross. And regime generalizability is limited — the 2019–2024 US large-cap universe is a specific confluence of liquidity, macro variety, and market structure; emerging markets, small caps, credit, or illiquid alternatives all require independent validation.

The relevant question for quant practitioners in 2026 is not whether to use LLMs in trading — they are already being used, by everyone, whether or not teams explicitly chose to. The question is whether the deployment is disciplined: Does the team measure live IC? Does it account for look-ahead bias in the pre-training window? Does it stress-test infrastructure against microstructure events? Does it have an explicit alpha-decay monitoring framework?

Cross-sectional equity signal generation on liquid names using public models on public text is a highly competitive domain where gross alpha is thin, the cost structure punishing, and the decay cycle fast. Alternative-data parsing, regime classification, macro narrative synthesis, and research orchestration are where LLM capabilities are genuinely differentiated — and where the infrastructure requirements, though demanding, are compatible with the signal's natural holding period. The practitioners who extract durable edge from this technology are not the ones who deployed fastest. They are the ones who measured most rigorously.

09 — The summary

Key takeaways.

Six rules to live by
  1. IC below 0.03 with turnover above 8× is almost certainly net-negative after costs — a hard gate in any signal-evaluation framework, computed before gross Sharpe reaches any investment committee.
  2. LLM signals are best deployed as factor overlays, not standalone strategies — the incremental IC is more durable and cost-efficient, because lower net turnover reduces execution drag.
  3. Treat any backtest inside the LLM's pre-training window as contaminated — figures are biased upward 15–30% in gross Sharpe; the only reliable evaluation is live out-of-sample.
  4. Regime detection and narrative classification are the highest-conviction LLM applications — macro narrative signals have longer holding periods, lower turnover, and stronger structural anchors than earnings sentiment.
  5. Infrastructure SLAs must be defined at the agent level, not the analytics level — point-in-time guarantees, tick completeness, and latency budgets under concurrent load should be contractual before go-live.
  6. Build explicit monitoring for silent degradation — rolling 63-day ICIR, IC trend, and live-vs-backtest IC comparison belong in standard risk reporting; this mode causes the largest losses precisely because it is undetected until the drawdown is already significant.
Factor foundations
Fama–French 3- and 5-factor papers; momentum (UMD); classical IC / ICIR literature on factor diversification and crowding dynamics.
Text-based macro
IMF and Bank of England research on NLP of corporate and central-bank communications; BIS (2025) Parsing the Pulse on decomposed macro sentiment.
Look-ahead bias
NeurIPS 2025 and SSRN benchmarks formalizing pre-training contamination in LLM-based financial models (15–30% gross Sharpe inflation).
Decay & deployment
IBKR Research (2026) on shortening alpha shelf-life; industry data on 70–95% production AI-agent failure rates driven by data quality and freshness.
What layer are you using LLMs at in your pipeline — and what does your live ICIR decomposition actually show? If you're running similar comparisons or hitting infrastructure bottlenecks in your agent pipelines, we would be genuinely interested to hear from teams running these experiments in productioncontact@nyquist.pro.
  About Nyquist

Power needs interpretation — built in, not bolted on.

A bitemporal ontology and a real-time cross-asset state layer feed 36 named agents and a domain SLM — ML where it adds edge, classical methods where they hold, and an interpretability and governance layer throughout. Built for decision-grade infrastructure, not leaderboard R².