Over the last few years a lot of effort has gone into adapting transformers and large language models (LLMs) to financial forecasting. The backtests often look excellent: high Sharpe ratios, shallow drawdowns, high win rates. Many of these models then disappoint once they trade live, sometimes badly, and often at the point where the market changes character.
This note sets out three reasons I think that happens: (1) errors compound when a model is rolled forward on its own predictions in a non-stationary market, (2) pre-trained text models carry information from the future into the backtest, and (3) nothing in the model constrains it when the regime changes. I then describe the safeguards we use. The mathematics here is meant as a way of organizing the argument. None of it is a proof.
1. Markets Are Not Stationary
The mistake I see most often from people coming to finance from machine learning is to treat price series as if they behaved like language, images or fluid flow.
The grammar of a language changes slowly, over generations. The equations that govern a fluid do not change at all. A network trained on text or on physical data is learning structure that will still be there tomorrow.
Markets are different. They are adaptive, adversarial games between many participants, and the distribution of returns moves over time. One way to write this down is to say that the transition probabilities between market states are not constant:
When a central bank moves from easing to tightening, or credit spreads widen sharply, the return distribution $\mathbf{F}_t(r)$ can change quickly. A transformer has no built-in notion of why prices move. It predicts the next step from recent context. After a long quiet period it will have learned that buying dips works. When liquidity dries up, it keeps applying the old pattern and loses money.
1.2 An Intuition for Why Generalization Guarantees Do Not Carry Over
The standard results that bound a model's out-of-sample error assume that training and test data come from the same distribution, which is exactly what equation (1) says we cannot rely on. They also tend to need the target to be reasonably well behaved. One common way to express that is a Lipschitz condition on the mapping $f: \mathcal{X} \to \mathcal{Y}$:
In words: similar inputs should give similar outputs. My argument is that around regime transitions markets do not behave this way, and the effective constant $L$ is very large or unbounded ($L \to \infty$). A small change in funding conditions, or a small policy surprise, can set off forced selling among leveraged participants, and the outcome is far out of proportion to the input.
If that is right, guarantees based on Vapnik-Chervonenkis (VC) dimension or Rademacher complexity tell us little here, and a model that fits its training period well can still be badly wrong out of sample. I offer this as intuition and not as a formal result. I have not tried to measure $L$, and I am not sure it could be measured.
2. Autoregressive Error Compounding
A decoder-only transformer generates a sequence one step at a time. Each prediction $\hat{y}_{t+1}$ is fed back into the context as if it were true and used to predict $\hat{y}_{t+2}$.
Every prediction carries some error $\epsilon_t \sim \mathcal{D}(0, \sigma^2)$. Over a forecast horizon $H$, the model is conditioning on its own earlier outputs:
A schematic way to write how the error builds up is:
I use equation (4) as a sketch of the mechanism, not as a derived bound. The point is that each step's error is carried into the later steps, and if the model amplifies its inputs at all, the accumulated error grows quickly with the horizon $k$. The model still produces a confident-looking path. The path just has less and less to do with what the market is doing.
2.2 Why Low Signal-to-Noise Makes This Worse
Define the signal-to-noise ratio of daily returns as:
The range in equation (5) is the rough working figure we use. It is not a measurement I am presenting here. In vision or language tasks the signal is far stronger relative to the noise, and a model rolled forward on its own outputs stays close to the truth for much longer.
As a back-of-envelope illustration: when $\text{SNR} \le 0.05$, noise variance exceeds signal variance by a factor on the order of $1 / \text{SNR}^2 \ge 400$. With that little signal, even a short rollout ($H \ge 3$) is conditioning mostly on its own earlier errors and very little on anything real.
3. Lookahead Leakage in Financial Embeddings
I think a good part of the backtest performance reported for financial language models comes from lookahead contamination: information from the future finding its way into the test. I cannot put a reliable number on how much. These are the three channels we watch for most closely.
Channel I: Tokenizer Vocabulary
Sub-word tokenizers are trained on filings and news that span the whole historical period (for example, 2000–2024). The vocabulary therefore contains tickers, company names and terms that did not exist at earlier dates. A model reading a 2011 earnings transcript with that tokenizer has been handed a small amount of information about which companies survived.
Channel II: Restated Filings
Scrapers that pull 10-K and 10-Q filings from EDGAR often pick up amended reports (10-K/A) or filings that include later restatements. The model then sees data labeled 2018 that was actually restated in 2021, which is knowledge nobody had at the time.
Channel III: Normalization Across Time
Batch normalization, layer normalization or rolling z-scores computed over a window that includes later data bring future variance into past observations. One effect is that historical drawdowns look smaller than they were.
Table 1: Sources of Lookahead Contamination
| Source | How It Leaks | What We Do About It |
|---|---|---|
| Tokenizer vocabulary | Future ticker merges, bankruptcy vocabulary | Point-in-time frozen subword vocabularies |
| SEC filing restatements | Ingesting 10-K/A amendments retroactively | Dual-timestamp immutable EDGAR log |
| Survivorship and delisting bias | Excluding bankrupt firms from training | Full historical CRSP/Compustat delisting universe |
| Two-sided smoothing filters | Centered rolling averages, HP filters | One-sided causal exponential kernels |
| Unpurged attention masks | Bidirectional attention across sequence boundaries | Causal triangular masking ($L_{ij} = 0 \text{ for } j > i$) |
4. Our Approach: Causal Constraints Around the Model
We do not try to make the model itself immune to regime change. We put constraints around it. Signal generation is kept separate from unconstrained curve-fitting, and every strategy has to operate inside a set of causal state-space constraints defined over a directed acyclic graph (DAG):
Three Constraints
- Regimes defined by market structure: We do not define regimes with moving averages or RSI thresholds. They are states in a graph built from the shape of the sovereign yield curve ($2\text{Y}/10\text{Y}$ and $10\text{Y}/30\text{Y}$ spreads), high-yield credit default swap (CDX) spreads and cross-asset volatility skew.
-
Fail-closed circuit breakers: When the estimated probability of a regime transition crosses a preset threshold:
the model is not allowed to add risk. Allocation is cut to cash or a market-neutral basis. We do not try to predict how a macro shock will resolve.$$ \mathbb{P}(\text{Regime Transition} \mid \mathbf{Z}_t) > \theta_{\text{critical}} $$(6)
- Deterministic risk checks (Blitz): Signals from a generative model are treated as untrusted proposals. Every order goes through the Blitz pre-trade risk engine, which enforces position limits, intraday drawdown limits and margin constraints before the order is routed.
4.2 The State-Space Model
The state evolves according to a transition equation whose parameters depend on the current regime:
Here $R_t$ is the discrete macro volatility regime, $\mathbf{m}_t$ is a vector of sovereign yield and liquidity shocks, and $\mathbf{s}_t$ is the latent market state. We run an adaptive Kalman filter alongside the graph and compute the Mahalanobis distance between live market data and the distribution the model was trained on:
If $D_M(\mathbf{x}_t) > \chi^2_{p, 0.999}$, we treat the live market as out of distribution. Control passes to capital-preservation logic straight away, without waiting for portfolio stop-losses to trigger.
5. Two Regime Shifts
Two recent episodes show the kind of change I mean. I am using them as illustrations. I do not have data on how any particular model performed through them.
Example I: Rates in 2022
For much of the period after 2008, equities and government bonds tended to move in opposite directions. A model trained on 2010–2021 data would have absorbed the assumption that bonds rise when equities sell off, so that a bond allocation acts as a hedge.
In 2022, central banks raised rates quickly in response to inflation, and equities and bonds fell together. The correlation between them turned positive. Any model that had learned the earlier relationship, whatever its architecture, would have been holding two losing positions where it expected one to offset the other.
Example II: The Yen Carry Trade in August 2024
At the end of July 2024 the Bank of Japan raised its policy rate. Over the following days the yen strengthened sharply against the dollar, and on 5 August the Nikkei 225 fell by about 12% in a single session. The move was widely attributed in part to the unwinding of yen-funded carry trades.
A model reading news headlines and company fundamentals had little to go on. US economic data had not changed much. The selling, which also reached large US technology stocks, came through funding and positioning. A language model with no view of cross-currency funding markets could not have seen it coming from text alone.
6. Checks We Run
These are the specific checks we apply to our own models. They are also reasonable questions to put to anyone running an LLM-based strategy, us included.
- Point-in-time tokenizer test: Re-run the model with a BPE tokenizer frozen on data from before 2015 and compare performance on 2016–2024. A large drop suggests the original result was leaning on vocabulary from the future.
- No multi-step rollouts: We do not let the model roll forward on its own predictions. Signals are single-step state-transition probabilities, evaluated independently within the DAG state machine and paired with fail-closed volatility limits.
- Dual-timestamped filings: We keep an immutable SEC EDGAR archive in which every filing has an event timestamp $T_{\text{event}}$ and an ingestion timestamp $T_{\text{knowledge}}$. An amended filing (Form 10-K/A) cannot appear in a backtest before the moment it was actually published.
- No LLM in the risk path: LLM output is probabilistic. Pre-trade risk controls need to be deterministic, so ours are compiled C++ state machines that run in static memory before an order is sent. A model's proposal is untrusted until Blitz has checked it.
- Combinatorial purged cross-validation: A rolling walk-forward test evaluates one historical path, and it is easy to overfit to that path through hyperparameter tuning. Combinatorial Purged Cross-Validation (CPCV) generates $\binom{N}{k}$ distinct train and test combinations, which gives a distribution of out-of-sample Sharpe ratios and drawdowns in place of a single number.
- Beta orthogonalization: Before we accept a factor, its raw signal is orthogonalized against five macro betas (equities, duration, the US Dollar index, crude oil and VIX). We admit only the residual, and only if its marginal information coefficient satisfies $\Delta \text{IC} > 0.03$.
7. Conclusion
Large language models are useful for reading text, extracting information and suggesting hypotheses. I think it is a mistake to use them as unconstrained end-to-end trading engines.
What has worked for us is treating the model as one component: it proposes, and causal constraints, held-out data and deterministic execution decide what happens next. That does not make a strategy safe from every regime change. It does mean that when the model is wrong, something other than the model is in charge of the risk.