Prism Data Lab

Every analysis here ships its dataset and the code that produced every figure. Nothing on this site is investment advice.

LLMs in Finance1,225 words, 6 minutes

Time-series forecasting with LLMs: claims, benchmarks, and sceptics

What 2026 papers report when language and time-series foundation models forecast returns, volatility, and factor rankings against simple baselines, by arXiv id.

The claim is that a model pretrained on text, or on a billion time series from other domains, can forecast financial series without being fitted to them. The claim has a general form and a financial form, and the evidence for the two is different enough that they should never be quoted together. This article reads six papers from the last year, three that make the case for pretrained forecasters and three that measure them against the simplest possible baselines in finance, and reports the numbers with the tables they come from. The short version: outside finance the argument is live; inside finance, on returns, nothing published beats the base rate, and on volatility one small model beats a 2009 regression by under two percent.

The general case

"Rethinking the Role of LLMs in Time Series Forecasting" (arXiv:2602.14744) is the strongest recent argument for the affirmative. It trains two GPT-2-based forecasters on 55 datasets from ten domains (over eight billion observations) and evaluates zero-shot on seven held-out datasets against time-series foundation models trained from scratch (Chronos, Moirai, UniTS). Table 1 reports the pre-alignment variant with the lowest MSE on several held-out sets, including a wind dataset at 1.015, a solar dataset at 0.228, and a NASDAQ dataset at 0.735, and Table 2's ablation finds that removing the pretrained language weights hurts on six of seven out-of-domain sets. The paper's claim is that pretrained language knowledge helps when the training corpus is large and the test domain is new.

Two things to hold onto. The datasets are mostly energy, traffic, and weather, where there is signal to find; and the one financial set, NASDAQ, is reported as an MSE with no baseline that a finance reader would recognise, such as a random walk or the unconditional mean. An MSE of 0.735 on normalised data does not say whether the model beat "tomorrow equals today".

The financial case, on returns

"When Directional Accuracy Lies" (arXiv:2607.12248) is the paper to read first if you have ever seen "80 percent directional accuracy" on a slide. It adapts TimesFM with LoRA to NASDAQ-100 and S&P 500 constituents (2005 to 2026), and reports every result as excess accuracy: the model's directional hit rate minus the hit rate of always predicting "up" on the same windows. Table 4 recreates the authors' own earlier setup and finds the always-up base rate at 0.704 and the fine-tuned model at 0.626, an excess of -0.078. On held-out stocks at a 128-day horizon (Table 5), the pooled adapter scores -0.081 on NASDAQ and -0.017 on the S&P 500. Per-sector adapters are significantly worse than a single pooled one (Table 6, Diebold-Mariano p below 0.001). A bull market makes any up-biased forecaster look skilled, and once the base rate is subtracted the skill is negative.

"Re(Visiting) Time Series Foundation Models in Finance" (arXiv:2511.18578) is broader: daily excess returns across 94 countries, 1990 to 2023, about two billion observations, with Chronos (five sizes) and TimesFM (four sizes) evaluated zero-shot against linear models, tree ensembles, and neural networks. The zero-shot numbers are Chronos-large at an out-of-sample R-squared of -1.37 percent with directional accuracy just above 51 percent, and TimesFM-500M at -2.80 percent with directional accuracy below 50 percent. The best fitted benchmark in Table 4, CatBoost, reaches an average out-of-sample R-squared of -0.10 percent. The paper reports a Sharpe ratio of 6.79 for a CatBoost portfolio at one window size; we quote it because it is in the table and we do not believe it describes anything achievable, for the reasons the search-aware paper below makes explicit. The paper's own conclusion is measured: off-the-shelf foundation models "perform weakly in zero-shot forecasting of daily excess returns" and need financial pretraining to compete.

The financial case, on volatility

Volatility is where a pretrained forecaster has a fair chance, because it is persistent. "Forecasting Realized Volatility with Time Series Foundation Models" (arXiv:2607.05291) tests nine foundation models zero-shot, with 1,000-day context windows, on 40 U.S. equities (2015 to 2026), five currency pairs, and five futures (2009 to 2026), against eight econometric specifications including the heterogeneous autoregressive family. Table 6 reports average QLIKE loss relative to the log-HAR model:

model h = 1 h = 5 h = 22
log-HAR (reference) 1.000 1.000 1.000
TTM 0.982 0.986 0.987
Sundial 0.998 1.084 1.182
HAR 0.998 1.009 1.070
ARFIMA 1.012 1.049 1.094

One model, TTM, beats the reference at every horizon, by 1.3 to 1.8 percent. Every other foundation model averages worse than a regression on three lagged averages of the same series. That is the most favourable published finance result for zero-shot foundation models we found in a year of arXiv, and it is a small win on the easiest target.

The LLM-in-the-loop case

Two papers test language models as reasoners rather than as sequence models, and both are careful about leakage in a way the field has not been. "Leakage-Aware Benchmarking of LLM Forecasting" (arXiv:2606.22719) has a 7B open model rank seven U.S. equity style factors each month from 2023-04 to 2026-03, using only information observable at decision time (lag-shifted FRED variables and an archived daily inflation nowcast). Table 1 and Table 3: the full pipeline's median monthly rank correlation is +0.154 with a mean of +0.131 and a long-short Sharpe of +0.71; a nearest-neighbour macro-analog model with no language model reaches a median of +0.161. The bootstrap 95 percent interval on the mean is [-0.02, +0.28] and the permutation p-value is 0.11, on 36 months. The authors' own reading is that the real-time inflation input and the analog retrieval explain most of the median signal.

"What survives honest evaluation?" (arXiv:2608.27734) has a frontier model discover trading strategies through a tool interface that excludes look-ahead features by construction, logs every evaluation, and deflates the reported Sharpe by the recorded trial count. On a 453-stock universe (2017 to 2025), Table 4 reports the agent's best strategy at a design-period Sharpe of 1.69 falling to 0.18 in evaluation, against buy-and-hold at 1.15 falling to 0.60. No model-discovered strategy across two frontier models, search budgets up to one hundred candidates, and five repeated runs passed the certification threshold; only passive benchmarks did. The experiment that should be pinned above every backtest is E1: a deliberately leaky oracle posting a Sharpe of 35 passed the deflated-Sharpe and backtest-overfitting tests completely. Statistical correction does not detect leakage; only the data pipeline can.

The sceptical reading, stated plainly

On daily returns, the published zero-shot results are at or below the base rate: negative excess accuracy in arXiv:2607.12248, negative R-squared in arXiv:2511.18578. On volatility, one model beats a 2009-vintage regression by under two percent (arXiv:2607.05291, Table 6). On macro factor ranking, the LLM pipeline's median is matched by a nearest-neighbour lookup and its mean is not distinguishable from zero on 36 months (arXiv:2606.22719). On strategy discovery under honest accounting, nothing survives (arXiv:2608.27734, Table 4).

None of that is a claim that language models are useless in financial forecasting. It is a claim about what has been shown. The general-domain results in arXiv:2602.14744 are real on wind and solar and say nothing about markets; the finance results say the simplest baselines are hard to beat and that any paper not printing the base rate, the leakage control, and the trial count is not yet evidence. When one does print all three and still wins, this article will be revised to say so.

This article is analysis and education, not investment, tax, or legal advice. Figures are cited to their source and dated; check them before relying on them.