Prism Data Lab

Every analysis here ships its dataset and the code that produced every figure. Nothing on this site is investment advice.

LLMs in Finance1,266 words, 6 minutes

Earnings-call sentiment: what the papers measured, FinBERT to LLMs

What financial-sentiment research reports, from the 2019 FinBERT paper to 2026 studies of call tone, evasion, KPI extraction, and lookahead bias, by arXiv id.

"Sentiment" is the oldest use of language models in finance and the one with the most published numbers. The numbers are worth reading carefully, because the field has moved through three distinct tasks that share one word: classifying a sentence as positive or negative, scoring a whole earnings call for tone and relating that tone to something in the market, and, lately, extracting specific claims and behaviours (a guided figure, an evasive answer) from the call. This article reads seven papers from the last year of arXiv plus the one they all descend from, and reports what each measured, with the table it came from.

Where the lineage starts

FinBERT (arXiv:1908.10063, Araci, submitted 2019-08-27) fine-tuned BERT on financial text and reported that it "outperforms state-of-the-art machine learning methods" on two financial sentiment datasets, one of them the Financial PhraseBank of labelled sentences from financial news. The abstract gives no accuracy figure, and the point that matters seven years on is not the score but the task definition: three classes, one sentence at a time, labelled by annotators. That is the benchmark almost every later paper still reports on, which is both a strength (comparability) and the source of the field's main blind spot (a sentence is not a call).

Two 2026 papers show how the sentence task has evolved and how little the ceiling has moved. RA-FinBERT (arXiv:2608.09834) adds four rule-derived features to FinBERT's representation with 1,024 extra trainable weights; on its held-out news test set it reports "69.89% accuracy and a macro F1 score of 0.634, compared with 63.44% and 0.526 for text-only FinBERT", with neutral-class recall rising "from 18.18% to 45.45%" (all from the abstract). Those are low absolute numbers for a three-class task, and the neutral class is the reason: the hardest sentences in financial text are the ones that are neither good nor bad news.

TriAgent (arXiv:2607.19794) routes each sentence through a lexicon, then FinBERT, then a small LLM, escalating only on disagreement. Table 3 of the paper gives the operating points on the Financial PhraseBank all-agree subset (4,838 sentences): F1 0.665 at $0.0001 per thousand sentences, 0.716 at $0.0006, 0.787 at $0.0054, against 0.809 for always calling the LLM at $0.0288. The paper's Table 4 also reports a 20-ticker backtest for 2023 to 2024 in which the always-FinBERT variant has a Sharpe of 1.36 and the always-LLM variant 0.11. We report that row because it is in the paper; a two-year, twenty-stock backtest is not evidence about anything except the fragility of turning a classifier into a strategy, and the article on backtest overfitting on this site explains why.

From sentences to calls

The paper that does the second task properly is "Corporate Earnings Calls and Analyst Beliefs" (arXiv:2511.15214). It takes 35,203 earnings-call transcripts from 5,370 U.S. companies, 2008 to 2023, masks every numeral in the text so the model cannot read the guidance figures directly, embeds the text with FinBERT, and asks whether the text improves predictions of analysts' one-, two-, and three-year earnings forecasts on top of 446 stock-characteristic features. Table 3 reports out-of-sample R-squared for the combined models of roughly 15 to 18 percent at one year, 12 to 15 at two, and 10 to 13 at three; Table 5's Clark-West tests put the improvement from adding text at about 6 to 13 percent of forecast mean squared error. The paper's most interesting result is not the fit but the asymmetry in Figures 8 and 9: analysts move their forecasts by about 35 basis points per unit of sentiment where realised earnings move by about 26, and by about 9 basis points per unit of uncertainty language where realised earnings move by about 41. Text in a call measures something analysts over-react to (tone) and something they under-react to (hedging).

Note what this design does and does not claim. It does not predict returns. It predicts what analysts will write, and it uses a model trained on data through 2017 to do so, which sidesteps a problem the last paper in this article makes explicit.

From sentiment to behaviour

EvasionBench (arXiv:2601.09142) changes the label. Instead of positive or negative, it asks whether an executive's answer to an analyst's question was direct, intermediate, or fully evasive. The corpus is 1.38 million transcripts from 8,081 companies (2002 to 2022), filtered to 11.27 million question-answer pairs; a 1,000-pair gold set was labelled by experts, with Cohen's kappa of 0.835 between annotators (Table 4). Table 5 reports macro-F1 on that gold set: a fine-tuned 4-billion-parameter model at 84.9, against 84.6 for Gemini 3 Flash, 84.4 for Claude Opus 4.5, and 80.9 for GPT-5.2, with the untuned 4B base model at 34.3. The intermediate class is the hardest for every system. The paper reports no market-reaction result of its own; it cites earlier work for the claim that evasion predicts later misses.

The KPI-extraction paper (arXiv:2605.03147) is the sobering one. It asks models to pull named metrics and their values out of 10,477 transcript chunks from 20 S&P 500 companies (2023 to 2024), with 587 chunks and 2,460 entities annotated. Table 3 gives exact-match F1 of 11.5 percent for the best LLM (Gemini 3 Pro) and 3.4 percent for Llama-3.3-70B; encoders trained on SEC filings scored 0 percent exact F1 on calls. Semantic F1, which forgives boundary and wording differences, rises to 61.6 and 51.5 percent. Human evaluation of the final system's extractions (Table 5) found 79.67 percent precision, and the human annotators themselves reached only a Krippendorff's alpha of 0.429. When people cannot agree on what the KPI in a sentence is, a model's score against those labels has a low ceiling by construction.

The problem underneath all of it

"Detecting Lookahead Bias in LLM Forecasts" (arXiv:2512.23847) asks a question every study using a modern LLM on historical text has to answer: does the model already know what happened? The authors define a lookahead propensity from a date-only query (firm, ticker, target date, no text), and test whether the model's predictive signal is stronger where that propensity is high. On 91,357 Bloomberg headlines (2012 to 2023) with Llama-3.3-70B, the interaction between the LLM signal and lookahead propensity has a coefficient of 0.162 (t = 3.64), meaning a one-standard-deviation rise in propensity amplifies the LLM's apparent predictive effect by about 32 percent; on 7,568 headlines from 2024, after the model's training cutoff, the interaction is insignificant (t about 1.06). For earnings calls and capital expenditure the pattern repeats at about 12 percent (t = 2.01), and again vanishes post-cutoff.

This is why the 2025 analyst-beliefs paper's choice of a 2017-vintage FinBERT with masked numbers is not a limitation but the design. Any earnings-call sentiment result computed with a 2025 model on 2015 calls, without a test like this one, may be measuring memory.

What to take from the numbers

Sentence classification on curated benchmarks sits around F1 0.8 with an LLM and 0.7 with FinBERT-class models, and the neutral class is where the errors live. Call-level tone carries information about what analysts will do, at a modest R-squared, and it carries it asymmetrically. Behavioural labels such as evasion can be detected at F1 around 85 with a small fine-tuned model, matching frontier models. Exact extraction of numbers from spoken text is a solved problem for nobody, at 11.5 percent exact F1 for the best system reported. And any of these results obtained with a model whose training data covers the test period needs a lookahead test before it means anything. None of this says what any stock will do; the papers that report Sharpe ratios do so on samples too short to count.

This article is analysis and education, not investment, tax, or legal advice. Figures are cited to their source and dated; check them before relying on them.