Prism Data Lab

Every analysis here ships its dataset and the code that produced every figure. Nothing on this site is investment advice.

LLMs in Finance1,174 words, 5 minutes

Language models reading financial statements: published accuracy

Four 2026 benchmarks measured LLMs computing ratios, verifying statements, answering from full 10-Ks, and summarising MD&A. Accuracies and failure modes, cited.

A financial statement is the most structured text a language model is ever asked to read, and the published record of how well models read it is more mixed than either the vendors or the sceptics say. Four benchmarks from the first half of 2026 test four different things: computing a financial indicator from statement data, catching an injected arithmetic error, answering a question from a 129,000-token annual report, and summarising a management discussion without changing the reader's conclusion. This article reports what each one found, with the table the number came from, and then the failure modes, which are more consistent across the four than the headline scores are.

Computing indicators from statements

The long-horizon test (arXiv:2607.28661) builds 68,307 samples from 829 listed companies, 384 financial indicators, and 28 reporting periods, and asks a model either for one number (a single-index task) or for a multi-metric, multi-period table (a table-index task). Table 3 gives the results, and the split that matters is whether the model was given the formula.

model (as named in the paper) single index, with formula table, with formula single index, no formula table, no formula
Gemini-3.1-Pro 79.61% 70.70% 64.90% 38.22%
Claude-Opus-4.8 78.46% 65.61% 65.79% 38.85%

The best open-weight model in the same table, Qwen3.7-Max, scores 78.53 percent on single indicators and 72.82 percent on tables with the formula given. Removing the formula hint costs the best models about 14 points on single indicators and about 32 points on tables. The paper calls this a "knowledge bottleneck": models pattern-match to a formula when it is on the page and do not reliably carry the accounting definition themselves. It also reports that domain-specific financial models did worse than the general ones (DianJin-R1-32B at 41.29 percent with hints on single indicators and 11.41 percent on tables), and that supervised fine-tuning of a 35B open model added 8.54 points on single indicators and 3.82 on tables (Table 4). Fine-tuning helps the easy task more than the hard one.

Verifying a statement

FinVerBench (arXiv:2605.29586) inverts the task. It takes XBRL-derived statements from 43 S&P 500 companies, injects one error per instance (arithmetic, cross-statement linkage, year-over-year consistency, or rounding) at magnitudes from 0.5 to 20 percent, and asks whether the statement contains an error. There are 1,985 instances. A deterministic rule-based checker with 15 accounting identities scores 100 percent precision and 52.8 percent recall (Table 4): it never cries wolf and it misses almost half the injected errors, because half of them are not violations of any identity it knows.

The LLM results on the observable subset (Table 6, 105 instances, 43 clean and 62 with errors) are the most instructive table in any of these papers. One model, Claude Sonnet 4, scored 100 percent on every column. Seven models (the paper lists GPT-5.5, GPT-5.4, GPT-5.2, Claude Opus 4.6, Claude Sonnet 4.6, DeepSeek V3.2, and MiniMax M2.5) flagged every statement as erroneous: 100 percent recall, 100 percent false-positive rate, 59.0 percent accuracy, which is just the share of instances that had an error. GPT-4.1 and DeepSeek R1 misclassified 41 of 43 clean statements. The authors read GPT-4.1's explanations and found that "95.1% of false positives flag real but benign numerical discrepancies": the model recomputed subtotals from the line items shown, found they did not sum, and did not recognise that a standardised statement omits lines.

Then Table 11. The one model that scored perfectly did so on statements rendered with unrounded decimals; when the same statements were rendered with the rounding a real filing uses, its recall fell from 100 to 79.0 percent. Part of what looked like arithmetic verification was detection of decimal artefacts left by the error injection. The benchmark's own validity is the paper's subject, and it earns the title.

Answering from a whole annual report

FinLongDocQA (arXiv:2604.03664) has 7,527 questions over 1,456 S&P 500 annual reports for fiscal 2022 to 2024, converted from EDGAR filings to Markdown, averaging about 129,000 tokens of input. Table 5 reports exact match, a tolerance-based accuracy, and F1:

method model exact match tolerance accuracy F1
no context GPT-4o-mini 0.12% 0.13% 0.56%
long context Gemini-3-Flash 17.79% 21.98% 26.03%
dense retrieval Gemini-3-Flash 33.03% 41.93% 49.90%
the paper's agent Gemini-3-Flash 41.34% 43.54% 51.29%

The no-context row is the baseline that every study like this should print: the model knows essentially none of these numbers on its own, so whatever it gets right it got from the document. Feeding the whole document into a long context window gets under 18 percent exact match; retrieval more than doubles it. By difficulty (Table 7), questions needing one table get 43.99 percent exact match, two tables 36.99, three or more 35.45.

Summarising without changing the answer

The fourth paper (arXiv:2606.29251) measures something no accuracy score captures. It takes 300 MD&A sections from 10-Qs and 297 earnings-call transcripts from S&P 100 companies (fiscal 2025), has a model form a bear, neutral, or bull view from the full text, compresses the text, and asks whether the view changes. Table 2: with plain prompted summarisation, the top view flipped for 33.0 percent of MD&A documents and 23.9 percent of calls; the paper's best method brought that to 20.3 and 18.5 percent. The control is essential: rereading the full document with no compression flipped the view 11.0 and 8.8 percent of the time on its own, so a fifth to a third of decisions changing under summarisation sits on top of a tenth changing from model noise.

The failure modes, side by side

Read together, the four papers describe the same machine from four angles.

Formula dependence. Given the definition, models compute; without it, accuracy on tables falls to under 40 percent (arXiv:2607.28661, Table 3). An accountant does not need the formula for a current ratio on the page.

Miscalibration toward "error". Nine of fourteen models in FinVerBench flagged nearly every clean statement (Table 6). A verifier that always says "wrong" has perfect recall and is useless; a rule-based checker with half the recall and no false alarms is more useful in practice.

Retrieval, not reasoning, is the bottleneck for long documents. FinLongDocQA's error analysis (Table 9) attributes 62 percent of its agent's failures to missing operands at retrieval, 19 percent to using the wrong page, 11 percent to reading the wrong cell, and 8 percent to arithmetic. The model can add; it cannot reliably find.

Rendering changes the score. Rounded versus unrounded numbers moved recall by 21 points for the best verifier (FinVerBench, Table 11). A benchmark's formatting choices are part of its result.

Compression changes conclusions. One document in five changed its implied stance when summarised by the best method tested (arXiv:2606.29251, Table 2).

None of these benchmarks report a result that would let anyone act on a model's reading of a filing without checking the number against the filing. The right comparison for the exact-match figures is not a human analyst but XBRL, which carries the tagged value at 100 percent exact match for the concepts a filer chose to tag, and that comparison is the subject of the next article in this category.

This article is analysis and education, not investment, tax, or legal advice. Figures are cited to their source and dated; check them before relying on them.