Prism Data Lab

Every analysis here ships its dataset and the code that produced every figure. Nothing on this site is investment advice.

LLMs in Finance1,110 words, 5 minutes

Extracting tables from filings: a model versus XBRL, when each wins

XBRL company facts for one filer, deduplicated, with the tag changes and restatements any extractor hits, and what 2026 table-extraction benchmarks report.

There are two ways to get the income statement out of a 10-K. One is to ask the SEC's XBRL API for the numbers the company tagged. The other is to point a model, or a parser, at the rendered document and extract the table as it appears on the page. People treat this as a choice between old and new; it is a choice between two different objects. The tagged facts are the company's own machine-readable assertions, exact but shaped by the taxonomy and by the filer's choices. The rendered table is what a human reader sees, complete but recovered with error. This article pulls one company's facts, shows the three things in them that trip up any extractor, and then reports what the 2026 benchmarks say about how well models recover tables from the page.

The facts, deduplicated

The script reads the company-facts JSON for Apple Inc. (CIK 320193) and keeps eight us-gaap concepts, fiscal-year rows only, from 10-K filings, with a duration check so quarterly slices do not slip in. Output of python code/llm-table-extraction-vs-xbrl.py on the file retrieved 2026-09-05:

concept fiscal years reported first period end last period end duplicate reports of the same period
CostOfGoodsAndServicesSold 19 2007-09-29 2025-09-27 32
EarningsPerShareDiluted 19 2007-09-29 2025-09-27 32
GrossProfit 19 2007-09-29 2025-09-27 32
NetIncomeLoss 19 2007-09-29 2025-09-27 32
OperatingIncomeLoss 19 2007-09-29 2025-09-27 32
RevenueFromContractWithCustomerExcludingAssessedTax 9 2017-09-30 2025-09-27 12
Revenues 3 2016-09-24 2018-09-29 0
SalesRevenueNet 11 2007-09-29 2017-09-30 16

306 fiscal-year facts reduce to 19 fiscal years because every 10-K restates two prior years. The duplicate count is the first trap: a naive read of the facts file triples the history.

df = df[(df["end"] - df["start"]).dt.days > 300]          # full fiscal years only (drops quarterly slices)
first = df.sort_values("filed").drop_duplicates(["concept", "start", "end"], keep="first")

Trap two: the concept changed name

The revenue line is the second trap. Apple's revenue for fiscal 2007 through 2017 is tagged SalesRevenueNet. Fiscal 2016 to 2018 also appears as Revenues, in the 10-K filed 2018-11-05. From the 10-K filed 2019-10-31, revenue is RevenueFromContractWithCustomerExcludingAssessedTax, which is the tag the revenue-recognition standard introduced. Fiscal 2017 revenue of 229.234 billion dollars exists under all three tags in three different filings (accessions 0000320193-17-000070, 0000320193-18-000145, 0000320193-19-000119). A query for Revenues alone returns three years; a query for the current tag returns nine; the full 19-year series needs all three and a rule for which wins when they overlap. A model reading the rendered table never sees this problem, because the row is labelled "Net sales" on the page in every year. That is the strongest argument for page extraction.

Trap three: restated values

The script checks whether a period's value is the same in every filing that reports it. For NetIncomeLoss, 17 periods appear in more than one 10-K, and in 2 of them the restated value differs from the first report. The dataset has both values, with accession numbers; the deduplicated table above keeps the first. For cost of goods sold the difference is visible in the shipped CSV: fiscal 2008 is 21.334 billion in the 10-K filed 2009-10-27 and 24.294 billion in the one filed 2010-10-27, a reclassification, not an error. A page extractor that reads only the latest 10-K gets the restated number and never learns that another one existed; XBRL carries both, dated. That is the strongest argument for the tags.

The deduplicated series the script prints (revenue and net income in billions of dollars, with the accession that first reported each year) ends: fiscal 2023, 383.285 and 96.995 (0000320193-23-000106); fiscal 2024, 391.035 and 93.736 (0000320193-24-000123); fiscal 2025, 416.161 and 112.010 (0000320193-25-000079). Every number carries its citation. That property is what XBRL gives you and no extractor can.

What the extraction benchmarks report

PulseBench-Tab (arXiv:2606.07534) is a table-extraction benchmark of 1,820 human-annotated tables from 380 source documents including "financial filings, government reports, corporate disclosures, and regulatory filings", in nine languages, with 48.1 percent of tables containing merged or spanning cells. Table 7 ranks nine systems on the benchmark's graph-based score: the top commercial parser at 0.935, Gemini 3.1 at 0.816, an agentic parsing service at 0.798, and the lowest open-source tool at 0.360. The score is structural (does the extracted grid match the annotated grid), not numerical, so a 0.8 does not mean 80 percent of cells are right; it means the shape is mostly recovered. For a financial statement, a merged header cell attached to the wrong column is a wrong number.

The Stanford EDGAR Filings Dataset (arXiv:2606.18192) is the large-scale attempt to do this for every filing: 18.5 million filings since 1994, parsed with a rule-based, layout-first method that reconstructs numbers split across prefix, value, and suffix cells and writes tables as MultiMarkdown with span notation. Its validation is a reconstruction test: a model is given the parsed representation and asked to rebuild the table as HTML, on which the dataset's format achieves 94.5 percent adjusted recall against 75.7 percent for a widely used open-source EDGAR tool. Human evaluation put structural and semantic accuracy above 99 percent, about one comprehension-breaking error per hundred pages. The paper notes the dataset does not validate its tables against XBRL; it preserves XBRL fields where present. One error per hundred pages is excellent for text and, for a 300-page 10-K, three errors of unknown location.

The FinLongDocQA error analysis (arXiv:2604.03664, Table 9) shows where the errors go once a model uses extracted tables to answer questions: 62 percent of failures were retrieval (the operand was never found), 19 percent evidence use (the wrong page), 11 percent value extraction (the wrong cell), and 8 percent arithmetic. Cell-level extraction is a small part of the failure; finding the right table is most of it.

When each is right

Use XBRL when the question is about a concept the taxonomy defines, for a filer that tags it consistently, and when you need the number with its accession, its period, and its restatement history. It is exact, dated, and free, and the script here is enough to get it.

Use page extraction when the table you need is not tagged (segment detail in a footnote, a non-GAAP reconciliation, anything in an older filing before XBRL), when the row labels matter more than the taxonomy names, or when you are reading the document as a human would. Expect the shape to be right most of the time and the cells to need checking, and expect no citation finer than the page.

Use both when the numbers have to be right. Extract from the page, then reconcile against the facts API for every concept that has a tag; the mismatches are either extractor errors or restatements, and both are worth knowing about. The dataset attached to this article is the XBRL side of that reconciliation for one company; the code runs for any CIK.

This article is analysis and education, not investment, tax, or legal advice. Figures are cited to their source and dated; check them before relying on them.