Data Sources1,087 words, 5 minutes
Reading SEC EDGAR in Python: submissions, full-text search, XBRL facts
The three free EDGAR JSON endpoints in one standard-library script: a company's filings by accession number, full-text search hits, and deduplicated XBRL facts.
EDGAR is the only free source of primary financial statements for U.S. public companies, and it has three JSON interfaces that need no key, no registration, and no library beyond urllib. One lists a company's filings with the accession numbers you cite. One searches the text of every filing since 2001. One serves the XBRL facts a company has tagged, which is the closest thing to a structured financial statement the public gets. The script attached to this article wraps all three in a hundred lines, and the article explains the two rules the SEC asks you to follow and the three traps the data sets for you.
The two rules
The SEC's "Accessing EDGAR Data" page states the request ceiling plainly: "Current max request rate: 10 requests/second." It also asks that you "declare your user agent in request headers", with a sample of the form Sample Company Name AdminContact@example-domain.com. Requests without a descriptive user agent are refused. The script reads the string from an EDGAR_UA environment variable and falls back to a default that names this site; set your own before you run it at any volume.
The APIs page adds two facts worth knowing before you build on them: "the submissions API is updated with a typical processing delay of less than a second; the xbrl APIs are updated with a typical processing delay of under a minute", and data.sec.gov "does not support Cross Origin Resource Scripting (CORS)", so these calls come from a script, not from a browser page.
1. Submissions: a company's filings with accession numbers
https://data.sec.gov/submissions/CIK##########.json
The CIK must be ten digits, zero-padded. Apple Inc. is CIK 320193, so the URL ends CIK0000320193.json. The response carries the company's name, tickers, fiscal year end, and a filings.recent block of parallel arrays: accessionNumber, form, filingDate, reportDate, primaryDocument. The script zips those arrays into records and filters on form type:
python code/sec-edgar-python-submissions-fts-xbrl.py --cik 320193 --filings 10-K --max 3
Run on 2026-09-06, the three newest 10-Ks:
| filed | period | accession | primary document |
|---|---|---|---|
| 2025-10-31 | 2025-09-27 | 0000320193-25-000079 | aapl-20250927.htm |
| 2024-11-01 | 2024-09-28 | 0000320193-24-000123 | aapl-20240928.htm |
| 2023-11-03 | 2023-09-30 | 0000320193-23-000106 | aapl-20230930.htm |
The accession number is the citation. It is unique across EDGAR, it never changes, and it builds the archive URL: strip the dashes, and the filing folder is https://www.sec.gov/Archives/edgar/data/320193/000032019325000079/. The primary document sits inside that folder. Older filings beyond the recent block are listed in separate files named in the filings.files array; the script does not follow them, and a company with a long history will need that extra step.
2. Full-text search: every filing since 2001
https://efts.sec.gov/LATEST/search-index?q="generative artificial intelligence"&forms=10-K&dateRange=custom&startdt=2026-01-01&enddt=2026-09-06
This is the JSON behind the EDGAR full-text search page. Quote the phrase for an exact match, restrict forms to a comma-separated list, and give a date range. The response has hits.total.value (the number of matching documents) and a hits.hits array whose _id is accession:filename and whose _source carries the filer's display name, CIKs, form, and file date.
python code/sec-edgar-python-submissions-fts-xbrl.py --fts "generative artificial intelligence" --forms 10-K --since 2026-01-01 --max 3
On 2026-09-06 that returned 604 documents. The three newest were Medical Exercise Inc. (10-K filed 2026-06-29, accession 0001213900-26-073032), Red Violet, Inc. (filed 2026-03-04, accession 0001193125-26-091708), and Inuvo, Inc. (filed 2026-03-05, accession 0001654954-26-001943). Notice that "documents" is not "filings": a 10-K with the phrase in both the main document and an exhibit counts twice. The count is a useful trend measure and a poor census.
3. XBRL company facts: the tagged numbers
https://data.sec.gov/api/xbrl/companyfacts/CIK##########.json
Every number a company has tagged in an XBRL filing, keyed by taxonomy (us-gaap, dei) and concept (Revenues, NetIncomeLoss), then by unit (USD, shares, USD/shares), as a list of facts. Each fact carries start and end dates, val, the fiscal year and period (fy, fp), the form, the accession number (accn), and the filing date. This is where the traps live.
Trap one: the same period appears many times. A 10-K reports three fiscal years, so the fiscal 2023 net income appears in the 2023, 2024, and 2025 10-Ks. Naively summing or plotting the facts for a concept triples every year. The script keeps the first filing that reported each (start, end) pair:
seen, rows = set(), []
for v in sorted(f["units"][unit], key=lambda v: (v["end"], v["filed"])):
if v.get("fp") != "FY" or v.get("form") != "10-K":
continue
key = (v.get("start"), v["end"])
if key in seen:
continue # later 10-Ks restate the same period; keep the first report
seen.add(key)
rows.append({...})
Keeping the first report gives you the number as originally published; keeping the last gives you the restated number. Both are legitimate and they can differ, which is the subject of a separate article on this site about model-based table extraction versus XBRL.
Trap two: quarterly and annual facts share a concept. NetIncomeLoss holds three-month, nine-month, and twelve-month values. The fp == "FY" filter and the form filter are the minimum; a duration check on end - start is safer, and the extraction article uses one.
Trap three: concept names change. Apple's revenue was tagged SalesRevenueNet through fiscal 2017, Revenues in the 2018 10-K, and RevenueFromContractWithCustomerExcludingAssessedTax from the 2019 10-K on, following the revenue-recognition standard change. Query one concept and you get a series with a hole where the company switched. The company-facts file lists every concept the filer has ever used, so the fix is to look at the keys before choosing.
python code/sec-edgar-python-submissions-fts-xbrl.py --cik 320193 --fact NetIncomeLoss
Deduplicated facts on 2026-09-06: 67 rows, the last four being fiscal years ending 2022-09-24 (99,803,000,000 USD, accession 0000320193-22-000108), 2023-09-30 (96,995,000,000, 0000320193-23-000106), 2024-09-28 (93,736,000,000, 0000320193-24-000123), and 2025-09-27 (112,010,000,000, 0000320193-25-000079). Each number carries the accession that first reported it, which is what a citation needs. There are 67 rows rather than 19 fiscal years because the raw fact list includes short-duration entries carrying fp == "FY" (quarterly slices reported inside an annual filing) that this plain script does not filter by duration; the duration check is the next thing to add, and the extraction article adds it.
What the endpoints do not give you
They do not give you the narrative. Risk factors, MD&A, and footnotes are in the primary document, which is HTML, and reading them means downloading the file from the archive folder and parsing it yourself. They do not give you standardised statements: XBRL tags are chosen by the filer within the taxonomy's rules, and two companies' "operating income" can be different concepts. They do not give you anything about a company's filings before it began tagging in XBRL, which for most large filers is around 2009. And the full-text index begins in 2001, so a phrase search says nothing about the 1990s.
Everything here is free, current, and citable by accession number. The script is a starting point, not a client library: it has no retry, no cache, and no pagination beyond the first page of search hits, all of which are easy to add once the shape of the data is familiar.