NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #4345 most downloaded on PyPI
Python library to access and analyze SEC Edgar filings, XBRL financial statements, 10-K, 10-Q, and 8-K reports
Last release 2 days ago
02 Oct 2026
Ships on a steady schedule
a new release about every 1 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
4 years old
447 releases · first in 2022
One column per quarter.
A gate that fails a pull request which stages a 6.0 change without updating the upgrade guide. Deprecations, will-raise warnings and "removed in 6.0"…
The last 5.x release before 6.0. It brings the ownership family to the top level (from edgar import Form4, Schedule13D), adds a timeout to the current-filings feed, and fixes 19 defects: most of them in BDC portfolio figures and Schedule 13D/G ownership, plus section text and grep.
6.0 is next (planned 2026-11-09). It drops Python 3.10 and turns several warnings you may already see in 5.x into errors. Running this release with warnings visible shows what will change for your code, and the upgrade guide says what to do about each one. Fixes made from now until 6.0 ship in 6.0.
Upgrading. Nothing is renamed or removed, but some values change because they were wrong:
portfolio_investments() no longer counts holdings twice, and joint-venture holdings, range facts and subtotals no longer inflate totals. Against the filed balance sheet, BCSF went from +90.5% to 0.0% and NMFC from +24.8% to +0.9%.NonAccrualResult.num_nonaccrual is None, not 0, when a filing gives only a portfolio-level non-accrual figure or none at all.total_shares does, instead of summing every reporting person. Some shares_change values change as a result.signatures section.Filing.grep(document="EX-10.1") searches only EX-10.1, not EX-10.10 through EX-10.19 as well. Counts that included those exhibits drop.edgar_fund returns AMBIGUOUS_BDC with up to five candidates when a name matches several BDCs, instead of silently picking one.get_current_filings(), iter_current_filings_pages() and get_all_current_filings() take a timeout. A slow feed page can get 90s without changing the 30s default for every other request, and a timed-out page still gets all 5 attempts where a process-wide 60s timeout allowed one. (GH #1391, bead edgartools-7xtb)from edgar import Form4, Schedule13D now works for the whole ownership family. Form3, Form4, Form5, Form144, Schedule13D and Schedule13G are top-level names, identical to their subpackage classes and loaded on first use, so import edgar time is unchanged (596 ms vs 602 ms). The filing.obj() import table in the API docs had four rows that raised ImportError; all thirteen now import, and a test keeps them that way.get_with_retry (sync and async); with vcr's httpcore patch removed, all three fail.docs/upgrade/6.0.md in the same PR; replayed over the 265 commits since the guide existed, it flags the four PRs that shipped staging with no guide entry, including #1201. Adds a PR template with the Definition of Done checklist.ReportingPerson.reported() and OwnershipComparison.reported_shares_change tell an unreported ownership figure from a reported 0. The 5.x fields still hold a placeholder 0 for a figure the filing leaves out and name it in ReportingPerson.unreported_fields; reported(), reported_shares_change and reported_percent_change return None for it, as the fields will in 6.0. (GH #1379)OwnershipComparison summed joint filers, while total_shares took the largest one. A fund and its adviser reporting one 10,000,000-share position compared as 20,000,000. The comparison now uses the same max, and the new aggregation_basis (single, identical, parent_sums_children, ambiguous) says when it may fall short: GAMCO's Nevro 13D is 1,322,950 by max and 2,123,900 by its own Item 5(a). (bead edgartools-qsk4)unfunded_commitments_fair_value (-1,448.1M) is added back, giving +0.3%. DERA summary_by_company() gains total_source (filed or summed). (bead edgartools-3vad)edgar_fund silently picked one BDC when a name matched several. "Golub" returned GOLUB CAPITAL BDC though four Golub BDCs score 99.4-100. A name now resolves only on one match equal to it (ignoring a suffix like Corp or Inc., so "Ares Capital" is still ARES CAPITAL CORP) or a clear top hit; otherwise the tool returns AMBIGUOUS_BDC with up to five candidates and their CIKs. (bead edgartools-1c6z)Filing.grep(document="EX-10.1") also searched EX-10.10 through EX-10.19. In SeeQC's S-1 it returned 378 matches for "agreement" where EX-10.1 holds 133. An exact filename or document type now wins, and a substring match applies only when nothing matches exactly, so document="EX-10" still selects every EX-10 exhibit. (bead edgartools-5qqy)grep returned shifted matches after a character whose lowercase is longer. Positions were found in the lowercased text and applied to the original, so four "İ" before "going concern" returned "g concern dou". Literal matches are now found in the original text, case-insensitively. (bead edgartools-6th2)BDCEntity.is_active would have turned False for BDCs the newest SEC report dropped, though they still file. Their rows come from the 2025 report, so Ares Capital's activity date stopped at 2025-05-29 and would leave the 18-month window around 2026-11-29. For these 22 BDCs, is_active now checks the company's own filings once the report date expires; the new in_latest_report flag marks them. (bead edgartools-huul)bdc_portfolio investment records named no borrower. Every record had only type, fair value, cost and rate because the builder read a name field PortfolioInvestment does not have. Records now carry company_name, identifier, principal, PIK rate, spread and percent of net assets, with None for a missing figure; MAIN's first record is "MSC Adviser I, LLC", $255.0M. (bead edgartools-rd23)NonAccrualResult.num_nonaccrual reported 0 when a filing gave only a portfolio-level non-accrual figure, or none. WhiteHorse Finance files $10.6M of non-accrual investments for 2025 with no per-investment detail and showed 0; the count is now None unless the filing itemizes, and the new evidence_level says which (investment, aggregate, none). (bead edgartools-vblr)HTMLParser now collapses whitespace runs inside a text node to one space, as a browser does, while <br> and <pre> still break lines. Item 7 of Google's FY2004 10-K had about 500 mid-sentence newlines. (#1370)portfolio_investments() no longer counts holdings twice. Company totals, tranche parents and a second schedule's copies of a row were read as holdings, so rows summed 2-48% above the filer's own total on every BDC checked; MAIN's FY2025 10-K summed to $6.30B against a reported $5.52B and now reconciles within 0.2%. Foreign-currency values no longer replace USD ones, and new reported_total_fair_value and reconciliation_gap show any remaining difference.TenK keeps both items of a combined heading such as "Items 1 and 2. Business and Properties". A TOC link reading "Items 1 and 2" (or "Items 1. and 2.", "Items 7. and 7A.") was dropped, so tenk.business was None on Viper, Devon, Freeport and Cheniere. The section is now keyed as the first item, lists both in Section.covered_items, and answers for either. Talos no longer shows a 1,311-character "Properties" fragment as Item 2. (GH #1382, #1383)OwnershipComparison.get_summary() warnings now point at the caller. It emitted two FutureWarnings attributed to amendments.py, so Python showed them once per process wherever it was called from; it now emits one, at the caller's line. A filing with an empty <reportingPersons> now displays its ownership as "not reported" instead of 0, matching the comparison. (GH #1379)OwnershipComparison.is_liquidating reported a sale against a Schedule 13D/G filing that gives no share figures. A 13D/A or 13G/A may omit them and a pre-2025 header-only filing has none, but they were read as 0: an amendment omitting them on the Aadi Bioscience 13D in the test data sold all 4,535,000 shares. The direction flags are now False when the change is unknown, and shares_change and percent_change warn that they return None in 6.0. (GH #1379)edgar.ownership.core.is_numeric() raised TypeError for pandas string and nullable numeric columns. The reported Arnaboldi Form 4 has UnderlyingShares of 0 [F1]; its value calculation now returns None instead of raising. Ownership share totals also return None for unreadable text rather than raising. (GH #1368, thanks @manantlerio)EntityFacts.get_total_liabilities() returned total assets for filers with no standalone liabilities total. Its last fallback was LiabilitiesAndStockholdersEquity, which always equals total assets: NIKE's FY2026 figure came back as $38.41 billion against $23.545 billion of liabilities, and 16 of 51 large US filers were affected. It now returns None there, matching Financials.get_total_liabilities(). (GH #1279)PortfolioInvestment.company_name returned the instrument (First Lien, Ordinary Shares) when the filer tagged the company as a member. On PSEC's FY2026 10-K (0001287032-26-000269) 142 of 248 identifiers named an instrument. A member candidate that fills the whole first pipe segment is now the company, not a prefix to strip. (GH #1373)PortfolioInvestment.company_name named a schedule heading or sub-sector instead of the company on category-led identifiers. On CGBD's Q2 2026 10-Q (0001544206-26-000055) 348 of 352 positions read Investment, Credit Fund or a sub-sector such as Durable; SLRC and PFX were hit too. Only the axis prefix is stripped now, and leading heading segments are skipped. (GH #1372)Form4.shares_traded raised AttributeError on a derivative-only filing. market_trades is None when the non-derivative table has no transactions, and the accessor read .Shares off it without the None/empty guard its neighbours already had: Vertex's 2020-10-19 Form 4 parsed and returned its derivative transaction but crashed on shares_traded, which now reports 0 market trades. (GH #1359)A patch release with two fixes. One is a regression from 5.59.0: ExxonMobil's 10-K Item 7 opened on a "Table of Contents" breadcrumb again. The other:
A patch release with two fixes. One is a regression from 5.59.0: ExxonMobil's 10-K Item 7 opened on a "Table of Contents" breadcrumb again. The other: edgar_read returned no 13F holdings section for any filing.
Upgrading. No action needed:
edgar_read returns a 13F holdings section. It lists the top 30 securities by value, with issuer, shares and value. The header counts securities ("29 of 29" for Berkshire), while the summary's Total Holdings counts information-table lines (89), so the two numbers can differ.edgar_read returned no holdings section for any 13F-HR. A filing whose summary reported Total Holdings: 211 came back with holdings set to None, because the section tested a DataFrame for truth and then looped over its column names. It now lists the top 30 rows by value (issuer, shares, value), and an extraction failure logs at warning level. (GH #1337)EDGAR_ACCESS_MODE and NORMAL / CAUTION / CRAWL emit DeprecationWarning . They never changed how edgartools talks to the SEC; use EDGAR_RATE_LIMIT_PER_…
The headline is filing sections. Seven fixes change how 10-K and 10-Q items are found through the table of contents. Items no longer return a neighbouring section, a summary, or each other's text: FirstEnergy's Item 7 no longer returns Item 8, Southern Co.'s Item 7A no longer duplicates Item 7, and Citi's FY2022 Item 7A no longer opens on Item 1. Tables in those sections render cell by cell instead of fusing numbers ("20217,294,800"). Alongside that: a 13F with a couple of holdings could be valued 1,000× too high without a warning, SGML downloads no longer cache an empty filing header when the SEC errors or refuses a request, and four XBRL fixes, including balance-sheet debt reported under combined concepts.
Upgrading. Several of these change values or behaviour you may rely on:
doc.text() keeps content it used to drop. Short figures such as "$85" and a cell's own label ahead of its <div>s are now included.Filing.sgml() and Filing.header raise on HTTP errors. A 5xx, 403 or 429 now propagates, where it used to return a header with every field None. The 429 keeps its retry_after. A later call on the same Filing retries.Ambiguous13FValueUnitWarning, so code that promotes warnings to errors may see it.sections['risk_factors'] works when the section is keyed part_i_item_1a, and 'mda' in sections is true when Item 7 exists.EDGAR_ACCESS_MODE and NORMAL/CAUTION/CRAWL emit DeprecationWarning. They never changed how edgartools talks to the SEC; use EDGAR_RATE_LIMIT_PER_SEC and EDGAR_HTTP_TIMEOUT. They are removed in 6.0.balance_sheet() gains debt lines for filers reporting under combined or short-term debt concepts.Notes.from_xbrl() without a FilingSummary now returns disclosure-category stems with their Tables, Policies and Details. The fallback builder only accepted note-category roles, and a family that hangs from a *DisclosureAbstract is classified disclosure, so XBRL.from_directory() callers got an empty concept index. gahc goes from 0 notes to 7, aapl from 0 to 16. (GH #1218)EDGAR_ACCESS_MODE and the NORMAL/CAUTION/CRAWL modes are deprecated and have no effect. They advertised a timeout, connection-limit and retry policy that was never wired into the HTTP client; nothing in the package read edgar_mode, so all 3 modes behaved identically. Use EDGAR_RATE_LIMIT_PER_SEC and EDGAR_HTTP_TIMEOUT. The names still import and now warn; removed in 6.0. (GH #1326)20217,294,800 for the year 2021 and a 7,294,800 share count; 144 such tokens across 70 fixtures are now 0. Tables render with the cells doc.text() shows, and doc.text() itself now keeps short figures such as Regions' "$85" and cell labels it used to drop. (bead edgartools-wzgu)Filing.sgml() and Filing.header now propagate every status-bearing HTTP error in both error modes, preserving the original 429 retry_after value. After an injected refusal, a later successful download on the same filing recovers Apple's 2024-09-28 report period; content-error homepage fallback is unchanged.obj["Item 7"] returned Item 8's financial statements (306,341 chars) instead of the MD&A. Such rows now keep their item label. Patch by @sf1tzp. (GH #1347)Item 7 and Item 7A sliced to one span; they are now separated by their headings, so Southern Co's obj['Item 7A'] returns its 347-char cross-reference instead of Item 7's 282,239 chars. Prospectus markdown() now also stops at the financial statements, as text() did. (GH #1345, bead edgartools-rc46)Item 7A opened three pages early on Item 1's human-capital section. The offset is now calibrated from the printed footers and is a no-op when they are aligned. (GH #1346)Filing.sgml() and Filing.header propagate HTTP 5xx errors in both default and strict error modes, leaving the same filing retryable. An injected 503 followed by Apple's checked-in 2024 10-K now recovers its 2024-09-28 report period on the second call.dei:LegalEntityAxis and a classification axis; the member hierarchy keyed rows by member alone, so the LegalEntity row — carrying the more precise $82.723B against the other's $82.7B — was dropped. A member on two axes is no longer reordered.Disclosure marker. UNP names most of one family DisclosureDebtDetails1 but two members DebtDetails6, so those surfaced as spurious top-level notes while the real Debt note lost two Details. The marker is now ignored when matching a stem, taking UNP from 20 notes to 18 and 318 reachable concepts to 336.Filing.html() raised AttributeError instead of returning None when a filing's homepage listed no primary document. The homepage property is optional and comes back None for some filings, but two call sites used it unguarded, so the scheduled build went red on a PDF-primary APP NTC filing. Both are guarded now.Tables and Details roles were returned by xbrl.notes() while their parent was returned by xbrl.disclosures(). In a filing that falls back to role names only the family stem hangs from a *DisclosureAbstract, so gahc's ConvertiblePromissoryNotesPayable moved alone and its 4 children stayed notes. A child now follows its nearest stem when that stem declares a disclosure; no other role in the 7 committed fixtures changes. (GH #1218)Statement.get_raw_data(view="detailed") dropped NVIDIA's reportable-segment revenue. On the FY2026 10-K (0001045810-26-000021) the default view kept Compute & Networking $193.479B and Graphics $22.459B; DETAILED returned neither and showed ProductOrService Compute/Networking instead. Member hierarchy keyed rows by the first axis, so two-axis facts collided once Data Center's children entered the set. Hierarchy now nests only single-axis members. (GH #1331)edgar could fail. Importing edgar once on the main thread before starting threads avoids it; this is now documented under Common Pitfalls. A structural fix through lazy submodule loading is planned for 6.0. (GH #1325)set_rate_limit() and enable_local_storage() appeared in the performance, SEC-compliance and Form 4 guides but are not in the package; the working names are EDGAR_RATE_LIMIT_PER_SEC (set before import; the default is 9, not 10) and use_local_storage(). The configuration page said submissions are cached up to 10 minutes; MAX_SUBMISSIONS_AGE_SECONDS has been 30 seconds since #471.PortfolioInvestments.from_xbrl() cut the first word off a borrower's name when the filer labelled a range member with it. BXSL's FY2025 10-K (0001736035-26-000004) returned Street Buyer, Inc. 2 with industry='High' for its four High Street Buyer, Inc. positions: BXSL labels srt:MaximumMember "High", and every member label was an industry candidate. The five srt:RangeAxis members no longer are; across 12 filings and 8,212 positions only those rows change. (GH #1324)balance_sheet() dropped debt reported under combined or short-term concepts. Filers using LongTermDebtAndCapitalLeaseObligations, DebtCurrent or ShortTermBorrowings lost the line entirely: CSX's balance sheet showed no debt, and now shows $18.165B long-term plus $708M current. (PR #1330, thanks @wittling)parse_investment_identifier() recompiled up to 933 regexes on every call. Six sites built patterns from INVESTMENT_TYPES per call, past the 512 re caches, so a piped Company | Type identifier never hit the cache. They are now compiled once and cached: parsing BXSL's 703 identifiers on its Q2 2026 10-Q (0001736035-26-000016) drops from 40.8 s to 1.1 s, every field identical. (GH #1342)to_dataframe(*columns) no longer builds the full fact width before projecting. Selecting columns now narrows the frame at construction instead of at the end, so a two-column call on the JPM 10-K fixture drops from 7.59 MiB to 0.71 MiB peak and from 78 ms to 8 ms, with a byte-identical frame. (GH #1181)The guide now requires Python 3.10, matching package metadata, and identifies cashflow_statement() as deprecated in favor of cash_flow_statement() .
The headline is the SEC's BDC data sets. When the SEC relocated them, the fair value column changed its heading and BDCDataset.summary_by_industry() quietly reported $0 for every industry. It now resolves fair value and industry from every header DERA has used, and portfolio investments carry an industry, where it came from, and a normalized sector. Alongside that: seven XBRL statement fixes, a warning when a 13F's value units are ambiguous, and a BDC register that stops losing companies each January.
Upgrading. Several of these change values you already read:
summary_by_industry() sums fair value from the current period only, counting a filer's industry subtotals rather than adding its line items on top, and gains a num_bdcs column. ARCC's 2024Q4 Software And Services exposure reads the filed $6.572B. summary_by_industry(by='sector') merges spellings such as "Software Sector" and "Software & Services".summary_by_company() reports fair value and counts investments, not rows. It gains total_fair_value, taken from the filer's own total where reported, and is sorted by it. num_investments is now the number of distinct positions at the period end.to_dataframe(clean=True) coalesces BDC columns. fair_value, cost, shares, principal and industry are single canonical columns, followed by industry_source and sector.get_bdc_list() combines the last three yearly reports, so is_bdc_cik() recognises Ares Capital again.[Axis], [Domain] and member nodes no longer appear as line items; rendered output is unchanged.decimals can be the string "INF" on assembled statement rows, as filed, where it used to be 0. Code that assumes an integer should allow for it.annual_report=True, and 20-F and 40-F annual reports are recognised.Credit Agency. Fannie Mae, Freddie Mac and Farmer Mac move off Operating Company, and is_financial_institution() counts them.XBRLS.from_filings() logs a warning naming any filing that fails to parse, instead of dropping it silently.edgar.dates. edgar.core re-exports them until 6.0.PortfolioInvestment.industry, .industry_source (axis, enumeration, identifier, peer) and .sector, with PortfolioInvestments.filter(industry=) and .by_industry(); the DERA path reads the same fallbacks and gains summary_by_industry(by='sector') and a summary_by_company() with fair value. Filings with a resolvable industry in 2025Q2 rise from 123 to 158 of 160. (bead edgartools-t1wr)get_period_views("CashFlowStatement") returned an empty list because the statement type had no entry in the period-view configuration, so to_dataframe(period_view=...) had no name to accept. Cash flow now offers the same named views as the income statement, which selects from the same periods. (GH #1253)edgar.dates. filing_date_to_year_quarters, current_year_and_quarter, is_start_of_quarter and parse_acceptance_datetime now live beside extract_dates instead of in edgar.core; edgar.core re-exports the same objects until 6.0 removes the shim. Continues the bounded extraction that produced edgar/settings.py, taking core.py from 579 to 514 lines.Path(__file__).parent, which breaks the moment a file moves directory; they now import TESTS_DIR, REPO_ROOT, FIXTURES_DIR, CASSETTES_DIR and DATA_DIR from a new tests/paths.py. Behaviour-preserving — test-fast reports the same 7,261 passed as before — and it is the prerequisite for the test-tree reorganisation.edgar/documents/utils/html_utils.py as text_content, text_stripped, text_joined and text_skipping_tables, with an html_to_text wrapper. Behaviour-preserving: the five characterization suites pass unchanged against their recorded bs4 baselines, 297 tests.docs/installation.md listed Python 3.8+ in two places (System Requirements and the venv setup section) while pyproject.toml requires >=3.10. Both now read 3.10, matching package metadata.cashflow_statement() as deprecated in favor of cash_flow_statement().ThirteenF could silently inflate ambiguous filing values by 1,000x. It now warns on ambiguous thousands conversions and exposes raw values and unit diagnostics. A verified value_unit='dollars' override reads Kahn Brothers' Q1 2022 holdings as $787,553,692; the SEC summary remains $1 higher. Automatic unit choices are unchanged.BDCDataset.summary_by_industry() reported $0 for every industry. The SEC's relocated BDC data sets head the fair value column Initial fair value of Investment and leave Investment Owned, Fair Value present but empty, so it summed an empty column. Fair value, cost, shares and industry now resolve from every header DERA has used, the summary counts only current-period rows and each filing's industry subtotals, and ARCC's Software And Services exposure comes back as the filed $6.572B.find_mutual_fund_cik() could answer from a stale ticker frame. The ticker-to-CIK dict had its own cache stacked on the cached frame, so clearing or replacing the frame left the lookup frozen for the rest of the process. The dict now follows the frame it was built from, and the fast test serves a 3-row slice of the SEC file instead of downloading it.entity_info compared the whole dei:DocumentType against 10-K, so Shopify's 10-K/A reported amendment=True alongside annual_report=False, contradicting the filing's own dei:DocumentAnnualReport. The amendment suffix is now stripped before classifying, and 20-F/40-F annual reports are recognised too. (GH #1226)Credit Agency category that is_financial_institution() counts; 6120 joins the bank codes as a depository. (GH #1120)get_bdc_list() read one year's CSV, and the 2026 report dropped 15 registrants the 2025 one carried — Ares Capital among them — so is_bdc_cik(1287750) answered False. The last three reports are now combined, restoring 22 registrants (212 to 234) with the most recent row per CIK. (GH #1146)point_in_time. Both lookups walked only the displayed columns, which are durations on a cash flow statement, so Apple's beginning and ending cash rows returned point_in_time=None and unit=NaN despite their metadata saying instant and usd. Period type and unit belong to the concept and are now read that way. (GH #1228)decimals="INF" was rewritten to integer 0 in assembled statements. Apple's exactly-filed $0.00001 par value reached get_statement() claiming to be rounded to the dollar, while the same fact still read INF through FactQuery. The sentinel is now preserved; scaling and formatting ask decimals_for_scaling for a number, so rendered statements are unchanged. (GH #1229)show_date_range=True was ignored by a stitched statement's normal rendering. The accessor stored the option, but StitchedStatement.render() defaulted the parameter to False and never read it, so repr() showed Dec 31, 2024 where an explicit render(show_date_range=True) showed Jan 1, 2024 - Dec 31, 2024. An explicit argument still wins. (GH #1176)get_statement() returned the presentation linkbase's [Axis], [Domain] and member nodes as line items — four empty automotive headings in Tesla's operations statement. Across the 8 committed fixtures this drops 2,390 such rows from 268 statements, with no value-bearing row lost and no change to rendered output. (GH #1224)Statement.to_markdown() emitted labels verbatim, so Global Arena Holding's "conversion of debt | shares" row carried four unescaped pipes where a two-column row has three, and Markdown parsers read it as a different shape than the header. Labels, headers and values are now escaped. (GH #1227)XBRLS with no trace. XBRLS.from_filings() swallowed every exception, so a batch that lost a filing looked identical to a complete one. The failure is now logged as a warning naming the accession number, form and exception; a filing with no XBRL at all is still skipped quietly. (GH #1174)Thirty-eight entries, all in the XBRL parser, the fact queries, stitching, and the Financials getters. The recurring shape is a lookup keyed too narro
Thirty-eight entries, all in the XBRL parser, the fact queries, stitching, and the Financials getters. The recurring shape is a lookup keyed too narrowly to tell filed facts apart — a context ID, a currency, a concept name, a label pattern — that then returned the first match or silently dropped the second. Every one of these now identifies the fact the filing actually reported, or returns None rather than a substitute.
Upgrading. Several of these change values you already read:
xbrl.us labels, statements selected by role URI, render(standard=False), and Statements.to_dataframe() all agree with Statement.to_dataframe(). Apple's FY2023 capex is −10,959,000,000 everywhere; it was positive on three of those paths.Financials scalar getters identify concepts, not labels. get_total_liabilities() no longer returns liabilities and equity (NIKE's debt_to_assets was exactly 1.0); where a filing reports no consolidated us-gaap:Liabilities — Coca-Cola, Amazon — it now returns None. get_revenue() returns the filed total rather than a component; get_shares_outstanding_basic() no longer stops at an abstract heading.FactQuery.to_dataframe(), EntityFacts, and stitched queries return the same columns and dtypes regardless of which rows match; stitched frames keep columns the matched rows leave empty. pivot_by_period() and _deduplicate_facts keep facts that collide on unit or currency — expect more rows.calculate_ratios() computes every ratio from one period; analyze_trends() works on a two-period 10-K; RenderedStatement.to_dict() is JSON-safe (Infinity → None).xbrl.axes / domains_for_role() keep an axis's second domain root and childless roots.FactQuery.sort_by() no longer reorders the shared cache, so independent queries stop changing each other's results.get_financial_metrics() renders each statement once instead of once per metric.aggregate() summed the filed strings and raised TypeError on ordinary numeric facts. It now aggregates numeric_value and skips facts that have none; Apple's FY2023 product revenue by srt:ProductOrServiceAxis returns totals. (GH #1277, bead edgartools-n93w.3)by_concept() could not match a concept whose name contains an underscore. Every _ was rewritten to :, so YUM's yum:YUM_LesseeOperatingLease… matched nothing in either spelling. The pattern is now matched against both spellings unaltered; 20,094 exact queries across 15 filings lose no row. (GH #1276, bead edgartools-n93w.4)facts_history(include_dimensions=True) returned the undimensioned series. The flag never reached the query. It does now, and the date column falls back to period_instant for instant facts — previously every balance-sheet concept returned an empty frame. (GH #1286, bead edgartools-n93w.5)sort_by() sorted the shared fact list in place, so an unrelated .limit(5) changed on 15 of 15 filings measured. The sort now builds a new list. (GH #1287, bead edgartools-n93w.1)sort_by() raised TypeError on most filings, and silently did nothing on a sparse column. None values broke the comparator, and the column was looked for in the first row only. The key is now a total order over mixed types and the column is found anywhere in the result. (GH #1288, bead edgartools-n93w.2)analyze_trends() returned {} for a 10-K reporting exactly two annual periods. It required a named three-period view; it now falls back to the periods the statement displays, so Auburn National's $603,000 and $598,000 of contract revenue are reported. (GH #1292, bead edgartools-hsgs.3)calculate_ratios() let each operand take the first value in its own dictionary, so Netflix's Q3 2024 net margin was 0.2767 against a correct 0.2406. Every operand now comes from one period, and a ratio missing an operand is omitted rather than borrowed. (GH #1280, bead edgartools-hsgs.1)RenderedStatement.to_dict() produced output that was not JSON. A prior-period value of zero yielded Infinity in the comparison, which 17 corpus filings hit. Non-finite comparisons now serialize as null. (GH #1290, bead edgartools-hsgs.4)efc_Year2038Member was absent from domains_for_role() though five facts use it. Every root is now registered; Axis.domain_id still names the first. (GH #1278, bead edgartools-xn1u.2)19-Sep-2024 failed all thirteen word-date formats. The tokenizer now accepts the registry's separators; 40,219 facts across 16 filings show no already-transformed value changing. (GH #1284, bead edgartools-xn1u.1)render(standard=False) displayed cash-flow outflows as positive, disagreeing with its own DataFrame. The statement type was stamped on rows only along the standardization path; it is now stamped regardless of the flag. Apple's FY2023 capex reads −10,959 in both. (GH #1289, bead edgartools-vaau.2)get_total_liabilities() returned liabilities and equity, so debt_to_assets came back as exactly 1.0. A Liabilities$ label pattern matched the combined total; the getter now identifies the concept and returns None where no consolidated total is filed (Coca-Cola, Amazon). NIKE's FY2026 ratio goes from 1.0 to the correct 0.613. (GH #1279, bead edgartools-y4yp.1)get_revenue() returned a component of revenue instead of the filed total. Contract revenue was tried before Revenues, understating Cato's FY2023 top line by $7.7M of other income. The order now matches Statement.REVENUE_CONCEPTS. (GH #1294, bead edgartools-y4yp.2)get_shares_outstanding_basic() returned None for shares present in the statement. The lookup stopped at the abstract heading above the populated row; it now skips abstract rows and iterates its matches, returning Auburn National's 3,498,030. (GH #1291, bead edgartools-y4yp.3)_1 collided with an internal duplicate-occurrence index, and the parser silently dropped one of the two facts. ClearOne's Q2 2025 10-Q lost 6 of 590 facts, including the quarter's own revenue. Key parts are now joined with |, which no NCName admits. (GH #1295, bead edgartools-jj77.1)unit was discarded as a structural element. ctso:NumberOfSharesInAunit was skipped by a bare-suffix test; structural elements are now matched by fully qualified name. (GH #1293, bead edgartools-jj77.2)by_statement_type('BalanceSheet') lost period_start on 5 of 5 stitched pairs. Columns and dtypes now follow the query's configuration per engineering/decisions/facts-dataframe-schema.md. (bead edgartools-trhe)pivot_by_period() on company facts discarded facts that collided on a key too narrow to tell them apart, and for an IFRS filer discarded nearly all of them. IFRS concepts carry no label upstream, so 765 of LPA's 768 facts shared one row. Unit is now part of fact identity and the pivot widens its key only for rows that collide; the filing pivots to 188 rows. (GH #1196, bead edgartools-6byc)source_filing_index was always None and by_filing_index() matched nothing. Rows now carry the filing the stitcher selected each period from; a two-filing ORCL stitch splits 115/100. (GH #1205, bead edgartools-qbm7)value flipped float64→int64 on 16 of 36 narrowing queries, and an empty result had no columns. FactQuery.to_dataframe() and EntityFacts.to_dataframe() now declare their columns and dtypes; fiscal_year stays int64 until 6.0. (bead edgartools-7wtj)Total* concepts, so a standardized statement returned several rows all claiming to be the same total. Citizens Financial's balance sheet had three "Total Assets" rows. A tag not marked is_total no longer resolves to a total (74 spellings), and IFRS tags answer only from IFRS entries. (GH #914, bead edgartools-gnx5)render(standard=True) returned a DataFrame with no standardization in it. The standard_concept column was never emitted. It is now present for every standardized statement and absent otherwise; filer labels are unchanged. (GH #1083, bead edgartools-t3zh)_merge_complementary_rows() matched dimensional rows on label alone, merging 213 different-concept pairs across the corpus. A dimensional row must now also match on concept; 2 merges remain. (GH #1206, bead edgartools-j7iz)xbrli:divide was unreachable, so usdPerShare was recorded as USD. Divided units are now parsed, and the unit-to-currency rule lives once in edgar.xbrl.core.unit_currency. (GH #1170, bead edgartools-uetp)PresentationTree now carries root_element_ids; 9 of 2,052 corpus roles are multi-root and gain 106 rows. (GH #1245, bead edgartools-0q0d)xbrl.us roles matched neither of two independent matchers, so Apple's FY2010 securities purchases showed +57,793M. Both paths now use is_negated_label_role; Apple FY2023 goes from 0 to 130 negated facts. (GH #1249, bead edgartools-hxnw)Statements.to_dataframe() and Statement.to_dataframe() returned different signs for the same line. The sign now lives in one function called by both; 13 rows of Apple's FY2023 cash flow change. to_dataframe(presentation=False) returns the raw filed values. (GH #1248, bead edgartools-5ztr)get_structured_statement(fiscal_period='FY') could return the prior-year comparative. A 10-K tags every instant fp=FY, and the filing-date tiebreak never fired. Period end now breaks the tie; Apple's FY2023 balance sheet reports AssetsNoncurrent 209,017,000,000 instead of the FY2022 217,350,000,000. (bead edgartools-i24a, GH discussion #946)select_periods(max_periods=4) returned 10 periods; balance-sheet views admitted empty incidental instants; parenthetical=True returned the face statement. The cap is applied, instants are fact-checked, and a filing with no parenthetical role raises StatementNotFoundError. (GH #1221, #1222, #1246, beads edgartools-rj4b, edgartools-zt8u, edgartools-z0y2)Statement helpers had never returned a correct result for any filing. analyze_trends() read a key that was never emitted, calculate_ratios() looked up reversed concept names, and to_dataframe(matrix=True) read every cell from standard_concept. All three now return filed values; Apple's FY2023 equity matrix goes from 33 null cells to its filed balances. (GH #1240, #1241, #1244, bead edgartools-ysr8)get_financial_metrics() rendered and converted the same three statements once per metric. The three getter helpers share one _render_statement_frame() with a per-call memo, so a call renders 3 times instead of 14–18. JPMorgan's warm median drops from 95.2 to 24.8 ms with every metric byte-identical. (GH #1283, bead edgartools-y4yp.4)It backs Filing.search() , and its old backend is deleted in 6.0, so it needed a replacement rather than a deprecation. Chunks are cut at headings wit…
Twenty-seven entries, led by the XBRL dimensional and calculation models. Both were keyed too coarsely — axes and domains by concept when they belong to a role, calculation edges by concept when a concept can roll up into two totals — so each returned one answer where the filing gives several.
Upgrading. Several of these change values you already read, all in the direction of returning more or refusing to guess:
xbrl.axes and xbrl.domains keep their shape but are now a union across roles rather than whichever role was parsed last. Use the new xbrl.axes_for_role(role_uri) / xbrl.domains_for_role(role_uri) wherever the role is known — the flat views cannot distinguish a domain's two-way breakdown on the income statement from its five-way one in a revenue note.calculation_linkbase() returns more rows: one per filed relationship rather than one per concept. Coca-Cola's FY2023 10-K goes from 278 rows to its filed 306. Code that assumed one row per concept will now see a concept twice, with a different parent and weight — which is what the filing says.get_revenue(unit=...) and get_gross_profit(unit=...) return None for a unit the company does not report in, where they previously returned a USD figure. If you passed unit='EUR' and got a number, that number was dollars.<div> is no longer dropped and nested text is no longer duplicated. related_fact_ids now hold fact ids rather than locator labels, so get_footnotes_for_fact() resolves on filings where it previously returned nothing.Report.get_dataframe() returns data instead of an empty DataFrame, and Document.chunks() no longer raises ModuleNotFoundError. Both were broken for every filing.CurrentReport.doc now warns; it was the last un-warned public route into edgar.files, which 6.0 deletes. Use .document, or .items / report[item].pip install --upgrade edgartoolsXBRL.axes_for_role(role_uri) and XBRL.domains_for_role(role_uri) return the axes and domains a single extended link role declares, keyed on element ID. Prefer them over xbrl.axes / xbrl.domains wherever the role is known: the same domain routinely carries different members in different roles, so the flat views can only offer the union. Axis and Domain gain role_uri, and XBRL.axes_by_role / XBRL.domains_by_role expose the underlying store.CalculationTree.all_arcs, one entry per filed calculation relationship. A concept may roll up into two totals with a different weight, and often a different sign, under each, which the one-node-per-concept all_nodes map cannot represent. all_nodes is unchanged and still the right thing for lookup and membership; all_arcs is the graph.ElementCatalog.typed_domain_ref and ElementCatalog.substitution_group, read from the element declaration, which is what lets Axis.is_typed_dimension report a dimension the filer declared as typed.get_revenue(unit=...) answered a request for one unit with a figure in another. When the unit filter correctly rejected every candidate fact, _get_standardized_concept_value fell through to its component calculation, and that calculation checked the two components against each other — which establishes only that they can be added — before normalising with no target unit at all. Apple's GrossProfit and CostOfGoodsAndServicesSold are both USD, so their sum was returned as the answer to unit='EUR' and to unit='shares' alike: 364,357,000,000 in both cases, and both now None. The default USD path is unchanged and still answers from the revenue fact itself. Strictness now reaches the fallbacks, and a derived figure has to answer the unit that was asked for — an exact match always does, a merely compatible one only where the caller did not pin the unit, which is the rule the direct path already used. get_gross_profit() was leaking the same way for a second reason — its own direct branch never asked for strict matching at all, so a compatible-but-different currency succeeded before the guard was even reached — and both fallbacks now answer by the same rule as the direct path. This was latent until bare concept names were indexed (GH #1202, same release): before that, get_fact('GrossProfit') returned None, the components were never found, and the fallback returned None for the wrong reason. (bead edgartools-885h)
Report.get_dataframe() returned an empty DataFrame for every filing. FinancialTableExtractor read table_node._processed, an attribute that exists only on the legacy TableNode, while the module parses with the modern one — so every call raised AttributeError into a bare except that returned an empty frame. Measured before the fix: 642 of 642 tables across 21 filings raised, and 0 of 42 Apple R-files produced data; after, 41 of 42, with Apple's Q2 FY2025 net sales reading the filed $95,359M and $219,659M. Two further defects surfaced once the code could run: a short period-header row was padded on the left, attributing every figure to the wrong period, and AUD without a word boundary matched inside "Unaudited", so an unaudited USD statement reported Australian dollars. The money scale now prefers the currency-qualified phrase, since "shares in Thousands, $ in Millions" was read as thousands. (bead edgartools-yq1l)
Document.chunks() raised ModuleNotFoundError on every call. It imported edgar.documents.extractors.chunk_extractor, a module that was never written, so this public retrieval API — token-budgeted chunks with overlap, the shape a RAG pipeline asks for — has been unusable since the parser rewrite, along with the DocumentChunk class that was only reachable through it. The module now exists; overlap >= chunk_size raises ValueError rather than advancing the window by zero and looping forever. (bead edgartools-vwtb)
Two taxonomies whose namespace URIs end in the same path segment became one concept. A concept is identified by its expanded name — namespace URI plus local name — and the prefix is only a display choice, but when a namespace was declared on the fact element rather than the instance root the parser fell back to deriving the prefix from the URI's final path segment. {http://example.com/l01/alpha/2024}Collision and {http://example.com/l01/beta/2024}Collision both became 2024:Collision: both facts survived, their technical identity did not, and grouping by concept silently merged concepts from different taxonomies. The prefix a document declares is now read wherever it declares it, and a derived segment already claimed by a different namespace is suffixed rather than shared. Real filings declare their namespaces on the instance root, so nothing changes for them — Apple FY2023, Microsoft FY2024 and Coca-Cola FY2023 return the same 1,640 concept strings over 6,138 facts as before, byte for byte. Separately, Axis.is_typed_dimension and Axis.typed_domain_ref were declared on the model and written by no code path at all, so every dimension read as explicit; xbrldt:typedDomainRef is now carried from the element declaration through the element catalog onto the axis. An axis declared only in an imported taxonomy still reads as explicit, because following xs:import into a remote taxonomy is a separate capability that remains open. (GH #1188, #1234, bead edgartools-0c1q.16)
Footnote-to-fact links resolved to nothing, and footnote text was truncated and duplicated at once. A footnoteArc's xlink:from names a link:loc whose href fragment is the fact id, and edgartools stored the locator label instead — so Coca-Cola's FY2011 10-K put six ids on its footnote that name no fact, no fact carried a footnote back, and get_footnotes_for_fact() returned nothing for a fact that has one. This resolved almost everywhere by coincidence, because 18,686 of the 18,699 footnote locators in the fixture corpus label a locator with the fact id verbatim; the 13 that do not are in Coca-Cola FY2011 and JPMorgan FY2012, and both are now fully linked. Footnote text was built from descendant <div>s whenever any existed, so direct-child and sibling text was dropped, and since findall('.//div') returns nested divs alongside their ancestors while itertext() already descends, nested text was emitted twice; a footnote in Golub Capital's Q1 FY2025 10-Q lost its opening sentence and its entire table while repeating one sentence, reading 187 words against 352. Text is now assembled in one ordered walk with a break after each block, so words no longer glue across boundaries. Footnote resources are also matched within their own footnoteLink rather than a document-wide dict, since two extended links may legitimately reuse a label — a collision returned a real footnote belonging to a different fact. Of the 123 footnotes across the three fixtures that exercise this, only the 3 with nested markup change; the rest are byte-identical. (GH #1169, #1190, #1230, beads edgartools-0c1q.15, edgartools-05gk)
calculation_linkbase() dropped a concept's second calculation parent, and the weight that came with it. Calculation relationships were stored one node per concept, so a concept rolling up into two totals in the same role kept only whichever edge the traversal reached last — and calculation_linkbase() documents its result as one row per parent-to-child relationship. Multiple parents are ordinary rather than exotic: Coca-Cola's FY2023 10-K returned 278 rows against 306 filed arcs, Apple's 194 against 215 and JPMorgan's 438 against 456, and each returns exactly its filed count now. The sign was the sharper half — Coca-Cola files OtherComprehensiveIncomeAvailableforsaleSecuritiesAdjustmentNetOfTaxPortionAttributableToNoncontrollingInterest into one total at +1 and another at -1, so the single row that survived reported a weight whose sign depended on traversal order. CalculationTree now carries all_arcs, one entry per filed edge, beside the all_nodes map that callers use for lookup and membership. The same error at the tree level is fixed with it: a role split across two calculationLink elements, which is legal XBRL, had its second element replace the whole tree built from the first, so relationships are now accumulated across every element and file before the trees are built. (GH #1184, #1238, bead edgartools-0c1q.14)
The dimensional model was not role-scoped, so a domain reused across roles kept only one role's members. Hypercubes were stored per extended link role, but axes and domains sat in dicts keyed on element ID alone with no role on either model, and the members of a domain were assigned unconditionally once per role. Apple's FY2023 10-K splits srt:ProductsAndServicesDomain two ways on the income statement and five ways in the revenue note; a single global object held whichever was parsed last, so one of those two roles was always answered with the other's breakdown. Axes and domains are now kept per role, reachable through xbrl.axes_for_role() and xbrl.domains_for_role(), and the existing xbrl.axes / xbrl.domains keep their shape as a union across roles rather than a truncation. Five further gaps in the same loop close with it: xbrldt:targetRole is followed across extended links, so a hypercube whose dimensions are declared in another role is no longer reported as having none; dimension-default arcs set Axis.default_member_id, without which a fact at the axis default cannot be told from an undimensioned one; Table.closed and Table.context_element are read from the all arc instead of being hardcoded; Domain.parent is written for a member that is itself a parent; and axes and members follow the filed order attribute, which was already being extracted and then discarded. Measured on Exxon's FY2022 10-K: 49 of 49 hypercubes now report the closed="true" they were filed with against 0 before, all 39 axes carry their default member against 0, and 37 domains carry a parent against 0. (GH #1171, #1194, #1219, #1225, #1235, #1236, bead edgartools-0c1q.13)
get_statement_facts() returned nothing for concepts that are demonstrably in the statement. A fact's statement membership was read from whichever presentation role the enrichment loop reached first and stored as a single scalar, but membership is a set — us-gaap:NetIncomeLoss is presented in the income statement, the cash flow statement, the statement of equity and more. Apple's FY2024 10-K returned 0 rows for NetIncomeLoss in CashFlowStatement, ComprehensiveIncome, StatementOfEquity, Notes and Disclosures, and 6 in each after the fix; statement row counts rise correspondingly, for instance StatementOfEquity from 25 to 63. The full membership is now accumulated once into an index — strictly less work than the per-fact tree scan it replaces — and by_statement_type() tests membership rather than equality. The scalar statement_type keeps its meaning as the primary statement, and the membership lists stay off the declared DataFrame schema. (GH #1242)
get_facts_with_dimensions() removed the dimension columns it selected rows on. The predicate matches a fact because it carries dim_* keys, then to_dataframe()'s projection dropped every one of them because include_dimensions was left falsy, so the caller got the right rows with the information that made them the right rows deleted. Intel's FY2024 10-K returned 1,135 correctly-selected rows and 0 dimension columns; it now returns the same rows with all 56. (GH #1243)
pivot_by_dimension() silently returned a table that was not a pivot. The lookup column was built by interpolating the caller's spelling, while the projected column is always normalised, so the natural QName form us-gaap:AwardTypeAxis missed and the method fell through to returning the plain unpivoted frame — a DataFrame of plausible shape that a caller who does not already know the pivoted layout cannot distinguish. All three spellings (dei_LegalEntityAxis, dei:LegalEntityAxis, LegalEntityAxis) now resolve to the same pivot, and an axis no fact carries raises ValidationError naming the available axes instead of returning an empty frame that reads as "no data". Both pivot methods also now keep colliding rows apart by unit where that is what distinguishes them, and report any collision they cannot resolve rather than quietly keeping one row of each. (GH #1223)
Unitless facts were classified as numeric, so numeric queries returned document metadata. float(value) ran on every fact unconditionally, so dei:DocumentFiscalYearFocus — a gYear with no unit — carried a numeric_value of 2024.0 and a query for facts valued between 2020 and 2030 returned it alongside real monetary facts. XBRL requires a unitRef on numeric items and forbids one on non-numeric items, which makes the unit an exact test where the element catalog is not reachable. Five to seven facts per filing are reclassified — CIK, fiscal-year focus, area code, auditor firm ID, postal code — with every genuine numeric fact untouched. numeric_value is what consumers read as "this is a numeric fact", so this reached well beyond by_value(). (GH #1220)
Inline XBRL facts from parse_html() carried no resolved context or unit. _extract_contexts and _extract_units ran namespace-aware XPath (//xbrli:context) against a tree built by lxml.html.fromstring(), whose elements carry the literal tag "xbrli:context" and no namespace URI, so the lookups matched nothing and every fact's .context and .unit came back None while the raw refs sat on the fact untouched. Matching on the local name instead resolves 180,747 of 180,747 context references and 169,255 of 169,255 unit references across the 60 iXBRL filings in the fixture corpus, where the previous count was zero of each. A <unit> built from xbrli:divide now resolves to USD/shares rather than to plain USD, and dimensional members are read from xbrli:scenario as well as xbrli:segment. (GH #1232)
A fact's value was the text before its first child, and nothing else. _get_fact_value read element.text, so descendant markup, tail text and the entire continuedAt chain were dropped and the truncated string looked like a complete value — the mechanism behind four separate reports. Relevant content is now assembled properly: descendants and their tails contribute, an ix:exclude subtree does not while the text following it does, escape="true" serializes child markup instead of flattening it, and continuation chains are followed in order with a cycle guard. ix:footnote and ix:continuation are resources rather than facts and are no longer routed through the fact extractor, which removes 451 blank fake facts from the same corpus. Since continuedAt is how issuers split long narrative disclosures, the truncation fell on exactly the text an LLM reads. (GH #1189, #1237, #1239)
Unsupported inline transformations passed human display text through as the fact value. The registry held five entries and matched them on an exact QName, so it missed even ixt:num-dot-decimal — the TR4 spelling that accounts for 68,880 of the 90,000 format attributes in the fixture corpus — and returned "Delaware" where the value is "DE" and "ten years" where it is "P10Y". It now implements Transformation Registry 1–4 and the SEC's ixt-sec extensions keyed on a version-independent name, covering every format those 60 filings use, and a format that cannot be applied is recorded on the fact and logged rather than silently yielding display text. Scale is computed with Decimal: applying it in binary float wrote the error into the lexical value, so a filed 0.7 at scale="-2" became 0.006999999999999999 and one filing reported 8199999999.999999 for $8.2 billion. (GH #1250, #1251)
Document.has_xbrl reported False on documents with thousands of extracted facts. The inline pre-process pass stores its facts on metadata.xbrl_data, while Document.xbrl_facts called a separate extractor that scanned the node tree for ix_tag metadata the inline pass never sets — two unconnected pipelines, so the public surface answered from one that was never populated. This is worse than missing data, because a working extraction looked like an absent feature and a caller correctly branched away from it. The property now reads the pass that is populated, and has_xbrl is True for all 60 iXBRL filings in the fixture corpus against zero before. (GH #1233)
Every XBRL duration was reported one day short. XBRL 2.1 reads a date-only endDate as the end of that day, so the context runs to 24:00 on it and the final day belongs to the period; reporting_periods[*]['days'] subtracted the two dates and dropped it. A calendar year of 2023-01-01 to 2023-12-31 measured 364, Apple's 53-week FY2023 measured 370 instead of 371, and a duration context covering a single day measured 0. The count now lives in one place, edgar.xbrl.core.duration_days, used for the stored field and for the two sites in period_selector that recompute a duration and compare it against that field. periods.py keeps its own exclusive count, because its bucket bounds were calibrated against it and a 53-week fiscal year sits exactly on the <= 370 bound - it no longer writes that local value back over the public field, which is what made the field mean different things depending on whether a statement had been rendered first. Measured over the 8 filings under data/xbrl/datafiles: 191 duration periods change their days value and nothing else changes - no period_type, fiscal_period, label or key, and no change to the periods selected for any statement. (GH #1247)
Table rows were dropped at random. <thead> rows were deduplicated by id(), but nothing kept the lxml element proxies alive — lxml frees a proxy once unreferenced and CPython reuses the address, so a tbody row could inherit a stale thead id, match it, and be skipped as "already processed". The row vanished from the table with no error. Measured on the ABRAMS Form 4 0000001923-04-000001: 1 to 22 divergences per 25 renders before, 0 in 100 after. Membership is now tested by identity against a list held for the whole function, so the collision is impossible by construction rather than dependent on GC timing. This is what made the regression lane fail at random. (bead edgartools-gf6v)
EntityFacts lookups reported that present data did not exist. The fact index was keyed by the qualified name a fact is tagged with (us-gaap:StockholdersEquity) and by the lowercased label, never by the bare concept name — so get_annual_fact('StockholdersEquity') returned None and warned "No fact found", while the value was present and correct. get_annual_fact('Assets') worked only because us-gaap:Assets happens to be labelled "Assets" and matched the label key by coincidence, which is what made this look intermittent rather than systematic: in Snowflake's company facts that coincidence holds for 4 of 339 concepts. The local name is now indexed too, in its filed case — a lowercased form would merge with existing label keys. Across AAPL, JPM and KO, 1,954 of 4,264 concept and label lookups that previously answered None now return a fact, with no existing answer changed or lost. get_fact() and available_periods() shared the defect and share the fix. (GH #1202)
A filtered or paged Reports collection returned reports whose content could not be read. Report.content fetches its R-file through reports._filing_summary._filing_sgml, and three of the four methods that return a new collection rebuilt it without that back-reference, so filter(), next() and previous() handed back reports with the right filenames and raised AttributeError: 'NoneType' object has no attribute '_filing_sgml' on .content. On Apple's FY2024 10-K a multi-row filter now reads R3.htm as the 79,891 characters get_by_category() always returned. Only the multi-row branch of filter() was affected — filtering to a single report returns a Report still attached to the original collection, which is why this stayed hidden. Deriving a collection now goes through one place, so the next method to return a subset carries the context by default. (GH #1191)
FactQuery.scale() was a silent no-op on every parsed fact. Transforms were applied to fact['value'], which holds the filed string, so scale()'s own isinstance(value, (int, float, Decimal)) guard rejected it and handed it straight back — while numeric_value, the float every consumer actually reads, was never transformed on any path. scale(1000), the example in its own docstring, changed nothing. A numeric fact is now transformed on its numeric value and both fields are written together so they cannot disagree; a non-numeric fact keeps being transformed on value, so text transforms still work and scale() leaves TextBlocks alone through the same guard. StitchedFactQuery duplicated the defect and shares the fix. (GH #1187)
FactQuery.to_dataframe() served a stale table after the query was narrowed. The cache was keyed on the column projection alone, so the first to_dataframe() answered every later one: after .by_concept('us-gaap:Assets'), execute() returned 2 rows while to_dataframe() still returned the original 1,075, with nothing to say the two disagreed. FactQuery is a fluent mutable builder whose methods return self, so one object describes a different population after each call, and the key now covers the query configuration as well as the projection. Identical queries are still served from the cache. StitchedFactQuery duplicated this one too. (GH #1186)
A disclosure whose subject is notes payable was routed to notes(). get_all_statements() falls back to keyword matching when a role carries no FilingSummary menu category, and that fallback tested the bare substring "note" before it tested "disclosure". "Note" names a financial-statement section and a debt instrument, so Oracle's four Disclosure - NOTES PAYABLE AND OTHER BORROWINGS roles (10-Q 0000950170-23-047713) reported type="Notes" / category="note", came back from xbrl.notes(), and were absent from xbrl.disclosures(). The ambiguous word no longer outranks a role that states what it is - by its Disclosure category marker, or by the concept it hangs from (us-gaap_DebtDisclosureAbstract) - unless the definition names the notes section itself, so Notes to Consolidated Financial Statements and Note 1 - Organization still classify as notes. Roles whose definition does not contain "note" are classified exactly as before. (GH #1207)
Filing.sections() is chunked by edgar.documents instead of the legacy parser. It backs Filing.search(), and its old backend is deleted in 6.0, so it needed a replacement rather than a deprecation. Chunks are cut at headings with each table its own chunk, capped so no chunk swallows a whole Item. Measured across 41 era-stratified fixtures the two chunkers index the same words to within 0.10%, and across 375 phrase queries mean recall@5 is unchanged at 99.2%. (bead edgartools-07lk.3)CurrentReport.doc warns; it was the last un-warned public route into edgar.files. It returns the legacy ChunkedDocument, while CompanyReport.doc on every other report class returns the modern edgar.documents document — so the same attribute name yields two unrelated types depending on the form, and a survey starting at the base class reads .doc as already migrated. Use .document for the parsed document, or .items / report[item] for item access. Measured across 133 era-stratified fixtures and 2,128 item lookups, no internal code path reaches the legacy parser through the report classes any more. (bead edgartools-07lk.3)6.0 deletes the edgar.files package, and these three were the user-reachable modules with no deprecation, so code importing them would have got no not…
Thirteen fixes, led by two clusters: period identity now comes from the duration a fact covers rather than the SEC's filing-focus label, and three typed fields on Form D / Form 144 that were silently wrong.
Upgrading: Form144.nothing_to_report is Optional[bool] now, so callers comparing it against the string "N" need updating; period_length is honoured rather than ignored, so callers who passed it were reading annual figures under a quarterly label and will see different numbers. edgar.effect and edgar.form144 still work but warn — they now live at edgar.offerings.effect and edgar.ownership.form144, and edgar.files.markdown / .tables / .text warn ahead of 6.0 removing them.
pip install --upgrade edgartoolsedgar.muniadvisors is a package rather than a single module. The MA-I parser moved to edgar/muniadvisors/core.py, matching edgar/ownership/ and edgar/entity/. Nothing changes for callers: a module converting to a package of the same name keeps the import path identical, and all 19 of its public names still resolve to the same objects — __all__ named only MunicipalAdvisorForm, so a naive re-export would have dropped the other 18. (bead edgartools-07lk.12.1)[tool.hatch.build].include opens with an unrestricted edgar/**/*.py, so every install also downloaded the evaluation harnesses, entity training scripts and demos under edgar/ — about 8,900 lines across ai/evaluation/, ai/examples/, entity/training/ and thirteenf/demo_comparison.py. They are excluded now; ai/exporters/ stays, since export_skill is public API. A regression test asserts that no shipped module imports any excluded path, which is the property that makes the exclusion safe rather than the exclusion itself.CurrencyConverter reported "no rates found" when extraction had crashed, and an MA-I summary whose disclosure block failed to parse read as a clean compliance record. All three now log a warning and carry the failure — XBRL.period_validation_unavailable, CurrencyConverter.extraction_warnings, and an explicit line in the advisor summary. No exception propagates that did not before._require_element() helper raising ValidationError with the element name and the namespace hint that explains most of these failures. ValidationError IS-A ValueError, so anything catching the old raise still catches this one.edgar.files.markdown, edgar.files.tables and edgar.files.text now warn. 6.0 deletes the edgar.files package, and these three were the user-reachable modules with no deprecation, so code importing them would have got no notice at all — just an ImportError on upgrade. They warn for external callers and stay silent for edgartools' own internal use. edgar.files.styles is deliberately left unwarned: parse_style runs thousands of times per document and the frame check is not free, and its entry points (Document, SECHTMLParser) already warn.edgar.effect and edgar.form144 moved, with the old paths kept until 6.0. Effect now lives in edgar.offerings.effect, beside the S-1/S-3/S-4, 424B and Form C/D parsers — a notice of effectiveness is a registration-lifecycle document. Form144 now lives in edgar.ownership.form144, beside the Forms 3/4/5 parsers — Form 144 is an affiliate's notice of a proposed sale. Neither name was in edgar.__all__ nor re-exported at the top level, so the module path was their whole public surface; both old paths still work and emit a DeprecationWarning naming the new one. The objects are the same objects, not copies, so isinstance checks against either path agree. filing.obj() routes to the new paths, so ordinary use never triggers the warning. (bead edgartools-07lk.12.1)period_length on EntityFacts.income_statement() and .cash_flow_statement() warns before 6.0 removes it. period= is the supported spelling and the only one that can also ask for 'ttm'. The parameter is honoured now rather than ignored — see Fixed below. (GH #1177)edgar/files/html_documents_id_parser.py. 687 lines and three classes (AssembleText, ParsedHtml10K, ParsedHtml10Q) with no callers anywhere in the library or the tests. The id_parse_document path it once served survives only in comments.edgar/tools/ (an empty package), edgar/analysis/ and edgar/xbrl/analysis/ (both already emptied of their modules, leaving only stale caches), and edgar/reference/financials.py (a seven-line scratch script with a __main__ block and no importers). Nothing in the library or the tests referenced any of them.cash_flow_statement(period='quarterly') reported Snowflake's PaymentsToAcquireInvestments for Q3 FY2022 as 3,042,396,000, the nine-month cumulative total, against a discrete quarter of 1,053,763,000 — a 2.89x overstatement. Two causes: comparative re-filings were discarded as inconsistent, which removed the six-month operand the quarter is derived from, and the quarterly path had no duration filter to mirror the annual one. (GH #1180)TTMCalculator.quarterize() returned periods that were not quarters. Deriving a quarter pairs a cumulative fact with an earlier one, and when the two came from different fiscal cycles the subtraction produced a real number over a nonsense interval: Snowflake's goodwill yielded -45,711,000 across 457 days labelled Q2 FY2025. Everything quarterize() returns is now checked to span a quarter. (GH #1197)latest_periods(1) returned three years of facts. Periods were keyed on (fiscal_year, fiscal_period), which is the SEC's filing focus rather than the period itself, so the 2022, 2023 and 2024 comparatives in Logistic Properties' 20-F — all tagged fy=2024, fp=FY — counted as one period. Periods are now keyed on the dates the facts cover, and annual-ness comes from the duration. (GH #1185)discrete_quarters=True, a concept present on the 12-month period but not the nine-month one had nothing to subtract and kept its cumulative value, while the period was relabelled a quarter regardless. Meta's FY2024 "Deferred income taxes" changes concept between the 10-K and the 10-Q, so $(4.738)B was shown as Q4. A quarter that cannot be derived is now dropped, so the cell reads empty rather than wrong. (GH #1179)recipient.address.zipcode was always None: the parser looked for a child element named 30361 — the sample ZIP from the code's own docstring, pasted in where the tag name belongs. The four recipients in the 1685 38th REIT notice now read 30361 as filed. Related-person addresses were never affected. (GH #1193)FormD.is_new was inverted, so every base notice rendered as FORMD/A and every amendment as FORMD. The field was assigned the filing's isAmendment answer, which asks the opposite question — AP Fund IV, a base D, reported is_new=False. to_context() now takes its heading from submission_type, so the heading cannot contradict the field beside it. (GH #1192)Form144.nothing_to_report returned a truthy "N". It carried the raw SEC flag against a bool annotation, and bool("N") is True, so if form.nothing_to_report: read a notice reporting a sale as reporting none — accession 0001958244-23-000454 has one prior-sale row. Y/N now parse to True/False, and an unanswered flag is None rather than a default. Callers comparing against the string need updating. (GH #1195)Document.parse(...).to_markdown() emitted four warnings, two naming modules the caller never touched. It now skips only the warning's own module.period_length was accepted, documented, and never read. EntityFacts.income_statement(period_length=3) and .cash_flow_statement(period_length=3) returned the annual statement with no warning — the parameter appeared nowhere in either method body and had not since the Facts API landed. It now selects the period it names, and a value contradicting period= or annual= raises ValidationError instead of being silently resolved. Numbers can change for callers who were passing it: they were reading annual figures under a quarterly label. (GH #1177)XBRLS.get_statement(use_optimal_periods=False) crashed instead of returning a statement. The non-optimal path appended the (statements, role, statement_type) tuple from XBRL.find_statement() to the list handed to StatementStitcher, which then called .get() on it and raised AttributeError: 'tuple' object has no attribute 'get'. It now calls get_statement_by_type(), the same accessor the optimal path uses, and skips filings that lack the statement type. (GH #1173)XBRLS.query(standardize=False) still standardized. The option was stored under the keyword standardize and read back as standard, so it never reached get_statement() and every query came back with standard-concept mapping applied. Both spellings are now accepted. (GH #1172)FactQuery.transform() and .scale() mutated the shared fact cache. Transformed values were written back into the row dictionaries returned by get_facts(), which come from the facts view's cache, so .scale(1000) scaled the cache itself and a second identical query returned values divided by a million. Rows are now copied before transformation. StitchedFactQuery had the same defect. (GH #1175)find() did not recognize tickers containing X. The ordinary-ticker pattern was ^[A-WYZ]{1,5}([.-][A-Z])?$, a character class excluding X in every position rather than only the trailing position that marks a mutual fund, so find("XOM") returned CompanySearchResults instead of the Company that Company("XOM") resolves. The ^[A-Z]{4}X$ fund pattern is now tested first and the ticker pattern admits the full alphabet. (GH #1178)$3. Many filers render a fee-waiver footnote as its own single-row table, and its prose contains "waive its management fees", so _classify_table labelled it operating_expenses on the keyword alone. That inflated the fee-table count, which routed the expense example to the single-column extractor, which parsed the digit out of the header string 3 Years — Ocean Park High Income ETF reported expense_1yr=3 against the filed $109/$381. Prose-only tables no longer fall through into the data classifiers, fee values are shape-validated without rejecting SEC footnote formats, and column alignment is preserved for indented operating-expense and shareholder-fee tables. Contributed by @sambai-dev. (GH #912)Six fixes, two of them contributed by @hashd1ve ( #1144 ) and @biautomator ( #1151 ). Thanks to both.
Six fixes, two of them contributed by @hashd1ve (#1144) and @biautomator (#1151). Thanks to both.
Upgrade note: FundSeries.get_filings() now returns that series' own filings rather than the whole trust's, so counts drop and filings[0] changes — the previous result could belong to a sibling fund. Fund.get_filings() without series_only is unchanged.
detect_page_breaks() and mark_page_breaks() warn before 6.0 removes them. No replacement: the edgar.documents parser treats page-break <hr>s and page-number containers as print chrome and discards them, so page-break positions are not part of the supported document model. Internal callers stay silent, so a suite running -W error is unaffected unless it calls these itself.get_financials() no longer goes quiet on a filing with no XBRL. It returned a truthy Financials whose income_statement(), balance_sheet() and cash_flow_statement() all answered None — Company(104599), Circuit City's 2008 10-K, is the case. It now emits a FutureWarning and will raise XBRLFilingWithNoXbrlData in 6.0; set EDGARTOOLS_STRICT_ERRORS=1 for that behaviour today. filing.xbrl() still returns None quietly, which is a true absence rather than a failure.get_revenue() returned a years-stale value after a company migrated its XBRL tag. The first variant with any matching fact won, regardless of period, so NVIDIA answered FY2022 (26,914,000,000) instead of FY2026 (215,938,000,000). Candidates are now gathered across all variants and ranked by recency. Every standardized getter built on the same helper was exposed. (GH #1149)get_revenue() returned the ASC-606 slice for insurers and banks. MetLife tags both the contract-revenue slice (2,436,000,000) and consolidated Revenues (77,084,000,000) for FY2025, and the slice ranked first — a 32x understatement, with an income statement showing less revenue than operating income. A same-period candidate that dwarfs the ranked pick is now taken as the consolidated total. Revenue only: net income's variants are not slices of one another. The statement builder held a second copy of the priority list and now shares one threshold with the getter.edgar_trends returned whatever concept contained the word. Concepts were matched as substrings, so "Revenue" also matched CostOfRevenue and DeferredRevenue; rows were truncated before the period-type filter, so the correct annual fact could be cut from the window; and quarterly facts labelled FY entered annual series under the same year label as the real annual figure. Apple's revenue read as falling from $383B to $7.7B. Concepts are now matched exactly, filtered before truncation, and period type comes from the reporting window rather than fiscal_period. Revenue trends take the same consolidated-total cross-check as get_revenue(). (GH #1138)FundSeries.get_filings() delegated to the fund company, and a trust files one report per series under one CIK, so every series returned the same list and filings[0] belonged to whichever sibling filed last: Vanguard's Extended Market Index fund, which excludes the S&P 500 by construction, was handed the 500 Index fund's portfolio. It now resolves its own series through SEC browse-edgar and returns empty rather than falling back to the trust. Fund.get_portfolio() and get_latest_report() took the same trust-wide path and are fixed with it. Fund.get_filings() without series_only still returns the whole trust. (GH #1143)year and quarter accept a list on the fund-series path. A list was untranslatable, silently dropped, and the caller received the series' entire unfiltered history. An unusable pair now raises rather than answering with everything. (GH #1143)edgar_ownership called a fund's unparsed positions "nan", and edgar_screen wrote a bare NaN into its JSON. A missing DataFrame cell is truthy, so if issuer: admitted it and str() rendered it; where nothing guarded it, the NaN reached to_json as a literal no strict parser reads back. One cell reader now lives in tools/base.py. (GH #1137)find_bdc("Ares") — the example in its own docstring — returned nothing. Names were scored with whole-string fuzz.ratio against a threshold of 50, so a short query against a longer registered name missed: ARES against ares core infrastructure scores 28.6. Candidates already come from an exact word index, so they are now scored with token_set_ratio. (GH #1145)edgar_fund(action="portfolio")'s total_value summed share counts, not dollars. None of the candidate columns existed on FundReport.investment_data() — the dollar column is value_usd — so it fell through to balance and summed raw units. VFINX reported a total_value of $14.4B against a top-5-holdings sum of ~$366B. Per-holding records were unaffected. (GH #1148)Both still accept a BeautifulSoup, and FilingHomepage(soup=...) still works, with a DeprecationWarning ; both go in 6.0. Neither is on the path you ta…
httpx is now capped below 0.29. httpx upstream has been dormant since 0.28.1 (December 2024) and stopped accepting issues in February 2026, so any future release on PyPI would be unexpected and should not be adopted automatically. Installs resolve to 0.28.1 exactly as before; the planned successor is the httpx2 migration in 6.0.
Filing homepage parsing is about 9x faster. The filing index page — the source of filing.attachments, filing.homepage.get_filers() and the filing dates — was parsed with BeautifulSoup's pure-Python html.parser; it now uses lxml, measured at 15.8ms to 1.8ms across the five tracked homepage fixtures. Output is unchanged: the whole parse of every fixture, down to each attachment's size and each filer's identification lines, is pinned against a baseline captured from the previous implementation. A blank or truncated index page still yields a homepage with nothing on it rather than raising.
FilingHomepage(...) and Attachments.load(...) take an lxml tree now, from the new edgar.attachments.parse_homepage_html(html). Both still accept a BeautifulSoup, and FilingHomepage(soup=...) still works, with a DeprecationWarning; both go in 6.0. Neither is on the path you take through filing.homepage.
R-file report rendering parses about 20x faster. filing.reports, TenK.reports and Report.view()/.text() parsed each R-file with BeautifulSoup's pure-Python html.parser; they now use lxml, measured at 438ms to 22ms over the 42 R-files of a tracked AAPL 10-Q. End to end the render is 1.8x faster (1217ms to 680ms) — the remaining time is table layout, not parsing. Output is unchanged: the rendered text of all 42 reports is pinned against a baseline captured from the previous implementation, covering both the ordinary single-table path and the embedded-table path added for issue #755.
Note text extraction is about 6x faster. The narrative and table text behind note.to_context() and notes.to_markdown() was parsed with BeautifulSoup's pure-Python html.parser; it now uses lxml, measured at 477ms to 80ms over 16 real note TextBlocks from three filers. The emitted text is unchanged, which matters more here than the speed: this is what an LLM reads. Narrative in both modes, the per-table markdown render, and every table's aligned plain text are pinned character-for-character against a baseline captured from the previous implementation.
DRS underlying-form detection is about 7x faster, and stops warning about the filings it handles. A DRS/DRS-A filing only says "DRS" in its metadata — the real form (S-1, F-1, 20-F, Form 10) has to be read off the cover page — and that read was done with BeautifulSoup; it now uses lxml, measured at 1481ms to 197ms over 12MB of real filings. Modern DRS filings are inline XBRL, and BeautifulSoup printed an XMLParsedAsHTMLWarning for every one of them: three of the eight filings in the new corpus warned before this change and none do now. Detection is unchanged, pinned against a baseline captured from the previous implementation over seven real filings — a genuine S-1, three 8-Ks from 2001, 2004 and 2008, a 20-F and two modern iXBRL filings — plus seventeen inputs written for the ways lxml and BeautifulSoup disagree about text.
Fund reference data resolves its download URL about 14x faster. get_fund_reference_data(), find_fund() and the class/series lookups behind them start by reading the SEC's investment-company series-and-class listing page to find the current CSV, and that page was parsed with BeautifulSoup's pure-Python html.parser; it now uses lxml, measured at 13.1ms to 0.9ms on the live 96KB page. The URL it resolves is unchanged, pinned against a baseline captured from the previous implementation over the real listing page as SEC served it plus forty-one inputs written for the ways lxml and BeautifulSoup disagree. One shape that the SEC page cannot have — a table written with no closing </th>, </td> or </tr> tags at all — now yields the CSV link that lxml recovers from it, where BeautifulSoup found nothing.
Registration fee tables parse about 8x faster, and stop warning about inline-XBRL exhibits. S1.fee_table, S3.fee_table, total_offering_amount, net_fee_due and securities all come from an EX-FILING FEES (Exhibit 107) attachment — or, for pre-2022 registration statements, from the "Calculation of Registration Fee" table inline in the document body — and both were parsed with BeautifulSoup; they now use lxml, measured at 281ms to 37ms over 2.2MB of real exhibits and filings. BeautifulSoup printed an XMLParsedAsHTMLWarning for every inline-XBRL exhibit, which the module suppressed by hand; that suppression is gone because the warning no longer exists. Output is unchanged, pinned against a baseline captured from the previous implementation over the eighteen documents the existing fee-table verification covers — fourteen Exhibit 107 attachments from 2022-2025, two of them genuine inline XBRL, and four whole pre-EX-107 registration statements from 2018-2021 — plus forty inputs written for the ways lxml and BeautifulSoup disagree about text.
Fund identifier resolution is about 8x faster. find_fund() given a class ID resolves it to a company CIK and then reads that company's series listing, and both pages were parsed with BeautifulSoup's pure-Python html.parser; they now use lxml, measured at 45ms to 5.7ms over 187KB of real browse-edgar pages. Output is unchanged: the full parse of twelve committed pages — five fund families from five to fifty series and up to seventy-three classes, a company with no series at all, and an identifier that matches nothing — is pinned against a baseline captured from the previous implementation, alongside thirty-six inputs written for the ways lxml and BeautifulSoup disagree about text.
Note markdown rendering is about 5.6x faster. The LLM-optimised markdown behind note.to_context() and notes.to_markdown() — the pass that merges a filer's lone $ or % cell back into the figure beside it, drops XBRL metadata tables, deduplicates repeated tables and picks each table's title — was built with BeautifulSoup's pure-Python html.parser; it now uses lxml, measured at 1585ms to 282ms over 2.2MB of real note TextBlocks. Output is unchanged, which matters more than the speed here because this is what an LLM reads: all sixteen notes from three filers render byte-for-byte as before, alongside 62 inputs written for the ways lxml and BeautifulSoup disagree about text, tree shape and malformed markup.
One shape changes, in your favour. A table whose <th>, <td> and <tr> are never closed — <table><tr><th>Item<th>Amount<tr><td>Revenue<td>$1,000 — was nested inside its own first cell by html.parser, so every cell swallowed the rest of the table and a two-by-two table rendered as a single column reading "Amount Revenue $1,000". lxml closes the tags the way a browser does and reads back the table the filer wrote. Note this is the opposite direction to the XML leniency issue: here lxml is the forgiving one.
An unreachable fund series/class parser was removed. get_series_and_classes_from_sec() had no callers anywhere in the library, was never exported, and appeared in no documentation; parse_series_and_classes_from_html() was reachable only from it. Instrumenting every public fund entry point — find_fund() by ticker, series ID and class ID, get_fund_with_filings(), and series.get_filings() — confirmed neither ever runs. The live series listing is parsed by _parse_series_table(), which is unaffected. Its test fixture went with it: a April 2025 capture of a page SEC has since restructured, which the surviving parser cannot read.
497K summary-prospectus extraction is about 7x faster. The fee tables, expense examples, performance tables and fund metadata behind Prospectus497K were read through BeautifulSoup; they now use lxml directly, measured at 458ms to 68ms over seventeen real 497K filings. BeautifulSoup was already parsing these with libxml2 underneath, so what goes is its Python object tree rather than the parse itself. Output is unchanged, pinned against a baseline captured from the previous implementation over that corpus — seventeen different fund families across 2012, 2016, 2020, 2023 and 2025 — plus thirty-five inputs written for the ways lxml and BeautifulSoup disagree about text.
Form 10-D header parsing is about 6x faster. TenD.issuing_entity, .depositor, .sponsors, .distribution_period and .security_classes were read with BeautifulSoup's pure-Python html.parser; they now use lxml, measured at 178ms to 28ms over nineteen real 10-D filings. Output is unchanged, pinned against a baseline captured from the previous implementation over that corpus — CMBS, auto lease, auto receivables, RMBS and structured-products trusts, from 2006 through 2025 — plus thirty-two inputs written for the ways lxml and BeautifulSoup disagree about text. One shape that no SEC 10-D has — a class table written with no closing </th>, </td> or </tr> tags at all — now yields the two class names the filer wrote, where BeautifulSoup nested the rows and ran them together.
Fund company and filing lookup is about 8x faster, and edgar/funds/data.py is off BeautifulSoup entirely. The browse-edgar company page behind get_fund_with_filings() and FundSeries.get_filings() — the fund's name, CIK, identifying information, addresses and each page of its filings — was parsed with BeautifulSoup's pure-Python html.parser; it now uses lxml, measured at 70.6ms to 8.5ms over 303KB of real company pages. Output is unchanged: the full parse of six real company pages, carrying up to a hundred filings each, is pinned against a baseline captured from the previous implementation, alongside thirty-one inputs written for the ways the two libraries disagree.
Three answers change. A page whose <td> and <tr> are never closed now reads the row the filer wrote rather than gluing the form type to the cell beside it (N-CSR, not N-CSRDocuments); libxml2 closes the tags where html.parser nested them. A company page with no companyInfo block, or none of the filings table, still raises AttributeError from the same place, but the message names lxml's method rather than BeautifulSoup's. Neither shape occurs on a page SEC actually serves.
Four remaining readers moved off BeautifulSoup. Subsidiary exhibits (TenK.subsidiaries), the SEC forms listing (list_forms()), the bulk-feed directory listing, and 40-F plain-text extraction now parse with lxml. Output is unchanged, pinned against a baseline captured from the previous implementation over the existing fixtures — 153 subsidiaries across three EX-21 exhibits, the forms page, and a real feed directory listing.
R-file concept extraction is about 13x faster. The concept-annotated rows behind ViewerReport.concept_rows — the tie between a rendered R-file row and its XBRL concept id — were parsed with BeautifulSoup's pure-Python html.parser; they now use lxml, measured at 408ms to 32ms over the 42 R-files of a tracked AAPL 10-Q. Output is unchanged: every field of all 620 rows across 44 real reports is pinned against a baseline captured from the previous implementation, including the column-position work behind issues #810, #812 and #818. Two shapes that cannot occur in an SEC-generated R-file — a table nested inside a row's label anchor, and an unclosed <th> — now parse the way lxml recovers rather than the way BeautifulSoup did; both are asserted explicitly.
Filing header parsing is about 7x faster. Filing.index_headers parsed the header page with BeautifulSoup's pure-Python html.parser; it now uses lxml, measured at 460µs to 66µs per header across the tracked header corpus. Output is unchanged — the full parsed model is pinned against a baseline captured from the previous implementation. Malformed input behaves as before, including the IndexError an empty page has always raised.
PressRelease.text() now reads through the modern parser. Output changes slightly, all of it in your favour: the old path leaked raw <img> markup into the text and left & undecoded as a literal "amp", both of which are gone. Across 12 real 8-K press releases no word is lost that is not one of those artifacts. Table cells no longer repeat the $ that filers put in their own column; the figures and the header's units are unchanged.
PressRelease.to_markdown() now renders through the modern parser too, and its images finally resolve. .text() moved to edgar.documents already; the markdown view was the last part of a press release still going through the legacy edgar.files stack, via MarkdownContent.from_html -> HtmlDocument. It delegates to Attachment.markdown() now, the same supported renderer Filing.markdown() uses. The visible win is images: the old path emitted a root-relative  that resolves to nothing, and you now get the absolute archive URL. Headings and bold survive the trip where they used to be flattened, and no content is lost — across four real press releases from two filers the word counts match to within half a percent, the difference being GFM escaping and the added structure. An attachment with no usable HTML renders an empty panel rather than raising the AttributeError the old path died with.
This was the last production caller of HtmlDocument, which is now referenced nowhere outside edgar/files/ itself, so edgar/_markdown.py no longer imports edgar.files at all — a step towards that package's removal in 6.0. What still reaches into the legacy stack is get_clean_html, at two call sites, both of them behind the already-deprecated include_page_breaks=True flag. A static check pins the edgar/_markdown.py half: if the dependency comes back, the suite fails.
SixK exhibit text now reads through the modern parser. _get_exhibit_content rendered 6-K exhibits with the legacy edgar.files document; it now uses parse_html(html).text(), the same renderer behind the modern document's own repr. Verified over fifteen real 6-K filings from 2025Q2 by comparing words rather than lengths: prose is equal or better, and the residue is dominated by legacy's own defects, which the modern parser does not have — words split at a line wrap ("forward-" / "looking") and glued neighbours ("visitwww.aunainvestors.com"). Nothing is truncated on even a 3.2MB exhibit. This landed deliberately after the <br> and wide-table fixes below, so cover-page line breaks and full table columns were already in place.
Pre-2013 13F TXT infotables silently lost rows to column bleed. When a filing's data lines do not honour the <S>/<C> marker-line offsets, the fixed-width CUSIP slice started past the true value or picked up trailing digits from 12-digit zero-padded values, and the 9-character gate then rejected every such row: Gilder Gagnon Howe's 2008-Q4 filing parsed 5 of its 217 declared entries. The marker spec is now treated as a hint — when the exact slice fails, a tolerance window is searched for a checksum-valid candidate, with token-boundary starts outranking mid-token windows. Recovered against cover-page Entry Totals confirmed at SEC: 192 (was 15), 194 (was 20), 215 of 217 (was 5), and Berkshire's 2008-Q4 regains its lost Wellpoint row, 108 of 108. The parser also reconciles its row count against the filing's own Entry Total and warns on shortfall, so any future silent-loss mode surfaces instead of emitting plausible wrong data. (GH #1072)
edgar_ownership(analysis_type="fund_portfolio") listed no positions for any fund. _get_fund_holdings iterated ThirteenF.holdings directly, and iterating a DataFrame yields its column names, so every attribute probe missed and nothing was appended. On Berkshire Hathaway's Q2 2026 13F-HR (0001193125-26-352200) the MCP returned holdings_count: 29 beside holdings: [], hiding Apple's 227,917,808 shares at $65,950,296,923. (GH #1136)
eightk["Item 9.99"], twentyf["Item 20"] and fortyf["No Such Section"] answered None in silence. A lookup for an item a filing does not have is supposed to say so: it emits a FutureWarning naming what edgartools 6.0 will do, and raises SectionNotFoundError today if you set EDGARTOOLS_STRICT_ERRORS=1. CompanyReport.__getitem__ does that, and TenK/TenQ do it — but CurrentReport (8-K), TwentyF (20-F) and FortyF (40-F) each overrode __getitem__, rewrote the miss path, and dropped the call. All three returned a bare None with no warning in either mode, so a 20-F or 8-K user got no migration notice at all, and report[item] would not have started raising in 6.0 for those forms as documented. All three now behave as the base class always did, verified against real filings of each form in both error modes.
.get() is unaffected and stays silent, as it must — it promises a default rather than a complaint, and it is the migration target. If a bare None is what you want from these lookups, report.get("Item 9.99") keeps giving it to you without the warning.
The three were independent mistakes with one cause: the base class's report_lookup_miss call is invisible at the override site, so each reimplementation of the miss path dropped it without anyone noticing.
A note holding an empty table took the whole note down. note.to_context() and notes.to_markdown() render through edgar.markdown.process_content when optimising for an LLM, and that raised TypeError: 'NoneType' object is not iterable whenever the disclosure contained a <table> with no usable rows — no <tr> at all, only width-grid layout rows, or cells outside any row. html_to_json() documents its first return value as a list of text blocks but handed back None in three places; it now returns an empty list, which is what its own docstring always promised, so the prose around the table renders instead of the note failing. Output for filings whose tables do have rows is unchanged, verified byte-for-byte across sixteen real Apple, JPMorgan and Coca-Cola note TextBlocks.
A filing's text stopped at the first deeply nested table, losing everything after it. lxml's parser discards anything nested deeper than 256 elements — silently, with no exception and nothing in its error log — and the parser this library builds did not lift that limit. 2000s-era filings nest layout tables that deep: a 2003 S-1 reaching depth 284 returned 137,419 words from filing.text() where 151,924 were present, about 10% of the document, all of it the tail. Sections, markdown and every other text consumer lost the same content. The limit is now lifted, which is what BeautifulSoup always did — html.parser has no depth limit and bs4's own lxml treebuilder lifts it too — so the readers moved to lxml for 6.0 are not lossy on these filings either. Across 282 filing fixtures this recovers text in one and changes nothing in the rest, at no cost in parse time.
FilingHomepage.get_filers() returned an empty list for every filing. It searched the filer block by id="filerDiv", but SEC emits class="filerDiv", so the selector never matched and the method returned before parsing anything — since 2024-05-28. The filer panel was missing from the homepage display for the same reason. Filer names, CIKs, identification lines and mailing/business addresses now come back, and a Form 4 correctly reports both its issuer and its reporting owner. A filer's name also no longer keeps its role suffix on some forms and not others: SEC writes an ordinary filer's role as plain text but a Form 4's as a link, and only the plain spelling was being stripped.
A 10-K item lookup spelled in capitals came back empty. tenk["ITEM 7"] and get_item_with_part("Part II", "ITEM 7") returned None while "Item 7" and "item 7" returned the section, because TenK matched the spelling case-sensitively in two places — deriving the item number, and mapping it to a friendly section name. The legacy parser underneath had been absorbing this, so removing it (below) made the gap visible. TenQ, TwentyF and EightK were already case-insensitive here. (GH #454)
A line break the filer wrote was dropped, gluing the words either side. A <br> between two inline elements — the shape of every 6-K and 8-K cover page — was pruned as an empty node, so UNITED STATES<br/>SECURITIES AND EXCHANGE COMMISSION came back as UNITED STATESSECURITIES AND EXCHANGE COMMISSION. <br> between bare text was never affected. Exelon's 8-K of 2005-04-27 regains two line breaks; its text is otherwise unchanged, at the same 2,629 characters.
A table's sparse label column, and the % or ) a filer put in a cell of its own, no longer vanish. Two rules in the same renderer. A column holding one real value in eight rows — a signature block's Date: June 30, 2025, an exhibit list's Exhibit/No. headers — scored below the spacing threshold and was discarded, so its text was in to_dataframe() and nowhere in text(). Separately, filers split the currency mark, the percent sign and the parenthesis closing a negative number into cells of their own; those scored as spacing too, and were dropped before they could be merged back, so (175,207) rendered as (175,207 and 93.55 % as 93.55. Affix cells now survive the filter and merge into the figure they belong to, in either direction.
Wide tables silently lost real columns. The text renderer scored each column of a table, sorted the scores highest-first, and kept only the top eight. Because of that ordering the discarded columns were not the right-hand edge but whichever scored lowest anywhere in the table, so a 21-column segment table rendered without its Corporate and unallocated and Total headers, and a 10-column voting table dropped the % Withheld figure altogether — 6.45 was in to_dataframe() and nowhere in text(). Every column that clears the content threshold is now rendered; per-column width is still bounded by table_max_col_width.
A word came back split across lines when the filer bolded one of its letters. Filers routinely put an acronym's initials in their own <font> runs; where those sit inside a <div> styled display:inline, Document.text() emitted each run as its own block. Aardvark Therapeutics' 8-K of 2025-04-01 rendered "(Hunger Elimination or Reduction Objective)" one letter to a line. Such a div now reads as a paragraph — unless it wraps a table or another block, which is how iXBRL containers hold whole statements.
edgar._markdown.fix_markdown(), edgar._markdown.html_to_markdown() and MarkdownContent.from_html(). fix_markdown repaired run-together Item headings in markdown and had no caller anywhere — not in the library, not internally, only in its own test. The other two went dead this release: MarkdownContent.from_html() was the sole caller of html_to_markdown(), and its own last caller was PressRelease.to_markdown(), which moved to Attachment.markdown() when press releases came off the legacy parser (above). edgar._markdown is private and appears in no documentation, so none of this is reachable through the public API. MarkdownContent itself, convert_table, markdown_to_rich and text_to_markdown are unaffected — construct MarkdownContent(markdown, title) directly, which is what the remaining callers do. Nothing that resolves today stops resolving.
edgar.datatools.table_tag_to_dataframe(). It converted a BeautifulSoup table tag to a DataFrame and had no callers anywhere in the library, the tests or the docs — it existed only to be migrated. edgar.datatools is not exported from edgar, so this is not reachable through the public API.
edgar.abs.distribution. DistributionReport, DistributionMetrics and ReportTable parsed the HTML distribution exhibit of a Form 10-D, and were never reachable: the module was not exported from edgar.abs, nothing in the library imported it, and it had no tests and no documentation. edgar/abs/__init__.py had recorded it as deferred at roughly 42% extraction accuracy and preserved "for future work" — work that has not happened, so it goes rather than being carried into 6.0. Its validation harness, scripts/validate_distribution_report.py, goes with it. TenD and everything else in edgar.abs are unaffected. Nothing that resolves today stops resolving.
The legacy-parser fallbacks under item lookup on 10-K, 10-Q, 20-F and 8-K. report["Item 7"] and get_item_with_part() used to try the modern parser and then, on a miss, read the filing again with edgar.files. Measured over ~2,110 lookups on filings from 2001 to 2025, those paths were reached 86 times and produced content on none of them, so they are gone. TenK.id_parse_document() and TenQ.id_parse_document() are removed with them. A lookup for an item a filing does not have no longer raises TypeError, which is what TenK did here; every report type now answers None and warns that 6.0 will raise SectionNotFoundError instead. Three report types were still answering None in silence and are fixed above. The deprecated public chunked_document property is unaffected and still available until 6.0.
The legacy parser's contribution to EightK.items. The item list unioned three detectors; the middle one read the filing again with edgar.files. Compared as item SETS — the only comparison that can catch a union quietly dropping a member — across 391 8-K filings from 1995 to 2026, it contributed a unique item to none of them: identical to the surviving detectors on 275, a strict subset on 101, and blind on 82. .items still unions the new parser's section tree with the text-based extractor, so item lists are unchanged, including on the pre-2004 and minimal-HTML filings the text extractor exists for.
No breaking changes. TenK.items , TenQ.items and TwentyF.items no longer consult the legacy parser, which changed the item list on zero of 115 corpus…
Installable: pip install edgartools==5.52.0
The modern parser now answers every item lookup that used to need the deprecated legacy one — across 10-K, 10-Q, 20-F and 8-K, and back to 1996 filings. On a 115-filing corpus, 15 report["Item N"] lookups and 4 part-qualified ones were still being rescued by ChunkedDocument; that number is now zero. This is the gate for removing edgar.files in 6.0.
Ten fixes, each found by measuring the corpus rather than by reading the code:
| what was wrong | effect |
|---|---|
| Item headers wrapping across two lines were unmatchable | a 2010 20-F recovers Items 5, 6, 11, 12 and 15–16F |
| Pre-2002 filings offered no header candidates at all | a 2001 10-K and 20-F recover their Item 7 |
Item 9A(T) hit a pattern with no ( in its vocabulary |
2007–2010 annual reports regain controls-and-procedures |
| Item titles over fifteen words were rejected as headers | Item 5's canonical seventeen-word title now matches |
| Items 4 and 14 have each carried two titles across eras | pre-2011 and pre-2003 annual reports keep both |
| A 10-Q's "PART II" marker styled on an inner element | Goldman Sachs regains its entire Part II |
| An item numbering its own subsections stopped at the first | a 1999 10-K regains its schedules and exhibit index |
| A TOC pointing two items at one anchor | P&G's Items 5 and 6 return their own 306 and 2,320 chars |
get_item_with_part() falling through to an older extractor |
222,536 characters replaced by the correct 975 |
| 10-Q size bands keyed on the bare item number | 38 of 65 corpus size warnings were false alarms |
That last one is worth calling out: a 10-Q has two Item 1s — Financial Statements in Part I at around 90,000 characters, and Legal Proceedings in Part II, often a pointer of a few hundred — and the second was being judged against the first's floor and reported as a truncated extraction. Size bands may now be written per Part.
Both returned a plausible value rather than raising, which is the kind that survives longest:
get_operating_cash_flow() returned None for Apple, taking get_free_cash_flow() with it. It matched five label patterns written against the standardized vocabulary, and Apple writes "Cash generated by operating activities". The filers it worked for worked by coincidence of house style. It now resolves by XBRL concept first, as get_revenue() already did; Apple's latest 10-Q returns $82,627,000,000. (#1083)Location is a bare path, and the redirect hops passed that header straight through, so httpx raised UnsupportedProtocol and the failure was swallowed — find_fund("KINCX").name returned "C000013712" instead of "Advisor Class C". Redirects now resolve relative Locations properly, and the swallow warns unless you are genuinely offline.Also: a section a filer answered by cross-reference is no longer reported as truncated. All five undersized Item 8s in the corpus are pointers, not truncations, so that diagnosis was wrong in every case it fired. (#927)
edgar.settings — the access modes (NORMAL, CAUTION, CRAWL), EdgarSettings, set_identity, get_identity and get_edgar_data_directory now have a home of their own instead of sitting in edgar.core beside quarter math and thread helpers. edgar.core re-exports every name as the same object, so nothing breaks now; those re-exports go away in 6.0.
No breaking changes. TenK.items, TenQ.items and TwentyF.items no longer consult the legacy parser, which changed the item list on zero of 115 corpus filings.
Full detail in CHANGELOG.md.
edgar.settings — a real home for connection settings and SEC identity. The access modes (NORMAL, CAUTION, CRAWL), EdgarSettings, set_identity, get_identity and get_edgar_data_directory now live in edgar.settings rather than in edgar.core beside quarter math, HTML sniffing and thread helpers. Nothing breaks: edgar.core re-exports every one of those names as the same objects, so isinstance and identity comparisons are unaffected, as is from edgar import set_identity, CAUTION. The edgar.core re-exports are removed in 6.0; see docs/upgrade/6.0.md.TenK.items, TenQ.items and TwentyF.items no longer fall back to the deprecated legacy parser, and now report exactly what the modern parser detects. Across a 115-filing corpus spanning 1996 to 2026 this changed the item list on zero filings, for all of 10-K, 10-Q, 20-F and 8-K. Item lookup — report["Item 7"], get_item_with_part() — still has the legacy parser as a fallback but no longer depends on it: every lookup in that corpus is now answered by the modern parser, down from 15 report["Item N"] and 4 part-qualified ones.
A fund's reference-data lookup now says when it failed, unless it failed by being offline. _build_hierarchy_from_mf_tickers falls back to the bare identifier when get_fund_reference_data() is unavailable, but caught every cause of that with except Exception: pass — so a network error, an SEC page restructure and a changed CSV shape were indistinguishable from "this class has no name". It now splits on is_unreachable(): offline degrades quietly, everything else degrades loudly with the exception type and message. The returned value is unchanged either way.
get_latest_bdc_report_year() no longer offers a hardcoded 2024 as though it were a finding. Both BDC probe loops ask whether a year or quarter exists by fetching a URL; a 404 comes back as a response, so reaching the except meant the probe got no answer at all — and each loop then stated a conclusion it had not earned, one returning the fallback year and the other [], which reads as "SEC published none". Both now split on is_unreachable() and say which case they are in, with return values unchanged. Live, the year resolves to 2026. Two reference modules also move to SEC's migrated data-research/sec-markets-data/* addresses, and a new network-marked contract suite watches the datasets themselves, naming the source that moved when one does.
get_operating_cash_flow() returned None for Apple, and took free cash flow down with it. It searched the cash flow statement for five label patterns written against the standardized vocabulary; Apple labels the line "Cash generated by operating activities", which matches none of them, so the method returned None for every Apple filing and get_free_cash_flow() followed it silently. The filers it did work for worked by coincidence of house style. Operating cash flow is now found by XBRL concept through the standardization store first, as get_revenue() and get_capital_expenditures() already were, with the label patterns kept as a fallback for unmapped or custom tags. Apple's latest 10-Q now returns $82,627,000,000, and the five issuers that already worked are unchanged. Reported in #1083.
Every fund series and class name came back as a bare identifier, and nothing said why. SEC turned a dataset page into a 301 whose Location is a bare path, which the standard permits (RFC 7231 §7.1.2), and every manual redirect hop passed that header straight into the next request — so httpx got a URL with no scheme and raised UnsupportedProtocol, which edgar/funds/data.py swallowed. find_fund("KINCX").name returned "C000013712" instead of "Advisor Class C". redirect_url() now resolves the header against the URL that produced it via httpx.URL.join, applied to all five hops, so relative and protocol-relative Locations both work. SEC is mid-migration, so this closes the class rather than the one address.
A 10-Q's Part II items were judged against Part I's expectations and told they were truncated. The size guardrail keyed its expectations on the bare item number, but a 10-Q has two Item 1s — Financial Statements in Part I at around 90,000 characters, and Legal Proceedings in Part II, often a pointer of a few hundred. Part II's was measured against Part I's floor of 18,009 and flagged as likely truncated: 38 of the 65 size warnings the fixture corpus produced were false alarms on correctly-extracted sections. Bands may now be written per Part; the 10-Q's are Part I's Items 1 and 2 and Part II's Exhibits, with the numbers unchanged, since they were measurements of Part I all along. Part II's Items 1 and 2 are left unenforced — they run from 74 to 12,718 characters, and no floor separates "the filer said nothing" from "the anchor missed the body". A warning's character count now matches what section.text() returns.
A section the filer answered with a cross-reference was reported as a truncated extraction. An Item 8 that says the financial statements are filed under Item 15 is a faithful extraction of a pointer, not a truncation, but the undersize warning sent callers to debug a parser that had done its job. All five undersized Item 8s in the fixture corpus are pointers — NVIDIA at 207 characters, Netflix 268, IBM 250, Oracle 158, CIK 915358 at 112 — so the diagnosis was wrong in every case. An undersized section is now tested for a deferral first and a pointer gets an incorporation-by-reference warning naming where the content lives. Confidence is still reduced, so what callers receive is unchanged; only what they are told about it. Reported in #927.
Two items of a 10-Q returned the identical text, with nothing to say which one was wrong. Procter & Gamble's table of contents points its Item 5 and Item 6 rows at the same body anchor, so both sliced to the same 2,628 characters and Item 5's "Other Information" answered with Item 6's exhibit list. The collision resolver re-points such an item at its own body heading, and here it had none to work with: P&G builds each header as a two-cell table row, which renders with nothing between the number and the title, while the body scan required a space. A period may now stand in for that separator, though not when a digit follows it, which keeps an 8-K's "Item 5.02" from reading as a bare Item 5. The two items now return their own 306 and 2,320 characters, and across a 115-filing corpus this is the only filing whose sections change.
get_item_with_part() returned the wrong text on some 10-Qs, from two separate faults. ExxonMobil puts Item 1 in a table and writes its other six as headings, but the strategy that reads item headers out of table cells only ran when fewer than half the form's items had been found — so it was held shut by the very headers that could never supply the missing one, and Part I Item 1 is the entire financial-statement section. Procter & Gamble's fault was different: the code filling in items a TOC omitted compared bare item numbers, so a TOC naming one of the two Item 1s was read as naming both. Both cases fell through to an older extractor that answered with 222,536 characters for a section of 975. Six sections across four filings are recovered, and no section changed its boundaries.
An item whose title runs long was not recognised as a heading at all. Header detection refuses to treat anything over fifteen words as a header, which is reasonable for unlabelled text and wrong for labelled text — several of the SEC's own canonical item titles are longer, Item 5's at seventeen words among them, so the cap rejected precisely the longest real headers. On one 2016 annual report Item 5 was the only item of twenty that went missing. A filing's own "Item N" or "PART N" label now waives the length test; unlabelled long text and prose cross-references are treated as before.
Older 10-Ks lost Items 4 and 14, because each has carried two titles and only the modern one was recognised. Item 4 was "Submission of Matters to a Vote of Security Holders" until Dodd-Frank gave it to mine safety in 2011; exhibits were Item 14 until the 2003 renumbering moved them to Item 15. Both older headers were found and then discarded for having the wrong title. Both titles are now recognised alongside the modern ones, and a modern filing's Items 14 and 15 are unaffected.
An item that numbers its own subsections no longer stops at the first of them. A filer dividing Item 14 with "Item 14(a)(1):", "Item 14 (a)(2):" and "Item 14 (a)(3):" headers was handing the extractor three things that look like item headers, and the item ended at one — on a 1999 annual report that dropped the schedules and the whole exhibit index. An item marker carrying a parenthesized sub-designation and no title is now read as a subdivision. Related: a bare "SIGNATURES" line now ends a section regardless of how the heading detector scored it, so seven sections across 10-K, 10-Q, 20-F and 8-K stop at the signature page instead of running to the end of the document.
Item 9A(T) was unmatchable, so a cohort of 2007–2010 annual reports had no controls-and-procedures section. 9A(T) was the SEC's transitional designation for a smaller reporting company's internal-control report, and ITEM 9A(T). CONTROLS AND PROCEDURES hit an item-header pattern whose punctuation vocabulary had no ( in it — so the match died after the number and the section was never created, though the header itself was found. An item number may now carry a parenthesized designation, across 10-K, 10-Q and 20-F alike.
A 10-Q whose "PART II" marker is styled on an inner element lost its whole Part II. Some filers, Goldman Sachs among them, render the marker as an unstyled paragraph wrapping a bold span, a shape the extractor only recognised for 10-K and 8-K. The damage was larger than one missing marker: each item header is assigned to the last part seen, so with no Part II boundary every later header stayed Part I and the Part II patterns then rejected their own headers — Items 5 and 6 were being found and thrown away. A bare "SIGNATURES" line is now recognised on 10-Q too, which stops Item 6 running past the end of the exhibit list; three further filings gained a correctly bounded Item 6 from that, ExxonMobil's among them.
Item headers that wrap across two lines are now found. A header reading <td>ITEM 5. OPERATING / AND FINANCIAL REVIEW AND PROSPECTS</td> carries a newline in the middle of its title. In HTML that is just whitespace, but the section patterns join words with .*, which does not cross one — so which items a filing lost came down to how each pattern happened to be written: Item 4's uses \s+ and matched, Item 5's uses .* and did not. Header text is now normalized before matching, recovering Items 5, 6, 11, 12 and 15–16F on a 2010 20-F, Items 6 and 11 on a 2016 20-F, and Item 7A on a 1999 10-K. A table-of-contents row can also no longer outrank the real body header for the same item.
Filings from before about 2002 returned no items at all from the modern parser. Those filings are preformatted text in minimal HTML, so they parse to a container-and-text tree with no headings and no paragraphs — and every header strategy drew its candidates from headings, section nodes, bold paragraphs or table cells, leaving no candidate source at all however well the patterns matched. Bare text nodes are now read as headers, using each node's first line, since one node carries both the heading and the body that follows it. A 2001 10-K and a 2001 20-F recover their Item 7; across a 121-filing corpus this was the last remaining difference between the modern and legacy parsers on these forms.
No breaking changes and no API changes — the lxml work is behavior-preserving by construction, and each port shipped with a before/after comparison ov…
Installable: pip install edgartools==5.51.0
Every SEC form-XML parser crossed from BeautifulSoup to lxml — 2.7x to 9.0x faster, with parsed output identical on both sides. The corpus comparisons that migration required also turned up nine bugs that had been returning wrong answers silently, several of them for years.
Nine parsers, each measured against a real filing corpus before and after, with every parsed field compared:
| what | speedup | measured over |
|---|---|---|
| EFFECT notices | 9.0x | 39 documents, four quarters |
| Form D notices | 8.8x | 42 D and D/A filings, 2022–2025 |
Report index (FilingSummary.xml) |
6.9x | 31 filings carrying 1,823 reports |
| Current-filings feed | 6.1x | a real 100-entry page |
| 13F cover pages | 5.7x | 36 filings, 2022–2025 |
| Schedule 13D/G | 4.2x | 125 filings, all of 2025 |
| Form 144 notices | 3.1x | 31 filings, 2022–2025 |
| MA-I municipal advisor | 2.9x | 40 filings, 2021–2026 |
| Forms 3, 4 and 5 | 2.7x | 69 filings, five quarters |
Where this shows depends on what you do. Interactive single-filing work is network-bound, so the parse was never the wait. Volume changes that: get_all_current_filings() pages through the whole feed, Form4.transactions across a portfolio of insider filings parses one document per filing, and local-storage batch jobs are parse-dominated end to end. On Forms 3/4/5 the XML layer itself is now 28x faster to parse and 4.8x faster to read, leaving DataFrame construction as the dominant cost.
Nine silent-wrong-output bugs — the kind that return a plausible answer rather than raising:
False, no matter what the filing said. 40 of the 45 disclosure booleans were read with child_value(), which looks for a <value> child, but MA-I carries the answer as the element's own text — so a filing disclosing a felony charge reported clean, indistinguishable from one that actually is.False because the parser asked for element names the SEC schema does not use — three of them the SEC's own misspellings (isIndependentRelatioship, isTrusteeApointed, isViloatedIndustryStandard), which is exactly why the mismatch was easy to miss.Form144.contact was None on every filing, from two independent mistakes: the parser looked for <contact> under <filer> where the SEC writes it under <filerInfo>, and it read child names the schema does not use.<previousName> spelling — 7 of 42 sampled filings, so a renamed company parsed as never having been renamed, returning an empty list rather than an error.associated_bd_name off a Form D sales-compensation recipient raised AttributeError — a colon where an equals belongs, so the value was parsed correctly, passed correctly, and dropped at the moment of assignment.BeautifulSoup(xml, "xml") parses with recover=True, so malformed filings had been read for years without anyone noticing; xmltools.parse_xml now recovers the same way.TypeError whenever they had no middle name, which municipal advisor filings hit constantly.str(person) printed the first name twice. repr() was always correct, which is why Rich tables looked right while string interpolation did not.SecForms.load() returned an object that raised a bewildering SyntaxError on use, having wrapped a SecForms in another SecForms.No breaking changes and no API changes — the lxml work is behavior-preserving by construction, and each port shipped with a before/after comparison over real filings as its acceptance test. This is groundwork for the 6.0 goal of shipping without a BeautifulSoup dependency (#931); the HTML-side parsers are next.
Full detail in CHANGELOG.md.
Parsing the current-filings feed is 6.1x faster, measured on a real 100-entry page (9.6ms to 1.6ms). Most of that is invisible behind the network round trip when you fetch one page, but get_all_current_filings() pages through the whole feed and pays it every time. Output is byte-identical; the entries, their order, and their fields are unchanged.
Parsing an EFFECT filing is 9.0x faster — 242µs to 27µs per submission, measured over 39 real EFFECT documents from four quarters, with every parsed field identical before and after. EFFECT notices are small, so the win only shows at volume; a day's worth of them is a few thousand filings.
Parsing a Form 3, 4 or 5 is 2.7x faster — 3.35ms to 1.25ms per filing, measured over 69 real ownership filings from five quarters, with every parsed field identical before and after: holdings, transactions, footnotes, signatures, issuer and all 126 reporting owners. The XML layer itself is 28x faster to parse and 4.8x faster to read; what is left is the DataFrame construction, which now dominates. Form4.transactions on a portfolio of insider filings is where this shows.
Reading a filing's report index is 6.9x faster — 16.2ms to 2.4ms per FilingSummary.xml, measured over 31 real filings (10-K, 10-Q, 8-K, 20-F, 2021 to 2025) carrying 1,823 reports between them, with every report, input file and supplemental file identical before and after. This is the parse behind filing.reports and behind the note lookup in TenK.notes.
Parsing a 13F cover page is 5.7x faster — 717µs to 125µs per primary document, measured over 36 real 13F-HR, 13F-HR/A and 13F-NT filings from 2022 to 2025, with every field identical before and after: manager, address, summary totals, other managers and amendment metadata. This is the parse behind ThirteenF.filing_manager, .total_value and .other_managers.
Parsing a Form 144 notice is 3.1x faster — 3.74ms to 1.20ms per notice, measured over 31 real 144 and 144/A filings from 2022 to 2025, with every field identical before and after: filer, issuer, address, both securities tables and the notice signature.
Parsing a Form D notice is 8.8x faster — 2.04ms to 0.23ms per notice, measured over 42 real D and D/A filings from 2022 to 2025, with every field identical before and after: issuer, all 145 related persons and their relationships, the offering sections, sales-compensation recipients and signatures.
Parsing an MA-I municipal advisor filing is 2.9x faster — 3.89ms to 1.33ms per filing, measured over 40 real MA-I and MA-I/A filings from 2021 to 2026, with every field identical before and after: filer, contact, notification addresses, applicant and other names, all advisory offices and their addresses, the full employment history and the signature. MA-I mixes three namespaces in one document, so more of the work stays in the local-name fallback than it does for the single-namespace forms.
Parsing a Schedule 13D or 13G is 4.2x faster — 2.41ms to 0.57ms per document, measured over 125 real SCHEDULE 13D, 13D/A, 13G and 13G/A filings from all four quarters of 2025, with every parsed field identical before and after: issuer and security info, all reporting persons with their voting and dispositive power, the 13D items 1-7, the 13G items 1-10 and the signatures. Structured XML for these forms only exists from the 2024-12-18 SEC mandate onward; older filings still come back from the SGML header with has_structured_data == False.
A Form 4 whose XML the SEC did not write quite correctly came back as raw markup instead of a rendered form. BeautifulSoup(xml, "xml") parses with recover=True, so for years edgartools read filings like AAR CORP's 2004-02-04 Form 4 — which carries a mangled <nonDerivativeTable ativeTable> attribute — without anyone noticing they were malformed. The move to lxml made parsing strict, and filing.sgml().text() on those filings went back to dumping <ownershipDocument> tags. xmltools.parse_xml now recovers exactly as bs4 did, which restores the behaviour for every form migrated so far, not just Forms 3/4/5. A document with no markup at all is still an error.
A person's name raised TypeError whenever they had no middle name. Name.full_name built the middle segment as (' ' + middle_name) or '', so the concatenation ran before the fallback could apply and a missing middle name crashed instead of being skipped. Municipal advisor filings hit this constantly — the parser feeds that field straight from the XML, which simply omits <middleName>. Michael NMN Tym Jr. still reads as Michael NMN Tym Jr., since NMN is data the SEC writes, not an absence.
str(person) printed the first name twice instead of the full name. repr() was always correct, which is why Rich tables and notebook output looked right while string interpolation did not.
Every disclosure answer on an MA-I municipal advisor filing read False, no matter what the filing said. 40 of the 45 disclosure booleans were read with child_value(), which looks for a <value> child element — but MA-I disclosure elements carry the answer as their own text, so the lookup always came back empty and a filing disclosing a felony charge reported clean, indistinguishable from one that actually is. Found by a 40-filing corpus comparison during the lxml migration and present long before it; every disclosure class and Disclosures.any() now reads the answer the SEC actually wrote.
Five more MA-I disclosure fields were hardwired False because the parser asked for element names the SEC schema does not use. Three of the five are the SEC's own misspellings — isIndependentRelatioship, isTrusteeApointed, isViloatedIndustryStandard — which is exactly why the mismatch was easy to miss; the parser searched for the correctly-spelled names and found nothing in any of 40 corpus documents. It now asks for what the SEC writes, with comments guarding each misspelling so nobody corrects them back into brokenness.
Form144.contact was None on every filing. The parser searched for <contact> under <filer> when the SEC writes it as a sibling under <filerInfo>, and even given the right element it read children named name, phone and email where the schema names them contactName, contactPhoneNumber and contactEmailAddress — two mistakes, each sufficient alone. All 31 corpus filings returned None; the SEC's own checked-in sample, whose contact block is populated, now parses.
A Form D issuer's previous names vanished when the SEC wrote them with the <previousName> spelling. The previous-name lists were read only through <value> children, but 7 of 42 sampled filings from 2022–2025 use <previousName> instead, so a renamed company like Shepherd's Finance, LLC parsed as never having been renamed — an empty list, never an error. Both spellings are now read, and the literal None placeholder the SEC writes for "no previous names" is filtered under either one.
Reading associated_bd_name off a Form D sales-compensation recipient raised AttributeError. The constructor line was self.associated_bd_name: associated_bd_name — a colon where an equals belongs, a bare annotation that binds nothing — so the value was parsed correctly, passed correctly, and dropped at the moment of assignment. The sibling CRD line was written correctly, which is what made the typo invisible.
SecForms.load() returned an object that raised a bewildering SyntaxError on use. It wrapped list_forms(), which already returns a SecForms, so the forms table ended up nested one level too deep and every read of it went through the wrong __getitem__ into a pandas query expression. SecForms.load().get_form("1-A") now returns Form 1-A, the Regulation A Offering Statement.
Installable: pip install edgartools==5.50.1
Installable: pip install edgartools==5.50.1
Single-fix patch release: cross-process HTTP caching actually works again.
The HTTP cache was wiped twice on every import edgar, so it never survived a process. The two "one-time" import-time cache clears (#457 locale fix, #672 empty-response fix) each kept their marker file inside _tcache — the directory the other one rmtrees — so each import deleted the other's marker and both clears fired forever, from v5.20.1 onward. Short-lived processes (CLI runs, cron jobs, serverless handlers) re-downloaded everything from SEC each run, and a wipe landing on a concurrent edgar process's in-flight cache write crashed it with FileNotFoundError. Markers now live in <edgar-data-dir>/.migrations, outside the cleared directory, and all migrations run as a single pass that clears at most once; the legacy in-cache marker is honored, so upgrading performs at most one final clear. (#1051)
Pairs with the data.sec.gov cache-rule fix in 5.50.0 (#989): that release made within-process caching work, this one makes it survive across processes. Thanks to @shahar-arbor for an exceptionally well-diagnosed report.
Installable: pip install edgartools==5.50.0
Installable: pip install edgartools==5.50.0
This release is about FilingSGML.text() and .html() holding up across the whole archive, 1993–2026. Five of the six fixes trace back to a full-corpus crawl by a user: PDF primaries no longer crash, XML-primary forms no longer come back as raw markup (one 143MB NPORT-P went from ~1.5 hours to 2.4 seconds), pre-1997 filings no longer leak the era's SGML table dialect, and a truncated submission now says so instead of silently parsing as zero documents.
Company.get_facts() re-downloaded companyfacts on every call, and the 30s /submissions TTL never took effect. Cache rules were keyed off SEC_BASE_URL alone, but httpxthrottlecache matches the request host against that key, and re.match(r'.*www\.sec\.gov', 'data.sec.gov') is None — so a fresh process paid full network cost every time. Keys now come from httpx.URL(...).host, one per host, matched exactly, which also restores caching for custom mirrors. Requires httpxthrottlecache>=0.6.1. (#989)
FilingSGML.html() crashed with UnicodeDecodeError on filings whose primary document is a PDF. The binary guard existed but never fired: is_binary() compared the raw extension against a lowercase list, so a file named .PDF answered False and its bytes went to a bare UTF-8 decode. Classification now goes through the attachment's normalized extension, and html() and xml() route through the decoder that already carried the right contract, including a NUL-byte sniff for what the extension table misses. (#1047)
FilingSGML.text() returned raw XML for XML-primary forms beyond ownership — and could take hours doing it. The HTML sniff asks whether <p>, <div or <span appears anywhere in the string, which inside a 143MB NPORT-P instance is a certainty, so one filing was walked node by node for about an hour and a half; a sniff miss on an X-17A-5 or 24F-2NT returned the markup verbatim. XML documents now get their own branch, keyed on both an XML declaration and a non-<html> root — an iXBRL 10-K opens with <?xml and then <html>, and a 1994 filing opens with <PAGE>, so both halves are load-bearing. The 143MB filing now answers in 2.4 seconds. (#1047)
FilingSGML.text() on pre-1997 filings leaked the era's SGML table dialect — <TABLE>, <CAPTION>, <S>, <C> and footnote tags — into the text. Dialect tags occupying whole lines are dropped whole, so no column moves; inline footnote references are rewritten width-neutrally, <F2> becoming [F2], which keeps every fixed-width table aligned. (#1047)
A truncated submission — a download that failed partway, a cut-off local file — parsed "successfully" as zero documents, and text() returned None. EDGAR always closes <DOCUMENT>, so an unterminated one is structural proof of a cut, and parsing now raises ValueError saying how many complete documents preceded the cut and to re-download or clear the cached copy. The header's PUBLIC DOCUMENT COUNT is also checked against what was parsed, with a warning on a marked deficit — tolerating the off-by-one that complete dissemination files routinely ship (Apple's full 10-K declares 103 and ships 102) and the pre-2004 header-only artifact that legitimately carries no documents. (#1049)
FilingSGML.text() on table-heavy filings peaks ~170MB lower and runs ~13% faster. Profiling a 25MB ABS-15G whose single table holds 66,929 rows and 1.6 million cells attributed the ~60x peak-memory amplification to the render path materialising the grid as millions of unslotted dataclass instances. Cell and MatrixCell now declare __slots__, and the parser only retains its copy of the original HTML when section detection is on. Peak RSS on the profiled filing drops from 1.51GB to 1.34GB with byte-identical output. (#1048)
Company.get_facts() re-downloaded companyfacts on every call, and the 30s /submissions TTL never took effect. Cache rules were keyed off SEC_BASE_URL alone, but httpxthrottlecache matches the request host against that key, and re.match(r'.*www\.sec\.gov', 'data.sec.gov') is None — so a fresh process logs No patterns matched data.sec.gov and pays full network cost every time. Keys now come from httpx.URL(...).host, one per host, matched exactly, which also restores caching for custom mirrors. Requires httpxthrottlecache>=0.6.1. (GH #989)
FilingSGML.html() crashed with UnicodeDecodeError on filings whose primary document is a PDF. The binary guard existed but never fired: is_binary() compared the raw extension against a lowercase list, so a file named .PDF answered False and its bytes went to a bare UTF-8 decode. Three ways to fail open lived in that one path — the case-sensitive compare, a malformed "png" entry that matched nothing, and the binary-extension table defined twice with the same typo in both copies. Classification now goes through the attachment's normalized extension, and html() and xml() route through the decoder that already carried the right contract, including a NUL-byte sniff for what the extension table misses. Ten filings in a 1993–2026 crawl hit this, every one a 40-17G, CERT or 40-24B2/A. (#1047)
FilingSGML.text() returned raw XML for XML-primary forms beyond ownership — and could take hours doing it. The HTML sniff asks whether <p>, <div or <span appears anywhere in the string, which inside a 143MB NPORT-P instance is a certainty, so one filing was walked node by node for about an hour and a half; a sniff miss on an X-17A-5 or 24F-2NT returned the markup verbatim, which is why the same bug looked different on different filings. XML documents now get their own branch, keyed on both an XML declaration and a non-<html> root — both halves load-bearing, since an iXBRL 10-K opens with <?xml and then <html>, and a 1994 filing opens with <PAGE>, which reads as a root element. The 143MB filing now answers in 2.4 seconds. What text() returns for unrenderable XML is unchanged: the document, verbatim. (#1047)
FilingSGML.text() on pre-1997 filings leaked the era's SGML table dialect — <TABLE>, <CAPTION>, <S>, <C> and footnote tags — into the text. Dialect tags occupying whole lines are dropped whole, so no column moves; inline footnote references are rewritten width-neutrally, <F2> becoming [F2], which keeps every fixed-width table aligned. The patterns are deliberately tight, because 1990s filings use a bare < as a less-than sign and blanket angle-bracket deletion would eat real content. (#1047)
A truncated submission — a download that failed partway, a cut-off local file — parsed "successfully" as zero documents, and text() returned None. EDGAR always closes <DOCUMENT>, so an unterminated one is structural proof of a cut, and parsing now raises ValueError saying how many complete documents preceded the cut and to re-download or clear the cached copy. The header's PUBLIC DOCUMENT COUNT is also checked against what was parsed, with a warning on a marked deficit. Both tolerances in that check are measured rather than assumed: complete dissemination files routinely ship one fewer <DOCUMENT> block than they declare — Apple's full 10-K declares 103 and ships 102 — so off-by-one stays silent, and a pre-2004 header-only artifact legitimately carries no documents at all. (#1049)
FilingSGML.text() on table-heavy filings peaks ~170MB lower and runs ~13% faster. Profiling a 25MB ABS-15G whose single table holds 66,929 rows and 1.6 million cells attributed the ~60x peak-memory amplification to the render path materialising the grid as millions of unslotted dataclass instances, each paying for a __dict__ it never uses. Cell and MatrixCell now declare __slots__, and the parser only retains its copy of the original HTML when section detection is on, since only the section extractors read it. Peak RSS on the profiled filing drops from 1.51GB to 1.34GB with byte-identical output. The remaining amplification is the render pipeline rebuilding the same grid several times over, which is tracked for the 6.0 performance pass. (#1048)Across our parity corpus the current parser moved from level with the deprecated ChunkedDocument to ahead of it — 10-K coverage went from +0.1% to +3.…
Installable: pip install edgartools==5.49.0
This release is about section extraction, and if you read 10-K items it is the one to take. Across our parity corpus the current parser moved from level with the deprecated ChunkedDocument to ahead of it — 10-K coverage went from +0.1% to +3.4%, and the items it was missing fell from 46 to 7.
A 10-K that writes Item 1: Business reported one section instead of fifteen.
The 10-K section vocabulary accepted a period between an item number and its title, or nothing at all — Item 1. Business and Item 1 Business — but not the colon, hyphen or em dash that filers also use. That is not one missing pattern. A filer picks a separator and uses it for the whole document, so all 23 item patterns failed together, and pattern matching is the last strategy tried, after the table of contents, the cross-reference index and heading detection have each declined. Real filings came back like this:
tenk.items # ['Item 8'] — the filing has fifteenEverything else was reachable only through the deprecated ChunkedDocument fallback, which 6.0 removes. The separator now lives in one place for 10-K, 10-Q and 20-F and admits . : ; - – —. Across the corpus this recovers 32 items on three filings and loses none.
Item 1C (Cybersecurity) and Item 16 were missing when the table of contents omitted them.
A TOC is something the filer wrote, not a manifest. Items go missing from it routinely — Part III when incorporated by reference from the proxy, Item 16 because it is optional and usually empty, Item 1C because it was new in 2023 and templates lagged. The parser already compensated by augmenting a successful TOC result with items found in the body, but the check deciding whether to do that work asked only whether Part III was complete. Part III is complete on nearly every filing, so the pass was skipped nearly always:
| filing | missing |
|---|---|
| Bank of America, JPMorgan, Tesla | Item 1C |
| American Express, Chevron, Johnson & Johnson | Item 16 |
On every one, the parser had already found the item and discarded it. Item 1C has been mandatory since December 2023, so this affected most modern 10-Ks. Sections are now merged by item rather than by section key, so the same item cannot arrive twice under the two naming conventions the detectors use.
A 20-F could report four sections where it has eighteen. Header detection runs in layers — semantic headings first, then bold paragraphs, table cells and plain paragraphs — and it stopped after the first layer as soon as any header mentioned an item. On one 2010 filing three stray headings, one of them a sentence reading "Please refer to Item 6.E…", suppressed the strategies that find its fifteen real item headers. The gate now asks whether the form's item structure was found rather than whether one item was named.
TwentyF.items read the deprecated parser first, and disagreed with twentyf['Item 5']. Item lookup has read the current parser for some time while the item list read the legacy one, so the two could describe the same filing differently. Items also come back deduplicated and in canonical SEC order — Item 4 before Item 4A before Item 5, Item 16A before Item 19.
report.items warned you about chunked_document, an attribute you never touched. The internal fallbacks read the public deprecated property, so a plain twentyf.items emitted a deprecation warning about a choice that was ours. Under -W error::DeprecationWarning that was not a warning but an exception. Relatedly, TenK, TenQ and CurrentReport had each overridden that property and silently lost its warning entirely — so the users who most needed notice got none.
BDC investment parsing recognizes structured Schedule of Investments labels and normalizes trailing numeric XBRL disambiguators out of investment_type, so grouping by type is stable while identifier still distinguishes tranches. Fourteen more tickers parse, and two label shapes that returned an industry or a share class where the company name belonged now return the issuer. Thanks to @HaCk3Dq. (GH #990)
Full detail in the CHANGELOG.
investment_type. The original label remains available in identifier, so separate tranches remain distinguishable while grouping by investment type is stable. Fourteen more BDC tickers parse their investment types, and two label shapes that previously yielded a company name of 'Specialty finance' or 'Class AA' now return the issuer. (GH #990)Item 1C (Cybersecurity) and Item 16 were missing from filings whose table of contents omitted them. A TOC is something the filer wrote, not a manifest, and items go missing from it routinely — Part III when it is incorporated by reference from the proxy, Item 16 because it is optional and usually empty, Item 1C because it was new in 2023 and templates lagged. The parser already handled this by augmenting a successful TOC result with items found in the body, but the check deciding whether that pass was worth running asked only whether Part III was complete. Part III is complete on nearly every filing, so the pass was skipped nearly always, and the items a TOC actually omits were the ones nobody got: TenK.items on Bank of America, JPMorgan and Tesla listed no Item 1C, and American Express, Chevron and Johnson & Johnson no Item 16, on filings where the parser had already found them and thrown the result away. Item 1C has been mandatory since December 2023, so this was most modern 10-Ks rather than a corner case. The check now asks whether the TOC named every item the form defines. Sections are also merged by item rather than by section key, so the same item cannot arrive twice under the two naming conventions the detectors use (part_ii_item_7 and mda are both Item 7). Across the parity corpus this closes seven 10-K filings and one 10-Q, with no filing losing an item and no measurable change in parse time.
A 10-K that writes Item 1: Business reported one section instead of fifteen. The 10-K section vocabulary accepted a period between an item number and its title, or nothing — Item 1. Business and Item 1 Business — but not the colon, hyphen or em dash that filers also use. That is not one missing pattern: a filer picks a separator and uses it for the whole document, so all 23 item patterns failed together, and the pattern extractor is the last strategy the detector tries, after the table of contents, the cross-reference index and headings have all declined. Two filings in our corpus came back with a single section each — TenK.items was ['Item 8'] where the legacy parser found 15 and 20 items — and everything else was reachable only through the deprecated ChunkedDocument fallback, so the 6.0 deletion would have taken the content with it. The separator now lives in one place for 10-K, 10-Q and 20-F and admits ., :, ;, -, – and —; 10-Q and 20-F already took the dash, and nothing was checking that 10-K agreed. Across the parity corpus this recovers 32 items on three filings and loses none, moving 10-K section coverage from +0.1% to +2.8% against the parser it replaces.
A 20-F could report four sections where it has eighteen. The section extractor finds headers in layers — semantic headings first, then bold paragraphs, table cells and plain paragraphs as fallbacks — and it stopped after the first layer as soon as any header mentioned an item. On filer-agent HTML where heading detection promotes a lot of styled text and few real headings, that test passed on almost nothing: one 2010 20-F was gated by three headings, one of which was a sentence reading "Please refer to Item 6.E…". The fallbacks that find its fifteen actual item headers never ran. The gate now asks whether the form's item structure has been found — a share of the items the form defines — rather than whether one item was mentioned anywhere. Across the parity corpus this closes 14 of the 26 sections a 20-F filing lost, with no filing losing any.
TwentyF.items read the deprecated parser first, and disagreed with twentyf['Item 5']. Item lookup has read the new parser for some time while the item list read the legacy ChunkedDocument, so the two could describe the same filing differently. .items now reads the new parser and falls back to the legacy one only when the new parser finds nothing, matching TenK, TenQ and CurrentReport. Items also come back deduplicated and in canonical SEC order (Item 4 before Item 4A before Item 5, Item 16A before Item 19) on both paths; previously the legacy path returned them in whatever order the document produced. On two filings in the corpus the new parser still finds fewer items than the legacy one, and those are tracked.
report.items warned you about chunked_document, an attribute you never touched. TenK, TenQ, TwentyF and CurrentReport try the new parser first and fall back to the legacy ChunkedDocument, and those fallbacks read the public deprecated property — so a plain twentyf.items emitted chunked_document is deprecated about a choice that was ours, not yours. 20-F got it on every call, because at the time 20-F took the legacy path first. If you run -W error::DeprecationWarning, that was not a warning but an exception. Internal paths now use a private accessor and say nothing; asking for chunked_document yourself still warns.
Three report classes had silently lost that deprecation entirely. TenK, TenQ and CurrentReport each overrode chunked_document to change how it was built, and an override that replaces the property also replaces the warnings.warn inside it — so their users got no notice that the attribute disappears in 6.0, which is the population the deprecation exists for. Construction now happens in _chunked_document, the warning lives in exactly one place, and a test asserts no subclass can take it away again.
Nothing raises differently than it did in 5.47.0 — the new behaviour is opt-in, and every deprecated name still resolves to the same object. What you…
Installable: pip install edgartools==5.48.0
This release stages the error-handling changes that 6.0 makes permanent. Nothing raises differently than it did in 5.47.0 — the new behaviour is opt-in, and every deprecated name still resolves to the same object. What you get now is warning of what changes, and a switch to try it early.
Four calls that answer None for a failure now say what they will raise. Each emits a FutureWarning naming the exception:
| Call | Today | 6.0 |
|---|---|---|
tenk["Item 99"] (no such item) |
warns, returns None |
raises SectionNotFoundError |
find("123456-99") (malformed accession) |
warns, returns None |
raises ValidationError |
filing.obj() on a modelled form with unreadable data |
warns, returns None |
raises DataObjectError |
tenk.document when the parser fails |
warns, returns None |
raises the parser's ParsingError |
The legitimate Nones are untouched. filing.obj() on a form edgartools does not model, and filing.xbrl() on a filing without XBRL, are answers about the world rather than failures — and that distinction is the entire point.
If an absent item is a normal outcome for your code, there is now a non-raising form beside the raising one. It never warns, today or in 6.0:
tenk["Item 7"] # the item you expect; raises in 6.0 if absent
tenk.get("Item 16", "") # the item that may not be thereEach of these warns once per line of your code, not once per filing — a loop over ten thousand filings gets one warning, not ten thousand.
Seven exception spellings are deprecated and removed in 6.0: StatementNotFound, NoCompanyFactsFound, SECFilingNotFoundError, InvalidDateException, IdentityNotSetException, TooManyRequestsException, DataObjectException. Each still resolves to the same object, so except/isinstance keep working.
edgar.exceptions — one exception vocabulary, four branches. There were 27 exception classes across ten packages with no shared base, of which exactly two were reachable from the top level, so except had to name a type from whichever module happened to raise. There is now a root, EdgarError, and four branches that answer the question a caller actually has:
EdgarError
├── TransportError we could not get an answer from SEC
├── NotFoundError you named a thing and it does not exist
├── ParsingError we got bytes and could not build the object
└── ValidationError your input was wrong before we asked
The first two being distinct is what matters most — an outage and an empty result must never arrive as the same value. The branches also inherit the builtin they replace (ValidationError is a ValueError, NotFoundError is a LookupError), so the except ValueError: you already wrote keeps working.
EDGARTOOLS_STRICT_ERRORS=1 runs 6.0's error behaviour today, so you can port before the break rather than after:
EDGARTOOLS_STRICT_ERRORS=1 pytestIt turns on both halves: the four conversions above, and wrapping any httpx failure that survives every retry into TransportError — which stops a dependency's exception types from being part of our public contract by accident.
edgartools ships a PEP 561 py.typed marker, so its type hints now reach your type checker. The README has said "type hints throughout" for a long time; it was true of the source and false of the installed package, because without the marker mypy refuses to look inside edgar at all and every symbol degrades to Any. Company(cik_or_ticker=[1, 2, 3]) type-checked clean against 5.47.0 and now reports the error. Pyright users saw types already; mypy and stub-strict configurations did not.
A missing-attachment lookup raises AttachmentNotFoundError rather than a bare KeyError. It is a KeyError, so existing handlers are unaffected.
resolve_accession() wrapped its fetch in except Exception: return None, and that None routes the caller to the quarterly index — correct for a pre-2001 accession, a wrong turn during an outage, with the only trace at DEBUG.NoCompanyFactsFound carried a message nobody could read — str(exc) was the empty string, so three raise sites had messages that never reached a traceback, a log line, or a user.ConnectionError. Nothing in the HTTP layer raises it, so that handler has never fired. It is TransportError.Full guide: Error handling · Upgrade path: docs/upgrade/6.0.md
edgar.exceptions — one exception vocabulary, four branches. There were 27 exception classes across ten packages with no shared base and no cross-package inheritance, of which exactly two were reachable from the top level, so except had to name a type from whichever module happened to raise. There is now a root, EdgarError, and four branches that answer the question a caller actually has: TransportError (we could not get an answer from SEC), NotFoundError (you named a thing and it does not exist), ParsingError (we got bytes and could not build the object), ValidationError (your input was wrong before we asked). The distinction between the first two is the one that matters most — an outage and an empty result must never arrive as the same value. Nothing changed about what is raised today: every existing class was re-based into the tree or kept as a deprecated alias for the same object, so except StatementNotFound: and pytest.raises(SECFilingNotFoundError) still work. The branches also inherit the builtin they replace — ValidationError is a ValueError, NotFoundError is a LookupError — so the except ValueError: you wrote against our 135 raw ValueError raises keeps working as those convert. Deprecated spellings warn and are removed in 6.0: StatementNotFound, NoCompanyFactsFound, SECFilingNotFoundError, InvalidDateException, IdentityNotSetException, TooManyRequestsException, DataObjectException.
A missing-attachment lookup raises AttachmentNotFoundError rather than a bare KeyError. It is a KeyError, so existing handlers are unaffected.
EDGARTOOLS_STRICT_ERRORS=1 runs 6.0's error behaviour today. The changes that would otherwise be a break are available behind the flag, so you can port before 6.0 lands rather than after. It turns on both halves of that change: the network wrap below, and the four silent-None conversions under Deprecated. Check it yourself with edgar.exceptions.strict_errors_enabled().
report.get(item, default) on every company report. The counterpart to report[item], and the reason that one can start raising in 6.0: a lookup whose only form raises leaves the "I'll take it if it's there" caller wrapping a one-liner in try/except. It never warns — it is the migration target, and warning there would give the users who took our advice the same noise as the users who ignored it.
Under strict, an httpx failure that survived every retry becomes a TransportError. httpx.ReadTimeout and friends were propagating verbatim out of get_with_retry, stream_with_retry, post_with_retry and inspect_response, which made a dependency's exception types part of our public contract by accident — and made any future HTTP-client change a breaking one for every except clause naming them. The wrap happens once at the boundary, carries .status_code (None when we never got an answer at all) and .url, and always chains the original as __cause__, so nothing is lost for debugging. This is a 6.0 flip because user code may be catching httpx.HTTPError around our calls; without the flag, nothing changes. TRANSPORT_ERRORS now catches both eras, so code written against it needs no revisiting.
edgartools ships a PEP 561 py.typed marker, so its type hints now reach your type checker. The README has said "type hints throughout" for a long time and it was true of the source and false of the installed package: without the marker, mypy refuses to look inside edgar at all — Skipping analyzing "edgar": module is installed, but missing library stubs or py.typed marker — and every symbol degrades to Any. Company(cik_or_ticker=[1, 2, 3]) type-checked clean against 5.47.0; it now reports Argument "cik_or_ticker" to "Company" has incompatible type "list[int]"; expected "str | int". Nothing in the library changed — this makes the annotations already there visible, and it is why the typing work behind them was worth doing. Pyright users saw types already, because it reads library source by default; mypy and stub-strict configurations did not.
None for a failure now say what they will raise in 6.0. Each emits a FutureWarning naming the exception, and raises it today under EDGARTOOLS_STRICT_ERRORS=1: tenk["Item 99"] becomes SectionNotFoundError, find("123456-99") on a malformed accession becomes ValidationError, filing.obj() on a form we model whose data will not read becomes DataObjectError, and TenK.document on a parse failure raises the parser's own ParsingError. The legitimate Nones are untouched: filing.obj() on a form edgartools does not model, and filing.xbrl() on a filing without XBRL, are answers about the world rather than failures — and that distinction is the entire point. TenK.document had collapsed "this filing has no HTML" and "this filing has HTML we could not read" into the same value; every sibling report class already let the parse failure through. Each of these warns once per call site, not once per filing — the per-filing detail (which accession, which items it does have) stays on the exception that strict mode raises, so a loop over a corpus of ten thousand filings gets one warning rather than ten thousand. See the new error-handling guide.docs/guides/current-filings.md told you to catch the builtin ConnectionError. Nothing in the HTTP layer raises it, so that handler has never fired. It is TransportError.
NoCompanyFactsFound carried a message nobody could read. Its __init__ called super().__init__() with no arguments and set self.message instead, so str(exc) was the empty string — three raise sites whose message never reached a traceback, a log line, or a user. It is now CompanyFactsNotFoundError and builds its message through the base class, which makes the empty case unrepresentable rather than merely fixed.
An EFTS outage was reported as "there is no filing at that accession". resolve_accession() wrapped its fetch in except Exception: return None, and None from that function routes the caller to the quarterly index — the correct answer for a pre-2001 accession and a wrong turn during an SEC outage, with the only trace at DEBUG. Transport failures now propagate; a malformed response still returns None, which is what that arm was for. This closes the sibling left open by the edgartools-tg7y fix, which fixed the same shape in edgar/funds/core.py.
company . cash_flow () # deprecated → removed in 6.0 financials . cashflow_statement () # deprecated → removed in 6.0 financials . cash_flow_statement…
Installable: pip install edgartools==5.47.0
Cash flow is cash_flow_statement() on every object. It had three spellings, and which one worked depended on which object you were holding — of the 16 classes exposing income_statement(), EntityFacts and CurrentPeriodView offered no cash_flow_statement() at all. It now matches income_statement() and balance_sheet() everywhere.
company.cash_flow() # deprecated → removed in 6.0
financials.cashflow_statement() # deprecated → removed in 6.0
financials.cash_flow_statement() # supportedBoth old spellings emit a DeprecationWarning. Upgrade path: docs/upgrade/6.0.md.
edgar now declares __all__ — 110 names. There was no way to tell an API from an accident: 141 names were reachable from the top level, including Optional and partial. from edgar import * is narrower by 31 names; direct imports are unaffected. Names left out stay importable in 5.x and become private in 6.0.
MICHAE|L P|., and the name heuristic matched that as a limited partnership. Six reported filers are affected (#1019). L L C had the same defect, and LLP, LLLP, PLLC plus non-US legal forms (AG, N.V., GMBH, A/S, S.P.A., …) were never recognised at all. Companies undetected by name in the ticker file: 313 → 163.Filing.markdown() dropped every image — it was the last public rendering method on the legacy pipeline, which has no image node. NVIDIA's FY2026 10-K now renders both, including the Item 5 stock performance graph that exists only as a chart.Filing.text() truncated long table cells at 200 characters with no ellipsis to show it. Rendering is also 4.5× faster.Document.to_markdown() dropped the exhibit index from 10-K filings.Filing.text(include_images=True) emits an [Image: …] placeholder per image. Off by default, so text feeding search and embeddings stays image-free.edgar.xbrl.analysis — unreachable and non-functional since it arrived; FinancialMetrics raised AttributeError on construction and nothing imported it.Full detail in CHANGELOG.md.
edgar now declares __all__ — 110 names — so the supported API is answerable. There was no way to tell an API from an accident: 141 names were reachable from the top level, including Optional and partial (imported for annotations) and Document, the legacy edgar.files parser, which is a different class from edgar.documents.Document and is removed in 6.0. Names left out stay importable; 6.0 makes them private. from edgar import * is narrower — it no longer yields those 31 names. Direct imports are unaffected.
Filing.text(include_images=True) and Document.text(include_images=True) emit an [Image: <alt or filename>] placeholder per image. This completes the text half of GH #886: a TextExtractor flag existed but never fired on a real filing, because SEC filers wrap <img> in a paragraph and the paragraph branch returned without descending to it. NVIDIA's 10-K now marks both its images, including the Item 5 stock performance graph whose five-year return comparison has no table beside it. Off by default, so the text feeding sections, search and embeddings stays image-free. (GH #886)
Cash flow is cash_flow_statement() on every object. It had three spellings — Company.cash_flow(), Financials.cashflow_statement(), xbrl.statements.cash_flow_statement() — and only Company accepted all three: of the 16 classes exposing income_statement(), EntityFacts and CurrentPeriodView had no cash_flow_statement() at all. It now matches income_statement() and balance_sheet() everywhere. The old spellings emit a DeprecationWarning and are removed in 6.0; see docs/upgrade/6.0.md. Three internal callers were still on the old name, which turned that warning into noise the caller could not act on — EntityFacts.cash_flow_statement(period='ttm'), spelled canonically, told the user to stop using cashflow_statement(), and every MCP filing read did the same. FilingViewer.compare_context() also disagreed with itself: it keyed its viewer-report table on the old spelling while looking the method up on the object, so the canonical name silently returned "(No viewer report found)". Both spellings now take the canonical path.
Filing.markdown() and Attachment.markdown() now render through edgar.documents. They were the last public rendering methods still on the legacy edgar.files pipeline, which has no image node at all — MarkdownRenderer.render handles text blocks, tables, headings and page breaks, and drops everything else. Every <img> in every filing vanished silently. NVIDIA's FY2026 10-K now renders both of its images, including the Item 5 stock performance graph whose five-year return comparison exists only as a chart. Relative src values are resolved against the filing's SEC archive directory, so the markdown carries working absolute links rather than bare sibling file names. Output also differs where the two renderers disagree on tables; a parity ratchet pins the numeric content of both. (GH #886)
Filing.text() truncated long table cells at 200 characters, with no ellipsis to show it had happened. It rendered its document with rich_to_text(document, width=500), but Document is not a rich renderable, so rich fell back to repr() — and Document.__repr__ is hardcoded text(table_max_col_width=200). The 500 never reached the table renderer. On Apple's FY2024 10-K that cost 1,434 characters across 8 chunks (whitespace collapsed, so wrapping does not account for it), including 300 characters of the unrecognized-tax-benefits disclosure: $22.0 billion, of which $10.8 billion, if recognized, would impact the Company's effective tax rate. Seven of eight 10-K fixtures lost content with no ellipsis marker. Filing.text() now calls document.text(table_max_col_width=500) directly, which also drops the rich round-trip — a 400 KB string was being rendered through a console and stripped of ANSI to recover text the extractor had already produced. Rendering is 4.5× faster (6.61s → 1.46s across 8 filings).
Image alt text leaked into Filing.text() as unlabelled prose. ImageNode.text() returned alt, which ParagraphNode.text() aggregates, so the bare string nvidialogoa10.jpg appeared inline in NVIDIA's 10-K text with nothing marking it as an image. alt describes an image rather than being text the filer wrote, and on SEC filings it is usually just the source file name; it no longer contributes to text. Callers that want images represented ask explicitly — see include_images below.
include_page_breaks and start_page_number on Filing.markdown() and Attachment.markdown(). Page-break rendering exists only in the legacy renderer, so passing include_page_breaks=True routes the whole document through it and forfeits images and the newer table rendering — the flag selects a renderer, not a feature. It still works and now emits a DeprecationWarning; both parameters go in 6.0 alongside edgar.files. There is no replacement: the edgar.documents builder treats page-break <hr>s and page-number containers as print chrome and discards them.edgar.xbrl.analysis — 2,098 lines of Altman Z, Beneish M, Piotroski F, Montier C and a ratios engine, none of it reachable and none of it working. FinancialMetrics.__init__ raised AttributeError: 'function' object has no attribute 'to_dataframe' on the first statement it touched: it read statements.balance_sheet without calling it, so the truthiness check always passed and .to_dataframe() ran against a bound method. All three statement loads had it, and statements.cash_flow was not a name that existed at all. Nothing imported the package — no __init__.py, fraud.py was the only consumer of metrics.py and nothing consumed fraud.py, ratios.py stood alone, and there were no tests, no docs and no entry in edgar.__all__. It arrived on 2025-04-12 in the XBRL2→xbrl rename and only ruff sweeps have touched it since. Removing a subtree that could not be constructed is not a behaviour change, which is why it does not wait for 6.0. If these metrics are wanted they should be built against the current Statement/Facts API rather than revived. (edgartools-07lk.12.1)LLP, LLLP and PLLC were missing from the entity name heuristic entirely. The strict keyword set carried LP, LLC and LTD but not the partnership forms beside them, so 340 LLP, 759 LLLP and 26 PLLC filers had no name signal. The 1,300-odd ALL PRO … LLLP real-estate partnerships were only ever detected by the accidental L P inside AL|L P|RO — removing that accident is what surfaced them.
Non-US legal forms went unrecognised by the entity name heuristic. SIEMENS AG, AIRBUS SE, ABN AMRO BANK N.V., ASTRAZENECA AB, COLOPLAST A/S, BREMBO S.P.A. and ADS-TEC HOLDING GMBH all read as no-signal names, because the keyword sets carried a US-centric handful and the strict path splits on \W+ — so it sees {S, A} where the name says S.A., and {N, V} where it says N.V.. Legal forms are now matched as the final token of the name, with punctuation stripped, so N.V., NV, A/S and S.P.A. resolve to one entry each: 313 of the 7,990 companies in the ticker file were undetected by name, now 163, and across SEC's full 1,054,270-name filer list 2,510 more filers are identified. Terminal-only is what makes the two-letter forms safe — anywhere else AS is English and SE is Spanish. ASA, KK, AD and PT were measured and deliberately left out: "Åsa" is a Nordic given name, so HEDIN ASA and ASK ASA are people in SEC's LAST FIRST order, and ROGERS HUGH A.D., GROOM BENJAMIN P.T. and NGAI ANTHONY K.K. are trailing initials. Five individuals in the whole filer list still read as companies through this path — all of the shape DAVIES JOHN A.B. — against 2,510 correct identifications.
Six individuals were classified as companies because their names contain "l p". The name heuristic carried the spaced legal suffix L P as a loose substring, and a substring does not care where a word ends: "Michael P." is MICHAE|L P|., and so are "Daniel Paul", "Jill P. Meyer", "O'NEIL PATRICK" and "MICHAEL PHILIP". All six reported filers came back with is_individual == False. L L C had the identical defect one keyword over and went unreported — "MICHAEL L COOPER" is MICHAE|L L C|OOPER — so both now match on whole words. The keyword set's own comment claimed its members were long or punctuated enough to be safe as substrings; these two never were. Removing the accidental L P match also removed the only name signal three real companies had, all of them S.A. — the strict path splits on \W+ and so sees {S, A} where the name says S.A. — so the punctuated form joins L.P. and L.L.C. in the set. Reported by @mwtarnowski. (GH #1019)
A table data row rendered only the first line of each cell, discarding the rest. Header rows were expanded line by line; data rows went through a formatter that did content.split('\n')[0]. Both text renderers had the same split — the Rich path paired overflow="fold" for headed columns with overflow="ellipsis" for headerless ones. Filers from the 1990s and 2000s wrap an entire document in a single-cell layout table, so the cut discarded the filing: Autoliv's 2001 DEFR14A rendered 20,523 characters instead of 83,316, losing its "DEAR STOCKHOLDER" cover letter from an 18,759-character cell. Modern filings lost less but still lost — Apple's FY2024 10-K gains 843 characters of non-whitespace content, AbbVie's 2,828. Exposed rather than caused by the exhibit-index fix, which reclassified three of Autoliv's layout tables from header rows to data rows and so moved them onto the truncating path. (edgartools-j8bs)
Table lines carried trailing padding to the last column's full width. Harmless when a tall cell paid it once; now that multi-line cells expand to one line each it is paid per line. Trimmed, which makes rendered text substantially smaller with identical content — Autoliv is 53,278 characters against 83,316 before the truncation bug existed, carrying 39,420 non-whitespace characters against 39,307.
A 424B cover-role label matched mid-sentence and named a firm from the next line. Placement Agent was matched anywhere in the cover slice, so Calidi's offering-table footnote — "...the exercise of any of the Common Warrants, or any of the Placement Agent / Warrants." — yielded a placement agent named "Warrants." The labels are cover-grid cells and must now start a line. The footnote was always in the filing; it only became visible once table cells stopped being truncated to their first line.
FilingSGML.text() truncated long table cells at 200 characters, exactly as Filing.text() did. Both rendered through rich_to_text(document, width=500), which goes via Document.__repr__ and its hardcoded table_max_col_width=200, so the 500 never reached the table renderer. Filing.text() was fixed first and this call site was missed, which left the two paths disagreeing — they are asserted equal by test_filing_text_baseline.test_both_paths_agree. Both now call document.text(table_max_col_width=500). On Apple's FY2023 10-K the SGML path recovers 6,394 characters, including a tax disclosure that had been cut mid-sentence at "gross unrecognized tax benefits was $19.5 billion, of which $9.5 billion, if recognized, would impact Apple"; Bank of America's 424B2 recovers 9,135, including its automatic-call terms.
A debt schedule laid out two-up lost almost all of its rows. UnitedHealth's FY2024 10-K prints two debt series side by side — $750 3.5%, Feb 2024 | 750 | $850 5.8%, Mar 2036 | 838 — so every data row names two different maturity years, and _is_header_row's multi-year branch read that as a 2024-vs-2036 comparison header. 35 of the table's 40 rows were classified as headers, leaving 3 data rows: the rendered filing kept 2 of 66 maturities and 16 of 66 coupon rates while the surrounding narrative stayed intact, so the output read as complete. Neither existing guard applied — the date-range guard needs a full March 1, 2024—March 31, 2024 span, and _has_prose_cell needs a cell of 100+ characters where these are short numeric cells. A row carrying currency amounts or thousands-grouped figures is now treated as data however many years it names; bare decimals deliberately do not count, since those appear in legitimate header labels where a dollar amount never does. All 66 maturities and coupon rates now render.
A row of decimals, and a row of ranges, were both read as table headers. Two more branches of _is_header_row were missing the kind of figures veto the two-up debt schedule needed, and each one collapsed a table that reads as complete without it. UnitedHealth's FY2024 10-K lost two of the three months in its Item 5 buyback table — November 30, 2024 | 0.9 | 593.39 | 0.9 | 38.7 — because a single date matches the period-header pattern and that branch's list of data indicators recognised $, thousands separators and parenthesised negatives but not a plain decimal; it was the only one of three near-identical patterns in the function to omit it, and October survived only because its price cell carried a stray $. The same filing lost three of the five rows of its stock-option assumptions table, including Expected volatility | 25.5% - 30.7% | 29.7% - 30.6% | 30.6% - 30.8%, because a cell holding a range was counted as text: the numeric test stripped $%,() and dropped . and -, leaving the spaces around the dash, and '255 307'.isdigit() is False. Forfeiture rate 5.0% survived precisely because a single value is not a range, so the rendered table kept the rows nobody asks about. The effect is far wider than the filing it was found on: across the parity corpus the numeric content lost against the legacy renderer falls 18%, with Microsoft's 10-Q going from 68 lost values to 8, its 10-K from 57 to 15, Amazon's from 34 to 13 and Meta's from 53 to 39, and Apple's 10-K now losing nothing at all. A third branch, year_cells >= 2, had the same gap and had never been seen to fire; a schedule laid out 2024 | $1,000 | 2025 | $2,000 fires it, so it now carries the veto too. (edgartools-y264)
Document.to_markdown() dropped the exhibit index from 10-K filings. Every exhibit row was classified as a header row — _is_header_row reads any date as evidence of a period column, and exhibit descriptions are full of them ("dated as of June 25, 2019") — leaving no data rows for the renderer to emit. On AbbVie's FY2024 10-K the index is back: 62 "incorporated by reference" rows, and the rendered document grows from 436,060 to 452,668 characters.
Table cell text gained a space wherever inline markup split a value. _extract_text inserted one between any two adjacent text fragments, so AbbVie's FY2024 10-K, which writes each exhibit number across two <a> tags, rendered 10.10 as 10.1 0 — present but unfindable. Separation now derives from block boundaries only, so <div>Exhibit</div><div>Number</div> still reads as two words.
On a multi-filer filing, `Filing.cik` and `Filing.company` now name the issuer rather than whichever filer the quarterly index listed first. Accession
On a multi-filer filing, Filing.cik and Filing.company now name the issuer rather than whichever filer the quarterly index listed first. Accession lookups go through EDGAR full-text search before falling back to the quarterly index, and the two order a filing's filers differently. find("0001918704-25-005439") was (70858, 'BANK OF AMERICA CORP /DE/') and is now (1682472, 'BofA Finance LLC'). all_ciks and all_entities still return every filer; single-filer filings are unaffected.
edgar.xbrl.facts.FactQuery.to_dataframe() returned a different column set depending on which rows matched. Columns now follow the query's configuration rather than its results. On Foot Locker's FY2024 10-K, .limit(5) returned five fewer columns than the same query unlimited, dropping balance, currency, decimals, unit_ref and weight because those rows were null. Unpopulated columns come back null, and an empty result carries the full column set. (GH #929)
httpxthrottlecache is no longer capped below 0.5.0. The cap existed because 0.5.0 briefly required the httpx2 fork; 0.6.0 makes httpx and httpx2 optional extras, so the dependency is now httpxthrottlecache[httpx]>=0.6.0 and edgartools stays on plain httpx. The extra is required rather than cosmetic — from 0.6.0 the package installs neither transport by default.
Removed the HttpxThrottleCache._get_httpx_transport_params monkeypatch. It was added when the upstream method dropped verify, leaving users behind SSL-inspecting proxies unable to disable verification. Upstream extracts verify itself now, so the patch was redundant — and a liability, since shadowing an upstream method silently reverts any later fix to it. A test pins the behaviour it protected.
EntityFactsParser._parse_date no longer swallows every exception type. It caught bare Exception and returned None; it now catches ValueError and TypeError, the two a date parse can legitimately raise. Malformed and non-string input still yields None, but an unanticipated failure surfaces instead of being silently absorbed once per fact. (GH #981, thanks @joseturegano)
The bundled CUSIP→ticker mapping carried 1,843 symbols with a literal XXXX appended, and the 13F parsers rendered them rather than dropping them. get_ticker_from_cusip("G3421J106") returned 'FERGXXXX' instead of 'FERG', so a Ferguson position in a 13F info table displayed a symbol that resolves to nothing — wrong data on the page, not a blank cell. The upstream dataset has been regenerated with all 1,843 recovered, and the merge that builds the bundled file now drops anything not shaped like a symbol. That also clears 534 corporate-action artifacts inherited from older vintages (Q999SPNOFF, 3977PAYRTS, **********, bare CUSIP fragments), none of which ever resolved. Coverage rises from 68,512 to 68,830 CUSIPs. (GH #978, thanks @MRileyLeBay)
Sections ended at the next item number rather than the next section in document order, so filings that group their items out of numeric order lost sections and mislabelled others. Morgan Stanley's FY2024 10-K (0000895421-25-000200) files Items 1B-5 behind the financial statements: Item 1A ran 673,015 characters, swallowing MD&A, the statements and the controls items at confidence 0.95, and Items 1B and 5 were absent. Item 1A is now 73,167 characters and both missing items are present.
Sections detected by heading patterns ran past their own end into the next item. A section ends at the next item's header, but filers wrap that header in a container that begins earlier, so the container was attached whole with everything it held. On Wells Fargo's FY2024 10-K (0000072971-25-000094) 20 of 23 sections carried a foreign item heading, and Item 8 came back as 3,329 characters spanning Items 9 through 9C; it is now 374.
Wells Fargo's 10-K published its exhibit list as the financial statements. The filing writes each item heading as a standalone one-row table, which the parser classifies as a header row, leaving .rows empty — and the table strategy scanned only .rows. Matching fell through to bare-title keywords, so the FINANCIAL STATEMENTS heading inside Item 15's exhibit list claimed the financial_statements key with 42,462 characters of wrong content. Header rows are now scanned too; the filing yields all 23 items.
Citigroup's 10-K items were unreachable by their canonical keys. Its FY2024 10-K (0000831001-25-000029) maps items to printed page ranges in a cross-reference index rather than labelling them in the body, and the parse stopped at the first </table> — recovering 14 of 22 items and silently losing all of Part III. The index now continues across adjacent tables and is used to build sections, so part_ii_item_7 opens on MD&A with 473,576 characters against 234,483 under the old mda key.
10-K items came back as raw HTML on filings that use a Cross Reference Index. TenK.__getitem__ returned CrossReferenceIndex.extract_item_content() straight through, and that method yields HTML by contract while every other branch of the same lookup returns text. On Citigroup's FY2024 10-K (0000831001-25-000029) obj['Item 1'] gave 1,685,461 characters of markup; it now gives 218,550 characters of text. A caller parsing the returned markup will need to stop. (GH #821)
Numbered index rows fabricated 10-K items. Freddie Mac's FY2025 10-K mapped every part_*_item_N onto an MD&A table caption at full confidence — obj['Item 11'] returned 30K characters opening on "Table 11 - Other Investments Portfolio" — because the bare row numbers of the MD&A's "List of Tables" fed the generic TOC scan as item numbers. Bare numbers are now ignored when the containing table's header names the numbering without an "Item" column. (GH #918)
A two-column table of contents read one column's page number as the other column's item number, inventing a section that truncated MD&A. Both columns share one HTML row, and the scan for an item label walked back across the gap to take the left column's page number. On Ambac's FY2022 10-K (0000874501-23-000040) that produced a phantom part_ii_item_10 anchored inside MD&A, cutting Item 7 from 158,411 to 149,459 characters. The scan now stops at the column boundary.
Document.to_dataframe() raised on every real annual and quarterly report. It failed on 10 of the 11 filings in the 6.0 performance corpus with three different pandas and numpy errors; the only one that worked was a single-table ABS-15G. A filing's tables do not share a schema — Meta's FY2024 10-K has 71 tables whose column indexes run 1 to 17 levels deep — which pandas cannot align. Tables are now flattened to single-level string columns before stacking.
A table whose first header text repeats built its row index out of every matching column, giving None-padded tuples instead of labels. TableNode.to_dataframe() moved the label column into the index by label, and filings repeat a header across spacer columns routinely, so a row label came back as ('Gross written premiums by line of business:', None, None). A two-dimensional match raised Index data must be 1-dimensional, breaking Tesla's FY2023 10-K. The first column is now taken by position.
Comparing a document node with a deep copy of itself raised RecursionError instead of returning an answer. Node was a plain dataclass, so the generated __eq__ recursed through parent and children. The per-instance id uuid normally decides the comparison first, but copy.deepcopy preserves it, so a == copy.deepcopy(a) walked back on itself. Nodes now compare by identity; nothing observable changes for existing callers, and nodes are hashable again.
A beneficial-ownership table was read as the underwriting syndicate, so lead_manager returned the column header 'Before Offering' and a roster of directors. Both tables have a name column beside a share-count column, so an ownership table satisfied the structural test. On Learn CW Investment Corp's S-1 (0001140361-21-010426) five directors were listed ahead of the one real underwriter; it now returns Evercore Group L.L.C. alone. Tables whose header region names beneficial ownership are skipped.
An underwriter filing as a division of its broker-dealer was rejected as parser junk, so lead_manager returned None on every filing it led. The name guard's allowlist of permitted lowercase tokens had no entry for division or the article a, so EF Hutton, division of Benchmark Investments, LLC failed on that one token. On Unicycive Therapeutics' S-1/A (0001213900-21-035033) the sole underwriter was the rejected name and the filing reported none; it now returns EF Hutton.
Section extraction is 4.4x faster on the benchmark corpus — 20.8s to 4.7s across ten filings, with every extracted section unchanged. Anchor resolution ran a full-document XPath per lookup: Morgan Stanley's 9.8MB 10-K made 92 lookups against one tree, 3,854ms and 60% of the stage, two-thirds of them re-resolving an id already looked up. One indexing pass now serves every lookup, and content collection starts at the anchor rather than the document root.
A filing was md5'd once per section to rebuild a cache key that could not have changed. Navigation-link filtering resolves its patterns through a cache keyed by an md5 of the entire filing. Extracting all 22 sections of Morgan Stanley's 9.8MB 10-K hashed it 44 times — 430.7MB hashed to answer a question about 9.8MB — and the answer was the same every time. Patterns are resolved once per document now: the sections stage drops 31%, 4.3s to 3.0s.
Company.get_facts() spent about 15% of its wall-clock re-deriving dates that were already ISO-8601. EntityFactsParser._parse_date tried datetime.strptime first, so the ISO fast path already sitting in the function was never reached — strptime takes a locale lock per call, ~3.6us against ~0.10us. WMT, HD, MCD, NKE, LOW and TGT make 364,523 date calls for 139,284 facts. ISO is tried first now. (GH #981, thanks @joseturegano)
filing.obj() on a registration statement downloaded its whole file-number family looking for a fee exhibit that could not be there. Each sibling was probed through filing.attachments, which fetches the entire .txt submission. Learn CW's S-1 (0001140361-21-010426) transferred 12.2MB for a 2.4MB filing and found nothing — Exhibit 107 postdates it. Siblings filed before that regime are no longer probed, and probing stops at the first match.
Three section-extraction fixes for 10-K item lookup, all reported against 5.44.x with reproductions. Each returned wrong content at full confidence wi
Three section-extraction fixes for 10-K item lookup, all reported against 5.44.x with reproductions. Each returned wrong content at full confidence with no warning, so a caller had no signal anything was off — if you cache Item text or section slices from 5.44.x or earlier, re-check them.
10-K item map shifted by one slot, so each item returned the previous item's body — Foot Locker's FY2024 10-K (0001437749-25-009620) gave Item 6's five-year financial data for obj['Item 7'] and the MD&A under Item 7A; Items 2 and 9B were missing. The body-header scan takes each item's nearest preceding anchor, but Novaworks nests the anchor inside the heading, so every item took the previous one's. It failed silently — every anchor was real. (GH #923)
10-K item map emitting codes that do not exist in Reg S-K — Foot Locker's FY2013 10-K (0001144204-14-019510) returned Item 2P, Item 3L, Item 8C and eleven more, each phantom splitting the real item's content; Item 1 ran to 275,435 chars, now 3,340. TOC rows split label and title across cells, so "Item 4Mine Safety Disclosures" read the title's initial as a suffix. A suffix letter is never followed by a lowercase letter, which now decides it. (GH #923)
Two-column tables of contents put items under the wrong Part — Ambac's FY2022 10-K (0000874501-23-000040) returned 538,701 chars for obj['Item 7'], roughly 70% of it Items 8 through 15, and dropped Item 4. Its TOC interleaves two columns, so one running part context saw both columns' Part headers, scrambling the order boundaries are sorted by. Columns are now read one at a time; Item 7 is 158,411 chars. (GH #924)
The three fixes above were all reported, with reproductions, by g-carmichael.
Full changelog: https://github.com/dgunning/edgartools/compare/v5.45.0...v5.45.1
It is removed; every document now takes one pipeline. streaming_threshold remains as a deprecated no-op.
Text extraction correctness in edgar/documents. Two things before upgrading:
Filing.text(), markdown and section slices will differ from 5.44.x. Almost all of it is repair, but cached text, stored hashes and embeddings will need regenerating.ParagraphNode.text() no longer spaces adjacent inline elements by tag name. An allowlist invented a space between any two adjacent inline elements, corrupting the Item headings section matchers key on (Item 1A. RI SK FACTORS). It is replaced by three signals that read the boundary: a CSS gap, a fixed-width marker box, a standalone list or checkbox glyph. Of 245 spaces removed across 57 fixtures, 222 are confirmed repairs; two elements with no signal at all are now joined, the deliberate cost.<div>s. Citigroup's FY2024 10-K returned 830K chars against 1.81M, missing MD&A, Risk Factors and Financial Statements entirely, and that path was slower than the one it relieved. It is removed; every document now takes one pipeline. streaming_threshold remains as a deprecated no-op.TheFederal Reserve, threereportable business segments, •MacBook Pro 16-in., Yes☒ on a dozen large-cap cover pages. Morgan Stanley's FY2024 10-K alone had ~490. Boundary whitespace is now collapsed and never deleted; genuinely unspaced splits stay glued.HeadingNode out of it, so 23.6% of all headings were glyphs or bare enumerators (Meta's FY2024 10-K: 180 of 296). They surfaced through doc.headings, through markdown as ### •, and through document search. Header detection now requires content that could be a heading.text() on very large tables taking hours — a dimension pass ran before the grid existed and rescanned an empty structure once per preceding row. Fannie Mae's 0000310522-18-000010 took 1h12m, now 24.1s, byte-identical..hdr.sgml artifact as a submission's text file (0000950123-96-000525), which was rejected outright. Behind it, Schedule 13D/G filers arrive under <FILED-BY> rather than <FILER>, so once routed the filing parsed "successfully" with zero filers and a None CIK.try: that turned a raising log handler, the standard way to locate a warning, into a parse failure. It now logs to edgar.documents.parser with the exception type and a content preview.`filing["Item 7"]` hung indefinitely on some 10-Ks — CrossReferenceIndex.has_index() matched the cross-reference heading, then probed for the index ta
filing["Item 7"] hung indefinitely on some 10-Ks — CrossReferenceIndex.has_index() matched the cross-reference heading, then probed for the index table with a single regex nesting six lazy quantifiers under DOTALL against the entire filing HTML. Where the heading matched but the table shape did not, it backtracked catastrophically: on ODP Corp's FY2025 10-K (5.6MB) the call did not finish within 45 seconds. A successful match returned instantly, so only the non-matching case was affected. It is reached from TenK.__getitem__, so any item lookup on an affected 10-K hung — and because re holds the GIL throughout, one such filing froze every other thread in the process, presenting as a whole-process hang rather than one slow filing. Detection is now anchored to the index table _find_index_table() already locates and scans it row by row, which is linear. Bounding the search window alone was not enough: on dense markup the old pattern exceeded 10s at 3.6K chars, so any fixed window stayed exploitable. GE and Citigroup still detect and parse unchanged. (GH #928)obj['PART II, Item 1'], opening mid-sentence inside the MD&A's cross-reference to "Item 1 'Business — Regulation' … in our 2023 Form 10-K" (same shape on CSCO and CTSH 10-Qs). The TOC anchor was correct, but the correctly-anchored Legal Proceedings stub is under 200 chars, which trips the short-section rescue; the rescue's on-heading check demanded the 10-K title for the item number (Item 1 → BUSINESS), misjudged the 10-Q text, and then regex-hunted the raw HTML document-wide, matching the cross-reference. A text opening with the item's own "ITEM N" heading now counts as correctly anchored whatever the form, and the rescue's search is bounded to the window between the section's start and end anchors — its premise is an anchor that landed just before the body, so the real heading is never behind the anchor. Thanks to @sf1tzp for the diagnosis and the fix. The second half of #918 — Freddie Mac's numbered "List of Tables" index rows fabricating part_*_item_N anchors from bare row numbers — is unfixed and that issue stays open. (GH #918)10-K Item 7A silently duplicating Item 7 — Regions Financial's FY2021 10-K (0001281761-22-000016) returned Item 7's 194K-char MD&A for obj['Item 7A'],
0001281761-22-000016) returned Item 7's 194K-char MD&A for obj['Item 7A'], at full confidence with no warning: the TOC links only page numbers, and both items start on page 41, so they collided on one anchor. Colliding items are now re-resolved from their own body headings, bounded so a boundary can't invert. (GH #920)TenK['Item 3'] (should be 1,165). Workiva gives a row's "Item N." label link a different href than its title link, and the label hrefs are broken, so items anchored in the wrong place or vanished when their title wasn't in the keyword vocabulary. Each TOC row is now resolved as a whole. (GH #915)edgartools[ai] installing an incompatible mcp — mcp 2.0.0 removed the decorator-based Server API that edgar/ai/mcp/server.py binds at import, so an unpinned resolve made edgartools-mcp fail to start. Capped at mcp>=1.12.3,<2.0.0 until the server is ported. (GH #917)FilingSGML.text() returning raw XML for ownership forms — Forms 3/4/5 have an <ownershipDocument> XML primary, which failed the is-it-HTML check and came back as tag soup on the SEC's highest-volume form type. Ownership XML is now detected by root element and rendered through Ownership.to_html(). Filing.text()/.html() are deliberately untouched — their contracts are load-bearing (see the d216a934 revert).<!--DOCTYPE HTML PUBLIC ...> (a typo for <!DOCTYPE), and lxml treats the stray <!-- as running to end of input, so the tree came back empty and text() raised HTMLParsingError (0001034670-01-500017). Unterminated comments are now closed at end-of-line; verified byte-identical across a 44-filing corpus.FilingSGML.text() returning mojibake for PDF-only filings — UPLOAD comment letters with a scanned-PDF primary were decoded as UTF-8, yielding pages of replacement characters that poisoned search and markdown downstream. Binary primaries are now detected before any decode: text() returns the SEC's TEXT-EXTRACT sibling when present, otherwise None..items dropping items — two causes. 2005-era filings glue an item header onto a horizontal rule (------Item 4.02 Non-Reliance...), which the line-anchored regex missed (GMAC 0000040729-05-000026); and when the HTML strategies returned a partial set, the text strategy was never consulted (Cimarex 0001047469-05-006981). Headers are now matched after a rule, and all three strategies are unioned. Validated on 66 modern 8-Ks with no spurious items.FILER block with a flat top-level section carried the old subheader into the new one, emitting Subheader ... not found warnings and dropping that section's values; on 1990s FORMER COMPANY entries the same leak raised KeyError. The subheader now resets per section, flat lines are captured, and the warning carries the accession number.available_quarters() hardcoded EDGAR's start as 1994 Q3, so get_by_accession_number() and get_filings() structurally rejected earlier periods even though SEC's full-index serves back to 1993 Q1 (verified; 1992 returns errors). The boundary and the open-ended filing_date default now start at 1993 Q1.colspan exhausting memory — filing 0001193125-06-185884 carries colspan="376967340"; the matrix allocated that many cells per row, reaching hundreds of GB (~700 GB reported from a 7M-filing crawl) before being killed. Spans are now clamped (colspan ≤ 1000, rowspan ≤ 10000) with a 2000-column cap and grid-bounded placement loops. Parses in ~4s at 0.3 GB..//tr, .//td), so a row nested N deep was processed once per ancestor: on 0000880195-09-000191 (8,207 tables, 26 deep) rows were processed 11.3× over, turning FilingSGML.text() into a 3-hour call and duplicating inner cells into outer tables. Traversal is now scoped to each table's own rows. Parses in ~7s.497K fee-waiver / net-expense mix-up — for fee tables using the standard "After Fee Waiver and Reimbursement" wording, ShareClassFees.fee_waiver held
ShareClassFees.fee_waiver held the net expense ratio and net_expenses was None (ProShares UltraPro QQQ: waiver=0.84 where the true waiver is 0.13). Net expenses are now matched first and the waiver normalized to a signed reduction, so total_annual_expenses + fee_waiver == net_expenses. (GH #912)"0.67 1" → 0.671), and "(0.37%)" came back None; every percentage field was affected, not just the waiver. The parser now unwraps parenthesised negatives and takes the leading numeric token before a marker, while dates, labels and period headers still parse as None. (GH #912)income_statement(period='ttm') gave GOOGL basic EPS of 9.45 for Q2 2026 instead of 20.16, varying with the requested period count: windows were labelled by SEC fiscal_year, which tags a re-filed comparative quarter with the filing's year, so two collided and the older won. Labels now derive from period_end. Per-share values also render with 2 decimals. (GH #910)series_only=True ignored for series and class IDs — Fund("S000026864").get_filings(..., series_only=True) returned the umbrella trust's 444 filings where Fund("VCLT") for the same fund returned 26; _target_series_id was set only on the ticker path, so the call fell through to trust-wide delegation. It is now backfilled from the resolved hierarchy, so ticker, series ID and class ID agree. (GH #909)TTMCalculator.quarterize() returned only Q1 for concepts like NetCashProvidedByUsedInInvestingActivities (12 quarters instead of ~48 for GOOGL), because _is_positive_concept substring-matched 'cash' and treated negative flows as data-quality errors. Cash flow and IncreaseDecreaseIn* lines are now signed; GOOGL FY2024 investing reconciles to the reported −45.536B. (GH #907)TenK['Item 7'] ran 257K chars to the director signatures with Items 7A/8/9A/10/11 embedded. Body-header recovery now union-merges the items the TOC missed (TOC wins conflicts), the bold check accepts split-span headers, and a guardrail flags any section embedding a later item's header. Standard filers parse byte-identically. (GH #904)part_i_item_6 (165K chars) at full confidence with no warning. The 10-Q FormSchema now declares per-part item ranges (Part I: 1–4, Part II: 1–6) and validation flags anything outside them, appending a warning and reducing confidence rather than dropping content. (GH #905)`RegistrationS3.sections` / `.section()` for S-3 shelf registrations — filing.obj() for an S-3 now exposes section-scoped Reg S-K access (risk_factors
RegistrationS3.sections / .section() for S-3 shelf registrations — filing.obj() for an S-3 now exposes section-scoped Reg S-K access (risk_factors, use_of_proceeds, plan_of_distribution, selling_stockholders, …), the same surface RegistrationS1 and Prospectus424B already provide. Previously S-3 fell through to the non-title-based default schema and could not be section-scoped, blocking RAG/NLP workflows over shelf registrations. A new S3_SCHEMA covering S-3, S-3/A, and S-3ASR reuses the shared S-1 prospectus vocabulary; a short filing with no resolvable titles falls back to a single full section (no content lost), and item-based forms (8-K, 10-K, …) are untouched. (GH #877)Company('GE').get_facts().balance_sheet() omitted Property, Plant & Equipment entirely for FY2021 onward. GE stopped reporting us-gaap:PropertyPlantAndEquipmentNet after FY2020 and now presents the net line only under a company-specific extension tag that the SEC companyfacts API does not expose, so the standardized statement built the row empty and dropped it. The builder now reconstructs standard 'Net' balance-sheet lines from component concepts the filer still reports (PP&E as PropertyPlantAndEquipmentGross − AccumulatedDepreciation…), matched to each displayed period by period_end — GE's components survive only as prior-year-end comparatives in later 10-Qs, tagged Q1–Q3 of the following fiscal year, never FY. A period already reporting the concept directly is untouched, and every component must be present or the period is skipped (no gross-as-net). The operating-lease ROU asset GE folds into its extension line renders on its own standardized row. The XBRL path (get_financials().balance_sheet()) was unaffected. (GH #894)filing.text() — the edgar.documents parser silently lost the text after an ix:nonfraction's closing tag, usually the unit word and the rest of the sentence, so "$95.2 billion, of which substantially all will be paid" rendered as a bare "$ 95.2". _get_element_text collected each child's text but never its lxml .tail; it now captures that trailing text (including after skipped ix:exclude children, whose content is still dropped). On NVIDIA's FY2026 10-K, scale words ('billion'/'million') and trailing clauses are recovered across dozens of facts; filings without the inline-container pattern render byte-for-byte unchanged. (GH #898)Unknown resource — the MCP SDK passes registered resource handlers a Pydantic AnyUrl, which did not compare equal to the string literals the handler matched against, so every listed resource URI failed to read even though it listed correctly. Incoming URIs are now normalized to strings before matching. (GH #897)parse_html() no longer eagerly renders the full document text on every parse. Document statistics (text_length, table_count, …) are computed lazily on first access to metadata.statistics rather than during post-processing, and two hot preprocessing regexes (repeated-<br> collapsing and sentence-spacing repair) were rewritten to avoid backtracking. Measured 583 ms vs 726 ms (Apple 10-K) through 1,534 ms vs 1,871 ms (Oracle 10-K); rendered output is unchanged, verified by SHA-256 digest of the full rendered text. (GH #900)`RegistrationS4` data object for S-4 / F-4 registrations — filing.obj() now returns a typed object for merger/acquisition/de-SPAC registrations, expos
RegistrationS4 data object for S-4 / F-4 registrations — filing.obj() now returns a typed object for merger/acquisition/de-SPAC registrations, exposing the standard registration field surface (cover_page, fee_table, total_offering, net_fee, securities, …) plus is_foreign (F-4) and an offering_type classifier. Covers S-4, S-4/A, F-4, F-4/A. (GH #876)GET /health liveness endpoint on the MCP HTTP server — when the MCP server runs in HTTP transport mode it now exposes an unauthenticated GET /health returning {"status": "ok", "version": <server version>}, giving container orchestrators (Docker HEALTHCHECK, Kubernetes liveness/readiness probes) a cheap liveness signal without the full MCP handshake that the /mcp endpoint requires. (GH #882)get_concept() stale-tag data — for companies that switched GAAP tags (e.g. NVDA/AMZN moving capex from PaymentsToAcquirePropertyPlantAndEquipment to PaymentsToAcquireProductiveAssets), get_concept() now picks the most recent fact across all synonyms instead of returning the first, stale one; return_metadata=True also carries the resolved period/period_end/filing_date. An explicit period= still resolves by priority. (GH #892)get_ttm_revenue() / get_ttm_net_income() stale values — these now evaluate every candidate concept and use the one whose TTM window ends most recently, instead of the first that resolves; tag-migrated companies are corrected (NVDA $10.9B→$253.5B, GOOG no longer a year behind). Adds a TTMMetric.is_stale flag (and warning) when the newest quarter lags the reference date. (GH #893).item / .part metadata — sections detected by the pattern extractor (e.g. small-cap 10-Ks that split the Item 7. fragment from its MD&A title) now resolve their part and item instead of returning None, so semantic keys like mda/business carry correct .item/.part. (GH #891)Part I / Item 4 heading leak — Controls and Procedures no longer absorbs the trailing "PART II — OTHER INFORMATION" heading; the legitimate "PART C — OTHER INFORMATION" heading of S-1/N-1A/N-2 filings is left intact. (GH #883)Fund() ticker/Class-ID collision — ETF tickers starting with "C" that have real SEC series/class registration (CIBR, CQQQ, COPX, CARZ, CALF, …) now resolve; fund.company/.series/.share_class return the object or None per their contract instead of raising AttributeError. (GH #889, #890)`Attachments.query()` no longer uses `eval()` — filter strings were passed to eval() with builtins reachable, allowing arbitrary code execution. Queri
Attachments.query() no longer uses eval() — filter strings were passed to eval() with builtins reachable, allowing arbitrary code execution. Queries are now evaluated against a restricted AST; disallowed input raises ValueError and legitimate queries are unchanged. (GH #884)Fund.get_filings(series_only=True) now actually filters to the fund's series — it silently returned the whole umbrella trust's filings (a sibling series' data). It now resolves the series via SEC browse-edgar, pushing the form filter server-side so even large funds (e.g. VOO) return reliably, and returns an empty Filings rather than the trust when nothing matches. (GH #888)obj() no longer crashes when a transaction has no <transactionCoding> — the missing Code column raised AttributeError, making the whole filing unreachable. Uncoded transactions now degrade to TransactionType=None in both the non-derivative and derivative tables. (GH #887)<img>) are no longer dropped from the new parser's markdown — Document.to_markdown() now renders images as  (resolving relative src against the document URL when known); TextExtractor gains an opt-in include_images placeholder. Filing.markdown() still uses the legacy parser and is tracked separately. (GH #886)get_revenue() / get_net_income() / get_operating_income() no longer return a prior-year value on multi-duration 10-Qs — the getters picked the period column positionally, but to_dataframe() columns aren't recency-ordered. They now order by period metadata (current reporting period first). Annual filings and balance-sheet getters are unaffected. (GH #885)Company.reit_subtype no longer mislabels net-lease equity REITs as mortgage — a stale/trivial interest line flipped equity REITs like WPC to mortgage. Classification now compares the magnitude of property income against net interest income (mortgage only when interest is at least 10% of property income). (GH #854)8-K last item no longer absorbs the SIGNATURES block — the SIGNATURES section following the last reported item (e.g. Item 5.02 on Meta's 0001628280-25
0001628280-25-058337) leaked into that item's text because _EIGHT_K_SECTION_PATTERNS had no 'signatures' entry, leaving the pattern extractor with no terminal boundary. Additionally, the bold-child header strategy (Strategy 3b) was gated to 10-K only, missing Workiva-style SIGNATURES headings rendered with font-weight:700 on a child <span> (not the paragraph itself), and a new strategy (5b) covers plain-text font-weight:400 SIGNATURES headings used by some filers (e.g. JPMorgan). The SIGNATURES block is now accessible as ek.document.sections.named("signatures") and is excluded from ek.items. (edgartools-papt, GH #879)TenK[item] no longer merges adjacent Part III items — on 10-Ks that incorporate Part III by reference with sparse markup (e.g. Tesla FY2022 0000950170-23-001409), Item 10 absorbed the ITEM 11. EXECUTIVE COMPENSATION header and body (687 chars) while Item 11 returned empty, and Part III items were missing from .items entirely. The section vocabulary now includes Part III Items 10–14 and Part IV Item 16, and a bold-child header strategy recognizes item headers rendered as bold text inside a paragraph (not a standalone heading) so each item's boundary is detected even when Part III is a compact "see proxy" stub. Items 10–14 are now independently extractable and listed in TenK.items; the pattern-merge step runs the same validation/guardrails as TOC sections and is gated so complete-TOC filings and non-10-K forms pay no cost. (edgartools-01x4, GH #880).sections no longer misattribute content across boundaries — three fixes to the title-based section engine that surfaced on Airbnb's IPO prospectuses (S-1 0001193125-20-294801, 424B4 0001193125-20-315318): (1) the boundary selector now requires a section's end anchor to be declared at-or-after it in the TOC, so an out-of-order sub-block anchor (Airbnb's "Glossary of Terms", listed before MD&A but anchored inside it) no longer truncates MD&A to ~100 chars; (2) 424B now shares S-1's full prospectus vocabulary — a final IPO prospectus repeats the entire S-1 body, so without the narrative sections (MD&A, Business, Management, …) the authoritative-TOC span split and dilution swallowed ~900KB; (3) a trailing-financials rescue clamps the last narrative section ("Experts" / "Dilution") at the untitled financial-statements (F-pages) block it previously absorbed. (edgartools-ti82, GH #878)Semantic section extraction reaches proxy statements, registration statements, and prospectuses — ProxyStatement, RegistrationS1, and Prospectus424B n
Semantic section extraction reaches proxy statements, registration statements, and prospectuses — ProxyStatement, RegistrationS1, and Prospectus424B now expose Reg S-K sections as section-scoped text over the shared title-based engine. Also adds 13F amendment-type and Form D relationship surfaces, reorganizes the offerings package (with back-compat shims), and fixes a batch of section-boundary and offerings-classification bugs.
ProxyStatement.sections (a dict of named ProxySections) and ProxyStatement.section(name) expose Schedule 14A / Reg S-K sections (proxy_summary, corporate_governance, compensation_discussion_and_analysis, pay_versus_performance, audit_matters, security_ownership, …) as section-scoped text. DEF 14A / PRE 14A now route through the title-based section engine; sections are labelled when heading/TOC anchors resolve and fall back to a single full section otherwise (no content loss). (edgartools-x341, GH #867)RegistrationS1.sections / Prospectus424B.sections and .section(name) expose Reg S-K sections (prospectus_summary, risk_factors, use_of_proceeds, mda, business, management, underwriting, …) over the shared title-based engine, same labelled-or-full contract. (edgartools-ybth, GH #866)Document — section objects now carry a kind, are reachable via named(), and named sections such as .signatures surface in document.sections. (edgartools-nqzc)Person — Person.relationships (e.g. ["Executive Officer", "Director", "Promoter"] from <relatedPersonRelationshipList>) and Person.relationship_clarification. The relationship is the analytically meaningful part of the related-persons section; previously only names and addresses were surfaced. Also shown in FormD rich rendering and to_context(). (edgartools-0dpz, GH #874)ThirteenF — is_amendment, amendment_type ("RESTATEMENT" | "NEW HOLDINGS" | None), amendment_number, and full amendment_info (confidential-treatment fields). The distinction is load-bearing for correctness: a NEW HOLDINGS 13F-HR/A discloses only previously-confidential positions and must be unioned with the original, whereas a RESTATEMENT replaces it — superseding the original by a NEW HOLDINGS amendment silently drops the real portfolio. (edgartools-preg, GH #872)crowdfunding / exempt / prospectus sub-packages (e.g. Form C, Form D, and the 424B/S-1 prospectus surfaces each split into their own modules). Existing from edgar.offerings.* import ... paths keep resolving via back-compat shims. (edgartools-n094)ProxyStatement.voting_proposals no longer emits text fragments as proposals on merger proxies — on some DEFM14A filings the extractor lifted numbered items out of an "Incorporation by Reference" / exhibit-list section (e.g. Veeco showed proposals numbered 2 and 3 whose text was 8-K item references). Proposal extraction is now anchored to the proxy's actual matters-to-be-voted-on structure. (edgartools-7pga, GH #875)_find_section_end only closed a section at a header that is a real section boundary (Item/PART/SIGNATURE/EXHIBIT/…); internal bold sub-headings (e.g. "Adoption of Fiscal Year 2027 Variable Compensation Plan") previously cut Item 5.02 short, dropping the item body. (edgartools-koq3, GH #871)RegistrationS1.underwriting.lead_manager no longer returns garbage table text — a "Shares Eligible for Future Sale" lock-up table was misclassified as an allocation table, leaking the row label "Earliest Date Available for Sale in the Public Market" as the lead underwriter; names are now validated at the source and in the S-1 consumer (ABNB resolves to "Morgan Stanley & Co. LLC"). (GH #868)PIPE_RESALE — the classifier now overrides to IPO inside the PIPE-resale branch when a 424B1/424B4 asserts its own offering is an IPO. (GH #869)A patch release with parser data-correctness fixes for historic SGML headers and EX-107 registration fee tables.
A patch release with parser data-correctness fixes for historic SGML headers and EX-107 registration fee tables.
CONFIRMING COPY: was misread as a section header, dropping every following field (incl. FILED AS OF DATE) so FilingSGML.filing_date returned None. (edgartools-sg9k)ShelfLifecycle.total_offering_capacity recovered from misparsed EX-107 fee tables — the parser picked the registration fee, or a column-misaligned cell, as total_offering_amount for a class of Exhibit 107 layouts, surfacing genuine shelves as null capacity. (edgartools-xn7e)$153.10 per $1,000,000 and $.0000927 no longer leave FeeTableSecurity.fee_rate orders of magnitude too large; dilution tables no longer double-prefix $ on values that already embed the sign.424B offering extraction now consumes the machine-readable EX-FILING FEES inline-XBRL exhibit it already parsed but previously ignored — wiring it int
424B offering extraction now consumes the machine-readable EX-FILING FEES inline-XBRL exhibit it already parsed but previously ignored — wiring it into deal sizing and offering-type classification, adding an IPO offering type, and exposing classifier provenance — plus an lxml rewrite of the exhibit parser and two fixes for offline / local-storage use of historic pre-HTML SGML filings.
Deal.gross_proceeds reads the authoritative EX-FILING FEES total (ffd:TtlOfferingAmt) when the cover-page and pricing-table text paths are missing or implausible, sizing the 424B2/424B5 debt/note and ATM shapes those paths miss. (edgartools-s9uo)ffd:OfferingSctyTp) before falling through to unknown: debt → debt_offering, equity → firm_commitment (low confidence), rights → rights_offering. The exhibit is only fetched on the otherwise-unknown path. (edgartools-2l2i)OfferingType.IPO — a 424B1/424B4 whose cover asserts "this is an initial public offering" now classifies as ipo instead of being folded into firm_commitment; shelf takedowns and follow-ons referencing a past IPO stay firm_commitment. (edgartools-ejk5)Prospectus424B and Deal — public offering_type_confidence, offering_type_signals (incl. xbrl_security_type:* markers), and offering_type_sub_type, also serialized by Deal.to_dict(), so consumers can tier values by how the type was determined. (edgartools-drzj)$1,000 per-note denomination artifact — a plausibility floor (calibrated: artifacts cluster at exactly $1,000, real deals are ≥ $100k) suppresses these, superseded by the XBRL total where an exhibit exists. (edgartools-s9uo)<span> security titles keep their whitespace — the lxml extractor no longer mashes titles spanning inline elements into runs like "Series APerpetual Stride".ShelfLifecycle.takedowns is returned in chronological order as its contract promised; previously a newest-first _related made avg_days_between_takedowns negative and days_since_last_takedown read the oldest takedown as the most recent. (edgartools-y22m)filing.text() works offline for historic pre-HTML text-only filings — a tightly-scoped path returns the primary <TEXT> body straight from locally-parsed SGML for that shape (a <FILENAME>-less, non-HTML/XML primary with no TEXT-EXTRACT sibling) instead of re-downloading; every other filing is unchanged. Also adds FilingSGML.text() for offline plain-text extraction of the primary document. (edgartools-0rvh)resolve_local_filing_path() now scans adjacent day-folders on read (accession numbers are globally unique), wired into Filing.sgml(), full_text_submission(), and the batch checkers; the local-miss log is downgraded to DEBUG. (edgartools-a3ej)A batch of offerings/prospectus extraction improvements — derived shelf-lifecycle signals, registration fee-table capacity recovery (including pre-202
A batch of offerings/prospectus extraction improvements — derived shelf-lifecycle signals, registration fee-table capacity recovery (including pre-2022 inline "Calculation of Registration Fee" tables), and more reliable lead-underwriter/placement-agent extraction on 424B prospectuses — plus an HTTP/1.1 transport default with a public HTTP/2 opt-in, 8-K item/date fixes, and a 13F portfolio fix.
h2.exceptions.InvalidBodyLengthError / httpx.RemoteProtocolError: ConnectionTerminated that crash long fan-out jobs. SEC's ~9 req/s rate limit means HTTP/2's multiplexing offers no real upside here.EightK.date_of_report / SixK.date_of_report now always return a datetime.date (or None when the filing header has no period of report), instead of a formatted string like 'December 20, 2024'. Consumers no longer need to parse mixed date/string types. (edgartools-83gh)HTTP_MGR.httpx_params. Set EDGAR_USE_HTTP2=true (env var) or call configure_http(http2=True) at runtime to opt back into HTTP/2; get_http_config() now reports the current http2 setting.extract_registration_fee_table() (and shelf/424B fee-capacity) returned None for every such filing. The inline body table is now read directly: the registered capacity is taken from the table's aggregate offering price, and indeterminate Rule 457(r) shelves resolve to a deferred fee. Verified across consumer, medical-device, biotech, energy, financial, tech, and REIT issuers. (edgartools-9q82)ShelfLifecycle — exposes signals consumers were re-deriving by hand: status (registered/effective/expired/withdrawn), is_effective / is_automatic_shelf / is_withdrawn / is_re_registered, program_mode and days_since_last_takedown, continuity (continuous/lapsed, computed gap-aware across generations, with a has_registration_gap data-quality flag), and program_age_days. The continuity logic distinguishes a Rule 415(a)(6) renewal (effective before the prior shelf expired) from a post-gap revival, so a revived shelf is never mislabelled as continuously registered. (edgartools-2w5y)EightK.items no longer under-reports items the new section parser silently missed. When the parser detects only some items (e.g. only Item 9.01 on a filing that also carries an Item 1.05 cybersecurity body), .items now unions the parser result with the chunked primary-document parser so the present item is listed. The eightk['1.05'] / eightk['Item 1.05'] accessor is also fixed to fall through to the text-based extractor instead of returning the chunked parser's None for a key-format mismatch. (edgartools-83gh)get_thirteenf_portfolio() now returns populated holdings instead of always an empty DataFrame. It contained a dead column-rename block and then sorted by a value_usd column that never existed, so the trailing KeyError was swallowed and an empty frame returned for every real filing. It now uses the single canonical PascalCase infotable schema (Issuer, Cusip, Value, …) shared by all parse paths, and adds a pct_value column. (edgartools-i5wx)ShelfLifecycle now derives the shelf's expiration from its current effective date (Rule 415's three-year window) rather than the original filing date, fixing mis-dated expiries that left recent takedowns stranded past a too-early expiry. (edgartools-fu3x)Prospectus424B underwriting/lead_manager now recovers the lead agent from 424B2 structured-note covers, best-efforts and ATM equity covers, and inline agency language ("engaged <firm> … as placement agent"); stitches firm names that wrap across cover-grid lines (e.g. Ladenburg → Ladenburg Thalmann); and filters out garbage table-of-contents/title text that previously leaked in as the underwriter name. (edgartools-2h4c, edgartools-zzr4)A batch of fixes across insider-ownership footnote handling and 10b5-1 plan detection, document search snippet highlighting, BDC/Form C/N-PORT data ac
A batch of fixes across insider-ownership footnote handling and 10b5-1 plan detection, document search snippet highlighting, BDC/Form C/N-PORT data access, TTM quarterly derivation, and date/currency formatting, plus an ISO 4217 currency column on XBRL facts and the Form C parser's migration to lxml.
currency column on the XBRL facts DataFrame — xbrl().facts.to_dataframe() now includes a currency column with each fact's ISO 4217 code (e.g. USD, HKD) resolved from its unit measure. Non-USD filers tag monetary facts with opaque unit ids such as UNIT_STANDARD_HKD_MNUSOXGRF0O9R60JINVDUQ, which were exposed verbatim in unit_ref and made currency-based filtering and display unreliable; the raw unit_ref is preserved, while currency gives a usable code. Per-share monetary units report their numerator currency, and non-monetary units (shares, pure, custom) resolve to None rather than a misleading value. (#850)<transactionCoding>, so TransactionActivity.footnote_ids / .footnotes_text came through empty for most filings. The extractor now gathers footnote IDs from the whole transaction (deduplicated), so per-transaction footnote reasoning — including is_10b5_1_plan — works directly rather than relying on the filing-wide fallback.SearchResult.snippet highlights context[start_offset:end_offset], but text, whole-word, and regex search stored absolute text offsets against a context string that was truncated (with … markers) around the match, so any match after the leading context window highlighted the wrong characters. Offsets are now computed relative to the returned context. (#860, #862)table:<term> search built a result whose end offset was the full table length against a context truncated to ~200 chars, so SearchResult.snippet wrapped the entire truncated context in **…** for any table longer than 200 characters. The match is now located within the table and a correct context window is produced, consistent with text and regex search.aff10b5One value before falling back to footnote text. When the checkbox is absent (e.g. pre-2023 filings), has_10b5_1_plan now scans the filing's full footnote set instead of relying solely on per-transaction footnote attribution, which is frequently empty. The fallback matcher no longer confuses the separate anti-fraud Rule 10b-5 with Rule 10b5-1 trading plans while recognizing common spacing and dash variants. (#863)edgar package logger no longer leaks log output in unconfigured applications — the package-root logger had no NullHandler, so when an application had not configured logging itself, edgartools warnings fell back to Python's logging.lastResort handler and were written to stderr. This was especially harmful in MCP / stdio environments, where stray stderr output corrupts the protocol stream. A NullHandler is now attached to the edgar logger per the Python logging HOWTO for libraries, keeping edgartools silent until the application opts into logging. (#856)/A) and filings with no XBRL attachments are normal, expected cases; the "no XBRL data" messages they produced are now logged at DEBUG rather than WARNING, so they no longer alarm users during routine processing. (#857)/files/structureddata/data/ to /files/datastandardsinnovation/data/, which made fetch_bdc_dataset() and related functions fail with 404; the base URL now points to the new location.FormC.filer_information.ccc no longer duplicates the CIK — the Form C parser read the filerCik element into both cik and ccc; ccc now reads the actual filerCcc element (always redacted to XXXXXXXX in disseminated filings) and is Optional, returning None when absent.FormC.filer_information.live_or_test is no longer always False — the parser looked for a testOrLive element under filer, but the schema places liveTestFlag under filerInfo; LIVE filings now correctly report True (older testOrLive documents remain supported).TTMCalculator derives Q4 as FY - (Q1+Q2+Q3). It previously selected the three input quarters by their fiscal_period label, but the SEC tags comparative facts in re-filings with the filing's fiscal period, so the same calendar quarter could appear labeled Q1, Q2 and Q3 across successive 10-Qs — producing a wrong, often negative Q4 (e.g. GAIN InvestmentCompanyDividendDistribution: 57.2M - 3×28.8M = -29.2M). Quarters are now selected by distinct calendar period (dedup by period_end, latest periodic filing wins), and derivation is skipped when a discrete Q4 is already reported. This affects quarterize(), TTM calculations, and quarterly statement views. (#848)format_currency_short rolls up to the next unit at magnitude boundaries — a value just under 1B (>= ~999.95M) rounded to 1,000.0 within the millions bucket and rendered as the nonsensical $1,000.0M instead of $1.0B; it now promotes to the billions unit when the millions mantissa rounds up to 1,000.datefmt no longer crashes on None or non-date values — the display-only date helper called value.strftime on its non-string branch unconditionally, so a None (e.g. a former name's open-ended to date, or a missing date_of_change) raised AttributeError and took down the whole header/former-name table render. It now returns "" for None, formats date/datetime objects, and degrades to str(value) for anything unexpected. (Complements the unrecognized-string pass-through fix in #859.)datefmt no longer crashes on unrecognized date strings — the helper parsed only YYYYMMDD, YYYYMMDDHHMMSS and YYYY-MM-DD strings and called str.strftime on everything else, raising AttributeError: 'str' object has no attribute 'strftime' for any other value (e.g. 2022/03/04, a non-zero-padded date, or an empty string). Unrecognized strings are now returned unchanged so date display in filing-header and former-name tables degrades gracefully instead of crashing.reverse_name handles a generational suffix between surname and given name — SEC's LAST SUFFIX FIRST ordering can place a suffix (III, Jr, …) immediately after the surname and before the given name (e.g. the PPG insider ROBERTS III CHRIS); reverse_name() treated it as part of the given name and produced Iii Chris Roberts. Leading suffix tokens are now pulled out of the given-name parts, so the name renders as Chris Roberts III.FormC.from_xml now uses the same lxml parsing pattern as the fund reports, with verified field-level parity across all Form C variants (C, C-U, C-AR, C-TR). Optional fields on the Form C models now declare explicit None defaults.A batch of robustness fixes across insider-ownership context, 6-K exhibit decoding, fund/N-PORT filing access, Schedule 13D/G, and XBRL depreciation s
A batch of robustness fixes across insider-ownership context, 6-K exhibit decoding, fund/N-PORT filing access, Schedule 13D/G, and XBRL depreciation standardization, plus an internal restructure of the ownership module. No public API changes.
form='N-PORT' resolves to NPORT-P — get_filings(form='N-PORT') and related queries now match the actual SEC form type (NPORT-P) via a form-name alias, so the intuitive name returns results instead of an empty set. (#843)Ownership.to_context() no longer crashes on string share values — Form 3/4/5 filings whose share amounts carried footnote references or other non-numeric text raised a TypeError when building the AI context string; the value is now coerced safely so to_context() always returns a string. (#846)SixK.text() no longer crashes on bytes exhibit content — 6-K exhibits whose Attachment.download() returns bytes raised TypeError: a bytes-like object is required, not 'str' in the legacy HTML parser's <TEXT> check; the parser now decodes bytes first, so bytes and str inputs parse identically. Non-UTF-8 exhibits (cp1252/latin-1, common in older filings) decode correctly via a cp1252→latin-1 fallback instead of emitting replacement characters. (#844)SixK.text() skips binary exhibits — .xlsx and .zip attachments are now classified as binary so SixK.text() no longer attempts to decode them as HTML/text. (#844)Fund(ticker).get_filings(series_only=True) now isolates the series — the flag previously returned filings beyond the requested series; series filtering is now applied correctly. (#843)N-PORT entry in FILER_TYPE_DOMESTIC_FORMS — the stale entry meant N-PORT filer-type filtering matched nothing. (#843)obj() returns a partial object instead of silent None — a parsing gap previously caused filing.obj() to return None for some 13D/G filings; it now returns a partial object so callers get the data that did parse rather than nothing, and the to_context() navigation hints for these filings were corrected. (#840, #841)OtherDepreciationAndAmortization no longer breaks standardized cash flow — filers reporting D&A under the OtherDepreciationAndAmortization concept had the primary D&A line dropped from the standardized cash-flow statement and the concept misclassified as non-operating income in XBRL standardization. The line is now retained and classified correctly, with the orphan-fold dedup hardened against duplicate facts. (#839)httpxthrottlecache <0.5.0 to avoid a breaking httpx2 fork in the 0.5.x line.edgar/ownership/ownershipforms.py split into focused submodules — the 2,279-line module was decomposed into models, core, tables, table_containers, owners, summary_records, summary, forms, and text_render (each under 600 lines). ownershipforms.py remains a backward-compatibility shim re-exporting every previously public name, and the edgar.ownership package surface is unchanged — verified by a new public-API guard test. Pure structural refactor with no behavior change.10-K section detection and agent TOC parsing receive two targeted fixes that close gaps introduced in 5.34.0.
10-K section detection and agent TOC parsing receive two targeted fixes that close gaps introduced in 5.34.0.
BDC non-accrual extraction no longer depends on a filer phrasing its footnotes exactly the way our whitelist expected, and a parsing gap is now surfac
BDC non-accrual extraction no longer depends on a filer phrasing its footnotes exactly the way our whitelist expected, and a parsing gap is now surfaced as a warning rather than read as a confirmed zero.
edgar.__version__ — the installed version is now exposed at the package root (import edgar; edgar.__version__), following the standard pkg.__version__ convention so downstream consumers can detect which version they have without reading edgar.__about__ or running pip show. (#794)NonAccrualResult.warnings — flags a portfolio that produced no non-accrual signal from any extraction layer, and recognized flags that resolved no investments, so an LLM consumer never mistakes a parsing gap for a confirmed zero. Surfaced in to_context, mirroring the Section.warnings pattern.Install: pip install -U edgartools
Full Changelog: https://github.com/dgunning/edgartools/compare/v5.34.0...v5.35.0
SEC section extraction is now form-aware by design: form structure is declarative data rather than 10-K-shaped heuristics, link-less-TOC bank filings
SEC section extraction is now form-aware by design: form structure is declarative data rather than 10-K-shaped heuristics, link-less-TOC bank filings (Goldman Sachs, Citigroup) extract their items correctly, and wrong-content sections are flagged instead of trusted.
Section.markdown() now works on TOC-detected sections — slices the section HTML and renders structure-preserving markdown (tables, lists) instead of falling back to flat text. Completes the Section.markdown() work from 5.32.0.form_schema.py) instead of branches in the TOC analyzer; supporting a new form is now a table entry.Section.warnings — flags sections whose content size is anomalous (truncated or over-captured) instead of returning them at high confidence.TenQ['Item 1'] returned Legal Proceedings instead of Financial Statements — pre-header 10-Q items were keyed without their Part prefix, so lookups fell through to Part II.get_company() silently returned None — SEC now types fund CIKs as numeric (225323.0), which broke key matching; CIKs are normalized through int so all forms key identically.TenK.items now returns canonical SEC order (1, 1A, … 16) on all paths, not detection order."Item 8" in sections still works.'part' no longer false-matches inside words like "counterparties" when inferring Part context.ct.pq (CUSIP→ticker, 13F rendering) refreshed from SEC Fails-to-Deliver and merged to preserve coverage (68,512 CUSIPs); company_tickers.parquet (ticker↔CIK resolution) refreshed as a clean mirror of SEC's current data (10,365 entries).pip install edgartools==5.34.0
`Filing.search()` highlights matched query terms in its output — search results now mark the terms that matched within each section, so you can see *w
Filing.search() highlights matched query terms in its output — search results now mark the terms that matched within each section, so you can see why a section was returned rather than just that it was. Complements the BM25/regex section-index fix shipped for the same issue in 5.32.0. (#765)Section.tables() returned each table up to ~24× on TOC-detected sections — Section._extract_section_html walked the section subtree with iterwalk and re-serialized every collected element via tostring(). Because tostring() already includes an element's full subtree, a <table> nested under collected ancestors was emitted once for itself plus once inside each ancestor, and _get_tables_from_toc_section then wrapped each copy as a distinct TableNode. AAPL's 10-K Item 8 returned 123 tables for 34 unique; deeply nested 20-F sections hit ~24× per table. Only top-level collected elements are serialized now (a parent's serialization already covers its descendants), so each table appears exactly once. Section.text() is byte-identical before and after — no content drift. (#826, reporter @HonzaCuhel)
XBRL statements rendered duplicate rows when a concept had repeated presentation arcs — duplicate presentation arcs pointing at the same concept produced repeated lines in render(). Arcs to the same concept are now de-duplicated, with roll-forward (beginning/ending balance) arcs exempted so cash-flow and equity roll-forwards still render both their opening and closing balance rows. (#825)
Embedded tables inside XBRL TextBlock report cells were dropped — SGML/HTML tables nested within a TextBlock disclosure are now rendered in the report cell instead of being silently omitted. (#755)
13F value-unit (thousands vs dollars) was inferred from a global filing-date cutoff — the thousands/dollars scale is now detected per-filing from the filing's own data rather than a date heuristic, fixing misscaled holding values for filings near the cutoff boundary.
`import edgar` emitted `DeprecationWarning` on every startup — the legacy HTML modules (edgar.files.html_documents, edgar.files.html, edgar.files.html…
xbrl.calculation_linkbase() DataFrame — exposes the per-filing calculation linkbase as one row per parent→child arc, with signed weight, role URI, taxonomy attribution (us-gaap vs filer extension), and SEC menucat classification. Enables external pipelines (e.g., bank revenue disaggregation, REIT rental income rollups) to build per-filer concept hierarchies without re-parsing _cal.xml. Layer 1 of the GH #766 implementation plan; the parser was already producing this data on CalculationTree/CalculationNode, this is a DataFrame projection over existing output. (#766)
Statement.extension_arcs() — surfaces filer-authored concepts that participate in a statement's calculation linkbase but are absent from its presentation tree, i.e. concepts that silently drop from render() output today. Opt-in via Statement.extension_arcs(include_values=False); default mode returns one ExtensionArc per concept (structural), include_values=True emits one per (concept, context) with the instance value attached. The existing render() path is untouched. Layer 2 of GH #766. Ground-truth verified on JPM FY2023 10-K cash flow (jpm:NetChangeInAdvancesToandInvestmentsInSubsidiaries, jpm:NetBorrowingsFromSubsidiaries — both calc-present, presentation-absent). (#766)
Section.markdown() accessor — closes the gap between Section.text() (item-aware but flattens tables and bullet lists) and Filing.markdown() (preserves structure but whole-document only). Per-item chunkers / RAG pipelines can now get structure-preserving markdown scoped to a single section. Pattern/heading-detected sections render the cached node tree via MarkdownRenderer; TOC-detected sections currently fall back to Section.text() to avoid corrupting adjacent-section markup (full TOC support tracked as a follow-up). Real-filing regression on AAPL 8-K Item 9.01 exhibit table locks in the pipe-table contract. (#833, contributor @HonzaCuhel)
StreamingParser dropped 20%+ of text from <span>-wrapped paragraphs on large filings — for SEC filings crossing the 10 MB streaming threshold (so most ~30–110 MB 10-Ks/20-Fs), filing.text() silently returned output 20%+ shorter than the non-streaming path. Two compounding bugs in the iterparse loop: elem.clear() ran on every event (both start and end), and ran on every element regardless of whether an enclosing structural element (<p>, <h1>–<h6>, <section>) had finished reading its children. Since SEC filings wrap virtually every word in <span style="…">, the inner <span>'s end event cleared .text/.tail before the enclosing <p> could read them — paragraphs came out empty, with no warning. Clearing now runs only on end events and is gated on a new _content_depth counter (mirroring the existing _table_depth gate). A separate gate prevents <p>/<h*>/<section> inside <td> from being emitted twice. (#830, contributor @kevinchiu)
HTTP_MGR had no default timeout — stalled requests could block workers indefinitely — the internal httpx client was constructed without a timeout, so a stalled upstream or slow TLS handshake could pin a worker on an uninterruptible socket read syscall. Downstream users observed processes running 50+ minutes past their job budget on a single request. get_http_mgr() now sets Timeout(30.0, connect=10.0) by default; EDGAR_HTTP_TIMEOUT (seconds) configures it statically and the existing configure_http(timeout=...) runtime API still works. Callers that need unbounded waits can opt out explicitly. (#831, contributor @kevinchiu)
13F-HR holdings merged Put/Call positions into the underlying equity row — ThirteenF.holdings grouped by CUSIP alone, so Put/Call rows aggregated into the same security's equity row and the PutCall column was lost on the merged result. Categories also used uppercase PUT/CALL while SEC XML emits title-case Put/Call, so the categorical conversion silently dropped those values too. Group key now includes PutCall when the column exists; category labels match SEC XML. Regression verified on SG Capital Management 13F-HR/A (3 distinct Put positions preserved in the aggregated view). (#824)
import edgar emitted DeprecationWarning on every startup — the legacy HTML modules (edgar.files.html_documents, edgar.files.html, edgar.files.htmltools) emitted warnings at module top, and edgartools' own startup cascade imports them, so the warnings fired on every fresh import. Downstream test suites running under -W error (a recommended pytest setup) had to install warning filters just to let import edgar succeed. The deprecation signal moved from module top to per-class __init__, so internal callers don't trip the warning while user-instantiated legacy classes still do. (#832, contributor @kevinchiu)
Filing.search() / Filing.grep() returned nothing on pre-2002 plain-text filings — Filing.search() raised AssertionError and Filing.grep() returned 0 matches on plain-text filings (e.g. PCG's 1999 10-K). Both relied on attachment iteration that finds nothing because SGML decomposition emits empty shells for text-only filings. sections() now falls back to chunking filing.text() on <PAGE> markers or blank lines when html() is None, and grep() falls back to filing.text() when no attachment yields usable text. (#819)
TOC analyzer fabricated phantom Items on 10-Q filings — TOCAnalyzer had three 10-K-shaped heuristics that fired regardless of form: it accepted any bare number 1–15 as an item identifier in preceding-<td> siblings (so a page-number cell like <td>8</td> became "Item 8"); it mapped any "financial statements" link to "Item 8" (correct for 10-K, wrong for 10-Q where Financial Statements is Part I, Item 1); and it sorted using a 10-K-shaped section-order table. All three heuristics are now form-guarded. (#827, contributor @HonzaCuhel)
SearchResults panel labels conflated BM25 rank with section index — SearchResults.__rich__ used the enumeration rank of the sorted display as the panel title, so the same numeric label meant different things in the BM25 and regex paths (BM25 sorts by score, regex preserves original order). "0" in BM25 output was the top-scoring section while "0" in regex output was the first section that matched, and the two were rarely the same. Panels now display DocSection.loc — the section's index in filing.sections() — consistently across search methods, so callers can index back into the corpus regardless of search mode. (#765)
calculation_linkbase() and Statement.extension_arcs() documented alongside Phase 1 and Phase 2 of the GH #766 implementation, including the difference from presentation linkbase and worked examples on real filings. (#766, Phase 3)Contributors: @HonzaCuhel, @kevinchiu, @0ywfe
Full changelog: v5.31.5...v5.32.0
`xbrl.facts.to_dataframe()` mislabeled Q2/Q3 as Q3/Q4 for 52/53-week fiscal-year filers (JNJ, PFE, AAPL, COST) — the XBRL instance parser's _quarter_f
xbrl.facts.to_dataframe() mislabeled Q2/Q3 as Q3/Q4 for 52/53-week fiscal-year filers (JNJ, PFE, AAPL, COST) — the XBRL instance parser's _quarter_for_date classified the fiscal quarter from the raw calendar month of the period end. 52/53-week issuers pin quarter ends to a weekday near the calendar quarter boundary, so the period_end can drift into the first days of the following month — JNJ Q2 2023 ended 2023-07-02, Q3 2023 ended 2023-10-01 — bucketing those facts into the next quarter. The EntityFacts layer already handled this via calculate_fiscal_year_for_label, but the XBRL parser has an independent fiscal classification path feeding xbrl.facts.to_dataframe() and query().by_fiscal_period(...), silently misclassifying quarterly data for any RAG / analytics pipeline reading raw facts. End dates in the first 7 days of a month are now treated as belonging to the previous month for quarter classification; the 7-day window covers max drift for Sunday-nearest (≤3 days), Saturday-nearest (≤1 day), and last-Sat/Sun (no drift) patterns with safety margin. (#816, reporter @kmatosli)Full Changelog: https://github.com/dgunning/edgartools/compare/v5.31.4...v5.31.5
Empty income statement on 16-week-quarter filers (CAVA, RRGB) — quarterly period selection bucketed durations as 80-100 days or 150-285 days, leaving
Empty income statement on 16-week-quarter filers (CAVA, RRGB) — quarterly period selection bucketed durations as 80-100 days or 150-285 days, leaving CAVA's 111-day Q1 in a dead zone. The selector now anchors on filing.period_of_report. (#822, reporter @mkdeak)
TenK.business silently returned Part II MD&A content on GS's 2025 10-K — the cross-Part lookup happily returned a mislabeled part_ii_item_1 key. Item lookup is now constrained to the SEC-canonical Part per item. (#821, reporter @FlorinAndrei)
Viewer ConceptRow.numeric_value returned wrong values on ADI 2019 and ADSK 2019 10-Ks — primary_period is now form-aware (annual forms prefer the longest \"X Months Ended\" duration), and class=\"th\" spacer cells are dropped from body rows so column positions align. (#818, reporter @mpreiss9)
Filing.search() raised a bare AssertionError on pre-2001 SGML/text filings — replaced with a descriptive ValueError pointing users at filing.text(). (#819, reporter @shenker)
`viewer.financial_statements` returned wrong income statement for filings with multi-row period headers (e.g. ADI 2019 10-K mislabeled annual columns
viewer.financial_statements returned wrong income statement for filings with multi-row period headers (e.g. ADI 2019 10-K mislabeled annual columns as quarterly). The R*.htm header parser was rewritten to walk <thead> row by row and filter footnote markers. Affected most 10-K/10-Q filings silently. (#812, reporter @mpreiss9)
Financials.get_net_income() returned wrong value (often wrong sign) for filers reporting a net loss with a separate noncontrolling-interest line — for Micron Q2 2013 returned +$2M (the NCI row) instead of -$286M. Also fixes IFRS 20-F filers whose row label isn't "Net income" (e.g. Barclays "Profit after tax"). Concept lookup is now exact and IFRS-aware. (#814, reporter @wei-jianlin)
`FundReport.options_data()` crashed with `TypeError: bad operand type for abs(): 'NoneType'` on N-PORT filings whose nested forwards had null USD amou
FundReport.options_data() crashed with TypeError: bad operand type for abs(): 'NoneType' on N-PORT filings whose nested forwards had null USD amounts — edgar/funds/reports.py:1011-1012 cast fwd.amount_sold / fwd.amount_purchased through abs() when the corresponding currency_* field equalled 'USD', but valid N-PORT XBRL can pair a stated USD currency with a null amount — every option-on-forward in such a filing tripped the crash before any data was returned. The documented public API was effectively unusable for any fund whose options referenced such a forward (reproducer: GOF NPORT-P). Both assignments now guard on amount_* is not None; the exchange-rate calculation just below was already safe via Python's short-circuiting. Defensive grep across the file confirmed lines 1011-1012 were the only unguarded abs() calls. (#811, reporter @HristoRaykov)
viewer.concept_rows[i].numeric_value silently returned a prior-year value when the primary reporting period had no fact for the row — ConceptRow.numeric_value (and the sibling Concept.value accessor on the concept graph) returned parse_numeric(next(iter(self.values.values()))) — the first entry of the values dict, which was populated only for periods that had a non-empty cell. When the primary (leftmost) reporting period had no value, the singular accessor silently returned whichever period happened to be first in the dict, masking missing-period as a prior-year value. Most visible on the ABT 2019 10-K income statement: concept_rows[16] (us-gaap_IncomeLossFromDiscontinuedOperationsNetOfTax) returned 34.0 (the 2018 value) because ABT had no 2019 discontinued-ops fact. Tracks primary_period on ConceptRow (populated by the R*.htm parser from period_headers[0]) and resolves numeric_value against it explicitly, returning None when the primary period has no value. Concept.value in concept_graph.py got the same fix — same antipattern, same underlying row data, user-visible via the concept graph's Rich/text rendering. (#810, reporter @mpreiss9)
FundFeeNotice crashed with AttributeError: 'list' object has no attribute 'get' on per-class 24F-2NT filings — xmltodict-style parsing returns repeated annualFilingInfo blocks as a list, but every typed accessor (fund_name, series, aggregate_sales, etc.) called .get() on the result. ~2% of recent 24F-2NT filings — including all five BNY Mellon family filings — file one block per share class, so the first call into the data object raised before any data was returned. The data model now iterates every annualFilingInfo block: typed financial properties (aggregate_sales, net_sales, redemptions_current_year, registration_fee, total_due, …) sum across blocks; metadata properties (fund_name, fiscal_year_end, investment_company_act_file_number) read from block[0] (identical across blocks); series deduplicates by seriesId. A new FundClassFee dataclass + is_per_class flag + class_fees list expose the per-share-class breakdown. The _parse_float helper now also handles accounting-parens notation (NNN) → -NNN, which appears in redemptionCreditsAvailableForUseInFutureYears. Backwards-compatible: every existing property keeps the same return shape; the fund total invariant aggregate_sales == sum(cf.aggregate_sales for cf in class_fees) is verified against BNY Mellon Research Growth Fund. (edgartools-8ohs)
viewer.concept_report.currency_scaling returned wrong scales for filers using non-Apple header formats — ConceptReport.currency_scaling was derived from a narrow text match on the R*.htm <th class='tl'> header ($ in millions / $in millions). Filers using In Millions, (in millions), USD ($) in Millions, or Dollars in Millions silently fell through to the default of 1, producing scaling that disagreed across statements within a single filing (ALGN balance sheet vs income statement) and wrong values for whole multi-year ranges (ABNB showing 1 for 2023/2024 when the actual scale is millions). ViewerReport.currency_scaling now derives the scale from the XBRL decimals attribute on monetary facts mapped to the report's role in the presentation linkbase — filer-mandated and uniform (-6 → millions, -3 → thousands, 0 → units). The text-match value is retained as a fallback when XBRL is unavailable. The resolved scale is mirrored back onto ConceptReport.currency_scaling so existing code reading it via the concept-report path also benefits. Same precedent as GH #799 (level enrichment from XBRL). (#807, reporter @mpreiss9)
Schedule 13D/13G silently dropped CUSIPs with the new ` ` wrapper — SEC began wrapping inside an container element on some Schedule 13D/13G filings (e
Schedule 13D/13G silently dropped CUSIPs with the new <issuerCusips> wrapper — SEC began wrapping <issuerCusipNumber> inside an <issuerCusips> container element on some Schedule 13D/13G filings (e.g. CIK 1906837 13D, CIK 1425851 13G). The parser's BS4 recursive=False lookup at the top-level only matched the flat layout, so subject_company.cusip came back as '' whenever the wrapper was present. Parsing now falls back to a recursive lookup when the flat probe misses, handling both wire formats. (#802, PR #803 by @HristoRaykov)
Schedule 13D/13G event-date attribute name mismatch — Schedule13D exposed the triggering-event date as date_of_event while Schedule13G exposed it as event_date, breaking duck-typing across a mixed list of 13D/13G filings and forcing callers to use getattr / hasattr. Both classes now accept either name; the underlying attribute is unchanged, so existing code keeps working. (#804, PR #805 by @0ywfe)
Spurious DocumentTooLargeError from StreamingParser on legitimate documents — The streaming HTML parser accumulated len(etree.tostring(elem)) on every lxml iterparse end event. Because tostring serializes the full subtree and end fires for every closing tag, nested elements were counted multiple times — large nested HTML could trip max_document_size even though the source document was under the limit. The per-event accumulator is also redundant: HTMLParser._parse already validates len(html.encode("utf-8")) against max_document_size before invoking streaming mode. The accumulator and its state are removed; size is now checked once at the top of StreamingParser.parse() and the same encoded bytes are reused for iterparse. (#806 by @kevinchiu)
Full Changelog: https://github.com/dgunning/edgartools/compare/v5.31.0...v5.31.1
`include_quarterly` parameter on stitched XBRLS statements — XBRLS.from_filings() previously emitted a single column per filing, preferring YTD/annual
include_quarterly parameter on stitched XBRLS statements — XBRLS.from_filings() previously emitted a single column per filing, preferring YTD/annual over the discrete-quarter period when both existed in the source XBRL (Issue #475 design). This created a parity gap with single-filing XBRL, which surfaces both. The new opt-in include_quarterly=False parameter on XBRLS.get_statement(), StitchedStatement, and statements.income_statement() / cashflow_statement() causes each 10-Q to contribute both a 90-day discrete column and the YTD column, and each 10-K to contribute both an annual column and its embedded Q4 column. Distinct from discrete_quarters (v5.30.3) which derives quarterly cash-flow values by subtraction; this surfaces facts already in the filing. Default behavior is preserved. Has no effect on Balance Sheet (instant periods only). (#780, reporter @AhmedShaker12)viewer.financial_statements silently dropped income statements miscategorized in FilingSummary.xml — AbbVie's 2021 10-K placed Consolidated Statements of Earnings under MenuCategory='Uncategorized' instead of 'Statements' — a filer mistake that EdgarTools faithfully reflected, so the income statement disappeared from viewer.financial_statements while comparable 2019/2020/2022-2025 filings worked fine. The viewer now returns the union of FilingSummary MenuCategory='Statements' and MetaLinks groupType='statement', deduplicated by HTML filename, in filing-position order. MetaLinks reflects XBRL taxonomy classification and is more reliable than filer-provided menu metadata. (#797, reporter @mpreiss9)
viewer.concept_rows[*].level always returned 0 — Modern SEC R*.htm files don't encode hierarchy in the rendered HTML — empirically verified across 10 diverse 2025 10-Ks (AAPL, ABT, JPM, WMT, XOM, VZ, MSFT, GS, PFE, BRK.B): zero `plN` class tokens on primary statements, almost no `padding-left` styles, no row nesting. The canonical source is the XBRL presentation linkbase, which the existing parser already loads as `xbrl.presentation_trees[role].all_nodes[concept_id].depth`. The viewer now lazy-loads the parsed XBRL on first `concept_rows` access and populates `ConceptRow.level` from the presentation tree, normalized so the smallest depth observed in a report becomes 0. For the issue's canary case (ABT balance sheet) the level distribution went from `{0: 45}` to `{0: 15, 1: 26, 2: 4}`. (#799, reporter @mpreiss9, investigation by @tjhub1983)
XBRLS.from_filings(list, filter_amendments=True) crashed with AttributeError — The signature accepts `Union[Filings, List[Filing]]` and defaults `filter_amendments=True`, but the implementation called `filings.filter()` unconditionally — raising `AttributeError: 'list' object has no attribute 'filter'` whenever a plain list was passed. The implementation now branches on whether the input has a `.filter` method; for plain lists it falls back to a form-suffix check that drops forms ending in `/A`. (edgartools-6k96)
Full Changelog: https://github.com/dgunning/edgartools/compare/v5.30.3...v5.31.0
`facts.time_series()` returned duplicate rows from fuzzy concept matching — When called with a fully-qualified XBRL concept like us-gaap:NetIncomeLoss
facts.time_series() returned duplicate rows from fuzzy concept matching — When called with a fully-qualified XBRL concept like us-gaap:NetIncomeLoss, the underlying by_concept query defaulted to fuzzy substring matching, so us-gaap:NetIncomeLossAvailableToCommonStockholdersBasic was silently included alongside it, producing duplicate rows for the same reporting period. time_series() now passes exact=':' in concept, so qualified names match exactly while bare names ('Revenue') retain fuzzy/label discovery. (#795, PR #798 by @tjhub1983)
Quarterly Q4 NetIncomeLoss off by ~1000× when proxy XBRL contained corrupt metadata — DX's 2026 DEF 14A disclosed historical NetIncomeLoss figures with fiscal_year=null, fiscal_period=null, including a single FY 2025 value with a 1000× scaling error (319,065 instead of 319,066,000). The corrupt fact entered the ANNUAL duration bucket in TTMCalculator._derive_q4_from_fy and won the period_end dedup because the proxy was filed after the 10-K, producing a Q4 2025 value of −133,387,935 instead of +185,359,000. The TTM calculator now (a) requires valid fiscal_period (FY/Q3/Q2) for inputs to each derivation method, and (b) prefers periodic-report sources (10-K/Q, 20-F, 40-F, 6-K and amendments) over proxy/registration forms when deduplicating. (#796)
DFIN-generated TOCs produced unprefixed item keys — TOCs from DFIN's filing tool (e.g., Microsoft 10-K) place PART I/PART II headers in text-only <tr> rows without anchor links. The previous parser iterated flat over <a> links and only updated current_part from a parsed link's text, so it never saw the part headers and produced keys like "Item 1" instead of "part_i_item_1". The parser now walks rows in document order so text-only rows update part context for item links that follow.
Non-standard duration stubs distorted "Three Recent Periods" view — PLTR's latest 10-Q exposes a 30-day stub context (duration_2026-03-01_2026-03-31, classified as 'Period') alongside the normal Q1 and FY durations. get_period_views() sorted all duration periods by end date and took the top 3, so the stub landed at period_keys[0] with no statement data and produced an all-null column. Period view generation now filters to durations whose classify_duration() bucket is a standard reporting period (Quarterly, Semi-Annual, Nine Months, Annual).
fix: prevent duplicate revenue rows from additional income statement by @ghedo44 in https://github.com/dgunning/edgartools/pull/790
Full Changelog: https://github.com/dgunning/edgartools/compare/v5.30.1...v5.30.2
get_filings(filing_date=(start, end)) crashed with TypeError — Entity.get_filings declared filing_date: Optional[Union[str, Tuple[str, str]]] but the underlying parser only handled the colon-separated string form. The tuple form crashed every CIK with TypeError: strptime() argument 1 must be str, not tuple before any HTTP request. extract_dates now accepts both (start, end) tuples and lists, with None in either slot meaning "open" (matching the existing "start:" / ":end" string-form semantics). (#794)
Duplicate revenue rows when RevenuesAbstract is an additional virtual-tree root — In ~35% of companies whose learned virtual trees contain RevenuesAbstract as an additional root alongside IncomeStatementAbstract, the rendered income statement showed Revenue twice — once promoted under IncomeStatementAbstract, once again as a child of the second root. The duplicate-root guard previously checked only top-level concepts; it now walks the tree recursively via _collect_concepts and prunes duplicate subtrees, preserving abstract containers only when they still hold unique descendants. (#789, PR #790 by @ghedo44)
Orphan section re-introduced Revenue under "Additional Financial Items" — As a follow-up to the #789 fix, the orphan dedup at the income-statement assembly layer was matching by display label only (existing_labels). When the canonical promotion produced an item with label "Total Revenue" while the orphan candidate fact carried the raw label "Revenue", the dedup missed the match and re-added Revenue under AdditionalItems. _collect_labels now tracks both labels and concepts so the orphan check (label or concept) in existing_labels matches by either form.
TTM income statement values labeled with wrong fiscal year for interim quarters — When SEC re-filed comparative facts in next year's 10-Q (e.g., AGNC'
TTM income statement values labeled with wrong fiscal year for interim quarters — When SEC re-filed comparative facts in next year's 10-Q (e.g., AGNC's Q1 2024 fact re-tagged with fiscal_year=2025 in a 2025 10-Q), _deduplicate_by_period_end kept the latest filing's version, and the TTM trend builder labeled the window with that comparative-shifted fiscal year. The result was duplicate column labels ("Q3 2025" appearing twice) that collided in the rendering layer's dict-keyed mapping, causing Company('AGNC').income_statement(periods=12, period='ttm') to display Q3 2024's TTM value under the "Q3 2025" column. The TTM calculator now derives the label fiscal year from period_end + FYE instead of the (potentially comparative-tagged) as_of_fact.fiscal_year. (#793)
Quarterly facts dropped for non-calendar FYE companies — Fixed regression introduced in 5.30.0 where the schedule-fact filter from #781 incorrectly rejected Q1/Q2/Q3 facts for companies with non-calendar fiscal year ends (ADSK, WMT, NVDA, CSCO, MSFT). Company('ADSK').income_statement(periods=4, annual=False) returned only Q4 across years instead of Q1–Q4 of the most recent fiscal year. The fiscal-year/period-end validator is now FYE-aware. (#779)
facts.time_series() returned indistinguishable rows for overlapping periods — When a company reported the same concept in both quarterly and YTD form (e.g., AGNC's NetIncomeLoss for period_end=2025-06-30 had a 3-month Q2 row and a 6-month H1 YTD row), time_series() returned both with identical period_end / fiscal_period / fiscal_year, leaving users no way to tell them apart. Output now includes period_start and a derived duration_days column. (#792)
download_submissions not importable from edgar.storage — Error messages in edgar/reference/company_dataset.py instructed users to run from edgar.storage import download_submissions, but the function was defined in edgar/storage/_local.py without being added to that module's __all__, so the star-import in edgar/storage/__init__.py did not re-export it. The advertised import path now works. (#791)
search_filings() — search_filings() now accepts an items parameter that is forwarded server-side to EFTS, enabling structured Item-based queries without falling back to client-side filtering (which previously lost the long tail to pagination caps). The query parameter is now optional when items is provided, supporting pure-structured lookups such as search_filings(forms="8-K", items="1.05", start_date="2023-12-01", end_date="2024-12-31") for cybersecurity disclosures.GrepResult repr/str unified via rich panel — GrepResult.__repr__ now renders the same Rich Panel as __repr_html__, replacing the old compact "GrepResult('pattern', N matches)" summary. __str__ has been removed; calling str(result) falls back to __repr__. Callers that want the prior plain-text dump should call result.to_context() explicitly.New ProxySeason and ProxyContest classes for grouping proxy filings by season and detecting contested director elections. Market-wide discovery is ava
New ProxySeason and ProxyContest classes for grouping proxy filings by season and detecting contested director elections. Market-wide discovery is available via proxy_contests(). (#773)
Structured extraction from DEF 14A proxy statements now covers Summary Compensation Tables (executive pay), CEO pay ratio with footnote cross-validation, voting proposals, beneficial ownership tables, director compensation tables, and audit fees by category.
EFTS search is enriched with relevance scores, aggregations, filtering, and pagination. A new .grep() method provides universal content search across filings.
viewer.search() — Restored section separator newlines inside Concept panels that were incorrectly removed in v5.29.0. (#776)Proxy season analysis — New ProxySeason and ProxyContest classes for grouping proxy filings by season and detecting contested elections. Market-wide discovery via proxy_contests() (#773)
Proxy HTML data extractors — Extract structured data from DEF 14A proxy statements:
Full-text search enhancements — Enriched EFTS search with relevance scores, aggregations, filtering, and pagination. New .grep() method for universal content search across filings
Fiscal year labels for non-calendar FYE companies — Statement period labels for companies with early fiscal year ends (Jan–Mar) now use the industry-standard convention, matching the SEC, Bloomberg, and company earnings releases. NVIDIA Q3 ending Oct 2025 is now labeled "Q3 2026" (FY2026), not "Q3 2025" (#779)
Empty statements from forward-looking schedule data — Companies like CLSK with XBRL-tagged footnote disclosures (expected amortization schedules) no longer produce phantom future periods that displace real quarterly data (#781)
Missing XBRL instance from SEC — Fetch XBRL instance directly from SEC when local feed file lacks it (#778)
XBRL parsing for bytes content — Hardened XBRL parser to handle bytes content and missing entity info without errors
Concept panel display in viewer.search() — Restored section separator newlines inside Concept panels that were incorrectly removed in v5.29.0 (#776)
`exact` parameter for `FactQuery.by_date_range()` — New exact=True option matches facts with period dates exactly equal to the specified date, instead
exact parameter for FactQuery.by_date_range() — New exact=True option matches facts with period dates exactly equal to the specified date, instead of the default <=/>= range behavior (#767)Company.reit_subtype property — Distinguishes equity REITs from mortgage REITs by checking for mortgage-related XBRL conceptsviewer.search() output (#768)business_category misclassifications across 4 patterns (#774)fiscal_period classification (#771)gaap_mappings defaulting to section totalsperiod_of_report triggering network calls for local storage usersFull Changelog: https://github.com/dgunning/edgartools/blob/main/CHANGELOG.md
exact parameter for FactQuery.by_date_range() — New exact=True option matches facts with period dates exactly equal to the specified date, instead of the default <=/>= range behavior (#767)
Company.reit_subtype property — New property distinguishes equity REITs from mortgage REITs by checking for mortgage-related XBRL concepts in the company's filings
Filing agent fingerprinting — Detect the filing agent (Donnelley, EDGAR Online, Workiva, Toppan Merrill) from HTML structure patterns, enabling agent-aware document parsing
Agent-aware TOC parsing — Table of contents section detection now uses agent-specific parsing strategies for the top 4 filing agents, improving section extraction accuracy
TOC section detection evaluation suite — Evaluation harness for measuring TOC section detection quality across a corpus of filings
Extra newlines in viewer.search() output — Removed spurious blank lines between sections in Concept panel display (#768)
business_category misclassifications — Corrected 4 classification patterns for more accurate company categorization (#774)
YTD periods missing fiscal_period classification — Year-to-date periods in XBRL facts now receive proper fiscal period labels (#771)
61 cash flow gaap_mappings defaulting to section totals — Corrected mappings that incorrectly pointed to section-level totals instead of specific line items
Duplicate facts in XBRL DataFrame — Deduplicate identical facts in facts.to_dataframe() output (#769)
period_of_report triggering network calls — Resolved unintended network requests when accessing period_of_report for local storage users
Fix HTML in `to_dataframe()` for disclosure TextBlock concepts (#762) — Disclosure/notes statements (e.g., segment tables) contained raw HTML markup i
to_dataframe() for disclosure TextBlock concepts (#762) — Disclosure/notes statements (e.g., segment tables) contained raw HTML markup in DataFrame cells. Now sanitized to plain text.DividendsEquity standard concept for equity statement dividends (#763) — GOOGL's dividend concept (AdjustmentsToAdditionalPaidInCapitalDividendsInExcessOfRetainedEarnings) had no standard_concept mapping on the equity statement. Added DividendsEquity to the equity vocabulary.Thanks to @BaraVaq for the detailed bug reports.
HTML markup in disclosure DataFrame output — to_dataframe() now strips HTML from XBRL TextBlock facts in disclosure/notes statements, producing clean plain text instead of raw markup. Uses the existing _is_html/html_to_text utilities. Includes regression test (#762)
Missing DividendsEquity standard concept for equity statement — Added DividendsEquity to the equity vocabulary (gaap_mappings.json, section_membership.json, display_names.json), fixing GOOGL's AdjustmentsToAdditionalPaidInCapitalDividendsInExcessOfRetainedEarnings being unmapped on the equity statement (#763)
Entity rich display alignment — Entity rich display now follows the same design language as Company, ensuring consistent visual presentation
fix: remove incorrect StockRepurchasesEquity mapping for tax withholding concept (#760) — Removed a wrong gaap_mappings.json entry that mapped Adjustm
fix: remove incorrect StockRepurchasesEquity mapping for tax withholding concept (#760) — Removed a wrong gaap_mappings.json entry that mapped AdjustmentsRelatedToTaxWithholdingForShareBasedCompensation to StockRepurchasesEquity (confidence 0.364). Tax withholding on RSU vests is not a stock repurchase — the misclassification caused duplicate rows in the statement of equity for companies like AAPL.
fix: add Q/YTD/FY period labels to equity and comprehensive income statements (#759) — Added StatementOfEquity and ComprehensiveIncome to the allow-list in rendering.py, so column headers now show friendly period labels (e.g., "Q1", "YTD") instead of raw date ranges, consistent with other financial statements.
Q/YTD/FY period labels missing from equity and comprehensive income — Equity and comprehensive income statements now receive the same Q1/Q2/Q3/Q4/YTD/FY column labels applied to income and cash flow statements (#759)
Incorrect StockRepurchasesEquity mapping — Removed erroneous StockRepurchasesEquity standard concept mapping for tax withholding on vested shares, which caused misclassification on equity statements (#760)
Schedule 13D/G total_shares and total_percent overcounting — Changed aggregation from sum() to max() to correctly represent reported totals rather than double-counting across rows
13F-HR TXT parser for pre-2013 filings — Rewrote the 13F-HR TXT parser to use column-position extraction, added regex fallback and decimal handling, achieving ~93% coverage of pre-2013 filings (#476)
Standard concept name misspellings — Corrected misspellings in standard concept names (#758)
Wrong quarter labels for non-calendar fiscal years — Quarter labels in financial statement columns now use the company's fiscal year end month instead
Wrong quarter labels for non-calendar fiscal years — Quarter labels in financial statement columns now use the company's fiscal year end month instead of hardcoded calendar months. Affects companies like AAPL (Sep FY), WMT (Jan FY), NKE (May FY) (#752)
Period-type suffixes always present on DataFrame columns — to_dataframe() now always includes period-type suffixes (Q1/Q2/Q3/Q4/YTD/FY) on all duration columns, not just when end dates collide (#753)
Incorrect Q4 fiscal year label for Jan-Mar FYE companies — Companies with fiscal years ending in January through March (e.g., NVDA, WMT, HD, CRM) now receive the correct Q4 label rather than a label belonging to the following calendar year (#754)
Capex extraction broken by label regex — Capital expenditure extraction now uses XBRL concept names instead of fragile label regex matching, fixing NVDA capex from $101M (wrong) to $6,042M (correct) (#756)
Wrong quarter labels for non-calendar fiscal years — Quarter labels in financial statement columns now use the company's fiscal year end month instead of hardcoded calendar months. Affects companies like AAPL (Sep FY), WMT (Jan FY), NKE (May FY) (#752)
Period-type suffixes always present on DataFrame columns — to_dataframe() now always includes period-type suffixes (Q1/Q2/Q3/Q4/YTD/FY) on all duration columns, not just when end dates collide (#753)
Incorrect Q4 fiscal year label for Jan-Mar FYE companies — Companies with fiscal years ending in January through March (e.g., WMT) now receive the correct Q4/FY label rather than a label belonging to the following calendar year (#754)
Capex extraction broken by label regex — Capital expenditure extraction now uses XBRL concept names (PaymentsToAcquirePropertyPlantAndEquipment, etc.) instead of fragile label regex matching, making it robust across filings with varied label text (#756)
business_category misclassifications — Fix ETFs, SPACs, commodity trusts, and BDCs being misclassified
business_category misclassifications — Fix ETFs, SPACs, commodity trusts, and BDCs being misclassified. Adds SPAC name pattern detection, "ETF" name check for crypto/commodity ETFs, SIC 6200s fund/trust heuristic, removes over-broad "CAPITAL CORP" BDC name pattern, and uses authoritative 814- file number for BDC detection (#561)
to_dataframe() missing columns — to_dataframe() now includes both quarterly and YTD columns when a filing contains both, instead of silently dropping one (#743)
13F values not normalized — Normalize 13F holdings values to dollars across all periods (#749)
obj() routing for Schedule 13D/G — obj() now correctly routes SC 13D/G forms to Schedule13D/13G parsers (#748)
find_ticker() wrong result — Fix wrong company ticker returned for CIK 1506307 (#745)
download_filings in Jupyter — Support download_filings in Jupyter notebook environments (#744)
reverse_name — Replace with improved implementation for more accurate name reversal
Punctuation normalization — Fix handling of digits and percent signs in text extraction
TOC section detection for split-link filings — Filings where TOC item labels and descriptive titles link to different anchors (e.g., TSLA 10-K) now va
TOC section detection for split-link filings — Filings where TOC item labels and descriptive titles link to different anchors (e.g., TSLA 10-K) now validate anchor targets against expected section headings, picking the correct anchor (#742)
Non-accrual extraction false positives — Footnotes that explicitly deny non-accrual status (e.g., "there were no investments on non-accrual status") are no longer treated as positive matches. Replaced naive substring matching with two-stage negation-then-affirmation classification
Non-accrual period resolution — extract_nonaccrual() now uses filing.period_of_report as anchor for period selection instead of picking the max instant date, which could resolve to filing dates or DEI dates instead of balance sheet dates
Full Changelog: https://github.com/dgunning/edgartools/compare/v5.28.0...v5.28.1
TOC section detection for split-link filings — Filings where TOC item labels and descriptive titles link to different anchors (e.g., TSLA 10-K) now validate anchor targets against expected section headings, picking the correct anchor (#742)
Non-accrual extraction false positives — Footnotes that explicitly deny non-accrual status (e.g., "there were no investments on non-accrual status") are no longer treated as positive matches. Replaced naive substring matching with two-stage negation-then-affirmation classification. Scored 50/50 on synthetic variations
Non-accrual period resolution — extract_nonaccrual() now uses filing.period_of_report as anchor for period selection instead of picking the max instant date, which could resolve to filing dates or DEI dates instead of balance sheet dates. ARCC now correctly resolves to 2025-12-31
New FilingViewer class provides access to the SEC's XBRL interactive data viewer for any filing — structured period headers, numeric values, scaling i
New FilingViewer class provides access to the SEC's XBRL interactive data viewer for any filing — structured period headers, numeric values, scaling information, and concept-level metadata via MetaLinks.json parsing.
New ConceptGraph builds a traversable graph of XBRL concepts and their relationships, enabling structured navigation across the taxonomy hierarchy.
New extract_nonaccrual() extracts non-accrual investment data from BDC filings using three layered strategies: XBRL footnotes (investment-level detail), custom XBRL concepts, and standard us-gaap fallback.
to_markdown() on Statement, Note, Notes, RenderedStatement, and StatementLineItem — GFM tables optimized for LLM context windowsto_context() expanded with detail parameter (minimal/standard/full) across all financial objectscompare_context() for LLM-based cross-validation of parsed values against SEC viewer outputU.S., D.C., e.g. no longer get split into U. S., D. C. in iXBRL filings (AAPL 10-K had 79 occurrences){self.tag} in plain-text output (#740)to_context(), to_markdown(), and compare_context() coverageFull Changelog: https://github.com/dgunning/edgartools/compare/v5.27.0...v5.28.0
FilingViewer — SEC Interactive Data Viewer — New FilingViewer class provides access to the SEC's interactive XBRL viewer for any filing. Parses MetaLinks.json for concept-level metadata, extracts R*.htm viewer reports, and exposes structured period headers, numeric values, and scaling information
ConceptGraph — navigable XBRL knowledge graph — New ConceptGraph class builds a traversable graph of XBRL concepts and their relationships, enabling structured navigation across the taxonomy hierarchy
BDC non-accrual extraction — New extract_nonaccrual() function in edgar.bdc.nonaccrual extracts non-accrual investment data from BDC XBRL filings using three layered strategies: XBRL footnotes (investment-level detail), custom XBRL concepts (rate only), and standard us-gaap aggregate fallback
to_markdown() for LLM drill-down — Notes, disclosures, and financial drill-down objects now expose to_markdown() for LLM-optimized output (#732)
compare_context() for LLM-based validation — New method on XBRL objects for cross-validating parsed values against SEC viewer output using an LLM judge
Cross-validation bridge between SEC Viewer and XBRL parser — FilingViewer and the XBRL parser can now be reconciled programmatically, with to_dataframe() and diagnostic outputs for systematic validation
MetaLinks.json parser — Full parser for the SEC XBRL viewer's MetaLinks.json metadata file, exposing concept-level role, label, and calculation arc data
Abbreviations and inline spacing preserved in iXBRL text extraction — Text extraction from iXBRL documents no longer splits abbreviations like U.S. into U. S. or D.C. into D. C.. Affects all inline XBRL filings (#734)
TOC part metadata parsing — Table of contents part metadata is now correctly extracted (#737) — contributed by external PR
Ruff code quality: 533 issues resolved — Full codebase pass fixing lint, f-string, and style issues including a LinkBlock.get_text() f-string bug (#740)
FilingViewer and ConceptGraphto_context() and to_markdown() coverageSixK — Form 6-K (Report of Foreign Private Issuer) with cover page metadata, exhibit access, and press release filtering
content_type property (earnings, cybersecurity, restructuring, etc.), is_amendment, get_exhibit()/get_exhibits(), and context-aware to_context()html() for XML-primary S-1/S-3 filings fixedpip install edgartools==5.27.0
Dedicated 6-K data object — New SixK class replaces the CurrentReport alias for Form 6-K (Report of Foreign Private Issuer). Extracts cover page metadata (commission file number, report month, annual report form, content description), provides exhibit access, press release filtering, and IFRS financials when present. Includes to_context() with cover page text at full detail level
S-1/F-1 registration statement data object — New RegistrationS1 class for S-1 and F-1 registration statements with cover page extraction, prospectus section access, and amendment support
DRS draft registration statement data object — New DraftRegistrationStatement class for confidential draft registration statements (DRS/DRS-A)
Generic XML filing data object — New XmlFiling class for XML+XSLT SEC forms (X-17A-5, TA-1, TA-2, SBSE, ATS-N-C, CFPORTAL, etc.) with automatic XSLT rendering
24F-2NT fund fee notice data object — New FundFeeNotice class for annual notices of securities sold by registered investment companies
497K fund summary prospectus data object — New Prospectus497K class for 497K fund summary prospectus filings
F-1/F-1A foreign registration support — RegistrationS1 now accepts F-1 and F-1/A forms for foreign private issuer IPO registrations
F-3 foreign shelf registration support — RegistrationS3 now accepts F-3, F-3/A, and F-3ASR forms
EightK improvements — New content_type property classifying 8-K filings (earnings, cybersecurity, restructuring, etc.), is_amendment property, get_exhibit() and get_exhibits() methods, and context-aware to_context() that adjusts available actions based on content type
8-K section boundary captures full body text — HTMLParser section detection now correctly extends section boundaries past table-wrapped item headings to include all body paragraphs until the next section (#733)
gaap_mappings: PaymentsToDevelopSoftware and PaymentsForSoftware — Both were incorrectly mapped to NetCashFromInvestingActivities (section total) instead of PurchaseOfIntangibleAssets (component line item) (#739)
Infinite recursion in html() for XML-primary filings — html() no longer recurses when the primary document of S-1/S-3 filings is XML rather than HTML
MunicipalAdvisorForm assert narrowed — Assert restricted to MA-I only; MA form now routes to XmlFiling
MCP outputSchema validation error — Claude Desktop rejected every MCP tool call with "outputSchema defined but no structured output returned." Removed
outputSchema from tool definitions, restoring full MCP functionality (#735)edgar_notes next-steps reference — referenced a non-existent tool name; corrected to valid tooledgar_screen state filter silently dropped — state filter was discarded on exchange-only queries, returning unfiltered resultsedgar_compare growth metrics broken — growth calculation failed due to insufficient time_series fetchFull Changelog: https://github.com/dgunning/edgartools/compare/v5.26.0...v5.26.1
MCP tool definitions: outputSchema removed — outputSchema was included in all MCP tool definitions, which is not part of the MCP protocol spec. Claude Desktop rejected every tool call, blocking all MCP usage entirely. Removing the field restores full MCP functionality (#735)
edgar_notes next-steps reference — edgar_notes referenced a non-existent tool name in its next_steps guidance; corrected to a valid tool
edgar_screen state filter silently dropped — State filter was silently discarded on queries that specified only an exchange (no SIC code), causing state-filtered screening to return unfiltered results
edgar_compare growth metrics broken — Growth metric calculation failed because time_series fetched insufficient periods; fetch count increased to ensure enough data points are available
ai-integration.md split into five focused pages (ai/index.md, ai/mcp-setup.md, ai/mcp-tools.md, ai/mcp-workflows.md, ai/skills.md) for easier navigation. Parameter defaults and required-field annotations corrected across all pagesS-3 shelf registration data object — New RegistrationS3 class with fee table extraction, shelf lifecycle tracking, prospectus section access, and auto
S-3 shelf registration data object — New RegistrationS3 class with fee table extraction, shelf lifecycle tracking, prospectus section access, and auto-shelf detection for well-known seasoned issuers (#728)
CORRESP/UPLOAD correspondence support — New Correspondence and CorrespondenceThread classes parse SEC correspondence filings with automatic classification and metadata extraction. Filing.correspondence() works on any filing type to find related SEC review threads
Point-in-Time mode for EntityFacts — EntityFacts.to_dataframe() now accepts a pit_mode parameter for lookahead-bias-free backtesting (#697)
TTM unification on EntityFacts — Unified TTM access with caching and fiscal year quarter labels (PR #721, @ghedo44)
TypeError in _get_statement_concepts when stmt type is NoneCORRESP/UPLOAD correspondence support — New Correspondence and CorrespondenceThread classes parse SEC correspondence filings with automatic classification (company_response, acceleration_request, sec_comment, review_complete, no_review) and metadata extraction (file number, referenced form, fiscal year). Filing.correspondence() works on any filing type to find related SEC review threads via file number
Point-in-Time mode for EntityFacts — EntityFacts.to_dataframe() now accepts a pit_mode parameter that includes filing_date and form_type columns, enabling lookahead-bias-free backtesting by filtering on filing_date <= as_of_date (#697)
S-3 shelf registration data object — New RegistrationS3 class with fee table extraction from EX-FILING FEES exhibits (Exhibit 107) supporting 5 HTML format variations, ShelfLifecycle with shelf capacity and offering capacity properties, prospectus section access with 16 section patterns, and auto-shelf detection for well-known seasoned issuers (#728)
TTM unification on EntityFacts — Unified TTM access on EntityFacts with streamlined Company delegation. TTM-ready facts are cached for performance. Quarter labels now use fiscal year (PR #721, ghedo44)
TOC named-anchor targets — Table-of-contents anchor matching centralized and now correctly resolves named-anchor targets (#727)
Revenue in income statement dedup — Revenue now included in the promoted income statement deduplication set
Shares concepts preserved in statements — Shares-denominated concepts (EPS, shares outstanding) are no longer dropped from income statements during unit filtering (PR #725, ghedo44)
TypeError in _get_statement_concepts — Fixed crash when statement type is None by using or '' fallback instead of relying on dict.get() default
Unit filter documentation — Docstrings updated to reflect native-unit filtering behavior
None, period index key caching added. Measured on AAPL: 20.5 MB → 15.0 MBBDC health metrics — PortfolioInvestments now exposes nonaccrual_fair_value, non_accrual_rate, pik_investments, pik_fair_value, and pik_exposure prope
PortfolioInvestments now exposes nonaccrual_fair_value, non_accrual_rate, pik_investments, pik_fair_value, and pik_exposure properties. Non-accrual data is extracted from the entity-level XBRL concept us-gaap:FairValueOptionLoansHeldAsAssetsAggregateAmountInNonaccrualStatus. Rich display shows color-coded non-accrual and PIK summary linesweakref with strong references in Note, StatementLineItem, FilingSummary, and WeakCache. Weak references caused pickle.dumps() to fail on these objects, breaking caching and multiprocessing workflowsFull Changelog: https://github.com/dgunning/edgartools/compare/v5.25.0...v5.25.1
Statement-to-note drill-down — Navigate from any financial statement line item to the note that explains it. balance_sheet['Cash and cash equivalents'
Statement-to-note drill-down — Navigate from any financial statement line item to the note that explains it. balance_sheet['Cash and cash equivalents'].note returns the related Note object via a lazy-built reverse index that maps XBRL concepts to notes — the same mechanism the SEC's own EDGAR viewer uses
Note and Notes classes — First-class objects for financial statement notes, built from FilingSummary.xml hierarchy. Access via tenk.notes or tenq.notes. Browse by number (notes[5]), title (notes['Debt']), or fuzzy search (notes.search('revenue')). Each note exposes .tables, .policies, .details, .text, .html, .expands (which statement lines it explains), and .to_context() for AI consumption
StatementLineItem — Lightweight wrapper returned by Statement.__getitem__ with .label, .concept, .note (most relevant note), .notes (all related), and .values. Uses __slots__ for minimal memory footprint
Statement.search() — Fuzzy search for statement line items with ranked results (exact > startswith > word match > substring). Complements the exact-match __getitem__. Consistent with the Notes.search() pattern
Statement.report property — Links to the FilingSummary Report for HTML table access. Enables note.tables[0].report.to_dataframe() for HTML-extracted DataFrames alongside the XBRL path
RenderedStatement.__getitem__ — Look up rows by exact label (case-insensitive) on rendered statements
edgar_notes MCP tool — New tool for AI agents to drill into notes and disclosures by company and topic. Returns structured note content, related statement lines, and child table data. Surfaces the detail behind financial statement numbers that no other SEC MCP server exposes
CompanyReport.notes — Cached property on TenK/TenQ providing hierarchical notes access from report objects
TenK.to_context(focus=...) / TenQ.to_context(focus=...) — Focus mode generates cross-cutting context for specific topics (e.g., focus='debt'), pulling statement line items, note content, and policies together
Role type definitions from schema — XBRL parser now extracts human-readable role definitions from taxonomy schemas, improving statement and note titles
XBRL memory optimizations — Label role URI strings are now interned via sys.intern(), eliminating ~10,000 duplicate URL string allocations per filing. comparison_data removed from RenderedStatement.metadata (was stored but never read back). Duplicate _collect_note_concepts tree walks eliminated in expands_statements
Statement.__getitem__ is now exact-match only — Previously used substring fallback that could silently return wrong rows for ambiguous queries like stmt['Total']. Now returns the correct match or None. Use stmt.search() for fuzzy lookups
stmt['Debt'].note before tenk.notes produced empty results because notes were built without FilingSummary. Now the XBRL object stores its FilingSummary during from_filing() so the lazy notes builder always gets the full hierarchyEvery filing.obj() class now has a to_context(detail='minimal'|'standard'|'full') method that produces LLM-optimized plain-text context. This is the f
to_context() on all filing typesEvery filing.obj() class now has a to_context(detail='minimal'|'standard'|'full') method that produces LLM-optimized plain-text context. This is the foundation for agentic SEC analysis workflows.
compose_context() — combine multiple objects into a single token-budgeted string with automatic detail downgradingHasContext Protocol — structural typing for any object with to_context()format_currency_short() — compact currency formatting ($394.3B, $18.6M) for context outputFinancials.get_currency_symbol() — detects reporting currency from XBRL unitsedgar_context → edgar_filing (primary filing examination), edgar_filing → edgar_read (section reader)edgar_filing now supports two paths: company+form (identifier="AAPL", form="10-K") or accession number/URLEarningsRelease.get_key_metrics(quarterly=True) — extracts EPS, revenue, net income using concept-level identification (unit='USD/shares' for EPS) with quarterly period preferencerow_types list fixes duplicate label collisiontext_content() for raw etree elementsFull Changelog: https://github.com/dgunning/edgartools/compare/v5.23.4...v5.24.0
Your coding agent can read these notes before it upgrades. Set up the MCP server →