Skip to content

About

Rebuilds PDFs as HTML with real headings, lists and tables — readable with a screen reader. Nothing leaves your browser. Chrome and Firefox, Manifest V3.

Resources

Stars

5 stars

Watchers

0 watching

Forks

Repository files navigation

Accessible PDF View

Read PDFs as accessible, structured HTML.

Official website: https://accessiblepdfview.org


What is Accessible PDF View?

Accessible PDF View is a browser extension that takes a PDF you are already looking at, analyses the document information it can recover, and rebuilds it as an ordinary semantic HTML document.

PDF
 └─ Accessible PDF View
     └─ Accessible Reader
         └─ semantic HTML
             └─ browser accessibility tree
                 └─ screen reader

It is a PDF-only tool. There are no plans to support other document formats.

Why, for screen reader users

A PDF is a description of marks on a page. Unless its author tagged it — and most authors do not — there is no heading structure, no list structure, no table structure and no reading order for assistive technology to work with. The text is there; the document is not.

This extension rebuilds the document as HTML, so that the things that work on every other web page work here too:

  • heading navigation (H, 1–6)
  • list navigation
  • table navigation, with real column and row headers
  • link navigation and the links list
  • the browser's own find

It uses native HTML elements wherever HTML has one, and reaches for ARIA only where HTML genuinely cannot express something.

Additional behaviour worth knowing:

  • Focus is never moved for you. When analysis finishes, the completion is announced through a live region and your focus stays where you put it.
  • Work that takes minutes says so while it runs. Reading the PDF, OCR on a page, a figure being described and the AI model downloading each report themselves through a role="status" region that is in the page from the start — a live region created at the same moment as its first message is frequently announced by nothing. The model download reports in ten steps rather than every percentage it is given, because a hundred announcements queued behind each other is not progress reporting.
  • A button that starts long work stays where it is. It is marked aria-disabled rather than disabled while the run goes: a disabled element cannot hold focus, so pressing one would drop you at the top of the document at the exact moment you asked for something.
  • The tab title says what the tab is doing — report.pdf (Analysing…), then the view it was left in. A tab working on its own is otherwise indistinguishable from one that has finished.
  • Every page state is stated in words. A page with no readable text says so; it is never reported as an error.
  • Figures are announced as figures, and the picture itself is shown. When the PDF supplies no alternative text, the Reader says exactly that.
  • A description is never presented as the author's words. Text the document supplied is read as "Image: …"; a description generated on your device is read as "Image (description generated by AI): …" — and in Japanese as 「画像: …」 and 「画像(AIによる自動生成の説明): …」. See Image descriptions below.

Installation

Accessible PDF View on the Chrome Web Store is the built extension for Chrome. Firefox is not listed yet — its build is below, and a listing on the add-on site is still to come.

To run it from source:

npm install
npm run build

Then load the unpacked extension:

  1. Open chrome://extensions.
  2. Turn on Developer mode.
  3. Click Load unpacked.
  4. Select .output/chrome-mv3.

For Firefox, npm run build:firefox produces .output/firefox-mv3; load manifest.json from it through about:debugging → This Firefox → Load Temporary Add-on. Both targets are Manifest V3.

For live-reloading development:

npm run dev

WXT launches a browser with the extension installed and rebuilds on change.

Usage

  1. Open a PDF in the browser.
  2. Click the Accessible PDF View toolbar button — or, from any page, right-click a link to a PDF and choose Open in Accessible PDF View. The right-click route works for links on the same site as the page you are on; for a link to another site, open it first and then use the toolbar button. By default the extension holds no permission to read another site on your behalf: when it cannot read one, the Reader says which site it could not reach and offers to ask for that site alone.
  3. A new tab opens with the Accessible Reader.

That button is the only way in unless you say otherwise: the browser's own viewer still opens every PDF you click. Opening PDFs in the settings panel changes it — an address ending in .pdf then opens in the Reader instead, and switching the setting back restores the browser's viewer. It asks for host access when you turn it on and hands it back when you turn it off; Permissions explains why it cannot work without it. It appears only where the browser offers the API, and reports the rule the browser has rather than the one it was asked for. It does not catch a local file or a PDF served from an address that does not end in .pdf — the toolbar button still opens both.

While it is on, the toolbar's Original PDF link opens the Reader as well: that link is a navigation to a .pdf address like any other, and the rule does not know who asked. The link says so while the setting is on — a new tab holding the same Reader is indistinguishable from a link that did nothing if you cannot see it. The Original view already shows the pages as they are, and switching the setting off gives the link back.

A toolbar across the top carries the view switch, the file name, print, a link to the original, and settings. Three views:

View What it shows
Reader The document rebuilt as semantic HTML.
Original The original PDF pages, rendered by PDF.js. Works even when there is no text to extract.
Markdown The document as copyable Markdown. Two sources: what the Reader is showing (including any generated image descriptions), or pdf-inspector's raw output.

The tab title says what the tab is doing: report.pdf (Analysing…) — Accessible PDF View while it works, then report.pdf (Reader) — Accessible PDF View, following the view switch from there. Progress is announced in a live region, which only reaches whoever is on the page; a tab left to work on its own used to say the extension's name and nothing else, which is indistinguishable from a tab that has finished. The document name stays in front, because a tab title truncates from the end and that is the part which tells one tab from another.

What the toolbar does not carry is the point. Page number, zoom percentage, fit-to-width and rotate are all controls for moving a fixed page image around in space. Once the document is HTML they have nothing to operate on — the browser's own zoom is better, and heading navigation replaces the page box. Copying Chrome's toolbar would mean making a screen reader user walk past a dozen dead controls to reach the document.

The bar is two things, and the seam between them is deliberate.

The actions — print, the link to the original, settings — carry role="toolbar", and with it the arrow keys: Left and Right move between them, Home and End go to the ends, both wrapping.

Not the roving tabindex the APG pairs with that role. Roving tabindex buys fewer Tab stops and sells controls that Tab cannot reach at all — and someone driving the page with Tab and Enter alone, which is what a two-switch setup gives you, would reach one of these three buttons and never learn the others exist. Arrow keys are a keyboard interface and satisfy WCAG 2.1.1, so the trade is conformant; it is still a control the user cannot get to. What it would have bought here is small: three actions, on a page that opens with a skip link. So every control keeps its own Tab stop, and the arrows work as well.

The view switch is a plain <select>, and it sits outside that toolbar, keeping its own Tab stop. The APG's toolbar pattern says to avoid controls whose operation needs the same arrow pair the toolbar navigates with, and a closed <select> answers Left and Right by changing its own value on Windows. Rather than replace it with a hand-built menu, or exile it to the end of the bar (the APG's own escape hatch), the toolbar boundary is drawn around the actions only. The pattern then applies with no exception to make, and the native element keeps every key it came with — Up and Down, Alt+Down, type-ahead, Home and End, touch, and whatever the platform's own option list does with a screen reader.

Suppressing the select's arrow keys with preventDefault and handing them to the toolbar would have kept it to a single Tab stop, and was rejected on evidence: Chrome honours it, Firefox does not reliably (1019630, 291082, 1428992). A control whose keyboard handling depends on the browser is worse than either version of it. The price paid instead is one extra Tab stop: two for the bar rather than one, against four before any of this.

The Reader does not replace the browser's built-in PDF viewer. It opens alongside it, and a link back to the original PDF is always in the toolbar.

Language

The interface is in English or Japanese, and which one is a setting rather than a consequence of your browser. browser.i18n cannot be switched at runtime — it follows the browser's own language and offers no way to say otherwise — so the Reader carries its own catalogue and the side panel offers 「自動 / English / 日本語」.

The extension's name and description, the ones the store lists, do follow the browser: nobody expects a store listing to obey a setting inside the thing being listed, and it cannot be read before installation. Those live in public/_locales/, with English as default_locale.

Three languages meet here and only the first is this setting:

What it is Where it comes from
Interface The Reader's own chrome this setting, or the browser
Document lang on the article the PDF's /Lang, detection, script
Description What the model writes in the document, your browser, or a setting

The names of the three views are katakana in Japanese — リーダー, オリジナル, マークダウン — and Latin in English. They are names, and a name in a Japanese interface has to be readable as Japanese: a screen reader meeting a bare Reader inside a Japanese sentence either switches voice for one word or mispronounces it. Every sentence that points at a view uses the same word, because one thing with two names is worse than either name.

English is the base language in the sense that matters for a contributor: lib/i18n/en.ts is the type, and ja.ts is declared as that type, so a key added to one and not the other does not compile. It is not the base language in the sense of being the original — the Reader was written in Japanese first, and the Japanese strings are the originals.

Settings

Settings in the toolbar opens the browser's side panel — Chrome's sidePanel, Firefox's sidebar. A side panel rather than a popup for one reason: a popup closes when it loses focus, and for anyone reading with a screen reader, a keyboard or a magnifier, moving focus is how you read.

In order: the interface language; text size, line spacing and column width; typeface; theme; the on-device AI and whether to get its download out of the way; which language AI-generated text is written in; which reading a tagged PDF opens with; which viewer opens a PDF; and resetting it all. The first is first because it changes the words of everything under it.

The theme choice selects the palette through a data-apv-theme attribute on the root element rather than through color-scheme and light-dark(). That was the original design and it did not work: with color-scheme: light computed on <html> over a dark OS setting, Chrome still resolved every light-dark() to its dark branch. color-scheme is still set — it is what paints scrollbars, form controls and the canvas behind the page — but it no longer picks the colours.

Two things about it are deliberate:

  • Every choice is a named step, not a slider. A step is something you can confirm without seeing it.
  • The settings become CSS custom properties, and the text size is a percentage of your browser's own default rather than a fixed size — so your browser setting still counts, and a user stylesheet still overrides everything.

Nothing is ever downloaded, and the two offered faces differ in what that means. Atkinson Hyperlegible Next is packaged inside the extension — 233 KB of variable font, upright and italic — so choosing it always works; it covers Latin only, and Japanese is drawn by the system face beside it. BIZ UDPGothic is Morisawa's and is used only if you have installed it, because a Japanese face is several megabytes rather than a few hundred kilobytes. The panel says which is which.

Development

npm run dev          # dev build with hot reload (Chrome)
npm run dev:firefox  # dev build for Firefox
npm run build        # production build -> .output/chrome-mv3
npm run build:firefox  # production build -> .output/firefox-mv3
npm run zip          # packaged build
npm run compile      # type-check only
npm run test         # Vitest

Project layout

entrypoints/
  background.ts          action click -> Reader tab
  reader/                the Reader page (React)
  sidepanel/             the settings panel
lib/
  ai/                    Chrome built-in model session lifecycle (shared)
  browser/               extension-facing glue: the tab, the handoff, the
                         per-site grant, the redirect, the side panel
  i18n/                  the interface catalogue — en.ts is the type, ja.ts
                         is declared as it
  pdf/
    source/              PdfSource implementations
    inspector/           pdf-inspector worker, protocol, adapter
    pdfjs/               metadata, struct-tree adapter, rendering, figures
    describe/            FigureDescriber interface and providers
    ocr/                 OcrProvider interface, providers, adapter, merge
    document-model.ts    the shared Document Model
    geometry.ts          rectangles and matrices, in PDF user space
  reader/                Document Model -> React/HTML
scripts/copy-assets.mjs  vendors PDF.js runtime assets into public/
tests/

See ARCHITECTURE.md for how the pieces fit together and where future features attach.

Tests

npm test needs nothing that is not in this repository. Where a suite needs a PDF it builds a small valid one in memory (tests/helpers/minimal-pdf.ts) and runs the real pdf-inspector WebAssembly module against it.

What the suite pins down is the promises: that a generated description is never readable as the author's own words; that "we could not read this page" is never written as "there is nothing here"; that a link is never split through; that every message exists in both languages. Those are what a contributor needs in order to change the code safely, and what a reader is owed.

The brief

prompt.md, the brief this was written from, is not in the repository. Code comments citing a section number (§15) mean that document. The design files the mark, the icons and the Reader's layout were worked out on are not here either; the icons they produce are, at public/icon/.

Image descriptions (Chrome only, optional)

Most PDFs contain images with no alternative text at all. Until you ask for something better, the Reader tells you the truth — that a figure is there and that nothing described it.

If you are using Chrome on a desktop that supports its built-in AI, the Reader can offer to describe those images on your device. Press the button in the header, above the rule — it sits with the document's own information and the OCR panel, and stays there whichever view is open; nothing happens automatically.

What you need to know about the result:

  • It is a guess made by a model, not the author's description. It may be wrong. The Reader labels every generated description as "Image (description generated by AI)" — 「画像(AIによる自動生成の説明)」 — wherever it is read, because someone relying on it cannot check it against the picture.

  • A description never replaces the author's own alternative text. Their words always win.

  • One exception, because it is not the author's: Word and PowerPoint can generate alternative text and write it into the PDF, disclaimer and all. That is detected, read as "Image (description generated by the word processor)" rather than as the author's, and offered for description again — it is usually a bare shape inventory ("タイムラインが含まれている画像") rather than a description.

  • Nothing is uploaded. Chrome's model runs locally, so the privacy model above is unchanged.

  • Descriptions are written in one of five languages — German, English, Spanish, French or Japanese. That is Chrome's list, not ours. A document in any other language still gets described, in English, and the panel says so before you start.

  • A PDF that declares the wrong language cannot be caught from inside the file — the file is the part that is wrong — so Language of AI-generated text in the settings panel switches both descriptions and OCR to the interface language instead. The panel then says that is what happened.

Requirements (Chrome's, not ours): Chrome on Windows 10/11, macOS 13+, Linux or ChromeOS; about 22 GB free storage; 16 GB RAM or a GPU with 4 GB+ of VRAM. The first run downloads a multi-gigabyte model. Where these are not met, the panel does not appear and the Reader behaves exactly as before.

The same on-device model is also registered as an experimental OCR provider for scanned pages. It is a small general-purpose model rather than a real OCR engine, and its accuracy on dense or low-quality Japanese scans is unverified — treat its output as a preview, not as the document.

Chrome's Language Detector is used in one narrow place: filling in the document language when the PDF declares none. A declared /Lang always wins, and a low-confidence guess is discarded — a wrong lang makes a screen reader mispronounce everything, which is worse than no lang at all.

Document information

The Reader reads the PDF's own description of itself with PDF.js — title, author, subject, keywords, creation and modification dates, producer, PDF version, and whether the document is tagged — and shows it in a collapsible Document information panel.

The field that matters most is /Lang. Without it the Reader cannot set a lang attribute, and a screen reader may read a Japanese document with an English voice. pdf-inspector does not report a language; PDF.js does, and the Reader applies it to the rendered article.

None of this is sent anywhere. It is read from the bytes already in the tab.

The Markdown view has two sources

Its selector sits in the header too, beside the structure picker and built the same way. The two questions are neighbours and easy to confuse — which reading of the PDF and whose Markdown — so the first option names the reading currently on screen rather than leaving the reader to hold it in their head.

They answer different questions, so both are available:

  • The same as the Reader — the document as rendered: the author's tag structure when the PDF is tagged, plus any figure descriptions that were generated. Generated descriptions keep their "generated by AI" label here too, so text copied out of the extension cannot be mistaken for the author's own words once it is somewhere this tool can no longer annotate it.
  • pdf-inspector's raw output — the parser's output, untouched. Still the right answer for "what did the parser make of this file".

GFM cannot express everything the Document Model holds — there is no syntax for a row header (<th scope="row">) — so the Reader, not the Markdown, is the accessible output.

Privacy

A normal, text-based PDF is processed entirely inside your browser.

  • The PDF bytes are read by the extension and handed to pdf-inspector (WebAssembly) and PDF.js. Both run locally.
  • Nothing is uploaded. There is no analytics, no telemetry, and no external request of any kind while opening a document.
  • Every asset the extension needs — the WebAssembly modules, the PDF.js worker, CMaps, fonts — is packaged inside the extension. Nothing is fetched from a CDN.

The one future exception is hosted OCR, described below. It is not implemented, and when it is, it will require explicit consent before anything is sent.

Permissions

Five on Chromium and four on Firefox — the difference is sidePanel, which Firefox does not need because its sidebar is a manifest key rather than a permission — and none of them is a host permission:

Permission Why
activeTab Grants a temporary permission for the tab you invoked the extension on, so it can read that tab's URL and fetch the PDF. It expires when you navigate away.
storage Holds your display settings — font size, line height, measure, typeface, theme — and your language choices. No document content is ever stored, which is why it can sync across your devices.
contextMenus Adds Open in Accessible PDF View when you right-click a link to a PDF.
sidePanel Opens the settings panel. Added automatically for the side-panel page.
declarativeNetRequestWithHostAccess Lets the Opening PDFs setting register its redirect rule — and only for sites you have granted access to, which is what the WithHostAccess half means.

Nothing above is a host permission, <all_urls> is never requested, and installing the extension grants no access to any site.

Host permissions are optional, and nothing asks for one until you do something that cannot be done without it. The manifest declares *://*/* as optional, because a pattern has to be declared before any part of it can be requested at runtime. Declaring it grants nothing.

Two things ask, for different amounts:

  • One site, when a PDF cannot be read. The extension normally reads a PDF by fetching it in the background while activeTab is live — your gesture standing in for a permission. That fails for a link to a PDF on another site, and on Firefox it fails for every cross-origin PDF, because Firefox's activeTab covers script injection and the tabs API but not a fetch. When it fails, the Reader says which site it could not reach and offers to ask for that one site. Granted access stays until you remove it from the browser's own extensions page; the offer says so.
  • Every site, for the Opening PDFs setting. Sending a PDF to the Reader means recognising its address before the page loads, which means standing in front of every navigation — and no browser allows that for sites an extension cannot already reach. There is no smaller version of that feature. It is asked for when you switch the setting on and handed back when you switch it off. The setting appears only where the browser offers the API; whether a given browser then honours the rule is its own business, and the panel reports what the browser says rather than what was asked for.

Leave both alone and nothing is ever asked for.

The rule itself is declarative: the browser matches it and acts on it, so the extension sees no request, runs no code per navigation, and learns nothing about what you visit.

One page is web-accessible: reader.html, and only that. A browser will not send a navigation to an extension page that is not declared that way — without it the redirect lands on ERR_BLOCKED_BY_CLIENT — and a manifest cannot make the declaration conditional on a setting. So the cost is paid whether or not you turn the setting on: any site can tell this extension is installed, by trying to load that one address. The address it is handed is fetched and read as a PDF, and the one place it becomes a link is sanitised the same way links taken out of a PDF are.

Because activeTab is granted to the extension's background at the moment you act, the PDF is fetched there and handed to the Reader tab through the extension's own Cache storage. Those bytes are dropped as soon as the Reader reads them, and any leftovers are cleared at browser startup.

OCR behaviour

Some PDFs contain no text at all — scans, photographed pages, image-only exports. pdf-inspector correctly returns nothing for those.

Empty output is not a parse error, and this extension never reports it as one. It is reported as what it is:

Some pages of this PDF contain no readable text. OCR is needed.

For those pages:

  • The Original view still shows you the page, rendered by PDF.js.
  • Screen reader users are told, per page, why the page produced no text — see below.
  • Only those pages are rasterised and read; mixed documents never process a page that already has text.

When a producer drops a page

"Empty parser output means OCR is needed, never that analysis failed" was a rule from the first day. One untagged document showed the half that was missing: "this page is empty" is a claim too, and it can be just as false.

pdf-inspector emitted nothing at all for two of its pages — no Markdown, not even their page markers — while the same result reported that both pages contain tables, and neither was flagged as needing OCR. PDF.js reads about 900 characters of ordinary Japanese from each. The Reader was telling readers those pages were empty.

So before a page may be called empty, PDF.js is asked. It is a decidable question — there is text or there is not — and it costs one getTextContent call on the pages that produced nothing.

A page recovered this way is marked origin: 'pdf-text' and says so where it is read: the words are the document's own, but the structure is gone. Headings, tables and lists were not recovered — the producer that should have supplied them is the thing that failed — so heading navigation will skip that page, and a reader navigating by heading needs to be told rather than left to conclude the page has none.

Two things it deliberately does not do. It never touches a requires-ocr page: that page has been looked at, OCR is the answer for it, and replacing the notice would take away the reader's only route to the content. And it ignores a page whose whole text is its folio — restoring "12" as a paragraph replaces one false statement with a more confusing one.

Why a page has no text

Two sources answer this, and the producer's own answer comes first.

pdf-inspector reports a cause per page — scanned, no_text, vector_text, suspected_garbled_text — while failing to read it. Three of those say what PDF.js would say after walking the page's content stream, so for those pages the walk is not run at all: on a 29-page scanned deck that is 26 content-stream walks skipped for an answer already in the result.

The fourth is why the two are kept apart. suspected_garbled_text is a page whose text is present and undecodable — a font with no usable ToUnicode map. Such a page is neither a picture nor outlines, so wording drawn from what it paints would send a reader looking for something that is not there. It gets its own sentence, and the first two pages of that same document are exactly this case.

The vocabulary is the producer's and may grow, so a code this build does not recognise is dropped rather than guessed at: the page falls back to being looked at directly, and the reader is told only what is certain.

What PDF.js contributes

"This page is displayed as an image" used to be said about every text-less page. It is often false. A slide deck exported with its text converted to outlines contains no image anywhere — every glyph is a vector path — so the message contradicted the figure detector, which correctly found nothing to describe. There is no way to settle that contradiction by looking at the page.

Each such page is therefore classified from its content stream (lib/pdf/pdfjs/page-content.ts) and the notice follows the answer:

Kind What is on the page What the Reader says
image Raster images are painted — a scan or a picture-only page stored as an image, so it cannot be extracted as text
vector Only paths, usually text converted to outlines drawn as shapes; there is no image on it either
blank Neither no readable content was found (OCR is not offered)

What the on-device model actually reads

Chrome's built-in model resizes every image input to 768×768, and says so in the console:

Image input (2480x3508) will be downscaled to 768x768. Dense spatial details
like small text may be lost.
Image input will be stretched from 4.21:1 to a square aspect ratio.

Two consequences, one addressed and one not.

Rendering more than that is pure waste, and it has stopped. A 300 DPI A4 page is 8.7 megapixels; the model reads 0.59 of them. The page render and the figure crop are now capped at what the consumer declares it uses (preferredImage on the provider), which for both Chrome providers is 768. A real OCR engine wants the resolution and simply leaves the cap unset. Nothing the model could see is lost — only the render, the PNG encode and the memory.

The squash is still there. A portrait page is stretched to a square before the model reads it, and so is a wide banner at 13:1. Feeding it a padded square instead would preserve the proportions and spend resolution on the padding, and which of the two the model reads better is not something this project can measure. Rather than guess, the geometry is left exactly as it was. It is a real limit on the on-device OCR's accuracy on dense pages, and it is a limit of the model's input format, not of the rendering.

OCR: Chrome's on-device model

Pages that need OCR get a panel offering to read them with Chrome's built-in model, on device. Per page as well as all at once — a page takes seconds, so a 29-page deck is minutes of work for a result the reader may only need one page of.

It sits in the header, above the rule and outside the view tabs, with the "OCR is needed" notice and the structure picker: which pages can be read at all is a fact about the whole document rather than a part of it, and the pages this offers to read are the ones a reader opens Original to look at.

A run of several pages shows each page as it is read, not at the end: the document fills in page by page while the rest is still being worked through, and Stop reading ends a long run without discarding what it has already read. A stopped run reports the pages it read and says the rest were not read — never that they could not be read, which would be a claim about pages nothing finished looking at.

It carries the same two rules as figure description: it never runs by itself, and a page it read is labelled "The text on this page was read by OCR (machine text recognition)" wherever that page is read. A machine's reading of a picture of text is not the document's text, and the person relying on it is the one who cannot check it.

A page OCR fails to read is left as it was. Replacing it with an empty page would turn "we could not read this" into "there is nothing here", which is the one claim this project refuses to make.

A picture on a page OCR reads survives it. The recognised text replaces the page's content, so a figure that was located in the PDF — with its image and any description already generated for it — would otherwise be deleted by the act of reading the page. Those figures are kept at the end of the page, and the page says so: "The 2 images on this page have been collected after the text. Recognised text carries no position information, so where they sat on the original page is not reflected here." Where they sat is not something this tool knows, and it does not pretend otherwise.

Requirements are the same as for figure description (Chrome desktop, ~22 GB free storage, 16 GB RAM or 4 GB VRAM). Where the model is unavailable the panel says so rather than disappearing — unlike figures, there is no honest placeholder to fall back on, so the reason has to be stated.

Bundled OCR engine

No bundled OCR engine ships in this version, and the Chrome model above is explicitly experimental — it is a small general-purpose model, not a purpose-built OCR engine, and Japanese vertical text and ruby annotations are exactly where it degrades.

In-browser OCR (Tesseract.js) was evaluated and deliberately deferred. The reasoning is recorded in full in lib/pdf/ocr/local-provider.ts; the short version is that it roughly doubles the extension's download size and is not accurate enough on the low-quality scans that bring people to this tool in the first place. Presenting that output as the document — to users who by definition cannot check it against the original — would be worse than saying nothing. When it lands it will be opt-in and labelled experimental.

Future hosted OCR

Hosted OCR (Azure AI Document Intelligence, Google Document AI, AWS Textract, or similar) may be offered later, possibly as a paid service. It is not implemented: there is no account system, no billing, no payment, no subscription and no backend in this codebase.

If it is added, OcrProvider.sendsDataExternally marks it, and the Reader will require explicit consent before anything leaves the device:

Running OCR sends this document, or its page images, to an external service.

Where structure comes from

The Reader has two sources of structure. It tells you which one you are looking at in the Document information panel — and for a tagged PDF, it lets you switch.

Tagged PDF (preferred). If the author tagged the document, PDF.js reads their structure tree and it is used as-is. This is the author's own answer, not a guess, and the difference is large: where inference sees a grid of text, the tags give a table with both column and row headers, and numbered sections as a properly numbered list.

A PDF can carry a structure tree without declaring itself Tagged PDF — the /MarkInfo /Marked true that ISO 32000 and PDF/UA require is missing. Its tree is read too, because it is still the author's structure, but only if it carries at least half of the text on the pages: without the declaration nothing says the tree is the whole document. The Document information panel says the declaration is missing, since other viewers and assistive technology may not use such tags.

Layout inference (fallback). For untagged documents, pdf-inspector infers structure from where the glyphs sit. It works well on straightforward documents, but it is a guess, and these are its observed limits, re-checked against the pdf-inspector build the extension ships (1.25.0 with the fix in vendor/pdf-inspector-wasm/) on the ten documents it is tested against:

  • The first row of a table is always a header row. GFM requires a delimiter row, so pdf-inspector emits one whether or not the table has headers — all nine tables in those documents, two of them label/value tables. For such a table this announces the wrong column header during table navigation. The extension keeps the producer's structure rather than guessing, because overriding it would break tables that genuinely do have headers.
  • The title may be read twice — once as the page's <h1>, and again in the body if the PDF repeats it in a form that was not recognised as a heading. Four of the seven documents with a title do.
  • Numbered sections are run into the paragraph after them. pdf-inspector reads 1. 背景および目的 and the body below it as one paragraph. The Reader splits such a heading back out where PDF.js shows the line break that ends it, which recovers all three in the one document that has them; a heading whose line break PDF.js cannot confirm stays in the body.

A limit listed here before, a closing 以上 read as a heading, no longer reproduces: the three documents that end with one read it as a paragraph.

When inference is in use, the Document information panel says so explicitly, so a reader knows how much to trust the headings and tables they are navigating.

Reading the same PDF both ways

A tagged PDF is read twice — once from the tag tree, once by inference — and How the structure was read switches between them. It sits in the header, under the document's own information and above the rule: which reading is on screen is a fact about the whole document, like its title and page count, not part of the text it changes. Which one a document opens with is a setting in the side panel, named with the same words as the control. It appears only when there really is more than one reading; an untagged file has one, and a control offering a choice that does not exist would be worse than no control.

A third reading, Structure tags + inferred headings, exists for one kind of document and is the default there: a tagged PDF whose tags carry no headings at all — which is what every tagged business document this was built against turned out to be, because their authors formatted headings by hand and Word had nothing to tag. It is the tag tree's reading with the headings the inferred reading can prove are there: an inferred heading is placed only where its text is a whole tagged paragraph or the start of one, and the words stay the author's. Each placed heading records that the claim was inferred, the picker says how many were placed, and the Document information panel says the rest is the author's structure. Where the author used heading styles at all, their outline stands and this reading is not offered.

It exists because the tag tree is the better answer and not always the right one. Tags can be stale, or applied by a tool that guessed; the way to find out is to read the document both ways. Switching changes every view at once — the Markdown view's "The same as the Reader" follows it, so the two never disagree about what the Reader is showing.

Two things about it are consequences of the design rather than choices:

  • The second reading is prepared when it is first asked for. Neither producer positions its figures; walking each page's content stream and rasterising what it finds is the expensive half of a load, and doing it twice up front would make every reader wait for a switch most will never make. The first switch says it is re-reading; after that it is instant.
  • Generated descriptions and OCR results do not cross. They stay with the reading they were applied to, and switching back finds them again. They do not follow, because the two readings do not agree on what a figure is: pdf-inspector's Nth image and the tag tree's Nth Figure are frequently not the same thing, and moving a checked description between them would attach it to the wrong picture.

The inferred reading of a tagged PDF is not just a curiosity — it is the one place the author's own alt text is absent, so it shows which pictures the document would leave undescribed if its tags were lost.

Current limitations

  • Local (file://) PDFs cannot be opened. That needs "Allow access to file URLs", which this version does not request. A clear message is shown instead.
  • PDFs displayed from a blob: URL cannot be read, because a blob URL is only valid inside the document that created it.
  • No OCR engine is bundled. Chrome's on-device model can read pages where it is available, but its accuracy on Japanese scans is unmeasured — treat it as experimental. Elsewhere, image-only PDFs can be viewed (Original) but not read as text.
  • Tagged PDF links inside the Reader are resolved by matching structure elements to link annotations by position. A link whose annotation rectangle does not overlap its text will fall back to plain text.
  • Tagged PDF Alt text is used when present, but many producers (Word among them) emit no Alt for figures. That is what the optional on-device description feature is for.
  • A /Alt is not proof a person wrote it. Word and PowerPoint generate alternative text and write it into the PDF verbatim, disclaimer included (「AI 生成コンテンツは誤りを含む可能性があります。」). Such text is detected, labelled "generated by the word processor" rather than presented as the author's, and — unlike the author's own words — offered for description again, because it is usually a shape inventory rather than a description. Detection matches only Microsoft's exact boilerplate: wrongly demoting text a person wrote is the more damaging error.
  • Figure positions come from the page content stream. In a tagged PDF the tags name the drawing operations each figure covers, so its position is exact — including for a figure drawn with vector paths rather than an embedded image.
  • In an untagged PDF there is no such link, so images are matched to figures by draw order, and a page that paints more images than it declares figures has the nearest ones merged (up to a gap of about a third of an inch). Both are heuristics and can mismatch.
  • Figure descriptions and OCR are written in the document's declared language when it has one. When it does not — very common — the browser's own language is used, and the panel says that is what happened. A document that declares the wrong language is not detectable at all, and the only answer to it is a setting the reader has to find. The model can only write German, English, Spanish, French and Japanese; anything else becomes English.
  • Chrome is the tested browser. Both targets are Manifest V3 and the code avoids chrome.* in favour of WXT's browser API, but Firefox has been run far less and Edge not at all. Two differences are known and handled: Firefox's activeTab does not cover a cross-origin fetch, so the route that reads a PDF without any host permission works on Chrome and not there — on Firefox every PDF from another origin ends at the offer to grant its site; and Firefox's Manifest V3 requires 'wasm-unsafe-eval' where its V2 refused it, which is now declared for both.
  • Image descriptions and OCR are Chrome's, and Firefox has no equivalent an extension can use. Firefox does have on-device inference — browser.trial.ml, the same runtime behind its own PDF alt-text feature — but on Release it is behind two about:config flags, its namespace promises no compatibility between versions, its captioning model writes English only, and it offers no OCR task at all. An accessibility feature that needs about:config is a feature for people who do not need this extension, so the panels leave themselves out there and say why.

Future direction

ROADMAP.md is the plan of record, and it is short on purpose: the next version of it should come from people who have used the extension rather than from the person who wrote it.

What is sketched in the code rather than promised there is in ARCHITECTURE.md, under Where future work attaches: a Structure view showing the readings side by side, a diff that says what they disagree about instead of leaving a reader to switch and compare, exact figure positions from /BBox, and the application/pdf MIME handler — the route that would catch a PDF served from an address not ending in .pdf, and that would need no host permission at all.

License

MIT. See LICENSE.

Third-party components and their licenses are listed in THIRD_PARTY_NOTICES.md.

About

Rebuilds PDFs as HTML with real headings, lists and tables — readable with a screen reader. Nothing leaves your browser. Chrome and Firefox, Manifest V3.

Resources

Stars

5 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages