Skip to content

Web articles

script and render accept public article URLs. Pages are fetched on this machine with a browser-like HTTP client, then extracted locally; no hosted reader receives the URL or page text.

sase-listen script https://example.com/article --json
sase-listen render https://example.com/article --dry-run

The first request stores the original HTML, repaired Markdown, extraction metadata, and deterministic verbatim script under $XDG_DATA_HOME/sase-listen/sources/. Later requests reuse that source, so a render can run offline and uses the same script. Pass --refresh to fetch and extract the URL again. script URL --json reports the source directory, word count, omissions, and outline headings found, restored, or missing.

The extractor restores H2 and H3 headings when their section text can be matched to the extracted Markdown. Some pages omit headings from extraction; the repair report makes that visible. Verbatim scripts normalize prose and report omitted content such as code, large tables, and math.

Only public HTML pages (text/html and application/xhtml+xml) and PDF documents are supported. Login pages and JavaScript-only pages are not extracted. For a page that requires browser access, save it and use --html FILE:

sase-listen script https://example.com/article --html saved-page.html

If a site blocks automated requests, the command explains how to retry with a saved HTML file. Pages shorter than 150 extracted words are rejected to catch login screens, abstracts, and error pages.

Brief and full editions

URL renders default to an AI-written brief adaptation. Choose full to cover the article's major sections in order, or verbatim for the deterministic article-text reading. A full edition is an adaptation within a 2,400-word budget, not a word-for-word transcript. Brief targets about 600 words.

Review a generated script before rendering it:

sase-listen script https://example.com/article -e full -o article_narration.md
sase-listen render article_narration.md

The writer uses Gemini and the existing Gemini engine credentials. Generated scripts and writer metadata are cached beside the stored source; a matching source, model, edition, and prompt version reuses the script. A cached script saved with lint findings is re-checked on the next run and reused once it lints clean, whether from rule fixes or hand edits. Pass --refresh to fetch and write again. Script writing uses Gemini API tokens and incurs model usage charges. render --dry-run writes the script and reports the later audio synthesis estimate.

PDF documents

PDF URLs work everywhere an article URL does (render, script, all three editions brief / full / verbatim, caching, --refresh, auto-publish). A URL is treated as a PDF when its content type is application/pdf (or application/x-pdf), or when an octet-stream or missing content type carries %PDF bytes:

sase-listen render https://arxiv.org/abs/2608.25174 -e full

Local .pdf files are valid sources too, cached by content hash so re-renders of the same file reuse the extraction:

sase-listen render paper.pdf -e full
sase-listen script paper.pdf -e verbatim --json

render URL --html saved.pdf keeps the URL identity for a PDF that had to be downloaded by hand.

Extraction runs locally with pdfminer.six. Headings come from the PDF bookmark outline when present, else from font sizes (large text, or short bold lines with section numbering). Dropped along the way: running headers and footers, page numbers, footnotes and other small text, table cells, numeric citation brackets like [12], the page-1 author/affiliation block, boilerplate (CCS Concepts, Keywords, copyright notices), and the references section. Kept metadata: the document title, a speakable author credit ("A and colleagues" for four or more authors), the arXiv date and arXiv site when the arXiv stamp is present, and the page count.

Limits: 64 MiB, 400 pages, no OCR — scanned PDFs without a text layer are rejected with "The extracted PDF text is too short". Editions and caching match articles, and script … --json reports the stored source_dir, word count, and outline diagnostics (found, restored, missing, plus the outline source: bookmarks, fonts, or none).

arXiv papers

arXiv paper URLs are recognized on arxiv.org, www.arxiv.org, or export.arxiv.org: /abs/<id>, /html/<id>, and /pdf/<id> (with or without a .pdf suffix). Old- and new-style IDs are accepted, any query or fragment is ignored, and the version suffix (for example v2) is kept. Each is fetched as https://arxiv.org/pdf/<id>, so all forms for one paper share one cached source, writer script, and episode. --html FILE keeps the URL as given. A failed PDF fetch is reported and does not fall back to the abstract page.