ReportWright
MIT licensed Get it

● Benchmarks · reproducible

PDF benchmarks, losses included.

ReportWright's engine against pdfmake, react-pdf, PDFKit and Puppeteer, on the same invoice and the same grouped table report, with the same data, font file, page size and margins. Every number on this page comes from results.json, written by a script in the repository that you can run yourself. Where another library is faster, smaller or lighter, the tables say so.

01

Method

What is measured and how. The full workload spec is bench/README.md; the shared data and layout numbers are bench/data.mjs.

Workloads

  • Invoice: one A4 page; seller, buyer, a 12-line item table, subtotal, VAT, total, a wrapping note.
  • Warm invoice: the invoice 50 times in one process; the median of the last 40.
  • Table report: 1,000, 10,000 and 100,000 rows; 7 fixed-width columns; the header repeated on every page; four region groups with a group header and a subtotal row; a grand total; "Page N of M".

Same inputs

  • The same generated rows, already in group order, and the same pre-formatted amounts.
  • Inter Regular and Bold, the same .ttf files for every library, embedded and subset.
  • A4, 36 pt margins, an 18 pt footer band: the same 523 × 752 pt body.
  • Libraries without grouping get the groups and sums from one shared function; they lay out the rows.

Measured

  • Render: building the document to the PDF on disk, inside the process. Data is made first.
  • Cold start: a fresh process to the first invoice on disk, timed from outside: Node start, import, browser launch, render, exit.
  • Peak RSS: Node's own peak, or the whole process tree sampled every 100 ms (Chrome counts).
  • Output size, install size (each library alone, with its browser), veraPDF verdicts.

Runs

  • Every run is a new process. The first run of each cell is a warmup and is dropped.
  • Then 5 runs (3 at 100,000 rows): the median is shown, with the min–max spread.
  • A run past the timeout is stopped and shown as "did not finish"; larger sizes of that library are then skipped.

Contestants and versions

LibraryPinnedHow it is used

Reproduce it yourself

git clone https://github.com/MrArun005/reportwright
cd reportwright/bench
npm ci                   # Puppeteer downloads its Chrome
node install-size.mjs    # out/install.json
node verapdf.mjs         # out/verapdf.json (needs veraPDF)
node run.mjs             # out/results.json, merging both

# or on GitHub: Actions › bench › Run workflow
# (.github/workflows/bench.yml, ubuntu-latest)

# a quick check without a browser, 100 rows:
node run.mjs --smoke --out=out/smoke.json

02

Results

Median of the measured runs; the thin line on each bar is the min–max spread. Lower is better in every column. The best value in a column is in bold, whichever library it belongs to.

Loading results.json…

03

Standards and capabilities

The PDF/A and PDF/UA columns are veraPDF's verdicts on each library's own output, not claims. A library without such an option is still validated: its ordinary file is what you would get.

How each library does the table report

04

Known biases

What could tilt these numbers, as far as we know. Corrections are welcome; a pull request that makes a competitor faster will be merged and the page rerun.

Who wrote it

The harness is written by ReportWright's author. Each implementation follows its library's documentation, but the author knows ReportWright best.

Unsure of best practice

  • react-pdf at 10,000+ rows: one wrapping Page is the documented pattern; splitting may be faster but breaks "Page N of M".
  • Puppeteer: a new page a document; reusing one page, or streaming HTML from a local server, may be faster.
  • PDFKit: manual layout as specified; 0.20 also has doc.table().

Different architectures

ReportWright streams the table (exportPdfStream) and reads the rows twice for "Page N of M". The others build the whole document in memory; PDFKit's bufferPages keeps every page until the end. Puppeteer's render time leaves out the browser launch, which is in its cold start.

Measurement

Every run is a fresh process, so render times include first-use costs (JIT, font parsing); the warm invoice shows the steady state. Process-tree memory is sampled every 100 ms and can miss a short peak. Line heights and paddings differ slightly by library, so page counts can differ; they are shown next to each result.

Machine

One shared GitHub Actions runner; its neighbours add noise, and the min–max spread shows how much. Your hardware will give other absolute numbers.

Not measured

  • Concurrency and throughput under load.
  • Charts, images, rich text and non-Latin scripts.
  • Browser (client-side) rendering.
  • Visual fidelity beyond the shared layout.