Tutorials

Python PDF Library: Production API Guide for Developers

Compare Python PDF libraries and APIs with real code, costs, failures, and PDFGeny production patterns for document generation.

M Mikel Rougstone · 29 September 2026 · 6 min read
Python PDF Library: Production API Guide for Developers

TL;DR

  • Python PDF libraries can generate files locally, but production systems often inherit browser, font, memory and security problems. PDFGeny renders PDFs through a hosted API with a median render time of 0.6 s and accepts HTML, URLs or 40 ready document templates.
  • Headless Chromium is the default PDFGeny engine, with WeasyPrint available as a second engine and Ghostscript used for PDF/A-2b output.
  • A cold headless Chromium process can cost about 7.8 s per request; keeping the browser warm reduces that pattern to about 0.65 s.
  • For small workloads, a local Python PDF library is often the better choice. For production apps handling invoices, certificates, reports or labels, removing browser operations can change the operational model.

A Python PDF library is the right answer for many scripts, internal tools and low-volume jobs. The production question changes when an application needs invoices, receipts, contracts, certificates, reports or labels on demand. PDFGeny processes a render request through POST https://pdfgeny.com/api/v1/render, with a median render time of 0.6 s, while developers avoid running their own Chromium process, browser pool or font installation workflow.

Python PDF libraries versus PDF APIs: the production trade-off

The first decision is not “which PDF library is fastest?” The first decision is where rendering complexity should live. A Python package keeps the whole pipeline inside your application. An API moves browser management, renderer updates and document workers outside your application boundary.

ApproachBest fitMain variable to manage
Local Python PDF libraryScripts, offline tools, small document countsRenderer behavior and deployment dependencies
Headless browser serviceHTML/CSS-heavy documentsBrowser memory, startup time and security
Hosted PDF APIApplication-generated documentsAPI design, authentication and job handling

Where libraries still win

For a handful of PDFs per day, a local library beats an API. A developer creating a weekly report generator or exporting a personal archive usually benefits from fewer moving parts. A dependency such as ReportLab can run without network access, accounts or external services.

The trade-off appears once HTML becomes the source format. Modern invoices often depend on CSS layouts, web fonts, charts and browser rendering behavior. At that point, the “PDF library” problem becomes a rendering environment problem.

The hidden cost of browser-based generation

Headless Chromium gives developers strong HTML compatibility, but production deployments must handle the browser lifecycle. A cold headless Chromium instance costs about 7.8 s per request; keeping the browser warm brings that pattern down to about 0.65 s.

This is why many teams end up maintaining browser pools, containers and worker queues instead of only writing PDF code. PDFGeny uses headless Chromium as the default engine and also supports WeasyPrint for different rendering needs.

Real Python code: calling a PDF generation API

A Python PDF API integration can be a small HTTP client. The important production variables are authentication, payload size, error handling and whether the job is synchronous or asynchronous.

import requests

API_KEY = "your_api_key"

payload = {
    "html": "<html><body><h1>Invoice #1042</h1></body></html>",
    "format": "pdf"
}

response = requests.post(
    "https://pdfgeny.com/api/v1/render",
    json=payload,
    headers={
        "Authorization": f"Bearer {API_KEY}"
    },
    timeout=30
)

response.raise_for_status()

with open("invoice.pdf", "wb") as file:
    file.write(response.content)

PDFGeny supports sync and async jobs, signed webhooks using HMAC, stored documents and batch requests of up to 100 documents in one call. Those features matter when a PDF request is part of a larger workflow rather than a one-off file download.

Developers comparing implementation patterns can also review Python Create PDF: Production API Guide for Developers and Python Generate PDF: Production API Guide for Developers.

Failure modes that break PDF generation in production

PDF generation failures are usually not caused by the final PDF command. They come from the environment around the renderer.

Failure mode: missing fonts and silent fallbacks

Web fonts silently fall back when a renderer finishes before the font has loaded. A document can look correct in a browser but change after conversion because the PDF engine captured the page before font assets were ready.

The fix depends on the architecture. Local Chromium setups need font installation, network access rules and wait conditions. A hosted renderer needs predictable asset loading behavior.

Failure mode: incorrect page breaks

CSS page breaks are a common issue for invoices and reports. A table row split across two pages can create unreadable documents even when the HTML looks correct.

Chrome HTML document conversion has its own operational concerns, including browser lifecycle management. Developers working through those trade-offs can compare approaches in Chrome HTML Document to PDF: Real Costs and Production Issues.

Failure mode: SSRF through URL rendering

A URL-to-PDF endpoint is an SSRF hole until the application resolves the host and rejects private, loopback and metadata addresses. This applies to both custom browser workers and hosted systems accepting arbitrary URLs.

A safe implementation validates destinations before the renderer requests content. Public URL fetching is useful, but it needs network controls.

Security teams commonly recommend reviewing URL fetching behavior against SSRF guidance from sources such as OWASP SSRF prevention guidance.

Why wkhtmltopdf is not dead, but the maintenance cost is real

wkhtmltopdf remains in many applications because it solved a real problem: converting HTML into PDFs with a simple command. The contrarian observation is that the tool is not “dead”; it is unmaintained, and the bigger cost is usually the CSS it never learned.

Teams often spend time adapting modern HTML and CSS back to older rendering engines. A browser-based renderer reduces that mismatch, but introduces browser operations.

Tool choiceOperational questionCommon limitation
wkhtmltopdfCan the HTML be written for its renderer?Older CSS support
Local ChromiumCan the team operate browsers in production?Memory and cold starts
PDF APICan external rendering fit the application?Network dependency

What We Got Wrong / What Surprised Us

The strongest non-obvious finding is that “more control” is not always better control. A local renderer gives direct access, but every dependency becomes part of the application: browser versions, fonts, system packages and security rules.

The surprising implementation detail is Django integration. Running sync_playwright inside Django raises SynchronousOnlyOperation unless the rendering happens on its own thread. A renderer that works in a command-line script can fail inside an async web framework.

The other surprise is that low volume changes the recommendation. A team generating only a few documents daily may spend more engineering time operating an API than running a local package. The correct choice depends on document volume, security requirements and deployment constraints.

Practical Takeaways

  • Measure your document workload. Time estimate: 30 minutes. Difficulty: Easy. Expected outcome: identify whether local generation or a service model fits your request volume.
  • Choose your rendering source. Time estimate: 1-2 hours. Difficulty: Medium. Expected outcome: decide whether your documents need HTML/CSS rendering, templates or programmatic drawing.
  • Test fonts and page breaks before launch. Time estimate: 2-4 hours. Difficulty: Medium. Expected outcome: catch layout failures before users receive incorrect PDFs.
  • Secure every URL renderer. Time estimate: 1-3 hours. Difficulty: Medium. Expected outcome: block SSRF paths by validating hosts and rejecting private, loopback and metadata addresses.
  • Add async processing for large batches. Time estimate: 1 day. Difficulty: Medium. Expected outcome: separate document creation from user-facing requests.

PDFGeny provides a hosted rendering endpoint for applications that need PDFs without maintaining Chromium workers. Send HTML, a URL or a template and get a finished PDF back in one API call.

Get a free API key

FAQ

What is the best Python PDF library for production?

The best choice depends on the document type. Local Python libraries fit offline and low-volume generation. Applications producing HTML-based invoices, certificates or reports often need a browser renderer or PDF API because CSS compatibility becomes the main requirement.

Can Python generate PDFs from HTML?

Yes. Python applications can send HTML to a renderer such as Chromium-based systems or use libraries that interpret document layouts. PDFGeny accepts HTML through POST https://pdfgeny.com/api/v1/render and returns a generated PDF.

How much does PDFGeny cost?

PDFGeny offers a free plan with 50 documents a month and no card required. Overage pricing is $0.009 per document. Batch requests support up to 100 documents in one call.

Should I replace wkhtmltopdf?

Not always. wkhtmltopdf can remain suitable for stable, simple HTML documents. Replacement becomes more attractive when teams need newer CSS behavior, modern fonts, browser rendering or fewer local renderer maintenance tasks.

:::

M

Mikel Rougstone

Founder, PDFGeny

I build and run PDFGeny — the API, the rendering fleet and the template catalog. Most of what I write here comes from something that broke in production first.