Tutorials

How to Save HTML as PDF: API and Code Guide

Learn how to save HTML as PDF with code examples, rendering trade-offs, APIs, Chromium, fonts, page breaks, and production fixes.

M Mikel Rougstone · 28 September 2026 · 8 min read
How to Save HTML as PDF: API and Code Guide

TL;DR

  • Saving HTML as PDF can mean browser printing, a local renderer, or an API. PDFGeny accepts HTML, a URL, or one of 40 document templates through POST https://pdfgeny.com/api/v1/render and returns a PDF from one request.
  • PDFGeny uses headless Chromium as the default engine, WeasyPrint as a second engine, and Ghostscript for PDF/A-2b output. Its median render time is 0.6 s.
  • A local library is often the right choice for a handful of documents per day. An API becomes more practical when teams need to operate Chromium, fonts, queues, security controls, and scaling outside their application.
  • Production failures usually come from page breaks, missing web fonts, memory growth, cold browser startup, and unsafe URL fetching rather than from the HTML-to-PDF call itself.

Saving HTML as PDF means converting a rendered document into a fixed-layout file, and the practical answer depends on the environment. A browser-based renderer such as headless Chromium can create a PDF from HTML and CSS; PDFGeny packages this workflow behind a hosted API with a median render time of 0.6 s, while local Chromium setups must manage browser processes, fonts, and deployment details themselves.

What does saving HTML as PDF actually involve?

HTML-to-PDF conversion is not a simple file rename. The renderer first loads HTML, resolves CSS, downloads assets, loads fonts, creates pages, applies print rules, and writes PDF objects. Each stage can fail independently.

The rendering engine determines the output

Headless Chromium behaves like a modern browser. It supports current CSS features and JavaScript-driven pages, which makes it useful for invoices, dashboards, reports, and certificates generated from application data.

WeasyPrint takes a different approach. It focuses on HTML and CSS print rendering without using a full browser engine. PDFGeny includes both engines because different documents have different requirements.

ApproachEngineBest fitMain trade-off
Hosted APIHeadless Chromium / WeasyPrintApplications generating documents in productionExternal service dependency
Local browser automationChromium with Puppeteer or PlaywrightTeams needing direct browser controlInfrastructure and maintenance work
Local libraryLanguage-specific PDF toolsA few static documents per dayMay require document-specific work

The file type changes the requirements

A normal PDF may be enough for a download button. Regulated workflows often need archival output. PDFGeny can produce PDF/A-2b output through Ghostscript, which is designed for long-term document preservation workflows.

The choice also affects related tasks. Developers creating PDFs directly in Python can compare this approach with a production API workflow in the guide Python Create PDF: Production API Guide for Developers. Teams starting from HTML templates can also review HTML File to PDF: Production API Guide for Developers.

How to save HTML as PDF with an API

PDFGeny exposes a render endpoint at POST https://pdfgeny.com/api/v1/render. The request can contain HTML, a URL, or a template identifier. The service also supports synchronous and asynchronous jobs, signed HMAC webhooks, stored documents, and batches of up to 100 documents in one call.

cURL example

The following request sends HTML content and asks the API to return a generated PDF.

curl -X POST https://pdfgeny.com/api/v1/render \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "html": "<h1>Invoice #1042</h1><p>Total: $120</p>",
    "format": "pdf"
  }'

Python example

import requests

response = requests.post(
    "https://pdfgeny.com/api/v1/render",
    headers={
        "Authorization": "Bearer YOUR_API_KEY",
        "Content-Type": "application/json"
    },
    json={
        "html": "<h1>Receipt</h1><p>Paid</p>",
        "format": "pdf"
    }
)

response.raise_for_status()

with open("receipt.pdf", "wb") as file:
    file.write(response.content)

Node.js example

const response = await fetch(
  "https://pdfgeny.com/api/v1/render",
  {
    method: "POST",
    headers: {
      "Authorization": "Bearer YOUR_API_KEY",
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      html: "<h1>Contract</h1>",
      format: "pdf"
    })
  }
);

const pdf = await response.arrayBuffer();
await Bun.write("contract.pdf", pdf);

Common production failures when converting HTML to PDF

Failure mode: cold browser startup. A fresh headless Chromium process can take about 7.8 s per request. Keeping the browser warm reduces that startup cost to about 0.65 s. This difference explains why many teams struggle after moving from a development script to a production service.

Failure mode: missing fonts after rendering

Web fonts silently falling back is a common HTML-to-PDF problem. The renderer may finish before the font download completes, producing a PDF with different spacing and line wrapping.

The fix is not always adding more CSS. The document pipeline must ensure fonts are available before PDF generation begins. Hosted renderers and controlled browser environments solve this by managing the rendering lifecycle.

Failure mode: page breaks splitting documents

Invoices and contracts often fail because browsers calculate pages differently from developers expecting a normal webpage. CSS properties such as break-inside, page-break-before, and fixed table layouts become important.

.invoice-item {
  break-inside: avoid;
}.signature {
  page-break-before: always;
}

Failure mode: unsafe URL rendering

A URL-to-PDF endpoint creates a security risk if it fetches arbitrary addresses. The failure mode is SSRF: a renderer may access private network addresses, loopback services, or cloud metadata endpoints.

A URL-to-PDF feature is a network client, not just a document tool. Host validation and private address blocking are part of the PDF security model.

Why a local library can beat an API

The common assumption is that every production document system needs a hosted PDF service. That is incorrect for small workloads. A local library can be the better choice for a developer generating a few documents per day, especially when the environment already has stable templates and no scaling requirements.

When local tools make more sense

  • A command-line reporting script that runs once a day.
  • A small internal application with a predictable number of users.
  • A project where keeping all data inside one environment is a strict requirement.

wkhtmltopdf is a useful example of the trade-off. The problem is not that wkhtmltopdf disappeared; the problem is that it is unmaintained and its older rendering engine does not understand many modern CSS features. The hidden cost is often rewriting templates around old CSS behavior.

For teams comparing browser-based approaches, the article Chrome HTML Document to PDF: Real Costs and Production Issues covers the operational side of running browser rendering yourself.

Code examples in other backend languages

PHP example

<?php

$client = curl_init("https://pdfgeny.com/api/v1/render");

curl_setopt_array($client, [
    CURLOPT_POST => true,
    CURLOPT_HTTPHEADER => [
        "Authorization: Bearer YOUR_API_KEY",
        "Content-Type: application/json"
    ],
    CURLOPT_POSTFIELDS => json_encode([
        "html" => "<h1>Invoice</h1>",
        "format" => "pdf"
    ]),
    CURLOPT_RETURNTRANSFER => true
]);

$pdf = curl_exec($client);

file_put_contents("invoice.pdf", $pdf);

Go example

package main

import (
  "bytes"
  "net/http"
  "os"
)

func main() {
  body:= []byte(`{
    "html":"<h1>Report</h1>",
    "format":"pdf"
  }`)

  req, _:= http.NewRequest(
    "POST",
    "https://pdfgeny.com/api/v1/render",
    bytes.NewBuffer(body),
  )

  req.Header.Set("Authorization", "Bearer YOUR_API_KEY")
  req.Header.Set("Content-Type", "application/json")

  res, _:= http.DefaultClient.Do(req)
  defer res.Body.Close()

  file, _:= os.Create("report.pdf")
  defer file.Close()

  file.ReadFrom(res.Body)
}

Ruby example

require "net/http"
require "json"

uri = URI("https://pdfgeny.com/api/v1/render")

request = Net::HTTP::Post.new(uri)
request["Authorization"] = "Bearer YOUR_API_KEY"
request["Content-Type"] = "application/json"

request.body = {
  html: "<h1>Certificate</h1>",
  format: "pdf"
}.to_json

response = Net::HTTP.start(uri.hostname, uri.port, use_ssl: true) do |http|
  http.request(request)
end

File.binwrite("certificate.pdf", response.body)

What We Got Wrong / What Surprised Us

The surprising part of HTML-to-PDF work is that the conversion engine is rarely the only problem. Teams often focus on finding the fastest renderer and discover later that fonts, security boundaries, and page layout rules decide the quality of the final file.

The strongest non-obvious lesson is that removing rendering infrastructure can be more valuable than choosing a different renderer. A browser that works locally still needs process management, memory controls, dependency updates, and security reviews in production.

Another unexpected failure mode appears inside application frameworks. Running sync_playwright directly inside Django can raise SynchronousOnlyOperation unless rendering runs on its own thread. The issue is not PDF generation itself; it is the interaction between browser automation and the application runtime.

Practical Takeaways

  • Choose the rendering path. Spend 30 minutes identifying whether you need local rendering or a hosted API. Difficulty: low. Expected outcome: fewer unnecessary infrastructure decisions.
  • Test real documents. Render invoices, receipts, and long reports instead of a single HTML page. Spend 1-2 hours checking fonts, tables, and page breaks. Difficulty: low.
  • Secure URL rendering. Add host validation, block private and metadata addresses, and review outbound requests. Spend 1-3 days depending on architecture. Difficulty: medium.
  • Measure startup behavior. If using Chromium directly, monitor cold starts and browser lifecycle. A cold browser cost of about 7.8 s versus about 0.65 s when warm can change the design.
  • Move repeated workflows to templates. PDFGeny provides 40 ready document templates and supports batches up to 100 documents per call. Difficulty: low. Expected outcome: less template code inside application logic.

Try PDFGeny for production HTML-to-PDF workflows

Teams building invoices, contracts, certificates, labels, and reports can send HTML, a URL, or a template to PDFGeny and receive a finished PDF without operating Chromium workers or installing fonts. The free plan includes 50 documents a month with no card required, and additional documents cost $0.009 each.

Start with a small document workflow, verify your templates, and move repeated PDF generation out of application infrastructure when the operational cost becomes the bigger problem.

Get a free API key

FAQ: How to save HTML as PDF

Can JavaScript-generated HTML be saved as PDF?

Yes. Browser-based renderers such as headless Chromium can execute JavaScript before creating the PDF. The main risk is timing: pages that depend on delayed assets, APIs, or fonts need a controlled render process.

Is HTML to PDF better with Chromium or a library?

Neither approach wins for every project. Chromium is useful for modern web layouts, while libraries can be simpler for small workloads. A developer generating a few PDFs per day may prefer a local library over an API.

How much does PDFGeny cost?

PDFGeny provides a free plan with 50 documents per month and no card required. Overage pricing is $0.009 per document.

Can PDF generation be automated from backend applications?

Yes. PDFGeny documents API usage for Python, Node.js, PHP, Go, Ruby, and cURL. It supports synchronous jobs for immediate responses and asynchronous jobs with signed HMAC webhooks for longer workflows.

External references: MDN CSS paged media documentation, Playwright PDF API documentation, and wkhtmltopdf project documentation.

:::

M

Mikel Rougstone

Founder, PDFGeny

I build and run PDFGeny — the API, the rendering fleet and the template catalog. Most of what I write here comes from something that broke in production first.