How to Save HTML as PDF: API and Code Guide
Learn how to save HTML as PDF with code examples, rendering trade-offs, APIs, Chromium, fonts, page breaks, and production fixes.
TL;DR
- Saving HTML as PDF can mean browser printing, a local renderer, or an API. PDFGeny accepts HTML, a URL, or one of 40 document templates through
POST https://pdfgeny.com/api/v1/renderand returns a PDF from one request.
- PDFGeny uses headless Chromium as the default engine, WeasyPrint as a second engine, and Ghostscript for PDF/A-2b output. Its median render time is 0.6 s.
- A local library is often the right choice for a handful of documents per day. An API becomes more practical when teams need to operate Chromium, fonts, queues, security controls, and scaling outside their application.
- Production failures usually come from page breaks, missing web fonts, memory growth, cold browser startup, and unsafe URL fetching rather than from the HTML-to-PDF call itself.
Saving HTML as PDF means converting a rendered document into a fixed-layout file, and the practical answer depends on the environment. A browser-based renderer such as headless Chromium can create a PDF from HTML and CSS; PDFGeny packages this workflow behind a hosted API with a median render time of 0.6 s, while local Chromium setups must manage browser processes, fonts, and deployment details themselves.
What does saving HTML as PDF actually involve?
HTML-to-PDF conversion is not a simple file rename. The renderer first loads HTML, resolves CSS, downloads assets, loads fonts, creates pages, applies print rules, and writes PDF objects. Each stage can fail independently.
The rendering engine determines the output
Headless Chromium behaves like a modern browser. It supports current CSS features and JavaScript-driven pages, which makes it useful for invoices, dashboards, reports, and certificates generated from application data.
WeasyPrint takes a different approach. It focuses on HTML and CSS print rendering without using a full browser engine. PDFGeny includes both engines because different documents have different requirements.
| Approach | Engine | Best fit | Main trade-off |
|---|---|---|---|
| Hosted API | Headless Chromium / WeasyPrint | Applications generating documents in production | External service dependency |
| Local browser automation | Chromium with Puppeteer or Playwright | Teams needing direct browser control | Infrastructure and maintenance work |
| Local library | Language-specific PDF tools | A few static documents per day | May require document-specific work |
The file type changes the requirements
A normal PDF may be enough for a download button. Regulated workflows often need archival output. PDFGeny can produce PDF/A-2b output through Ghostscript, which is designed for long-term document preservation workflows.
The choice also affects related tasks. Developers creating PDFs directly in Python can compare this approach with a production API workflow in the guide Python Create PDF: Production API Guide for Developers. Teams starting from HTML templates can also review HTML File to PDF: Production API Guide for Developers.
How to save HTML as PDF with an API
PDFGeny exposes a render endpoint at POST https://pdfgeny.com/api/v1/render. The request can contain HTML, a URL, or a template identifier. The service also supports synchronous and asynchronous jobs, signed HMAC webhooks, stored documents, and batches of up to 100 documents in one call.
cURL example
The following request sends HTML content and asks the API to return a generated PDF.
curl -X POST https://pdfgeny.com/api/v1/render \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"html": "<h1>Invoice #1042</h1><p>Total: $120</p>",
"format": "pdf"
}'
Python example
import requests
response = requests.post(
"https://pdfgeny.com/api/v1/render",
headers={
"Authorization": "Bearer YOUR_API_KEY",
"Content-Type": "application/json"
},
json={
"html": "<h1>Receipt</h1><p>Paid</p>",
"format": "pdf"
}
)
response.raise_for_status()
with open("receipt.pdf", "wb") as file:
file.write(response.content)
Node.js example
const response = await fetch(
"https://pdfgeny.com/api/v1/render",
{
method: "POST",
headers: {
"Authorization": "Bearer YOUR_API_KEY",
"Content-Type": "application/json"
},
body: JSON.stringify({
html: "<h1>Contract</h1>",
format: "pdf"
})
}
);
const pdf = await response.arrayBuffer();
await Bun.write("contract.pdf", pdf);
Common production failures when converting HTML to PDF
Failure mode: cold browser startup. A fresh headless Chromium process can take about 7.8 s per request. Keeping the browser warm reduces that startup cost to about 0.65 s. This difference explains why many teams struggle after moving from a development script to a production service.
Failure mode: missing fonts after rendering
Web fonts silently falling back is a common HTML-to-PDF problem. The renderer may finish before the font download completes, producing a PDF with different spacing and line wrapping.
The fix is not always adding more CSS. The document pipeline must ensure fonts are available before PDF generation begins. Hosted renderers and controlled browser environments solve this by managing the rendering lifecycle.
Failure mode: page breaks splitting documents
Invoices and contracts often fail because browsers calculate pages differently from developers expecting a normal webpage. CSS properties such as break-inside, page-break-before, and fixed table layouts become important.
.invoice-item {
break-inside: avoid;
}.signature {
page-break-before: always;
}
Failure mode: unsafe URL rendering
A URL-to-PDF endpoint creates a security risk if it fetches arbitrary addresses. The failure mode is SSRF: a renderer may access private network addresses, loopback services, or cloud metadata endpoints.
A URL-to-PDF feature is a network client, not just a document tool. Host validation and private address blocking are part of the PDF security model.
Why a local library can beat an API
The common assumption is that every production document system needs a hosted PDF service. That is incorrect for small workloads. A local library can be the better choice for a developer generating a few documents per day, especially when the environment already has stable templates and no scaling requirements.
When local tools make more sense
- A command-line reporting script that runs once a day.
- A small internal application with a predictable number of users.
- A project where keeping all data inside one environment is a strict requirement.
wkhtmltopdf is a useful example of the trade-off. The problem is not that wkhtmltopdf disappeared; the problem is that it is unmaintained and its older rendering engine does not understand many modern CSS features. The hidden cost is often rewriting templates around old CSS behavior.
For teams comparing browser-based approaches, the article Chrome HTML Document to PDF: Real Costs and Production Issues covers the operational side of running browser rendering yourself.
Code examples in other backend languages
PHP example
<?php
$client = curl_init("https://pdfgeny.com/api/v1/render");
curl_setopt_array($client, [
CURLOPT_POST => true,
CURLOPT_HTTPHEADER => [
"Authorization: Bearer YOUR_API_KEY",
"Content-Type: application/json"
],
CURLOPT_POSTFIELDS => json_encode([
"html" => "<h1>Invoice</h1>",
"format" => "pdf"
]),
CURLOPT_RETURNTRANSFER => true
]);
$pdf = curl_exec($client);
file_put_contents("invoice.pdf", $pdf);
Go example
package main
import (
"bytes"
"net/http"
"os"
)
func main() {
body:= []byte(`{
"html":"<h1>Report</h1>",
"format":"pdf"
}`)
req, _:= http.NewRequest(
"POST",
"https://pdfgeny.com/api/v1/render",
bytes.NewBuffer(body),
)
req.Header.Set("Authorization", "Bearer YOUR_API_KEY")
req.Header.Set("Content-Type", "application/json")
res, _:= http.DefaultClient.Do(req)
defer res.Body.Close()
file, _:= os.Create("report.pdf")
defer file.Close()
file.ReadFrom(res.Body)
}
Ruby example
require "net/http"
require "json"
uri = URI("https://pdfgeny.com/api/v1/render")
request = Net::HTTP::Post.new(uri)
request["Authorization"] = "Bearer YOUR_API_KEY"
request["Content-Type"] = "application/json"
request.body = {
html: "<h1>Certificate</h1>",
format: "pdf"
}.to_json
response = Net::HTTP.start(uri.hostname, uri.port, use_ssl: true) do |http|
http.request(request)
end
File.binwrite("certificate.pdf", response.body)
What We Got Wrong / What Surprised Us
The surprising part of HTML-to-PDF work is that the conversion engine is rarely the only problem. Teams often focus on finding the fastest renderer and discover later that fonts, security boundaries, and page layout rules decide the quality of the final file.
The strongest non-obvious lesson is that removing rendering infrastructure can be more valuable than choosing a different renderer. A browser that works locally still needs process management, memory controls, dependency updates, and security reviews in production.
Another unexpected failure mode appears inside application frameworks. Running sync_playwright directly inside Django can raise SynchronousOnlyOperation unless rendering runs on its own thread. The issue is not PDF generation itself; it is the interaction between browser automation and the application runtime.
Practical Takeaways
- Choose the rendering path. Spend 30 minutes identifying whether you need local rendering or a hosted API. Difficulty: low. Expected outcome: fewer unnecessary infrastructure decisions.
- Test real documents. Render invoices, receipts, and long reports instead of a single HTML page. Spend 1-2 hours checking fonts, tables, and page breaks. Difficulty: low.
- Secure URL rendering. Add host validation, block private and metadata addresses, and review outbound requests. Spend 1-3 days depending on architecture. Difficulty: medium.
- Measure startup behavior. If using Chromium directly, monitor cold starts and browser lifecycle. A cold browser cost of about 7.8 s versus about 0.65 s when warm can change the design.
- Move repeated workflows to templates. PDFGeny provides 40 ready document templates and supports batches up to 100 documents per call. Difficulty: low. Expected outcome: less template code inside application logic.
Try PDFGeny for production HTML-to-PDF workflows
Teams building invoices, contracts, certificates, labels, and reports can send HTML, a URL, or a template to PDFGeny and receive a finished PDF without operating Chromium workers or installing fonts. The free plan includes 50 documents a month with no card required, and additional documents cost $0.009 each.
Start with a small document workflow, verify your templates, and move repeated PDF generation out of application infrastructure when the operational cost becomes the bigger problem.
FAQ: How to save HTML as PDF
Can JavaScript-generated HTML be saved as PDF?
Yes. Browser-based renderers such as headless Chromium can execute JavaScript before creating the PDF. The main risk is timing: pages that depend on delayed assets, APIs, or fonts need a controlled render process.
Is HTML to PDF better with Chromium or a library?
Neither approach wins for every project. Chromium is useful for modern web layouts, while libraries can be simpler for small workloads. A developer generating a few PDFs per day may prefer a local library over an API.
How much does PDFGeny cost?
PDFGeny provides a free plan with 50 documents per month and no card required. Overage pricing is $0.009 per document.
Can PDF generation be automated from backend applications?
Yes. PDFGeny documents API usage for Python, Node.js, PHP, Go, Ruby, and cURL. It supports synchronous jobs for immediate responses and asynchronous jobs with signed HMAC webhooks for longer workflows.
External references: MDN CSS paged media documentation, Playwright PDF API documentation, and wkhtmltopdf project documentation.
:::