Linux Commands Examples

A great documentation place for Linux commands


program to convert PDF files into HTML, XML and PNG images

see also : pdfdetach - pdffonts - pdfimages - pdfinfo - pdftocairo - pdftoppm - pdftops - pdftotext


pdftohtml [options] <PDF-file> [<HTML-file> <XML-file>]

add an example, a script, a trick and tips

: email address (won't be displayed)
: name

Step 2

Thanks for this example ! - It will be moderated and published shortly.

Feel free to post other examples
Oops ! There is a tiny cockup. A damn 404 cockup. Please contact the loosy team who maintains and develops this wonderful site by clicking in the mighty feedback button on the side of the page. Say what happened. Thanks!


no example yet ...

... Feel free to add your own example above to help other Linux-lovers !


This manual page documents briefly the pdftohtml command. This manual page was written for the Debian GNU/Linux distribution because the original program does not have a manual page.

pdftohtml is a program that converts PDF documents into HTML. It generates its output in the current working directory.


A summary of options are included below.
-h, -help

Show summary of options.

-f <int>

first page to print

-l <int>

last page to print


do not print any messages or errors


print copyright and version info


exchange .pdf links with .html


generate complex output


generate single HTML that includes all pages


ignore images


generate no frames. Not supported in complex output mode.


use standard output

-zoom <fp>

zoom the PDF document (default 1.5)


output for XML post-processing

-enc <string>

output text encoding name

-opw <string>

owner password (for encrypted files)

-upw <string>

user password (for encrypted files)


force hidden text extraction


output device name for Ghostscript (png16m, jpeg etc). Unless this option is specified, Splash will be used


image file format for Splash output (png or jpg). If complex is selected, but neither -fmt or -dev are specified, -fmt png will be assumed


do not merge paragraphs


override document DRM settings

-wbt <fp>

adjust the word break threshold percent. Default is 10. Word break occurs when distance between two adjacent characters is greater than this percent of character height.

see also

pdfdetach , pdffonts , pdfimages , pdfinfo , pdftocairo , pdftoppm , pdftops , pdftotext


Pdftohtml was developed by Gueorgui Ovtcharov and Rainer Dorsch. It is based and benefits a lot from Derek Noonburg’s xpdf package.

This manual page was written by Søren Boll Overgaard <boll[:at:]debian[:dot:]org>, for the Debian GNU/Linux system (but may be used by others).

How can this site be more helpful to YOU ?

give  feedback