html2markdown.py3(1) [debian man page]

HTML2MARKDOWN(1)						   User Commands						  HTML2MARKDOWN(1)

NAME

       html2markdown - converts a page of HTML into markdown.

SYNOPSIS

       html2markdown [options...] [(filename|url) [encoding]]

DESCRIPTION

       html2markdown downloads the specified HTML page, and converts it to text marked up with markdown.  The source HTML page may be a local file
       or remote URL.  If not specified, it will be read from standard input.  The output is printed to standard output.

       If an encoding is specified, it will override any encoding information provided by the HTTP Server.  When not specified,  python-feedparser
       (if available) will be used to determine the source encoding.  If not available, or when reading local files, the encoding is assumed to be
       UTF-8.

OPTIONS

       --ignore-emphasis
	      Don't include any formatting for emphasis.

       --ignore-links
	      Don't include any formatting for links.

       --ignore-images
	      Don't include any formatting for images.

       -g, --google-doc
	      Convert an html-exported Google Document.

       -d, --dash-unordered-list
	      Use a dash rather than a star for unordered list items.

       -b BODY_WIDTH, --body-width=BODY_WIDTH
	      Number of characters per output line, 0 for no wrap.

       -i LIST_INDENT, --google-list-indent=LIST_INDENT
	      Number of pixels Google indents nested lists.

       -s, --hide-strikethrough
	      Hide strike-through text. Only relevant when -g is specified as well.

       --version
	      Show program's version number and exit.

       -h, --help
	      Show a help message and exit.

AUTHOR

       This manpage was written for Debian, by Stefano Rivera <stefanor@debian.org>.

html2markdown 3.200.1						   January 2012 						  HTML2MARKDOWN(1)

Check Out this Related Man Page

PDFTOHTML(1)						      General Commands Manual						      PDFTOHTML(1)

NAME

       pdftohtml - program to convert PDF files into HTML, XML and PNG images

SYNOPSIS

       pdftohtml [options] <PDF-file> [<HTML-file> <XML-file>]

DESCRIPTION

       This  manual  page documents briefly the pdftohtml command.  This manual page was written for the Debian GNU/Linux distribution because the
       original program does not have a manual page.

       pdftohtml is a program that converts PDF documents into HTML. It generates its output in the current working directory.

OPTIONS

       A summary of options are included below.

       -h, -help
	      Show summary of options.

       -f <int>
	      first page to print

       -l <int>
	      last page to print

       -q     do not print any messages or errors

       -v     print copyright and version info

       -p     exchange .pdf links with .html

       -c     generate complex output

       -s     generate single HTML that includes all pages

       -i     ignore images

       -noframes
	      generate no frames. Not supported in complex output mode.

       -stdout
	      use standard output

       -zoom <fp>
	      zoom the PDF document (default 1.5)

       -xml   output for XML post-processing

       -enc <string>
	      output text encoding name

       -opw <string>
	      owner password (for encrypted files)

       -upw <string>
	      user password (for encrypted files)

       -hidden
	      force hidden text extraction

       -dev   output device name for Ghostscript (png16m, jpeg etc).  Unless this option is specified, Splash will be used

       -fmt   image file format for Splash output (png or jpg).  If complex is selected, but neither -fmt or -dev are specified, -fmt png will	be
	      assumed

       -nomerge
	      do not merge paragraphs

       -nodrm override document DRM settings

AUTHOR

       Pdftohtml was developed by Gueorgui Ovtcharov and Rainer Dorsch. It is based and benefits a lot from Derek Noonburg's xpdf package.

       This manual page was written by Soren Boll Overgaard <boll@debian.org>, for the Debian GNU/Linux system (but may be used by others).

SEE ALSO

       pdffonts(1), pdfimages(1), pdfinfo(1), pdftocairo(1), pdftoppm(1), pdftops(1), pdftotext(1)

																      PDFTOHTML(1)

Linux and UNIX Man Pages

html2markdown.py3(1) [debian man page]

Check Out this Related Man Page