Parse 'HTML' with a Bundled 'Gumbo' Parser

Parses real-world 'HTML' with a bundled copy of the 'Gumbo' parser (< https://codeberg.org/gumbo-parser/gumbo-parser>), which follows the 'WHATWG' parsing algorithm, so that no system library is required. Documents become immutable trees navigated with a documented subset of 'CSS' selectors. Attributes, text, lists, tables, links, forms and page metadata ('JSON-LD', microdata) are extracted into ordinary character vectors, lists and data frames, and nodes convert to 'Markdown'. Input is a string, raw bytes, a file, a URL or a connection, and raw input is decoded as browsers decode it, from a byte-order mark or a '' declaration. Parsing is bounded by limits on input size, native memory and nesting depth.


zuhtml

R-CMD-check coverage

zuhtml parses real-world HTML the way a browser does and turns it into ordinary R values: character vectors, lists and data frames. It bundles the Gumbo HTML5 parser, so it needs no system library, and it has no hard dependencies.

  • Malformed markup is repaired by the HTML parsing algorithm, not rejected. html_problems() lists what was repaired.
  • Select elements with a documented subset of CSS. Anything outside the subset is an error, never a partial match.
  • Extract text, attributes, links, lists and tables. Tables handle row and column spans and keep every column as character unless you ask for conversion, so "0012" stays "0012".
  • Read page metadata (title, <meta> tags, JSON-LD, microdata) and forms, and convert any node to Markdown with html_markdown().
  • Raw bytes are decoded as a browser decodes them: from a byte-order mark, your encoding, or the page's <meta charset>.
  • Every call runs under explicit limits on input size, native memory and nesting depth, and every error is a classed condition.

html_read() reads a file, a URL or any R connection. zuhtml has no HTTP client of its own: a URL goes through base R's url(), and a fetcher that needs headers or authentication hands it the body. zuhtml does not run JavaScript or sanitize HTML.

Installation

install.packages("zuhtml")

The development version, from GitHub:

# install.packages("pak")
pak::pak("pedrobtz/zuhtml")

Example

library(zuhtml)

doc <- html_parse('
  <div class=product><h2>Sencha</h2><span class=price>3.50</span>
    <a href="sencha.html">details</a></div>
  <div class=product><h2>Genmaicha</h2>
    <a href="genmaicha.html">details</a></div>',
  base_url = "https://example.org/shop/"
)

cards <- html_elements(doc, ".product")
data.frame(
  name  = html_text_clean(html_element(cards, "h2")),
  price = html_text_clean(html_element(cards, ".price")),
  url   = html_url(html_element(cards, "a"))
)
#>        name price                                     url
#> 1    Sencha  3.50    https://example.org/shop/sencha.html
#> 2 Genmaicha  <NA> https://example.org/shop/genmaicha.html

html_element() returns one result per card, with a missing value where a card has no price, so the columns stay aligned.

html_read() reads a file, a connection or a URL. A URL becomes the document's base URL, so relative links resolve against the page:

doc <- html_read("https://cran.r-project.org/web/views/")
html_title(doc)
#> [1] "CRAN Task Views"

rows <- html_elements(doc, "table tr")
views <- data.frame(
  topic = html_text_clean(html_element(rows, "td:nth-child(2)")),
  url   = html_url(html_element(rows, "a"))
)
head(views, 3)
#>                  topic                                                        url
#> 1    Actuarial Science https://cran.r-project.org/web/views/ActuarialScience.html
#> 2 Agricultural Science      https://cran.r-project.org/web/views/Agriculture.html
#> 3    Anomaly Detection https://cran.r-project.org/web/views/AnomalyDetection.html

The getting started guide walks through a complete extraction. There are also guides to selectors, tables and lists, and limits, encodings and safety.

Licence

zuhtml is MIT-licensed. The bundled Gumbo parser is Apache-2.0, and LICENSE.note explains how the two apply.

Reference manual

It appears you don't have a PDF plugin for this browser. You can click here to download the reference manual.

install.packages("zuhtml")

0.1.0 by Pedro Baltazar, 10 hours ago


https://github.com/pedrobtz/zuhtml, https://pedrobtz.github.io/zuhtml/


Report a bug at https://github.com/pedrobtz/zuhtml/issues


Browse source code at https://github.com/cran/zuhtml


Authors: Pedro Baltazar [aut, cre, cph] , Google Inc. [cph] (Gumbo , bundled in src/vendor/gumbo) , Bjoern Hoehrmann [cph] (UTF-8 decoder in src/vendor/gumbo/utf8.c)


Documentation:   PDF Manual  


MIT + file LICENSE license


Suggests jsonlite, knitr, rmarkdown, testthat, withr


See at CRAN