Parses real-world 'HTML' with a bundled copy of the 'Gumbo' parser (< https://codeberg.org/gumbo-parser/gumbo-parser>), which follows the 'WHATWG' parsing algorithm, so that no system library is required. Documents become immutable trees navigated with a documented subset of 'CSS' selectors. Attributes, text, lists, tables, links, forms and page metadata ('JSON-LD', microdata) are extracted into ordinary character vectors, lists and data frames, and nodes convert to 'Markdown'. Input is a string, raw bytes, a file, a URL or a connection, and raw input is decoded as browsers decode it, from a byte-order mark or a '' declaration. Parsing is bounded by limits on input size, native memory and nesting depth.
zuhtml parses real-world HTML the way a browser does and turns it into ordinary R values: character vectors, lists and data frames. It bundles the Gumbo HTML5 parser, so it needs no system library, and it has no hard dependencies.
html_problems() lists what was repaired."0012" stays "0012".<meta> tags, JSON-LD, microdata) and forms,
and convert any node to Markdown with html_markdown().encoding, or the page's <meta charset>.html_read() reads a file, a URL or any R connection. zuhtml has no HTTP
client of its own: a URL goes through base R's url(), and a fetcher
that needs headers or authentication hands it the body. zuhtml does not
run JavaScript or sanitize HTML.
install.packages("zuhtml")
The development version, from GitHub:
# install.packages("pak")
pak::pak("pedrobtz/zuhtml")
library(zuhtml)
doc <- html_parse('
<div class=product><h2>Sencha</h2><span class=price>3.50</span>
<a href="sencha.html">details</a></div>
<div class=product><h2>Genmaicha</h2>
<a href="genmaicha.html">details</a></div>',
base_url = "https://example.org/shop/"
)
cards <- html_elements(doc, ".product")
data.frame(
name = html_text_clean(html_element(cards, "h2")),
price = html_text_clean(html_element(cards, ".price")),
url = html_url(html_element(cards, "a"))
)
#> name price url
#> 1 Sencha 3.50 https://example.org/shop/sencha.html
#> 2 Genmaicha <NA> https://example.org/shop/genmaicha.html
html_element() returns one result per card, with a missing value where a
card has no price, so the columns stay aligned.
html_read() reads a file, a connection or a URL. A URL becomes the
document's base URL, so relative links resolve against the page:
doc <- html_read("https://cran.r-project.org/web/views/")
html_title(doc)
#> [1] "CRAN Task Views"
rows <- html_elements(doc, "table tr")
views <- data.frame(
topic = html_text_clean(html_element(rows, "td:nth-child(2)")),
url = html_url(html_element(rows, "a"))
)
head(views, 3)
#> topic url
#> 1 Actuarial Science https://cran.r-project.org/web/views/ActuarialScience.html
#> 2 Agricultural Science https://cran.r-project.org/web/views/Agriculture.html
#> 3 Anomaly Detection https://cran.r-project.org/web/views/AnomalyDetection.html
The getting started guide walks through a complete extraction. There are also guides to selectors, tables and lists, and limits, encodings and safety.
zuhtml is MIT-licensed. The bundled Gumbo parser is Apache-2.0, and
LICENSE.note
explains how the two apply.