Summarize Text by Ranking Sentences and Finding Keywords

The 'textrank' algorithm is an extension of the 'Pagerank' algorithm for text. The algorithm allows to summarize text by calculating how sentences are related to one another. This is done by looking at overlapping terminology used in sentences in order to set up links between sentences. The resulting sentence network is next plugged into the 'Pagerank' algorithm which identifies the most important sentences in your text and ranks them. In a similar way 'textrank' can also be used to extract keywords. A word network is constructed by looking if words are following one another. On top of that network the 'Pagerank' algorithm is applied to extract relevant words after which relevant words which are following one another are combined to get keywords. More information can be found in the paper from Mihalcea, Rada & Tarau, Paul (2004) < http://www.aclweb.org/anthology/W04-3252>.


This repository contains an R package which handles summarizing text by using textrank.

For ranking sentences, this algorithm basically consists of.

  • Finding links between sentences by looking for overlapping terminology
  • Using Google Pagerank on the sentence network to rank sentences in order of importance

For finding keywords, this algorithm basically consists of.

  • Extract words following one another to construct a word network
  • Using Google Pagerank on the word network to rank words in order of importance
  • Constructing keywords - which are the combination of relevant words identified by the Pagerank algorithm which follow each other

Installation & License

The package is available under the Mozilla Public License Version 2.0. Installation can be done as follows. Please visit the package documentation and package vignette for further details.

install.packages("textrank")
vignette("textrank", package = "textrank")

For installing the development version of this package: devtools::install_github("bnosac/textrank", build_vignettes = TRUE)

Support in text mining

Need support in text mining? Contact BNOSAC: http://www.bnosac.be

News

CHANGES IN textrank VERSION 0.3.0

  • Speedup textrank_candidates_all (thanks to @emillykkejensen)
  • Make sure in the help a sentence identifier is used explicitely when doing data preparation for textrank_sentences

CHANGES IN textrank VERSION 0.2.0

  • Allow to extract keywords for summarizing text by adding textrank_keywords
  • textrank is now called textrank_sentences

CHANGES IN udpipe VERSION 0.1.0

  • Allow to rank sentences for summarizing text

Reference manual

It appears you don't have a PDF plugin for this browser. You can click here to download the reference manual.

install.packages("textrank")

0.3.0 by Jan Wijffels, 3 months ago


https://github.com/bnosac/textrank


Browse source code at https://github.com/cran/textrank


Authors: Jan Wijffels [aut, cre, cph] , BNOSAC [cph]


Documentation:   PDF Manual  


Task views: Natural Language Processing


MPL-2.0 license


Imports utils, data.table, igraph, digest

Suggests textreuse, knitr, udpipe


See at CRAN