Powered by OPF
OPF WIKI STATIC ARCHIVE
2,681 pages · 153 spaces · 776 tags · 4,025 history records · 96.2% of the original wiki recovered
Archived copy. This page was recovered from the Internet Archive snapshot of /display/SPR/Extraction of keywords (and images) from large collections of text based files taken on 2013-07-16. The original wiki at wiki.opf-labs.org is being decommissioned.

Extraction of keywords (and images) from large collections of text based files

Created by Paul Wheatley on May 08, 2012 · last edited by Paul Wheatley · on May 08, 2012 (view history) · 2 versions
Title
Extraction of keywords (and images) from large collections of text based files
Detailed description To facilitate a rapid initial categorisation of large hetergoneous collections of primarily text-based digital files/documents, a tool which parsed the documents, and presented a summary of (eg) the top 5 keywords (by wordcount) from each document, along with thumbnails of any images embedded in the document.
Issue champion Richard Freeston
Other interested parties
 
Possible Solution approaches Solution from previous mashup event: Analysis of Lucene Index Word Frequency
Context
Lessons Learned
Datasets
Solutions