Powered by OPF
OPF WIKI STATIC ARCHIVE
2,681 pages · 153 spaces · 776 tags · 4,025 history records · 96.2% of the original wiki recovered
Archived copy. This page was recovered from the Internet Archive snapshot of /pages/viewpage.action?pageId=16713702 taken on 2021-10-23. The original wiki at wiki.opf-labs.org is being decommissioned.

Extracting and aggregating metadata with Apache Tika

Added by Thom Carter · last edited by Thom Carter · on Sep 20, 2012

Extracting and aggregating metadata with Tika

Apache Tika was used with a custom wrapper to extract metadata (e.g. author, title, extent, dates and file formats) and content (text) from collection files. A report was produced that summarised the metadata and content from the collection. This information will be used to inform collection management decisions and identify potential preservation issues.

A detailed description of the Solution. Feel free to include links to further information (eg. OPF blog posts!). Note that a Solution is a specific digital preservation application of a software tool or tools to a particular Issue with a particular Dataset. It might for example be a scripted tool, or a myExperiment workflow

Solution Champion
Thom Carter, Rebecca Webster

Corresponding Issue(s)
Produce a report summarising collection metadata and content
Sorting, appraising and metadata creation for deposited personal collections

Tool/code link
A link to code on Git hub or a corresponding myExperiment if applicable

Tool Registry Link
Add an entry to the OPF Tool Registry, and provide a link to it here.

Evaluation
Any notes or links on how the solution performed.