Powered by OPF
OPF WIKI STATIC ARCHIVE
2,681 pages · 153 spaces · 776 tags · 4,025 history records · 96.2% of the original wiki recovered
Archived copy. This page was recovered from the Internet Archive snapshot of /pages/viewpage.action?pageId=16713846 taken on 2021-10-23. The original wiki at wiki.opf-labs.org is being decommissioned.

Extracting and aggregating metadata with Apache Tika

Added by Thom Carter · last edited by Thom Carter · on Sep 20, 2012

Extracting and aggregating metadata with Tika

Apache Tika was used with a custom wrapper to extract metadata (e.g. author, title, extent, dates and file formats) and content (text) from files in two large digital archive collections. A Java script was then used to produce a report that summarised the metadata and content across the collection. This information will be used to inform collection management decisions and identify potential preservation issues.

Solution Champion
Thom Carter, Rebecca Webster

Corresponding Issue(s)
Produce a report summarising collection metadata and content
Sorting, appraising and metadata creation for deposited personal collections

Tool/code link
[Link to Pete's code]

Tool Registry Link
Apache Tika

Evaluation
Any notes or links on how the solution performed.