Powered by OPF
OPF WIKI STATIC ARCHIVE
2,681 pages · 153 spaces · 776 tags · 4,025 history records · 96.2% of the original wiki recovered
Archived copy. This page was recovered from the Internet Archive snapshot of /display/SP/Detecting duplicates on large collections of digitized book pages taken on 2016-09-11. The original wiki at wiki.opf-labs.org is being decommissioned.

Detecting duplicates on large collections of digitized book pages

Created by Miguel Ferreira on Feb 27, 2013 · last edited by Miguel Ferreira · on Feb 27, 2013 (view history) · 8 versions

Detecting duplicates on large collections of digitized book pages

Context

Europe has invested millions of euros in mass digitization projects. Such projects are error prone in the sense that sometimes book pages are digitized more than once (due to the particularities of the digitization process) or entire books appear duplicated because they belong to multiple book collections at the same time.

Duplicates are not exact copies of each other as they result from distinct digitization processes. Images may be skewed, different color tones, rotated, cropped, etc. 

Additionally, post processing jobs on the digitized images produce results that are do not comply to the quality standards of the collection owner or eliminate images in the process.

This means that traditional duplicate image detection techniques will not work on these scenarios. OCR based techniques will not work either because books are sometimes handwritten or written in a deprecated language.

Manual approaches tend to be imprecise and time consuming so a tool that automates this process in an accurate way is necessary.

SCAPE has developed MatchBox, which uses an innovative approach to address this problem. MatchBox applies state of the art image processing approaches to the domain of digital preservation and quality assurance. Based on computer vision algorithms selecting key characteristics of the images and extracting information from their content, large digitized collections can be processed in a scalable way. 

Preservation issues

Solution