Powered by OPF
OPF WIKI STATIC ARCHIVE
2,681 pages · 153 spaces · 776 tags · 4,025 history records · 96.2% of the original wiki recovered
Archived copy. This page was recovered from the Internet Archive snapshot of /display/AQuA/Finding duplicate images taken on 2013-03-25. The original wiki at wiki.opf-labs.org is being decommissioned.

Finding duplicate images

Created by Toby Atkin-Wright on Jun 13, 2011 · last edited by Julie Allinson · on Jul 27, 2011 (view history) · 5 versions
One line summary How do we ensure that duplicate data is not archived?
Detailed description Duplicate images and data can exists for various reasons. Images may be scanned twice, may be duplicated inadvertently during processing, or the original archive may include duplicate documents. How do we weed these out of our digital archive?
Issue champion Toby Atkin-Wright
Possible approaches Currently the Brightsolid project ensures that each issue date for each newspaper is unique, so if the metadata is correct, there should be no duplicates.
It also checks that each of the delivered JP2 and ALTO files have a unique SHA256 fingerprint.
Suggested enhancements include using fuzzy OCR to compare page content, and match any pages that appear to have similar content. This could be applied just to headlines throughout the newspaper issues, as these are higher quality data. (The headlines are all manually QCed after OCR, so are the best quality data in the pages.)
Context  
AQuA Solutions Perceptual Image Diff comparison
java image blocks comparison
ssdeep for duplicate image detection
Collections Brightsolid digitisation of British Library newspapers