Powered by OPF
OPF WIKI STATIC ARCHIVE
2,681 pages · 153 spaces · 776 tags · 4,025 history records · 96.2% of the original wiki recovered
Archived copy. This page was recovered from the Internet Archive snapshot of /display/AQuA/Use of OCR metadata taken on 2013-07-27. The original wiki at wiki.opf-labs.org is being decommissioned.

Use of OCR metadata

Created by Toby Atkin-Wright on Jun 13, 2011 · last edited by Toby Atkin-Wright · on Jun 15, 2011 (view history) · 4 versions
One line summary How can we use OCR metadata to identify pages for human QC investigation?
Detailed description The ABBYY FineReader 9 engine outputs various OCR stats, and these are expressed in the ALTO files. For each page there is a predicted word accuracy percentage, a suspicious character count, a word count, and a suspicious word count. For each OCRed word, there is also a word confidence (0 to 1, where 1 is good) and a character confidence (0 to 9, where 0 is good).
Issue champion Toby Atkin-Wright
Possible approaches Currently the Brightsolid project makes use of the predicted word accuracy (PWA) for each page, calculates the mean PWA across each year of each newspaper, and calculatse the median absolute deviation (MAD). It then marks for manual investigation all pages that have a predicted word accuracy less than (mean - 3x MAD). However, most of these pages are fine, and the variations in OCR quality can be explained by content changes or physical page damage. Is there a better way to use the OCR metadata to find pages that may have questionable scan quality?
Context  
AQuA Solutions  
Collections Brightsolid digitisation of British Library newspapers