Powered by OPF
OPF WIKI STATIC ARCHIVE
2,681 pages · 153 spaces · 776 tags · 4,025 history records · 96.2% of the original wiki recovered
Archived copy. This page was recovered from the Internet Archive snapshot of /display/AQuA/Using METS data to inform analysis taken on 2013-07-27. The original wiki at wiki.opf-labs.org is being decommissioned.

Using METS data to inform analysis

Added by Toby Atkin-Wright · last edited by Toby Atkin-Wright · on Jun 15, 2011 (view change)
One line summary Can we use metadata in METS files to help us target QC analysis of the OCRed text?
Detailed description METS files describe structure of documents, listing the pages (with links to their ALTO and image files), and showing what type of data is included in each page. Examples of data types could be headlines, articles, illustrations, family notices, and adverts.
Can we use this structure to target our QC analysis of the OCR text?
Issue champion Toby Atkin-Wright
Possible approaches Perform statistical analysis of the article text in each issue, ignoring other content types. The article text will better match expected English usage than other text on the page.
Context  
AQuA Solutions  
Collections Brightsolid digitisation of British Library newspapers