Powered by OPF
OPF WIKI STATIC ARCHIVE
2,681 pages · 153 spaces · 776 tags · 4,025 history records · 96.2% of the original wiki recovered
Archived copy. This page was recovered from the Internet Archive snapshot of /display/SP/IS26 Dealing with difficult identification cases taken on 2012-11-27. The original wiki at wiki.opf-labs.org is being decommissioned.

IS26 Dealing with difficult identification cases

Created by Andrew Jackson on Oct 20, 2011 · last edited by Bjarne Andersen · on Sep 26, 2012 (view history) · 16 versions
Title
Dealing with difficult identification cases
Detailed description Identification Requirements, Format Languages, Requirements and Difficult Cases. Mutants and wild types. Strains. See below for specific examples.
Scalability Challenge
The solution must be able to identify and describe the large number of formats in our collections, and their complexity.
Issue champion Maureen Pennock (BL)
Other interested parties
Any other parties who are also interested in applying Issue Solutions to their Datasets. Identify the party with a link to their contact page on the SCAPE Sharepoint site, as well as identifying their institution in brackets. Eg: Schlarb Sven (ONB)
Possible Solution approaches We need to collect concrete examples of difficult format cases, ones that we need to identify for preservation purposes, but which current tools do not adequately describe. Hopefully we can generate a richer language for describing format, and ensure that it covers the cases we need. 

We can also compare the results from different identification systems to expose where the format is poorly understood or poorly described, and to drive improvements in format coverage.
Context We have a lot of stuff we can't identify, and a lot of stuff that is only identified at a coarse level.
Lessons Learned Notes on Lessons Learned from tackling this Issue that might be useful to inform the development of Future Additional Best Practices, Task 8 (SCAPE TU.WP.1 Dissemination and Promotion of Best Practices)
Training Needs Is there a need for providing training for the Solution(s) associated with this Issue? Notes added here will provide guidance to the SCAPE TU.WP.3 Sustainability WP.
Datasets All datasets! :-)
Solutions SO3 Comparing identification tools

Evaluation

Objectives Which scape objectives does this issues and a future solution relate to? e.g. scaleability, rubustness, reliability, coverage, preciseness, automation
Success criteria Describe the success criteria for solving this issue - what are you able to do? - what does the world look like?
Automatic measures What automated measures would you like the solution to give to evaluate the solution for this specific issue? which measures are important?
If possible specify very specific measures and your goal - e.g.
 * process 50 documents per second
 * handle 80Gb files without crashing
 * identify 99.5% of the content correctly
Manual assessment Apart from automated measures that you would like to get do you foresee any necessary manual assessment to evaluate the solution of this issue?
If possible specify measures and your goal - e.g.
 * Solution installable with basic linux system administration skills
 * User interface understandable by non developer curators
Actual evaluations links to acutual evaluations of this Issue/Scenario

Specific examples