Powered by OPF
OPF WIKI STATIC ARCHIVE
2,681 pages · 153 spaces · 776 tags · 4,025 history records · 96.2% of the original wiki recovered
Archived copy. This page was recovered from the Internet Archive snapshot of /display/SP/PC.CC Tasks and Ideas taken on 2016-09-11. The original wiki at wiki.opf-labs.org is being decommissioned.

PC.CC Tasks and Ideas

Created by Per Møldrup-Dalum on Mar 29, 2012 · last edited by Luis Faria · on May 14, 2012 (view history) · 18 versions

This page and its children are--for starters--rough and unstructured info. Consider this a drop box for issues and ideas

Tools we'll work with:
File formats:

Format naming schema:

We'll improve the tool evaluation framework. We'll also look into integrating it with REF. In the long run we'll have an automatic tool evaluator for on the fly evaluation of our tools.

We need to discuss and involve the Evaluation of Results work package.

In a year we will repeat the (automatic) tool evaluation of the at that time newest versions of DROID and FIDO to compare them to the present versions.

We need discussion regarding FITS and JHOVE2. What value can they add to this WP? We will evaluate the FITS output format for use for communication between sub projects.

File format -> application mapping (creation software? rendering software?)

How are we going to characterise composite objects?
How are we going to detect and characterise DRM?
Other stuff
A concrete goal to reach before June:
  1. Run identification/characterisation on a set of files
  2. Load the result into REF
  3. Compare the result with existing results in REF
Hadoop

Some personal notes on a basic Hadoop experiment using web archive meta data:

       Moved to and maintained at: Web Content Testbed - Next steps

Taverna
Regarding FITS or similar:

(just outlining some thoughts from PW, I will add more structured info soon, if you consider this helpful)

Web Content Characterization:

Tasks and intended checkpoints