Powered by OPF
OPF WIKI STATIC ARCHIVE
2,681 pages · 153 spaces · 776 tags · 4,025 history records · 96.2% of the original wiki recovered
Archived copy. This page was recovered from the Internet Archive snapshot of /display/SP/File Format Identification and Characterisation of Web Archives taken on 2016-09-11. The original wiki at wiki.opf-labs.org is being decommissioned.

File Format Identification and Characterisation of Web Archives

Created by Peter Cliff on Feb 07, 2013 · last edited by William Palmer · on Nov 06, 2013 (view history) · 9 versions

Status

Active

This story is associated with a success story: http://wiki.opf-labs.org/display/SP/QA+and+Characterisation+of+Web+Content

Contact

Per Møldrup-Dalum, SB, [email protected]

William Palmer, BL (william (.) palmer (@)) bl (.) uk)

User Story

As a Web Archive I need a Digital Preservation System that can process both ARC and WARC files and identify file formats/characterize of items contained so that I can assess preservation risks and plan which tools will be required for access to those formats.

User Requirements/Components

  1. A tool that can efficiently work through the content of an ARC file and identify the type of files found.
    1. Must provide a report in a usable format - where this is a large dataset, this could be a database of some sort
    2. Ideally the tool will also perform file format identification on files within container formats - media streams, zips and other compressed files, etc.
  2. Look up of an appropriate access tool/software would be a bonus!

Experiments

Create experiments as child pages and they should appear automatically here

Developer Notes

Space for discussion, suggested solutions, links to other scenarios, etc.

Related Documents

Scenarios, case studies, etc. that provide background to this story.