Powered by OPF
OPF WIKI STATIC ARCHIVE
2,681 pages · 153 spaces · 776 tags · 4,025 history records · 96.2% of the original wiki recovered
Archived copy. This page was recovered from the Internet Archive snapshot of /display/REQ/Identifying web content taken on 2013-07-27. The original wiki at wiki.opf-labs.org is being decommissioned.

Identifying web content

Created by Louise Fauduet on Sep 27, 2011 · last edited by Paul Wheatley · on May 09, 2012 (view history) · 12 versions
Title
Identifying web content
Detailed description The Web archives team at BnF has long suspected that the MIME types declared by the server for their pages' content were not accurate, which is a problem for preservation and could impede future emulation, for instance.Thus, when expanding BnF's preservation system (SPAR) to ingest our Web archives collection, we had an ARC module developed for JHOVE 2 (http://bitbucket.org/lbihanic/jhove2-bnf _ this fork will be integrated to the general release of JHOVE2 in the coming year). This module produces a report of the characteristics of the ARC files, including the declared MIME types of the content files, and an identification of those same files using the FILE utility. When comparing the results during the initial tests of the web archives ingest process, we realized the results differed, especially when scripts and softwares were concerned.We would like to be able to evaluate the accuracy of those reports and correctly identify the content of our web archives. The problem is compounded by the huge size of the collection and the vast array of file formats present.
Issue champion Louise Fauduet
Other interested parties
Any other parties who are also interested in applying Issue Solutions to their Datasets
Possible Solution approaches Brief brainstorm of possible approaches to solving the Issue. Each approach should be described in a single sentence as part of a bulleted list
Context Bibliothèque nationale de France (National Library of France, BnF)
Lessons Learned Notes on Lessons Learned from tackling this Issue that might be useful to inform digital preservation best practice
Datasets French Web Archives
Solutions Server MIME Type Correction