Powered by OPF
OPF WIKI STATIC ARCHIVE
2,681 pages · 153 spaces · 776 tags · 4,025 history records · 96.2% of the original wiki recovered
Archived copy. This page was recovered from the Internet Archive snapshot of /display/SPR/Identification of file formats with incorrect file extensions taken on 2013-07-27. The original wiki at wiki.opf-labs.org is being decommissioned.

Identification of file formats with incorrect file extensions

Added by Hannah Green · last edited by Paul Wheatley · on May 08, 2012 (view change)
Title Identification of file formats with incorrect file extensions
Detailed description Electronic documents and image files are deposited on a variety of media, including floppy disc, CD-R and memory sticks.  In copying process, file extensions can be lost, or period marks in file names result in everything after period mark being read as a file extension, resulting in unreadable files because correct file association has been lost. Time consuming to identify where unreadable file extensions are genuine, but unusual file types, or are incorrect extensions.  Time consuming, hit-and-miss process currently used to try and identify file types and access content, with potential impact on authenticity & integrity of files in the process.
Issue champion Hannah Green
Other interested parties Rebecca Nielsen 
Richard Freeston
Possible Solution approaches
  • use Apache Tika to identify file types & extract metadata
  • develop script which will run over directory of files at once, allowing for quick identification of large quantity of files
Context Details of the institutional context to the Issue. (May be expanded at a later date)
Lessons Learned Notes on Lessons Learned from tackling this Issue that might be useful to inform digital preservation best practice
Datasets Seven Stories author & illustrator files
Solutions Tika Batch File Identification