Powered by OPF
OPF WIKI STATIC ARCHIVE
2,681 pages · 153 spaces · 776 tags · 4,025 history records · 96.2% of the original wiki recovered
Archived copy. This page was recovered from the Internet Archive snapshot of /display/KB/Identifying Aggregations of Duplicates in a Dataset taken on 2013-06-21. The original wiki at wiki.opf-labs.org is being decommissioned.

Identifying Aggregations of Duplicates in a Dataset

Added by Michael Olson · last edited by Seth Shaw · on Jun 05, 2013 (view change)

Title
Identifying Aggregations of Duplicates in a Dataset

Detailed description
I am an archivist and my collection consists of multiple computers or hand held media.  I would like to determine if there are duplicates and the degree of duplication across the media in the collection.  The degree of duplication is important for the following reasons:

Note: File deduplication is a well known space-saving technique in production by many storage tools and as stand-alone utilities.This case differs in that we are interested in identifying clusters of duplicated files either for appraisal, prioritization, or identifying relationships between groups of materials. No deduplication tools we identified visualized the locations and prevalence of duplication. The visualization is the heart of this use case.

Issue champions

Heather Gendron
Seth Shaw
Meg Tuomala
Michael Olson

Other interested parties
Digital Archivists, researchers, curatorial staff negotiating acquisitions

Possible Solution approaches

Possible Treemap tools:

Draft base workflow:

  1. Generate checksum list of files
  2. Process checksum list into JSON file-tree structure with node variables:
    • name
    • file-count
    • dup-count
    • dup-locations[]
  3. load into visualization

Context

Will allow archivists and collection curators to determine the degree of duplication within and across datasets (computers, hand held media) and generate reports on duplication

Visualization of the duplication across the dataset will be useful for an archivist in determining how they apply their limited processing resources

Researchers might be interested in the degree of duplication.

Lessons Learned
Notes on Lessons Learned from tackling this Issue that might be useful to inform digital preservation best practice

Datasets
Environmental Artists Datasets

Program on Public Life administrative records and director email

Solutions
CSV listing of Aggregations of Duplicates in a Dataset