Powered by OPF
OPF WIKI STATIC ARCHIVE
2,681 pages · 153 spaces · 776 tags · 4,025 history records · 96.2% of the original wiki recovered
Archived copy. This page was recovered from the Internet Archive snapshot of /display/SP/Data Processing Cluster taken on 2012-11-25. The original wiki at wiki.opf-labs.org is being decommissioned.

Data Processing Cluster

Added by Rainer Schmidt · last edited by Rainer Schmidt · on Nov 05, 2012 (view change)

About
The SCAPE Execution Platform provides a massively parallel environment for executing and orchestrating preservation tools and workflows. The SCAPE Central Instance (CI) provides a shared infrastructure for deploying the Platform software components as well as other project outcomes. Following shared infrastructures are presently being set up:

Shared Infrastructures
The Platform Concept Release (M14) comprises two shared data processing clusters hosted at IMF and AIT. Both clusters are based on Apache Hadoop. The setup of the IMF infrastructure hosting the Central Instance differs from the development infrastructure hosted by AIT. Most notably, the Central Instance at the IMF data center is hosted using low consumption nodes and no visualization layer is introduced. AIT provides a development cluster that is hosted on top of a virtualized private cloud environment. Both deployments presently support MapReduce, HDFS, HBase, and Hoop. Preservation tools can be installed on-demand. A Fedora Commons-based repository is presently being added to the AIT infrastructure.

Documentation

Software and Downloads
The SCAPE Platform basically comprises of (1) an integrated system that extend the execution platform with a number of software components developed within the projects, and (2) tools, scripts, and applications that enable users to execute preservation workflows on to of this environment.

Further Reading