Powered by OPF
OPF WIKI STATIC ARCHIVE
2,681 pages · 153 spaces · 776 tags · 4,025 history records · 96.2% of the original wiki recovered
Archived copy. This page was recovered from the Internet Archive snapshot of /display/SP/SCAPEdev1 Workshop - Experimenting with Hadoop taken on 2016-09-12. The original wiki at wiki.opf-labs.org is being decommissioned.

SCAPEdev1 Workshop - Experimenting with Hadoop

Added by Andrew Jackson · last edited by Andrew Jackson · on Jun 06, 2011 (view change)

Introduction

First we watched the first half of Mapreduce & HDFS (PDF Slides)

See also:

The exercises use The Cloudera Hadoop Demo VM and The Cloudera Training Material.

Exercise 1 - Getting Familiar with Hadoop

Exercise 2 - Running a Map Reduce Job

Exercise 3 - Advanced notions

Here's a few ideas for more advanced things to do.

So why do we need HBase?

We need something like HBase because HDFS does not cope well with lots of 'small' files (due to the HDFS block size). See http://www.cloudera.com/blog/2009/02/the-small-files-problem/ for information and some alternative solutions.