Showing posts with label Hadoop. Show all posts
Showing posts with label Hadoop. Show all posts

Thursday, 17 October 2013


How to Learn Hadoop

             Big data is one the hottest area in IT, so everyone is wanting to learn Hadoop.   I am constantly being asked how to learn Hadoop.  So I want to share an approach I've been recommending and have gotten a lot of positive feedback on.  It's often hard to learn a new technology because a lot of the terminology, concepts and architecture approaches are new. Books, white papers, blogs usually try to teach Hadoop from a perspective of already knowing it.  These resources use terms, concepts and context that a newbie does not understand so it can make it very hard to learn a completely new subject.   So here is my recommendation for a way to learn Hadoop.

Learn the Basic Concepts First
            Everyone gets in a hurry to learn a new technology, so they are trying to learn all the tricks and fancy stuff right away and do not build a solid base first.

            Big data books.  Hadoop is all about the data.  Learn big data concepts before looking at Hadoop in any depth.  These books will build core data concepts around Hadoop.

  • Disruptive Possibilities: How Big Data Changes Everything,- This is a must read for anyone getting started in Big Data.
  • Big Data Now, 2012 Edition:  Easy read and good insights on Big Data.  Some of this content on companies is out of date, but there is a lot of valuable information here so this is still a must read.
  • Big Data, by Nathan Marx - The book does a great job of teaching core concepts, fundamentals and provides a great perspective of Big Data.  This book will build a solid foundation and helps you understand the Lambda architecture.  You may need to get this book from MEAP if it has not released yet. (http://www.manning.com/marz/)

               A good way to learn basic concepts and terminology before looking at technology in more depth.  Both are short reads.   By reading these basic books, they will gently introduce you so when you read the more technical books you will understand them better.

  • Hadoop for Dummies, by Tim Jones - Easy introduction to learn basic concepts and terms.   
  • Big Data for Dummies, - This is a very gentle introduction to Big Data, concepts and technologies surrounding it.

Three Defining Whitepapers to Read
           These papers are excellent papers to build fundamental knowledge around Hadoop and Hive.  Even though they are a few years old, the concepts and perspective discussed are excellent.  They will provide foundational insights into Hadoop.


Professional Training
            Professional training is the quickest and easiest way to learn core concepts, fundamentals and get some hands on experience.   I do work at Hortonworks, but there are some specific reasons I  recommend Hortonworks University.  The reason is Hortonworks is all open source so you are not learning someone's proprietary or open-proprietary distribution.   By learning from 100% open source at Hortonworks you can learn from the open source base, which is then applicable to any distribution.  Also, Hadoop 2 has YARN which is a key foundational component and Hortonworks is driving the innovation and roadmap around YARN.

Additional Resources
Once you get the fundamental concepts down you will be wanting to learn in more detail.  The two books below are good for taking that next step.  However, I recommend reading them in parallel and bouncing back and forth.  The reason is each has areas that I believe they do a better job on.  Each book has sections that I prefer and using them together was very helpful for me.

  • Apache Hadoop Yarn (not released yet), by Arun Murthy, Jeffrey Markham, Vinod Vavilapalli, Doug Eadline
  • Hadoop The Definitive Guide (3rd Edition), by Tom White
  • Hadoop Operations, by Eric Sammer

Getting Hands on Experience and Learning Hadoop in Detail
            A great way to start getting hands on experience and learning Hadoop through tutorials, videos and demonstrations is with the Hortonworks Sandbox.   The Hortonworks sandbox is designed for beginners, so it is an excellent platform for learning and skill development.   The tutorials, videos and demonstrations will be updated on a regular basis.   The sandbox is available in a Virtualbox or VMware virtual machine.  An additional 4GB of RAM and 2GB of storage is recommended for either of the virtual machines.  If you have a laptop that does not have a lot of memory you can go to the VM settings and cut the RAM for the VM down to about 1.5 - 2GB of RAM.  This is  likely to impact performance of the VM but it will help it at least run on a minimal configured laptop.

Other books to consider:

  • Programming Hive, by Edward Capriolo, Dean Wampler, ...
  • Programming Pig, by Alan Gates

Engineering Blogs:

  • http://engineering.linkedln.com/hadoop
  • http://engineering.twitter.com

Hadoop Ecosystem


What is Hadoop?


Data


Hadoop Distributions

Below are the companies offering commercial implementations and/or providing support for Apache Hadoop, which is the base for all the below.


  • Cloudera offers CDH (Cloudera's Distribution including Apache Hadoop) and Cloudera Enterprise.
  • Hortonworks (formed by Yahoo and Benchmark Capital), whose focus is on making Hadoop more robust and easier to install, manage and use for enterprise users. Hortonworks provides Hortonworks Data Platform (HDP).
  • MapR Technologies offers distributed filesystem and MapReduce engine, the MapR Distribution for Apache Hadoop.
  • Oracle announced the Big Data Appliance, which integrates Cloudera's Distribution Including Apache Hadoop (CDH).
  • IBM offers InfoSphere BigInsights based on Hadoop in both a basic and enterprise edition.
  • Greenplum, A Division of EMC, offers Hadoop in Community and Enterprise editions.
  • Intel - the Intel Distribution for Apache Hadoop is the product includes the Intel Manager for Apache Hadoop for managing a cluster.
  • Amazon Web Services - Amazon offers a version of Apache Hadoop on their EC2 infrastructure, sold as Amazon Elastic MapReduce.
  • VMware - Initiate Open Source project and product to enable easily and efficiently deploy and use Hadoop on virtual infrastructure.
  • Bigtop - project for the development of packaging and tests of the Apache Hadoop ecosystem.
  • DataStax - DataStax provides a product of Hadoop which fully integrates Apache Hadoop with Apache Cassandra and Apache Solr in its DataStax Enterprise platform.
  • Cascading - A popular feature-rich API for defining and executing complex and fault tolerant data processing workflows on a Apache Hadoop cluster. 
  • Mahout - Apache project using Hadoop to build scalable machine learning algorithms like canopy clustering, k-means and many more.
  • Cloudspace - uses Apache Hadoop to scale client and internal projects on Amazon's EC2 and bare metal architectures.
  • Datameer - Datameer Analytics Solution (DAS) is a Hadoop-based solution for big data analytics that includes data source integration, storage, an analytics engine and visualization.
  • Data Mine Lab - Developing solutions based on Hadoop, Mahout, HBase and Amazon Web Services.
  • Debian - A Debian package of Apache Hadoop is available.
  • HStreaming - offers real-time stream processing and continuous advanced analytics built into Hadoop, available as free community edition, enterprise edition, and cloud service.
  • Impetus
  • Karmasphere - Distributes Karmasphere Studio for Hadoop, which allows cross-version development and management of Apache Hadoop jobs.
  • Nutch - Apache Nutch, flexible web search engine software.
  • NGDATA - Makes available Lily Open Source that builds upon Hadoop, HBase and SOLR. Distributes Lily Enterprise.
  • Pentaho – Pentaho provides a complete, end-to-end open-source BI and offers an easy-to-use, graphical ETL tool that is integrated with Apache Hadoop for managing data and coordinating Hadoop related tasks in the broader context of ETL and Business Intelligence workflow.
  • Pervasive Software - Provides Pervasive DataRush, a parallel dataflow framework which improves performance of Apache Hadoop and MapReduce jobs by exploiting fine-grained parallelism on multicore servers.
  • Platform Computing - Provides an Enterprise Class MapReduce solution for Big Data Analytics with high scalability and fault tolerance. Platform MapReduce provides unique scheduling capabilities and its architecture is based on almost two decades of distributed computing research and development.
  • Sematext International - Provides consulting services around Apache Hadoop and Apache HBase, along with large-scale search using Apache Lucene, Apache Solr, and Elastic Search.
  • Talend - Talend Platform for Big Data includes support and management tools for all the major Apache Hadoop distributions. Talend Open Studio for Big Data is an Apache License Eclipse IDE, which provides a set of graphical components for HDFS, HBase, Pig, Sqoop and Hive.
  • Think Big Analytics - Offers expert consulting services specializing in Apache Hadoop, MapReduce and related data processing architectures.
  • Tresata - Financial Industry's first software platform architected from the ground up on Hadoop. Data storage, processing, analytics and visualization all done on Hadoop.
  • WANdisco is a committed member & sponsor of the Apache Software community and has active committers on several projects including Apache Hadoop

What is Hadoop


         Apache Hadoop is, an open-source software framework, written in Java, by Doug Cutting and Michael J. Cafarella, that supports data-intensive distributed applications, licensed under the Apache v2 license. It supports the running of applications on large clusters of commodity hardware. Hadoop was derived from Google's MapReduce and Google File System (GFS) papers.    

         The Hadoop framework transparently provides both reliability and data motion to applications. Hadoop implements a computational paradigm named MapReduce, where the application is divided into many small fragments of work, each of which may be executed or re-executed on any node in the cluster. It provides a distributed file system that stores data on the compute nodes, providing very high aggregate bandwidth across the cluster. Both map/reduce and the distributed file system are designed so that node failures are automatically handled by the framework. It enables applications to work with thousands of computation-independent computers and petabytes of data. The entire Apache Hadoop platform is commonly considered to consist of the Hadoop kernel, MapReduce and Hadoop Distributed File System (HDFS), and number of related projects including Apache Hive, Apache HBase, Apache Pig, Zookeeper etc.