3. Big Data Roadmap
Big Data – What?
◦ Timeline – Big Data Predictions
◦ Data Explosion
Big Data Myths
Big Data
5Vs of Big Data
Why Big Data
Data as Data Science
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
5. Timeline – Big Data
Predictions
1944- Yale Library in 2040 will have “approximately
200,000,000 Volumes
1961- Scientific Journals will grow exponentially rather than
linearly, doubling every fifteen years and increasing
by a factor of ten during every half-century.
1975- Ministry of Posts and Telecommunications in Japan
introduced words as unifying unit of measurement
1997- First article published by Michael Cox and David
Ellsworth in in the ACM digital library to the term
“Big data.”
Big Data evolved in 1997 and exploded to greater heights in
2010 and become popular in 2012
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
6. DATA EXPLOSION & ATTENTION
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
7. Big Data Explosion
12+ TBs
of tweet data
every day
25+ TBs
of
log data
every day
?TBsof
dataevery
day
2+
billion
people
on the
Web by
end 2011
30 billion RFID
tags today
(1.3B in 2005)
4.6
billion
camera
phones
world
wide
100s of
millions
of GPS
enabled
devices
sold
annually
76 million smart
meters in 2009…
200M by 2014
8. Data Growth – in Units
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
10. Desktop
Hobbyist
The Future?
Internet
Big Data
Byte : one grain of rice
Kilobyte : cup of rice
Megabyte : 8 bags of rice
Gigabyte : 3 Semi trucks with rice
Terabyte : 2 Container Ships
Petabyte : State full of rice bag
Exabytes : States filled with rice bag
Zettabyte : Fills the Pacific Ocean
Yottabyte : A EARTH SIZE RICE BALL!
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
11. BIG DATA FACTS
Every 2 days we create as much information
as we did from the beginning of time until
2003
Over 90% of all the data in the world was
created in the past 2 years.
It is expected that by 2020 the amount of
digital information in existence will have
grown from 3.2 zettabytes today to 40
zettabytes.
Every minute we send 204 million emails,
generate 1.8 million Facebook likes, send
278 thousand Tweets, and up-load 200,000Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
12. Big Data – Popularity
2010 – Social Networking – Post Internet
era – Facebook – Big Data
2012 – Election of US President – Hot
After 2012 – Industries – Hadoop
Big Data in India
Election process – BJP Govt
Current electioneering scenario
Election Campaign through social media
requires permission - Election
commission
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
13. BIG DATA GOT SO BIG ?
Big data is a commodity as compared
to Gold.
Big data is Hot!Now What?
Businesses Freak Out Over Big Data
– Information Week
Big Data Grows up – Forbes
2012 : The Year of Big
Big Data Powers Revolution in
Decision Making – Wall Street Journal
Business Opportunities in Big Data –
Inc
THE HINDU 2015
www.thehindu.com THE WORLD’S FAVOURITE NEWSPAPER - Since 1879
5
14. BIG DATA MYTHS
Big Data
• New
• Only About Massive Data Volume
• Means Hadoop
• Need A Data Warehouse
• Means Unstructured Data
• for Social Media & Sentiment
Analysis
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
15. What is Big Data?
Lets Us Clarify
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
16. Big Data
Big Data is
A complete subject with tools, techniques
and frameworks.
Technology which deals with large and
complex dataset which are varied in data
format and structures, does not fit into
the memory.
Not about huge volume of data; provide
an opportunity to find new insight into the
existing data and guidelines to capture
and analyze future data
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
17. Big Data : A Definition
Big data is the realization of greater
business intelligence by storing,
processing, and analyzing data that
was previously ignored due to the
limitations of traditional data
management technologies
:Source: Harness the Power of Big Data: The IBM Big Data Platform
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
18. BIG DATA as Platform
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
Source: IBM
20. 4 V‘s of Big Data
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
21. 6Vs of Big Data
Volume
Velocity
Variety
Veracity
Value
Validity
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
22. What is in Big Data ?
Why Big Data Analytics?
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
23. Big Data
Exploration
Find, visualize,
understand all big
data to improve
decision making
Enhanced 360o View
of the Customer
Extend existing customer
views (MDM, CRM, etc) by
incorporating additional
internal and external
information sources
Security/Intelligence
Extension
Lower risk, detect fraud
and monitor cyber security
in real-time
Data Warehouse Augmentation
Integrate big data and data warehouse
capabilities to increase operational
efficiency
Operations Analysis
Analyze a variety of machine
data for improved business results
The 5 Key Big Data Use Cases
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
27. Conventional approaches
RDBMS
OS FILE SYSTEM
SQL QUERIES
CUSTOM FRAMEWORK
* C / C++
* PERL
* PYTHON
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
28. ISSUES IN LEGACY SYSTEMS
Limited Storage Capacity
Limited Processing Capacity
No Scalability
Single point of Failure
Sequential Processing
RDBMSs can handle Structured Data
Requires preprocessing of Data
Information is collected according to
current business needs
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
25
29. Mr. HADOOP says he has a solution to
our BIG problem !
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
27
30. Mr. HADOOP says he has a solution
to our BIG problem !
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
27
31. What is
Apache Hadoop Is A Framework That Allows For The
Distributed Processing Of Large Datasets Across Clusters Of
Commodity Computers Using A Simple Programming Model.
Concept
Moving computation is more efficient than moving
large data
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
31
37. WHAT IS HDFS
Hadoop Distributed File System
Highly Fault tolerant , distributed , reliable , scalable
file system for data storage.
Stores multiple copies of data on different nodes
A File is split up into blocks and stored on multiple
machines
Hadoop cluster typically has a single namenode and no.
of data nodes to form a hadoop cluster.
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
38. HDFS BLOCKS
• Files are broken in to large blocks.
Typically 128 MB block size
Blocks are replicated for reliability
One replica on local node
Another replica on a remote rack
Third replica on local rack,
Additional replicas are randomly placed
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
41
39. HDFS BLOCKS contd.,
ADVANTAGES OF HDFS BLOCKS
Fixed Size
Chunk of file < block size : Only needed space
is used.
Eg : 420 MB file is split as
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
42
58. Potential Talent Pool -Big
Data
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
India will require a minimum of 1 lakh data scientists in the next couple
of years in addition to data analysts and data managers to support the
Big Data space.
63. SBI
State Bank of India (SBI) ran its newly
acquired data-mining software recently to
check for purity of data.
Made an interesting find - close to one crore
accountholders have not provided any
nomination for their savings accounts. What
is worse, over half of them are senior
citizens.
To analyse trends in Banks, SBI has hired a
whole team of statisticians and economists.
Identify default patterns, high value
customers.
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
64. Big Data Applications - India
Big Data – Elections
SBI uses big data mining to check
defaults
Karnataka Govt – Identify water
leakage
Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-
65. Big Data Challenges
Privacy Protection
All Big data stages collect, store, process,
knowledge
Integration with enterprise landscape
All systems store data in rdbms,DW
Does not support bulk loading to Big data store
Limited number of analytics from Mahout
Big data technologies lack visualization support
and deliverable methods
Leveraging cloud computing for big data applications
Addressing Real time needs with varied format
and volume Dr.V.Bhuvaneswari, Asst.Professor, Dept. of Comp. Appll., Bharathiar University,-