SlideShare una empresa de Scribd logo
1 de 11
Descargar para leer sin conexión
Strategic  Advisory
Big  Data  – Cloud   -­‐ Analytics
Info
Strategy
Fishing  in  the  
big  data  lake
DATA  EXPLORATION  AND  DISCOVERY  ANALYTICS  
FOR  DEEPER  BUSINESS  INSIGHTS
InfoStrategy
What  is  a  “data  lake”
data  lake (plural data  lakes)
A  massive,  easily  accessible  data  repository  
built  on  (relatively)  inexpensive  computer  
hardware  for  storing  "big  data".  Unlike  data  marts,  
which  are  optimized  for  data  analysis  by  storing  only  some  
attributes  and  dropping  data  below  the  level  aggregation,  a  
data  lake  is  designed  to  retain  all  attributes,  
especially  so  when  you  do  not  yet  know  what  the  
scope  of  data  or  its  use  will  be.
http://en.wiktionary.org/wiki/data_lake
…  Enterprise  Data  Hub  sounds  too  boring   !
InfoStrategy
Optimise  business  through  insights
Insight
Action
Optimise
Move  a  metric
Change  a  product
Change  behaviour/process
Hindsight
Realtime
Foresight
Trusted  information
Act  on  insights  gained
Execute  theories
Measure
Outcomes
Sentiment
Feedback
Explore  datasets,  discover  correlations,  patterns.
Undiscovered  facts
Information  Value
Data  Volumes
Forecasting,  planning  &  trending
Statistical  Analysis
Operational  reporting,  SCADA  control
Alerts  &  Events
Historical  reporting, Proof  of  operation
Regulatory,  statutory,  financial
Uncover  previously  
unknown  facts  
from  enriched  data  
in  the  data  lake
InfoStrategy
Future  state  of  analytics
Strategic  Intent
To  improve  BI  and  Analytical  capabilities  to  a  level  where  organisations  are  able  to  
access  and  analyse  information  in  a  secure,  timely  and  cost-­‐effective  manner.
Gain  key  insights  to  optimise  the  operations  of  your  business,  predict  the  best  
possible  outcomes  for  growth,  new  opportunities,   and  competitive  advantage  
across  all  business  lines.
Mission  Statement
“Providing  advanced  analytics  capability  across  all  business  units,  empowering  our  
people  with  the    processes  and  supporting  technologies  to  exploit  our  information  
assets  for  business  benefit.”
Target  Operating  Model  will  deliver:
Rapid  access  to  data  to  uncover  new  facts  via  advanced  data  exploration  and  
discovery  analytics.
Clarity  of  who  is  responsible  and  accountable  for  maintaining  critical  information  
assets  via  a  well  structured  governance  and  engagement  model.
A  trusted  and  highly  secure  source  of  data  for  all  analytical  information  requirements  
via  a  data  quality  assurance  program.
Trawling  for  value  in  the  big  data  lake
InfoStrategy
‘Fish  stocks’  are  replenished  from  existing  and  future  
operational  systems  plus  external  sources
Core  
Transactional  Data  
“operational”
Management  
Reporting
Unstructured  &  
External  Data
“contextual”
Enterprise  Dashboards
Reporting
Consolidation
Data  ScientistsBusiness  AnalystsBusiness  UsersCustomers
Data  Extraction
Discovery  Analytics  
Platform
Visualisation
Analysis
Data  Preparation
Data  Collection
Operational  
Reporting
Operational  Dashboards
Real-­‐time  Reports
Alerts  &  Exceptions
Embedded  BI
Production   Data  Repository
“Data  Lake”
Information  Governance
Data  Management
Supplier  &  
Industry  Data
“comparative”
InfoStrategy
Consolidated
Management
Reporting
Operational
Supporting
Capability
Discovery
Analytics
To  meet  the  demand  for  rapid  access  to  information  
users  must  adopt  a  flexible  multi-­‐platform   architecture  
What  reporting  does  for  established  operations  …  discovery  analytics  does  for  new  business  development.
The  trend  within  industry  is  to  move  away  from  the  single-­‐platform  monolithic  data  warehouses  towards  a  physically  distributed  environment  
for  information  delivery.  Many  businesses  are  extending  their  data  warehouse  environments  to  include  new  standalone  data  platforms  that  
are  conducive  to  discovery  analytics.  A  holistic  view  is  maintained  via  a  common,  single  replicated  dataset  and  an  enterprise information  
management  program,  governing  delivery  and  access  to  key  information  (data  lake).
Source   Applications
ERP
CRM
HR
Finance
Telemetry
Geospatial  GIS
Documents
Email
Files
Real-­time  Data  
Capture
Cleansing
Loading
Data  Warehouse
Modelling
Relational  DW
Data  Marts
Analysis  Cubes
Analytics Delivery
Cloud-­based    Service  Model
Actuarial  
Applications
Event-­Based  
Applications
Reporting
Production  
Reporting
OLAP  Analytics
Ad  Hoc  Query
External
Data
Exploration  &  
Discovery
Metadata  Integration
Event  Processing Results
Detailed  Datasets Results  
Collection  and  blending Insights
Portal
PDF
Desktop
Guided  
Visualisation
Mobile  BI
Active  
Dashboards
Data  Replication
Historical Data  Preparation
Storytelling
Information  Governance
Operational  Reporting  
Dimensional  
Modelling
ProductioniseInsights
InfoStrategy
Principles:  Easier  access  information   to  discover  new  
facts  about  the  business.
◦ Described  as  a  ‘sandpit’  environment,  providing  the  ability  to  explore  and  discover  new  
facts  about  the  business,  it’s  members  and  customers,  partners  and  competitive  
pressures.
◦ Also  used  for  testing  a  hypothesis  or  running  scenarios  across  the  data
◦ Getting  answers  to  ‘one-­‐off’  questions  which  are  not  addressed  through  the  normal  
published,  scheduled  operational  reporting  channels
◦ Data  is  replicated  from  all  operational  systems  into  a  single  landing  area,  ensuring  
traceability  and  reconciliation  to  all  consuming  applications,  such  as  the  data  warehouse,  
analytical  application,  and  other  business  applications.
◦ Clearly  defined  critical  business  entities/records  are  synchronised  (or  Mastered)  across  
all  applications  eliminating  duplication  and  confusion.  Data  quality  attributes  are  defined  
and  managed  for  each  critical  business  entity.
◦ A  fully  integrated  Member/Customer  view  is  established  across  both  analytical  and  
transactional  applications.
◦ Using  the  replicated  data  to  build  more  dynamic  analytical  data  structures  for  scheduled  
production  reporting  and  ah-­‐hoc  analysis
◦ Provide  users  with  the  tools  to  access    and  analyse data,  freely  explore  current  and  new  
datasets,  and  visualise patterns  and  discoveries  to  gain  deep  insights.
Providing  business  users  with  direct  
access  to  data  to  meet  immediate  
information  needs  where  the  
accuracy  of  the  data  is  not  the  
primary  objective.  
Having  a  single  source  of  truth  
across  all  business  applications  at  
detailed  level  from  which  all  
information  requests  are  satisfied.
Improved  environment  for  more  
cost  effective  and  faster  business  
intelligence  delivery.
Provide  business   users  with  the  ability  to  access  production  information  directly,  collect  it  as  needed,  and  
prepare  the  data  for  analysis.  Exploring  the  data  to  uncover  previously   unknown  facts  about  the  business,   and  
sharing  those  facts  visually  with  others.  Enrich  production  data  with  external  “context”  to  extend  insights.
Key  Principles Description
InfoStrategy
Benefits  of  Discovery  Analytics  versus  traditional   data  
warehousing
Classic  Data  Warehouse  Issues Discovery  Analytics Benefit
Lengthy  IT  Backlog  and  lack  of  resources  to  extend the  
EDW  to  support  new  business  requirements.
Data  can  be  explored  and  analysed  outside  of the  EDW  
environment  before  it  is  put  into  production  use.
High  costs  of  supporting increasing  data  volumes  and  
new  types  of  data.
Data  can  be  filtered  and  transformed  before  it  is  loaded  
into  the  EDW
Lack  of  flexibility  in  the  EDW  data  model  to  support  
constantly changing  business  requirements.
Data  discovery  support  dynamic  schema  on  read  
approach  which reduces  the  need  for  detailed  up-­‐front  
modelling.
Need  to  have  data  quality  and  governance  processes  in  
place  before  user  can  access  the  EDW  data.
The  investigative  nature  of data  discovery  has  lower  data  
quality  and  governance  requirements
Growing  use  of  personal  data  marts to  overcome  IT  
barriers  and  the  performance  overheads  of  ad  hoc  
processing
The  flexibility  and  performance  of  data  discovery  
encourages  shared  use  of  data  and  analytics.
Recent  proof  of  concept  for  Discovery  Analytics  in  the  cloud  (AWS),  has  provided  some  
considerable  cost  &  time  savings  in  infrastructure  and  hosting,  viz.:
$55  per  day  to  host  a  960GB  data  warehouse  
$32  per  day  to  host  a  Data  Integration  server  AND  a  BI  server.
2.5  weeks  to  setup  POC  environment  and  start  analysis  and  visualising  results.
InfoStrategy
Discovery  Analytics  Target  POC  Architecture
Structured  
Data
Unstructured  
Data
ERP
Telemetry
Web/External
Replication  of  corporate  data,  enriched  with  external  data  and  
content,  available  in  a  centrally  available  and  scalable  repository  
ready  for  exploration,  discovery  and  predictive  analysis  to  gain  
deep  insights  and  actionable  results.
InfoStrategy
Fishing  safely  with  the  appropriate  life  vests  is  
important  too.
Security  and  data  management  standards  are  available
International  
Standard  on  
Assurance  
Engagements
Service  Organisation  
Control  framework
Federal  Information  
Management  
Security  Act
Payment  Card  
Industry  –Data  
Security  Standard
Federal  Information  
Processing  Standard
International  Standards  
Organisation  –
Information  Security  
Standard
Source:  Amazon  Web  Services
Info
Strategy
To  learn  more  about  how  InfoStrategy
can  help  you  develop  your  big  data  
strategy  to  solve  your  big  business  
problems,  or  to  arrange  a  Proof  of  
Concept,  please  contact  us  today  using  
the  details  below.
InfoStrategy Pty  Ltd
246  Oxford  St,  Balmoral
Queensland  4171
Australia
Tel:  +61  7  3151  2021
Email:  
contactus@infostrategy.com.au

Más contenido relacionado

La actualidad más candente

Introduction to Data Engineering
Introduction to Data EngineeringIntroduction to Data Engineering
Introduction to Data EngineeringHadi Fadlallah
 
Date warehousing concepts
Date warehousing conceptsDate warehousing concepts
Date warehousing conceptspcherukumalla
 
Data Lakehouse, Data Mesh, and Data Fabric (r1)
Data Lakehouse, Data Mesh, and Data Fabric (r1)Data Lakehouse, Data Mesh, and Data Fabric (r1)
Data Lakehouse, Data Mesh, and Data Fabric (r1)James Serra
 
Data Lakehouse, Data Mesh, and Data Fabric (r2)
Data Lakehouse, Data Mesh, and Data Fabric (r2)Data Lakehouse, Data Mesh, and Data Fabric (r2)
Data Lakehouse, Data Mesh, and Data Fabric (r2)James Serra
 
Demystifying data engineering
Demystifying data engineeringDemystifying data engineering
Demystifying data engineeringThang Bui (Bob)
 
Building Lakehouses on Delta Lake with SQL Analytics Primer
Building Lakehouses on Delta Lake with SQL Analytics PrimerBuilding Lakehouses on Delta Lake with SQL Analytics Primer
Building Lakehouses on Delta Lake with SQL Analytics PrimerDatabricks
 
Achieving Lakehouse Models with Spark 3.0
Achieving Lakehouse Models with Spark 3.0Achieving Lakehouse Models with Spark 3.0
Achieving Lakehouse Models with Spark 3.0Databricks
 
Business Intelligence (BI) and Data Management Basics
Business Intelligence (BI) and Data Management  Basics Business Intelligence (BI) and Data Management  Basics
Business Intelligence (BI) and Data Management Basics amorshed
 
Building an Effective Data Warehouse Architecture
Building an Effective Data Warehouse ArchitectureBuilding an Effective Data Warehouse Architecture
Building an Effective Data Warehouse ArchitectureJames Serra
 
Building Modern Data Platform with Microsoft Azure
Building Modern Data Platform with Microsoft AzureBuilding Modern Data Platform with Microsoft Azure
Building Modern Data Platform with Microsoft AzureDmitry Anoshin
 
DW Migration Webinar-March 2022.pptx
DW Migration Webinar-March 2022.pptxDW Migration Webinar-March 2022.pptx
DW Migration Webinar-March 2022.pptxDatabricks
 
Building a modern data warehouse
Building a modern data warehouseBuilding a modern data warehouse
Building a modern data warehouseJames Serra
 
Pipelines and Data Flows: Introduction to Data Integration in Azure Synapse A...
Pipelines and Data Flows: Introduction to Data Integration in Azure Synapse A...Pipelines and Data Flows: Introduction to Data Integration in Azure Synapse A...
Pipelines and Data Flows: Introduction to Data Integration in Azure Synapse A...Cathrine Wilhelmsen
 
Databricks + Snowflake: Catalyzing Data and AI Initiatives
Databricks + Snowflake: Catalyzing Data and AI InitiativesDatabricks + Snowflake: Catalyzing Data and AI Initiatives
Databricks + Snowflake: Catalyzing Data and AI InitiativesDatabricks
 
Owning Your Own (Data) Lake House
Owning Your Own (Data) Lake HouseOwning Your Own (Data) Lake House
Owning Your Own (Data) Lake HouseData Con LA
 
Modernizing to a Cloud Data Architecture
Modernizing to a Cloud Data ArchitectureModernizing to a Cloud Data Architecture
Modernizing to a Cloud Data ArchitectureDatabricks
 
Introducing Databricks Delta
Introducing Databricks DeltaIntroducing Databricks Delta
Introducing Databricks DeltaDatabricks
 

La actualidad más candente (20)

Introduction to Data Engineering
Introduction to Data EngineeringIntroduction to Data Engineering
Introduction to Data Engineering
 
Date warehousing concepts
Date warehousing conceptsDate warehousing concepts
Date warehousing concepts
 
Data Lakehouse, Data Mesh, and Data Fabric (r1)
Data Lakehouse, Data Mesh, and Data Fabric (r1)Data Lakehouse, Data Mesh, and Data Fabric (r1)
Data Lakehouse, Data Mesh, and Data Fabric (r1)
 
Data Lake,beyond the Data Warehouse
Data Lake,beyond the Data WarehouseData Lake,beyond the Data Warehouse
Data Lake,beyond the Data Warehouse
 
Data Lakehouse, Data Mesh, and Data Fabric (r2)
Data Lakehouse, Data Mesh, and Data Fabric (r2)Data Lakehouse, Data Mesh, and Data Fabric (r2)
Data Lakehouse, Data Mesh, and Data Fabric (r2)
 
Lakehouse in Azure
Lakehouse in AzureLakehouse in Azure
Lakehouse in Azure
 
Demystifying data engineering
Demystifying data engineeringDemystifying data engineering
Demystifying data engineering
 
Building Lakehouses on Delta Lake with SQL Analytics Primer
Building Lakehouses on Delta Lake with SQL Analytics PrimerBuilding Lakehouses on Delta Lake with SQL Analytics Primer
Building Lakehouses on Delta Lake with SQL Analytics Primer
 
Achieving Lakehouse Models with Spark 3.0
Achieving Lakehouse Models with Spark 3.0Achieving Lakehouse Models with Spark 3.0
Achieving Lakehouse Models with Spark 3.0
 
Business Intelligence (BI) and Data Management Basics
Business Intelligence (BI) and Data Management  Basics Business Intelligence (BI) and Data Management  Basics
Business Intelligence (BI) and Data Management Basics
 
Building an Effective Data Warehouse Architecture
Building an Effective Data Warehouse ArchitectureBuilding an Effective Data Warehouse Architecture
Building an Effective Data Warehouse Architecture
 
Building Modern Data Platform with Microsoft Azure
Building Modern Data Platform with Microsoft AzureBuilding Modern Data Platform with Microsoft Azure
Building Modern Data Platform with Microsoft Azure
 
DW Migration Webinar-March 2022.pptx
DW Migration Webinar-March 2022.pptxDW Migration Webinar-March 2022.pptx
DW Migration Webinar-March 2022.pptx
 
Building a modern data warehouse
Building a modern data warehouseBuilding a modern data warehouse
Building a modern data warehouse
 
Pipelines and Data Flows: Introduction to Data Integration in Azure Synapse A...
Pipelines and Data Flows: Introduction to Data Integration in Azure Synapse A...Pipelines and Data Flows: Introduction to Data Integration in Azure Synapse A...
Pipelines and Data Flows: Introduction to Data Integration in Azure Synapse A...
 
Databricks + Snowflake: Catalyzing Data and AI Initiatives
Databricks + Snowflake: Catalyzing Data and AI InitiativesDatabricks + Snowflake: Catalyzing Data and AI Initiatives
Databricks + Snowflake: Catalyzing Data and AI Initiatives
 
Owning Your Own (Data) Lake House
Owning Your Own (Data) Lake HouseOwning Your Own (Data) Lake House
Owning Your Own (Data) Lake House
 
Modernizing to a Cloud Data Architecture
Modernizing to a Cloud Data ArchitectureModernizing to a Cloud Data Architecture
Modernizing to a Cloud Data Architecture
 
Introducing Databricks Delta
Introducing Databricks DeltaIntroducing Databricks Delta
Introducing Databricks Delta
 
Data Engineering Basics
Data Engineering BasicsData Engineering Basics
Data Engineering Basics
 

Similar a Data lake benefits

intelligent-data-lake_executive-brief
intelligent-data-lake_executive-briefintelligent-data-lake_executive-brief
intelligent-data-lake_executive-briefLindy-Anne Botha
 
What Data Do You Have and Where is It?
What Data Do You Have and Where is It? What Data Do You Have and Where is It?
What Data Do You Have and Where is It? Caserta
 
BI Masterclass slides (Reference Architecture v3)
BI Masterclass slides (Reference Architecture v3)BI Masterclass slides (Reference Architecture v3)
BI Masterclass slides (Reference Architecture v3)Syaifuddin Ismail
 
Setting Up the Data Lake
Setting Up the Data LakeSetting Up the Data Lake
Setting Up the Data LakeCaserta
 
Datawarehousing
DatawarehousingDatawarehousing
Datawarehousingwork
 
Derfor skal du bruge en DataLake
Derfor skal du bruge en DataLakeDerfor skal du bruge en DataLake
Derfor skal du bruge en DataLakeMicrosoft
 
Why Your Data Science Architecture Should Include a Data Virtualization Tool ...
Why Your Data Science Architecture Should Include a Data Virtualization Tool ...Why Your Data Science Architecture Should Include a Data Virtualization Tool ...
Why Your Data Science Architecture Should Include a Data Virtualization Tool ...Denodo
 
BAR360 open data platform presentation at DAMA, Sydney
BAR360 open data platform presentation at DAMA, SydneyBAR360 open data platform presentation at DAMA, Sydney
BAR360 open data platform presentation at DAMA, SydneySai Paravastu
 
CS8091_BDA_Unit_I_Analytical_Architecture
CS8091_BDA_Unit_I_Analytical_ArchitectureCS8091_BDA_Unit_I_Analytical_Architecture
CS8091_BDA_Unit_I_Analytical_ArchitecturePalani Kumar
 
Big Data's Impact on the Enterprise
Big Data's Impact on the EnterpriseBig Data's Impact on the Enterprise
Big Data's Impact on the EnterpriseCaserta
 
How Data Virtualization Puts Enterprise Machine Learning Programs into Produc...
How Data Virtualization Puts Enterprise Machine Learning Programs into Produc...How Data Virtualization Puts Enterprise Machine Learning Programs into Produc...
How Data Virtualization Puts Enterprise Machine Learning Programs into Produc...Denodo
 
Dataware housing
Dataware housingDataware housing
Dataware housingwork
 
Data Virtualization. An Introduction (ASEAN)
Data Virtualization. An Introduction (ASEAN)Data Virtualization. An Introduction (ASEAN)
Data Virtualization. An Introduction (ASEAN)Denodo
 
[DSC Europe 23] Milos Solujic - Data Lakehouse Revolutionizing Data Managemen...
[DSC Europe 23] Milos Solujic - Data Lakehouse Revolutionizing Data Managemen...[DSC Europe 23] Milos Solujic - Data Lakehouse Revolutionizing Data Managemen...
[DSC Europe 23] Milos Solujic - Data Lakehouse Revolutionizing Data Managemen...DataScienceConferenc1
 
Overview of Business Intelligence
Overview of Business IntelligenceOverview of Business Intelligence
Overview of Business IntelligenceParthiv Dixit
 
Big Data Analytics and Machine Learning Document.docx
Big Data Analytics and Machine Learning Document.docxBig Data Analytics and Machine Learning Document.docx
Big Data Analytics and Machine Learning Document.docxZitin Technologies PVT LTD
 
What's New in Pentaho 7.0?
What's New in Pentaho 7.0?What's New in Pentaho 7.0?
What's New in Pentaho 7.0?Xpand IT
 
Big data journey to the cloud maz chaudhri 5.30.18
Big data journey to the cloud   maz chaudhri 5.30.18Big data journey to the cloud   maz chaudhri 5.30.18
Big data journey to the cloud maz chaudhri 5.30.18Cloudera, Inc.
 

Similar a Data lake benefits (20)

intelligent-data-lake_executive-brief
intelligent-data-lake_executive-briefintelligent-data-lake_executive-brief
intelligent-data-lake_executive-brief
 
What Data Do You Have and Where is It?
What Data Do You Have and Where is It? What Data Do You Have and Where is It?
What Data Do You Have and Where is It?
 
BI Masterclass slides (Reference Architecture v3)
BI Masterclass slides (Reference Architecture v3)BI Masterclass slides (Reference Architecture v3)
BI Masterclass slides (Reference Architecture v3)
 
Setting Up the Data Lake
Setting Up the Data LakeSetting Up the Data Lake
Setting Up the Data Lake
 
Datawarehousing
DatawarehousingDatawarehousing
Datawarehousing
 
Derfor skal du bruge en DataLake
Derfor skal du bruge en DataLakeDerfor skal du bruge en DataLake
Derfor skal du bruge en DataLake
 
Why Your Data Science Architecture Should Include a Data Virtualization Tool ...
Why Your Data Science Architecture Should Include a Data Virtualization Tool ...Why Your Data Science Architecture Should Include a Data Virtualization Tool ...
Why Your Data Science Architecture Should Include a Data Virtualization Tool ...
 
Big data and oracle
Big data and oracleBig data and oracle
Big data and oracle
 
BAR360 open data platform presentation at DAMA, Sydney
BAR360 open data platform presentation at DAMA, SydneyBAR360 open data platform presentation at DAMA, Sydney
BAR360 open data platform presentation at DAMA, Sydney
 
CS8091_BDA_Unit_I_Analytical_Architecture
CS8091_BDA_Unit_I_Analytical_ArchitectureCS8091_BDA_Unit_I_Analytical_Architecture
CS8091_BDA_Unit_I_Analytical_Architecture
 
Big Data's Impact on the Enterprise
Big Data's Impact on the EnterpriseBig Data's Impact on the Enterprise
Big Data's Impact on the Enterprise
 
How Data Virtualization Puts Enterprise Machine Learning Programs into Produc...
How Data Virtualization Puts Enterprise Machine Learning Programs into Produc...How Data Virtualization Puts Enterprise Machine Learning Programs into Produc...
How Data Virtualization Puts Enterprise Machine Learning Programs into Produc...
 
Dataware housing
Dataware housingDataware housing
Dataware housing
 
Data Virtualization. An Introduction (ASEAN)
Data Virtualization. An Introduction (ASEAN)Data Virtualization. An Introduction (ASEAN)
Data Virtualization. An Introduction (ASEAN)
 
[DSC Europe 23] Milos Solujic - Data Lakehouse Revolutionizing Data Managemen...
[DSC Europe 23] Milos Solujic - Data Lakehouse Revolutionizing Data Managemen...[DSC Europe 23] Milos Solujic - Data Lakehouse Revolutionizing Data Managemen...
[DSC Europe 23] Milos Solujic - Data Lakehouse Revolutionizing Data Managemen...
 
Overview of Business Intelligence
Overview of Business IntelligenceOverview of Business Intelligence
Overview of Business Intelligence
 
Big Data Analytics and Machine Learning Document.docx
Big Data Analytics and Machine Learning Document.docxBig Data Analytics and Machine Learning Document.docx
Big Data Analytics and Machine Learning Document.docx
 
Machine Data Analytics
Machine Data AnalyticsMachine Data Analytics
Machine Data Analytics
 
What's New in Pentaho 7.0?
What's New in Pentaho 7.0?What's New in Pentaho 7.0?
What's New in Pentaho 7.0?
 
Big data journey to the cloud maz chaudhri 5.30.18
Big data journey to the cloud   maz chaudhri 5.30.18Big data journey to the cloud   maz chaudhri 5.30.18
Big data journey to the cloud maz chaudhri 5.30.18
 

Último

Call me @ 9892124323 Cheap Rate Call Girls in Vashi with Real Photo 100% Secure
Call me @ 9892124323  Cheap Rate Call Girls in Vashi with Real Photo 100% SecureCall me @ 9892124323  Cheap Rate Call Girls in Vashi with Real Photo 100% Secure
Call me @ 9892124323 Cheap Rate Call Girls in Vashi with Real Photo 100% SecurePooja Nehwal
 
Zuja dropshipping via API with DroFx.pptx
Zuja dropshipping via API with DroFx.pptxZuja dropshipping via API with DroFx.pptx
Zuja dropshipping via API with DroFx.pptxolyaivanovalion
 
BDSM⚡Call Girls in Mandawali Delhi >༒8448380779 Escort Service
BDSM⚡Call Girls in Mandawali Delhi >༒8448380779 Escort ServiceBDSM⚡Call Girls in Mandawali Delhi >༒8448380779 Escort Service
BDSM⚡Call Girls in Mandawali Delhi >༒8448380779 Escort ServiceDelhi Call girls
 
Al Barsha Escorts $#$ O565212860 $#$ Escort Service In Al Barsha
Al Barsha Escorts $#$ O565212860 $#$ Escort Service In Al BarshaAl Barsha Escorts $#$ O565212860 $#$ Escort Service In Al Barsha
Al Barsha Escorts $#$ O565212860 $#$ Escort Service In Al BarshaAroojKhan71
 
Log Analysis using OSSEC sasoasasasas.pptx
Log Analysis using OSSEC sasoasasasas.pptxLog Analysis using OSSEC sasoasasasas.pptx
Log Analysis using OSSEC sasoasasasas.pptxJohnnyPlasten
 
Delhi Call Girls Punjabi Bagh 9711199171 ☎✔👌✔ Whatsapp Hard And Sexy Vip Call
Delhi Call Girls Punjabi Bagh 9711199171 ☎✔👌✔ Whatsapp Hard And Sexy Vip CallDelhi Call Girls Punjabi Bagh 9711199171 ☎✔👌✔ Whatsapp Hard And Sexy Vip Call
Delhi Call Girls Punjabi Bagh 9711199171 ☎✔👌✔ Whatsapp Hard And Sexy Vip Callshivangimorya083
 
Smarteg dropshipping via API with DroFx.pptx
Smarteg dropshipping via API with DroFx.pptxSmarteg dropshipping via API with DroFx.pptx
Smarteg dropshipping via API with DroFx.pptxolyaivanovalion
 
Call Girls Bannerghatta Road Just Call 👗 7737669865 👗 Top Class Call Girl Ser...
Call Girls Bannerghatta Road Just Call 👗 7737669865 👗 Top Class Call Girl Ser...Call Girls Bannerghatta Road Just Call 👗 7737669865 👗 Top Class Call Girl Ser...
Call Girls Bannerghatta Road Just Call 👗 7737669865 👗 Top Class Call Girl Ser...amitlee9823
 
Mature dropshipping via API with DroFx.pptx
Mature dropshipping via API with DroFx.pptxMature dropshipping via API with DroFx.pptx
Mature dropshipping via API with DroFx.pptxolyaivanovalion
 
BabyOno dropshipping via API with DroFx.pptx
BabyOno dropshipping via API with DroFx.pptxBabyOno dropshipping via API with DroFx.pptx
BabyOno dropshipping via API with DroFx.pptxolyaivanovalion
 
BPAC WITH UFSBI GENERAL PRESENTATION 18_05_2017-1.pptx
BPAC WITH UFSBI GENERAL PRESENTATION 18_05_2017-1.pptxBPAC WITH UFSBI GENERAL PRESENTATION 18_05_2017-1.pptx
BPAC WITH UFSBI GENERAL PRESENTATION 18_05_2017-1.pptxMohammedJunaid861692
 
VIP Model Call Girls Hinjewadi ( Pune ) Call ON 8005736733 Starting From 5K t...
VIP Model Call Girls Hinjewadi ( Pune ) Call ON 8005736733 Starting From 5K t...VIP Model Call Girls Hinjewadi ( Pune ) Call ON 8005736733 Starting From 5K t...
VIP Model Call Girls Hinjewadi ( Pune ) Call ON 8005736733 Starting From 5K t...SUHANI PANDEY
 
Week-01-2.ppt BBB human Computer interaction
Week-01-2.ppt BBB human Computer interactionWeek-01-2.ppt BBB human Computer interaction
Week-01-2.ppt BBB human Computer interactionfulawalesam
 
Best VIP Call Girls Noida Sector 22 Call Me: 8448380779
Best VIP Call Girls Noida Sector 22 Call Me: 8448380779Best VIP Call Girls Noida Sector 22 Call Me: 8448380779
Best VIP Call Girls Noida Sector 22 Call Me: 8448380779Delhi Call girls
 
Market Analysis in the 5 Largest Economic Countries in Southeast Asia.pdf
Market Analysis in the 5 Largest Economic Countries in Southeast Asia.pdfMarket Analysis in the 5 Largest Economic Countries in Southeast Asia.pdf
Market Analysis in the 5 Largest Economic Countries in Southeast Asia.pdfRachmat Ramadhan H
 
Vip Model Call Girls (Delhi) Karol Bagh 9711199171✔️Body to body massage wit...
Vip Model  Call Girls (Delhi) Karol Bagh 9711199171✔️Body to body massage wit...Vip Model  Call Girls (Delhi) Karol Bagh 9711199171✔️Body to body massage wit...
Vip Model Call Girls (Delhi) Karol Bagh 9711199171✔️Body to body massage wit...shivangimorya083
 
CebaBaby dropshipping via API with DroFX.pptx
CebaBaby dropshipping via API with DroFX.pptxCebaBaby dropshipping via API with DroFX.pptx
CebaBaby dropshipping via API with DroFX.pptxolyaivanovalion
 
Digital Advertising Lecture for Advanced Digital & Social Media Strategy at U...
Digital Advertising Lecture for Advanced Digital & Social Media Strategy at U...Digital Advertising Lecture for Advanced Digital & Social Media Strategy at U...
Digital Advertising Lecture for Advanced Digital & Social Media Strategy at U...Valters Lauzums
 
Determinants of health, dimensions of health, positive health and spectrum of...
Determinants of health, dimensions of health, positive health and spectrum of...Determinants of health, dimensions of health, positive health and spectrum of...
Determinants of health, dimensions of health, positive health and spectrum of...shambhavirathore45
 

Último (20)

Call me @ 9892124323 Cheap Rate Call Girls in Vashi with Real Photo 100% Secure
Call me @ 9892124323  Cheap Rate Call Girls in Vashi with Real Photo 100% SecureCall me @ 9892124323  Cheap Rate Call Girls in Vashi with Real Photo 100% Secure
Call me @ 9892124323 Cheap Rate Call Girls in Vashi with Real Photo 100% Secure
 
Zuja dropshipping via API with DroFx.pptx
Zuja dropshipping via API with DroFx.pptxZuja dropshipping via API with DroFx.pptx
Zuja dropshipping via API with DroFx.pptx
 
BDSM⚡Call Girls in Mandawali Delhi >༒8448380779 Escort Service
BDSM⚡Call Girls in Mandawali Delhi >༒8448380779 Escort ServiceBDSM⚡Call Girls in Mandawali Delhi >༒8448380779 Escort Service
BDSM⚡Call Girls in Mandawali Delhi >༒8448380779 Escort Service
 
Al Barsha Escorts $#$ O565212860 $#$ Escort Service In Al Barsha
Al Barsha Escorts $#$ O565212860 $#$ Escort Service In Al BarshaAl Barsha Escorts $#$ O565212860 $#$ Escort Service In Al Barsha
Al Barsha Escorts $#$ O565212860 $#$ Escort Service In Al Barsha
 
Log Analysis using OSSEC sasoasasasas.pptx
Log Analysis using OSSEC sasoasasasas.pptxLog Analysis using OSSEC sasoasasasas.pptx
Log Analysis using OSSEC sasoasasasas.pptx
 
Delhi Call Girls Punjabi Bagh 9711199171 ☎✔👌✔ Whatsapp Hard And Sexy Vip Call
Delhi Call Girls Punjabi Bagh 9711199171 ☎✔👌✔ Whatsapp Hard And Sexy Vip CallDelhi Call Girls Punjabi Bagh 9711199171 ☎✔👌✔ Whatsapp Hard And Sexy Vip Call
Delhi Call Girls Punjabi Bagh 9711199171 ☎✔👌✔ Whatsapp Hard And Sexy Vip Call
 
Smarteg dropshipping via API with DroFx.pptx
Smarteg dropshipping via API with DroFx.pptxSmarteg dropshipping via API with DroFx.pptx
Smarteg dropshipping via API with DroFx.pptx
 
Call Girls Bannerghatta Road Just Call 👗 7737669865 👗 Top Class Call Girl Ser...
Call Girls Bannerghatta Road Just Call 👗 7737669865 👗 Top Class Call Girl Ser...Call Girls Bannerghatta Road Just Call 👗 7737669865 👗 Top Class Call Girl Ser...
Call Girls Bannerghatta Road Just Call 👗 7737669865 👗 Top Class Call Girl Ser...
 
Mature dropshipping via API with DroFx.pptx
Mature dropshipping via API with DroFx.pptxMature dropshipping via API with DroFx.pptx
Mature dropshipping via API with DroFx.pptx
 
BabyOno dropshipping via API with DroFx.pptx
BabyOno dropshipping via API with DroFx.pptxBabyOno dropshipping via API with DroFx.pptx
BabyOno dropshipping via API with DroFx.pptx
 
BPAC WITH UFSBI GENERAL PRESENTATION 18_05_2017-1.pptx
BPAC WITH UFSBI GENERAL PRESENTATION 18_05_2017-1.pptxBPAC WITH UFSBI GENERAL PRESENTATION 18_05_2017-1.pptx
BPAC WITH UFSBI GENERAL PRESENTATION 18_05_2017-1.pptx
 
VIP Model Call Girls Hinjewadi ( Pune ) Call ON 8005736733 Starting From 5K t...
VIP Model Call Girls Hinjewadi ( Pune ) Call ON 8005736733 Starting From 5K t...VIP Model Call Girls Hinjewadi ( Pune ) Call ON 8005736733 Starting From 5K t...
VIP Model Call Girls Hinjewadi ( Pune ) Call ON 8005736733 Starting From 5K t...
 
Week-01-2.ppt BBB human Computer interaction
Week-01-2.ppt BBB human Computer interactionWeek-01-2.ppt BBB human Computer interaction
Week-01-2.ppt BBB human Computer interaction
 
CHEAP Call Girls in Saket (-DELHI )🔝 9953056974🔝(=)/CALL GIRLS SERVICE
CHEAP Call Girls in Saket (-DELHI )🔝 9953056974🔝(=)/CALL GIRLS SERVICECHEAP Call Girls in Saket (-DELHI )🔝 9953056974🔝(=)/CALL GIRLS SERVICE
CHEAP Call Girls in Saket (-DELHI )🔝 9953056974🔝(=)/CALL GIRLS SERVICE
 
Best VIP Call Girls Noida Sector 22 Call Me: 8448380779
Best VIP Call Girls Noida Sector 22 Call Me: 8448380779Best VIP Call Girls Noida Sector 22 Call Me: 8448380779
Best VIP Call Girls Noida Sector 22 Call Me: 8448380779
 
Market Analysis in the 5 Largest Economic Countries in Southeast Asia.pdf
Market Analysis in the 5 Largest Economic Countries in Southeast Asia.pdfMarket Analysis in the 5 Largest Economic Countries in Southeast Asia.pdf
Market Analysis in the 5 Largest Economic Countries in Southeast Asia.pdf
 
Vip Model Call Girls (Delhi) Karol Bagh 9711199171✔️Body to body massage wit...
Vip Model  Call Girls (Delhi) Karol Bagh 9711199171✔️Body to body massage wit...Vip Model  Call Girls (Delhi) Karol Bagh 9711199171✔️Body to body massage wit...
Vip Model Call Girls (Delhi) Karol Bagh 9711199171✔️Body to body massage wit...
 
CebaBaby dropshipping via API with DroFX.pptx
CebaBaby dropshipping via API with DroFX.pptxCebaBaby dropshipping via API with DroFX.pptx
CebaBaby dropshipping via API with DroFX.pptx
 
Digital Advertising Lecture for Advanced Digital & Social Media Strategy at U...
Digital Advertising Lecture for Advanced Digital & Social Media Strategy at U...Digital Advertising Lecture for Advanced Digital & Social Media Strategy at U...
Digital Advertising Lecture for Advanced Digital & Social Media Strategy at U...
 
Determinants of health, dimensions of health, positive health and spectrum of...
Determinants of health, dimensions of health, positive health and spectrum of...Determinants of health, dimensions of health, positive health and spectrum of...
Determinants of health, dimensions of health, positive health and spectrum of...
 

Data lake benefits

  • 1. Strategic  Advisory Big  Data  – Cloud   -­‐ Analytics Info Strategy Fishing  in  the   big  data  lake DATA  EXPLORATION  AND  DISCOVERY  ANALYTICS   FOR  DEEPER  BUSINESS  INSIGHTS
  • 2. InfoStrategy What  is  a  “data  lake” data  lake (plural data  lakes) A  massive,  easily  accessible  data  repository   built  on  (relatively)  inexpensive  computer   hardware  for  storing  "big  data".  Unlike  data  marts,   which  are  optimized  for  data  analysis  by  storing  only  some   attributes  and  dropping  data  below  the  level  aggregation,  a   data  lake  is  designed  to  retain  all  attributes,   especially  so  when  you  do  not  yet  know  what  the   scope  of  data  or  its  use  will  be. http://en.wiktionary.org/wiki/data_lake …  Enterprise  Data  Hub  sounds  too  boring   !
  • 3. InfoStrategy Optimise  business  through  insights Insight Action Optimise Move  a  metric Change  a  product Change  behaviour/process Hindsight Realtime Foresight Trusted  information Act  on  insights  gained Execute  theories Measure Outcomes Sentiment Feedback Explore  datasets,  discover  correlations,  patterns. Undiscovered  facts Information  Value Data  Volumes Forecasting,  planning  &  trending Statistical  Analysis Operational  reporting,  SCADA  control Alerts  &  Events Historical  reporting, Proof  of  operation Regulatory,  statutory,  financial Uncover  previously   unknown  facts   from  enriched  data   in  the  data  lake
  • 4. InfoStrategy Future  state  of  analytics Strategic  Intent To  improve  BI  and  Analytical  capabilities  to  a  level  where  organisations  are  able  to   access  and  analyse  information  in  a  secure,  timely  and  cost-­‐effective  manner. Gain  key  insights  to  optimise  the  operations  of  your  business,  predict  the  best   possible  outcomes  for  growth,  new  opportunities,   and  competitive  advantage   across  all  business  lines. Mission  Statement “Providing  advanced  analytics  capability  across  all  business  units,  empowering  our   people  with  the    processes  and  supporting  technologies  to  exploit  our  information   assets  for  business  benefit.” Target  Operating  Model  will  deliver: Rapid  access  to  data  to  uncover  new  facts  via  advanced  data  exploration  and   discovery  analytics. Clarity  of  who  is  responsible  and  accountable  for  maintaining  critical  information   assets  via  a  well  structured  governance  and  engagement  model. A  trusted  and  highly  secure  source  of  data  for  all  analytical  information  requirements   via  a  data  quality  assurance  program. Trawling  for  value  in  the  big  data  lake
  • 5. InfoStrategy ‘Fish  stocks’  are  replenished  from  existing  and  future   operational  systems  plus  external  sources Core   Transactional  Data   “operational” Management   Reporting Unstructured  &   External  Data “contextual” Enterprise  Dashboards Reporting Consolidation Data  ScientistsBusiness  AnalystsBusiness  UsersCustomers Data  Extraction Discovery  Analytics   Platform Visualisation Analysis Data  Preparation Data  Collection Operational   Reporting Operational  Dashboards Real-­‐time  Reports Alerts  &  Exceptions Embedded  BI Production   Data  Repository “Data  Lake” Information  Governance Data  Management Supplier  &   Industry  Data “comparative”
  • 6. InfoStrategy Consolidated Management Reporting Operational Supporting Capability Discovery Analytics To  meet  the  demand  for  rapid  access  to  information   users  must  adopt  a  flexible  multi-­‐platform   architecture   What  reporting  does  for  established  operations  …  discovery  analytics  does  for  new  business  development. The  trend  within  industry  is  to  move  away  from  the  single-­‐platform  monolithic  data  warehouses  towards  a  physically  distributed  environment   for  information  delivery.  Many  businesses  are  extending  their  data  warehouse  environments  to  include  new  standalone  data  platforms  that   are  conducive  to  discovery  analytics.  A  holistic  view  is  maintained  via  a  common,  single  replicated  dataset  and  an  enterprise information   management  program,  governing  delivery  and  access  to  key  information  (data  lake). Source   Applications ERP CRM HR Finance Telemetry Geospatial  GIS Documents Email Files Real-­time  Data   Capture Cleansing Loading Data  Warehouse Modelling Relational  DW Data  Marts Analysis  Cubes Analytics Delivery Cloud-­based    Service  Model Actuarial   Applications Event-­Based   Applications Reporting Production   Reporting OLAP  Analytics Ad  Hoc  Query External Data Exploration  &   Discovery Metadata  Integration Event  Processing Results Detailed  Datasets Results   Collection  and  blending Insights Portal PDF Desktop Guided   Visualisation Mobile  BI Active   Dashboards Data  Replication Historical Data  Preparation Storytelling Information  Governance Operational  Reporting   Dimensional   Modelling ProductioniseInsights
  • 7. InfoStrategy Principles:  Easier  access  information   to  discover  new   facts  about  the  business. ◦ Described  as  a  ‘sandpit’  environment,  providing  the  ability  to  explore  and  discover  new   facts  about  the  business,  it’s  members  and  customers,  partners  and  competitive   pressures. ◦ Also  used  for  testing  a  hypothesis  or  running  scenarios  across  the  data ◦ Getting  answers  to  ‘one-­‐off’  questions  which  are  not  addressed  through  the  normal   published,  scheduled  operational  reporting  channels ◦ Data  is  replicated  from  all  operational  systems  into  a  single  landing  area,  ensuring   traceability  and  reconciliation  to  all  consuming  applications,  such  as  the  data  warehouse,   analytical  application,  and  other  business  applications. ◦ Clearly  defined  critical  business  entities/records  are  synchronised  (or  Mastered)  across   all  applications  eliminating  duplication  and  confusion.  Data  quality  attributes  are  defined   and  managed  for  each  critical  business  entity. ◦ A  fully  integrated  Member/Customer  view  is  established  across  both  analytical  and   transactional  applications. ◦ Using  the  replicated  data  to  build  more  dynamic  analytical  data  structures  for  scheduled   production  reporting  and  ah-­‐hoc  analysis ◦ Provide  users  with  the  tools  to  access    and  analyse data,  freely  explore  current  and  new   datasets,  and  visualise patterns  and  discoveries  to  gain  deep  insights. Providing  business  users  with  direct   access  to  data  to  meet  immediate   information  needs  where  the   accuracy  of  the  data  is  not  the   primary  objective.   Having  a  single  source  of  truth   across  all  business  applications  at   detailed  level  from  which  all   information  requests  are  satisfied. Improved  environment  for  more   cost  effective  and  faster  business   intelligence  delivery. Provide  business   users  with  the  ability  to  access  production  information  directly,  collect  it  as  needed,  and   prepare  the  data  for  analysis.  Exploring  the  data  to  uncover  previously   unknown  facts  about  the  business,   and   sharing  those  facts  visually  with  others.  Enrich  production  data  with  external  “context”  to  extend  insights. Key  Principles Description
  • 8. InfoStrategy Benefits  of  Discovery  Analytics  versus  traditional   data   warehousing Classic  Data  Warehouse  Issues Discovery  Analytics Benefit Lengthy  IT  Backlog  and  lack  of  resources  to  extend the   EDW  to  support  new  business  requirements. Data  can  be  explored  and  analysed  outside  of the  EDW   environment  before  it  is  put  into  production  use. High  costs  of  supporting increasing  data  volumes  and   new  types  of  data. Data  can  be  filtered  and  transformed  before  it  is  loaded   into  the  EDW Lack  of  flexibility  in  the  EDW  data  model  to  support   constantly changing  business  requirements. Data  discovery  support  dynamic  schema  on  read   approach  which reduces  the  need  for  detailed  up-­‐front   modelling. Need  to  have  data  quality  and  governance  processes  in   place  before  user  can  access  the  EDW  data. The  investigative  nature  of data  discovery  has  lower  data   quality  and  governance  requirements Growing  use  of  personal  data  marts to  overcome  IT   barriers  and  the  performance  overheads  of  ad  hoc   processing The  flexibility  and  performance  of  data  discovery   encourages  shared  use  of  data  and  analytics. Recent  proof  of  concept  for  Discovery  Analytics  in  the  cloud  (AWS),  has  provided  some   considerable  cost  &  time  savings  in  infrastructure  and  hosting,  viz.: $55  per  day  to  host  a  960GB  data  warehouse   $32  per  day  to  host  a  Data  Integration  server  AND  a  BI  server. 2.5  weeks  to  setup  POC  environment  and  start  analysis  and  visualising  results.
  • 9. InfoStrategy Discovery  Analytics  Target  POC  Architecture Structured   Data Unstructured   Data ERP Telemetry Web/External Replication  of  corporate  data,  enriched  with  external  data  and   content,  available  in  a  centrally  available  and  scalable  repository   ready  for  exploration,  discovery  and  predictive  analysis  to  gain   deep  insights  and  actionable  results.
  • 10. InfoStrategy Fishing  safely  with  the  appropriate  life  vests  is   important  too. Security  and  data  management  standards  are  available International   Standard  on   Assurance   Engagements Service  Organisation   Control  framework Federal  Information   Management   Security  Act Payment  Card   Industry  –Data   Security  Standard Federal  Information   Processing  Standard International  Standards   Organisation  – Information  Security   Standard Source:  Amazon  Web  Services
  • 11. Info Strategy To  learn  more  about  how  InfoStrategy can  help  you  develop  your  big  data   strategy  to  solve  your  big  business   problems,  or  to  arrange  a  Proof  of   Concept,  please  contact  us  today  using   the  details  below. InfoStrategy Pty  Ltd 246  Oxford  St,  Balmoral Queensland  4171 Australia Tel:  +61  7  3151  2021 Email:   contactus@infostrategy.com.au