Data Science at Scale using Spark and Hadoop DSSH training course Kuala Lumpur Malaysia
Duration : 3 Days
Who should attend
- Developers
- Data analysts
- Statisticians
Prerequisites
- Proficiency in a scripting language
- Python is strongly preferred
- Perl or Ruby is sufficient
- Basic knowledge of Apache Hadoop
- Experience working in Linux environments
Course Objectives
After completing this class, you will learn:
- How to identify potential business use cases where data science can provide impactful results
- How to obtain, clean and combine disparate data sources to create a coherent picture for analysis
- What statistical methods to leverage for data exploration that will provide critical insight into your data
- Where and when to leverage Hadoop streaming and Apache Spark for data science pipelines
- What machine learning technique to use for a particular data science project
- How to implement and manage recommenders using Spark’s MLlib, and how to set up and evaluate data experiments
- What are the pitfalls of deploying new analytics projects to production, at scale
Course Outline
Data Science Overview
- What Is Data Science?
- The Growing Need for Data Science
- The Role of a Data Scientist
- Finance
- Retail
- Advertising
- Defense and Intelligence
- Telecommunications and Utilities
- Healthcare and Pharmaceuticals
- Steps in the Project Lifecycle
- Lab Scenario Explanation
- Where to Source Data
- Acquisition Techniques
- Data Formats
- Data Quantity
- Data Quality
- File Format Conversion
- Joining Data Sets
- Anonymization
- Relationship Between Statistics and Probability
- Descriptive Statistics
- Inferential Statistics
- Vectors and Matrices
- Overview
- The Three C’s of Machine Learning
- Importance of Data and Algorithms
- Spotlight: Naive Bayes Classifiers
- What is a Recommender System?
- Types of Collaborative Filtering
- Limitations of Recommender Systems
- Fundamental Concepts
- What is Apache Spark?
- Comparison to MapReduce
- Fundamentals of Apache Spark
- Spark’s MLlib Package
- Overview of ALS Method for Latent Factor Recommenders
- Hyperparameters for ALS Recommenders
- Building a Recommender in MLlib
- Tuning Hyperparameters
- Weighting
- Designing Effective Experiments
- Conducting an Effective Experiment
- User Interfaces for Recommenders
- Deploying to Production
- Tips and Techniques for Working at Scale
- Summarizing and Visualizing Results
- Considerations for Improvement
- Next Steps for Recommenders