WEBVTT

00:00:00.000 --> 00:00:14.070 align:middle line:90%


00:00:14.070 --> 00:00:16.570 align:middle line:84%
I'm Lorry Hardt, the
global product marketing

00:00:16.570 --> 00:00:19.330 align:middle line:84%
manager for Advanced
Analytics, and joining me today

00:00:19.330 --> 00:00:21.400 align:middle line:84%
is Jonathan Wexler,
the principal product

00:00:21.400 --> 00:00:23.590 align:middle line:84%
manager for SAS's
Machine Learning suite.

00:00:23.590 --> 00:00:25.670 align:middle line:84%
Jonathan, thanks for
talking with me today.

00:00:25.670 --> 00:00:26.540 align:middle line:84%
JONATHAN WEXLER: Glad
to be here, Lorry,

00:00:26.540 --> 00:00:28.450 align:middle line:90%
to discuss our new release.

00:00:28.450 --> 00:00:29.500 align:middle line:84%
LORRY HARDT:
Justifiably, there's

00:00:29.500 --> 00:00:31.090 align:middle line:84%
a lot of hype that
surrounds supplying

00:00:31.090 --> 00:00:32.860 align:middle line:84%
more advanced
modeling techniques,

00:00:32.860 --> 00:00:34.330 align:middle line:90%
such as machine learning.

00:00:34.330 --> 00:00:36.960 align:middle line:84%
This is coming from the C level
enterprise architecture leaders

00:00:36.960 --> 00:00:38.486 align:middle line:90%
to chief analytic officers.

00:00:38.486 --> 00:00:40.430 align:middle line:84%
This has really
become a hot topic.

00:00:40.430 --> 00:00:42.152 align:middle line:90%
What are your thoughts on this?

00:00:42.152 --> 00:00:42.690 align:middle line:90%
Why now?

00:00:42.690 --> 00:00:43.720 align:middle line:84%
JONATHAN WEXLER:
Lorry, to be frank,

00:00:43.720 --> 00:00:45.940 align:middle line:84%
machine learning is
not a new concept.

00:00:45.940 --> 00:00:47.920 align:middle line:84%
Many of these techniques,
such as neural networks,

00:00:47.920 --> 00:00:49.880 align:middle line:90%
have been around for many years.

00:00:49.880 --> 00:00:51.640 align:middle line:84%
It's just that the
technology has finally

00:00:51.640 --> 00:00:53.980 align:middle line:84%
caught up with the
computational need for analyzing

00:00:53.980 --> 00:00:56.830 align:middle line:84%
large complex data sources,
including suspicious trading

00:00:56.830 --> 00:00:58.840 align:middle line:84%
activity,
recommendation engines,

00:00:58.840 --> 00:01:00.970 align:middle line:90%
to analyzing complex images.

00:01:00.970 --> 00:01:01.930 align:middle line:84%
LORRY HARDT: One
of the key areas

00:01:01.930 --> 00:01:03.370 align:middle line:84%
that when we discuss
machine learning

00:01:03.370 --> 00:01:05.770 align:middle line:84%
is the enabling non-expert
data scientists or even

00:01:05.770 --> 00:01:07.090 align:middle line:84%
business analysts
with the ability

00:01:07.090 --> 00:01:09.250 align:middle line:84%
to rapidly explore
and model data.

00:01:09.250 --> 00:01:12.010 align:middle line:84%
In SAS's latest release of SAS
Visual Data Mining and Machine

00:01:12.010 --> 00:01:13.930 align:middle line:84%
Learning, how do you see
this product addressing

00:01:13.930 --> 00:01:15.100 align:middle line:90%
those types of tasks?

00:01:15.100 --> 00:01:16.760 align:middle line:84%
JONATHAN WEXLER: Lorry,
SAS Visual Data Mining

00:01:16.760 --> 00:01:18.130 align:middle line:84%
and Machine Learning
enables users

00:01:18.130 --> 00:01:20.470 align:middle line:84%
to rapidly expand
their models in a very

00:01:20.470 --> 00:01:22.390 align:middle line:90%
visual interactive way.

00:01:22.390 --> 00:01:25.900 align:middle line:84%
This solution enables users to
analyze data in the experience

00:01:25.900 --> 00:01:27.430 align:middle line:90%
that they feel comfortable with.

00:01:27.430 --> 00:01:30.652 align:middle line:84%
They can quickly go from data
prep to interactive modeling

00:01:30.652 --> 00:01:32.597 align:middle line:84%
to advanced
pipelining, and finally

00:01:32.597 --> 00:01:34.763 align:middle line:84%
and most importantly,
to deployment all

00:01:34.763 --> 00:01:36.200 align:middle line:90%
within this one solution.

00:01:36.200 --> 00:01:39.755 align:middle line:84%
Let's see SAS Visual Data Mining
and Machine Learning in action.

00:01:39.755 --> 00:01:41.729 align:middle line:90%


00:01:41.729 --> 00:01:43.840 align:middle line:84%
For my particular
use case, I'm trying

00:01:43.840 --> 00:01:46.006 align:middle line:84%
to analyze what
chemical factors helped

00:01:46.006 --> 00:01:48.040 align:middle line:90%
me predict the quality of wine.

00:01:48.040 --> 00:01:50.290 align:middle line:84%
I've already prepared my
data using a data preparation

00:01:50.290 --> 00:01:53.050 align:middle line:84%
area of SAS Visual Data
Mining and Machine Learning,

00:01:53.050 --> 00:01:55.994 align:middle line:84%
and now I'm in the process
of exploring and expanding

00:01:55.994 --> 00:02:00.050 align:middle line:84%
my analysis by using advanced
machine learning techniques.

00:02:00.050 --> 00:02:06.655 align:middle line:84%
So I can very quickly
add in a neural network,

00:02:06.655 --> 00:02:08.710 align:middle line:84%
and before I build
my neural network,

00:02:08.710 --> 00:02:10.120 align:middle line:84%
the system gives
me an indication

00:02:10.120 --> 00:02:11.200 align:middle line:90%
of what I'm about to build.

00:02:11.200 --> 00:02:13.330 align:middle line:84%
So if I'm not familiar
with the neural network,

00:02:13.330 --> 00:02:16.480 align:middle line:84%
it shows me that it's about
to add in a network diagram

00:02:16.480 --> 00:02:18.490 align:middle line:90%
and various assessment metrics.

00:02:18.490 --> 00:02:23.220 align:middle line:84%
And it's very easy to
build this neural network.

00:02:23.220 --> 00:02:26.490 align:middle line:84%
I can just select my attributes
and the neural network

00:02:26.490 --> 00:02:28.490 align:middle line:84%
will build itself
in near real-time.

00:02:28.490 --> 00:02:30.839 align:middle line:90%


00:02:30.839 --> 00:02:32.950 align:middle line:84%
I can interact with
the neural network

00:02:32.950 --> 00:02:34.894 align:middle line:84%
by adding additional
hidden layers.

00:02:34.894 --> 00:02:39.222 align:middle line:90%


00:02:39.222 --> 00:02:41.110 align:middle line:84%
I can change any
of the activation

00:02:41.110 --> 00:02:43.585 align:middle line:84%
functions of the hidden
layers themselves.

00:02:43.585 --> 00:02:46.560 align:middle line:90%


00:02:46.560 --> 00:02:49.260 align:middle line:84%
I can also edit any of
the additional options.

00:02:49.260 --> 00:02:52.502 align:middle line:84%
I can even take advantage of
SAS's advanced auto-tuning

00:02:52.502 --> 00:02:54.390 align:middle line:84%
where the model
will automatically

00:02:54.390 --> 00:02:57.980 align:middle line:84%
find the best performing
attributes for me.

00:02:57.980 --> 00:03:00.320 align:middle line:84%
I can also interpret
my results by looking

00:03:00.320 --> 00:03:02.034 align:middle line:90%
at a relative importance plot.

00:03:02.034 --> 00:03:04.700 align:middle line:84%
Typically, machine learning
models are black box

00:03:04.700 --> 00:03:06.588 align:middle line:84%
and you're not
given an indication

00:03:06.588 --> 00:03:10.020 align:middle line:84%
of the most important factors
that drive your model.

00:03:10.020 --> 00:03:10.980 align:middle line:84%
LORRY HARDT: Jonathan,
you mentioned

00:03:10.980 --> 00:03:14.250 align:middle line:84%
in this latest release users can
rapidly expand their analyzes.

00:03:14.250 --> 00:03:17.100 align:middle line:84%
Can you give an example of that
type of analytic expansion?

00:03:17.100 --> 00:03:19.230 align:middle line:84%
JONATHAN WEXLER: Sure, Lorry,
within SAS Visual Data Mining

00:03:19.230 --> 00:03:21.210 align:middle line:84%
and Machine Learning, we
give users the ability

00:03:21.210 --> 00:03:23.820 align:middle line:84%
to enhance their interactive
models by mixing and matching

00:03:23.820 --> 00:03:26.370 align:middle line:84%
different techniques, applying
more advanced methods,

00:03:26.370 --> 00:03:29.314 align:middle line:84%
and even integrating
programming into their analyzes.

00:03:29.314 --> 00:03:33.170 align:middle line:90%


00:03:33.170 --> 00:03:35.226 align:middle line:84%
Let's go back to
that neural network.

00:03:35.226 --> 00:03:37.892 align:middle line:84%
What SAS Visual Data
Mining and Machine Learning

00:03:37.892 --> 00:03:40.660 align:middle line:84%
allows you to do is
expand this neural network

00:03:40.660 --> 00:03:44.540 align:middle line:84%
and build upon it by
creating a pipeline.

00:03:44.540 --> 00:03:46.820 align:middle line:84%
And by creating a
pipeline, this enables

00:03:46.820 --> 00:03:50.544 align:middle line:84%
me to try out different advanced
machine learning techniques.

00:03:50.544 --> 00:03:53.210 align:middle line:84%
And you'll see here that
the pipeline represents

00:03:53.210 --> 00:03:55.460 align:middle line:84%
all of the steps
that I took inside

00:03:55.460 --> 00:03:57.440 align:middle line:90%
the exploratory environment.

00:03:57.440 --> 00:03:59.930 align:middle line:84%
But now that I'm in
this more advanced view,

00:03:59.930 --> 00:04:02.596 align:middle line:84%
I have more advanced
techniques available to me.

00:04:02.596 --> 00:04:05.900 align:middle line:84%
So for example, I can
try anomaly detection

00:04:05.900 --> 00:04:09.240 align:middle line:84%
to either remove outliers
or model them separately.

00:04:09.240 --> 00:04:12.170 align:middle line:84%
I can also add in other
advanced techniques.

00:04:12.170 --> 00:04:14.810 align:middle line:84%
I could add in a
decision tree node.

00:04:14.810 --> 00:04:19.496 align:middle line:84%
I can also add in, for example,
a gradient boosting node.

00:04:19.496 --> 00:04:21.329 align:middle line:84%
And the system
will automatically

00:04:21.329 --> 00:04:23.460 align:middle line:84%
pick the best
performing model for me.

00:04:23.460 --> 00:04:25.293 align:middle line:84%
Let's take a look
at the results.

00:04:25.293 --> 00:04:29.390 align:middle line:90%


00:04:29.390 --> 00:04:31.390 align:middle line:84%
In this case, it
picked the gradient

00:04:31.390 --> 00:04:34.030 align:middle line:84%
boosting model with the
highest KS statistic,

00:04:34.030 --> 00:04:37.418 align:middle line:84%
but it also compared itself to
the interactive neural network

00:04:37.418 --> 00:04:40.570 align:middle line:84%
that I built and a decision
tree that I added in.

00:04:40.570 --> 00:04:43.960 align:middle line:84%
The system automatically
picked the best model for me.

00:04:43.960 --> 00:04:47.440 align:middle line:84%
I can also customize this
pipeline by changing options.

00:04:47.440 --> 00:04:51.170 align:middle line:84%
I have complete flexibility
to enhance this analysis.

00:04:51.170 --> 00:04:53.170 align:middle line:84%
LORRY HARDT: Jonathan,
you mentioned

00:04:53.170 --> 00:04:55.240 align:middle line:84%
that users can quickly
collaborate and share

00:04:55.240 --> 00:04:56.050 align:middle line:90%
their analyzes.

00:04:56.050 --> 00:04:59.110 align:middle line:84%
How can users take advantage of
these best practice templates?

00:04:59.110 --> 00:05:00.010 align:middle line:84%
JONATHAN WEXLER:
Out of the box, SAS

00:05:00.010 --> 00:05:02.920 align:middle line:84%
will provide a series of
best practice templates.

00:05:02.920 --> 00:05:05.127 align:middle line:84%
These are stored
in the SAS toolbox.

00:05:05.127 --> 00:05:07.960 align:middle line:84%
Data scientists can also
create their own pipelines

00:05:07.960 --> 00:05:09.760 align:middle line:84%
and share them with
the community, which

00:05:09.760 --> 00:05:11.920 align:middle line:84%
in turn, a business
analysts is empowered

00:05:11.920 --> 00:05:13.370 align:middle line:90%
to produce rapid results.

00:05:13.370 --> 00:05:16.251 align:middle line:90%


00:05:16.251 --> 00:05:19.250 align:middle line:84%
If we go back to the pipeline
that I've already built,

00:05:19.250 --> 00:05:23.170 align:middle line:84%
I can access the
SAS toolbox, which

00:05:23.170 --> 00:05:27.220 align:middle line:84%
allows me to take advantage
of either SAS best practice

00:05:27.220 --> 00:05:29.650 align:middle line:84%
pipelines, or even
add in a pipeline

00:05:29.650 --> 00:05:34.297 align:middle line:84%
from one of my coworkers,
a proven best practice.

00:05:34.297 --> 00:05:37.130 align:middle line:84%
Now this best practice
pipeline has a mix and match

00:05:37.130 --> 00:05:40.150 align:middle line:84%
of different advanced
techniques from gradient

00:05:40.150 --> 00:05:42.800 align:middle line:84%
boosting to using our
auto tuning to even

00:05:42.800 --> 00:05:44.150 align:middle line:90%
a deep neural network.

00:05:44.150 --> 00:05:47.668 align:middle line:84%
It even has the ability
to enable SAS code.

00:05:47.668 --> 00:05:48.390 align:middle line:90%
Let's run it.

00:05:48.390 --> 00:05:56.516 align:middle line:90%


00:05:56.516 --> 00:05:59.460 align:middle line:84%
And you can see it still
picked the gradient boosting

00:05:59.460 --> 00:06:02.190 align:middle line:84%
model, which had the
highest KS statistic.

00:06:02.190 --> 00:06:04.540 align:middle line:84%
It also had an ensemble
model in there.

00:06:04.540 --> 00:06:07.357 align:middle line:84%
It also did auto tuning
for my decision tree,

00:06:07.357 --> 00:06:09.690 align:middle line:84%
and it also added in
a logistic regression

00:06:09.690 --> 00:06:12.578 align:middle line:84%
and a deep neural network,
but the gradient boosting

00:06:12.578 --> 00:06:15.330 align:middle line:90%
was still the champion model.

00:06:15.330 --> 00:06:17.880 align:middle line:84%
Now I mentioned this
pipeline also enables

00:06:17.880 --> 00:06:20.560 align:middle line:90%
users to write their own code.

00:06:20.560 --> 00:06:24.490 align:middle line:84%
With this system, you can not
only use pre-built SAS nodes,

00:06:24.490 --> 00:06:26.530 align:middle line:84%
but you can also
add in SAS code,

00:06:26.530 --> 00:06:28.812 align:middle line:84%
and you have the
complete SAS language

00:06:28.812 --> 00:06:30.700 align:middle line:84%
available at your
fingertips to be

00:06:30.700 --> 00:06:33.700 align:middle line:84%
able to add in procedures,
macros, data step,

00:06:33.700 --> 00:06:34.460 align:middle line:90%
and so forth.

00:06:34.460 --> 00:06:36.580 align:middle line:84%
So you can really
build the best model

00:06:36.580 --> 00:06:38.408 align:middle line:90%
that makes sense for your data.

00:06:38.408 --> 00:06:40.630 align:middle line:84%
LORRY HARDT: One other
thought from you.

00:06:40.630 --> 00:06:42.630 align:middle line:84%
The analytic lifecycle
is no longer a debate.

00:06:42.630 --> 00:06:44.420 align:middle line:90%
It is necessary to support.

00:06:44.420 --> 00:06:46.870 align:middle line:84%
No model's complete unless
you get it into action,

00:06:46.870 --> 00:06:48.925 align:middle line:84%
enable businesses
to respond quickly,

00:06:48.925 --> 00:06:50.380 align:middle line:90%
and take action if needed.

00:06:50.380 --> 00:06:52.460 align:middle line:84%
Likewise, we live
in an open world.

00:06:52.460 --> 00:06:53.830 align:middle line:84%
Can you talk about
the ability to get

00:06:53.830 --> 00:06:56.140 align:middle line:84%
these models in production
internally and externally

00:06:56.140 --> 00:06:57.640 align:middle line:90%
within the organization?

00:06:57.640 --> 00:06:58.750 align:middle line:84%
JONATHAN WEXLER:
Absolutely, Lorry.

00:06:58.750 --> 00:07:00.430 align:middle line:84%
As the scale and
demand increases

00:07:00.430 --> 00:07:02.830 align:middle line:84%
for these properly
tuned analytic models,

00:07:02.830 --> 00:07:05.941 align:middle line:84%
having an effective deployment
strategy and architecture

00:07:05.941 --> 00:07:08.650 align:middle line:84%
is a key component in
the analytics lifecycle.

00:07:08.650 --> 00:07:10.650 align:middle line:84%
Let me take you
through a few steps.

00:07:10.650 --> 00:07:12.760 align:middle line:90%


00:07:12.760 --> 00:07:14.940 align:middle line:84%
So if I go back to
my pipeline, let's

00:07:14.940 --> 00:07:18.310 align:middle line:84%
take a look at how we can deploy
these models very quickly.

00:07:18.310 --> 00:07:21.950 align:middle line:84%
So if I go over to the
pipeline comparison area,

00:07:21.950 --> 00:07:24.469 align:middle line:84%
this screen will
compare the models

00:07:24.469 --> 00:07:27.080 align:middle line:84%
across the different
pipelines that I've built.

00:07:27.080 --> 00:07:29.420 align:middle line:84%
It will automatically pick
the best model for me.

00:07:29.420 --> 00:07:32.023 align:middle line:84%
Now I can override this
pipeline comparison,

00:07:32.023 --> 00:07:34.300 align:middle line:84%
I can even add in
challengers, but once I

00:07:34.300 --> 00:07:37.070 align:middle line:84%
have decided which model
or models I want to deploy,

00:07:37.070 --> 00:07:39.958 align:middle line:84%
the system enables you to
deploy these in the method

00:07:39.958 --> 00:07:41.681 align:middle line:90%
that you see fit.

00:07:41.681 --> 00:07:44.680 align:middle line:84%
You can publish these models
to databases or to Hadoop

00:07:44.680 --> 00:07:47.202 align:middle line:90%
with one click with no rewrite.

00:07:47.202 --> 00:07:49.090 align:middle line:84%
I can instantly
score these models

00:07:49.090 --> 00:07:51.350 align:middle line:84%
once they're pushed
to the database.

00:07:51.350 --> 00:07:54.127 align:middle line:84%
This allows for tremendous
flexibility in building

00:07:54.127 --> 00:07:56.020 align:middle line:90%
and testing models instantly.

00:07:56.020 --> 00:07:58.635 align:middle line:84%
I can also download
a scoring API.

00:07:58.635 --> 00:08:02.890 align:middle line:84%
This API call is given to
a user in multiple ways.

00:08:02.890 --> 00:08:05.140 align:middle line:84%
It's given to them
in a SAS wrapper,

00:08:05.140 --> 00:08:07.720 align:middle line:84%
but it's also
wrapped with Python.

00:08:07.720 --> 00:08:09.610 align:middle line:84%
So users have the
ability to call

00:08:09.610 --> 00:08:12.610 align:middle line:84%
these methods and the
score code in the method

00:08:12.610 --> 00:08:14.287 align:middle line:90%
that they see fit.

00:08:14.287 --> 00:08:16.120 align:middle line:84%
And if you had a
web application,

00:08:16.120 --> 00:08:19.105 align:middle line:84%
we'll even give you
the REST API call.

00:08:19.105 --> 00:08:22.160 align:middle line:84%
So you have the ability to
call these models in the way

00:08:22.160 --> 00:08:23.648 align:middle line:90%
that you see fit.

00:08:23.648 --> 00:08:26.370 align:middle line:84%
LORRY HARDT: Jonathan,
thanks so much for sharing

00:08:26.370 --> 00:08:29.200 align:middle line:84%
SAS Visual Data Mining and
Machine Learning with us today.

00:08:29.200 --> 00:08:31.320 align:middle line:84%
JONATHAN WEXLER: You're
very welcome, Lorry.

00:08:31.320 --> 00:08:33.170 align:middle line:84%
LORRY HARDT: Thanks
for watching.

00:08:33.170 --> 00:08:39.659 align:middle line:90%