Записки программиста, обо всем и ни о чем. Но, наверное, больше профессионального.

2015-11-02

Week 4, Categorical data and OHE

Курс Scalable Machine Learning. Hadoop, Apache Spark, Python, ML -- вот это всё.

Продолжаю конспектировать пройденный курс. Неделя 4.
В прошлый раз было про применение логистической регрессии для получения вероятностных оценок.
Теперь займемся нечисловыми фичами (наблюдениями) и кодированием таких фич.

CATEGORICAL DATA AND ONE-HOT-ENCODING
In this segment, we'll talk about how
we can deal with categorical data when training a model.
In particular, we'll introduce one-hot encoding,
which is a standard technique for converting
categorical features into numerical ones.

В чем проблема: делая модели для предсказания, тренируя эти модели, мы имеем дело с числами. Перемножение матриц, да.
Но часто данные предметной области выражены не числами а, скажем, словами.
Выходит, перед тем как строить модели, надо исходный набор данных трансформировать в набор чисел. Вот об этом и поговорим.


Как вариант, можно пользоваться методами, работающими с нечисловыми данными (Desicion Trees, Naive Bayes). Но это ограничивает наши возможности.

Посмотрим.
Есть два основных нечисловых вида фичей:
– категорные (пол, страна проживания, язык, …)
– ординарные (плохо, хорошо, отлично; низко, средне, высоко; …), отличаются наличием внутреннего порядка
Про ординарные не будем, это отдельная тема. Поговорим про категорные.

First, we have categorical features.
These are features that contain two or more categories,
and these categories have no intrinsic ordering.
The second common class of non-numeric features
are ordinal features.
Ordinal features, again, have two or more categories.
But here, these categories do contain
some ranking information but only
relative ordering information.

Specifically, we could represent a non-numeric feature
with a single numeric feature by representing each category
with a distinct number.
According to this strategy, we would
assign each of the four categories
to integer values ranging from 1 to 4.
At first glance, this seems like a reasonable strategy,
and, in fact, using a single numerical feature
does preserve the ordering for ordinal features.
However, doing this introduces a notion
of closeness between the ordinal categories that
didn't previously exist.

Назначая числовые значения для категорных/ординарных фич, мы дезинформируем модель, добавляя несуществовавший ранее параметр «близости» между значениями.


Так не годится.

Входит One-hot-encoding.

Идея такая:
для каждой категории добавляется переменная, которая принимает значение 0 или 1 в зависимости от значения исходной фичи. Понять это проще на примере.
Фича «Страны», может принимать значения «Аргентина», «Франция», «США».
Имеем три разных значения для фичи.
Значит, создаем три переменные (фактически вектор) и задаем им значения в зависимости от значений исходной фичи.
Допустим, в первой записи наших данных указана страна Аргентина. Тогда наш вектор будет выглядеть так: [1, 0, 0].
Для Франции это будет [0, 1, 0].

Понятно, почему One-hot-encoding?


COMPUTING AND STORING OHE FEATURES
In this segment, we'll go through an example of how
to compute one-hot encoded, or OHE features,
and we'll then discuss how to efficiently store
these OHE features.

Рассмотрим пример, датасет из трех фич (животное, цвет, диета) и трех записей (мышка, черная, -), (сошка, полосатая, мышка), (медведь, черный, рыбка).


Creating OHE features is a two-step process
where the first step involves creating an OHE dictionary.

Первым делом надо посчитать количество кодируемых категорий. Три вида животных + два цвета + две диеты = семь категорий.
Это означает, что итоговая запись будет содержать 7 фич вместо 3 в исходном наборе.
Составим словарик — соответствие, пронумеруем значения: (животное, медведь) = 0; (животное, кошка) = 1; … (цвет, черный) = 3; ...
Итоговая запись кодируется просто: берем вектор из 7 ноликов и меняем 0 на 1 там, где присутствует одна из семи категорий.


Как можно заметить, данные получаются разреженными, сплошные нолики.

Note that for a given categorical feature, at most
a single OHE dummy feature is non-zero.
In other words, most of our OHE features are equal to 0,
or our OHE features are sparse.

Это дает возможность оптимизировать хранение/обработку (Sparse Vectors).

an alternative way to store our features
is to store the indices and values for all non-zero entries
and assume that all other entries equal 0.

So back to our A1 example, this would
mean that we would store two tuples or a total of four
values.
In each of these tuples, the first component
is the feature index, and the second component
is the value of this feature.


При использовании Sparse Vectors экономия на хранении может быть весьма значительной, до двух порядков. Считать (умножать матрицы) тоже можно быстрее


Всё просто замечательно. Но только до некоторого предела.
В какой-то момент производных фич становится слишком много. Ибо в исходных данных слишком много категорий.
Это, кстати, характерно для задачи Click-through Rating.

However, a consequence of OHE features
is that they drastically increase the dimensionality
of our data, as the number of features in an OHE
representation equals the total number of categories
in our data.
This is a major issue in the context of clickthrough rate
prediction, as our initial user, ad,
and publisher features include many names and a lot of text.

OHE features can be problematic whenever
we're working with text data.
For instance, as we discussed earlier in the course,
we often use a bag of words representation
to represent documents with a vocabulary of words.
There are over one million words in the English language alone,
and thus the size of our vocabulary
can be quite large, resulting in high dimensional
representation.
Additionally, we sometimes would like
to create a richer representation of our text data
by considering all pairs of adjacent words,
using a so-called bigram representation.
This bigram representation is similar in spirit
to the idea of quadratic features discussed earlier
in the course.
Bigram representations of text documents
result in even larger feature dimensionality.

Используя OHE, для работы с текстом понадобятся векторы размером в миллионы! И это не считая биграмм, не говоря уж о триграммах.

Высокая размерность фич приводит к статистическим проблемам: нам надо больше записей для обучения, зависимости в данных теряют выраженность, размываются по размерностям.
Вычислительные проблемы тоже надо учитывать: нужно больше хранилища, больше процессорного времени, толще каналы связи.

Входит Feature Hashing

Feature hashing provides an alternative method
for reducing the dimension of OHE features.
By using hashing principles, we can
create a lower dimensional feature representation
and completely skip the process of creating an OHE dictionary.

Feature hashing preserves the sparsity of OHE features,
and is motivated by theoretical arguments.
It's also worth noting that the feature hashing can
be interpreted as an unsupervised learning
method, which is performed as a pre-processing step
for a downstream supervised learning task.


Все, что нам надо знать про хеш в данном контексте, это то, что мы имеем фиксированное количество корзин (buckets) и хеш-функция кладет объект в одну из этих корзин.

In the context of feature hashing,
each feature category is an object,
and we have fewer buckets than feature categories.
So in contrast to the OHE dictionary,
which is a one to one mapping between feature categories
and dummy features, feature hashing uses a many to one
mapping between feature categories and buckets.

Пример, про животных.

Suppose the number of buckets equals m equals 4,
and that we've already selected a hash function to map future
categories to buckets.


Принцип похож на OHE, только вместо словаря, отражающего категории 1:1 в индексы результирующего вектора, у нас хеш-функция, возвращающаяя этот индекс для всех категорий. Могут быть коллизии, не без этого.
Семь категорий кодируются в четыре бита.

It turns out that in spite of the slightly strange behavior,
feature hashing has nice theoretical properties.
In particular, many learning methods
can be interpreted as relying on training data
only in terms of pairwise inner products between the data
points.
And it turns out that under certain conditions,
hashed features lead to good approximations
of these inner products, theoretically.
Moreover, hashed features have been used in practice
and have resulted in good practical performance
on various text classification tasks.

And so in summary, hash features are a reasonable alternative
to OHE features.

Что хорошо, хеширование (препроцессинг данных) хорошо ложится на распределенную среду. Локальные вычисления не требуют коммуникаций с другими нодами, результат можно сохранить как SparseVector-ы.


Вот и вся теория четвертой недели. Далее будет лабораторка.






original post http://vasnake.blogspot.com/2015/11/week-4-categorical-data-and-ohe.html

Movement'alizer

Как говорится, была бы охота.

За 12 человекодней три молодца одинаковы с лица сделали прототип, программулину, которая определяет, что делает человек: стоит, идет или танцует.

We quickly assembled our dream team consisting of Bjørn Remseth, who took a class or two in machine learning; Audun Torp, a mathematician; and Jan Jongboom, a certified data gatherer. And with that we were off to work, having 36 hours to start building our classifier…

Идея такая: взять смартфон/планшет, снять данные с гироскопа и акселерометра, обучить (Machine Learning) модель. Применить обученную модель в аппликухе распознавания движения.

На видео показано, что белыми буковками «s» маркируется состояние «сидит/стоит», зелеными «w» состояние «идёт» и красными «d» состояние «танцует».
Веб-технологии, Python, JavaScript





original post http://vasnake.blogspot.com/2015/11/movementalizer.html

2015-11-01

Week 4, Logistic Regression probabilistic interpretation

Курс Scalable Machine Learning. Hadoop, Apache Spark, Python, ML -- вот это всё.

Продолжаю конспектировать пройденный курс. Неделя 4.
В прошлый раз было описание задачи (предсказание клика по рекламному банеру) и модели для решения (логистическая регрессия). Продолжим.

Вероятностная интерпретация логистической регресии.
In this segment, we'll introduce the probabilistic
interpretation of logistic regression.
And we'll show how this probabilistic interpretation
relates to our previous discussion on linear decision
boundaries.

Рассматривая логистическую регрессию, мы остановились на том, что модель выдавала нам да/нет в зависимости от знака результата.
Но что, если мы хотим получить значение вероятности? К примеру от 0 до 1?

However, what if we want more granular information
and, instead, want a model of a conditional probability
that a label equals 1 given some set of predictive features?

Например, вероятность дождичка в четверг


Возвращаясь к нашим рекламным банерам.
Вероятность клика в 10% это очень высокая вероятность. Учитывая контекст задачи.

If we're considering an ad that has
good historical performance, a user that has a high click rate
frequency, and a highly relevant publisher page,
we might expect a probability to be around 0.1.
And this would suggest that there's a 10% chance
that a click event will occur.
Although this number's low on an absolute scale,
recall that click events are rare,
and thus, the probability of 0.1 is a relatively high
probability in our context.

So if we want to work with a linear model,
we need to squash its output so that it'll
be in the appropriate range.
We can do this using the logistic, or sigmoid function,
which takes as input the dot product between our model
and our features and returns a value between 0 and 1
which we can interpret as a probability.


Логистическая функция, сигмоид, вгоняет результат линейной модели в рамки от 0 до 1.


В итоге, мы можем интерпретировать результат логистической регрессии как использование логистической функции для моделирования вероятности


Если же нам надо получить бинарную классификацию, типа да/нет, мы можем использовать пороговое значение, threshold.

It turns out that using this thresholding rule leads to a very natural connection with our previous discussion about decision boundaries


Когда значение сигмоид = 0.5, пороговое значение, это эквивалентно тому, что wT dot x = 0, что есть decision boundaries.

Поговорим о использовании порогового значения и предсказании вероятности.
USING PROBABILISTIC PREDICTIONS

we don't necessarily
have to set our threshold to 0.5.
Let's consider a spam detection example to see this point.

Now, imagine that we've made an incorrect spam prediction.
We can potentially make two types of mistakes.
The first type, which we call a false positive,
occurs when we classify a legitimate email as spam.
The second type, which we call a false negative,
occurs when we classify a spam email as a legitimate one.

В случае спама, false positive вреднее, чем false negative. Пропавшее письмо не все догадаются поискать в папке «спам».

These two types of errors are typically
at odds with each other.
Or, in other words, we often trade off one type of error
for the other.

Можно сдвинуть порог с 0.5 на другое значение. Но как определить оптимальный порог?

Надо попробовать набор значения порога и построить график

One natural way to do this is to use a ROC plot, which
visualizes the trade-offs we make
as we change our threshold.
In particular, a ROC plot focuses on two standard metrics
that are often at odds with one another, namely
the false positive rate, or FPR, and the true positive rate,
or TPR.

Using our spam detection application
as a running example, we can interpret
FPR as the percentage of legitimate emails
that are incorrectly predicted as spam.
And we can view TPR as the percentage of spam emails
that are correctly predicted as spam.

В идеальном случае мы должны получить FPR = 0%, TPR = 100%
В случае, когда решение спам/не спам принимается произвольно, мы получим FPR = 50%, TPR = 50%

we can use our classifier to generate
conditional probabilities for observations in our validation
set. And once we have these probabilities,
we can generate a ROC plot by varying
the threshold we use to convert these probabilities to class
predictions, and computing FPR and TPR
at each of these thresholds.


Допустим, мы готовы мириться с FPR = 10%, что означает, 10% нужных писем могут попасть в спам.

Но в некоторых случаях нам не нужна бинарная классификация а надо использовать напрямую значения вероятностей
As we know, click events are rare.
And as a result, our predicted click probabilities
will be uniformly low.

We may also want to combine these predictions
with other information and thus don't want a threshold,
as we'll lose this fine-grained information by doing so.




original post http://vasnake.blogspot.com/2015/11/week-4-logistic-regression_1.html

Week 4, Logistic Regression

Курс Scalable Machine Learning. Hadoop, Apache Spark, Python, ML -- вот это всё.

Продолжаю конспектировать пройденный курс. В прошлый раз было описание лабораторки Lab 3. Продолжим.

Неделя 4: Логистическая регрессия;
предсказание кликов по рекламным банерам. В смысле, щелкнет посетитель по банеру или нет.

Logistic Regression and Click-through Rate Prediction.
Topics: Online advertising, linear classification, logistic regression, working with probabilistic predictions, categorical data and one-hot-encoding, feature hashing for dimensionality reduction.

Онлайн реклама как задача для large-scale ML комплекса.

Характеристики задачи таковы:
– миллиарды показов банеров миллионам людей
– менее 1% захода по рекламе
– легко собрать много labeled data

Основные действующие лица:
– паблишеры, показывают рекламу и берет деньги с рекламодателей
– рекламодатели, платят за рекламу и ждут привлечения покупателей
– сводня, прослойка между паблишером и рекламодателем, решает задачу, какой именно банер показать конкретному пользователю на конкретном сайте.

Как показатель успеха (за что платит рекламодатель) часто принимается клик по банеру.

The most common metric for success
is clickthrough rate prediction.
And this is what we'll be focusing on

В итоге, сводня должна постоянно решать (в режиме реального времени) задачу:
оценить вероятность того, что данный посетитель (со всей своей историей браузинга, поиска и поведения в целом) щелкнет по конкретному банеру (у конкретного паблишера, от конкретного рекламодателя, по конкретному товару, со всеми их характеристиками).

This problem involves estimating a conditional probability,
or the probability of a click event
occurring given a set of predictive features

Данных много, но
However, the data is high-dimensional, sparse,
and skewed.
Additionally, only a small fraction of these ads
actually get clicked on by users,
and thus we have a heavily biased label training set.

In summary, the goal of the clickthrough rate prediction
problem is to use large amounts of data
to estimate the conditional probability that a user will
click on an ad, given features about the user, the ad,
and the publisher.

С задачей разобрались. Теперь посмотрим на инструмент её решения.

LINEAR CLASSIFICATION AND LOGISTIC REGRESSION

In this segment, we'll talk about linear classification
and relate it to the problem of linear regression.
We'll also introduce logistic regression,
which is a particular type of linear classification model.

Click-through rate prediction is another example (классификации, после примера с электропочтой – спам/не спам).
Here, our observations are (user, ad, publisher) triples.
Our labels are not-click and click.
And given a set of labeled observations,
we want to predict whether a new (user ad publisher) triple
will result in a click.

Точнее, оценка нужна не в виде кликнет/не кликнет а в виде вероятность(кликнет) в диапазоне 0 до 1.

В качестве модели будем рассматривать одну из простейших моделей линейной классификации.
Для начала вспомним линейную регрессию.

In linear regression, our key assumption
is that there's a linear mapping between the features
and the label.


Обратите внимание на доп. коэффициент w0 и как, добавив 1 к фичам, мы записали сумму произведений в матричном виде.

Почему линейное уравнение?
– это просто
– часто неплохо работает
– можно произвольно усложнять добавляя в уравнение фичи высших порядков

Осталось придумать, как сделать классификатор на линейном уравнении.
Нам нужен бинарный ответ: щелкнет/не щелкнет.

To achieve this goal, we can't directly
use the linear regression model since linear regression
predictions are real values rather than class predictions.

Ответ заключается в слове threshold




What this rule says is that a w transpose x is positive,
then our prediction y hat is 1.
And a w transpose x is negative, our prediction is negative 1.
And when w transpose x equals 0, we
can arbitrarily choose the prediction
to be 1 or negative 1.
Another way to say this is that w transpose x equals 0
is our decision boundary.

То есть, линия, рисуемая по полученному уравнению, это граница между да и нет.



This decision boundary is defined by the equation w transpose dot x equals 0.

ОК, глядя на знак результата, мы получаем ответ да/нет.
Возникает следующий вопрос: как оценить точность результата?
В линейной регрессии мы использовали среднеквадратичное отклонение – чем больше разница между лейблом и вычисленным результатом, тем больше ошибка/штраф.
Как быть с бинарной классификацией?
Допустим, так: угадал – штраф 0; не угадал – штраф 1. Тоже бинарненько.

Тогда можно сделать так: определим z как произведение лабели y на вычисленное значение wT * x
И определим штрафную функцию как 1 если z < 0 и 0 если z >= 0.


Теперь, когда с моделью и ее оценкой мы определились, можно подумать о том, как тренировать модель.
Суть тренировки в том, чтобы подобрать такой вектор параметров w, который дает минимальную ошибку (сумму значений штрафной функции) на тренировочном наборе данных x, y.

Задача оптимизации.


Но тут есть грабли: штрафная функция не convex

Unfortunately, this is not a convex optimization problem
as a 0-1 loss is not a convex function.
And thus, this is a hard problem to solve.

Надо искать штрафную функцию такую, для которой можно будет найти глобальный минимум, решая задачу оптимизации.

But all hope isn't lost, because instead of using the 0-1 loss
in our optimization problem, we can use a surrogate loss that
is convex and also provides a good approximation to 0-1 loss


For instance, the hinge loss is used to define the SVM model.
The exponential loss is used to describe the Adaboost model.
And the logistic loss is used to define the logistic regression
model.
In particular, we'll focus on logistic regression

Мы сосредоточимся на функции логистической регрессии, хотя SVM выглядит более обещающе.


Как результат применения convex функции, мы можем использовать стандартный метод оптимизации: gradient descent.


И, для борьбы с оверфиттингом, добавляем регуляризацию:

Finally, similar to ridge regression,
we can use a regularized version of logistic regression.
In this regularized version, we're
trading off between two terms.
The first is the log loss on the training data.
And the second is the model complexity.



we introduce a user specified hyper parameter, lambda,
to determine the trade-off between these two terms

В общем-то, это всё, что вам необходимо знать про логистическую регрессию.

В следующий раз продолжим конспектировать материалы 4-ой недели. Будет интерпретация логистической регрессии с точки зрения вероятности.




original post http://vasnake.blogspot.com/2015/11/week-4-logistic-regression.html

2015-10-09

Coffee

Про CoffeeScript

Verbally Readable !== Quicker Comprehension
Verbally readable is not exactly the same as quicker comprehension.
See what I did there? The main argument for using CoffeeScript is that it’s more readable and should therefore be easier to write, debug and maintain. However, reading something and comprehending something are very different

Статья не новая, но стоит прочтения.
В основном, в ней про темную сторону CoffeeScript.
Понятно, не был бы JavaScript так ужасен, не было бы CoffeeScript. Но иногда кое-кому лучше продолжать говнокодить на JS.

Лично мне очень понравилась часть статьи, где подробно разбирается постулат, что легкочитаемый код вовсе не означает его легкопонимаемость.
Вот уж тема достойная дискуссии, особенно в контексте JS.


И на закуску:



original post http://vasnake.blogspot.com/2015/10/coffee.html

2015-10-03

Week 3, lab


В прошлый раз я рассказал, что входило в лекции по теории, третья неделя:
Linear Regression and Distributed Machine Learning Principles 

Теперь посмотрим на лабораторку lab3

This lab covers a common supervised learning pipeline, using a subset of the Million Song Dataset from theUCI Machine Learning Repository. Our goal is to train a linear regression model to predict the release year of a song given a set of audio features.

Lab 3: Millionsong Regression Pipeline. Develop an end-to-end linear regression pipeline to predict the release year of a song given a set of audio features. You will implement a gradient descent solver for linear regression, use Spark's machine Learning library ( mllib) to train additional models, tune models via grid search, improve accuracy using quadratic features, and visualize various intermediate results to build intuition

Построим модель (линейная регрессия, gradient descent), потренируем ее и будем угадывать год выхода песни по ее свойствам. Это в принципе. А в кожухе, в лабе будет еще много всякого, типа извлечения дополнительный фичей, построения альтернативных моделей, поиск и тюнинг гиперпараметров для ridge regression и проч.

The notebook can be found on GitHub at the following links:

ОК, скачиваем нотебук, запускаем виртмашину
valik@snafu:~$  pushd ~/sparkvagrant/
valik@snafu:~/sparkvagrant$  vagrant up
И понеслась http://localhost:8001/

Part 1: Read and parse the initial dataset
Visualization 1: Features
Visualization 2: Shifting labels
Part 2: Create and evaluate a baseline model
Visualization 3: Predicted vs. actual
Part 3: Train (via gradient descent) and evaluate a linear regression model
Visualization 4: Training error
Part 4: Train using MLlib and tune hyperparameters via grid search
Visualization 5: Best model's predictions
Visualization 6: Hyperparameter heat map
Part 5: Add interactions between features

Документация к используемым API:

1. Загрузка и парсинг датасета.

На входе у нас текстовый файл, CSV, где в первом поле указан год (label) и остальные поля содержат свойства музыкального трека, в виде чисел.
numPartitions = 2
rawData = sc.textFile(fileName, numPartitions)
numPoints = rawData.count()
print numPoints
samplePoints = rawData.take(5)
print samplePoints
6724
[u'2001.0,0.884123733793,0.610454259079,0.600498416968,0.474669212493,0.247232680947,0.357306088914,0.344136412234,0.339641227335,0.600858840135,0.425704689024,0.60491501652,0.419193351817', u'2001.0,0.854411946129,0.604124786151,0.593634078776,0.495885413963,0.266307830936,0.261472105188,0.506387076327,0.464453565511,0.665798573683,0.542968988766,0.58044428577,0.445219373624', ...

rawData это будет RDD с сырыми данными. Почти 7 тыщ записей.

Распарсим данные и создадим RDD содержащий объекты LabeledPoint.
Сначала напишем функцию конвертации строки из исходного датасета в LabeledPoint
from pyspark.mllib.regression import LabeledPoint
import numpy as np
def parsePoint(line):
    """Converts a comma separated unicode string into a `LabeledPoint`.

    Args:
        line (unicode): Comma separated unicode string where the first element is the label and the
            remaining elements are features.

    Returns:
        LabeledPoint: The line is converted into a `LabeledPoint`, which consists of a label and
            features.
    """
    label = подсказывать нельзя
    features = ???
    res = LabeledPoint(label, features)
    return res

parsedSamplePoints = [parsePoint(x) for x in samplePoints]
firstPointFeatures = parsedSamplePoints[0].features
firstPointLabel = parsedSamplePoints[0].label
print firstPointFeatures, firstPointLabel

Для осознания, какие данные у нас в датасете, был построен график:
Then we will look at the raw features for 50 data points by generating a heatmap that visualizes each feature on a grey-scale and shows the variation of each feature across the 50 sample data points. The features are all between 0 and 1, with values closer to 1 represented via darker shades of grey.

Продолжаем играться с данными. Найдем диапазон лет (лабелей) для песен из датасета. Применим функцию parsePoint для создания нового RDD, выцепим оттуда labels и возьмем min, max:
parsedDataInit = rawData.map(<FILL IN>)
onlyLabels = parsedDataInit.map(<FILL IN>)
minYear = <FILL IN>
maxYear = <FILL IN>
print maxYear, minYear
2011.0 1922.0

Диапазон от 1922 года до 2011. Для обучения модели с этим надо что-то сделать. Нормализовать, чтобы минимальное значение было 0.
parsedData = parsedDataInit.<FILL IN>
# Should be a LabeledPoint
print type(parsedData.take(1)[0])
# View the first point
print '\n{0}'.format(parsedData.take(1))
Блин, с этим договором не давать подсказок, фак! Невозможно ничего показать.
В общем, тут создается новый RDD (через map), где из каждой лабели вычтено минимальное значение.

Теперь наш датасет надо разделить на сеты: тренировочный, кросс-валидации и тестовый. Для разделения используем
Датасеты закешируем в оперативке, ибо обращаться к ним мы будем постоянно.
weights = [.8, .1, .1]
seed = 42
parsedTrainData, parsedValData, parsedTestData = parsedData.randomSplit(weights, seed)
parsedTrainData.cache()
parsedValData.cache()
parsedTestData.cache()
nTrain = parsedTrainData.count()
nVal = parsedValData.count()
nTest = parsedTestData.count()
print nTrain, nVal, nTest, nTrain + nVal + nTest
print parsedData.count()
5371 682 671 6724
6724


2. Создание и оценка референсной модели.

Референсная (baseline) модель, это модель самая простая, нужная для того, чтобы иметь ориентир хоть какой-то.
Самым простым вариантом была признана константа, равная среднему арифметическому
averageTrainYear = (parsedTrainData
    .map(lambda x: x.label)
    .mean()
)
print averageTrainYear
53.9316700801

Напишем пару функций для оценки модели методом среднеквадратичной ошибки
def squaredError(label, prediction):
    """Calculates the the squared error for a single prediction.

    Args:
        label (float): The correct value for this observation.
        prediction (float): The predicted value for this observation.

    Returns:
        float: The difference between the `label` and `prediction` squared.
    """
    import math
    res = math.pow(label - prediction, 2)
    return res

def calcRMSE(labelsAndPreds):
    """Calculates the root mean squared error for an `RDD` of (label, prediction) tuples.

    Args:
        labelsAndPred (RDD of (float, float)): An `RDD` consisting of (label, prediction) tuples.

    Returns:
        float: The square root of the mean of the squared errors.
    """
    res = (labelsAndPreds
           .map(lambda (x, y): квадрат ошибки)
           .reduce(lambda x, y: просуммировать)
    )
    import math
    return math.sqrt(поделить сумму на количество)

labelsAndPreds = sc.parallelize([(3., 1.), (1., 2.), (2., 2.)])
# RMSE = sqrt[((3-1)^2 + (1-2)^2 + (2-2)^2) / 3] = 1.291
exampleRMSE = calcRMSE(labelsAndPreds)
print exampleRMSE
1.29099444874

Посмотрим теперь, какую ошибку в годах дает наша базовая модель для песенного датасета (трех: тренировочного, кросс-валидации и тестового).
labelsAndPredsTrain = parsedTrainData.map(lambda x: (актуальная лабель и предсказание))
rmseTrainBase = calcRMSE(labelsAndPredsTrain)
labelsAndPredsVal = parsedValData.map(lambda x: (лабель и предсказание))
rmseValBase = calcRMSE(labelsAndPredsVal)
labelsAndPredsTest = parsedTestData.map(lambda x: (лабель и предсказание))
rmseTestBase = calcRMSE(labelsAndPredsTest)
print 'Baseline Train RMSE = {0:.3f}'.format(rmseTrainBase)
print 'Baseline Validation RMSE = {0:.3f}'.format(rmseValBase)
print 'Baseline Test RMSE = {0:.3f}'.format(rmseTestBase)
Baseline Train RMSE = 21.306
Baseline Validation RMSE = 21.586
Baseline Test RMSE = 22.137

3. Тренировка и оценка самопальной модели линейной регресии.

Размялись? Теперь пора занятся настоящей моделью, линейная регрессия обученная через gradient descent.

Напишем функцию вычисления апдейта для одной итерации
def gradientSummand(weights, lp):
    """Calculates the gradient summand for a given weight and `LabeledPoint`.

    Note:
        `DenseVector` behaves similarly to a `numpy.ndarray` and they can be used interchangably
        within this function.  For example, they both implement the `dot` method.

    Args:
        weights (DenseVector): An array of model weights (betas).
        lp (LabeledPoint): The `LabeledPoint` for a single observation.

    Returns:
        DenseVector: An array of values the same length as `weights`.  The gradient summand.
    """
    w = транспонированные веса
    x = фичи лабельпойнта
    y = лабель лабельпойнта
    res = (w.dot(x) - y) * x
    return res

exampleW = DenseVector([1, 1, 1])
exampleLP = LabeledPoint(2.0, [3, 1, 4])
# gradientSummand = (dot([1 1 1], [3 1 4]) - 2) * [3 1 4] = (8 - 2) * [3 1 4] = [18 6 24]
summandOne = gradientSummand(exampleW, exampleLP)
print summandOne
[18.0,6.0,24.0]

Теперь нужна функция вычисляющая предсказание. Нужно просто умножить два вектора – веса и фичи, dot product
def getLabeledPrediction(weights, observation):
    """Calculates predictions and returns a (label, prediction) tuple.

    Note:
        The labels should remain unchanged as we'll use this information to calculate prediction
        error later.

    Args:
        weights (np.ndarray): An array with one weight for each features in `trainData`.
        observation (LabeledPoint): A `LabeledPoint` that contain the correct label and the
            features for the data point.

    Returns:
        tuple: A (label, prediction) tuple.
    """
    label = лабель наблюдения
    w = транспонированные веса
    x = фичи наблюдения
    prediction = веса дотпродакт фичи
    res = (label, prediction)
    return res

weights = np.array([1.0, 1.5])
predictionExample = sc.parallelize([LabeledPoint(2, np.array([1.0, .5])),
                                    LabeledPoint(1.5, np.array([.5, .5]))])
labelsAndPredsExample = predictionExample.map(lambda lp: getLabeledPrediction(weights, lp))
print labelsAndPredsExample.collect()
[(2.0, 1.75), (1.5, 1.25)]


Теперь самое вкусное, gradient descent
def linregGradientDescent(trainData, numIters):
    """Calculates the weights and error for a linear regression model trained with gradient descent.

    Note:
        `DenseVector` behaves similarly to a `numpy.ndarray` and they can be used interchangably
        within this function.  For example, they both implement the `dot` method.

    Args:
        trainData (RDD of LabeledPoint): The labeled data for use in training the model.
        numIters (int): The number of iterations of gradient descent to perform.

    Returns:
        (np.ndarray, np.ndarray): A tuple of (weights, training errors).  Weights will be the
            final weights (one weight per feature) for the model, and training errors will contain
            an error (RMSE) for each iteration of the algorithm.
    """
    # The length of the training data
    n = trainData.count()
    # The number of features in the training data
    d = len(trainData.take(1)[0].features)
    w = np.zeros(d)
    alpha = 1.0
    # We will compute and store the training error after each iteration
    errorTrain = np.zeros(numIters)
    for i in range(numIters):
        # Use getLabeledPrediction from (3b) with trainData to obtain an RDD of (label, prediction)
        # tuples.  Note that the weights all equal 0 for the first iteration, so the predictions will
        # have large errors to start.
        labelsAndPredsTrain = trainData.map getLabeledPrediction
        errorTrain[i] = calcRMSE(labelsAndPredsTrain)

        # Calculate the `gradient`.  Make use of the `gradientSummand` function you wrote in (3a).
        # Note that `gradient` sould be a `DenseVector` of length `d`.
        gradient = trainData.map gradientSummand reduce просуммировать

        # Update the weights
        alpha_i = alpha / (n * np.sqrt(i+1))
        w -= альфа помноженная на градиент
    return w, errorTrain

# create a toy dataset with n = 10, d = 3, and then run 5 iterations of gradient descent
# note: the resulting model will not be useful; the goal here is to verify that
# linregGradientDescent is working properly
exampleN = 10
exampleD = 3
exampleData = (sc
               .parallelize(parsedTrainData.take(exampleN))
               .map(lambda lp: LabeledPoint(lp.label, lp.features[0:exampleD])))
print exampleData.take(2)
exampleNumIters = 5
exampleWeights, exampleErrorTrain = linregGradientDescent(exampleData, exampleNumIters)
print exampleWeights
[LabeledPoint(79.0, [0.884123733793,0.610454259079,0.600498416968]), LabeledPoint(79.0, [0.854411946129,0.604124786151,0.593634078776])]
[ 48.88110449 36.01144093 30.25350092]


Теперь применим эти функции для тренировки линрег модели, получим веса. И проверим точность модели на датасете кросс-валидации.
numIters = 50
weightsLR0, errorTrainLR0 = linregGradientDescent(распарсенные трен.данные, numIters)
labelsAndPreds = parsedValData.map(lambda x: getLabeledPrediction(weightsLR0, x))
rmseValLR0 = calcRMSE(labelsAndPreds)
print 'Validation RMSE:\n\tBaseline = {0:.3f}\n\tLR0 = {1:.3f}'.format(rmseValBase,
                                                                       rmseValLR0)
Validation RMSE:
Baseline = 21.586
LR0 = 19.192

Фиговенькая точность, хотя и лучше чем у базовой модели.


4. Тренировка моделей в/из MLlib, поиск гиперпараметров через grid search.

Теперь воспользуемся библиотечной функциональностью
для построения модели линейной регрессии с регуляризацией и доп.коэффициентом/смещением (intercept).
from pyspark.mllib.regression import LinearRegressionWithSGD
# Values to use when training the linear regression model
numIters = 500  # iterations
alpha = 1.0  # step
miniBatchFrac = 1.0  # miniBatchFraction
reg = 1e-1  # regParam
regType = 'l2'  # regType
useIntercept = True  # intercept
firstModel = LinearRegressionWithSGD.тренировать(
    распарсенные трен.данные,
    iterations=numIters,
    step=alpha,
    miniBatchFraction=miniBatchFrac,
    regParam=reg,
    regType=regType,
    intercept=useIntercept
)
# weightsLR1 stores the model weights; interceptLR1 stores the model intercept
weightsLR1 = firstModel.weights
interceptLR1 = firstModel.intercept
print weightsLR1, interceptLR1
[16.682292427,14.7439059559,-0.0935105608897,6.22080088829,4.01454261926,-3.30214858535,11.0403027232,2.67190962854,7.18925791279,4.46093254586,8.14950409475,2.75135810882] 13.3335907631

Воспользуемся для предсказания библиотечной функцией
samplePoint = parsedTrainData.take(1)[0]
samplePrediction = firstModel.предсказать(samplePoint.features)
print samplePrediction
56.8013380112

Модель есть, натренирована. Проверим, насколько она ошибается.
labelsAndPreds = распарсенные валид.данные map(lambda x: (x.label, firstModel.предсказать(x.features)))
rmseValLR1 = calcRMSE(labelsAndPreds)
print ('Validation RMSE:\n\tBaseline = {0:.3f}\n\tLR0 = {1:.3f}' +
       '\n\tLR1 = {2:.3f}').format(rmseValBase, rmseValLR0, rmseValLR1)
Validation RMSE:
Baseline = 21.586
LR0 = 19.192
LR1 = 19.691

Забавно, ошибка больше, чем у модели обученной на меньшем количестве итераций. Очевидно, регуляризация была подобрана неудачно.

Поэтому что? Правильно, надо занятся подбором гиперпараметров.
Методом Grid Search подберем размер регуляризатора

bestRMSE = rmseValLR1
bestRegParam = reg
bestModel = firstModel

numIters = 500
alpha = 1.0
miniBatchFrac = 1.0
for reg in (1e-10, 1e-5, 1):
    model = LinearRegressionWithSGD.train(parsedTrainData, numIters, alpha,
                                          miniBatchFrac, regParam=reg,
                                          regType='l2', intercept=True)
    labelsAndPreds = parsedValData.map(lambda lp: (lp.label, model.predict(lp.features)))
    rmseValGrid = calcRMSE(labelsAndPreds)
    print rmseValGrid

    if rmseValGrid < bestRMSE:
        bestRMSE = rmseValGrid
        bestRegParam = reg
        bestModel = model
rmseValLRGrid = bestRMSE

print ('Validation RMSE:\n\tBaseline = {0:.3f}\n\tLR0 = {1:.3f}\n\tLR1 = {2:.3f}\n' +
       '\tLRGrid = {3:.3f}').format(rmseValBase, rmseValLR0, rmseValLR1, rmseValLRGrid)
17.0171700716
17.0175981807
23.8007746698
Validation RMSE:
Baseline = 21.586
LR0 = 19.192
LR1 = 19.691
LRGrid = 17.017

Вот, другое дело. Теперь предсказания точнее на два с лихуем года. А по сравнению с базовой моделью, так вообще, 4+


Теперь проверим, что будет при других значениях альфа (шаг gradient descent)
reg = bestRegParam
modelRMSEs = []

for alpha in (1e-5, 10):
    for numIters in (5, 500):
        model = LinearRegressionWithSGD.train(parsedTrainData, numIters, alpha,
                                              miniBatchFrac, regParam=reg,
                                              regType='l2', intercept=True)
        labelsAndPreds = parsedValData.map(lambda lp: (lp.label, model.predict(lp.features)))
        rmseVal = calcRMSE(labelsAndPreds)
        print 'alpha = {0:.0e}, numIters = {1}, RMSE = {2:.3f}'.format(alpha, numIters, rmseVal)
        modelRMSEs.append(rmseVal)
alpha = 1e-05, numIters = 5, RMSE = 56.970
alpha = 1e-05, numIters = 500, RMSE = 56.893
alpha = 1e+01, numIters = 5, RMSE = 355124752.221
alpha = 1e+01, numIters = 500, RMSE = 33110728225678989137251839534941311488306376280786874812632607716020409015644103457295574688229365257796099669135380709376.000

Ну нихрена ж себе! Видали, что будет если шаг слишком большой?


5. Добавим в модель производные фичи, попарно связав исходные.

Now, we will add features that capture the two-way interactions between our existing features. Write a function twoWayInteractions that takes in a LabeledPoint and generates a new LabeledPoint that contains the old features and the two-way interactions between them. Note that a dataset with three features would have nine ( 3^2 ) two-way interactions

import itertools

def twoWayInteractions(lp):
    """Creates a new `LabeledPoint` that includes two-way interactions.

    Note:
        For features [x, y] the two-way interactions would be [x^2, x*y, y*x, y^2] and these
        would be appended to the original [x, y] feature list.

    Args:
        lp (LabeledPoint): The label and features for this observation.

    Returns:
        LabeledPoint: The new `LabeledPoint` should have the same label as `lp`.  Its features
            should include the features from `lp` followed by the two-way interaction features.
    """
    newfeatures = itertools.перемножить(lp.features, repeat=2)
    newfeatures = [x*y for x,y in newfeatures]
    label = lp.label
    features = np.горизонтальный стек((lp.features, newfeatures))
    res = LabeledPoint(label, features)
    return res

print twoWayInteractions(LabeledPoint(0.0, [2, 3]))

# Transform the existing train, validation, and test sets to include two-way interactions.
trainDataInteract = parsedTrainData.для каждого lp(lambda x: twoWayInteractions(x))
valDataInteract = parsedValData.для каждого lp(lambda x: twoWayInteractions(x))
testDataInteract = parsedTestData.для каждого lp(lambda x: twoWayInteractions(x))
(0.0,[2.0,3.0,4.0,6.0,6.0,9.0])

Посмотрим, что даст тренировка модели на нашем раздутом датасете. Предварительный поиск гиперпараметров был сделан за кадром.
numIters = 500
alpha = 1.0
miniBatchFrac = 1.0
reg = 1e-10

modelInteract = LinearRegressionWithSGD.train(trainDataInteract, numIters, alpha,
                                              miniBatchFrac, regParam=reg,
                                              regType='l2', intercept=True)
labelsAndPredsInteract = valDataInteract.map(lambda lp: (lp.label, modelInteract.predict(lp.features)))
rmseValInteract = calcRMSE(labelsAndPredsInteract)

print ('Validation RMSE:\n\tBaseline = {0:.3f}\n\tLR0 = {1:.3f}\n\tLR1 = {2:.3f}\n\tLRGrid = ' +
       '{3:.3f}\n\tLRInteract = {4:.3f}').format(rmseValBase, rmseValLR0, rmseValLR1,
                                                 rmseValLRGrid, rmseValInteract)
Validation RMSE:
Baseline = 21.586
LR0 = 19.192
LR1 = 19.691
LRGrid = 17.017
LRInteract = 15.690

Опаньки, отыграли еще два года точности, более 10%. Отлично!

Остается только проверить последнюю, самую точную модель на тестовом наборе данных

labelsAndPredsTest = testDataInteract.map(lambda x: (x.label, modelInteract.predict(x.features)))
rmseTestInteract = calcRMSE(labelsAndPredsTest)

print ('Test RMSE:\n\tBaseline = {0:.3f}\n\tLRInteract = {1:.3f}'
       .format(rmseTestBase, rmseTestInteract))
Test RMSE:
Baseline = 22.137
LRInteract = 16.327

Видно, что ошибка больше, чем на кросс-валидационном наборе. Догадываетесь почему? Правильно, кросс-валидационный набор был использован в подборе гиперпараметров, что привело к некоторому оверфиттингу для этого набора.


Вот и вся третья лабораторка. Энджой.




original post http://vasnake.blogspot.com/2015/10/week-3-lab.html

Архив блога

Ярлыки

linux (241) python (191) citation (186) web-develop (170) gov.ru (159) video (124) бытовуха (115) sysadm (100) GIS (97) Zope(Plone) (88) бурчалки (84) Book (83) programming (82) грабли (77) Fun (76) development (73) windsurfing (72) Microsoft (64) hiload (62) internet provider (57) opensource (57) security (57) опыт (55) movie (52) Wisdom (51) ML (47) driving (45) hardware (45) language (45) money (42) JS (41) curse (40) bigdata (39) DBMS (38) ArcGIS (34) history (31) PDA (30) howto (30) holyday (29) Google (27) Oracle (27) tourism (27) virtbox (27) health (26) vacation (24) AI (23) Autodesk (23) SQL (23) humor (23) Java (22) knowledge (22) translate (20) CSS (19) cheatsheet (19) hack (19) Apache (16) Klaipeda (15) Manager (15) web-browser (15) Никонов (15) functional programming (14) happiness (14) music (14) todo (14) PHP (13) course (13) scala (13) weapon (13) HTTP. Apache (12) SSH (12) frameworks (12) hero (12) im (12) settings (12) HTML (11) SciTE (11) USA (11) crypto (11) game (11) map (11) HTTPD (9) ODF (9) Photo (9) купи/продай (9) benchmark (8) documentation (8) 3D (7) CS (7) DNS (7) NoSQL (7) cloud (7) django (7) gun (7) matroska (7) telephony (7) Microsoft Office (6) VCS (6) bluetooth (6) pidgin (6) proxy (6) Donald Knuth (5) ETL (5) NVIDIA (5) Palanga (5) REST (5) bash (5) flash (5) keyboard (5) price (5) samba (5) CGI (4) LISP (4) RoR (4) cache (4) car (4) display (4) holywar (4) nginx (4) pistol (4) spark (4) xml (4) Лебедев (4) IDE (3) IE8 (3) J2EE (3) NTFS (3) RDP (3) holiday (3) mount (3) Гоблин (3) кухня (3) урюк (3) AMQP (2) Baltic (2) ERP (2) IE7 (2) NAS (2) Naudoc (2) PDF (2) address (2) air (2) british (2) coffee (2) fitness (2) font (2) ftp (2) fuckup (2) messaging (2) notify (2) sharepoint (2) ssl/tls (2) stardict (2) tests (2) tunnel (2) udev (2) APT (1) CRUD (1) Canyonlands (1) Cyprus (1) DVDShrink (1) Jabber (1) K9Copy (1) Matlab (1) Portugal (1) VBA (1) WD My Book (1) autoit (1) bike (1) cannabis (1) chat (1) concurrent (1) dbf (1) ext4 (1) idioten (1) join (1) krusader (1) license (1) life (1) migration (1) mindmap (1) navitel (1) pneumatic weapon (1) quiz (1) regexp (1) robot (1) science (1) seaside (1) serialization (1) shore (1) spatial (1) tie (1) vim (1) Науру (1) крысы (1) налоги (1) пианино (1)