Improving Home Depot Search Relevance

Amy (Yujing) Ma, Brett Amdur and Christopher Redino

Posted on Mar 22, 2016

Contributed by Amy(Yujing) Ma, Brett Amdur, Christopher Redino. They are currently in the NYC Data Science Academy 12 week full time Data Science Bootcamp program taking place between January 11th to April 1st, 2016. This post is based on their machine learning project (due on the 8th week of the program).

Given only raw text as input, our goal is to predict the relevancy of products to search results at the Home Depot website. Our strategy is a little different from most other teams in this Kaggle competition, where we generated a workflow that starts with text cleaning, passes through feature engineering and ends with model selection and parameter tuning in the attempt to stand out among thousands of competitors.

Feature Engineering

One interesting aspect of this project was that "feature engineering" here was essentially equivalent to "feature creation." That's because the data set that Home Depot provided contained no actual features that we could use as inputs to a model. Instead, our task was to take the data provided (search queries and product titles/descriptions/attributes) and use that data to derive all the features to use as predictors.

From the very beginning of the feature engineering process, our primary challenge was relatively clear: fix the upper left problem. The upper left problem refers to a recurring issue: any single feature we used as a predictor during our simple exploratory analysis performed reasonably well at higher values, but abysmally at lower values. In other words, the upper left of a correlation plot was always too heavily populated. The plot on the right is an example. Using training set data, the x axis shows the number of words in the search term of an observation that match the product title, and the y axis shows the associated relevance score for that observation. It is not surprising that higher match count scores generate higher relevance scores. What might be surprising is that the opposite is not true: lower match scores were just as likely to generate low relevance scores as high ones. We surmised that success in this competition might depend our ability to find features (or sets of features) that didn't have such wide dispersion in their outputs at lower values of the feature.

From the very beginning of the feature engineering process, our primary challenge was relatively clear: fix the upper left problem.

Ultimately, the features we fed into our model fell into four categories, shown at left. "Direct Match" features are relatively straightforward. They track "hits": search term words and phrases that matched words in the target variables (i.e. title, description, and brand names). Ratio features use the percentage of words that are hits (in, for example, the product description), and Length features refer to the number of words in the variable's content.

The last category of features is probably worth some explanation. Certain features we designed were related only to data in the training set, and were therefore "disconnected" from the test set. For example, we devised a methodology for assigning a "word power" score to words contained in search queries. Specifically, for every word in a training set search term (after the data cleansing performed in the first phase, of course), we looked at the average relevancy score for observations where it appeared. This allowed us to create a dictionary with search word - scores as the key-value pair. We then applied this dictionary to the test set. That is, we applied the word power score for each word in the training set search queries to each word in the test set search queries. We used the sum of these word scores to create a word power score for each search in the test set.

One last point about our approach to feature engineering might be worth noting. We used R's tm package, but not for the tf-idf (term frequency - inverse document frequency) calculations for which it is often used. Instead, we found it to be an efficient tool for performing word lookups for word score calculations. Its document term matrix provided a convenient (and relatively fast) way to identify the words in the search term dictionary that also appeared in product titles. From there, it was a straightforward process to calculate the sum of word scores for each observation.

[slideshare id=59901142&doc=kagglepresentationv3-160322195719&w=650&h=350]

The Python 3 code for best model is shown below:

About Authors

Amy (Yujing) Ma

View all posts by Amy (Yujing) Ma >

Brett Amdur

Brett has spent his career at the intersection of technology, analytics, business and law. As a Fellow at NYC Data Science Academy, he is applying this diverse experience to helping organizations maximize the impact of data driven decisions....

View all posts by Brett Amdur >

Christopher Redino

The common thread through all of Christopher's endeavors is his love of problem solving, with his usual methods being analytical and computational in nature. Having learned coding at an early age, Christopher picks up new programming languages quickly...

View all posts by Christopher Redino >

Data Analysis

Injury Analysis of Soccer Players with Python

Machine Learning

Ames House Prices Predictions

Python

US Honey Production Analysis With Python (1998-2012)

Machine Learning

The Ames Data Set: Sales Price Tackled With Diverse Models

Python

EDA and machine learning Ames housing price prediction project

Cancel reply

You must be logged in to post a comment.

Kaggle Competition "Home Depot Product Search Relevance" | YoutubePro September 9, 2017

[…] Learn more: http://nycdatascience.edu/blog/studen… […]

Trista April 15, 2017

Hi friends, how is everything, and what you desire to say regarding this article, in my view its actually remarkable for me.

Improving Home Depot Search Relevance

Feature Engineering

About Authors

Amy (Yujing) Ma

Brett Amdur

Christopher Redino

Related Articles

Leave a Comment

Cancel reply

View Posts by Categories

Our Recent Popular Posts

View Posts by Tags

NYC Data Science Academy

Get detailed curriculum information about our
amazing bootcamp!

Offerings

About

SOCIAL MEDIA

Improving Home Depot Search Relevance

Feature Engineering

About Authors

Amy (Yujing) Ma

Brett Amdur

Christopher Redino

Related Articles

Leave a Comment

Cancel reply

View Posts by Categories

Our Recent Popular Posts

View Posts by Tags

NYC Data Science Academy

Get detailed curriculum information about our amazing bootcamp!

Offerings

About

SOCIAL MEDIA

Get detailed curriculum information about our
amazing bootcamp!