Scraping Data from Indeed.com

Posted on Feb 21, 2017
The skills the author demoed here can be learned through taking Data Science with Machine Learning bootcamp with NYC Data Science Academy.

Audience

This blog is intended for prospective job searchers to have a macro look at common job listings through indeed.com as well as examine the data on what are common pros and cons with popular companies such as Grubhub or The Washington Post.

Overview

Indeed.com is popular among many data scientist as the go to site for finding employment and company scouting.  Unsurprisingly, they are a popular site for data science practitioners to scrape.   Prospective students and  job searchers will be able to find insight on best paying career paths as well as desirable companies according to company reviews.   The blog will also look into what are features that can predict salary of a position as well as company job titles.

Finally there will be wordclouds of pros and cons from reviews written regarding popular companies as well as a sentiment analysis of polarity scores among reviews for The Washington Post.  The data was collected via scrapy spiders and code for the spider can be accessed below as well as the compiled csv's from said data.

Data on Best Pay

As of February 2017, the best paying position from indeed.com is Radiologist followed by Senior Vice President and Vice President of Sales.  There is variance among how much Vice President of Sales pays as demonstrated by the black variance bar.

Scraping Data from Indeed.com

LabCorp has the highest mean salary offers for their listings.  However, that doesn't necessarily mean they are paying the best for each position, simply that they have high paying position offers.  Open Systems Technologies, Inc has a large spread on their salary postings which indicates they likely have a wide range of position offerings.  Mitchell/Martin Inc. is interesting here as they also appear among the companies with top review scores by their employers.

Scraping Data from Indeed.com

Data on Best Sum of Review Scores

The review sum is taken from the sum of the mean score of compensation, management, work life balance, culture and job security.  There are noticeable statements from this graph.  First, there are positions where there is some variance among the positions with the highest review sum scores.  Across the board, those most satisfied with their position overall come from analysts and engineers.

BestJobReviews

Visualized below, universities and staffing agencies are enjoy the highest scores among compensation, management, work life balance, culture and job security.

Scraping Data from Indeed.com

Exploratory Data Analysis

A heatmap can be used to display that review scores among work life balance, management, job security and culture all strong correlations with each other.  Compensation is the feature which displays some correlation with all five other categories.

HeatMap

A pair plot can be created to display this visualization better especially when the data is separated by salary quantiles.  Note that the gigantic spread among salaries in the highest quantile.

pairplotquantiles

Unsurprisingly, higher ratings in compensation tend to indicate higher salary.  Also higher paying jobs tend to have lower job security.  Surprisingly those with higher salary also tend to give higher scores for company culture.  This could mean that those in better paid position tend to "buy in" their company culture and there could be some hypothesis that this could be some mercenary attitude in that loyalty can be bought.

SalaryPredictionCoefficient

Also as expected, if feature importance were to be examined on what is most likely to be an accurate predictor of job title.  Salary is by far the strongest predictor among the data set with work life, job security and culture having almost no impact.

FeaturePredictionForestV2

What's being said

A wordcloud can be constructed to illustrate what is being said among well known companies within the data set.  Among the pros, these are commonly cited words among the following companies.

WordCloudPros

The cons tend to have stronger words.  Oddly enough, there are some strange polarities with grubhub with "acceptance" being a popular word among positive reviews and "fascist" being used in a negative one.  Humorously, it has also been verified by many employees that the coffee in Bloomberg is indeed, terrible.

WordCloudCons

Finally, a sentiment analysis can be created for reviews of various companies such as the one below for The Washington Post.

SubjectivePolarityNLP

Conclusion

While this is overall a summary of a multi-facet picture of the data scraped from Indeed.com, there is a still significant amount left to explore.  As a result, the collection of this project along with accompanying code can be found in the github below for a user to peruse further analysis on their own.

https://github.com/DraceZhan/Indeeedscrapy/tree/master/WebScrapeProj

About Author

Drace Zhan

Drace Zhan has honed the bulk of his communication skills by teaching math and reading skills to high school and college graduates since 2007. A whiz at translating abstruse concepts to easily understood terms, he entered the field...
View all posts by Drace Zhan >

Related Articles

Leave a Comment

Molly June 27, 2017
This is really useful, thanks.

View Posts by Categories


Our Recent Popular Posts


View Posts by Tags

#python #trainwithnycdsa 2019 2020 Revenue 3-points agriculture air quality airbnb airline alcohol Alex Baransky algorithm alumni Alumni Interview Alumni Reviews Alumni Spotlight alumni story Alumnus ames dataset ames housing dataset apartment rent API Application artist aws bank loans beautiful soup Best Bootcamp Best Data Science 2019 Best Data Science Bootcamp Best Data Science Bootcamp 2020 Best Ranked Big Data Book Launch Book-Signing bootcamp Bootcamp Alumni Bootcamp Prep boston safety Bundles cake recipe California Cancer Research capstone car price Career Career Day citibike classic cars classpass clustering Coding Course Demo Course Report covid 19 credit credit card crime frequency crops D3.js data data analysis Data Analyst data analytics data for tripadvisor reviews data science Data Science Academy Data Science Bootcamp Data science jobs Data Science Reviews Data Scientist Data Scientist Jobs data visualization database Deep Learning Demo Day Discount disney dplyr drug data e-commerce economy employee employee burnout employer networking environment feature engineering Finance Financial Data Science fitness studio Flask flight delay gbm Get Hired ggplot2 googleVis H20 Hadoop hallmark holiday movie happiness healthcare frauds higgs boson Hiring hiring partner events Hiring Partners hotels housing housing data housing predictions housing price hy-vee Income Industry Experts Injuries Instructor Blog Instructor Interview insurance italki Job Job Placement Jobs Jon Krohn JP Morgan Chase Kaggle Kickstarter las vegas airport lasso regression Lead Data Scienctist Lead Data Scientist leaflet league linear regression Logistic Regression machine learning Maps market matplotlib Medical Research Meet the team meetup methal health miami beach movie music Napoli NBA netflix Networking neural network Neural networks New Courses NHL nlp NYC NYC Data Science nyc data science academy NYC Open Data nyc property NYCDSA NYCDSA Alumni Online Online Bootcamp Online Training Open Data painter pandas Part-time performance phoenix pollutants Portfolio Development precision measurement prediction Prework Programming public safety PwC python Python Data Analysis python machine learning python scrapy python web scraping python webscraping Python Workshop R R Data Analysis R language R Programming R Shiny r studio R Visualization R Workshop R-bloggers random forest Ranking recommendation recommendation system regression Remote remote data science bootcamp Scrapy scrapy visualization seaborn seafood type Selenium sentiment analysis sentiment classification Shiny Shiny Dashboard Spark Special Special Summer Sports statistics streaming Student Interview Student Showcase SVM Switchup Tableau teachers team team performance TensorFlow Testimonial tf-idf Top Data Science Bootcamp Top manufacturing companies Transfers tweets twitter videos visualization wallstreet wallstreetbets web scraping Weekend Course What to expect whiskey whiskeyadvocate wildfire word cloud word2vec XGBoost yelp youtube trending ZORI