Sampling Job Listing Data from Indeed.com

Drace Zhan
Posted on Feb 21, 2017

Audience

This blog is intended for prospective job searchers to have a macro look at common job listings through indeed.com as well as examine what are common pros and cons with popular companies such as Grubhub or The Washington Post.

Overview

Indeed.com is popular among many data scientist as the go to site for finding employment and company scouting.  Unsurprisingly, they are a popular site for data science practitioners to scrape.   Prospective students and  job searchers will be able to find insight on best paying career paths as well as desirable companies according to company reviews.   The blog will also look into what are features that can predict salary of a position as well as company job titles.  Finally there will be wordclouds of pros and cons from reviews written regarding popular companies as well as a sentiment analysis of polarity scores among reviews for The Washington Post.  The data was collected via scrapy spiders and code for the spider can be accessed below as well as the compiled csv's from said data.

Best Pay

As of February 2017, the best paying position from indeed.com is Radiologist followed by Senior Vice President and Vice President of Sales.  There is variance among how much Vice President of Sales pays as demonstrated by the black variance bar.

BestPayingJob

LabCorp has the highest mean salary offers for their listings.  However, that doesn't necessarily mean they are paying the best for each position, simply that they have high paying position offers.  Open Systems Technologies, Inc has a large spread on their salary postings which indicates they likely have a wide range of position offerings.  Mitchell/Martin Inc. is interesting here as they also appear among the companies with top review scores by their employers.

BestPayingCompany

Best Sum of Review Scores

The review sum is taken from the sum of the mean score of compensation, management, work life balance, culture and job security.  There are noticeable statements from this graph.  First, there are positions where there is some variance among the positions with the highest review sum scores.  Across the board, those most satisfied with their position overall come from analysts and engineers.

BestJobReviews

Visualized below, universities and staffing agencies are enjoy the highest scores among compensation, management, work life balance, culture and job security.

BestWorkPlace

Exploratory Data Analysis

A heatmap can be used to display that review scores among work life balance, management, job security and culture all strong correlations with each other.  Compensation is the feature which displays some correlation with all five other categories.

HeatMap

A pair plot can be created to display this visualization better especially when the data is separated by salary quantiles.  Note that the gigantic spread among salaries in the highest quantile.

pairplotquantiles

Unsurprisingly, higher ratings in compensation tend to indicate higher salary.  Also higher paying jobs tend to have lower job security.  Surprisingly those with higher salary also tend to give higher scores for company culture.  This could mean that those in better paid position tend to "buy in" their company culture and there could be some hypothesis that this could be some mercenary attitude in that loyalty can be bought.

SalaryPredictionCoefficient

Also as expected, if feature importance were to be examined on what is most likely to be an accurate predictor of job title.  Salary is by far the strongest predictor among the data set with work life, job security and culture having almost no impact.

FeaturePredictionForestV2

What's being said

A wordcloud can be constructed to illustrate what is being said among well known companies within the data set.  Among the pros, these are commonly cited words among the following companies.

WordCloudPros

The cons tend to have stronger words.  Oddly enough, there are some strange polarities with grubhub with "acceptance" being a popular word among positive reviews and "fascist" being used in a negative one.  Humorously, it has also been verified by many employees that the coffee in Bloomberg is indeed, terrible.

WordCloudCons

Finally, a sentiment analysis can be created for reviews of various companies such as the one below for The Washington Post.

SubjectivePolarityNLP

Conclusion

While this is overall a summary of a multi-facet picture of the data scraped from Indeed.com, there is a still significant amount left to explore.  As a result, the collection of this project along with accompanying code can be found in the github below for a user to peruse further analysis on their own.

https://github.com/DraceZhan/Indeeedscrapy/tree/master/WebScrapeProj

About Author

Drace Zhan

Drace Zhan

Drace Zhan has honed the bulk of his communication skills by teaching math and reading skills to high school and college graduates since 2007. A whiz at translating abstruse concepts to easily understood terms, he entered the field...
View all posts by Drace Zhan >

Related Articles

Leave a Comment

Avatar
Molly June 27, 2017
This is really useful, thanks.

View Posts by Categories


Our Recent Popular Posts


View Posts by Tags

#python #trainwithnycdsa 2019 airbnb Alex Baransky alumni Alumni Interview Alumni Reviews Alumni Spotlight alumni story Alumnus API Application artist aws beautiful soup Best Bootcamp Best Data Science 2019 Best Data Science Bootcamp Best Data Science Bootcamp 2020 Best Ranked Big Data Book Launch Book-Signing bootcamp Bootcamp Alumni Bootcamp Prep Bundles California Cancer Research capstone Career Career Day citibike clustering Coding Course Demo Course Report D3.js data Data Analyst data science Data Science Academy Data Science Bootcamp Data science jobs Data Science Reviews Data Scientist Data Scientist Jobs data visualization Deep Learning Demo Day Discount dplyr employer networking feature engineering Finance Financial Data Science Flask gbm Get Hired ggplot2 googleVis Hadoop higgs boson Hiring hiring partner events Hiring Partners Industry Experts Instructor Blog Instructor Interview Job Job Placement Jobs Jon Krohn JP Morgan Chase Kaggle Kickstarter lasso regression Lead Data Scienctist Lead Data Scientist leaflet linear regression Logistic Regression machine learning Maps matplotlib Medical Research Meet the team meetup Networking neural network Neural networks New Courses nlp NYC NYC Data Science nyc data science academy NYC Open Data NYCDSA NYCDSA Alumni Online Online Bootcamp Open Data painter pandas Part-time Portfolio Development prediction Prework Programming PwC python python machine learning python scrapy python web scraping python webscraping Python Workshop R R language R Programming R Shiny r studio R Visualization R Workshop R-bloggers random forest Ranking recommendation recommendation system regression Remote remote data science bootcamp Scrapy scrapy visualization seaborn Selenium sentiment analysis Shiny Shiny Dashboard Spark Special Special Summer Sports statistics streaming Student Interview Student Showcase SVM Switchup Tableau team TensorFlow Testimonial tf-idf Top Data Science Bootcamp twitter visualization web scraping Weekend Course What to expect word cloud word2vec XGBoost yelp