The idea behind this dataset is to try to predict whether a particular comment would be highly up-voted or down-voted gi