This corpus is divided into training, validation and evaluation. All of them contains tokens extracted using ChemTok app