The documents of requirements are normalized by standard pre-processing techniques including splitting identifiers, special token elimination, stemming, and stop word removal.
The DiDi corpus has an overall size of around 600.000 Tokens gathered from 136 South Tyrolean Facebook users who participated in the DiDi project. It consists of 11.102 Facebook...