We propose a feature vector approach to characterize the variation in large data sets of biological sequences. Each candidate sequence produces a single feature vector constructed with the number and
This data accompanies a paper for the 2nd Annual Conference for Computational Literary Studies: "The Authorship of Stephen King’s Books Written Under the Pseudonym 'Richard Bachman': A Stylometric Ana