This corpus includes Lithuanian texts (mostly newspapers but also fiction, non-fiction, and specialised magazines) published between 1990 and 2008. The corpus is encoded in TEI. Non-linguistic metadat
The ICE-AUS corpus is a 1 m.word corpus of transcribed spoken and written Australian English from 1992-1995. Its internal structure with 500 samples (60% speech, 40%writing) matches that of other ICE
The corpus was manually post-edited to correct the PoS tags automatically assigned by CLAWS. The corpus is available for online querying via CQPWeb (registration required) for download from the Oxford