geco_augmented.md
March 29, 2021 ยท View on GitHub
GECO Augmented
Files from the GECO dataset Cop et al. 2017, dowloaded from https://expsy.ugent.be/downloads/geco/:
SubjectInformation.xlsxL2ReadingData.xlsxMonolingualReadingData.xlsx
The last two files (renamed to end with Augmented) have the following additional fields:
WORD_NORMlowercased word, without punctuation. Numbers reprepresented asNUM.WORD_LENword length, excluding punctuation.FREQ-BLLIP-log2(word frequency) in BLLIP (Charniak et al. 2000). BLLIP vocabulary size 229,538 words.FREQ-SUBTLEX-log2(word frequency) based on SUBTLEX-US Brysbaert and New 2009. SUBTLEX-US vocabulary size 74,286 words.FREQ-WEB-log2(word frequency) based on the Kaggle web word frequency list. Kaggle vocabulary size 333,332 words.OOV-BLLIP1 if the word out of vocabulary in BLLIP, 0 otherwise.OOV-SUBTLEX1 if the word out of vocabulary in SUBTLEX-US, 0 otherwise.OOV-WEB1 if the word out of vocabulary in Kaggle, 0 otherwise.SURP-GPT2word surprisal according to the GPT2 language model Radford et al. 2019.