Dear colleagues and corpus users,
The beta version of GICR 2.0 is now available on the RNC platform. It includes texts from the VKontakte social network covering the period from 2007 to early 2022, with a total volume of 11.3 billion words.
The morphological annotation of GICR 2.0 was produced using an integrated morphosyntactic parser developed by Daniil Anastasyev, with improved lemmatization. More information about these improvements is available here. The resulting annotation, originally in the Universal Dependencies format, was then converted into the RNC morphological standard.
The metatextual annotation includes the year and month when the text was written, the text type (post or comment), and the author’s year of birth, gender, and region, as extracted from the author’s social media profile.
The RNC interface allows users to search by lemma, word form, and grammatical features, including with regular expressions; define subcorpora based on metatextual features; view results in Concordance and KWIC formats with sorting by a selected parameter; build charts based on publication time; and obtain statistics by sociolinguistic parameters.
We will be happy to answer your questions about GICR 2.0: please write to us at info@ruscorpora.ru. If you find an error in the RNC system, please report it using the form.
Поделитесь с коллегами!