About GICR

The General Internet-Corpus of Russian (GICR) is a megacorpus (over 20 billion words) created with a fully automated technology of collecting and tagging texts from Russian Internet and based on the latest achievements of computational linguistics.

As of summer 2026, two versions of the corpus exist.

Version 2.0 contains an increased volume of texts with high-quality deduplication, is equipped with modern morphological and metatext markup, and is based on the next-generation RNC software platform. However, search is currently only available for the VKontakte segment. For questions, please contact us at info@ruscorpora.ru.

Version 1.0 contains materials from the VKontakte social network, LiveJournal blogs, texts from the Журнальный Зал and news from 1999-2013. The search is implemented in the old interface (to access version 1.0, please email us at geekrya@gmail.com).

The project has the status of an educational and scientific one, and  students of the Department of Computational Linguistics of RSUH and of MIPT participate in its realization, as well as MSU and the University of Leeds (UK).

The project is open to external researchers (at the moment, with some limitations related to the fact that the project is in active development and testing).

The project is accompanied by scientific seminars, which are open to all who are interested to contribute to the creation of GICR or to conduct linguistic experiments on it.