Projects in the field of Spoken Language Corpora
- Extension of CGN with speech of children, non-natives, elderly and human-machine interaction (JASMIN-CGN)
- Voice technologies for support of information society (MGACR)
- Corpus of Spoken Lithuanian
- Multilevel Annotation, Tools Engineering (MATE)
- Network for Euro-Mediterranean LAnguage Resources (NEMLAR)
- HCRC MapTask Corpus
- HCRC Maingrant
- Dialogue computer database
- DCIEM Map Task Corpus
- European Cultural Heritage Online (ECHO)
- Integrated Reference Corpora for Spoken Romance Languages (C-ORAL-ROM)
- Welsh speech database
- BAS Infrastructures for Technical Speech Processing (BITS)
- Data and Annotations for Socio Linguistics (DASL)
- Word Order, Prosody, and Information Structure (WOPIS)
- MALTILEX
- SpeechDat in Eastern Europe (SPEECHDAT-EAST)
- Multilingual Personalised Information Objects (M-PIRO)
- Visualization and Analysis of a very large real world language corpus (ROPA)
- The International Corpus of English (ICE)
- Louvain International Database of Spoken English Interlanguage (LINDSEI)
- Prague Dependency Treebank of Spoken Language (PDTSL)
- Spoken language resources and speech technology






