LT World

Sections
Personal tools
Log in

Skip to content. | Skip to navigation

Supporters

provided by

dfki logo

with support by

eu star logofp7 logo

through

meta logo
clarin logo

as well as by

bmbf logo

through

take logo

You are here: Home kb Resources & Tools Language Data TELRI Multilingual Plato corpus

TELRI Multilingual Plato corpus


  • Multilingual

morphosyntactically

syntactic dependencies, POS

  • Trans-European Language Resources Infrastructure (TELRI)

  • POS-tagged Text Corpus

  • Written

The TELRI Concerted Action has released a CD-ROM containing multilingual language resources, mainly corpora, and tools for language engineering. This product represents the concrete results of the joint research aspect of the project, which brought together partners across Europe to work together on a close and practical level. The CD-ROM provides standardised resources for a large number of languages, mainly from non-EU countries, for which such resources still tend to be scarce. The making of these resources served to foster the use of existing conventions and recommendations such as TEI, EAGLES, MULTEXT in Central and Eastern European countries.

 

Each volume of the CD ROM contains an annotated multilingual corpus. The two corpora differ in composition and encoding, reflecting the different practices adopted in their creation.

The first volume contains a parallel corpus, comprising Plato's "Republic" in twenty one languages. This corpus grew out of the work within the TELRI working groups WG7 Joint Research, WG5 TELRI Service Pool, and WG4 Lingware Availability. The "Joint Research" WG was formed with the aim of encouraging collaboration between as many of the TELRI partners as possible. For this reason, it was decided to focus on building a sample parallel corpus of translations of one text. Such corpora are able to furnish researchers with considerable information about language patterning in general, for instance on translational equivalents, collocations and phraseological units. At least initially, the size of this corpus was of less import than the coverage across the languages represented in TELRI, and the active involvement of as many partners as possible.

 

Finally, a small parallel speech corpus is also included: it comprises six translations of forty passages from EUROM/SAM, and has been recorded and digitised for four of the languages, namely Estonian, Hungarian, Romanian, and Slovene.


http://nl.ijs.si/ME/CD/docs/CDdesc.html

  • Audio CD

  • Ann Lawson
  • Laurent Romary

  • Trans-European Language Resources Infrastructure (TELRI)