How tight is your language? A semantic typology based on Mutual Information

Publikace

Abstrakt

Languages differ in the degree of semantic flexibility of their syntactic roles. For example, Eng-lish and Indonesian are considered more flexible with regard to the semantics of subjects, whereas German and Japanese are less flexible.

In Hawkins’ classification, more flexible lan-guages are said to have a loose fit, and less flexible ones are those that have a tight fit. This classification has been based on manual inspection of example sentences.

The present paper proposes a new, quantitative approach to deriving the measures of looseness and tightness from corpora. We use corpora of online news from the Leipzig Corpora Collection in thirty typolog-ically and genealogically diverse languages and parse them syntactically with the help of the Universal Dependencies annotation software.

Next, we compute Mutual Information scores for each language using the matrices of lexical lemmas and four syntactic dependencies (intransi-tive subjects, transitive subject, objects and obliques). The new approach allows us not only to reproduce the results of previous investigations, but also to extend the typology to new lan-guages.

We also demonstrate that verb-final languages tend to have a tighter relationship be-tween lexemes and syntactic roles, which helps language users to recognize thematic roles early during comprehension.

Klíčová slova

Mutual Information scores Universal Dependencies language typology semantic and syntactic roles