In:
Linguistic Issues in Language Technology, University of Colorado at Boulder, Vol. 1 ( 2008-06-01)
Kurzfassung:
We present a general approach to formally modelling corpora with multi-layered annotation in a typed logical representation language, OWL DL. By defining abstractions over the corpus data, we can generalise from a large set of individual corpus annotations, thereby inducing a lexicon model. The resulting combined corpus and lexicon model can be interpreted as a graph structure that offers flexible querying functionality beyond current XML-based query languages. Its powerful methods for characterising and checking consistency can be used for incremental model refinement. In addition, the formalisation in a graph-based structure offers the means of defining flexible lexicon views over the corpus data. These views can be tailored for linguistic inspection or to define clean interfaces with other linguistic resources. We illustrate our approach by applying it to the syntactically and semantically annotated SALSA/TIGER corpus, a collection of German newspaper text.
Materialart:
Online-Ressource
ISSN:
1945-3604
DOI:
10.33011/lilt.v1i.1191
Sprache:
Unbekannt
Verlag:
University of Colorado at Boulder
Publikationsdatum:
2008
ZDB Id:
2434947-1