janispagel.de

Source code for https://janispagel.de
git clone git://git.janispagel.de/janispagel.de.git
Log | Files | Refs | README | LICENSE

pagel2020a.md (1328B)


      1 ---
      2 layout: pub
      3 type: inproceedings
      4 title: "GerDraCor-Coref: A Coreference Corpus for Dramatic Texts in German"
      5 author:
      6 - Janis Pagel
      7 - Nils Reiter
      8 location: "Marseille, France"
      9 month: 5
     10 spage: 55
     11 epage: 64
     12 booktitle: Proceedings of the Language Resources and Evaluation Conference (LREC)
     13 iurl: "http://www.lrec-conf.org/proceedings/lrec2020/pdf/2020.lrec-1.7.pdf"
     14 year: 2020
     15 ---
     16 
     17 ## Abstract
     18 
     19 Dramatic texts are a highly structured literary text type. Their quantitative analysis so far has relied on analysing structural properties (e.g., in the form of networks). Resolving coreferences is crucial for an analysis of the content of the character speech, but developing automatic coreference resolution (CR) systems depends on the existence of annotated corpora. In this paper, we present an annotated corpus of German dramatic texts, a preliminary analysis of the corpus as well as some baseline experiments on automatic CR. The analysis shows that with respect to the reference structure, dramatic texts are very different from news texts, but more similar to other dialogical text types such as interviews. Baseline experiments show a performance of 28.8 CoNLL score achieved by the rule-based CR system CorZu. In the future, we plan to integrate the (partial) information given in the dramatis personae into the CR model.