Computer Science – Computation and Language
Scientific paper
1996-03-18
Computer Science
Computation and Language
12 pages, revised version of paper presented at 1st ACL-Workshop on Very Large Corpora, Columbus, Ohio, June 1993
Scientific paper
The usefulness of a statistical approach suggested by Church et al. (1991) is evaluated for the extraction of verb-noun (V-N) collocations from German text corpora. Some problematic issues of that method arising from properties of the German language are discussed and various modifications of the method are considered that might improve extraction results for German. The precision and recall of all variant methods is evaluated for V-N collocations containing support verbs, and the consequences for further work on the extraction of collocations from German corpora are discussed. With a sufficiently large corpus (>= 6 mio. word-tokens), the average error rate of wrong extractions can be reduced to 2.2% (97.8% precision) with the most restrictive method, however with a loss in data of almost 50% compared to a less restrictive method with still 87.6% precision. Depending on the goal to be achieved, emphasis can be put on a high recall for lexicographic purposes or on high precision for automatic lexical acquisition, in each case unfortunately leading to a decrease of the corresponding other variable. Low recall can still be acceptable if very large corpora (i.e. 50 - 100 million words) are available or if corpora for special domains are used in addition to the data found in machine readable (collocation) dictionaries.
No associations
LandOfFree
Extraction of V-N-Collocations from Text Corpora: A Feasibility Study for German does not yet have a rating. At this time, there are no reviews or comments for this scientific paper.
If you have personal experience with Extraction of V-N-Collocations from Text Corpora: A Feasibility Study for German, we encourage you to share that experience with our LandOfFree.com community. Your opinion is very important and Extraction of V-N-Collocations from Text Corpora: A Feasibility Study for German will most certainly appreciate the feedback.
Profile ID: LFWR-SCP-O-613130