Extraction of V-N-Collocations from Text Corpora: A Feasibility Study for German

Computer Science – Computation and Language

Scientific paper

Rate now

  [ 0.00 ] – not rated yet Voters 0   Comments 0

Details

12 pages, revised version of paper presented at 1st ACL-Workshop on Very Large Corpora, Columbus, Ohio, June 1993

Scientific paper

The usefulness of a statistical approach suggested by Church et al. (1991) is evaluated for the extraction of verb-noun (V-N) collocations from German text corpora. Some problematic issues of that method arising from properties of the German language are discussed and various modifications of the method are considered that might improve extraction results for German. The precision and recall of all variant methods is evaluated for V-N collocations containing support verbs, and the consequences for further work on the extraction of collocations from German corpora are discussed. With a sufficiently large corpus (>= 6 mio. word-tokens), the average error rate of wrong extractions can be reduced to 2.2% (97.8% precision) with the most restrictive method, however with a loss in data of almost 50% compared to a less restrictive method with still 87.6% precision. Depending on the goal to be achieved, emphasis can be put on a high recall for lexicographic purposes or on high precision for automatic lexical acquisition, in each case unfortunately leading to a decrease of the corresponding other variable. Low recall can still be acceptable if very large corpora (i.e. 50 - 100 million words) are available or if corpora for special domains are used in addition to the data found in machine readable (collocation) dictionaries.

No associations

LandOfFree

Say what you really think

Search LandOfFree.com for scientists and scientific papers. Rate them and share your experience with other people.

Rating

Extraction of V-N-Collocations from Text Corpora: A Feasibility Study for German does not yet have a rating. At this time, there are no reviews or comments for this scientific paper.

If you have personal experience with Extraction of V-N-Collocations from Text Corpora: A Feasibility Study for German, we encourage you to share that experience with our LandOfFree.com community. Your opinion is very important and Extraction of V-N-Collocations from Text Corpora: A Feasibility Study for German will most certainly appreciate the feedback.

Rate now

     

Profile ID: LFWR-SCP-O-613130

  Search
All data on this website is collected from public sources. Our data reflects the most accurate information available at the time of publication.