Physics – Physics and Society
Scientific paper
2011-12-28
Physics
Physics and Society
Scientific paper
10.1088/1367-2630/13/12/123024
Many features from texts and languages can now be inferred from statistical analyses using concepts from complex networks and dynamical systems. In this paper we quantify how topological properties of word co-occurrence networks and intermittency (or burstiness) in word distribution depend on the style of authors. Our database contains 40 books from 8 authors who lived in the 19th and 20th centuries, for which the following network measurements were obtained: clustering coefficient, average shortest path lengths, and betweenness. We found that the two factors with stronger dependency on the authors were the skewness in the distribution of word intermittency and the average shortest paths. Other factors such as the betweeness and the Zipf's law exponent show only weak dependency on authorship. Also assessed was the contribution from each measurement to authorship recognition using three machine learning methods. The best performance was a ca. 65 % accuracy upon combining complex network and intermittency features with the nearest neighbor algorithm. From a detailed analysis of the interdependence of the various metrics it is concluded that the methods used here are complementary for providing short- and long-scale perspectives of texts, which are useful for applications such as identification of topical words and information retrieval.
Altmann Eduardo G.
Amancio Diego R.
Costa Luciano da F.
Oliveira Osvaldo N. Jr.
No associations
LandOfFree
Comparing intermittency and network measurements of words and their dependency on authorship does not yet have a rating. At this time, there are no reviews or comments for this scientific paper.
If you have personal experience with Comparing intermittency and network measurements of words and their dependency on authorship, we encourage you to share that experience with our LandOfFree.com community. Your opinion is very important and Comparing intermittency and network measurements of words and their dependency on authorship will most certainly appreciate the feedback.
Profile ID: LFWR-SCP-O-32073