2007 InformationGenealogy
Jump to navigation
Jump to search
- (Shaparenko & Joachims, 2007) ⇒ Benyah Shaparenko, Thorsten Joachims. (2007). “Information Geneology: Uncovering the flow of ideas in non-hyperlinked document databases.” In: Proceedings of the 13th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD-2007). doi:10.1145/1281192.1281259
Subject Headings: citation inference, flow of ideas, information genealogy, language models, temporal data, text mining
Notes
Cited By
~ 10 http://scholar.google.com/scholar?cites=10658038528366543569
Quotes
Abstract
- We now have incrementally-grown databases of text documents ranging back for over a decade in areas ranging from personal email, to news-articles and conference proceedings. While accessing individual documents is easy, methods for overviewing and understanding these collections as a whole are lacking in number and in scope. In this paper, we address one such global analysis task, namely the problem of automatically uncovering how ideas spread through the collection over time. We refer to this problem as Information Genealogy. In contrast to bibliometric methods that are limited to collections with explicit citation structure, we investigate content-based methods requiring only the text and timestamps of the documents. In particular, we propose a language-modeling approach and a likelihood ratio test to detect influence between documents in a statistically well-founded way. Furthermore, we show how this method can be used to infer citation graphs and to identify the most influential documents in the collection. Experiments on the NIPS conference proceedings and the Physics ArXiv show that our method is more effective than methods based on document similarity.
References
,
Author | volume | Date Value | title | type | journal | titleUrl | doi | note | year | |
---|---|---|---|---|---|---|---|---|---|---|
2007 InformationGenealogy | Thorsten Joachims Benyah Shaparenko | Information Geneology: Uncovering the flow of ideas in non-hyperlinked document databases | KDD-2007 Proceedings | http://www.cs.cornell.edu/People/tj/publications/shaparenko joachims 07a.pdf | 10.1145/1281192.1281259 | 2007 |