WIPIVERSE

Co-citation Proximity Analysis

Co-citation Proximity Analysis (CPA) is a document similarity measure that uses citation analysis to assess semantic similarity between documents at both the global document level and at individual section-level. It builds on traditional co-citation analysis but differs by exploiting the information implied in the placement of citations within the full texts of documents.

Overview

CPA was conceived by Bela Gipp in 2006 and first published by Gipp and Beel in 2009. The measure rests on the assumption that within a document's full text, documents cited close to each other tend to be more strongly related than those cited farther apart. For example, if two references appear within the same sentence, they are considered more closely related than two references separated by several paragraphs.

The key advantage of CPA over other citation analysis approaches (such as Bibliographic Coupling, traditional Co-Citation, or the Amsler measure) is improved precision. These earlier approaches do not account for the location or proximity of citations within documents. CPA enables a more granular automatic classification of documents and can identify not only related documents but also the specific sections within texts that are most related.

Method of Calculation

CPA calculates a Citation Proximity Index (CPI) for each set of documents cited by an examined document. Cited documents are assigned a weight of $ \frac{1}{2^n} $, where $ n $ represents the number of levels between citations. Beginning at the lowest level, levels may be defined as:

  • Citation groups
  • Sentences
  • Paragraphs
  • Chapters
  • The entire document or journal

Algorithm Variants

  • Basic-CPA – The fundamental concept as described above.
  • Extended-CPA – Considers the tree structure and order of citations within citation groups.
  • Multidimensional-CPA – Incorporates additional information such as impact factor.
  • Hybrid-CPA – Combines the CPI with other similarity measures (e.g., text-based measures) to boost performance, especially for documents with insufficient citation information.

Performance

CPA has been found to outperform traditional co-citation analysis, particularly when documents contain extensive bibliographies or have not been frequently cited together (i.e., have a low co-citation score). Liu and Chen (2011) found that sentence-level co-citations are potentially more efficient markers for co-citation analysis compared to loosely coupled article-level co-citations, as they preserve the essential structure of the traditional co-citation network while forming a much smaller subset of all co-citation instances.

An analysis by Schwarzer et al. (2016) showed that citation-based measures (CPA and co-citation analysis) have complementary strengths compared to text-based similarity measures. Text-based approaches reliably identified more narrowly similar articles (e.g., articles sharing identical terms), while CPA outperformed co-citation analysis at identifying more broadly related articles and more popular articles, which the authors suggest are likely also of higher quality.

Applications

CPA was developed with two primary applications in mind:

  1. Recommender systems – An improved, more fine-grained measure of document similarity can significantly improve the relevance of academic literature recommendations.
  2. Clustering – A more granular similarity measure enables more precise clustering algorithms for academic literature.

See Also

  • CITREC – An evaluation framework for citation-based similarity measures such as Bibliographic coupling, Co-citation, Co-citation Proximity Analysis, and others.
  • Bibliographic coupling
  • Co-citation

References

  1. Bela Gipp and Joeran Beel (2009). "Citation Proximity Analysis (CPA) – A new approach for identifying related work based on Co-Citation Analysis." Proceedings of the 12th International Conference on Scientometrics and Informetrics (ISSI'09), vol. 2, pp. 571–575, Rio de Janeiro, Brazil.
  2. Bela Gipp and Joeran Beel. "Method and system for detecting a similarity of documents." Patent Application, Oct 27, 2011. 2011/0264672 A1.
  3. Bela Gipp (2006). "Doctoral Proposal: (Co-)Citation Proximity Analysis – A Measure to Identify Related Work."
  4. M. Schwarzer, M. Schubotz, N. Meuschke, C. Breitinger, V. Markl, and B. Gipp (2016). "Evaluating Link-based Recommendations for Wikipedia." Proceedings of the 16th ACM/IEEE-CS Joint Conference on Digital Libraries (JCDL), pp. 191–200.
  5. Shengbo Liu and Chaomei Chen (2011). "The Effects of Co-citation Proximity on Co-citation Analysis." The 13th Conference of the International Society for Scientometrics and Informetrics (ISSI), Durban, South Africa.
  6. Bela Gipp, Norman Meuschke & Mario Lipinski (2015). "CITREC: An Evaluation Framework for Citation-Based Similarity Measures based on TREC Genomics and PubMed Central." Proceedings of the iConference 2015, Newport Beach, California.
Browse

More topics to explore

    Browse all articles