WIPIVERSE

Concordancer

A concordancer is a computer program that automatically constructs a concordance—an alphabetised index of every occurrence of a word or phrase in a body of text, each entry displayed with its surrounding context. Concordancers are primary tools in corpus linguistics, lexicography, computer-assisted translation, and language teaching. The most common display format is the key word in context (KWIC) layout, in which each hit appears centred on a line with a fixed span of words to its left and right, enabling rapid scanning of usage patterns across many occurrences.

History

Pre-computational concordances

The compilation of concordances predates computers by many centuries. Around 1230, the French Dominican cardinal Hugh of Saint-Cher directed a team of friars in assembling a concordance of the Latin Vulgate Bible, generally regarded as the first systematic concordance of any text. Later milestones include a Hebrew Old Testament concordance compiled by Rabbi Mordecai Nathan (1448), Alexander Cruden's Complete Concordance to the Holy Scriptures (1737), and the manuscript Asaf ha-Mazkir, an unfinished concordance to the Babylonian Talmud compiled by Moses Rigotz around the turn of the 19th century.

First computer concordance

The first concordance produced with computing assistance was the Index Thomisticus, a comprehensive lexical index of the writings of and around Thomas Aquinas, totalling approximately 10.6 million Latin words. The Italian Jesuit priest Roberto Busa conceived the project in 1946 and secured the sponsorship of IBM in 1949. Keypunch operators in Gallarate, Italy, encoded the texts onto punched cards from around 1950. The full 56-volume printed edition was completed around 1980, followed by a CD-ROM edition in 1989 and a web-accessible version in 2005.

The KWIC format

The key word in context (KWIC) display was formalised as a computational technique by Hans Peter Luhn, a researcher at IBM, in a 1960 paper in American Documentation. In KWIC output, each instance of the search term (the node word) is centred on a line with a fixed window of words to each side; sorting the resulting lines alphabetically by the immediately adjacent word reveals collocational and phraseological patterns at a glance.

COCOA

One of the first dedicated concordancing programs was COCOA (COunt and COncordance Generation on Atlas), created in 1965 by D. B. Russell at University College London and the Atlas Computer Laboratory in Harwell, Oxfordshire. Written in approximately 4,000 cards of FORTRAN, it processed text annotated with flat, non-hierarchical markup tags and could produce word counts and concordances in multiple languages.

Oxford Concordance Program

The Oxford Concordance Program (OCP) was designed and written in FORTRAN by Susan Hockey and Ian Marriott at Oxford University Computing Services between 1979 and 1980 and first released in 1981. By the mid-1980s it had been licensed to approximately 240 institutions in 23 countries. A personal computer version, Micro-OCP, was developed for the IBM PC and sold by Oxford University Press from the late 1980s.

Personal computer era

The availability of affordable personal computers in the 1980s and 1990s enabled standalone concordancing applications. MicroConcord, developed by Mike Scott and Tim Johns and published by Oxford University Press in 1993 for MS-DOS, was among the first concordancers designed specifically for classroom language teaching. WordSmith Tools, also developed by Mike Scott, was first released in 1996 and became one of the most widely used corpus analysis suites in academic linguistics research.

Web-based concordancers

From the late 1990s onwards, web-based concordancers hosted on remote servers gave researchers browser access to large preloaded corpora. The Sketch Engine, developed by Adam Kilgarriff and Pavel Rychlý, was launched commercially in July 2003 and introduced word sketches—automatically generated one-page profiles of a word's typical grammatical relations and collocations. AntConc, created by Laurence Anthony at Waseda University, Tokyo, was first released in 2002 as freeware for Windows, macOS, and Linux.

Features

Modern concordancers typically offer a range of analytical functions beyond basic KWIC display, including:

  • KWIC display with the node word centred and context words in aligned columns, sortable by adjacent words
  • Concordance plots, visualising the distribution of hits across each text in the corpus
  • Frequency and word lists, both alphabetical and ranked by frequency
  • Collocation statistics, identifying words that co-occur with the search term more often than chance, quantified by measures such as mutual information, the t-score, or log-likelihood
  • Keyword analysis, comparing word frequencies between a study corpus and a reference corpus to identify statistically distinctive items
  • N-gram analysis, finding frequently recurring word sequences of a specified length
  • Part-of-speech tagging integration, allowing searches filtered to particular grammatical categories
  • Unicode support for multilingual text

Bilingual and parallel concordancers additionally display aligned text in two or more languages side by side, enabling comparison of translation equivalents across language pairs.

Notable concordancers

  • WordSmith Tools – Created by Mike Scott and first released in 1996, a Windows corpus analysis suite with three core modules: Concord, WordList, and Keywords.
  • AntConc – A freeware, multiplatform concordancing toolkit created by Laurence Anthony, first released in 2002, widely used in linguistics courses and independent research.
  • Sketch Engine – A corpus management and query system launched in 2003, providing browser-based access to over 800 corpora in more than 100 languages, used by major publishers for lexicographic research.
  • Wmatrix – A web-based corpus processing environment developed by Paul Rayson at Lancaster University, integrating CLAWS part-of-speech tagging and the USAS semantic tagger.
  • ParaConc – A Windows concordancer for parallel (multilingual) corpora, accepting up to four aligned texts in different languages.
  • LancsBox – A free, cross-platform corpus analysis tool developed at Lancaster University, released in 2015, supporting more than 15 languages.

Applications

Corpus linguistics

Concordancers are the primary analytical instrument in corpus linguistics, providing systematic access to patterns of use across large samples of authentic text. Common research uses include studying collocations and phraseology, analysing semantic prosody, comparing language varieties, and tracking lexical and grammatical change over time.

Lexicography

John Sinclair at the University of Birmingham pioneered the systematic use of concordance data in dictionary making through the COBUILD project, funded by Collins from the early 1980s. The project produced the Collins COBUILD English Language Dictionary (1987), generally considered the first major English dictionary compiled entirely from corpus evidence. Corpus-driven methods have since become standard practice in commercial lexicography.

Computer-assisted translation

In computer-assisted translation (CAT) software, a concordancer search allows translators to query a translation memory for all previously translated instances of a word or phrase in context. Bilingual concordancers are also used to locate translation equivalents in existing translated texts. Web-based bilingual concordancers such as Linguee and Reverso Context extend this capability to large publicly accessible multilingual corpora.

Language teaching

Tim Johns at the University of Birmingham coined the term data-driven learning (DDL) around 1990 to describe a pedagogical approach in which language learners use concordancers to explore corpus evidence and discover grammatical and lexical patterns inductively. Johns and Mike Scott developed MicroConcord (1993) specifically for classroom use. DDL has been found to support learner autonomy and awareness of collocational patterns.

Browse

More topics to explore

    Browse all articles