How seed term visualisation and metadata analysis offer a better understanding of newspaper corpora in large-scale corpus-based discourse analysis

Written by Helen Caple

In large-scale corpus-based discourse analysis, newspaper corpora are typically constructed from databases like Factiva or Nexis using ‘seed terms’ related to the topic of interest. In addition, the articles retrieved from these databases often contain metadata information concerning different sections or genres (e.g. Analysis, Business, News). Such corpora may contain millions, if not billions of words, which makes it impossible to know the corpus and its included genres. In a recent chapter (Caple & Bednarek 2026), we studied the ways in which seed term visualisation and metadata analysis can be used to gain better knowledge of the corpus, the diversity of included materials, and the differentiation of news and non-news genres. The corpus that we analysed was available in Sketch Engine (Kilgarriff et al. 2004) and covered news representations of violence against women in three British newspapers (The Daily Telegraph, The Guardian, The Times) over twenty years. It was created by the NEWSGEN project,[i] using seed terms such as domestic violence, sexual violence, and battered women.

1. How did we use seed term visualisation?

Using the data visualisation tool Kaleidographic (Caple and Bednarek 2017; Caple et al 2019), we created two dynamic visualisations of how seed terms occur. Our first visualisation (VAW UK Corpus v1) demonstrated the dominance of two seed terms, domestic violence and sexual violence, across the entire twenty-year period and in all three newspapers (static snapshot in Figure 1). The dominance of these two seed terms made it impossible to observe the occurrence of any other seed terms.

A screenshot of the dynamic visualisation showing that the seed terms "sexual violence" and "domestic violence" are more frequent than the other seed terms in the three newspapers
Figure 1 A static snapshot showing the dominance of two seed terms

We therefore produced a second visualisation (VAW UK Corpus v3) to account only for the remaining lower frequency seed terms. In this second visualisation, we observed that battered women and crime of passion occur most frequently, and most often in The Daily Telegraph. We further made use of other functionalities in the visualization tool (blocking out segments) to focus attention. In so doing, we observed that from 2014 onwards, the seed term battered women falls off dramatically in use in all newspapers. Similarly, we observed that from 2010 onwards, crime of passion falls out of use in The Guardian, while the seed term femicide (also used mainly by The Guardian) increases in use from 2015 onwards. Engaging in this way with the chronological unfolding of the results also showed us that some of the seed terms are very infrequent in all newspapers (sexist violence, gender violence and femicide). These appear to be marked choices, as we discuss further in relation to femicide in our chapter.

The outcome: Dynamically visualising the occurrence of seed terms provided a better understanding of the corpus and led us to unexpected avenues for further investigation.

2. How did we use metadata analysis?

Another goal of our study was to find a way of reducing the amount of non-news genres in the corpus, for a focus on true news reporting as opposed to non-news genres like opinion, letters, and analysis. To do so, we reviewed corpus annotations which were themselves based on metadata from Factiva (captured by the NEWSGEN team during corpus building). Sketch Engine provided the interface for this review.

Our workflow started by excluding articles identified through the corpus annotations as Arts, Careers, Commentary, Letters, Motoring, Reviews, and Society which comprised mainly analysis, opinion, portraits, explainers, interviews, etc, or items written in the first person by the journalist/columnist/author. These were clearly non-news genres and were excluded from the corpus. We then independently viewed the remaining sections, which were less clear, and read a random sample of articles from each section. Rejection or retention of the section was based on consensus. Finally, we used the Sketch Engine interface for creating a targeted sub-corpus containing only the retained sections (i.e. excluding non-news genres).

The outcome: This workflow reduced the number of sections retained from 130 to 47, and the corpus size by 22 per cent. While this approach does not eliminate all non-news, it does make the corpus much more about news reporting and by including targeted review of corpus contents, this further increases analysts’ understanding of the corpus. It also highlights the importance of considering news genres in corpus design, construction, and comparison. This is important, given that they have distinct social purposes and linguistic characteristics, which may influence corpus results.

While the two approaches that we trialled in this study are promising, we encourage further experimentation that facilitates more refined corpus-based discourse analyses of big, heterogenous corpora from newspaper databases.

References

Caple, Helen & Monika Bednarek. 2026. Exploring corpora through seed term visualisation and metadata analysis: A case study of a newspaper corpus covering violence against women. In Laura Mercé & Sergio Maruenda-Bataller (eds) Discourse Approaches to Gender-based Violence: Deconstructing Social Inequality through Linguistic Inquiry. De Gruyter Mouton: 135-160.

Caple, Helen, Lawrence Anthony & Monika Bednarek. 2019. Kaleidographic: A data visualization tool. International Journal of Corpus Linguistics 24(2). 246–263.

Caple, Helen & Monika Bednarek. 2017. Kaleidographic (Computer Software). http://Kaleidographic.org/

Kilgarriff, Adam, Pavel Rychlý, Pavel Smrz & David Tugwell. 2004. The sketch engine. Proceedings of the 11th EURALEX International Congress: 105–116. http://www.Sketchengine.eu


[i] The NEWSGEN project is a corpus linguistic research project at the University of Valencia focusing on public discourses around social and gender inequalities in the digital press funded by the Spanish Ministry of Science and Innovation (NEWSGEN – Cod. PID2019- 110863GB-I00).