How seed term visualisation and metadata analysis offer a better understanding of newspaper corpora in large-scale corpus-based discourse analysis

Written by Helen Caple In large-scale corpus-based discourse analysis, newspaper corpora are typically constructed from databases like Factiva or Nexis using ‘seed terms’ related to the topic of interest. In addition, the articles retrieved from these databases often contain metadata information concerning different sections or genres (e.g. Analysis, Business, News). Such corpora may contain millions,…

Human rights in British parliamentary debates: a computational approach

Written by Marco Duranti Introduction: The symbolic force of human rights The last half century has witnessed the ascendancy of human rights as a language of legal, moral, and political claim-making across much of the globe. The idiom of human rights has gained particular symbolic force (Neves 2007) and is regarded by some as ‘the…

Representations of obesity in the news

Written by Monika Bednarek and Gavin Brookes Note: This post was simultaneously published by the Sydney Corpus Lab and by the ESRC Centre for Corpus Approaches to Social Science at Lancaster University. It is published under a Creative Commons — Attribution Noncommercial license. If you want to republish it, please follow the relevant licensing guidelines….

Large language models (LLMs) in corpus linguistics – Using GenAI with corpora

written by Monika Bednarek The Sydney Corpus Lab recently published a post containing a synthesis of how large language models (LLMs) and generative artificial intelligence (GenAI) tools have been incorporated into corpus linguistic research. This blog post is intended as a companion to that much longer post. It presents the main take-aways that researchers may…

2025: The year in review for the Sydney Corpus Lab

written by Monika Bednarek 2025 was a slightly less active year for the Sydney Corpus Lab, as I was on long service leave during semester 1. Nevertheless, we continued working on various projects before and after my leave, including our national collaboration on the Language Data Commons of Australia (LDaCA), which is a project led…

Generative AI in corpus linguistics: A synthesis

Written by Kelvin Lee The advent of large language models (LLMs) and generative artificial intelligence (GenAI) tools such as ChatGPT has led to AI being used in many facets of everyday life as well as in education and research. In this blog post, I will explore how AI has been incorporated into (primarily English) corpus…