Search

3 results

4 - The Data
Benedikt Szmrecsanyi, KU Leuven, Belgium, Jason Grafmiller, University of Birmingham
Book:

Comparative Variation Analysis

Published online:

24 August 2023

Print publication:

07 September 2023, pp 56-81
- Chapter
- - Get access
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation
Summary

This chapter begins with a general discussion of potential data types in variationist linguistics. Next, we present the two main data sources we use in the study: the International Corpus of English (ICE) and the Global Corpus of Web-Based English (GloWbE). The former comprises a set of parallel, balanced corpora representative of language usage across a wide range of standard national varieties. Each ICE corpus contains 500 texts of 2000 words each, sampled from twelve spoken and written genres/registers, totaling approx. 1 million words. GloWbE contains data collected from 1.8 million English language websites – both blogs and general web pages – from twenty different countries (approx. 1.8 billion words in all). Discussion of the corpora is followed by a detailed description of the data collection, identification, and annotation procedures for our three alternations. Here we carefully define the variable context for each alternation, and outline the methods for coding various linguistic constraints that are included in our analyses.

9 - Comparing Methods for the Evaluation of Cluster Structures in Multidimensional Analyses
from Part III - Perspectives on Multifactorial Methods
- By Ole Schützler
Edited by Ole Schützler, Universität Leipzig, Julia Schlüter
Book:

Data and Methods in Corpus Linguistics

Published online:

06 May 2022

Print publication:

26 May 2022, pp 259-288
- Chapter
- - Get access
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation
Summary

This chapter sets out by discussing the way in which multidimensional techniques and visualizations have been used to analyse linguistic data. While, for instance, multidimensional scaling and unrooted phenograms (or NeighborNets) have primarily been designed for exploratory purposes, the author argues that they are in fact regularly used to put linguistic assumptions or hypotheses to the test. Cluster goodness (in terms of internal coherence and external distance from other clusters) in such approaches are typically evaluated based on a two-dimensional visualization. The author compares the affordances and limitations of visual inspection with a quantitative set of metrics that directly relates to visual displays but adds a degree of precision not attained by the human eye. The empirical part of the paper applies both approaches to a study of concessive constructions in six varieties of English, based on spoken and written material from the International Corpus of English. The author suggests that the new metrics can be usefully applied to a variety of multidimensional techniques to endow them with a measure of objectivity.

Chapter 1 - Introduction
- By Tobias Bernaisch
Edited by Tobias Bernaisch, Justus-Liebig-Universität Giessen, Germany
Book:

Gender in World Englishes

Published online:

11 December 2020

Print publication:

07 January 2021, pp 1-22
- Chapter
- - Get access
    
    Check if you have access via personal or institutional login
    
    Log in Register
- Export citation
Summary

Setting the agenda for the volume, this introduction amalgamates the so far relatively isolated strands of research into genderlectal variation and World Englishes, relying on state-of-the-art empirical approaches. As they apply to the vast majority of speakers of English around the world, the notions of English as a second language and English as a foreign language are introduced and – in this light – recent attempts at bridging this paradigm gap between these two speaker groups as well as the models employed in these attempts are briefly discussed. For the study of gender and language, the central pillars of its most prominent theoretical waves – the dominance, the difference and the social construct framework – are presented and the corresponding methodological approaches critically appraised. Against this background, it is concluded that responsible explorations of genderlectal variation in World Englishes need to be based on transparent empirical foundations – both in terms of datasets and statistical modelling. For this reason, the tenets of corpus linguistics are explored and the benefits of multifactorial statistical techniques as consistently applied in this volume are illustrated. After previews of the individual chapters in the volume, the introduction ends with summarising remarks including the moderator function of gender in World Englishes.

Search Results

Refine search

Refine search

Actions for selected content:

3 results

4 - The Data

Summary

9 - Comparing Methods for the Evaluation of Cluster Structures in Multidimensional Analyses

Summary

Chapter 1 - Introduction

Summary

Search Results

Refine search

Refine search

Actions for selected content:

Save Search

3 results

4 - The Data

Summary

9 - Comparing Methods for the Evaluation of Cluster Structures in Multidimensional Analyses

Summary

Chapter 1 - Introduction

Summary