Reliable Auditing of Linguistic Variation in Large Language Models
Loading...
Files
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
University of Waterloo
Abstract
A language model may respond differently to the same underlying request depending on how it is phrased. Dialect, register, and grammatical form vary across users, even when the intended meaning remains unchanged. As a result, models may behave inconsistently across the populations they encounter when they respond to these linguistic features. The composition of an evaluation set is therefore an important part of a fairness audit. A test set written entirely in standardized English cannot reveal behavior that other forms of language may elicit. Paraphrasing is a common way to introduce linguistic variation, yet an unlabeled rewrite only shows that the model is sensitive to a change in wording without identifying which linguistic change produced the effect. This thesis examines the value of specifying the linguistic change prior to paraphrasing, the cost of doing so, and whether the same linguistic contrast can be identified within a model’s internal representations.
The first study introduces AUGMENT, a framework that restricts each paraphrase to a single, explicitly defined linguistic transformation, and applies it to nine models across two benchmarks. Aggregate differences in accuracy remain below two percentage points. When the same results are separated by transformation type, the effects are several times larger and often move in opposite directions across dataset subsets. Unconstrained paraphrasing does not recover the full size of these effects and sometimes fails to recover their direction. These results show that the main value of controlled paraphrasing lies in attribution, since it allows changes in model behavior to be associated with specific linguistic properties. The second study examines the cost of achieving this level of control. Using a reference set of human-annotated paraphrases, it finds that the validity of unfiltered generation varies substantially depending on the requested transformation. Annotator agreement, threshold-based filters, and LLM judges also vary widely in reliability. The amount of validation required is driven primarily by the structural and social complexity of the transformation rather than by evaluation-set size. Dialectal and syntactic variation are among the most demanding cases, even though they are especially important for fairness auditing.
The third study examines the internal representations of an open-weight model. A dialect direction estimated from matched Standard American English and African American English pairs remains stable across all layers of Llama-3.1-8B-Instruct and transfers to naturally occurring AAE. When used as a steering intervention, it also reproduces the direction of dialect-conditioned behavioral effects on three of four tasks, although the magnitude is not always preserved. These findings suggest that representational evidence can help identify where a linguistic contrast is encoded and whether that contrast is connected to model behavior. Constructing such evidence still requires grounded and carefully validated pairs. Taken together, the studies frame linguistic fairness auditing as a problem of resource allocation. Auditors must define the relevant forms of variation, establish appropriate validation requirements, and decide how broadly the audit should reach while limiting error. Making these choices explicit clarifies the audit’s scope and the extent to which its conclusions can reasonably be generalized.