UWSpace
UWSpace is the University of Waterloo’s institutional repository for the free, secure, and long-term home of research produced by faculty, students, and staff.
Depositing Theses/Dissertations or Research to UWSpace
Are you a Graduate Student depositing your thesis to UWSpace? See our Thesis Deposit Help and UWSpace Thesis FAQ pages to learn more.
Are you a Faculty or Staff member depositing research to UWSpace? See our Waterloo Research Deposit Help and Self-Archiving pages to learn more.

Communities in UWSpace
Select a community to browse its collections.
- The University of Waterloo institution-wide UWSpace community.
Recent Submissions
Item type: Item , Reliable Auditing of Linguistic Variation in Large Language Models(University of Waterloo, 2026-09-30) Ma, RebeccaA language model may respond differently to the same underlying request depending on how it is phrased. Dialect, register, and grammatical form vary across users, even when the intended meaning remains unchanged. As a result, models may behave inconsistently across the populations they encounter when they respond to these linguistic features. The composition of an evaluation set is therefore an important part of a fairness audit. A test set written entirely in standardized English cannot reveal behavior that other forms of language may elicit. Paraphrasing is a common way to introduce linguistic variation, yet an unlabeled rewrite only shows that the model is sensitive to a change in wording without identifying which linguistic change produced the effect. This thesis examines the value of specifying the linguistic change prior to paraphrasing, the cost of doing so, and whether the same linguistic contrast can be identified within a model’s internal representations. The first study introduces AUGMENT, a framework that restricts each paraphrase to a single, explicitly defined linguistic transformation, and applies it to nine models across two benchmarks. Aggregate differences in accuracy remain below two percentage points. When the same results are separated by transformation type, the effects are several times larger and often move in opposite directions across dataset subsets. Unconstrained paraphrasing does not recover the full size of these effects and sometimes fails to recover their direction. These results show that the main value of controlled paraphrasing lies in attribution, since it allows changes in model behavior to be associated with specific linguistic properties. The second study examines the cost of achieving this level of control. Using a reference set of human-annotated paraphrases, it finds that the validity of unfiltered generation varies substantially depending on the requested transformation. Annotator agreement, threshold-based filters, and LLM judges also vary widely in reliability. The amount of validation required is driven primarily by the structural and social complexity of the transformation rather than by evaluation-set size. Dialectal and syntactic variation are among the most demanding cases, even though they are especially important for fairness auditing. The third study examines the internal representations of an open-weight model. A dialect direction estimated from matched Standard American English and African American English pairs remains stable across all layers of Llama-3.1-8B-Instruct and transfers to naturally occurring AAE. When used as a steering intervention, it also reproduces the direction of dialect-conditioned behavioral effects on three of four tasks, although the magnitude is not always preserved. These findings suggest that representational evidence can help identify where a linguistic contrast is encoded and whether that contrast is connected to model behavior. Constructing such evidence still requires grounded and carefully validated pairs. Taken together, the studies frame linguistic fairness auditing as a problem of resource allocation. Auditors must define the relevant forms of variation, establish appropriate validation requirements, and decide how broadly the audit should reach while limiting error. Making these choices explicit clarifies the audit’s scope and the extent to which its conclusions can reasonably be generalized.Item type: Item , Explicit Relation Structures for Visual Content Generation and Quality-Aware Representation Learning(University of Waterloo, 2026-09-30) Naseri, MahdiModern visual learning systems often depend on relations that are not fully captured by a single label, prompt, image, or scalar quality score. A sparse scene description leaves many plausible objects and relations unstated. A graph self-supervised learner needs guidance about which nodes should share representations beyond raw augmentation pairs. An image-quality representation must account for the structured interaction between content, distortion type, and distortion severity. This thesis studies explicit relation structures as an intermediate language for making such dependencies available to learning systems. The central argument is that relation structures are useful not only as mathematical objects, but also as interfaces for inspecting, training, and evaluating visual models when the relevant dependency is broader than a single paired sample. The thesis develops this argument through three technical studies. First, Generated Contents Enrichment (GCE) formulates visual content enrichment as an explicit scene-graph problem. A sparse scene graph is enriched by predicting additional object and relation semantics before image generation, so the system exposes what content has been added before the image is rendered. This separation lets the work ask whether graph-level enrichment provides inspectable semantic content that supports the rendered image. The chapter therefore evaluates graph recovery, image-side behavior, qualitative examples, and human judgments as complementary evidence about the enrichment process. Second, Explicitly-Generated Relation Graph for Self-Supervised Representation Learning (ExGRG) studies explicit relation construction for node-level graph Self-Supervised Learning (SSL). Standard augmentation-based graph SSL uses paired views of the same node, but this pairwise identity signal does not fully describe which nodes should share representation structure. ExGRG constructs a compositional relation graph from multiple cues, including augmentation identity, graph structure, positional and structural encodings, representation-space neighborhoods, and online clustering. The resulting relation graph weights the self-supervised training signal and provides a concrete object for ablation and training analysis. This chapter shows how explicitly generated relation structure can improve graph representation learning under the tested node-classification protocol while allowing the relation cues behind the invariance signal to be inspected. Third, SHAped Modeling of Implicit Structural Associations for Self-Supervised No-Reference Image Quality Assessment (SHAMISA) applies the same relation-structure perspective to self-supervised No-Reference Image Quality Assessment (NR-IQA). In image quality assessment, perceptual similarity depends on content, distortion family, distortion severity, and the interaction among multiple distortions. SHAMISA constructs metadata-driven and feature-driven relation graphs over distorted images and uses these graphs to shape a quality-aware self-supervised representation before any human quality labels are used for downstream evaluation. The method is evaluated through within-dataset and cross-dataset quality prediction, ablation studies, and representation diagnostics, emphasizing that quality-aware representation learning requires more than generic image-level invariance. Under the reported protocols, GCE raises held-out object accuracy from 15.19% to 20.60% and available-edge recall from 36.64% to 75.18% relative to SceneGraphGen+, while triplet F1 increases from 3.73% to 7.73%. ExGRG has the highest reported node-classification point estimate on all nine tested graph datasets, including 97.87% on Cora and 89.68% on CiteSeer. SHAMISA obtains the strongest six-dataset average among the SSL methods in the within-dataset Image Quality Assessment (IQA) comparison, with average SRCC/PLCC of 0.886/0.904. Across the three studies, the thesis uses relation structures in different roles. They serve as predicted semantic content for generation, as supervisory graphs for self-supervised representation learning, and as explicit structures whose effects can be examined through ablations, visualizations, and stress tests. The common design principle is to make the relevant dependency explicit enough that it can be inspected, modified, and tested. At the same time, the thesis keeps the limits of explicit structure in view. Scene-graph annotations are incomplete, generated relation graphs depend on the quality of their sources, and auxiliary visualizations do not replace quantitative evaluation. The contribution is a set of methods and evaluations showing that carefully constructed relation structures can provide useful control, supervision, and analysis when visual tasks depend on structured dependencies that are otherwise hidden.Item type: Item , Laser Direct Writing of Copper-Graphene Heterostructures for Flexible Devices(University of Waterloo, 2026-09-29) Rathod, ShasvatCopper conductors, essential building blocks for modern electronics, degrade in performance as devices become thinner and flexible, due to oxidation, grain boundary scattering, poor substrate adhesion, and mechanical fatigue. The hybridization of graphene with metals and metal oxides such as copper-graphene (Cu-Gr) holds promise for enhancing flexible devices by overcoming these challenges. However, current fabrication techniques involve complex, multi-step mixing processes and lack a clear understanding of how graphene integrates with copper. Additionally, the effects of these fabrication methods on copper nucleation, the graphene-copper interface, and microstructure remain unclear, as do their impacts on electrical, thermal, environmental, and mechanical properties. This thesis investigates these topics through three detailed studies using laser direct writing (LDW). First, to simplify the fabrication process, a minimalist LDW technique was developed to create graphene-metal heterostructures and flexible devices by layered fabrication of laser-reduced graphene oxide (LrGO) followed by reduction of CuOx, ZnOx, and FeOx nanomaterials from metal-ion precursors. Supplied laser energy during fabrication, controlled through laser processing parameters, tuned the oxygen functional groups on the LrGO surface and determined the metal oxide composition, which enabled the process to program sensor and junction responses. The sensors showed ranged tunability in: normalized current gains from −2.7 to 3.5, response times of 0.02 and 15 s, and recovery times of 0.04 and 6 s. Additionally, LDW produced LrGO/CuOₓ PN junctions and bipolar transistors with rectification ratios up to 160 and common-emitter current gains of 35.5–38.2. Second, to overcome the weak bonding and voids characteristic of planar interfaces in layer by layer assembled composites, LDW was adapted for simultaneous fabrication of graphene and copper. By tailoring plasma plume physics through a confinement mechanism, the local energy input was controlled to toggle between keyhole and conduction irradiation modes, which respectively governed graphene formation and copper reduction. Modified LDW to initiate keyhole and conduction simultaneously produced distinct Cu-Gr nanocomposite structures, including copper-coated graphene and copper nanoparticles embedded within graphene. The resulting interconnects reached a resistivity of 9.37 × 10⁻⁸ Ω·m and breakdown current density of 1.61 × 10⁸ A·cm⁻², approaching annealed copper. Third, graphene flakes were dispersed in the copper precursor prior to laser irradiation, facilitating the in situ growth of copper on graphene. This growth pathway promotes intimate interfacial contact and uniform distribution while reducing the gaps, contamination, and agglomeration commonly associated with layered assembly of composites, or post-synthesis mixing. Graphene flakes, which substantially changed copper growth mechanisms under laser irradiation, acted as preferential nucleation sites that lowered the minimum laser energy for copper nucleation from approximately 1.5 to 0.4 J mm⁻³, producing a higher density of copper nanoparticles that sintered into a continuous network encapsulating the graphene. The resulting dense composite reached approximately 98% relative density, thermal conductivity up to 1095 W m⁻¹ K⁻¹, and sheet resistance as low as 0.15 Ω sq⁻¹. In summary, this thesis establishes LDW as a process with precise control over graphene and copper formation, progressively increasing graphene integration in Cu-Gr nanocomposites for flexible conductors, sensors, and thermally conductive films.Item type: Item , “We Still Have the Land, Right?”: Catholic Agricultural Sites as Hubs for Social and Ecological Teaching in Canada(University of Waterloo, 2026-09-29) Szoller, BenThis ethnographic research examines the influence of Roman Catholic ecological agro-activism, that is, ecological activism that involves growing food, in Canadian society. For over a century, Catholic social teaching has informed the Roman Catholic Church’s response to various social issues and injustices, especially those brought about by the industrial revolution. More recently, this body of work has incorporated concern for ecological issues such as climate change, over-consumption, and industrial agriculture, and was popularized through the work of Pope Francis and his landmark encyclical, Laudato ‘Si, or Care for Our Common Home. At the same time, local and regional Catholic communities have developed their own unique blend of Catholic social and ecological programming to meet the needs of local parishes, clergy, institutions, and lay people. These top-down and bottom-up activities can appear at times seamlessly connected, disjointed, or even in conflict. As a result, questions remain about the degree to which these principles actually influence the attitudes and behaviours of Catholics. To assess the impact of Catholic ecological agro-activism in public discourse in Canada, this project examines three levels of Catholic activity: regional programming, individual beliefs and practices, and official Church doctrine. I conducted ethnographic research at three Catholic sites to examine how they articulate the Church’s social and ecological teaching through various agricultural programs and advocacy efforts. Methods included participant observation and extensive interviews with regional Catholic leaders, as well as interviews with non-Catholic participants and leaders from other Catholic organizations. I develop my analysis according to four primary themes—places, people, participation, and politics. Key findings include a strong correlation between religious vocations and the formation of ecological “narratives,” the proliferation of partnerships with non-Catholic stakeholders, and the surprising role the COVID-19 pandemic played in bringing greater visibility for the congregations and their skills-training programs. In the end, I argue that these sites act as “hubs” that help develop and deploy a unique blend of Catholic social and ecological teaching to the public through various agricultural projects (organic growing workshops, community gardens, etc.) and advocacy efforts. These hubs promote pro-environmental narratives and influence public discourse through organizational partnerships, various political activities, and the day-to-day interactions of Catholic parishioners and leaders. Such findings help to highlight both the challenges faced by Catholic congregations today and the creativity with which they navigate the Church’s complex structure. Moreover, describing the rural congregations at the heart of these activities helps not only to better understand the unique experiences of rural Catholics, but also the religious landscape in Canada more broadly.Item type: Item , How Well Do Large Language Models Detect Bugs in Code Changes?(University of Waterloo, 2026-09-29) Yakubu, AyindeThis thesis evaluates how well general-purpose open-weight large language models detect bugs in software code changes. The evaluation uses historical development data from the Apache Kafka project obtained through the ApacheJIT dataset. From approximately 12,000 commit records, the dataset was filtered to obtain 524 one-to-one bug-inducing commit (BIC) and bug-fixing commit (BFC) relationships and 530 non-bug-inducing commits. An automated framework was developed to retrieve commit patches, submit code changes for LLM-based review, and record predictions and review comments. Three open-weight LLMs—gpt-oss-120b, gemma-4-31B-it, and Qwen3.6-35B-A3B—were evaluated under a common zero-shot prompting strategy across three repeated experimental runs. Performance was measured using precision, accuracy, recall, F1-score, balanced accuracy, Matthews correlation coefficient, and processing coverage. In addition, an LLM-as-a-Judge procedure assessed whether generated defect reports were semantically consistent with evidence from corresponding bug-fixing commits and Apache Kafka JIRA issue records. The results show that the evaluated LLMs have limited reliability as autonomous defect detectors. Although the models identified subsets of historically labelled bug-inducing changes, substantial numbers of false positives and false negatives were observed. The first gpt-oss-120b run achieved the highest reported recall of 0.5163, while the highest individual-run accuracy was 0.4872. However, comparison with a trivial always-NOBUG baseline showed that model accuracies did not exceed the corresponding baseline accuracies on successfully processed records. Across the reported runs, balanced accuracy remained below 0.5 and Matthews correlation coefficient (MCC) remained negative, indicating weak overall discrimination between BIC and non-BIC benchmark examples. The results further demonstrate that conventional classification metrics alone provide an incomplete characterisation of LLM bug-detection reliability, because useful review requires semantic correctness, actionable explanations, and sufficient project context. Classification metrics alone do not establish explanation quality. The assigned judges rated 16–23% of first-run true-positive explanations as matching the historically documented defect. The findings suggest that assistant-style use is a more appropriate direction for further evaluation than autonomous defect detection. The thesis contributes a real-world evaluation framework, a comparative empirical assessment of three open-weight LLMs, and an evidence-based methodology for assessing generated explanations against historical defect evidence. The results also highlight repository context, semantic grounding, and hallucination reduction as important directions for improving future LLM-based bug detection systems.