Explicit Relation Structures for Visual Content Generation and Quality-Aware Representation Learning
| dc.contributor.author | Naseri, Mahdi | |
| dc.date.accessioned | 2026-09-30T13:41:58Z | |
| dc.date.issued | 2026-09-30 | |
| dc.date.submitted | 2026-09-28 | |
| dc.description.abstract | Modern visual learning systems often depend on relations that are not fully captured by a single label, prompt, image, or scalar quality score. A sparse scene description leaves many plausible objects and relations unstated. A graph self-supervised learner needs guidance about which nodes should share representations beyond raw augmentation pairs. An image-quality representation must account for the structured interaction between content, distortion type, and distortion severity. This thesis studies explicit relation structures as an intermediate language for making such dependencies available to learning systems. The central argument is that relation structures are useful not only as mathematical objects, but also as interfaces for inspecting, training, and evaluating visual models when the relevant dependency is broader than a single paired sample. The thesis develops this argument through three technical studies. First, Generated Contents Enrichment (GCE) formulates visual content enrichment as an explicit scene-graph problem. A sparse scene graph is enriched by predicting additional object and relation semantics before image generation, so the system exposes what content has been added before the image is rendered. This separation lets the work ask whether graph-level enrichment provides inspectable semantic content that supports the rendered image. The chapter therefore evaluates graph recovery, image-side behavior, qualitative examples, and human judgments as complementary evidence about the enrichment process. Second, Explicitly-Generated Relation Graph for Self-Supervised Representation Learning (ExGRG) studies explicit relation construction for node-level graph Self-Supervised Learning (SSL). Standard augmentation-based graph SSL uses paired views of the same node, but this pairwise identity signal does not fully describe which nodes should share representation structure. ExGRG constructs a compositional relation graph from multiple cues, including augmentation identity, graph structure, positional and structural encodings, representation-space neighborhoods, and online clustering. The resulting relation graph weights the self-supervised training signal and provides a concrete object for ablation and training analysis. This chapter shows how explicitly generated relation structure can improve graph representation learning under the tested node-classification protocol while allowing the relation cues behind the invariance signal to be inspected. Third, SHAped Modeling of Implicit Structural Associations for Self-Supervised No-Reference Image Quality Assessment (SHAMISA) applies the same relation-structure perspective to self-supervised No-Reference Image Quality Assessment (NR-IQA). In image quality assessment, perceptual similarity depends on content, distortion family, distortion severity, and the interaction among multiple distortions. SHAMISA constructs metadata-driven and feature-driven relation graphs over distorted images and uses these graphs to shape a quality-aware self-supervised representation before any human quality labels are used for downstream evaluation. The method is evaluated through within-dataset and cross-dataset quality prediction, ablation studies, and representation diagnostics, emphasizing that quality-aware representation learning requires more than generic image-level invariance. Under the reported protocols, GCE raises held-out object accuracy from 15.19% to 20.60% and available-edge recall from 36.64% to 75.18% relative to SceneGraphGen+, while triplet F1 increases from 3.73% to 7.73%. ExGRG has the highest reported node-classification point estimate on all nine tested graph datasets, including 97.87% on Cora and 89.68% on CiteSeer. SHAMISA obtains the strongest six-dataset average among the SSL methods in the within-dataset Image Quality Assessment (IQA) comparison, with average SRCC/PLCC of 0.886/0.904. Across the three studies, the thesis uses relation structures in different roles. They serve as predicted semantic content for generation, as supervisory graphs for self-supervised representation learning, and as explicit structures whose effects can be examined through ablations, visualizations, and stress tests. The common design principle is to make the relevant dependency explicit enough that it can be inspected, modified, and tested. At the same time, the thesis keeps the limits of explicit structure in view. Scene-graph annotations are incomplete, generated relation graphs depend on the quality of their sources, and auxiliary visualizations do not replace quantitative evaluation. The contribution is a set of methods and evaluations showing that carefully constructed relation structures can provide useful control, supervision, and analysis when visual tasks depend on structured dependencies that are otherwise hidden. | |
| dc.identifier.uri | https://hdl.handle.net/10012/24445 | |
| dc.language.iso | en | |
| dc.pending | false | |
| dc.publisher | University of Waterloo | en |
| dc.relation.uri | https://github.com/Mahdi-Naseri/GCE | |
| dc.relation.uri | https://github.com/Mahdi-Naseri/SHAMISA | |
| dc.subject | explicit relation structures | |
| dc.subject | scene graphs | |
| dc.subject | visual content generation | |
| dc.subject | graph self-supervised learning | |
| dc.subject | no-reference image quality assessment | |
| dc.subject | quality-aware representation learning | |
| dc.title | Explicit Relation Structures for Visual Content Generation and Quality-Aware Representation Learning | |
| dc.type | Doctoral Thesis | |
| uws-etd.degree | Doctor of Philosophy | |
| uws-etd.degree.department | Electrical and Computer Engineering | |
| uws-etd.degree.discipline | Electrical and Computer Engineering | |
| uws-etd.degree.grantor | University of Waterloo | en |
| uws-etd.embargo.terms | 0 | |
| uws.contributor.advisor | Wang, Zhou | |
| uws.contributor.affiliation1 | Faculty of Engineering | |
| uws.peerReviewStatus | Unreviewed | en |
| uws.published.city | Waterloo | en |
| uws.published.country | Canada | en |
| uws.published.province | Ontario | en |
| uws.scholarLevel | Graduate | en |
| uws.typeOfResource | Text | en |