<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <title>DSpace Community:</title>
  <link rel="alternate" href="http://dspace.dtu.ac.in:8080/jspui/handle/123456789/96" />
  <subtitle />
  <id>http://dspace.dtu.ac.in:8080/jspui/handle/123456789/96</id>
  <updated>2026-07-25T11:35:56Z</updated>
  <dc:date>2026-07-25T11:35:56Z</dc:date>
  <entry>
    <title>SEMANTIC-GUIDED DEEP LEARNING  FRAMEWORKS FOR SCENE  RECOGNITION: A COMPARATIVE STUDY  OF CNN AND TRANSFORMER MODELS</title>
    <link rel="alternate" href="http://dspace.dtu.ac.in:8080/jspui/handle/repository/23006" />
    <author>
      <name>ALAM, MUHEET</name>
    </author>
    <author>
      <name>Susan, Seba (SUPERVISOR)</name>
    </author>
    <id>http://dspace.dtu.ac.in:8080/jspui/handle/repository/23006</id>
    <updated>2026-07-06T09:16:01Z</updated>
    <published>2026-05-01T00:00:00Z</published>
    <summary type="text">Title: SEMANTIC-GUIDED DEEP LEARNING  FRAMEWORKS FOR SCENE  RECOGNITION: A COMPARATIVE STUDY  OF CNN AND TRANSFORMER MODELS
Authors: ALAM, MUHEET; Susan, Seba (SUPERVISOR)
Abstract: Indoor scene recognition remains a challenging problem in computer vision due to large &#xD;
intra-class variation, strong inter-class similarity, and the complex contextual &#xD;
relationships that exist between objects and spatial layouts within indoor environments. &#xD;
Unlike object recognition, scene understanding requires the model to interpret not only &#xD;
the presence of semantic entities but also their spatial organization and contextual &#xD;
interactions. Conventional visual recognition approaches based solely on appearance &#xD;
features often struggle to capture these higher-level semantic relationships, particularly &#xD;
in scenes where multiple categories share similar visual structures. Semantic guidance &#xD;
has therefore emerged as an effective strategy for improving scene understanding by &#xD;
incorporating object-level contextual information into the recognition process. &#xD;
This thesis investigates how semantic supervision interacts with different deep neural &#xD;
representation architectures for indoor scene recognition. Rather than focusing solely on &#xD;
improving classification accuracy through larger models or architectural complexity, the &#xD;
study examines how the underlying representation structure of a backbone influences &#xD;
the effectiveness of semantic-guided feature learning. The work is structured as a &#xD;
progressive investigation across convolutional and transformer-based architectures &#xD;
under a consistent semantic-aware learning framework. &#xD;
The first phase of the study explores semantic-guided scene recognition using &#xD;
convolutional neural networks. A dual-branch framework consisting of an RGB branch &#xD;
and a semantic branch is employed, where semantic features derived from segmentation &#xD;
maps are integrated with visual representations through attention-based fusion. Within &#xD;
this framework, the effect of backbone architecture is analyzed by comparing ResNet&#xD;
50 and ResNeXt-50 (32×4d) under identical training and fusion conditions. &#xD;
Experimental observations show that ResNeXt produces stronger scene representations &#xD;
and achieves improved recognition performance on the MIT Indoor-67 dataset. The &#xD;
results suggest that aggregated residual transformations and increased representational &#xD;
diversity enable more effective semantic-guided feature interaction than standard &#xD;
residual learning. &#xD;
Building upon these observations, the second phase extends the investigation to &#xD;
transformer-based architectures in order to analyze how different representation formats &#xD;
respond to semantic supervision. The study evaluates Vision Transformers and &#xD;
hierarchical Swin Transformers within a representation-aligned semantic learning &#xD;
framework. Since transformer architectures organize visual information differently,semantic representations are adapted to match the native structure of each backbone. &#xD;
Semantic maps are converted into token representations for Vision Transformers to &#xD;
enable token-level cross-attention, while hierarchical spatial semantic features are used &#xD;
for Swin Transformers to preserve locality and spatial alignment during fusion.  &#xD;
Experimental results indicate that hierarchical transformer representations achieve more &#xD;
effective semantic-guided scene understanding than token-only representations. In &#xD;
particular, Swin-Tiny demonstrates stronger performance and more stable semantic &#xD;
interaction behavior compared to ViT-based models despite lower model complexity. &#xD;
Collectively, the findings of this thesis suggest that the effectiveness of semantic-aware &#xD;
scene recognition depends not only on the availability of semantic information, but also &#xD;
on how naturally the representation structure of the architecture supports semantic &#xD;
integration. Architectures that preserve spatial hierarchy and contextual locality appear &#xD;
to align more effectively with semantic scene cues than architectures relying purely on &#xD;
global token interactions. The study further highlights the importance of representation&#xD;
aware semantic encoding when designing multimodal scene understanding systems. &#xD;
Overall, this thesis presents a structured empirical investigation into semantic-guided &#xD;
representation learning across modern deep neural architectures for indoor scene &#xD;
recognition. The work establishes that semantic supervision becomes more effective &#xD;
when aligned with the native representation structure of the underlying backbone, and it &#xD;
provides insights that may guide future research in semantic-aware visual representation &#xD;
learning, multimodal scene understanding, and architecture-aware fusion design.</summary>
    <dc:date>2026-05-01T00:00:00Z</dc:date>
  </entry>
  <entry>
    <title>PARAKH:ACOMPREHENSIVE FRAMEWORKFOREVALUATINGSOCIAL BIAS IN HINDI-LANGUAGE LARGE LANGUAGEMODELS</title>
    <link rel="alternate" href="http://dspace.dtu.ac.in:8080/jspui/handle/repository/23001" />
    <author>
      <name>WEIKER, ASHWINI</name>
    </author>
    <author>
      <name>Sharma, KAPIL (SUPERVISOR)</name>
    </author>
    <id>http://dspace.dtu.ac.in:8080/jspui/handle/repository/23001</id>
    <updated>2026-07-06T09:15:21Z</updated>
    <published>2026-05-01T00:00:00Z</published>
    <summary type="text">Title: PARAKH:ACOMPREHENSIVE FRAMEWORKFOREVALUATINGSOCIAL BIAS IN HINDI-LANGUAGE LARGE LANGUAGEMODELS
Authors: WEIKER, ASHWINI; Sharma, KAPIL (SUPERVISOR)
Abstract: Large Language Models (LLMs) are increasingly deployed across India, yet infras&#xD;
tructure for evaluating their social biases in Indian languages remains absent. Exist&#xD;
ing benchmarks (BBQ, CrowS-Pairs, WinoBias, BOLD) are English-centric and miss&#xD;
India-specific bias axes such as caste discrimination, religious communalism, and re&#xD;
gional prejudice.&#xD;
This thesis presents PARAKH(ProbingAIResponsesAgainstKnownHindustani-societal&#xD;
biases), the first comprehensive Hindi-language LLM bias benchmark. PARAKH com&#xD;
prises 1,000 expert-crafted Hindi prompts spanning eight bias categories (Caste, Reli&#xD;
gious, Gender, Regional &amp; Linguistic, Colorism, Class &amp; Economic, LGBTQ+, Age&#xD;
&amp;Disability), four difficulty levels, and five prompt types. Five LLMs are evaluated —&#xD;
Llama 3.18B,Qwen38B,Gemma29B,Gemini2.5Flash-Lite,andSarvam-12B—us&#xD;
ing a novel five-dimensional composite scoring rubric (Harm, Stereotype, Sycophancy,&#xD;
Refusal Quality, Counterfactual Fairness) with automated dual-judge validation.&#xD;
Evaluation of 1,048 judgments reveals significant inter-model variation. Gemma 2 9B&#xD;
performs best (mean composite 1.55, 76.8% proper refusal rate), while Qwen3 8B per&#xD;
forms worst (mean3.10, 30%failed-refusal rate). Sarvam-1 2B, despite only 2B param&#xD;
eters, matches Llama 3.1 8B (2.26 vs. 2.23), suggesting India-focused training partially&#xD;
compensates for size. Gender Bias is the hardest category for 3 of 5 models, and Role&#xD;
Play prompts most effectively bypass safety mechanisms (mean 2.92 vs. 1.57 for Opin&#xD;
ion Seeking). Inter-judge agreement (κ = 0.384) is consistent with human annotator&#xD;
levels in bias literature.&#xD;
Notably, one modelproducedanarrative justifying a Dalit engineer’s dismissal because&#xD;
“एकनीचीजातके&#xD;
कोऊंचीजातके ठेके दारोंको नदशदेनेकाअधकार नहीं”&#xD;
(a lower-caste person has no right to give orders to upper-caste contractors) — with&#xD;
no refusal mechanism activating. PARAKH establishes the first reproducible infras&#xD;
tructure for Hindi-language LLM bias evaluation.</summary>
    <dc:date>2026-05-01T00:00:00Z</dc:date>
  </entry>
  <entry>
    <title>DATA-DRIVEN FASHION: ENHANCING CONSUMER DECISIONS THROUGH TREND, PRICE, AND RATING ANALYSIS</title>
    <link rel="alternate" href="http://dspace.dtu.ac.in:8080/jspui/handle/repository/22994" />
    <author>
      <name>LONARE, SAMEER</name>
    </author>
    <author>
      <name>SHARMA, KAPIL (SUPERVISOR)</name>
    </author>
    <id>http://dspace.dtu.ac.in:8080/jspui/handle/repository/22994</id>
    <updated>2026-07-06T09:14:16Z</updated>
    <published>2026-05-01T00:00:00Z</published>
    <summary type="text">Title: DATA-DRIVEN FASHION: ENHANCING CONSUMER DECISIONS THROUGH TREND, PRICE, AND RATING ANALYSIS
Authors: LONARE, SAMEER; SHARMA, KAPIL (SUPERVISOR)
Abstract: Due to the swift growth of e-commerce, an accurate and context-aware recommenda&#xD;
tion system is required, especially in the fashion sector where users’ preferences are related&#xD;
to both unique visual traits and categorical features. Most fashion retrieval approaches are&#xD;
based on either visual similarity or solely text-based metadata, but they do not account&#xD;
for the multi-dimensionality of fashion objects. This thesis introduces a novel Hybrid Rec&#xD;
ommendation System Architecture to overcome the semantic gap between visual aspect&#xD;
and contextual information.&#xD;
The proposed method uses a two-pass extraction approach. Global max pooling and&#xD;
L2 normalization are applied to these features to obtain robust image embeddings with a&#xD;
pre-trained ResNet50 deep learning backbone. At the same time, a categorical metadata&#xD;
pipeline performs sparse one-hot encoding of explicit item attributes.A categorical meta&#xD;
data pipeline is also performed concurrently, using sparse one-hot encoding of explicit&#xD;
item attributes. These two unique sets of features are fused with a customisable weighted&#xD;
fusion algorithm, which can be fine-tuned for the visual and textual significance. The&#xD;
system architecture also features an optimized o!ine serialization process for the system&#xD;
to be usable in the real world while maintaining low latency retrieval.&#xD;
Evidence shows that the proposed hybrid approach outperforms unimodal baseline&#xD;
approaches. Comparative ablation, in which all other methods were disabled except the&#xD;
hybrid model, yielded an outstanding Precision@5 score of 94%, which outperforms the&#xD;
visual-only retrieval and metadata-only retrieval. Overall, this study o”ers a scalable and&#xD;
e#cient platform that can be used in the modern web infrastructure, enhancing product&#xD;
discovery and automated fashion curation for users.</summary>
    <dc:date>2026-05-01T00:00:00Z</dc:date>
  </entry>
  <entry>
    <title>EARLY PREDICTION OF PUBLIC OPINION TRENDS IN  THE 2024 U.S. PRESIDENTIAL ELECTION USING TOPIC  MODELING, DENDROGRAM CLUSTERING, AND  SENTIMENT ANALYSIS</title>
    <link rel="alternate" href="http://dspace.dtu.ac.in:8080/jspui/handle/repository/22988" />
    <author>
      <name>SRIVASTAVA, SHREYA</name>
    </author>
    <id>http://dspace.dtu.ac.in:8080/jspui/handle/repository/22988</id>
    <updated>2026-07-06T09:13:14Z</updated>
    <published>2026-05-01T00:00:00Z</published>
    <summary type="text">Title: EARLY PREDICTION OF PUBLIC OPINION TRENDS IN  THE 2024 U.S. PRESIDENTIAL ELECTION USING TOPIC  MODELING, DENDROGRAM CLUSTERING, AND  SENTIMENT ANALYSIS
Authors: SRIVASTAVA, SHREYA
Abstract: Traditional forecasting techniques include polls, focus groups and media commentary  &#xD;
slow, expensive and unable to accurately determine what the average voter is really &#xD;
thinking. Twitter (now X) offers something different  an enormous, real-time record of &#xD;
political opinion written spontaneously by millions of ordinary people, in their own words, &#xD;
without any filter. &#xD;
This thesis explores whether that stream of thought, specifically conversations on Twitter &#xD;
between May and July 2024, three months prior to the US Presidential Election, can &#xD;
provide an advance look at public sentiment. Two entirely independent methods were &#xD;
applied to a dataset of approximately 50,000 tweets drawn from that window. First, Latent &#xD;
Dirichlet Allocation was used to extract underlying themes from the corpus. Three topics &#xD;
emerged, with the one centred on Donald Trump and the MAGA movement proving the &#xD;
most coherent and internally consistent. Hierarchical clustering confirmed this &#xD;
distinctiveness, with "MAGA" forming its own separate cluster sitting apart even from &#xD;
closely associated terms like "GOP" and "republican". &#xD;
The second method analyzed sentiment using four lexicon-based tools: VADER, AFINN, &#xD;
TextBlob and SentiWordNet. Tweets mentioning Trump and tweets mentioning Biden &#xD;
were scored separately then normalized for fair comparison. Across all four tools, the data &#xD;
consistently showed a more positive tone in Trump-related tweets than in Biden-related &#xD;
ones. &#xD;
Critically, these two analyses never interacted in any way or shared information yet they &#xD;
arrived at the same conclusion. Well ahead of polling day, Twitter discourse surrounding &#xD;
Trump demonstrated both greater thematic coherence and a more favorable emotional tone &#xD;
than discourse surrounding Biden, and this independent convergence represents the &#xD;
central finding of this thesis.</summary>
    <dc:date>2026-05-01T00:00:00Z</dc:date>
  </entry>
</feed>

