Data Science Seminar

Logic based classification for Low Resource Setting

Vivek Gupta

<<< All talks

Logic based classification for Low Resource Setting

When Friday, March 19, 2021, 2:00 PM – 3:00 PM (MT)

Slides

Abstract

An NLP model’s ability to reason should be independent of language. Previous works utilize Natural Language Inference(NLI) to understand the reasoning ability of models, mostly focusing on high resource languages like English. To address scarcity of data in low-resource languages such as Hindi, we use data recasting to create four NLI datasets from existing four text classification datasets in Hindi language. Through experiments, we show that our recasted dataset is devoid of statistical irregularities and spurious patterns. We study the consistency in predictions of the textual entailment models and propose a consistency regulariser to remove pairwise-inconsistencies in predictions. Furthermore, we propose a novel two-step classification method which uses textual-entailment predictions for classification tasks. We further improve the classification performance by jointly training the classification and textual entailment tasks together. We therefore highlight the benefits of data recasting and our approach with supporting experimental results. You can access the dataset and paper here: https://www.aclweb.org/anthology/2020.aacl-main.71.pdf (https://www.aclweb.org/anthology/2020.aacl-main.71.pdf). Joint work with BloomBerg AI.

Speaker

Vivek Gupta

University of Utah

aclweb.org

Tags: machine learning natural language processing


Part of the Data Science Seminar. Something wrong on this page? Edit _data/talks/2021-03-19-vivek-gupta.toml.