Data Science Seminar
Logic based classification for Low Resource Setting
Vivek Gupta
Logic based classification for Low Resource Setting
| When | Friday, March 19, 2021, 2:00 PM – 3:00 PM (MT) |
|---|
Abstract
An NLP model’s ability to reason should be independent of language. Previous works utilize Natural Language Inference(NLI) to understand the reasoning ability of models, mostly focusing on high resource languages like English. To address scarcity of data in low-resource languages such as Hindi, we use data recasting to create four NLI datasets from existing four text classification datasets in Hindi language. Through experiments, we show that our recasted dataset is devoid of statistical irregularities and spurious patterns. We study the consistency in predictions of the textual entailment models and propose a consistency regulariser to remove pairwise-inconsistencies in predictions. Furthermore, we propose a novel two-step classification method which uses textual-entailment predictions for classification tasks. We further improve the classification performance by jointly training the classification and textual entailment tasks together. We therefore highlight the benefits of data recasting and our approach with supporting experimental results. You can access the dataset and paper here: https://www.aclweb.org/anthology/2020.aacl-main.71.pdf (https://www.aclweb.org/anthology/2020.aacl-main.71.pdf). Joint work with BloomBerg AI.
Speaker
Tags: machine learning natural language processing
Part of the Data Science Seminar. Something wrong on this page? Edit _data/talks/2021-03-19-vivek-gupta.toml.