Data Science Seminar

InfoTabS: Inference on Tables as Semi-Structured Data

Vivek Gupta

<<< All talks

InfoTabS: Inference on Tables as Semi-Structured Data

When Friday, August 28, 2020, 11:50 AM – 1:10 PM (MT)

Recording

Abstract

Experience of the everyday language indicates the use of complicated reasonings both for people and the AI systems. Natural Language Inference (NLI) is the process of reasoning about inferential relationships, meaning to establish whether a hypothesis is a true (entailment), false (contradiction), or undetermined (neutral) given a premise. Previous works have generated inference corpora, such as the SNLI and the MNLI, which comprise only unstructured representations of text in the form of sentences in which relationships between words are explicitly expressed, and often need information extraction. However, text can also occur universally in other structured forms like tables, graphs, and databases.Building upon previous work on large-scale datasets for inference, we introduce a new dataset called INFOTABS, comprising of human-written textual hypotheses based on premises that are tables extracted from Wikipedia info-boxes. Our analysis shows that the semi-structured, multi-domain, and heterogeneous nature of the premises admits complex, multi-faceted reasoning. Experiments reveal that, while human annotators agree on the relationships between a table-hypothesis pair, several standard modeling strategies are unsuccessful at the task, suggesting that reasoning about tables can pose a new modeling challenge. For more details on InfoTabS visit http://infotabs.github.io (http://infotabs.github.io/)

Speaker

Vivek Gupta

UoU

vgupta123.github.io

Vivek is a Ph.D. student at the School of Computing, University of Utah. Previously, he was working as a Research Fellow in Microsoft Research Lab, India, in Machine Learning and Natural Language Processing group. He graduated as a dual degree student in the Department of Computer Science and Engineering at IIT Kanpur in 2016. He is broadly interested in research in the field of Machine Learning and Natural Language Processing. To know more about his current research interest, you can visit https://vgupta123.github.io/ (https://www.youtube.com/redirect?q=https%3A%2F%2Fvgupta123.github.io%2F&v=YhfU1BON8EI&event=video_description&redir_token=QUFFLUhqbFVhUUxOS2RDSHBPYWJPMjNqaFpoalpzY2E3d3xBQ3Jtc0tuWlZRQ0tfSTdPNkxsTzFuWDN2Y3lvdXdwRXJrOTJ6bGI3NXhrc2x4QVYxalNTTUVXWE10V3VncklsZ2Q1RnltN1VaWkpsek9OY0EtdWwtQWZNVHR0Q2k4Rng3eVE4N3FWYTk5dUtETV9aZXRSaGlEcw%3D%3D)

Tags: data management education machine learning natural language processing


Part of the Data Science Seminar. Something wrong on this page? Edit _data/talks/2020-08-28-vivek-gupta.toml.