BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//Utah Center for Data Science//Talks//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:Utah Data Science & AI Lecture Series
X-WR-CALDESC:Talks in the Data Science & AI Lecture Series at the Utah Cent
 er for Data Science. Generated from the talk records at datascience.utah.e
 du/talks/.
X-WR-TIMEZONE:America/Denver
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/Denver
BEGIN:DAYLIGHT
TZOFFSETFROM:-0700
TZOFFSETTO:-0600
TZNAME:MDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0600
TZOFFSETTO:-0700
TZNAME:MST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
UID:2020-01-09-chris-musco@datascience.utah.edu
DTSTAMP:20200109T191500Z
DTSTART;TZID=America/Denver:20200109T121500
DTEND;TZID=America/Denver:20200109T133000
SUMMARY:Randomized FunctionalAnalysis - Chris Musco
DESCRIPTION:Chris Musco (NYU)\n\nSketching and subsampling are central algo
 rithmic tools in scaling statisticalmethods to very large datasets. These 
 techniques seek to quickly compress datadown to a compact set of informati
 ve features or examples\, which can then beprocessed in place of the origi
 nal data\, at much lower computational cost. Thecentral question of this t
 alk is what sketching methods can teach us abouteffective machine learning
  and data analysis in the small data regime. Inapplications where high qua
 lity data examples remain a rare luxury\, can ourknowledge of data sketchi
 ng guide more efficient initial data collection?\n\nWe study this problem 
 by focusing specifically on techniques for large matrixcomputations. In th
 e field of randomized numerical linear algebra\, importancesampling has em
 erged as an important tool for dataset compression. Statisticalleverage sc
 ores and related measures are used to judge the importance of rowsor colum
 ns in a matrix\, which are then non-uniformly subsampled\, leading tofaste
 r algorithms for regression\, low-rank approximation\, kernel methods\, an
 dmany other data problems.\n\nI will introduce a simple generalization of 
 leverage score sampling to infinitedimensional linear operators and show t
 he potential of this generalization indeveloping sample efficient algorith
 ms for small data applications.Specifically\, I will survey a number of re
 cent results on robust polynomialcurve fitting\, bandlimited function inte
 rpolation\, off-grid sparse Fouriertransforms\, and sample efficient covar
 iance estimation. I will illustrateconnections between these new results a
 nd classical tools in approximationtheory and signal processing\, and will
  discuss several open researchdirections.\n\nhttps://datascience.utah.edu/
 talks/2020-01-09-chris-musco/
URL:https://datascience.utah.edu/talks/2020-01-09-chris-musco/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147
CATEGORIES:algorithms & theory,machine learning,statistics
END:VEVENT
BEGIN:VEVENT
UID:2020-01-16-alexander-lex@datascience.utah.edu
DTSTAMP:20200116T191500Z
DTSTART;TZID=America/Denver:20200116T121500
DTEND;TZID=America/Denver:20200116T133000
SUMMARY:Literate Visualization: Making Visual Analysis Sessions Reproducibl
 e and Reusable - Alexander Lex
DESCRIPTION:Alexander Lex\n\nInteractive visualization is an important part
  of the data science process. It enables analysts to directly interact wit
 h the data\, exploring it with minimal effort. Unlike code\, however\, an 
 interactive visualization session is ephemeral and can’t be easily share
 d\, revisited\, or reused. Computational notebooks\, such as Jupyter Noteb
 ooks\, R Markdown\, or Observable are a perfect match for many data scienc
 e applications. They are also the most popular embodiment of Knuth’s “
 Literate Programming”\, where the logic of a program is explained in nat
 ural language\, figures\, and equations. In this talk\, I will sketch appr
 oaches to “Literate Visualization”. I will show how we can leverage pr
 ovenance data of an analysis session to create well-documented and annotat
 ed visualization stories that enable reproducibility and sharing. I will a
 lso introduce early work on semi-automatically inferring mid-level analysi
 s goals\, which allows us to understand the analysis process at a higher l
 evel. Understanding analysis goals enables us to speed up interactions and
  even re-used visual analysis processes.\n\nhttps://datascience.utah.edu/t
 alks/2020-01-16-alexander-lex/
URL:https://datascience.utah.edu/talks/2020-01-16-alexander-lex/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147
CATEGORIES:visualization
END:VEVENT
BEGIN:VEVENT
UID:2020-01-23-harish-maringanti@datascience.utah.edu
DTSTAMP:20200123T191500Z
DTSTART;TZID=America/Denver:20200123T121500
DTEND;TZID=America/Denver:20200123T133000
SUMMARY:Data Science projects in Marriott Library - Harish Maringanti
DESCRIPTION:Harish Maringanti (Associate Dean for IT & Digital Library Serv
 ices\, Marriott Library)\n\nAt research intensive universities\, libraries
  have traditionally supported data science activities in various ways incl
 uding acquiring datasets that researchers need\, hosting workshops and tra
 ining sessions on data science tools\, and offering data support tools for
  creation of persistent identifiers(dois)\, etc. At Marriott Library\, in 
 addition to supporting data science programs on campus\, we have embarked 
 on a suite of data science projects to add value to our culturally-rich co
 llections. Our efforts are focused on enriching our collection data\, and 
 making this collection data available for computational use (collections a
 s data [1]) so that developers\, scientists\, and digital humanists can pr
 ogrammatically interact with the data in myriad ways and undertake project
 s related to data mining & text analysis\, advanced visualizations\, and g
 eospatial analysis. In this presentation\, I will talk about two specific 
 projects - Utah Digital Newspapers [2] and machine learning meets archives
  [3] - to highlight these efforts.\n\nUtah Digital Newspapers (UDN): Marri
 ott Library was an early pioneer in digitizing newspapers and making the c
 ontent available to historians\, researchers\, and lifelong learners. UDN 
 program has been operating since 2002 and is recognized as one of the lead
 ers in newspaper digitization in the United States. We have continued to p
 artner with universities\, colleges\, state agencies\, county and city lib
 raries\, and other agencies to digitize\, deliver\, and archive historical
  newspaper collections\; As of 2019\, UDN has well over 22.5 million newsp
 aper articles and 3.5 million pages in the repository platform. In this pr
 esentation\, we will talk about the importance of looking at collections a
 s data\, our API work with UDN and demonstrate the usefulness of this appr
 oach with specific examples.\n\nMachine learning meets archives: Metadata 
 is the bedrock of library archives and Digital Library systems\, as it hel
 ps in users discovering the unique content in various collections housed i
 n digital libraries. But creating metadata is a time-intensive process. We
  are working with machine learning algorithms to generate descriptive meta
 data for digital images. I will share the results of our work\, and also l
 essons learned from working with digital library data.\n\nhttps://datascie
 nce.utah.edu/talks/2020-01-23-harish-maringanti/
URL:https://datascience.utah.edu/talks/2020-01-23-harish-maringanti/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147
CATEGORIES:machine learning
END:VEVENT
BEGIN:VEVENT
UID:2020-01-30-qingyao-ai@datascience.utah.edu
DTSTAMP:20200130T191500Z
DTSTART;TZID=America/Denver:20200130T121500
DTEND;TZID=America/Denver:20200130T133000
SUMMARY:Unbiased Learning to Rank: Theory and Practice - Qingyao Ai
DESCRIPTION:Qingyao Ai (Utah SoC)\n\nImplicit feedback (e.g.\, user clicks)
  is an important source of data for modern search engines. While heavily b
 iased\, it is cheap to collect and particularly useful for user-centric re
 trieval applications such as search ranking. Therefore\, a learning-to-ran
 k algorithm that can effectively learn from implicit user feedback without
  affected by its inherent biases could fundamentally change the design of 
 ranking systems and significantly improve the quality of modern search eng
 ines. To develop an unbiased learning-to-rank system with biased feedback\
 , previous studies have focused on constructing probabilistic graphical mo
 dels (e.g.\, click models) with user behavior hypothesis to extract and tr
 ain ranking systems with unbiased relevance signals. Recently\, a novel co
 unterfactual learning framework that estimates and adopts examination prop
 ensity for unbiased learning to rank has attracted much attention. In this
  talk\, we aim to provide an overview of the fundamental mechanism for unb
 iased learning to rank. We describe the theory behind existing frameworks\
 , and give instructions on how to conduct unbiased learning to rank in pra
 ctice.\n\nhttps://datascience.utah.edu/talks/2020-01-30-qingyao-ai/
URL:https://datascience.utah.edu/talks/2020-01-30-qingyao-ai/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147
CATEGORIES:fairness & ethics
END:VEVENT
BEGIN:VEVENT
UID:2020-02-06-gail-zasowski@datascience.utah.edu
DTSTAMP:20200206T191500Z
DTSTART;TZID=America/Denver:20200206T121500
DTEND;TZID=America/Denver:20200206T133000
SUMMARY:Big Data\, Big Universe: Data-Driven Discoveries in Astrophysics - 
 Gail Zasowski
DESCRIPTION:Gail Zasowski (Utah Physics & Astronomy)\n\nThe stars in the ni
 ght sky have inspired questions about our place in the Universe throughout
  history. The development of telescopes showed us that the stars visible t
 o the naked eye are but a tiny fraction of their vast numbers within our o
 wn Galaxy\, and revealed energy signatures invisible to the human senses. 
 We now know that there are billions of stars in our galaxy\, billions of g
 alaxies in our Universe\, and nearly 14 billion years of cosmic evolution 
 that have led to where and what we are today. As the volume of astronomica
 l data grows at an ever quickening rate\, new discoveries increasingly com
 e from careful mining and analysis of existing data\, often used in unfore
 seen ways. This talk will describe some of the major unanswered questions 
 in astrophysics\, and how new data-driven analysis techniques are uncoveri
 ng new insights into solving them.\n\nhttps://datascience.utah.edu/talks/2
 020-02-06-gail-zasowski/
URL:https://datascience.utah.edu/talks/2020-02-06-gail-zasowski/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147
CATEGORIES:physics & astronomy
END:VEVENT
BEGIN:VEVENT
UID:2020-02-20-bei-wang@datascience.utah.edu
DTSTAMP:20200220T191500Z
DTSTART;TZID=America/Denver:20200220T121500
DTEND;TZID=America/Denver:20200220T133000
SUMMARY:TopoAct: Exploring the Shape of Activations in Deep Learning - Bei 
 Wang
DESCRIPTION:Bei Wang (Utah SoC\, SCI)\n\nDeep neural networks such as GoogL
 eNet and ResNet have achieved superhuman performance in tasks like image c
 lassification. To understand how such superior performance is achieved\, w
 e can probe a trained deep neural network by studying neuron activations\,
  that is\, combinations of neuron firings\, at any layer of the network in
  response to a particular input. With a large set of input images\, we aim
  to obtain a global view of what neurons detect by studying their activati
 ons. We ask the following questions: What is the shape of the space of act
 ivations? That is\, what is the organizational principle behind neuron act
 ivations\, and how are the activations related within a layer and across l
 ayers? Applying tools from topological data analysis\, we present TopoAct\
 , a visual exploration system used to study topological summaries of activ
 ation vectors for a single layer as well as the evolution of such summarie
 s across multiple layers. We present visual exploration scenarios using To
 poAct that provide valuable insights towards learned representations of an
  image classifier.\n\nhttps://datascience.utah.edu/talks/2020-02-20-bei-wa
 ng/
URL:https://datascience.utah.edu/talks/2020-02-20-bei-wang/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147
CATEGORIES:deep learning,machine learning
END:VEVENT
BEGIN:VEVENT
UID:2020-02-27-taylor-sparks@datascience.utah.edu
DTSTAMP:20200227T191500Z
DTSTART;TZID=America/Denver:20200227T121500
DTEND;TZID=America/Denver:20200227T133000
SUMMARY:New Algorithms\, Descriptors\, and Machine Learning Techniques Tail
 ored to the Challenges of Materials Informatics - Taylor Sparks
DESCRIPTION:Taylor Sparks (Utah Materials Science & Engineering)\n\nNew mat
 erials are required to address many of the energy\, environmental\, and te
 chnological needs of the present and future. Materials Informatics\, or th
 e application of data science techniques to solve materials research chall
 enges\, is expected to play a key role in materials development and discov
 ery given the infinite palette available for new materials. Interestingly\
 , the requirements and tasks of Materials Informatics do not always overla
 p with general machine learning. Therefore\, adopting existing data scienc
 e tools including visualization\, algorithms\, featurization schemas etc m
 ay not provide the ideal outcomes for the specific needs of Materials Info
 rmatics.\n\nIn this talk I will focus on some of our recent work to bring 
 tailored data science approaches to actual materials research problems. Sp
 ecifically\, I will introduce how we have created an attention-based neura
 l network architecture for the prediction of materials properties. We show
  that this novel algorithm outperforms other methods in the absence of che
 mical information\, even when the statistical and ensemble learning techni
 ques are given domain-specific chemical knowledge about the materials.\n\n
 Dr. Sparks is an Associate Professor and Associate Chair of the Materials 
 Science and Engineering Department at the University of Utah. He is origin
 ally from Utah and an alumni of the department he now teaches in. Before g
 raduate school he worked at Ceramatec Inc. He did his MS in Materials at U
 CSB and his PhD in Applied Physics at Harvard University in David Clarke
 ’s laboratory and then did a postdoc with Ram Seshadri in the Materials 
 Research Laboratory at UCSB. His current research centers on the discovery
 \, synthesis\, characterization\, and properties of new materials for ener
 gy applications. He is a pioneer in the emerging field of materials inform
 atics whereby big data\, data mining\, and machine learning are leveraged 
 to solve challenges in materials science. He also hosts a podcast entitled
  “Materialism” where he discusses the past\, present\, and future of M
 aterials Science.\n\nhttps://datascience.utah.edu/talks/2020-02-27-taylor-
 sparks/
URL:https://datascience.utah.edu/talks/2020-02-27-taylor-sparks/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147
CATEGORIES:algorithms & theory,machine learning,physics & astronomy
END:VEVENT
BEGIN:VEVENT
UID:2020-03-05-john-horel@datascience.utah.edu
DTSTAMP:20200305T191500Z
DTSTART;TZID=America/Denver:20200305T121500
DTEND;TZID=America/Denver:20200305T133000
SUMMARY:The Promise and Perils of Big Data in the Cloud - Examples from the
  Atmospheric Sciences - John Horel
DESCRIPTION:John Horel (Professor\, Chair\, Department of Atmospheric Scien
 ces)\n\nFrom the inception of numerical weather prediction in the 1950’s
 \, atmospheric scientists have stretched the envelope on the hardware and 
 procedures available to write\, store\, and use data on mass storage syste
 ms. The opportunities now to rely on cloud resources to process\, access\,
  and disseminate environmental data offer improved capabilities for data s
 cience applications in the atmospheric sciences but also introduce complex
 ities for university researchers.\n\nThe Big Data Project of the National 
 Oceanographic and Atmospheric Administration is assessing the potential be
 nefits of storing in the cloud observations and weather and climate model 
 output that are generating petabytes of data daily. Retrieving\, archiving
 \, analyzing\, and disseminating only a fraction of this environmental inf
 ormation has required us to move beyond computational approaches tradition
 ally used within the atmospheric science community. For example\, hundreds
  of users rely on a 140+ Tbyte archive we maintain of High Resolution Rapi
 d Refresh (HRRR) model output on the Pando system of the University’s Ce
 nter for High Performance Computing. Computing resources available nationw
 ide as part of the Open Science Grid- a high-throughput computing resource
 - have been used to analyze wildland fire events. Data analytical methods 
 are being explored that rely on Zarr data compression to provide classes a
 nd functions for working with N-dimensional arrays.\n\nhttps://datascience
 .utah.edu/talks/2020-03-05-john-horel/
URL:https://datascience.utah.edu/talks/2020-03-05-john-horel/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147
CATEGORIES:climate & environment
END:VEVENT
BEGIN:VEVENT
UID:2020-08-28-vivek-gupta@datascience.utah.edu
DTSTAMP:20200828T175000Z
DTSTART;TZID=America/Denver:20200828T115000
DTEND;TZID=America/Denver:20200828T131000
SUMMARY:InfoTabS: Inference on Tables as Semi-Structured Data - Vivek Gupta
DESCRIPTION:Vivek Gupta (UoU)\n\nExperience of the everyday language indica
 tes the use of complicated reasonings both for people and the AI systems. 
 Natural Language Inference (NLI) is the process of reasoning about inferen
 tial relationships\, meaning to establish whether a hypothesis is a true (
 entailment)\, false (contradiction)\, or undetermined (neutral) given a pr
 emise. Previous works have generated inference corpora\, such as the SNLI 
 and the MNLI\, which comprise only unstructured representations of text in
  the form of sentences in which relationships between words are explicitly
  expressed\, and often need information extraction. However\, text can als
 o occur universally in other structured forms like tables\, graphs\, and d
 atabases.Building upon previous work on large-scale datasets for inference
 \, we introduce a new dataset called INFOTABS\, comprising of human-writte
 n textual hypotheses based on premises that are tables extracted from Wiki
 pedia info-boxes. Our analysis shows that the semi-structured\, multi-doma
 in\, and heterogeneous nature of the premises admits complex\, multi-facet
 ed reasoning. Experiments reveal that\, while human annotators agree on th
 e relationships between a table-hypothesis pair\, several standard modelin
 g strategies are unsuccessful at the task\, suggesting that reasoning abou
 t tables can pose a new modeling challenge. For more details on InfoTabS v
 isit http://infotabs.github.io (http://infotabs.github.io/)\n\nRecording: 
 https://www.youtube.com/redirect?q=https%3A%2F%2Fvgupta123.github.io%2F&v=
 YhfU1BON8EI&event=video_description&redir_token=QUFFLUhqbFVhUUxOS2RDSHBPYW
 JPMjNqaFpoalpzY2E3d3xBQ3Jtc0tuWlZRQ0tfSTdPNkxsTzFuWDN2Y3lvdXdwRXJrOTJ6bGI3
 NXhrc2x4QVYxalNTTUVXWE10V3VncklsZ2Q1RnltN1VaWkpsek9OY0EtdWwtQWZNVHR0Q2k4Rn
 g3eVE4N3FWYTk5dUtETV9aZXRSaGlEcw%3D%3D\n\nhttps://datascience.utah.edu/tal
 ks/2020-08-28-vivek-gupta/
URL:https://datascience.utah.edu/talks/2020-08-28-vivek-gupta/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://us02web.zoom.us/j/83503994251?pwd=Wm01bk40UWxva2pSV0dTbFZH
 bGtLQT09
CATEGORIES:data management,education,machine learning,natural language proc
 essing
END:VEVENT
BEGIN:VEVENT
UID:2020-09-04-dheeraj-mekala@datascience.utah.edu
DTSTAMP:20200904T175000Z
DTSTART;TZID=America/Denver:20200904T115000
DTEND;TZID=America/Denver:20200904T131000
SUMMARY:Contextualized Weak Supervision for Text Classification - Dheeraj M
 ekala
DESCRIPTION:Dheeraj Mekala (UCSD)\n\nWeakly supervised text classification 
 based on a few user-provided seed words has recently attracted much attent
 ion from researchers. Existing methods mainly generate pseudo-labels in a 
 context-free manner (e.g.\, string matching)\, therefore\, the ambiguous\,
  context-dependent nature of human language has been long overlooked. In t
 his paper\, we propose a novel framework ConWea\, providing contextualized
  weak supervision for text classification. Specifically\, we leverage cont
 extualized representations of word occurrences and seed word information t
 o automatically differentiate multiple interpretations of the same word\, 
 and thus create a contextualized corpus. This contextualized corpus is fur
 ther utilized to train the classifier and expand seed words in an iterativ
 e manner. This process not only adds new contextualized\, highly label-ind
 icative keywords but also disambiguates initial seed words\, making our we
 ak supervision fully contextualized. Extensive experiments and case studie
 s on real-world datasets demonstrate the necessity and significant advanta
 ges of using contextualized weak supervision\, especially when the class l
 abels are fine-grained.\n\nhttps://datascience.utah.edu/talks/2020-09-04-d
 heeraj-mekala/
URL:https://datascience.utah.edu/talks/2020-09-04-dheeraj-mekala/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://us02web.zoom.us/j/83503994251?pwd=Wm01bk40UWxva2pSV0dTbFZH
 bGtLQT09
CATEGORIES:machine learning,natural language processing
END:VEVENT
BEGIN:VEVENT
UID:2020-09-11-parthe-pandit@datascience.utah.edu
DTSTAMP:20200911T175000Z
DTSTART;TZID=America/Denver:20200911T115000
DTEND;TZID=America/Denver:20200911T131000
SUMMARY:Characterizing the asymptotic performance of inverse problems over 
 Deep Networks - Parthe Pandit
DESCRIPTION:Parthe Pandit (UCLA)\n\nAt the heart of Machine Learning lies t
 he question of generalizability of learned models over previously unseen d
 ata. While over-parameterized models based on neural networks are now ubiq
 uitous in machine learning applications\, our understanding of their gener
 alization capabilities remains incomplete. This task is made harder by the
  non-convexity of the underlying estimation/learning problems.\n\nThe solu
 tions to some of these estimation problems can be analyzed using a class o
 f algorithms called Approximate Message Passing\, even in the presence of 
 the non-convexity. The dynamics of this algorithm follow a simplified macr
 oscopic description often called the State Evolution. This analytical tool
  allows us to provide some insights into two broad classes of problems rel
 ated to Neural Networks:\n\n1. Generalization error in 1 and 2-layer Netwo
 rks\n2. Image Reconstruction error with Deep Image Priors\n\nhttps://datas
 cience.utah.edu/talks/2020-09-11-parthe-pandit/
URL:https://datascience.utah.edu/talks/2020-09-11-parthe-pandit/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://us02web.zoom.us/j/83503994251?pwd=Wm01bk40UWxva2pSV0dTbFZH
 bGtLQT09
CATEGORIES:algorithms & theory,deep learning,machine learning,optimization
END:VEVENT
BEGIN:VEVENT
UID:2020-09-18-vishnu-lokhande@datascience.utah.edu
DTSTAMP:20200918T175000Z
DTSTART;TZID=America/Denver:20200918T115000
DTEND;TZID=America/Denver:20200918T131000
SUMMARY:Optimization methods for imposing Fairness in Computer Vision Model
 s - Vishnu Lokhande
DESCRIPTION:Vishnu Lokhande (University of Wisconsin-Madison)\n\nIn this ta
 lk\, we will study a mechanism to impose fairness in computer vision model
 s concurrently while training the model and informed by standard fairness 
 measures. While existing fairness based approaches in vision have largely 
 relied on training adversarial modules together with the primary classific
 ation/regression task\, in an effort to remove the influence of the protec
 ted attribute or variable\, we will discuss how ideas based on well-known 
 optimization concepts can provide a simpler alternative. In our proposed s
 cheme\, imposing fairness just requires specifying the protected attribute
  and utilizing our optimization routine. We will discuss experiments\, tha
 t are interpretable\, demonstrating that several fairness measures from th
 e literature can be reliably imposed on standard vision tasks. We will als
 o discuss technical analysis on the convergence guarantees of the said opt
 imization routine.\n\nhttps://datascience.utah.edu/talks/2020-09-18-vishnu
 -lokhande/
URL:https://datascience.utah.edu/talks/2020-09-18-vishnu-lokhande/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://us02web.zoom.us/j/83503994251?pwd=Wm01bk40UWxva2pSV0dTbFZH
 bGtLQT09
CATEGORIES:computer vision,fairness & ethics,machine learning,optimization
END:VEVENT
BEGIN:VEVENT
UID:2020-09-25-matthew-h-samore-jeffrey-humpherys@datascience.utah.edu
DTSTAMP:20200925T175000Z
DTSTART;TZID=America/Denver:20200925T115000
DTEND;TZID=America/Denver:20200925T131000
SUMMARY:Learning to care - Matthew H. Samore\, Jeffrey Humpherys
DESCRIPTION:Matthew H. Samore (Utah Epidemiology)\, Jeffrey Humpherys (Utah
  Epidemiology)\n\nhttps://datascience.utah.edu/talks/2020-09-25-matthew-h-
 samore-jeffrey-humpherys/
URL:https://datascience.utah.edu/talks/2020-09-25-matthew-h-samore-jeffrey-
 humpherys/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://us02web.zoom.us/j/83503994251?pwd=Wm01bk40UWxva2pSV0dTbFZH
 bGtLQT09
END:VEVENT
BEGIN:VEVENT
UID:2020-10-02-ellen-riloff@datascience.utah.edu
DTSTAMP:20201002T175000Z
DTSTART;TZID=America/Denver:20201002T115000
DTEND;TZID=America/Denver:20201002T131000
SUMMARY:Identifying Affective Events and the Reasons for their Polarity - E
 llen Riloff
DESCRIPTION:Ellen Riloff (Utah Computer Science)\n\nRecognizing affective s
 tates is essential for narrative text\nunderstanding and for applications 
 such as conversational dialogue\,\nsummarization\, and sarcasm recognition
 . Many tools have been developed\nto recognize explicit expressions of sen
 timent\, but affective states\ncan also be inferred from events. This talk
  will focus on "affective\nevents"\, which are generally desirable or unde
 sirable experiences that\nimplicitly suggest an affective state for the ex
 periencer. For\nexample\, buying a home is usually desirable and associate
 d with a\npositive affective state\, but being laid off is undesirable and
 \nassociated with a negative state. First\, we will describe a weakly\nsup
 ervised learning method to induce affective events from a text\ncorpus by 
 optimizing for semantic consistency. Second\, we aim to\ncharacterize affe
 ctive events based on Human Needs Categories\, which\noften explain people
 's motivations\, goals\, and desires. We will\npresent a co-training model
  for Human Needs categorization that uses\nan event expression classifier 
 and an event context classifier to\nlearn from both labeled and unlabeled 
 texts.\n\nhttps://datascience.utah.edu/talks/2020-10-02-ellen-riloff/
URL:https://datascience.utah.edu/talks/2020-10-02-ellen-riloff/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://us02web.zoom.us/j/83503994251?pwd=Wm01bk40UWxva2pSV0dTbFZH
 bGtLQT09
CATEGORIES:machine learning,natural language processing
END:VEVENT
BEGIN:VEVENT
UID:2020-10-09-swaroop-mishra@datascience.utah.edu
DTSTAMP:20201009T175000Z
DTSTART;TZID=America/Denver:20201009T115000
DTEND;TZID=America/Denver:20201009T131000
SUMMARY:DQI: Measuring Data Quality in NLP - Swaroop Mishra
DESCRIPTION:Swaroop Mishra (Arizona State University)\n\nNeural language mo
 dels have achieved human-level performance across several NLP datasets. Ho
 wever\, recent studies have shown that these models are not truly learning
  the desired task\; rather\, their high performance is attributed to overf
 itting using spurious biases\, which suggests that the capabilities of AI 
 systems have been over-estimated. We introduce a generic formula for Data 
 Quality Index (DQI) to help dataset creators create datasets with minimal 
 unwanted biases. We propose a new data creation paradigm using DQI to crea
 te higher quality data. The data creation paradigm consists of several dat
 a visualizations to help data creators (i) understand the quality of data 
 and (ii) visualize the impact of the created data instance on the overall 
 quality. It also has a couple of automation methods to (i) assist data cre
 ators and (ii) make the model more robust to adversarial attacks. We use D
 QI along with these automation methods to renovate biased examples in SNLI
 . We show that models trained on the renovated SNLI dataset generalize bet
 ter to out of distribution tasks. Renovation results in reduced model perf
 ormance\, exposing a large gap with respect to human performance. DQI syst
 ematically helps in creating harder benchmarks using active learning. Our 
 work takes the process of dynamic dataset creation forward\, wherein datas
 ets evolve together with the evolving state of the art\, therefore serving
  as a means of benchmarking the true progress of AI. Finally\, we also sho
 w that DQI helps in pruning a dataset without compromising IID and OOD per
 formance significantly.\n\nhttps://datascience.utah.edu/talks/2020-10-09-s
 waroop-mishra/
URL:https://datascience.utah.edu/talks/2020-10-09-swaroop-mishra/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://us02web.zoom.us/j/83503994251?pwd=Wm01bk40UWxva2pSV0dTbFZH
 bGtLQT09
CATEGORIES:data management,fairness & ethics,natural language processing
END:VEVENT
BEGIN:VEVENT
UID:2020-10-16-varun-gangal@datascience.utah.edu
DTSTAMP:20201016T175000Z
DTSTART;TZID=America/Denver:20201016T115000
DTEND;TZID=America/Denver:20201016T131000
SUMMARY:Examining Extra Sentential Abilities of Contextual Embeddings - Var
 un Gangal
DESCRIPTION:Varun Gangal (LTI\, CMU)\n\nIn the first third of our talk\, we
  try to understand what and how much does BERT already know about event ar
 guments (including cross-sentence ones)?. We observe that BERT's attention
  heads have modest but well above-chance ability to spot event arguments s
 ans any training. Furthermore\, we investigate how our methods do for cros
 s-sentence event arguments\, proposing a procedure to isolate "best heads"
  for cross-sentence argument detection separately of those for intra-sente
 nce arguments. In the second third\, we take a closer look at the infillin
 g abilities of BERT. We know BERT is good at Word-level Infilling (obvious
 ly!) . How good is it though at Sentence-level Infilling a.k.a Cloze ? We 
 introduce a human-created sentence cloze dataset\, collected from public s
 chool English examinations. Our task requires a model to fill up multiple 
 blanks in a passage from a shared candidate set with distractors designed 
 by English teachers. Our experiments show a significant performance gap be
 tween BERT (72%) and humans (87%)\, encouraging future models to bridge th
 is gap.We conclude our talk by going through a somewhat unrelated\, recent
  foray into data augmentation for finetuning pretrained generators on low 
 resource domains.\n\nhttps://datascience.utah.edu/talks/2020-10-16-varun-g
 angal/
URL:https://datascience.utah.edu/talks/2020-10-16-varun-gangal/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://us02web.zoom.us/j/83503994251?pwd=Wm01bk40UWxva2pSV0dTbFZH
 bGtLQT09
CATEGORIES:deep learning
END:VEVENT
BEGIN:VEVENT
UID:2020-10-23-nancy-wang@datascience.utah.edu
DTSTAMP:20201023T175000Z
DTSTART;TZID=America/Denver:20201023T115000
DTEND;TZID=America/Denver:20201023T131000
SUMMARY:Global Table Extractor (GTE): A Framework for Joint Table Identific
 ation and Cell Structure Recognition Using Visual Context - Nancy Wang
DESCRIPTION:Nancy Wang (IBM)\n\nDocuments are often the format of choice fo
 r knowledge sharing and preservation in business and science\, within whic
 h are tables that capture most of the critical data. Unfortunately\, most 
 documents are stored and distributed as PDF or scanned images\, which fail
  to preserve table formatting.\nRecent vision-based deep learning approach
 es have been proposed to address this gap\, but most still cannot achieve 
 state-of-the-art results.\n\n We present Global Table Extractor (GTE)\, a 
 vision-guided systematic framework for joint table detection and cell stru
 ctured recognition\, which could be built on top of any object detection m
 odel. With GTE-Table\, we invent a new penalty based on the natural cell c
 ontainment constraint of tables to train our table network aided by cell l
 ocation predictions. GTE-Cell is a new hierarchical cell detection network
  that leverages table styles. Further\, we design a method to automaticall
 y label table and cell structure in existing documents to cheaply create a
  large corpus of training and test data. We use this to enhance PubTabNet 
 with cell labels and create FinTabNet\, real-world and complex scientific 
 and financial datasets with detailed table structure annotations to help t
 rain and test structure recognition.\n\n Our deep learning framework surpa
 sses previous state-of-the-art results on the ICDAR 2013 and ICDAR 2019 ta
 ble competition test dataset in both table detection and cell structure re
 cognition. Further experiments demonstrate a greater than 45% improvement 
 in cell structure recognition when compared to a vanilla RetinaNet object 
 detection model in our new financial dataset (FinTabNet).\n\nhttps://datas
 cience.utah.edu/talks/2020-10-23-nancy-wang/
URL:https://datascience.utah.edu/talks/2020-10-23-nancy-wang/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://us02web.zoom.us/j/83503994251?pwd=Wm01bk40UWxva2pSV0dTbFZH
 bGtLQT09
CATEGORIES:biology & genomics,computer vision,data management,deep learning
END:VEVENT
BEGIN:VEVENT
UID:2020-10-30-daniel-scharfstein@datascience.utah.edu
DTSTAMP:20201030T175000Z
DTSTART;TZID=America/Denver:20201030T115000
DTEND;TZID=America/Denver:20201030T131000
SUMMARY:Semiparametrics: A Biostatistician’s Toolbox - Daniel Scharfstein
DESCRIPTION:Daniel Scharfstein (Utah\, Population Health Sciences)\n\nIn th
 is talk\, I will discuss the theory of semiparametrics that I use to estim
 ate causal effects at root-n rates. Estimators of these effects depend on 
 estimators of nuisance parameters that can be estimated at rates slower th
 an root-n\; I provide sufficient conditions for these rates. I will seek a
 dvice on the machine learning estimation techniques that satisfy these con
 ditions. I will illustrate the theory in the context of estimating the cau
 sal contrast of two competing treatments based on data from a comprehensiv
 e cohort study in which clinically eligible individuals are ﬁrst asked t
 o enroll in a randomized trial and\, if they refuse\, are then asked to en
 roll in a parallel observational study in which they can choose treatment 
 according to their own preference.\n\nhttps://datascience.utah.edu/talks/2
 020-10-30-daniel-scharfstein/
URL:https://datascience.utah.edu/talks/2020-10-30-daniel-scharfstein/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://us02web.zoom.us/j/83503994251?pwd=Wm01bk40UWxva2pSV0dTbFZH
 bGtLQT09
CATEGORIES:causal inference,health & medicine,statistics
END:VEVENT
BEGIN:VEVENT
UID:2020-10-30-alberto-cairo@datascience.utah.edu
DTSTAMP:20201030T213000Z
DTSTART;TZID=America/Denver:20201030T153000
DTEND;TZID=America/Denver:20201030T163000
SUMMARY:Data Visualization: How to Make Good Decisions - Alberto Cairo
DESCRIPTION:Alberto Cairo\n\nData visualization\, the display of data throu
 gh graphs\, charts\, maps\, and diagrams\, is a skill in great demand in m
 any disciplines\, from the sciences to communication or business analytics
 . However\, visualization is often misunderstood. For instance\, it's ofte
 n taught as the application of a series of strict rules. This talk argues 
 that visualization is more akin to writing: yes\, we do need to understand
  visualization's grammar but\, beyond that\, visualization design is flexi
 ble\, and can't be based on rules that are set in stone. Instead\, designe
 rs need to develop a good decision-making framework based on asking themse
 lves a series of questions.\n\nhttps://datascience.utah.edu/talks/2020-10-
 30-alberto-cairo/
URL:https://datascience.utah.edu/talks/2020-10-30-alberto-cairo/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
CATEGORIES:human-centered computing,networks & graphs,visualization
END:VEVENT
BEGIN:VEVENT
UID:2020-11-06-bhargavi-paranjape@datascience.utah.edu
DTSTAMP:20201106T185000Z
DTSTART;TZID=America/Denver:20201106T115000
DTEND;TZID=America/Denver:20201106T131000
SUMMARY:An Information Bottleneck Approach for Rationale Extraction - Bharg
 avi Paranjape
DESCRIPTION:Bhargavi Paranjape (University of Washington)\n\nDecisions of c
 omplex models for language understanding can be explained by limiting the 
 inputs they are provided to a relevant sub-sequence of the original text 
 — a rationale. Models that condition predictions on a concise rationale\
 , while being more interpretable\, tend to be less accurate than models th
 at are able to use the entire context. In this paper\, we show that it is 
 possible to better manage the trade-off between concise explanations and h
 igh task accuracy by optimizing abound on the Information Bottleneck (IB) 
 objective. Our approach jointly learns an explainer that predicts sparse b
 inary masks over input sentences without explicit supervision and an end-t
 ask predictor that considers only the residual sentences. Using IB\, we de
 rive a learning objective that allows direct control of mask sparsity leve
 ls through a tunable sparse prior. Experiments on the ERASER benchmark dem
 onstrate significant gains over previous work for both task performance an
 d agreement with human rationales\n\nhttps://datascience.utah.edu/talks/20
 20-11-06-bhargavi-paranjape/
URL:https://datascience.utah.edu/talks/2020-11-06-bhargavi-paranjape/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://us02web.zoom.us/j/83503994251?pwd=Wm01bk40UWxva2pSV0dTbFZH
 bGtLQT09
CATEGORIES:education,fairness & ethics,natural language processing
END:VEVENT
BEGIN:VEVENT
UID:2020-11-13-akanksha-atrey@datascience.utah.edu
DTSTAMP:20201113T185000Z
DTSTART;TZID=America/Denver:20201113T115000
DTEND;TZID=America/Denver:20201113T131000
SUMMARY:Towards High-Performance Machine Learning on the Edge - Akanksha At
 rey
DESCRIPTION:Akanksha Atrey (UMass)\n\nModern day distributed technologies\,
  such as mobile systems and the Internet of Things (IoT)\, enable the glob
 al integration of heterogeneous smart devices via wireless networks. A com
 mon characteristic across these technologies is their ability to collect a
 nd communicate continuously streaming data. The generation of high bandwid
 th data makes machine learning (ML) and artificial intelligence (AI) appea
 ling for processing\, reasoning\, and predicting about the environment\, b
 ut low network latency requirements make offloading intelligence to the cl
 oud undesirable. This raises an important question: how can we design\, de
 velop and evaluate ML algorithms that perform beyond predictive accuracy (
 e.g.\, generalizable\, explainable\, privacy-aware\, and efficient) while 
 being accessible and scalable in resource-constrained edge environments? I
 n this talk\, I will cover two aspects\, explainability and privacy\, when
  deploying such ML models in modern distributed technologies. The focus of
  the talk will be two-fold: (1) counterfactual evaluation of the explanati
 ons generated using saliency maps in deep reinforcement learning for appli
 cations such as autonomous vehicles\, and (2) privacy implications of pers
 onalized ML models in context-aware mobility applications. The talk will b
 e concluded with a discussion on where the future of ML lies in evolving d
 istributed technologies.\n\nhttps://datascience.utah.edu/talks/2020-11-13-
 akanksha-atrey/
URL:https://datascience.utah.edu/talks/2020-11-13-akanksha-atrey/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://us02web.zoom.us/j/83503994251?pwd=Wm01bk40UWxva2pSV0dTbFZH
 bGtLQT09
CATEGORIES:algorithms & theory,machine learning,privacy & security
END:VEVENT
BEGIN:VEVENT
UID:2020-11-20-akhil-arora@datascience.utah.edu
DTSTAMP:20201120T185000Z
DTSTART;TZID=America/Denver:20201120T115000
DTEND;TZID=America/Denver:20201120T131000
SUMMARY:Low-rank Subspaces for Unsupervised Entity Linking - Akhil Arora
DESCRIPTION:Akhil Arora (EPFL)\n\nEntity linking is an important problem wi
 th many applications. Most previous solutions were designed for settings w
 here annotated training data is available\, which is\, however\, not the c
 ase in numerous domains. We propose a light-weight and scalable entity lin
 king method\, Eigenthemes\, that relies solely on the availability of enti
 ty names and a referent knowledge base. Eigenthemes exploits the fact that
  the entities that are truly mentioned in a document (the ``gold entities'
 ') tend to form a semantically dense subset of the set of all candidate en
 tities in the document. Geometrically speaking\, when representing entitie
 s as vectors via some given embedding\, the gold entities tend to lie in a
  low-rank subspace of the full embedding space. Eigenthemes identifies thi
 s subspace using the singular value decomposition and scores candidate ent
 ities according to their proximity to the subspace. Extensive experiments 
 on benchmark datasets from a variety of real-world domains showcase the ef
 fectiveness of our approach.\n\nhttps://datascience.utah.edu/talks/2020-11
 -20-akhil-arora/
URL:https://datascience.utah.edu/talks/2020-11-20-akhil-arora/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://us02web.zoom.us/j/83503994251?pwd=Wm01bk40UWxva2pSV0dTbFZH
 bGtLQT09
CATEGORIES:deep learning,machine learning
END:VEVENT
BEGIN:VEVENT
UID:2020-12-04-danish-pruthi@datascience.utah.edu
DTSTAMP:20201204T185000Z
DTSTART;TZID=America/Denver:20201204T115000
DTEND;TZID=America/Denver:20201204T131000
SUMMARY:A Tale of Evidence and Explanations - Danish Pruthi
DESCRIPTION:Danish Pruthi (LTI\, CMU)\n\nI would present a brief overview o
 f the state of research in explainability and its evaluation (or lack ther
 eof). Then\, I would offer a new lens into explanations\, viewing them as 
 a communication channel between a teacher and a student. This view enables
  us to quantitatively evaluate different attribution methods in a principl
 ed way at scale. Shifting gears\, in the second part of the talk\, I would
  introduce new techniques to supplement predictions with evidence to enabl
 e stakeholders to verify the outcomes readily.\n\nhttps://datascience.utah
 .edu/talks/2020-12-04-danish-pruthi/
URL:https://datascience.utah.edu/talks/2020-12-04-danish-pruthi/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://us02web.zoom.us/j/83503994251?pwd=Wm01bk40UWxva2pSV0dTbFZH
 bGtLQT09
CATEGORIES:education
END:VEVENT
BEGIN:VEVENT
UID:2021-01-22-grad-student-spotlights@datascience.utah.edu
DTSTAMP:20210122T210000Z
DTSTART;TZID=America/Denver:20210122T140000
DTEND;TZID=America/Denver:20210122T150000
SUMMARY:1. Introduction and Logistics - Grad Student Spotlights
DESCRIPTION:Grad Student Spotlights (10-minute talks by 4 current graduate 
 students)\n\nArchit Rathore: Exploring the Shape of Activations - A BERT c
 ase studyBenwei Shi: At-the-time and Back-in-time Persistent Sketches\nJoe
  Vinu: What all it takes for Performant Deep Nets to be Reliable and Fair?
 \nBrian Lavallee: Rounding Out Structural Rounding\nVivek Gupta: Inference
  on Tables as Semi-Structured Data\n\nhttps://datascience.utah.edu/talks/2
 021-01-22-grad-student-spotlights/
URL:https://datascience.utah.edu/talks/2021-01-22-grad-student-spotlights/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://us02web.zoom.us/j/87538638627?pwd=WGZrYmIyNjFsVEtwWk5pemRu
 M0JsZz09
CATEGORIES:data management,fairness & ethics
END:VEVENT
BEGIN:VEVENT
UID:2021-01-29-sanghamitra-dutta@datascience.utah.edu
DTSTAMP:20210129T210000Z
DTSTART;TZID=America/Denver:20210129T140000
DTEND;TZID=America/Denver:20210129T150000
SUMMARY:A Systematic Understanding of Exempt and Non-Exempt Algorithmic Bia
 ses - Sanghamitra Dutta
DESCRIPTION:Sanghamitra Dutta (CMU)\n\nWith the growing use of machine lear
 ning algorithms in highly consequential domains\, the quantification and r
 emoval of bias with respect to gender\, race\, etc.\, is becoming increasi
 ngly important. While quantifying bias is essential\, sometimes the needs 
 of a business (e.g.\, hiring) may require the use of certain features that
  are critical in a way that any bias that can be explained by them might n
 eed to be exempted (inspired from the business necessity defense of Title 
 VII of Civil Rights Act). For instance\, in hiring a software engineer\, a
  standardized coding-test score may be a critical feature that is weighed 
 strongly in the decision even if it introduces bias\, whereas other featur
 es\, such as name\, zip code\, or reference letters may be used to improve
  decision-making\, but only to the extent that they do not introduce bias.
  In this work\, we propose a novel information-theoretic measure of non-ex
 empt bias\, which quantifies the part of the bias that cannot be accounted
  for by the critical features. This measure can be applied for (i) Auditin
 g trained models to check if the bias arose purely due to the critical fea
 tures\; and also for (ii) Training with selective removal of the non-exemp
 t bias if desired. We arrive at this decomposition through canonical examp
 les that lead to a set of desirable properties (axioms) that any measure o
 f non-exempt bias should satisfy. We then propose a causal measure of non-
 exempt bias that satisfies all of them. We also propose observational meas
 ures that only satisfy some of these properties (including an impossibilit
 y result on observational measures being able to satisfy all properties). 
 Then\, we perform case studies using them to show how one can train models
  while reducing non-exempt bias. Our quantification bridges ideas of causa
 lity\, Simpson's paradox\, and a body of work from information theory call
 ed Partial Information Decomposition (PID). The talk will be fairly access
 ible\, no knowledge of information theory is required.\n\nhttps://datascie
 nce.utah.edu/talks/2021-01-29-sanghamitra-dutta/
URL:https://datascience.utah.edu/talks/2021-01-29-sanghamitra-dutta/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://us02web.zoom.us/j/87538638627?pwd=WGZrYmIyNjFsVEtwWk5pemRu
 M0JsZz09
CATEGORIES:algorithms & theory,causal inference,fairness & ethics,machine l
 earning
END:VEVENT
BEGIN:VEVENT
UID:2021-02-05-michal-moshkovitz@datascience.utah.edu
DTSTAMP:20210205T210000Z
DTSTART;TZID=America/Denver:20210205T140000
DTEND;TZID=America/Denver:20210205T150000
SUMMARY:Unexpected Effects of Online no-Substitution k-means Clustering - M
 ichal Moshkovitz
DESCRIPTION:Michal Moshkovitz (UCSD)\n\nOffline k-means clustering was stud
 ied extensively and algorithms with a constant approximation are available
 . However\, online clustering is still uncharted. New factors come into pl
 ay: the ordering of the dataset and whether the number of points\, n\, is 
 known in advance or not. Their exact effects are unknown. In this work\, w
 e focus on the online setting where the decisions are irreversible: after 
 a point arrives the algorithm needs to decide whether to take the point as
  a center or not\, and this decision is final. How many centers are needed
  and sufficient to achieve constant approximation in this setting? We show
  upper and lower bounds for all the different cases. These bounds are exac
 tly the same up to a constant\, thus achieving optimal bounds. For example
 \, for k-means cost with constant k>1 and random order\, Θ(logn) centers 
 are enough to achieve a constant approximation\, while the mere a priori k
 nowledge of n reduces the number of centers to a constant. These bounds ho
 ld for any distance function that obeys a triangle-type inequality.\n\nhtt
 ps://datascience.utah.edu/talks/2021-02-05-michal-moshkovitz/
URL:https://datascience.utah.edu/talks/2021-02-05-michal-moshkovitz/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://us02web.zoom.us/j/87538638627?pwd=WGZrYmIyNjFsVEtwWk5pemRu
 M0JsZz09
CATEGORIES:algorithms & theory
END:VEVENT
BEGIN:VEVENT
UID:2021-02-12-fritz-lekschas@datascience.utah.edu
DTSTAMP:20210212T210000Z
DTSTART;TZID=America/Denver:20210212T140000
DTEND;TZID=America/Denver:20210212T150000
SUMMARY:Visual Pattern Exploration At and Across Scales - Fritz Lekschas
DESCRIPTION:Fritz Lekschas (Harvard)\n\nVisually exploring data is a powerf
 ul approach to discover\,\nunderstand\, and interpret novel or not-well de
 fined patterns. It\nallows us to gain insights and generate hypotheses for
  subsequent\nanalyses. However\, visual exploration can become challenging
  when the\npatterns of interest are sparsely-distributed\, several orders 
 of\nmagnitude smaller than the entire dataset\, or detected with high\nunc
 ertainty. In this talk\, I will discuss challenges in visually\nexploring 
 multi-modal and multi-scale data\, and present new\nvisualization systems 
 for efficiently browsing\, comparing\, and finding\npatterns in the contex
 t of genomic\, geospatial\, and time-series data.\nSpecifically\, I will d
 escribe a web platform for browsing multi-modal\nand multi-scale datasets\
 , as well as their guided navigation. I will\npresent a generalized framew
 ork and toolkit for interactively\narranging\, grouping\, and aggregating 
 thousands of pattern instances.\nAnd I will demonstrate how interactive vi
 sual machine learning can\nenhance our ability to find patterns effectivel
 y.\n\nhttps://datascience.utah.edu/talks/2021-02-12-fritz-lekschas/
URL:https://datascience.utah.edu/talks/2021-02-12-fritz-lekschas/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://us02web.zoom.us/j/87538638627?pwd=WGZrYmIyNjFsVEtwWk5pemRu
 M0JsZz09
CATEGORIES:biology & genomics,geospatial,machine learning,visualization
END:VEVENT
BEGIN:VEVENT
UID:2021-02-19-samson-zhou@datascience.utah.edu
DTSTAMP:20210219T210000Z
DTSTART;TZID=America/Denver:20210219T140000
DTEND;TZID=America/Denver:20210219T150000
SUMMARY:Tight Bounds for Adversarially Robust Streams and Sliding Windows v
 ia Difference Estimators - Samson Zhou
DESCRIPTION:Samson Zhou (CMU)\n\nWe introduce difference estimators for dat
 a stream computation\, which provide approximations to F(v)-F(u) for frequ
 ency vectors v\,u and a given function F. We show how to use such estimato
 rs to carefully trade error for memory in an iterative manner. The functio
 n F is generally non-linear\, and we give the first difference estimators 
 for the frequency moments F_p for p between 0 and 2\, as well as for integ
 ers p>2. Using these\, we resolve a number of central open questions in ad
 versarial robust streaming and sliding window models.\n\nFor both models\,
  we obtain algorithms for norm estimation whose dependence on epsilon is 1
 /epsilon^2\, which shows\, up to logarithmic factors\, that there is no ov
 erhead over the standard insertion-only data stream model for these proble
 ms.\n\nhttps://datascience.utah.edu/talks/2021-02-19-samson-zhou/
URL:https://datascience.utah.edu/talks/2021-02-19-samson-zhou/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
CATEGORIES:algorithms & theory,statistics
END:VEVENT
BEGIN:VEVENT
UID:2021-02-26-rajesh-jayaram@datascience.utah.edu
DTSTAMP:20210226T210000Z
DTSTART;TZID=America/Denver:20210226T140000
DTEND;TZID=America/Denver:20210226T150000
SUMMARY:An Improved Analysis of the Quadtree for High Dimensional EMD - Raj
 esh Jayaram
DESCRIPTION:Rajesh Jayaram (CMU)\n\nThe Earth Mover Distance (EMD) between 
 two multi-sets A\,B in R^d of size s is the min-cost of bipartite matching
 s between points in A and B\, where cost is measured by distance between p
 oints. In this talk\, we discuss a classic divide-and-conquer algorithm kn
 own as Quadtree for approximating EMD.\nWe give a new analysis of the Quad
 tree\, showing that it gives a Õ(log s) approximation. This improves on t
 he previous known O(min{log s \, log d} * log s)-approximation of Andoni\,
  Indyk\, and Krauthgamer [SODA 08]\, and Backurs\, Dong\, Indyk\, Razensht
 eyn\, and Wagner (ICML 20).\n\nWe also give new space efficient sketching 
 and streaming algorithms for estimating EMD with the improved approximatio
 n factor. The main conceptual contribution is an analytical framework for 
 studying the Quadtree which goes beyond worst-case distortion of randomize
 d tree embeddings.\n\nBased on a joint work with Xi Chen\, Amit Levi\, and
  Erik Waingarten.\n\nPaper: https://rajeshjayaram.com/EarthMoverCJLW.pdf (
 https://rajeshjayaram.com/EarthMoverCJLW.pdf)\n\nSlides: https://rajeshjay
 aram.com/EarthMoverCJLW.pdf\n\nhttps://datascience.utah.edu/talks/2021-02-
 26-rajesh-jayaram/
URL:https://datascience.utah.edu/talks/2021-02-26-rajesh-jayaram/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://us02web.zoom.us/j/87538638627?pwd=WGZrYmIyNjFsVEtwWk5pemRu
 M0JsZz09
CATEGORIES:algorithms & theory,data management,machine learning
END:VEVENT
BEGIN:VEVENT
UID:2021-03-12-yaoqing-yang@datascience.utah.edu
DTSTAMP:20210312T210000Z
DTSTART;TZID=America/Denver:20210312T140000
DTEND;TZID=America/Denver:20210312T150000
SUMMARY:Boundary thickness and robustness in learning models - Yaoqing Yang
DESCRIPTION:Yaoqing Yang (UC Berkeley)\n\nRobustness of machine learning mo
 dels to various adversarial and non-adversarial corruptions continues to b
 e of interest. In this talk\, we present the notion of the "boundary thick
 ness" of a classifier\, and we describe its connection with and usefulness
  for model robustness. Thick decision boundaries lead to improved performa
 nce\, while thin decision boundaries lead to overfitting (e.g.\, measured 
 by the robust generalization gap between training and testing) and lower r
 obustness. We show that a thicker boundary helps improve robustness agains
 t adversarial examples (e.g.\, improving the robust test accuracy of adver
 sarial training) as well as so-called out-of-distribution (OOD) transforms
 \, and we show that many commonly-used regularization and data augmentatio
 n procedures can increase boundary thickness. On the theoretical side\, we
  establish that maximizing boundary thickness during training is akin to t
 he so-called mixup training procedure. Using these observations\, we show 
 that noise-augmentation on mixup training further increases boundary thick
 ness\, thereby combating vulnerability to various forms of adversarial att
 acks and OOD transforms. We can also show that the performance improvement
  in several lines of recent work happens in conjunction with a thicker bou
 ndary.\n\nhttps://datascience.utah.edu/talks/2021-03-12-yaoqing-yang/
URL:https://datascience.utah.edu/talks/2021-03-12-yaoqing-yang/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://us02web.zoom.us/j/87538638627?pwd=WGZrYmIyNjFsVEtwWk5pemRu
 M0JsZz09
CATEGORIES:machine learning,privacy & security
END:VEVENT
BEGIN:VEVENT
UID:2021-03-19-vivek-gupta@datascience.utah.edu
DTSTAMP:20210319T200000Z
DTSTART;TZID=America/Denver:20210319T140000
DTEND;TZID=America/Denver:20210319T150000
SUMMARY:Logic based classification for Low Resource Setting - Vivek Gupta
DESCRIPTION:Vivek Gupta (University of Utah)\n\nAn NLP model’s ability to
  reason should be independent of language. Previous works utilize Natural 
 Language Inference(NLI) to understand the reasoning ability of models\, mo
 stly focusing on high resource languages like English. To address scarcity
  of data in low-resource languages such as Hindi\, we use data recasting t
 o create four NLI datasets from existing four text classification datasets
  in Hindi language. Through experiments\, we show that our recasted datase
 t is devoid of statistical irregularities and spurious patterns. We study 
 the consistency in predictions of the textual entailment models and propos
 e a consistency regulariser to remove pairwise-inconsistencies in predicti
 ons. Furthermore\, we propose a novel two-step classification method which
  uses textual-entailment predictions for classification tasks. We further 
 improve the classification performance by jointly training the classificat
 ion and textual entailment tasks together. We therefore highlight the bene
 fits of data recasting and our approach with supporting experimental resul
 ts. You can access the dataset and paper here: https://www.aclweb.org/anth
 ology/2020.aacl-main.71.pdf (https://www.aclweb.org/anthology/2020.aacl-ma
 in.71.pdf). Joint work with BloomBerg AI.\n\nSlides: https://www.aclweb.or
 g/anthology/2020.aacl-main.71.pdf\n\nhttps://datascience.utah.edu/talks/20
 21-03-19-vivek-gupta/
URL:https://datascience.utah.edu/talks/2021-03-19-vivek-gupta/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://us02web.zoom.us/j/87538638627?pwd=WGZrYmIyNjFsVEtwWk5pemRu
 M0JsZz09
CATEGORIES:machine learning,natural language processing
END:VEVENT
BEGIN:VEVENT
UID:2021-04-09-emily-beth-wall@datascience.utah.edu
DTSTAMP:20210409T200000Z
DTSTART;TZID=America/Denver:20210409T140000
DTEND;TZID=America/Denver:20210409T150000
SUMMARY:As We Are: Detecting and Mitigating Human Bias in Visual Analytics 
 - Emily Beth Wall
DESCRIPTION:Emily Beth Wall (Emory)\n\nVisual Analytics combines the comple
 mentary strengths of humans (perception and sensemaking capabilities) and 
 machines (fast and accurate information processing). However\, people are 
 susceptible to inherent limitations and biases\, including cognitive biase
 s (e.g.\, anchoring bias)\, social biases borne of cultural stereotypes an
 d prejudices (e.g.\, gender bias)\, and perceptual biases (e.g.\, illusion
 s). These biases can impact data analysis and decision making in critical 
 ways\, leading to inaccurate or inefficient choices\, or even propagating 
 long-standing institutional and systemic biases.\n\nGiven our knowledge of
  these biases and the increased use of data visualization to support decis
 ion making in data science\, the goal of this research is to detect and mi
 tigate human biases in visual data analysis. In this talk\, I describe (1)
  which types of bias are particularly relevant in the process of visual da
 ta analysis\, (2) how user interactions with data can be used to approxima
 te human biases\, and (3) how visualization systems can be designed to inc
 rease user awareness of potentially unconscious or implicit biases. By cre
 ating systems that promote real-time awareness of bias\, people can reflec
 t on their behavior and decision making and ultimately engage in a less-bi
 ased analysis and decision making process.\n\nhttps://datascience.utah.edu
 /talks/2021-04-09-emily-beth-wall/
URL:https://datascience.utah.edu/talks/2021-04-09-emily-beth-wall/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:Virtual | https://us02web.zoom.us/j/87538638627?pwd=WGZrYmIyNjFsVE
 twWk5pemRuM0JsZz09
CATEGORIES:fairness & ethics,visualization
END:VEVENT
BEGIN:VEVENT
UID:2021-04-16-arun-sai-suggala@datascience.utah.edu
DTSTAMP:20210416T200000Z
DTSTART;TZID=America/Denver:20210416T140000
DTEND;TZID=America/Denver:20210416T150000
SUMMARY:Game Theoretic Statistics - Arun Sai Suggala
DESCRIPTION:Arun Sai Suggala (CMU)\n\nGame theory and statistics are often 
 regarded as disparate research areas. This is because typical statistical 
 estimation settings are non-adversarial\, and the samples are assumed to b
 e generated by some stationary non-reactive source. However\, there is a g
 reat degree of commonality between the two fields. Classically\, the mathe
 matical philosophy of statistics\, particularly frequentist statistics\, p
 osits that the source of samples is potentially adversarial. This resulted
  in the rich theory of minimax statistical games and estimation. Boosting 
 algorithms\, which are often regarded as best off-the-shelf classifiers\, 
 can be viewed as playing a zero-sum game against a weak learner. To allow 
 for various departures of ``test environment'' from ``train environments''
 \, the emerging field of robust machine learning allows for adversarial ma
 nipulation of the train or test environments. Finally\, an emerging class 
 of density estimators (GANs) in modern machine learning use an adversarial
  ``critic'' of the density estimator to improve the final density estimati
 on. The common theme among these classical and modern developments is an i
 nterplay between statistical estimation and two player games.\n\nIn this t
 alk\, I will present some of my recent work at the intersection of statist
 ics and game theory and show how game theory can help statistics. In parti
 cular\, I will present my work on minimax statistical estimation\, where w
 e develop algorithmic techniques for constructing minimax estimators. Our 
 algorithms rely on tools from online nonconvex learning and help us constr
 uct minimax estimators for fundamental estimation problems such as covaria
 nce estimation and entropy estimation.\n\nhttps://datascience.utah.edu/tal
 ks/2021-04-16-arun-sai-suggala/
URL:https://datascience.utah.edu/talks/2021-04-16-arun-sai-suggala/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:Virtual | https://us02web.zoom.us/j/87538638627?pwd=WGZrYmIyNjFsVE
 twWk5pemRuM0JsZz09
CATEGORIES:algorithms & theory,machine learning,optimization,statistics
END:VEVENT
BEGIN:VEVENT
UID:2021-04-23-lizzie-kumar@datascience.utah.edu
DTSTAMP:20210423T200000Z
DTSTART;TZID=America/Denver:20210423T140000
DTEND;TZID=America/Denver:20210423T150000
SUMMARY:Epistemic values in feature importance methods: Lessons from femini
 st epistemology - Lizzie Kumar
DESCRIPTION:Lizzie Kumar (Utah)\n\nAs the public seeks greater accountabili
 ty and transparency from machine learning algorithms\, the research litera
 ture on methods to explain algorithms and their outputs has rapidly expand
 ed. Feature importance\, or the practice of assigning quantitative importa
 nce values to the input features of a machine learning model\, form a popu
 lar class of such methods. Much of the research on feature importance rest
 s on formalizations that attempt to capture universally desirable properti
 es. We investigate the ways in which epistemic values are implicitly embed
 ded in these methods and analyze the ways in which they conflict with idea
 s from feminist philosophy. We offer some suggestions on how to conduct re
 search on explanations that respects feminist epistemic values\, taking in
 to account the importance of social context\, the epistemic privileges of 
 subjugated knowers\, and adopting more interactional ways of knowing.\n\nh
 ttps://datascience.utah.edu/talks/2021-04-23-lizzie-kumar/
URL:https://datascience.utah.edu/talks/2021-04-23-lizzie-kumar/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:Virtual | https://us02web.zoom.us/j/87538638627?pwd=WGZrYmIyNjFsVE
 twWk5pemRuM0JsZz09
CATEGORIES:algorithms & theory,fairness & ethics,machine learning
END:VEVENT
BEGIN:VEVENT
UID:2021-08-27-jeff-phillips@datascience.utah.edu
DTSTAMP:20210827T200000Z
DTSTART;TZID=America/Denver:20210827T140000
DTEND;TZID=America/Denver:20210827T150000
SUMMARY:A Visual tour of Bias Mitigation - Jeff Phillips
DESCRIPTION:Jeff Phillips (University of Utah)\n\nWord vector embeddings ha
 ve been shown to contain and amplify biases in data they are extracted fro
 m. Consequently\, many techniques have been proposed to identify\, mitigat
 e\, and attenuate these biases in word representations. In this talk\, I w
 ill review a collection of state-of-the-art debiasing techniques. To aid t
 his\, we provide an open source web-based visualization tool VERB (Visuali
 zation of Embedding Representations for deBiasing) and offer hands-on expe
 rience in exploring the effects of these debiasing techniques on the geome
 try of high-dimensional word vectors. To help understand how various debia
 sing techniques change the underlying geometry\, I will show how to decomp
 ose each technique into interpretable sequences of primitive operations an
 d study their effect on the word vectors using dimensionality reduction an
 d interactive visual exploration.\n\nhttps://datascience.utah.edu/talks/20
 21-08-27-jeff-phillips/
URL:https://datascience.utah.edu/talks/2021-08-27-jeff-phillips/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 | https://utah.zoom.us/j/93778940103?pwd=TStQRWhWVjRxd0hG
 V1hTK05SUFZwUT09
CATEGORIES:deep learning,fairness & ethics,natural language processing,visu
 alization
END:VEVENT
BEGIN:VEVENT
UID:2021-09-03-shandian-zhe@datascience.utah.edu
DTSTAMP:20210903T200000Z
DTSTART;TZID=America/Denver:20210903T140000
DTEND;TZID=America/Denver:20210903T150000
SUMMARY:Multi-fidelity Learning and Optimization for Physical Simulation an
 d AutoML - Shandian Zhe
DESCRIPTION:Shandian Zhe (Utah SoC)\n\nMulti-fidelity learning involves usi
 ng training examples at different fidelities or resolutions. High-fidelity
  examples are of high-quality but often are much more costly to collect th
 an inaccurate\, low-fidelity examples. How to retrieve and leverage exampl
 es at multiple fidelities is the key to reduce the learning cost while max
 imizing efficiency.\n\nThis talk will introduce our recent work in multi-f
 idelity learning and optimization. First\, I will introduce our deep auto-
 regressive models that can capture complex correlations across the fidelit
 ies to integrate examples of high-dimensional outputs. These are common in
  applications of physical simulation. Second\, I will introduce our work o
 f deep multi-fidelity active learning and Bayesian optimization that can i
 mprove the learning and optimization efficiency while reducing the cost of
  generating training examples\, namely maximizing the benefit-cost ratio. 
 Finally\, a batch version of the active learning and optimization techniqu
 e will be presented\, which can reduce the query redundancy\, improve dive
 rsity\, and further boost the benefit-cost ratio. I will showcase the adva
 ntage of our methods in standard benchmarks of physical simulation\, topol
 ogy structure optimization\, and typical tasks in hyper-parameter tuning/A
 utoML.\n\nhttps://datascience.utah.edu/talks/2021-09-03-shandian-zhe/
URL:https://datascience.utah.edu/talks/2021-09-03-shandian-zhe/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147
CATEGORIES:optimization
END:VEVENT
BEGIN:VEVENT
UID:2021-09-10-pierre-lermusiaux@datascience.utah.edu
DTSTAMP:20210910T200000Z
DTSTART;TZID=America/Denver:20210910T140000
DTEND;TZID=America/Denver:20210910T150000
SUMMARY:Neural Closure Models for Dynamical Systems - Pierre Lermusiaux
DESCRIPTION:Pierre Lermusiaux\n\nComplex dynamical systems are used for pre
 dictions in many domains. Because of computational costs\, models are trun
 cated\, coarsened or aggregated. As the neglected and unresolved terms bec
 ome important\, the utility of model predictions diminishes. We develop a 
 novel\, versatile and rigorous methodology to learn non-Markovian closure 
 parametrizations for known-physics/low-fidelity models using data from hig
 h-fidelity simulations. The new neural closure models augment low-fidelity
  models with neural delay differential equations (nDDEs)\, motivated by th
 e Mori–Zwanzig formulation and the inherent delays in complex dynamical 
 systems. We demonstrate that neural closures efficiently account for trunc
 ated modes in reduced-order models\, capture the effects of subgrid-scale 
 processes in coarse models and augment the simplification of complex biolo
 gical and physical-biogeochemical models. We find that using non-Markovian
  over Markovian closures improves long-term prediction accuracy and requir
 es smaller networks. We derive adjoint equations and network architectures
  needed to efficiently implement the new discrete and distributed nDDEs\, 
 for any time-integration schemes and allowing non-uniformly spaced tempora
 l training data. The performance of discrete over-distributed delays in cl
 osure models is explained using information theory\, and we find an optima
 l amount of past information for a specified architecture. Finally\, we an
 alyze computational complexity and explain the limited additional cost due
  to neural closure models.\n\nPaper: https://royalsocietypublishing.org/do
 i/10.1098/rspa.2020.1004 (https://royalsocietypublishing.org/doi/10.1098/r
 spa.2020.1004?fbclid=IwAR1DZy5lHuf68lYwkk5Z_kUYzbHXdgRi0zUhoy5bohGy-w0BjZu
 wjqPhytQ)\n\nhttps://datascience.utah.edu/talks/2021-09-10-pierre-lermusia
 ux/
URL:https://datascience.utah.edu/talks/2021-09-10-pierre-lermusiaux/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://utah.zoom.us/j/93778940103?pwd=TStQRWhWVjRxd0hGV1hTK05SUFZ
 wUT09
CATEGORIES:algorithms & theory,machine learning,statistics
END:VEVENT
BEGIN:VEVENT
UID:2021-09-17-ross-whitaker@datascience.utah.edu
DTSTAMP:20210917T200000Z
DTSTART;TZID=America/Denver:20210917T140000
DTEND;TZID=America/Denver:20210917T150000
SUMMARY:Air Quality Mapping Using Sensor Networks and Statistical Regressio
 n - Ross Whitaker
DESCRIPTION:Ross Whitaker (Utah SoC & SCI)\n\nThis talk describes an air qu
 ality mapping system that is currently deployed in the Salt Lake Valley. W
 e begin with the motivations for the system and then briefly describe the 
 cyberphysical infrastructure. We then review the Gaussian process modeling
  approach we are using and discuss several important practical considerati
 ons around challenges of data wrangling and numerical implementations. We 
 present some examples of AQ estimates associated with specific events of b
 ad air quality. Finally we talk about current directions in research assoc
 iated with this approach.\n\nCOI Disclaimer: Ross Whitaker has a financial
  interest in the company Tetrad\, which has business interests related to 
 the technologies in this talk.\n\nMEB 3147\n\nhttps://datascience.utah.edu
 /talks/2021-09-17-ross-whitaker/
URL:https://datascience.utah.edu/talks/2021-09-17-ross-whitaker/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147
CATEGORIES:climate & environment,geospatial,statistics
END:VEVENT
BEGIN:VEVENT
UID:2021-09-24-c-seshadhri@datascience.utah.edu
DTSTAMP:20210924T200000Z
DTSTART;TZID=America/Denver:20210924T140000
DTEND;TZID=America/Denver:20210924T150000
SUMMARY:Studying the (in)effectiveness of low dimensional graph embeddings 
 - C. Seshadhri
DESCRIPTION:C. Seshadhri (UC Santa Cruz)\n\nLow dimensional graph embedding
 s are a fundamental and popular tool used for machine learning on graphs. 
 Given a graph\, the basic idea is to produce a low-dimensional vector for 
 each vertex\, such that "similarity" in geometric space corresponds to "pr
 oximity" in the graph. These vectors can then be used as features in a ple
 thora of machine learning tasks\, such as link prediction\, community labe
 ling\, recommendations\, etc. Despite many results emerging in this area o
 ver the past few years\, there is less study on the core premise of these 
 embeddings. Can such low-dimensional embeddings effectively capture the st
 ructure of real-world (such as social) networks? Contrary to common wisdom
 \, we mathematically prove and empirically demonstrate that popular low-di
 mensional graph embeddings do not capture salient properties of real-world
  networks. We mathematically prove that common low-dimensional embeddings 
 cannot generate graphs with both low average degree and large clustering c
 oefficients\, which have been widely established to be empirically true fo
 r real-world networks. Empirically\, we observe that the embeddings genera
 ted by popular methods fail to recreate the triangle structure of real-wor
 ld networks\, and do not perform well on certain community labeling tasks.
 \n\nJoint work with Ashish Goel\, Caleb Levy\, Aneesh Sharma\, and Andrew 
 Stolman\n\nhttps://datascience.utah.edu/talks/2021-09-24-c-seshadhri/
URL:https://datascience.utah.edu/talks/2021-09-24-c-seshadhri/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://utah.zoom.us/j/93778940103?pwd=TStQRWhWVjRxd0hGV1hTK05SUFZ
 wUT09
CATEGORIES:algorithms & theory,deep learning,machine learning,networks & gr
 aphs
END:VEVENT
BEGIN:VEVENT
UID:2021-10-01-tony-h-grubesic@datascience.utah.edu
DTSTAMP:20211001T200000Z
DTSTART;TZID=America/Denver:20211001T140000
DTEND;TZID=America/Denver:20211001T150000
SUMMARY:Estimating Potential Oil Spill Trajectories and Coastal Impacts fro
 m Near-Shore Storage Facilities: A Case Study of FSO Nabarima and the Gulf
  of Paria - Tony H. Grubesic
DESCRIPTION:Tony H. Grubesic (UT Austin)\n\nThe FSO Nabarima is a floating 
 storage facility and offloading vessel in the Gulf of Paria\, between Vene
 zuela and the island of Trinidad. During the latter half of 2020\, the Nab
 arima was disabled\, holding approximately 1.3 million barrels (55 million
  gallons) of crude oil on board. In October of 2020\, the vessel was tilti
 ng and potentially at risk of spilling its payload into open water. Althou
 gh all of the oil on the Nabarima was successfully offloaded by April 2021
 \, the threat of large crude oil releases is ubiquitous and persistent in 
 many coastal regions\, threatening local ecosystems and livelihoods in coa
 stal communities. The purpose of this presentation is to highlight a geoco
 mputational framework for evaluating the potential spatial vulnerability o
 f coastlines should oil be released from near-shore storage facilities. We
  use the Nabarima as a broadly representative case study\, discuss potenti
 al spill cleanup and mitigation strategies\, and highlight the challenges 
 of coordinating cross-national responses to these types of spill scenarios
 .\n\nhttps://datascience.utah.edu/talks/2021-10-01-tony-h-grubesic/
URL:https://datascience.utah.edu/talks/2021-10-01-tony-h-grubesic/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://utah.zoom.us/j/93778940103?pwd=TStQRWhWVjRxd0hGV1hTK05SUFZ
 wUT09
CATEGORIES:society & policy
END:VEVENT
BEGIN:VEVENT
UID:2021-10-08-anna-little@datascience.utah.edu
DTSTAMP:20211008T200000Z
DTSTART;TZID=America/Denver:20211008T140000
DTEND;TZID=America/Denver:20211008T150000
SUMMARY:The Mathematics of the Signal-to-Noise Ratio and Insights for Data 
 Science - Anna Little
DESCRIPTION:Anna Little (Utah Math)\n\nDespite the huge variety of data typ
 es and goals in data science\, there is usually an underlying signal-to-no
 ise ratio that governs the difficulty of the data science task. Analysis o
 f this signal-to-noise ratio leads to important insights about sample size
  requirements and the impact of the data dimension\, which are consistent 
 across various data models and tasks. This talk will illustrate these univ
 ersal insights in three specific contexts: (1) density-based clustering vi
 a graph embeddings\, (2) clustering mixture models via multidimensional sc
 aling\, and (3) signal recovery from noisy data. The underlying data model
 s are motivated by applications such as imaging\, particle physics\, singl
 e-cell RNA sequencing\, and cryo-electron microscopy.\n\nhttps://datascien
 ce.utah.edu/talks/2021-10-08-anna-little/
URL:https://datascience.utah.edu/talks/2021-10-08-anna-little/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 | https://utah.zoom.us/j/93778940103?pwd=TStQRWhWVjRxd0hG
 V1hTK05SUFZwUT09
CATEGORIES:algorithms & theory,biology & genomics
END:VEVENT
BEGIN:VEVENT
UID:2021-10-22-erin-wolf-chambers@datascience.utah.edu
DTSTAMP:20211022T200000Z
DTSTART;TZID=America/Denver:20211022T140000
DTEND;TZID=America/Denver:20211022T150000
SUMMARY:Applications of topology and geometry to root analysis - Erin Wolf 
 Chambers
DESCRIPTION:Erin Wolf Chambers (St. Loius University)\n\nAnalysis of 3d sha
 pes is a core problem in many fields\, and there are\nmany tools from topo
 logy and geometry that can provide insight and\nunderstanding. In this tal
 k\, we focus on developing significance\nmeasures for 3d plant structures\
 , primarily root systems of plants.\nOur measures are based on the medial 
 axis transform\, which plays a\nfundamental role in shape matching and ana
 lysis\, but is widely known\nto be unstable to even small boundary perturb
 ations. Methods for\npruning the medial axis are usually guided by some me
 asure of\nsignificance\, with considerable work done for both 2- and\n3-di
 mensional shapes. Such significance measures can be used for\nidentifying 
 salient features\, and hence are useful for simplification\,\ncomparison\,
  and alignment. In this talk\, we will present theoretical\ninsights and p
 roperties of commonly used significance measures\,\nfocusing on those in 2
 D and 3D that are both shape-revealing and\ntopology-preserving\, as well 
 as being robust to noise on the boundary.\nWe'll then discuss several meth
 ods that de-noise a shape and identify\ntopologically and geometrically pr
 ominent features\, using both the\nmedial axis and other measures commonly
  used in topological data\nanalysis. Our methods are quite successful comp
 ared to the state of\nthe art\, and are available in the package TopoRoot\
 , an automatic\npipeline for plant architectural analysis from 3D Imaging.
 \n\nhttps://datascience.utah.edu/talks/2021-10-22-erin-wolf-chambers/
URL:https://datascience.utah.edu/talks/2021-10-22-erin-wolf-chambers/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 | https://utah.zoom.us/j/93778940103?pwd=TStQRWhWVjRxd0hG
 V1hTK05SUFZwUT09
END:VEVENT
BEGIN:VEVENT
UID:2021-10-29-bao-wang@datascience.utah.edu
DTSTAMP:20211029T200000Z
DTSTART;TZID=America/Denver:20211029T140000
DTEND;TZID=America/Denver:20211029T150000
SUMMARY:How Differential Equations and Random Graph Insights Benefit Deep L
 earning - Bao Wang
DESCRIPTION:Bao Wang (Utah Math\, SCI)\n\nWe will present recent results on
  developing new deep learning algorithms leveraging differential equations
  and random graph insights.\nFirst\, we will present a new class of contin
 uous-depth deep neural networks that were motivated by the ODE limit of th
 e classical momentum method\, named heavy-ball neural ODEs (HBNODEs). HBNO
 DEs enjoy two properties that imply practical advantages over NODEs: (i) T
 he adjoint state of an HBNODE also satisfies an HBNODE\, accelerating both
  forward and backward ODE solvers\, thus significantly accelerate learning
  and improve the utility of the trained models. (ii) The spectrum of HBNOD
 Es is well structured\, enabling effective learning of long-term dependenc
 ies from complex sequential data.\nSecond\, we will extend HBNODE to graph
  learning leveraging diffusion on graphs\, resulting in new algorithms for
  deep graph learning. The new algorithms are more accurate than existing d
 eep graph learning algorithms and more scalable to deep architectures\, an
 d also suitable for learning at low labeling rate regimes. Moreover\, we w
 ill present a fast multipole method-based efficient attention mechanism fo
 r modeling graph nodes interactions.\nThird\, if time permits\, we will di
 scuss building an efficient and reliable overlay network for decentralized
  federated learning based on the random graph theory.\n\nhttps://datascien
 ce.utah.edu/talks/2021-10-29-bao-wang/
URL:https://datascience.utah.edu/talks/2021-10-29-bao-wang/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 | https://utah.zoom.us/j/93778940103?pwd=TStQRWhWVjRxd0hG
 V1hTK05SUFZwUT09
CATEGORIES:algorithms & theory,deep learning,networks & graphs
END:VEVENT
BEGIN:VEVENT
UID:2021-11-05-marina-kogan@datascience.utah.edu
DTSTAMP:20211105T200000Z
DTSTART;TZID=America/Denver:20211105T140000
DTEND;TZID=America/Denver:20211105T150000
SUMMARY:Sequence-based approaches as human-centered data science methods fo
 r crisis informatics - Marina Kogan
DESCRIPTION:Marina Kogan (Utah SoC)\n\nSocial media platforms have been inc
 reasingly used by the public in crisis situations\, partly because they up
 end the traditional top-down broadcasting model of risk communication. Ins
 tead\, social media platforms facilitate a two-way information exchange be
 tween the official response channels and the general public\, enabling mor
 e participatory crisis communication\, as well as coordination and self-or
 ganization among the public. In this more complex information ecosystem\, 
 understanding the flow of information is crucial to supporting those affec
 ted and preventing malicious actors from capitalizing on the uncertainty. 
 However\, the study of such information flows is challenging\, because the
  high-tempo\, high-volume convergent nature of crisis events produces vast
  amounts of social media data\, necessitating the use of the data science 
 methods. On the other hand\, to glean meaningful insight from the crisis-r
 elated social media activity\, it is necessary to use methods that account
  for the complex social context of the user activity. In this talk I will 
 show how the Human-Centered Data Science (HCDS) provides methodological ap
 proaches that both harness the power of computational methods and account 
 for the highly situated nature of social media activity in disruption. I w
 ill focus on sequence-based approaches as examples of HCDS methods in two 
 empirical studies: analysis of attention-garnering information during a na
 tural disaster and investigation of behavioral signatures in coordinated i
 nformation operations.\n\nhttps://datascience.utah.edu/talks/2021-11-05-ma
 rina-kogan/
URL:https://datascience.utah.edu/talks/2021-11-05-marina-kogan/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 | https://utah.zoom.us/j/93778940103?pwd=TStQRWhWVjRxd0hG
 V1hTK05SUFZwUT09
CATEGORIES:human-centered computing,society & policy
END:VEVENT
BEGIN:VEVENT
UID:2021-11-12-mikhail-belkin@datascience.utah.edu
DTSTAMP:20211112T210000Z
DTSTART;TZID=America/Denver:20211112T140000
DTEND;TZID=America/Denver:20211112T150000
SUMMARY:From classical statistics to modern deep learning - Mikhail Belkin
DESCRIPTION:Mikhail Belkin\n\nRecent empirical successes of deep learning h
 ave exposed significant gaps in our\nfundamental understanding of learning
  and optimization mechanisms.\nModern best practices for model selection a
 re in direct contradiction to the methodologies\nsuggested by classical an
 alyses. Similarly\, the efficiency of SGD-based local methods\nused in tra
 ining modern models\, appeared at odds with the standard intuitions on opt
 imization.\n\nFirst\, I will present evidence\, empirical and mathematical
 \, that necessitates\nrevisiting classical statistical notions\, such as o
 ver-fitting. I will continue to discuss the emerging\nunderstanding of gen
 eralization\, and\, in particular\, the "double descent" risk curve\, whic
 h extends\nthe classical U-shaped generalization curve beyond the point of
  interpolation.\n\nSecond\, I will discuss why the landscapes of over-para
 meterized neural networks are\ngenerically never convex\, even locally. In
 stead they satisfy the Polyak-Lojasiewicz (PL)\ncondition across most of t
 he parameter space instead\, presents an powerful framework for optimizati
 on in general over-parameterized models and allows SGD-type methods to con
 verge to a global minimum.\n\nWhile our understanding has significantly gr
 own in the last few years\, a key piece of the puzzle remains -- how does 
 optimization align with statistics to form the complete mathematical pictu
 re of modern ML?\n\nhttps://datascience.utah.edu/talks/2021-11-12-mikhail-
 belkin/
URL:https://datascience.utah.edu/talks/2021-11-12-mikhail-belkin/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://utah.zoom.us/j/93778940103?pwd=TStQRWhWVjRxd0hGV1hTK05SUFZ
 wUT09
CATEGORIES:deep learning,machine learning,optimization,statistics
END:VEVENT
BEGIN:VEVENT
UID:2021-11-19-michael-yeh@datascience.utah.edu
DTSTAMP:20211119T210000Z
DTSTART;TZID=America/Denver:20211119T140000
DTEND;TZID=America/Denver:20211119T150000
SUMMARY:Towards a Near Universal Time Series Data Mining Tool: Introducing 
 the Matrix Profile - Michael Yeh
DESCRIPTION:Michael Yeh (Visa Research)\n\nMatrix profile is a data structu
 re that annotates a time series by recording the location of and the dista
 nce to the nearest neighbors of each subsequences in the time series. The 
 matrix profile stores such information in an efficient and easy-to-access 
 fashion and can be used in a variety of data mining tasks like motif/disco
 rd discovery\, semantic segmentation\, and clustering. In this talk\, I wi
 ll 1) introduce what matrix profile is\, 2) discuss the computational chal
 lenge associated with matrix profile\, and 3) show how it can be used in d
 ifferent time series data mining tasks.\n\nhttps://datascience.utah.edu/ta
 lks/2021-11-19-michael-yeh/
URL:https://datascience.utah.edu/talks/2021-11-19-michael-yeh/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://utah.zoom.us/j/93778940103?pwd=TStQRWhWVjRxd0hGV1hTK05SUFZ
 wUT09
CATEGORIES:algorithms & theory,computer vision
END:VEVENT
BEGIN:VEVENT
UID:2021-12-03-julia-silge@datascience.utah.edu
DTSTAMP:20211203T210000Z
DTSTART;TZID=America/Denver:20211203T140000
DTEND;TZID=America/Denver:20211203T150000
SUMMARY:Data visualization for machine learning practitioners - Julia Silge
DESCRIPTION:Julia Silge (RStudio)\n\nVisual representations of data inform 
 how machine learning practitioners think\, understand\, and decide. Before
  charts are ever used for outward communication about a ML system\, they a
 re used by the system designers and operators themselves as a tool to make
  better modeling choices. Practitioners use visualization\, from very fami
 liar statistical graphics to creative and less standard plots\, at the poi
 nts of most important human decisions when other ways to validate those de
 cisions can be difficult. Visualization approaches are used to understand 
 both the data that serves as input for machine learning and the models tha
 t practitioners create. In this talk\, learn about the process of building
  a ML model in the real world\, how and when practitioners use visualizati
 on to make more effective choices\, and considerations for ML visualizatio
 n tooling.\n\nhttps://datascience.utah.edu/talks/2021-12-03-julia-silge/
URL:https://datascience.utah.edu/talks/2021-12-03-julia-silge/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 and Zoom | https://utah.zoom.us/j/93778940103?pwd=TStQRWh
 WVjRxd0hGV1hTK05SUFZwUT09
CATEGORIES:human-centered computing,machine learning,visualization
END:VEVENT
BEGIN:VEVENT
UID:2021-12-10-sameer-singh@datascience.utah.edu
DTSTAMP:20211210T210000Z
DTSTART;TZID=America/Denver:20211210T140000
DTEND;TZID=America/Denver:20211210T150000
SUMMARY:Evaluating and Testing Natural Language Processing Models - Sameer 
 Singh
DESCRIPTION:Sameer Singh (UC Irvine)\n\nCurrent evaluation of the generaliz
 ation of natural language processing (NLP) systems\, and much of machine l
 earning\, primarily consists of measuring the accuracy on held-out instanc
 es of the dataset. Since the held-out instances are often gathered using s
 imilar annotation process as the training data\, they include the same bia
 ses that act as shortcuts for machine learning models\, allowing them to a
 chieve accurate results without requiring actual natural language understa
 nding. Thus held-out accuracy is often a poor proxy for measuring generali
 zation. Further\, aggregate metrics have little to say about where the pro
 blems may lie\, and how to address them.\nIn this talk\, I will introduce 
 a number of approaches we are investigating to perform a more thorough eva
 luation of NLP systems. I will first provide a quick overview of automated
  techniques for perturbing instances in the dataset that identify loophole
 s and shortcuts in NLP models\, including semantic adversaries and univers
 al triggers. I will then describe recent work on creating comprehensive an
 d thorough tests and evaluation benchmarks for NLP using CheckList\, that 
 aim to directly evaluate comprehension and understanding capabilities. The
  talk will include a number of NLP tasks\, such as sentiment analysis\, te
 xtual entailment\, paraphrase detection\, and question answering.\n\nDr. S
 ameer Singh is an Associate Professor of Computer Science at the Universit
 y of California\, Irvine (UCI) and an Allen AI Fellow at Allen Institute f
 or AI. He is working primarily on robustness and interpretability of machi
 ne learning algorithms\, along with models that reason with text and struc
 ture for natural language processing. Sameer was a postdoctoral researcher
  at the University of Washington and received his PhD from the University 
 of Massachusetts\, Amherst. He has received the NSF CAREER award\, selecte
 d as a DARPA Riser\, UCI Distinguished Early Career Faculty award\, and th
 e Hellman Faculty Fellowship. His group has received funding from Allen In
 stitute for AI\, Amazon\, NSF\, DARPA\, Adobe Research\, Hasso Plattner In
 stitute\, NEC\, Base 11\, and FICO. Sameer has published extensively at ma
 chine learning and natural language processing venues and received confere
 nce paper awards at KDD 2016\, ACL 2018\, EMNLP 2019\, AKBC 2020\, and ACL
  2020. (https://sameersingh.org/)\n\nhttps://datascience.utah.edu/talks/20
 21-12-10-sameer-singh/
URL:https://datascience.utah.edu/talks/2021-12-10-sameer-singh/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://utah.zoom.us/j/93778940103?pwd=TStQRWhWVjRxd0hGV1hTK05SUFZ
 wUT09
CATEGORIES:algorithms & theory,machine learning,natural language processing
END:VEVENT
BEGIN:VEVENT
UID:2022-01-14-yi-zhou@datascience.utah.edu
DTSTAMP:20220114T220000Z
DTSTART;TZID=America/Denver:20220114T150000
DTEND;TZID=America/Denver:20220114T160000
SUMMARY:Understanding the Convergence of Optimization Algorithms for Minima
 x Machine Learning - Yi Zhou
DESCRIPTION:Yi Zhou (Utah ECE)\n\nThe past decade has witnessed the great s
 uccess of deep learning in broad societal and commercial applications. How
 ever\, conventional deep learning relies on wildly fitting data with neura
 l networks\, which is known to produce models that lack resilience. For in
 stance\, models used in facial recognition and healthcare are known to be 
 biased toward people of a certain race or gender. Models used in autonomou
 s driving are vulnerable to malicious attacks\, i.e.\, putting an art stic
 ker on a stop sign may force the model to classify it as a speed limit sig
 n. Therefore\, the next-generation deep learning paradigm aims to deliver 
 resilient models that promote robustness to malicious attacks\, fairness a
 mong users\, and privacy preservation\, and this can be realized by levera
 ging the emerging minimax machine learning framework. In this talk\, I wil
 l present three gradient-descent-ascent (GDA) type of optimization algorit
 hms for solving different classes of nonconvex minimax machine learning pr
 oblems. Then\, I will present a principled nonconvex minimax optimization 
 theory that establishes the global convergence and convergence rates of th
 ese algorithms.\n\nhttps://datascience.utah.edu/talks/2022-01-14-yi-zhou/
URL:https://datascience.utah.edu/talks/2022-01-14-yi-zhou/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 1250 | https://utah.zoom.us/j/94615915833?pwd=TW50Mk5GeFpQN0lB
 MEZQQ2Z1ZUdFUT09
CATEGORIES:algorithms & theory,deep learning,machine learning,optimization
END:VEVENT
BEGIN:VEVENT
UID:2022-01-21-chinmay-hedge@datascience.utah.edu
DTSTAMP:20220121T220000Z
DTSTART;TZID=America/Denver:20220121T150000
DTEND;TZID=America/Denver:20220121T160000
SUMMARY:Designing Neural Networks for Efficient Encrypted Inference - Chinm
 ay Hedge
DESCRIPTION:Chinmay Hedge (NYU)\n\nAs deep neural networks become ever more
  pervasive\, so too are concerns surrounding users' data privacy. Curiousl
 y\, standard cryptographic encryption approaches for guaranteeing data pri
 vacy do not interact well with traditional neural network models. In this 
 talk\, I will (a) outline why standard networks are not encryption-efficie
 nt\, (b) suggest two new approaches for designing deep networks that do su
 pport efficient and secure inference\, and (c) show results instantiating 
 these approaches on real-world use cases.\n\nhttps://datascience.utah.edu/
 talks/2022-01-21-chinmay-hedge/
URL:https://datascience.utah.edu/talks/2022-01-21-chinmay-hedge/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://utah.zoom.us/j/94615915833?pwd=TW50Mk5GeFpQN0lBMEZQQ2Z1ZUd
 FUT09
CATEGORIES:deep learning,privacy & security
END:VEVENT
BEGIN:VEVENT
UID:2022-01-28-swaroop-mishra@datascience.utah.edu
DTSTAMP:20220128T220000Z
DTSTART;TZID=America/Denver:20220128T150000
DTEND;TZID=America/Denver:20220128T160000
SUMMARY:Towards the Development of Models that Learn New Tasks from Instruc
 tions - Swaroop Mishra
DESCRIPTION:Swaroop Mishra (ASU\, Ph.D. Student)\n\nhttps://datascience.uta
 h.edu/talks/2022-01-28-swaroop-mishra/
URL:https://datascience.utah.edu/talks/2022-01-28-swaroop-mishra/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 (LCR) | https://utah.zoom.us/j/94615915833?pwd=TW50Mk5GeF
 pQN0lBMEZQQ2Z1ZUdFUT09
CATEGORIES:education
END:VEVENT
BEGIN:VEVENT
UID:2022-02-04-tuhin-chakrabarty@datascience.utah.edu
DTSTAMP:20220204T220000Z
DTSTART;TZID=America/Denver:20220204T150000
DTEND;TZID=America/Denver:20220204T160000
SUMMARY:The Curious Case of Figurative Language - Tuhin Chakrabarty
DESCRIPTION:Tuhin Chakrabarty (Columbia University\, Ph.D Student)\n\nDespi
 te the ubiquity of figurative language across various forms of speech and 
 writing\, the vast majority of NLP research focuses primarily on literal l
 anguage. Figurative language is challenging because of its implicit nature
 . In this talk\, I will specifically focus on the following questions: 1) 
 Can large language models understand/interpret them? 2) Can computers gene
 rate figurative language?\n\nhttps://datascience.utah.edu/talks/2022-02-04
 -tuhin-chakrabarty/
URL:https://datascience.utah.edu/talks/2022-02-04-tuhin-chakrabarty/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 (LCR) | https://utah.zoom.us/j/94615915833?pwd=TW50Mk5GeF
 pQN0lBMEZQQ2Z1ZUdFUT09
CATEGORIES:natural language processing
END:VEVENT
BEGIN:VEVENT
UID:2022-02-11-vaggos-chatziafratis@datascience.utah.edu
DTSTAMP:20220211T220000Z
DTSTART;TZID=America/Denver:20220211T150000
DTEND;TZID=America/Denver:20220211T160000
SUMMARY:Neural Networks Expressivity through the lens of Dynamical Systems 
 - Vaggos Chatziafratis
DESCRIPTION:Vaggos Chatziafratis (UC Santa Cruz)\n\nGiven a target function
  f\, how large must a neural network be in order to approximate f?\nUnders
 tanding the representational power of Deep Neural Networks (DNNs) and how 
 their structural properties (e.g.\, depth\, width\, type of activation uni
 t) affect the functions they can compute\, has been an important yet chall
 enging question in approximation theory and deep learning even in the earl
 y days of AI.\n\nIn this talk\, I want to tell you about some recent progr
 ess on this topic that uses ideas from dynamical systems. The main results
  are exponential depth-width trade-offs for DNNs representing certain fami
 lies of functions. Our techniques rely on a generalized notion of fixed po
 ints\, called periodic points that have played a major role in chaos theor
 y (Li-Yorke chaos and Sharkovsky's theorem).\n\nBased on three recent work
 s:\n- with Ioannis Panageas\, Sai Ganesh Nagarajan and Xiao Wang from ICLR
 '20 (spotlight): https://arxiv.org/abs/1912.04378 (https://arxiv.org/abs/1
 912.04378)\n- with Ioannis Panageas and Sai Ganesh Nagarajan from ICML'20:
  https://arxiv.org/abs/2003.00777 (https://arxiv.org/abs/2003.00777)\n- wi
 th Clayton Sanford from AISTATS'22: https://arxiv.org/abs/2110.10295 (http
 s://arxiv.org/abs/2110.10295)\n\nhttps://datascience.utah.edu/talks/2022-0
 2-11-vaggos-chatziafratis/
URL:https://datascience.utah.edu/talks/2022-02-11-vaggos-chatziafratis/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://utah.zoom.us/j/94615915833?pwd=TW50Mk5GeFpQN0lBMEZQQ2Z1ZUd
 FUT09
CATEGORIES:deep learning
END:VEVENT
BEGIN:VEVENT
UID:2022-02-18-jeff-phillips@datascience.utah.edu
DTSTAMP:20220218T220000Z
DTSTART;TZID=America/Denver:20220218T150000
DTEND;TZID=America/Denver:20220218T160000
SUMMARY:Some Very Basic Theory of Classification (in low dimensions - Jeff 
 Phillips
DESCRIPTION:Jeff Phillips (Utah SoC)\n\nI plan to talk about some new resul
 ts on some very fundamental (but overlooked) theory questions in classific
 ation. First\, how fast can you find a linear classifier that eps-approxim
 ates the optimal one in terms of miss-classification? Second\, if you want
  to preserve a Euclidean margin between perfectly classified points for a 
 polynomial classifier\, how many samples do you need?\n\nhttps://datascien
 ce.utah.edu/talks/2022-02-18-jeff-phillips/
URL:https://datascience.utah.edu/talks/2022-02-18-jeff-phillips/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 1250
CATEGORIES:machine learning
END:VEVENT
BEGIN:VEVENT
UID:2022-02-25-sunipa-dev@datascience.utah.edu
DTSTAMP:20220225T220000Z
DTSTART;TZID=America/Denver:20220225T150000
DTEND;TZID=America/Denver:20220225T160000
SUMMARY:Towards Inclusive and Socially Aware Language Technologies - Sunipa
  Dev
DESCRIPTION:Sunipa Dev (Google Research)\n\nLarge language models are commo
 nly used in different paradigms of natural language processing and machine
  learning\, and are known for their efficiency as well as their overall la
 ck of interpretability. Their data driven approach for emulating human lan
 guage often results in human biases being encoded and even amplified\, pot
 entially leading to cyclic propagation of representational and allocationa
 l harm. We discuss in this talk some aspects of detecting\, evaluating\, a
 nd mitigating biases and associated harms in a holistic\, inclusive\, and 
 culturally-aware manner. In particular\, we discuss the disparate impact o
 n society of common language tools that are not inclusive of all gender id
 entities.\n\nhttps://datascience.utah.edu/talks/2022-02-25-sunipa-dev/
URL:https://datascience.utah.edu/talks/2022-02-25-sunipa-dev/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://utah.zoom.us/j/94615915833?pwd=TW50Mk5GeFpQN0lBMEZQQ2Z1ZUd
 FUT09
CATEGORIES:fairness & ethics,natural language processing
END:VEVENT
BEGIN:VEVENT
UID:2022-03-04-anirudh-goyal-of-montreal@datascience.utah.edu
DTSTAMP:20220304T220000Z
DTSTART;TZID=America/Denver:20220304T150000
DTEND;TZID=America/Denver:20220304T160000
SUMMARY:From Specialists to Generalists: Inductive Biases of Deep Learning 
 for Higher Level Cognition - Anirudh Goyal of Montreal
DESCRIPTION:Anirudh Goyal of Montreal\n\nA fascinating hypothesis is that h
 uman and animal intelligence could be explained by a few principles (rathe
 r than an encyclopedic list of heuristics). If that hypothesis was correct
 \, we could more easily both understand our own intelligence and build int
 elligent machines. Just like in physics\, the principles themselves would 
 not be sufficient to predict the behavior of complex systems like brains\,
  and substantial computation might be needed to simulate human-like intell
 igence. This hypothesis would suggest that studying the kind of inductive 
 biases that humans and animals exploit could help both clarify these princ
 iples and provide inspiration for AI research and neuroscience theories. D
 eep learning already exploits several key inductive biases\, and my work c
 onsiders a larger list\, focusing on those which concern mostly higher-lev
 el and sequential conscious processing. The objective of clarifying these 
 particular principles is that they could potentially help us build AI syst
 ems benefiting from humans' abilities in terms of flexible out-of-distribu
 tion and systematic generalization\, which is currently an area where a la
 rge gap exists between state-of-the-art machine learning and human intelli
 gence.\n\nhttps://datascience.utah.edu/talks/2022-03-04-anirudh-goyal-of-m
 ontreal/
URL:https://datascience.utah.edu/talks/2022-03-04-anirudh-goyal-of-montreal
 /
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 (LCR) | https://utah.zoom.us/j/94615915833?pwd=TW50Mk5GeF
 pQN0lBMEZQQ2Z1ZUdFUT09
CATEGORIES:deep learning,fairness & ethics,machine learning,statistics
END:VEVENT
BEGIN:VEVENT
UID:2022-03-18-greg-herschlag-jon-mattingly@datascience.utah.edu
DTSTAMP:20220318T210000Z
DTSTART;TZID=America/Denver:20220318T150000
DTEND;TZID=America/Denver:20220318T160000
SUMMARY:Quantifying Gerrymandering: Hearing the Will of the People - Greg H
 erschlag\, Jon Mattingly
DESCRIPTION:Greg Herschlag (Duke University)\, Jon Mattingly (Duke Universi
 ty)\n\nGerrymandering is the process of manipulating political\ndistricts 
 either to amplify the power of a political group or suppress\nthe represen
 tation of certain demographic groups. Although we have seen\nincreasingly 
 precise and effective gerrymanders\, a number of\nmathematicians\, politic
 al scientists\, and lawyers are developing\neffective methodologies at unc
 overing and understanding the intent and\neffects of gerrymandered distric
 ts.\n\nThe basic idea behind these methods is to compare a given set of\nd
 istricts to a large collection of neutrally drawn plans. The process\nreli
 es on three distinct components: First\, we determine rules for\ncompliant
  redistricting plans along with codifying preferences between\nthese plans
 \; next\, we sample the space of compliant redistricting plans\n(according
  to our preferences) and generate a large collection of\nnon-partisan alte
 rnatives\; finally\, we compare the collection of plans\nto a particular p
 lan of interest.\n\nIn this talk\, we will discuss how our research group 
 at Duke has\nanalyzed gerrymandering. We will discuss the sampling methods
  we employ\,\nincluding several recent algorithmic advances. We will also 
 discuss how\nrecent advances in North Carolina's legal theory have brought
  partisan\nsymmetry considerations back into focus.\n\nhttps://datascience
 .utah.edu/talks/2022-03-18-greg-herschlag-jon-mattingly/
URL:https://datascience.utah.edu/talks/2022-03-18-greg-herschlag-jon-mattin
 gly/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 (LCR) | https://utah.zoom.us/j/94615915833?pwd=TW50Mk5GeF
 pQN0lBMEZQQ2Z1ZUdFUT09
CATEGORIES:algorithms & theory,statistics
END:VEVENT
BEGIN:VEVENT
UID:2022-03-25-chad-topaz-jude-higdon@datascience.utah.edu
DTSTAMP:20220325T210000Z
DTSTART;TZID=America/Denver:20220325T150000
DTEND;TZID=America/Denver:20220325T160000
SUMMARY:Quantitative Approaches to Social Justice - Chad Topaz\, Jude Higdo
 n
DESCRIPTION:Chad Topaz (QSIDE)\, Jude Higdon (QSIDE)\n\nCivil rights leader
 \, educator\, and investigative journalist Ida B. Wells said that "the way
  to right wrongs is to shine the light of truth upon them." This talk will
  demonstrate how mathematical\, statistical\, and computational approaches
  can shine a light on social injustices and help build solutions to remedy
  them. We will present research-to-action projects on diversity in art mus
 eums\, inclusion in STEM\, equity in criminal sentencing\, and other topic
 s. The tools engaged include crowdsourcing\, data cleaning\, clustering\, 
 hypothesis testing\, statistical modeling\, Markov chains\, data visualiza
 tion\, and much more. Overall\, we hope that this talk leaves you informed
  about the breadth of social justice applications that one can tackle usin
 g quantitative tools in careful collaboration with other scholars and acti
 vists.\n\nhttps://datascience.utah.edu/talks/2022-03-25-chad-topaz-jude-hi
 gdon/
URL:https://datascience.utah.edu/talks/2022-03-25-chad-topaz-jude-higdon/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 1250 | https://utah.zoom.us/j/94615915833?pwd=TW50Mk5GeFpQN0lB
 MEZQQ2Z1ZUdFUT09
CATEGORIES:fairness & ethics,statistics
END:VEVENT
BEGIN:VEVENT
UID:2022-04-01-debanjan-mahata@datascience.utah.edu
DTSTAMP:20220401T210000Z
DTSTART;TZID=America/Denver:20220401T150000
DTEND;TZID=America/Denver:20220401T160000
SUMMARY:Identifying Keyphrases from Text Documents - From Heuristics to Lan
 guage Models - Debanjan Mahata
DESCRIPTION:Debanjan Mahata (Moody Analytics)\n\nAutomatic identification o
 f keyphrases from text documents is an extreme summarization problem that 
 lies at the intersection of the areas of natural language processing (NLP)
  and information retrieval (IR). Keyphrases aid in capturing the most sali
 ent topics from the input text and are useful in multiple downstream tasks
  such as classification\, clustering\, summarization\, document recommenda
 tion\, query expansion\, interactive document retrieval\, semantic and fac
 eted search. Despite the ground-breaking advancements triggered by deep ne
 ural networks\, automatically identifying keyphrases from text using machi
 ne learning techniques is still a challenging problem that hasn't been exp
 lored as much as other related and popular tasks such as named entity extr
 action\, question answering\, summarization. This talk will provide an ove
 rview of the advances made in keyphrase extraction and generation from tex
 t documents and present the approaches\, datasets\, evaluation strategies\
 , and the associated challenges. It will also dive into the topic of how l
 anguage models have been effective in pushing state-of-the-art performance
 s in this domain. Lastly\, the speaker will present the current trends and
  future directions for research in this domain.\n\nhttps://datascience.uta
 h.edu/talks/2022-04-01-debanjan-mahata/
URL:https://datascience.utah.edu/talks/2022-04-01-debanjan-mahata/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:" | https://utah.zoom.us/j/94615915833?pwd=TW50Mk5GeFpQN0lBMEZQQ2Z
 1ZUdFUT09
CATEGORIES:deep learning,machine learning,natural language processing,socie
 ty & policy
END:VEVENT
BEGIN:VEVENT
UID:2022-04-08-tao-li@datascience.utah.edu
DTSTAMP:20220408T210000Z
DTSTART;TZID=America/Denver:20220408T150000
DTEND;TZID=America/Denver:20220408T160000
SUMMARY:Improving Data Efficiency of Neural Models using Logic - Tao Li
DESCRIPTION:Tao Li (Google Research)\n\nIn this talk\, we will focus on a s
 imple approach that uses logic to improve neural model performance for nat
 ural language processing (NLP) tasks. Many downstream NLP tasks involve do
 main knowledge that can be easily stated in logical forms. We argue that w
 e can use such knowledge to improve model learning. This results in better
  data efficiency\, i.e.\, a model that performs better with less annotatio
 n. To this end\, we propose frameworks that integrate domain knowledge\, e
 xpressed as declarative constraints\, with neural models. We show that suc
 h integration substantially improves state-of-the-art neural models in a v
 ariety of NLP tasks. To facilitate using our frameworks\, we will also pro
 pose a PyTorch library that unifies differentiable tensor operations and l
 ogical operations.\n\nhttps://datascience.utah.edu/talks/2022-04-08-tao-li
 /
URL:https://datascience.utah.edu/talks/2022-04-08-tao-li/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 (LCR) | https://utah.zoom.us/j/94615915833?pwd=TW50Mk5GeF
 pQN0lBMEZQQ2Z1ZUdFUT09
CATEGORIES:natural language processing
END:VEVENT
BEGIN:VEVENT
UID:2022-04-15-abhinav-kumar@datascience.utah.edu
DTSTAMP:20220415T210000Z
DTSTART;TZID=America/Denver:20220415T150000
DTEND;TZID=America/Denver:20220415T160000
SUMMARY:Mathematical Modeling for Landmark and 3D Object Detection - Abhina
 v Kumar
DESCRIPTION:Abhinav Kumar\n\nModern computer vision models have excelled on
  several tasks and beaten several benchmarks. However\, many of these mode
 ls have avoided the mathematical and principled approaches to computer vis
 ion\, resulting in sub-optimal performance and issues with interpretabilit
 y. This talk will re-introduce mathematical modeling for computer vision t
 asks such as facial landmark localization and monocular 3D object detectio
 n for autonomous driving. In particular\, we discuss the joint estimation 
 of location\, uncertainty\, and visibility for facial landmark detection a
 nd introduce a mathematically differentiable NMS for monocular 3D detectio
 n. The mathematical modeling enables end-to-end learning in these tasks\, 
 resulting in improved performance.\n\nhttps://datascience.utah.edu/talks/2
 022-04-15-abhinav-kumar/
URL:https://datascience.utah.edu/talks/2022-04-15-abhinav-kumar/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 (LCR) | https://utah.zoom.us/j/94615915833?pwd=TW50Mk5GeF
 pQN0lBMEZQQ2Z1ZUdFUT09
CATEGORIES:computer vision,robotics
END:VEVENT
BEGIN:VEVENT
UID:2022-04-22-khyati-chandu@datascience.utah.edu
DTSTAMP:20220422T210000Z
DTSTART;TZID=America/Denver:20220422T150000
DTEND;TZID=America/Denver:20220422T160000
SUMMARY:Anchoring Multimodal Narrative Generation - Khyati Chandu
DESCRIPTION:Khyati Chandu (Meta AI)\n\nHumans inherently learn from and int
 eract with multiple views of information\, be it various modalities or lan
 guages. So\, the expectations from contemporary and 21st-century technolog
 y are a testimony to the increasing need to model these multiview contexts
  better. Natural language generation plays a pivotal role in communicating
  these contexts in human-understandable languages. This talk brings togeth
 er both of these transformative technologies to make strides toward a long
 standing dream of human-like multiview narrative generation. The critical 
 challenge is identifying the natural-sounding properties of long-form text
 s and modeling them in tandem with visual contexts. I present anchors for 
 grounding three such properties including content (relevance)\, structure 
 (coherence)\, and surface form realization (expression)\, and anchors them
  with relevant visual contexts. These anchors also provide us with human i
 nterpretable handles for controlling these properties.\n\nDetails: To illu
 strate the effectiveness of the anchors for each of the three properties\,
  I present: Starting with content: In situated multimodal contexts\, relev
 ance is the concept of the elements in one modality being connected to the
  other modality that makes this context informative and complementary. I p
 resent visual infilling with curriculum learning as a global objective for
  content and hierarchically attending over entity skeletons as a local obj
 ective for content\, to generate visual stories and procedures. To improve
  the controllability and transferability in English and five other languag
 es\, I also introduce a dual-stage model with weakly supervised skeletons 
 and a text-as-side attention mechanism to denoise the content in an image 
 caption. Moving onto structure: The alignment of descriptions in language 
 to the corresponding visual inputs is crucial to generating a logical and 
 coherent narrative. I present a scaffolding technique as a local objective
  for structure by extracting a layout from vast amounts of unsupervised te
 xt to incorporate structure into cooking recipes generated from images. Fi
 nally\, surface form: The crux of naturalness to automatic generation come
 s by incorporating individualized and personalized ways of expressing the 
 same content. I present a locally guided\, weakly supervised model for gen
 erating persona-based visual stories and a dual-staged adversarial techniq
 ue to generate mixed views from non-parallel data. All of the above work m
 ainly focuses on static multimodal narratives\, and finally\, I present a 
 case to highlight the significance of transitioning to dynamic grounding. 
 I conclude by presenting the shortcomings of the current approaches in the
  NLP domain to the grounding problem and offer recommendations along with 
 executable actions for course correction to bridge this gap and enable gro
 unding for machines to resemble human communication.\n\nhttps://datascienc
 e.utah.edu/talks/2022-04-22-khyati-chandu/
URL:https://datascience.utah.edu/talks/2022-04-22-khyati-chandu/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 1250 and Zoom | https://utah.zoom.us/j/94615915833?pwd=TW50Mk5
 GeFpQN0lBMEZQQ2Z1ZUdFUT09
CATEGORIES:computer vision,education,machine learning,natural language proc
 essing
END:VEVENT
BEGIN:VEVENT
UID:2022-04-29-james-brundage@datascience.utah.edu
DTSTAMP:20220429T210000Z
DTSTART;TZID=America/Denver:20220429T150000
DTEND;TZID=America/Denver:20220429T160000
SUMMARY:Leveraging Unlabeled Data for Machine Learning in the Electrocardio
 gram - James Brundage
DESCRIPTION:James Brundage (UU HSC)\n\nSupervised deep learning (DL) has be
 come an increasingly common tool for advanced analysis of the electrocardi
 ogram (ECG). These methods rely heavily on labeled datasets\, in which the
 re is a clinical annotation for each ECG. However\, real world ECG dataset
 s may not contain enough labeled recordings to facilitate robust feature e
 xtraction\, preventing DL analysis for clinical problems with small datase
 ts. Self-supervised learning (SSL) seeks to utilize cheaply labeled or unl
 abeled data to improve performance in a supervised learning task. This pro
 cess consists of first training a model on a primary task with cheap data 
 labels\, followed by a second training process which attempts to learn the
  downstream task by initializing with weights learned from the first. Whil
 e SSL has become a popular tool in many machine learning domains\, it is o
 nly starting to be used in ECG based machine learning. Here\, we demonstra
 te the progress we have made in applying SSL approaches to detect low left
  ventricular ejection fraction\, a complex ECG detection task\, using data
  extracted from the University of Utah.\n\nhttps://datascience.utah.edu/ta
 lks/2022-04-29-james-brundage/
URL:https://datascience.utah.edu/talks/2022-04-29-james-brundage/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 (LCR) | https://utah.zoom.us/j/94615915833?pwd=TW50Mk5GeF
 pQN0lBMEZQQ2Z1ZUdFUT09
CATEGORIES:health & medicine,machine learning
END:VEVENT
BEGIN:VEVENT
UID:2022-08-24-bei-wang-phillips@datascience.utah.edu
DTSTAMP:20220824T163000Z
DTSTART;TZID=America/Denver:20220824T103000
DTEND;TZID=America/Denver:20220824T114500
SUMMARY:On Hypergraph Analysis and Visualization - Bei Wang Phillips
DESCRIPTION:Bei Wang Phillips (Utah SoC/SCI)\n\nHypergraphs capture multi-w
 ay relationships in data\, and they have\nconsequently seen a number of ap
 plications in higher-order network\nanalysis\, computer vision\, geometry 
 processing\, and machine learning.\nIn this talk\, I will discuss hypergra
 ph analysis and visualization.\nIn particular\, I will focus on developing
  the theoretical foundations\nin studying the space\nof hypergraphs using 
 ingredients from optimal transport.\n\nThis talk is\nbased on joint works\
 nwith Youjia Zhou\, Archit Rathore\, Emilie Purvine\, Samir Chowdhury\, To
 m\nNeedham\, and Ethan Semrad.\n\nhttps://datascience.utah.edu/talks/2022-
 08-24-bei-wang-phillips/
URL:https://datascience.utah.edu/talks/2022-08-24-bei-wang-phillips/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780
CATEGORIES:networks & graphs,visualization
END:VEVENT
BEGIN:VEVENT
UID:2022-08-31-casey-greene@datascience.utah.edu
DTSTAMP:20220831T163000Z
DTSTART;TZID=America/Denver:20220831T103000
DTEND;TZID=America/Denver:20220831T114500
SUMMARY:Talk by Casey Greene - Casey Greene
DESCRIPTION:Casey Greene (CU Anschutz)\n\nhttps://datascience.utah.edu/talk
 s/2022-08-31-casey-greene/
URL:https://datascience.utah.edu/talks/2022-08-31-casey-greene/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780 | https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4
 R0RiSUpsb3kvdz09
END:VEVENT
BEGIN:VEVENT
UID:2022-09-07-kevin-moon@datascience.utah.edu
DTSTAMP:20220907T163000Z
DTSTART;TZID=America/Denver:20220907T103000
DTEND;TZID=America/Denver:20220907T114500
SUMMARY:Scalable supervised manifold learning with random forests and neura
 l networks - Kevin Moon
DESCRIPTION:Kevin Moon (USU)\n\nThe manifold assumption has been used in ma
 ny machine learning applications to combat the curse of dimensionality. Mo
 st manifold learning methods are unsupervised and typically focus on prese
 rving the dominant structure and variation in the data. In many cases\, we
  wish to analyze the data in a supervised setting with respect to expert-p
 rovided data labels. Most supervised manifold learning methods exaggerate 
 the separation between data points of different classes\, distorting the t
 rue structure of the data. In this talk\, I will present RF-PHATE\, a supe
 rvised dimensionality reduction method that preserves the true structure o
 f the variables that are relevant for the supervised task. RF-PHATE is bas
 ed upon a diffusion process applied to random forest proximities and is we
 ll-suited for data visualization. I will then show how to improve the scal
 ability of RF-PHATE and any other manifold learning algorithm and perform 
 out of sample extension using geometry regularized autoencoders (GRAE).\n\
 nhttps://datascience.utah.edu/talks/2022-09-07-kevin-moon/
URL:https://datascience.utah.edu/talks/2022-09-07-kevin-moon/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780 | https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4
 R0RiSUpsb3kvdz09
CATEGORIES:deep learning,machine learning,visualization
END:VEVENT
BEGIN:VEVENT
UID:2022-09-14-elliot-smith@datascience.utah.edu
DTSTAMP:20220914T163000Z
DTSTART;TZID=America/Denver:20220914T103000
DTEND;TZID=America/Denver:20220914T114500
SUMMARY:Human neuronal population encoding of temporal difference learning 
 variables during risky choices - Elliot Smith
DESCRIPTION:Elliot Smith (Utah Neurology)\n\nRecent research in AI showed t
 hat agents designed to predict the full distribution of potential rewards\
 , rather than a central estimate of that distribution\, generate richer le
 arning distributions that allow them to perform better\, especially on ris
 ky tasks. Such distributional reinforcement learning (distRL) was also dis
 covered in dopamine neurons in the rodent ventral tegmental area. In this 
 nanosymposium presentation\, I will discuss recent work from direct brain 
 recordings in neurosurgical patients undergoing monitoring for treatment o
 f medically refractory epilepsy who performed a risky decision making task
  called the Balloon Analog Risk Task. Results from two studies will be pre
 sented: In the first study\, we examined neuronal population recordings (1
 57 neurons) from microelectrodes implanted in the anterior cingulate\, orb
 itofrontal and temporal cortices (15 participants)\, finding that human pr
 efrontal and mesial temporal neurons exhibited signatures of distRL: corre
 lated diverse optimism in reward coding and diverse asymmetric scaling of 
 reward prediction error. In the second study\, we examined correlations be
 tween broadband high frequency local field potentials (an established corr
 elate of population neuronal firing) and variables from temporal differenc
 e learning models for reward and risk while 37 participants made risky cho
 ices during BART. We found differences in which brain areas (3199 stereoel
 ectroencephalography or electrocorticography contacts sampling frontal\, t
 emporal\, and parietal lobes) encoded temporal difference learning model v
 ariables between participants who were more or less risk averse in their c
 hoices during BART. These areas included the left dorsolateral prefrontal\
 , anterior cingulate\, and orbitofrontal cortices. The results from these 
 studies shed light on the neural underpinnings of human value learning in 
 uncertain environments.\n\nhttps://datascience.utah.edu/talks/2022-09-14-e
 lliot-smith/
URL:https://datascience.utah.edu/talks/2022-09-14-elliot-smith/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780 | https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4
 R0RiSUpsb3kvdz09
CATEGORIES:biology & genomics,health & medicine,human-centered computing
END:VEVENT
BEGIN:VEVENT
UID:2022-09-21-jes-ford@datascience.utah.edu
DTSTAMP:20220921T163000Z
DTSTART;TZID=America/Denver:20220921T103000
DTEND;TZID=America/Denver:20220921T114500
SUMMARY:Model Review: Improving Transparency\, Reproducibility\, & Knowledg
 e Sharing using MLflow - Jes Ford
DESCRIPTION:Jes Ford (Cash App)\n\nCode Review is an integral part of softw
 are development\, but many teams don’t have similar processes in place f
 or the development and deployment of Machine Learning (ML) models. I will 
 motivate the decision to create a Model Review process\, starting from the
  principles of transparency\, reproducibility\, and knowledge sharing. MLf
 low is a useful Python package to help simplify and automate much of the t
 racking necessary to create detailed records of machine learning experimen
 ts. Much of this talk will be spent introducing this tool\, and demonstrat
 ing the core MLflow Tracking functionality. I’ll discuss how my team is 
 currently running a Model Review process for any ML models that we push to
  production\, and how we use MLflow to streamline this work and learn from
  each other.\n\nhttps://datascience.utah.edu/talks/2022-09-21-jes-ford/
URL:https://datascience.utah.edu/talks/2022-09-21-jes-ford/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780
CATEGORIES:machine learning
END:VEVENT
BEGIN:VEVENT
UID:2022-09-28-prashant-pandey@datascience.utah.edu
DTSTAMP:20220928T163000Z
DTSTART;TZID=America/Denver:20220928T103000
DTEND;TZID=America/Denver:20220928T114500
SUMMARY:Scalability Challenges in Large-Scale Sequence Search - Prashant Pa
 ndey
DESCRIPTION:Prashant Pandey (Utah SoC)\n\nSequence-level searches on large 
 collections of RNA sequencing experiments\, such as the NCBI Sequence Read
  Archive (SRA)\, would enable one to ask many questions about the expressi
 on or variation of a given transcript in a population. Building an efficie
 nt and scalable sequence search index at the scale of SRA data is a challe
 nging task and requires fundamental innovations in compression and scalabl
 e indexing. Recently\, several tools have been proposed to index and searc
 h through SRA data but they offer various trade-offs in terms of space\, s
 peed\, updatability\, and accuracy. In this talk\, I will present Mantis\,
  a fast\, exact\, and updatable sequence search index. Mantis uses recent 
 advancements in fast and compact hash tables\, domain-specific data compre
 ssion techniques\, and scalable indexing to build a scalable and updatable
  index and supports fast sequence searches on ~40K experiments (>100TB) in
  size from SRA.\n\nhttps://datascience.utah.edu/talks/2022-09-28-prashant-
 pandey/
URL:https://datascience.utah.edu/talks/2022-09-28-prashant-pandey/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780 | https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4
 R0RiSUpsb3kvdz09
CATEGORIES:biology & genomics,data management
END:VEVENT
BEGIN:VEVENT
UID:2022-10-05-jessica-shi@datascience.utah.edu
DTSTAMP:20221005T163000Z
DTSTART;TZID=America/Denver:20221005T103000
DTEND;TZID=America/Denver:20221005T114500
SUMMARY:Theoretically and Practically Efficient Parallel Nucleus Decomposit
 ion - Jessica Shi
DESCRIPTION:Jessica Shi (MIT)\n\nWe study the nucleus decomposition problem
 \, which has been shown to be useful in finding dense substructures in gra
 phs. We present a novel parallel algorithm that is efficient both in theor
 y and in practice. Our algorithm achieves a work complexity matching the b
 est sequential algorithm while also having low depth (parallel running tim
 e)\, which significantly improves upon the only existing parallel nucleus 
 decomposition algorithm (Sariyuce et al.\, PVLDB 2018). The key to the the
 oretical efficiency of our algorithm is a new lemma that bounds the amount
  of work done when peeling cliques from the graph\, combined with the use 
 of theoretically-efficient parallel algorithms for clique listing and buck
 eting.\n\nWe introduce several new practical optimizations\, including a n
 ew multi-level hash table structure to store information on cliques space-
 efficiently and a technique for traversing this structure cache-efficientl
 y. On a 30-core machine with two-way hyper-threading on real-world graphs\
 , we achieve up to a 55x speedup over the state-of-the-art parallel nucleu
 s decomposition algorithm by Sariyuce et al.\, and up to a 40x self-relati
 ve parallel speedup. We are able to efficiently compute larger nucleus dec
 ompositions than prior work on several million-scale graphs for the first 
 time.\n\nhttps://datascience.utah.edu/talks/2022-10-05-jessica-shi/
URL:https://datascience.utah.edu/talks/2022-10-05-jessica-shi/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780 | https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4
 R0RiSUpsb3kvdz09
CATEGORIES:algorithms & theory,education,networks & graphs
END:VEVENT
BEGIN:VEVENT
UID:2022-10-19-jie-zhang@datascience.utah.edu
DTSTAMP:20221019T163000Z
DTSTART;TZID=America/Denver:20221019T103000
DTEND;TZID=America/Denver:20221019T114500
SUMMARY:Active Sampling for Min-Max Fairness - Jie Zhang
DESCRIPTION:Jie Zhang (U Washington)\n\nModels satisfying Min-max fairness 
 minimizes maximum group specific losses\, so that the model has a more equ
 itable performance over all groups. Benefit of min-max fair models over ot
 her fairness notions and models include it levels up: meaning it only degr
 ades performance of a group if the degradation improves performance on the
  worst off group.\n\nIn this talk\, I will briefly mention the prior works
  that defined min-max fairness. I will mainly focus on our paper “Active
  Sampling for Min-Max Fairness” to introduce two algorithms that find mi
 n-max fair models with convergence guarantees.\n\nhttps://datascience.utah
 .edu/talks/2022-10-19-jie-zhang/
URL:https://datascience.utah.edu/talks/2022-10-19-jie-zhang/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4R0RiSUpsb3k
 vdz09
CATEGORIES:fairness & ethics,statistics
END:VEVENT
BEGIN:VEVENT
UID:2022-11-02-alex-chin@datascience.utah.edu
DTSTAMP:20221102T163000Z
DTSTART;TZID=America/Denver:20221102T103000
DTEND;TZID=America/Denver:20221102T114500
SUMMARY:Quantifying supply-demand imbalance in ridesharing systems - Alex C
 hin
DESCRIPTION:Alex Chin (Lyft)\n\nThe status of the rider and driver distribu
 tions and how they interact in a two-sided marketplace has implications fo
 r market efficiency and policy-making. As such\, accurately characterizing
  the supply-demand state of Lyft's marketplace is a crucial task. I will d
 escribe how this task can be framed in terms of an asymmetric optimal tran
 sport problem that yields a multi-resolution view of the supply-demand sta
 te and how it varies both spatially and temporally. I will then discuss ho
 w this approach can be incorporated into applications such as policy optim
 ization and machine learning prediction problems. Time permitting I will a
 lso discuss other science efforts we have at Lyft.\n\nhttps://datascience.
 utah.edu/talks/2022-11-02-alex-chin/
URL:https://datascience.utah.edu/talks/2022-11-02-alex-chin/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780 | https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4
 R0RiSUpsb3kvdz09
CATEGORIES:society & policy
END:VEVENT
BEGIN:VEVENT
UID:2022-11-09-shireen-elhabian@datascience.utah.edu
DTSTAMP:20221109T173000Z
DTSTART;TZID=America/Denver:20221109T103000
DTEND;TZID=America/Denver:20221109T114500
SUMMARY:Data-driven Shape Analysis: Methods\, Applications\, and Future - S
 hireen Elhabian
DESCRIPTION:Shireen Elhabian (Utah CS/SCI)\n\nQuantitative analysis of shap
 es is contingent upon defining a metric in the space of shapes to compare 
 shapes and perform shape statistics. A growing consensus in the field that
  such a metric should be adapted to the specific population under investig
 ation\, begging for learning such a metric in a data-driven manner. This t
 alk will cover a state-of-art data-driven approach for statistical shape m
 odeling that provides unbiased\, objective\, and intuitive evaluation of g
 eometric shapes\, emphasizing anatomical structures reconstructed from vol
 umetric images. I will talk about how we applied shape analysis to clinica
 l and scientific questions in various ways. I will also highlight the role
  of machine learning in mitigating critical bottlenecks and significant ba
 rriers to making shape modeling a robust tool for on-demand clinical diagn
 ostics and streamlining its adoption in research and practice. I will end 
 with a future outlook for shape modeling to enable more complex and divers
 e modeling scenarios.\n\nhttps://datascience.utah.edu/talks/2022-11-09-shi
 reen-elhabian/
URL:https://datascience.utah.edu/talks/2022-11-09-shireen-elhabian/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780 | https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4
 R0RiSUpsb3kvdz09
CATEGORIES:health & medicine,statistics
END:VEVENT
BEGIN:VEVENT
UID:2022-11-16-bernadette-stolz@datascience.utah.edu
DTSTAMP:20221116T173000Z
DTSTART;TZID=America/Denver:20221116T103000
DTEND;TZID=America/Denver:20221116T114500
SUMMARY:Applications of global and local persistent homology for the shape 
 of biological data - Bernadette Stolz
DESCRIPTION:Bernadette Stolz (Oxford)\n\nIn the first part of this talk\, I
  will showcase how persistent homology can be used to spatially characteri
 se structural abnormality in tumour blood vessel networks. More specifical
 ly\, I will show that the number of vessel loops and their distribution in
  these networks change over time when tumours undergo treatment with vascu
 lar targeting agents and radiation therapy. In the second part of the talk
 \, I will speak about applications of local persistent homology. I will sh
 ow how local persistent homology can be used to select landmarks from larg
 e and noisy data sets. In contrast to existing methods\, this subsampling 
 process is robust to outliers and is developed specifically for persistent
  homology. I will further introduce a novel method that can detect geometr
 ic anomalies\, such as intersections or boundaries\, in point cloud data s
 ampled from intersecting surfaces. This detection is based on the computat
 ion of persistent homology in local annular neighbourhoods around points a
 nd is less sensitive to the size of the local neighbourhood and surface cu
 rvature than an existing method.\n\nhttps://datascience.utah.edu/talks/202
 2-11-16-bernadette-stolz/
URL:https://datascience.utah.edu/talks/2022-11-16-bernadette-stolz/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4R0RiSUpsb3k
 vdz09
CATEGORIES:biology & genomics
END:VEVENT
BEGIN:VEVENT
UID:2022-11-30-aaron-quinlan@datascience.utah.edu
DTSTAMP:20221130T173000Z
DTSTART;TZID=America/Denver:20221130T103000
DTEND;TZID=America/Denver:20221130T114500
SUMMARY:Finding signals of genome mutation in the noise of DNA sequencing e
 rror - Aaron Quinlan
DESCRIPTION:Aaron Quinlan (Utah Human Genetics)\n\nThe research in our labo
 ratory is focused on the application of computational methods to develop a
  deeper understanding of genetic variation in diverse contexts. Modern exp
 erimental methods allow us to examine entire genomes with exquisite detail
 . Perhaps not surprisingly\, staggering complexity is revealed as we look 
 more closely at how genetic variation (both inherited and somatic) contrib
 utes to phenotypes. Modern genomic technologies necessitate efficient appr
 oaches for exploring\, manipulating and comparing large genomic datasets. 
 We develop such methods so that we and others may apply them to experiment
 s investigating the impact of genetic variation on human disease\, evoluti
 on\, and somatic differentiation. Genome research is difficult - we strive
  to develop computational means that make it easier.\n\nhttps://datascienc
 e.utah.edu/talks/2022-11-30-aaron-quinlan/
URL:https://datascience.utah.edu/talks/2022-11-30-aaron-quinlan/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780 | https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4
 R0RiSUpsb3kvdz09
CATEGORIES:biology & genomics
END:VEVENT
BEGIN:VEVENT
UID:2022-12-07-tao-yang@datascience.utah.edu
DTSTAMP:20221207T173000Z
DTSTART;TZID=America/Denver:20221207T103000
DTEND;TZID=America/Denver:20221207T114500
SUMMARY:Optimizing Ranking Effectiveness and Fairness - Tao Yang
DESCRIPTION:Tao Yang (Utah SoC)\n\nAdvanced ranking techniques have led to 
 improvements in AI-powered information services that significantly changed
  people's lives. For example\, search engines that rank information accord
 ing to their utilities to use's queries have helped billions of people bet
 ter finish their tasks in daily work\; recommendation systems that rank pr
 oducts/movies/news according to the user's interests have completely chang
 ed the way people discover information everyday. Therefore\, how to constr
 uct and optimize ranking systems is one of the most important research pro
 blems in the field of Information Retrieval (IR). When optimizing ranking 
 systems\, there are two important criteria to measure the quality of resul
 t rankings in IR systems. The first criterion is ranking effectiveness\, w
 hich refers to the ability of a ranking system to effectively present resu
 lts based on their relevance to the users' needs. The second criterion is 
 ranking fairness\, which refers to the ability of a ranking system to pres
 ent results fairly. For example\, in job recommendation\, if a ranking sys
 tem only considers ranking effectiveness and ranks items solely according 
 to relevance\, a small number of top candidates will always be exposed to 
 users and dominate users' attention as users usually only examine the top 
 ranks. In such case\, other candidates will rarely have the chance to be h
 ired even when they are highly qualified for the job. Therefore\, it is im
 portant to balance the effectiveness of ranked lists with the fairness in 
 ranking optimization.\nIn this talk\, I will present my recent works on ra
 nking effectiveness and fairness optimization. The talk will be two parts.
  In the first part of this talk\, I will introduce works on sole-effective
 ness optimization where I propose uncertainty-aware rank systems based on 
 Bayes modelling. In the second part of this talk\, I will introduce works 
 on fairness-effectiveness joint optimization.\n\nhttps://datascience.utah.
 edu/talks/2022-12-07-tao-yang/
URL:https://datascience.utah.edu/talks/2022-12-07-tao-yang/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780 | https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4
 R0RiSUpsb3kvdz09
CATEGORIES:fairness & ethics,optimization
END:VEVENT
BEGIN:VEVENT
UID:2023-01-11-george-vega-yon@datascience.utah.edu
DTSTAMP:20230111T174500Z
DTSTART;TZID=America/Denver:20230111T104500
DTEND;TZID=America/Denver:20230111T114500
SUMMARY:Prediction of Gene Functions by Leveraging Biological Insights with
  Mechanistic Machine Learning - George Vega Yon
DESCRIPTION:George Vega Yon (Utah Epidemiology)\n\nBiomedical sciences\, in
  particular\, bioinformaticians and computational biologists\, are in a ra
 ce to annotate the immense number of genes and gene products we are still 
 learning from. In this talk\, I present a new method for predicting gene f
 unctions using mechanistic machine learning. Mechanistic machine learning 
 is a relatively new area of research where predictive ML-based algorithms 
 are improved by incorporating domain knowledge via mechanistic models. Her
 e\, we use a theoretically-funded function evolution model that relies sol
 ely on phylogenetic trees from PantherDB and annotations from the Gene Ont
 ology (GO) to make high-quality predictions. I will illustrate how combini
 ng the mentioned model with a large gene expression database (Bgee) in an 
 ML model significantly improves prediction quality.\n\nhttps://datascience
 .utah.edu/talks/2023-01-11-george-vega-yon/
URL:https://datascience.utah.edu/talks/2023-01-11-george-vega-yon/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 | https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4
 R0RiSUpsb3kvdz09
CATEGORIES:biology & genomics,machine learning,networks & graphs,statistics
END:VEVENT
BEGIN:VEVENT
UID:2023-01-18-echo-warner@datascience.utah.edu
DTSTAMP:20230118T174500Z
DTSTART;TZID=America/Denver:20230118T104500
DTEND;TZID=America/Denver:20230118T120000
SUMMARY:College of Nursing\, Division of Acute and Chronic Care - Echo Warn
 er
DESCRIPTION:Echo Warner (Utah Nursing\, HCI)\n\nUnproven health claims on t
 he internet may have substantial influence on patient behaviors and decisi
 on making. For example\, cancer patients with curable disease who pursue u
 nproven cancer treatment in lieu of evidence-based approaches demonstrate 
 2-4 times higher mortality than patients who avoid unproven cancer treatme
 nt. Interventions to mitigate the impact of online health misinformation a
 re desperately needed. The few interventions that have been tested do not 
 consider the extent of exposure to misinformation online\, primarily becau
 se individual estimates of exposure are based on self-report and are consi
 dered highly unreliable. Web-monitoring software is typically used by busi
 nesses to monitor remote employee productivity. Our paradigm shifting appr
 oach applies web-monitoring to quantify online cancer misinformation expos
 ure. We will discuss the feasibility of using web-monitoring software to q
 uantify exposure to online health information with a special focus on canc
 er symptom management and unproven cancer treatment misinformation. Specif
 ically\, we will review 1) characteristics of online cancer health misinfo
 rmation and the impacts this exposure may have on cancer patient health ou
 tcomes\, relationships\, and finances 2) Ethical considerations of web-mon
 itoring\, and 3) methodological rigor and reproducibility of web-monitorin
 g approaches for studying exposure to other types of health misinformation
  online.\n\nhttps://datascience.utah.edu/talks/2023-01-18-echo-warner/
URL:https://datascience.utah.edu/talks/2023-01-18-echo-warner/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:Virtual: | https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4
 R0RiSUpsb3kvdz09
CATEGORIES:health & medicine,society & policy
END:VEVENT
BEGIN:VEVENT
UID:2023-01-25-shweta-jain@datascience.utah.edu
DTSTAMP:20230125T174500Z
DTSTART;TZID=America/Denver:20230125T104500
DTEND;TZID=America/Denver:20230125T114500
SUMMARY:Putting Parameterization into Practice - Shweta Jain
DESCRIPTION:Shweta Jain (Utah SoC)\n\nGraphs are everywhere: social network
 s\, protein interaction networks\, citation networks\, epidemic spread net
 works. Applications that use these graphs have progressively become more s
 ophisticated\, relying on getting fast and accurate solutions to graph-the
 oretic problems\, many of which are NP-Hard. However\, the sizes of today'
 s graphs easily run into millions of vertices and edges\, if not more. As 
 a result\, many classic algorithms are infeasible for such graphs.\n\nFort
 unately\, real-world graphs across different domains show a lot of common 
 characteristics such as an abundance of triangles\, low average distance b
 etween vertices (small-world property)\, low degeneracy (a measure of the 
 sparsity of edges) etc. which we can leverage to design algorithms that ar
 e provably efficient given those parameters. I will demonstrate this in th
 e context of clique counting and decomposition of graphs\, which has appli
 cations in community detection\, spam detection\, fraud detection\, gene m
 odule detection etc. The new algorithms not only massively improve the per
 formance in practice but also improve our theoretical understanding of rea
 l-world graphs and help to bridge the gap between theory and practice.\n\n
 https://datascience.utah.edu/talks/2023-01-25-shweta-jain/
URL:https://datascience.utah.edu/talks/2023-01-25-shweta-jain/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 | https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4
 R0RiSUpsb3kvdz09
CATEGORIES:algorithms & theory,biology & genomics,networks & graphs
END:VEVENT
BEGIN:VEVENT
UID:2023-02-01-ana-marsovic@datascience.utah.edu
DTSTAMP:20230201T174500Z
DTSTART;TZID=America/Denver:20230201T104500
DTEND;TZID=America/Denver:20230201T114500
SUMMARY:“AI” That Masters Language Could Reason about Negation - Ana Ma
 rsovic
DESCRIPTION:Ana Marsovic (Utah SoC)\n\nWe have experienced firsthand the gr
 owing impact that AI technologies like ChatGPT have. It is agreed upon tha
 t risks involving such technologies must be managed. This is at a glance a
 kin to how people handle safety-critical systems in\, e.g.\, aviation. How
 ever\, the rigorous principles of safety engineering are not easily applic
 able to AI technologies. AI-backed solutions are obscured in AI’s intern
 als\, and the exact requirements for AI safety are unverifiable since they
  are neither defined nor regulated. One technique that has emerged as a st
 ep to measure an aspect of AI safety is to construct data that represents 
 a specific phenomenon that a safe model must handle well\, and that has no
  spurious correlations. Low performance on the dataset is undesired.In thi
 s talk\, I’ll focus on one common linguistic phenomenon\, negation\, wit
 hout which the full power of human language-based communication cannot be 
 realized. I will show how we carefully constructed a question-answering da
 taset\, CondaQA\, to study how well current models (described by the New Y
 ork Time Magazine as “mastering language”) reason about negated statem
 ents. An InstructGPT model (the latest GPT model before Nov 28\, 2022) com
 bined with chain-of-thought prompting achieves 66.28% accuracy and 27.28% 
 consistency on CondaQA\, way behind human accuracy of 91.94% and consisten
 cy of 81.58%.\n\nhttps://datascience.utah.edu/talks/2023-02-01-ana-marsovi
 c/
URL:https://datascience.utah.edu/talks/2023-02-01-ana-marsovic/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 | https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4
 R0RiSUpsb3kvdz09
CATEGORIES:large language models,natural language processing
END:VEVENT
BEGIN:VEVENT
UID:2023-02-08-aaron-clauset@datascience.utah.edu
DTSTAMP:20230208T174500Z
DTSTART;TZID=America/Denver:20230208T104500
DTEND;TZID=America/Denver:20230208T114500
SUMMARY:Meritocracy or systemic bias? Untangling the drivers the productivi
 ty and prominence among scientists - Aaron Clauset
DESCRIPTION:Aaron Clauset (UC Boulder)\n\nSimple measures of scholarly prod
 uctivity and prominence vary enormously across both individual scientists 
 and institutions -- but to what degree do these inequalities represent gen
 uine meritocratic differences vs. systemic biases that limit scientific pr
 ogress?\n\nIn this talk\, I'll describe a sequence of results that substan
 tially untangle the underlying systemic drivers of productivity and promin
 ence among scientists. First\, I'll show that productivity and prominence 
 are\, to a significant degree\, environmental variables such that the pres
 tige of a scientist's working environment drives their individual producti
 vity\, largely by providing larger research groups to elite scientists. Se
 cond\, I'll describe a network-based generative model of individual produc
 tivity and prominence that untangles these measures from their underlying 
 collaboration networks. These models corroborate the labor-advantage hypot
 hesis of elite institutions\, and also reveal both that gendered differenc
 es in the productivity and prominence of mid-career researchers can be lar
 gely explained by gendered differences in coauthorship networks\, and that
  these networks are partially transferable from senior to junior collabora
 tors. Hence\, collaboration networks\, and the systemic factors that shape
  them\, play a critical role in driving scholarly inequalities in science\
 , and suggest that these networks are an important form of unequally distr
 ibuted social capital that shapes who makes what scientific discoveries. I
 'll close with a discussion of policies that could potentially mitigate th
 e unequal distribution of this social capital and help both diversify the 
 academy and broaden its contributions to society.\n\nhttps://datascience.u
 tah.edu/talks/2023-02-08-aaron-clauset/
URL:https://datascience.utah.edu/talks/2023-02-08-aaron-clauset/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 | https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4
 R0RiSUpsb3kvdz09
CATEGORIES:fairness & ethics,networks & graphs,society & policy
END:VEVENT
BEGIN:VEVENT
UID:2023-02-15-orly-alter@datascience.utah.edu
DTSTAMP:20230215T174500Z
DTSTART;TZID=America/Denver:20230215T104500
DTEND;TZID=America/Denver:20230215T114500
SUMMARY:Solving Cancer with Data: Mathematical Discovery and Computational 
 and Experimental Validation of Whole-Genome Genotype–Survival and Respon
 se to Treatment Phenotype Relationships in Cancer - Orly Alter
DESCRIPTION:Orly Alter (Utah BME\, SCI)\n\n1/2 of men and 1/3 of women will
  face cancer\, a disease of the whole 3B-nucleotide genome. But\, despite 
 the availability of open-source data and the $100/1-hour genome\, genetic 
 tests remain limited to one to a few hundred genes. Therefore\, the progno
 sis\, diagnosis\, and treatment of cancer remain unchanged. This is due to
  the lack of suitable AI/ML. I will describe work in my lab inventing AI/M
 L that connects the whole genome with a patient’s survival and response 
 to treatment. Our algorithms discover accurate\, precise\, and interpretab
 le predictors\, applicable to the general population\, from as few as 50
 –100 patients. Our predictors outperform all other indicators\, where th
 ey exist. All other methods miss them. I will describe my international re
 trospective clinical trial\, which validated a genome-wide pattern in tumo
 rs from glioblastoma brain cancer patients as the best predictor of life e
 xpectancy and response to standard of care. We discovered this\, and predi
 ctors in\, e.g.\, adult lung\, ovarian\, and uterine adenocarcinoma tumors
  and pediatric nerve neuroblastoma tumors\, in public data\, proving that 
 the algorithms and predictors are uniquely suited to personalized medicine
 . I will also describe work translating the algorithms and predictors to t
 he clinic.\n\nhttps://datascience.utah.edu/talks/2023-02-15-orly-alter/
URL:https://datascience.utah.edu/talks/2023-02-15-orly-alter/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780 | https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4
 R0RiSUpsb3kvdz09
CATEGORIES:algorithms & theory,biology & genomics,health & medicine
END:VEVENT
BEGIN:VEVENT
UID:2023-02-22-vivek-gupta@datascience.utah.edu
DTSTAMP:20230222T174500Z
DTSTART;TZID=America/Denver:20230222T104500
DTEND;TZID=America/Denver:20230222T114500
SUMMARY:Inference and Reasoning for Semi-structured Tables - Vivek Gupta
DESCRIPTION:Vivek Gupta (Utah SoC)\n\nUnderstanding semi-structured tabular
  data\, which is ubiquitous in the real world\, requires an understanding 
 of the meaning of text fragments and the implicit connections between them
 . We believe such data could be used to investigate how individuals and ma
 chines reason about semi-structured data. First\, we present the InfoTabS 
 dataset\, which consists of human-written textual predictions based on tab
 les collected from Wikipedia's infoboxes. Our research demonstrates that t
 he semi-structured\, multi-domain\, and heterogeneous nature of the premis
 es prompts complicated\, multi-faceted reasoning\, offering a modeling cha
 llenge for traditional modeling techniques. Second\, we analyzed these cha
 llenges in-depth and developed simple\, effective preprocessing strategies
  to overcome them. Thirdly\, despite accurate NLI prediction\, we demonstr
 ate through rigorous probing that the existing model does not reason with 
 the provided tabular facts. To address this\, we suggest a two-stage evide
 nce extraction and tabular inference technique for enhancing model reasoni
 ng and interpretability. We also investigate efficient methods for enhanci
 ng tabular inference datasets with semi-automatic data augmentation and pa
 ttern-based pre-training. Lastly\, to ensure that tabular reasoning models
  work in more than one language\, we introduce XInfoTabS\, a unique proble
 m of bilingual tabular inference\, and a cost-effective pipeline for trans
 lating tables. In the near future\, we plan to test the tabular reasoning 
 model for temporal changes\, especially for dynamic tables where informati
 on changes over time.\n\nhttps://datascience.utah.edu/talks/2023-02-22-viv
 ek-gupta/
URL:https://datascience.utah.edu/talks/2023-02-22-vivek-gupta/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 | https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4
 R0RiSUpsb3kvdz09
CATEGORIES:data management,machine learning,natural language processing
END:VEVENT
BEGIN:VEVENT
UID:2023-03-01-emily-hadley@datascience.utah.edu
DTSTAMP:20230301T174500Z
DTSTART;TZID=America/Denver:20230301T104500
DTEND;TZID=America/Denver:20230301T114500
SUMMARY:Applied Strategies for Advancing Racial Equity and Addressing Bias 
 in Big Data Research - Emily Hadley
DESCRIPTION:Emily Hadley (RTI International)\n\nIn the last decade\, big da
 ta research studies have proliferated and\, in some cases\, offered consid
 erable promise for humanity. Yet\, numerous incidents have documented that
  without safeguards\, this same research can reproduce and amplify existin
 g societal biases and disparities. Given the increased public awareness of
  structural racism\, bias based on race and ethnicity in big data research
  is particularly concerning. Big data researchers can mitigate the risk of
  perpetuating bias based on race and ethnicity by intentionally incorporat
 ing best practices that reduce or eliminate racial biases. We synthesize k
 ey findings and applied recommendations from over 140 sources for addressi
 ng race and ethnicity bias in big data research. We discuss considerations
  when planning a big data project\, including identifying study motivation
 s\, centering participatory involvement\, and addressing key concerns rega
 rding race and ethnicity in big data collection and quality. We detail iss
 ues related to proxy discrimination\, algorithmic audits\, data completene
 ss\, and the Big Data Paradox. We provide real-world examples of advancing
  racial equity and addressing bias and share recommendations for additiona
 l resources and opportunities for further investigation.\n\nhttps://datasc
 ience.utah.edu/talks/2023-03-01-emily-hadley/
URL:https://datascience.utah.edu/talks/2023-03-01-emily-hadley/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 | https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4
 R0RiSUpsb3kvdz09
CATEGORIES:fairness & ethics,society & policy
END:VEVENT
BEGIN:VEVENT
UID:2023-03-15-nate-veldt@datascience.utah.edu
DTSTAMP:20230315T164500Z
DTSTART;TZID=America/Denver:20230315T104500
DTEND;TZID=America/Denver:20230315T114500
SUMMARY:Measuring homophily in group interactions: hypergraph models and co
 mbinatorial impossibilities - Nate Veldt
DESCRIPTION:Nate Veldt (Texas A&M)\n\nHomophily is the well-known sociologi
 cal principle that people tend to connect with others who are similar to t
 hem\, or more informally: "birds of a feather flock together." Although ma
 ny social interactions occur in groups\, homophily is typically measured u
 sing a graph\, which only accounts for interactions involving two individu
 als. This talk will present a new hypergraph framework for more directly m
 easuring homophily in group settings. Our measures highlight natural patte
 rns in group homophily that appear with gender in scientific collaboration
  and political affiliation in legislative bill cosponsorship\, and also re
 veal distinctive gender distributions in group photographs\, all of which 
 cannot be fully captured by graph-based measures. We will also discuss sub
 tle combinatorial limits and impossibilities that arise when measuring hom
 ophily in hypergraphs\, which are completely independent of human behavior
  and must be properly accounted for in order to understand how homophily c
 an (and cannot) be manifested in group interactions.\n\nhttps://datascienc
 e.utah.edu/talks/2023-03-15-nate-veldt/
URL:https://datascience.utah.edu/talks/2023-03-15-nate-veldt/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 | https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4
 R0RiSUpsb3kvdz09
CATEGORIES:networks & graphs
END:VEVENT
BEGIN:VEVENT
UID:2023-03-22-bailey-fosdick@datascience.utah.edu
DTSTAMP:20230322T164500Z
DTSTART;TZID=America/Denver:20230322T104500
DTEND;TZID=America/Denver:20230322T114500
SUMMARY:Modeling Infection Fatality Rates to Assess the Burden of COVID-19 
 in Developing Countries - Bailey Fosdick
DESCRIPTION:Bailey Fosdick (CU Anschutz)\n\nCOVID-19 spread quickly around 
 the world after first being discovered in China in late 2019. It has had d
 evastating impacts\, however its impacts\, both in terms of infection prev
 alence and fatalities\, have been non-uniformly distributed worldwide. Whi
 le early studies focused on COVID-19 infection and fatality rates in high-
 income countries\, much less attention has been given to the impacts of CO
 VID-19 in developing countries. In this work\, we systematically reviewed 
 the literature to identify all COVID-19 serology studies conducted by earl
 y 2021 using population representative samples. We developed a Bayesian hi
 erarchical model for simultaneously modeling serology and death data to ma
 ke inference on age-specific seroprevalence and age-specific infection fat
 ality rates. This model directly accounts for conventional sampling uncert
 ainty\, as well as uncertainty about the serological test assay sensitivit
 y and specificity. Through a careful analysis of data from over thirty dev
 eloping countries\, we found seroprevalence in many developing country loc
 ations was markedly higher than in high-income countries early in the pand
 emic and age-specific infection fatality rates were roughly twice as high 
 as that in high-income countries.\n\nhttps://datascience.utah.edu/talks/20
 23-03-22-bailey-fosdick/
URL:https://datascience.utah.edu/talks/2023-03-22-bailey-fosdick/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 | https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4
 R0RiSUpsb3kvdz09
CATEGORIES:statistics
END:VEVENT
BEGIN:VEVENT
UID:2023-03-29-justin-baker@datascience.utah.edu
DTSTAMP:20230329T164500Z
DTSTART;TZID=America/Denver:20230329T104500
DTEND;TZID=America/Denver:20230329T114500
SUMMARY:Monotone Implicit Graph Neural Networks for Long-Range Dependency L
 earning - Justin Baker
DESCRIPTION:Justin Baker (Utah Math & SCI)\n\nFrom social networks to chemi
 cal engineering\, deep graph neural networks play an instrumental role in 
 advancing our industrial and scientific frontier. Of particular interest a
 re networks which can learn long range dependencies in a scalable and expr
 essive manner. In this talk\, we will delve into the power of deep learnin
 g on graphs\, with a focus on implicit graph neural networks (IGNNs) and t
 heir scalability. We will also discuss how monotone operator theory enhanc
 es the expressivity of IGNNs\, overcoming a crucial obstacle to learning l
 ong-range dependencies. By doing so\, monotone IGNNs can significantly imp
 rove graph learning and have the potential to make breakthroughs in variou
 s fields.\n\nhttps://datascience.utah.edu/talks/2023-03-29-justin-baker/
URL:https://datascience.utah.edu/talks/2023-03-29-justin-baker/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4R0RiSUpsb3k
 vdz09
CATEGORIES:deep learning,networks & graphs
END:VEVENT
BEGIN:VEVENT
UID:2023-04-05-yao-yaun-mao@datascience.utah.edu
DTSTAMP:20230405T164500Z
DTSTART;TZID=America/Denver:20230405T104500
DTEND;TZID=America/Denver:20230405T114500
SUMMARY:Searching for dwarf (small) galaxies in astronomical surveys - Yao-
 Yaun Mao
DESCRIPTION:Yao-Yaun Mao (Utah Astro)\n\nDwarf galaxies are small fuzzy gal
 axies that contain much fewer stars than Milky Way-mass galaxies. Observat
 ions of these little galaxies can enhance our understanding of galaxy form
 ation and the nature of dark matter. Finding these dwarf galaxies is\, how
 ever\, not an easy task because they are faint and dim from our perspectiv
 e. I will describe the challenges and recent efforts on the search for nea
 rby dwarf galaxies in astronomical surveys. One particular challenge is to
  identify potential dwarf galaxies with only image data that do not contai
 n distance information. I will discuss a few traditional methods that are 
 used to identify dwarf galaxies and obtain their astronomical distances\, 
 and why these methods are mostly used to find dwarf galaxies in specific p
 atches of the sky. I will then discuss how the distance information can be
  used as training data in machine learning algorithms\, such as convolutio
 nal neural networks\, to derive distance information from just astronomica
 l images. Finally\, I will discuss the remaining challenges we have\, and 
 how we may improve the methods in preparation for future observations from
  the Rubin Observatory Legacy Survey of Space and Time (LSST) and Roman Sp
 ace Telescope.\n\nhttps://datascience.utah.edu/talks/2023-04-05-yao-yaun-m
 ao/
URL:https://datascience.utah.edu/talks/2023-04-05-yao-yaun-mao/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 | https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4
 R0RiSUpsb3kvdz09
CATEGORIES:deep learning,machine learning,physics & astronomy
END:VEVENT
BEGIN:VEVENT
UID:2023-04-12-titus-brown@datascience.utah.edu
DTSTAMP:20230412T164500Z
DTSTART;TZID=America/Denver:20230412T104500
DTEND;TZID=America/Denver:20230412T114500
SUMMARY:Is everything everywhere all at once? Asking questions of all publi
 c microbiome shotgun data - Titus Brown
DESCRIPTION:Titus Brown (UC Davis)\n\nPublic sequence data offers many oppo
 rtunities for reuse\, exploration\, and discovery. What happens if you mak
 e it really\, really easy to search the content of all the public microbio
 me data sets? It turns out you can enable some interesting science\, but y
 ou also need to address many technical\, social\, and policy issues. In th
 is talk I’ll showcase some of our results from being able to search ever
 ything\, everywhere\, all at once\; describe our current efforts\; and dis
 cuss the possible opportunities and challenges of petabyte-scale sequence 
 search.\n\nhttps://datascience.utah.edu/talks/2023-04-12-titus-brown/
URL:https://datascience.utah.edu/talks/2023-04-12-titus-brown/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 | https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4
 R0RiSUpsb3kvdz09
CATEGORIES:health & medicine,society & policy,statistics
END:VEVENT
BEGIN:VEVENT
UID:2023-04-19-sumana-basu@datascience.utah.edu
DTSTAMP:20230419T164500Z
DTSTART;TZID=America/Denver:20230419T104500
DTEND;TZID=America/Denver:20230419T114500
SUMMARY:Towards Reinforcement Learning for Precision Drug Dosing - Sumana B
 asu
DESCRIPTION:Sumana Basu (McGill University)\n\nDrug dosing is an important 
 application of AI\, which can be formulated as a Reinforcement Learning (R
 L) problem\, since every individual’s drug dosing requirement is differe
 nt. In this talk\, we will talk about two major challenges of using RL for
  drug dosing: delayed and prolonged effects of medications\, which break t
 he Markov assumption of the RL framework. We will talk about an approach t
 o solve this problem in a model free reinforcement learning setting\, talk
  further about the challenges of deploying it in real life and sketch the 
 outline of a more realistic Model Based Reinforcement Learning (MBRL) solu
 tion to it.\n\nhttps://datascience.utah.edu/talks/2023-04-19-sumana-basu/
URL:https://datascience.utah.edu/talks/2023-04-19-sumana-basu/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://utah.zoom.us/j/92180148411?pwd=dG1OaXlZTTQ1d0M4R0RiSUpsb3k
 vdz09
CATEGORIES:health & medicine,machine learning
END:VEVENT
BEGIN:VEVENT
UID:2023-05-30-bodhisattwa-majumder@datascience.utah.edu
DTSTAMP:20230530T210000Z
DTSTART;TZID=America/Denver:20230530T150000
DTEND;TZID=America/Denver:20230530T163000
SUMMARY:User-centric Natural Language Processing - Bodhisattwa Majumder
DESCRIPTION:Bodhisattwa Majumder (UCSD)\n\nArtificial intelligence (AI) has
  shown remarkable effectiveness in knowledge-seeking applications (e.g.\, 
 for recommendations and explanations). However\, the increasing expectatio
 n of more trust\, accessibility\, and anthropomorphism in these AI systems
  requires the underlying components (dialog models\, LLMs\, classifiers) t
 o be adaptive and adequately knowledge grounded. In reality\, the outputs 
 of the constituent models often lack commonsense\, explanations\, and subj
 ectivity\, which motivates us to ask the question: what can we achieve by 
 redesigning AI systems to start with individual needs?\n\nIdeally\, an ass
 istive AI system must be aware of the surrounding world\, produce faithful
  explanations\, and align with the user's preferences. In this talk\, I wi
 ll discuss a post-hoc knowledge-injection technique that enriches the dial
 og responses at the decoding time and promotes achieving conversational go
 als. Then\, I will explore how to elevate existing AI systems using a unif
 ied framework to map low-level and abstractive explanations by background 
 knowledge. Finally\, I will hint at a user-centric interventionist approac
 h that can help users obtain more equitable predictions backed by faithful
  explanations as compared to a black-box counterpart. I will conclude with
  the future possibilities and societal impacts of next-generation user-cen
 tric systems.\n\nbodhi_flyer.png\n\nhttps://datascience.utah.edu/talks/202
 3-05-30-bodhisattwa-majumder/
URL:https://datascience.utah.edu/talks/2023-05-30-bodhisattwa-majumder/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:Where: MEB 3147 Large Conference Room Simcast:  (Meeting ID: 965 3
 805 0936\, Passcode: 404653) | https://utah.zoom.us/j/96538050936
CATEGORIES:fairness & ethics,natural language processing
END:VEVENT
BEGIN:VEVENT
UID:2023-08-30-blair-sullivan@datascience.utah.edu
DTSTAMP:20230830T163000Z
DTSTART;TZID=America/Denver:20230830T103000
DTEND;TZID=America/Denver:20230830T114500
SUMMARY:Improving Fairness of Information Access in Networks - Blair Sulliv
 an
DESCRIPTION:Blair Sullivan (Utah)\n\nIn social networks\, node position is 
 a form of social capital which enables faster and more reliable access to 
 diverse information. Structural biases often arise from network formation 
 and can lead to significant disparities in information access based on pos
 ition. We discuss ways to quantify this social capital through the lens of
  information flow in the network\, focusing on the setting where each node
  may be a source of distinct desirable information. We define several meas
 ures of access advantage\, and consider the problem of improving equity by
  making interventions in the network\, focusing on the case of edge augmen
 tation. We describe several heuristic strategies for budgeted intervention
 \, and present the results of an empirical evaluation on a corpus of real-
 world networks.\n\nhttps://datascience.utah.edu/talks/2023-08-30-blair-sul
 livan/
URL:https://datascience.utah.edu/talks/2023-08-30-blair-sullivan/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:FASB 295
CATEGORIES:algorithms & theory,fairness & ethics,networks & graphs
END:VEVENT
BEGIN:VEVENT
UID:2023-09-06-anna-fariha@datascience.utah.edu
DTSTAMP:20230906T163000Z
DTSTART;TZID=America/Denver:20230906T103000
DTEND;TZID=America/Denver:20230906T114500
SUMMARY:Blame the data\, not the system: how data constraints can help in t
 rustworthy machine learning and explain causes of data-system malfunction 
 - Anna Fariha
DESCRIPTION:Anna Fariha (Utah)\n\nThe core of modern data-driven systems co
 mprises models learned from large datasets\, and they are usually optimize
 d to target particular data and workloads. While these data-driven systems
  have seen wide adoption and success\, their reliability and proper functi
 on hinge on the data's continued conformance to the systems initial settin
 gs and assumptions. My research focuses on designing mechanisms to assess 
 the trustworthiness of a system's inferences and explain causes of system 
 malfunction due to data nonconformance. The key idea here is that since da
 ta is central to data-driven systems\, it can guide us to determine whethe
 r predictions made by an ML model can be trusted\, and to expose the cause
  of a system's unexpected behavior. In this talk\, I will talk about mecha
 nisms and explanation frameworks to facilitate trusting and understanding 
 outcomes involving data and data systems.\n\nhttps://datascience.utah.edu/
 talks/2023-09-06-anna-fariha/
URL:https://datascience.utah.edu/talks/2023-09-06-anna-fariha/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:FASB 295
CATEGORIES:machine learning
END:VEVENT
BEGIN:VEVENT
UID:2023-09-13-david-tench@datascience.utah.edu
DTSTAMP:20230913T163000Z
DTSTART;TZID=America/Denver:20230913T103000
DTEND;TZID=America/Denver:20230913T114500
SUMMARY:Dynamic Graph Sketching: To Infinity And Beyond - David Tench
DESCRIPTION:David Tench (Berkeley Lab)\n\nExisting graph stream processing 
 systems must store the graph explicitly in RAM which limits the scale of g
 raphs they can process. The graph semi-streaming literature offers algorit
 hms which avoid this limitation via linear sketching data structures that 
 use small (sublinear) space\, but these algorithms have not seen use in pr
 actice to date. In this talk I will explore what is needed to make graph s
 ketching algorithms practically useful\, and as a case study present a ske
 tching algorithm for connected components and a corresponding high-perform
 ance implementation. Finally\, I will give an overview of the many open pr
 oblems in this area\, focusing on improving query performance of graph ske
 tching algorithms.\n\nhttps://datascience.utah.edu/talks/2023-09-13-david-
 tench/
URL:https://datascience.utah.edu/talks/2023-09-13-david-tench/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:FASB 295
CATEGORIES:algorithms & theory,networks & graphs
END:VEVENT
BEGIN:VEVENT
UID:2023-09-20-el-kindi-rezig@datascience.utah.edu
DTSTAMP:20230920T163000Z
DTSTART;TZID=America/Denver:20230920T103000
DTEND;TZID=America/Denver:20230920T114500
SUMMARY:Data Preparation: The Biggest Roadblock in Data Science - El Kindi 
 Rezig
DESCRIPTION:El Kindi Rezig (Utah)\n\nWhen building Machine learning (ML) mo
 dels\, data scientists face a significant hurdle: data preparation. ML mod
 els are exactly as good as the data we train them on. Unfortunately\, data
  preparation is tedious and laborious because it often requires human judg
 ment on how to proceed. In fact\, data scientists spend at least 80% of th
 eir time locating the datasets they want to analyze\, integrating them tog
 ether\, and cleaning the result.In this talk\, I will present my key contr
 ibutions in data preparation for data science\, which address the followin
 g problems: (1) data discovery: how to discover data of interest from a la
 rge collection of heterogeneous tables (e.g.\, data lakes)\; (2) error det
 ection: how to find errors in the input and intermediate data in complex d
 ata workflows\; and (3) data repairing: how to repair data errors with min
 imal human intervention. The developed systems are specifically designed t
 o support data science development which poses particular requirements suc
 h as interactivity and modularity.\n\nhttps://datascience.utah.edu/talks/2
 023-09-20-el-kindi-rezig/
URL:https://datascience.utah.edu/talks/2023-09-20-el-kindi-rezig/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:FASB 295
CATEGORIES:data management
END:VEVENT
BEGIN:VEVENT
UID:2023-09-27-parikshit-gopalan@datascience.utah.edu
DTSTAMP:20230927T163000Z
DTSTART;TZID=America/Denver:20230927T103000
DTEND;TZID=America/Denver:20230927T114500
SUMMARY:Loss Minimization and Multi-group Fairness - Parikshit Gopalan
DESCRIPTION:Parikshit Gopalan (Apple Research)\n\nTraining a predictor to m
 inimize a loss function fixed in advance is the dominant paradigm in machi
 ne learning. However\, loss minimization by itself might not guarantee des
 iderata like fairness and accuracy that one could reasonably expect from a
  predictor. In contrast\, various group-fairness notions have been propsoe
 d that constrain the predictor to share certain statistical properties of 
 the data\, even when conditioned on a rich family of subgroups. There is n
 o explicit attempt at loss minimization.In this talk\, we will explore som
 e recently discovered connections between loss minimization and notions of
  multi-group fairness. We will see settings where one can lead to the othe
 r\, and other settings where this is unlikely.\n\nhttps://datascience.utah
 .edu/talks/2023-09-27-parikshit-gopalan/
URL:https://datascience.utah.edu/talks/2023-09-27-parikshit-gopalan/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:FASB 295
CATEGORIES:algorithms & theory,fairness & ethics,machine learning,society &
  policy
END:VEVENT
BEGIN:VEVENT
UID:2023-10-04-bryan-perozzi@datascience.utah.edu
DTSTAMP:20231004T163000Z
DTSTART;TZID=America/Denver:20231004T103000
DTEND;TZID=America/Denver:20231004T114500
SUMMARY:Leveraging the Structure of Data - Bryan Perozzi
DESCRIPTION:Bryan Perozzi (Google Research)\n\nAlthough predictions from ma
 chine learning models influence more and more of our lives\, the standard 
 way of posing a ML problem has remained relatively unchanged for decades. 
 In the search for better models\, a new and popular family of techniques (
 sometimes called Graph Machine Learning) has emerged. These techniques rel
 y on expanding beyond the features of an individual entity and instead loo
 k to pull information from its relationships. The methods offer a tantaliz
 ing way of improving task performance by leveraging previously unused info
 rmation. However\, it is not a free lunch\, as these models can be more co
 mplex\, difficult to train\, and may have challenges in interpretability. 
 This talk will discuss the fundamentals of graph machine learning\, a few 
 models\, and some insights from years of real-world applications.\n\nhttps
 ://datascience.utah.edu/talks/2023-10-04-bryan-perozzi/
URL:https://datascience.utah.edu/talks/2023-10-04-bryan-perozzi/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:FASB 295
CATEGORIES:algorithms & theory,machine learning,networks & graphs,optimizat
 ion
END:VEVENT
BEGIN:VEVENT
UID:2023-10-18-mihai-budiu@datascience.utah.edu
DTSTAMP:20231018T163000Z
DTSTART;TZID=America/Denver:20231018T103000
DTEND;TZID=America/Denver:20231018T114500
SUMMARY:DBSP: A formal model for streaming computation and its applications
  to incremental computations and databases - Mihai Budiu
DESCRIPTION:Mihai Budiu (Feldera)\n\nDBSP is a simple streaming programming
  language inspired by Digital Signal Processing [DSP]. DBSP can be used to
  give a precise definition of incremental computations -- operating on cha
 nges (deltas\, diffs). Moreover\, given a DBSP program\, a simple algorith
 m can convert it to a DBSP program that computes on changes. All practical
  database query operators (the relational algebra\, group-by\, aggregation
 s\, fixed-points\, etc) can be expressed in DBSP. As a consequence we obta
 in an algorithm which can incrementalize essentially any database query. T
 he DBSP theory has been formally verified using a theorem prover\, making 
 it the first verified theory of incremental view maintenance.\n\nThe DBSP 
 paper has received the best paper award at the 2023 conference on Very Lar
 ge Databases [VLDB].\n\nhttps://datascience.utah.edu/talks/2023-10-18-miha
 i-budiu/
URL:https://datascience.utah.edu/talks/2023-10-18-mihai-budiu/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:FASB 295 | https://utah.zoom.us/j/91737198805pwd=Z1o2SzE4OVRodVhDW
 WExOTdVcUs5Zz09
CATEGORIES:algorithms & theory,data management
END:VEVENT
BEGIN:VEVENT
UID:2023-10-25-aydin-buluc@datascience.utah.edu
DTSTAMP:20231025T163000Z
DTSTART;TZID=America/Denver:20231025T103000
DTEND;TZID=America/Denver:20231025T114500
SUMMARY:Computational journeys in a sparse universe - Aydin Buluc
DESCRIPTION:Aydin Buluc (Berkeley Lab)\n\nSparsity is a fundamental assumpt
 ion that allows us to compute efficiently on and find parsimonious solutio
 ns to science and engineering problems. Sparsity exists in all basic scien
 ces such as physics\, biology\, and chemistry. I am going to give a sampli
 ng of recent work we have done on sparse computations. My talk will travel
  across diverse problem domains including randomized linear algebra\, grap
 h neural networks\, protein family and structure discovery from metagenomi
 c data\, and tensor computations. The underlying theme will be the challen
 ges posed by sparsity and the computational techniques we employ to overco
 me these challenges.\n\nhttps://datascience.utah.edu/talks/2023-10-25-aydi
 n-buluc/
URL:https://datascience.utah.edu/talks/2023-10-25-aydin-buluc/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:FASB 295 | https://utah.zoom.us/j/91737198805?pwd=Z1o2SzE4OVRodVhD
 WWExOTdVcUs5Zz09
CATEGORIES:biology & genomics,networks & graphs
END:VEVENT
BEGIN:VEVENT
UID:2023-11-01-william-kuszmaul@datascience.utah.edu
DTSTAMP:20231101T163000Z
DTSTART;TZID=America/Denver:20231101T103000
DTEND;TZID=America/Denver:20231101T114500
SUMMARY:Linear Probing Revisited: How to Get Rid of Clustering - William Ku
 szmaul
DESCRIPTION:William Kuszmaul (MIT)\n\nThe linear-probing hash table is one 
 of the oldest and most widely used data structures in computer science. Ho
 wever\, linear probing also famously comes with a major drawback: as soon 
 as the hash table reaches a high memory utilization\, elements within the 
 hash table begin to cluster together\, causing insertions to become slow. 
 This clustering phenomenon\, which was first discovered by Donald Knuth in
  1962\, increases the expected time per insertion to $\\Theta(x^2)$ (rathe
 r than the more desirable $\\Theta(x)$) in a hash table that is a $1 - 1/x
 $ fraction full.\nA natural question is whether one can somehow reduce clu
 stering. In this talk\, we establish an even stronger statement: the class
 ical linear-probing hash table (even as it was first implemented in the 19
 50s) already has less clustering than the classical results would seem to 
 suggest. As insertions and deletions are performed over time\, the tombsto
 nes left behind by deletions cause the combinatorial structure of the hash
  table to stabilize in a way that eliminates clustering. This means that\,
  for some versions of linear probing\, the amortized expected time per ope
 ration is actually $\\tilde{O}(x)$. We also present a new version of linea
 r probing that avoids clustering entirely\, achieving $O(x)$ expected time
  per operation.\n\nhttps://datascience.utah.edu/talks/2023-11-01-william-k
 uszmaul/
URL:https://datascience.utah.edu/talks/2023-11-01-william-kuszmaul/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:FASB 295 | https://utah.zoom.us/j/91737198805?pwd=Z1o2SzE4OVRodVhD
 WWExOTdVcUs5Zz09
CATEGORIES:algorithms & theory,data management
END:VEVENT
BEGIN:VEVENT
UID:2023-11-22-bryan-perozzi@datascience.utah.edu
DTSTAMP:20231122T173000Z
DTSTART;TZID=America/Denver:20231122T103000
DTEND;TZID=America/Denver:20231122T114500
SUMMARY:“Leveraging the Structure of Data“ - Bryan Perozzi
DESCRIPTION:Bryan Perozzi (Google Research)\n\nAlthough predictions from ma
 chine learning models influence more and more of our lives\, the standard 
 way of posing a ML problem has remained relatively unchanged for decades. 
 In the search for better models\, a new and popular family of techniques (
 sometimes called Graph Machine Learning) has emerged. These techniques rel
 y on expanding beyond the features of an individual entity and instead loo
 k to pull information from its relationships. The methods offer a tantaliz
 ing way of improving task performance by leveraging previously unused info
 rmation. However\, it is not a free lunch\, as these models can be more co
 mplex\, difficult to train\, and may have challenges in interpretability. 
 This talk will discuss the fundamentals of graph machine learning\, a few 
 models\, and some insights from years of real-world applications.\n\nhttps
 ://datascience.utah.edu/talks/2023-11-22-bryan-perozzi/
URL:https://datascience.utah.edu/talks/2023-11-22-bryan-perozzi/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:FASB 295
CATEGORIES:algorithms & theory,machine learning,networks & graphs,optimizat
 ion
END:VEVENT
BEGIN:VEVENT
UID:2023-11-29-shangdi-yu@datascience.utah.edu
DTSTAMP:20231129T173000Z
DTSTART;TZID=America/Denver:20231129T103000
DTEND;TZID=America/Denver:20231129T114500
SUMMARY:Framework for Parallel Hierarchical Agglomerative Clustering - Shan
 gdi Yu
DESCRIPTION:Shangdi Yu (MIT)\n\nWe study the hierarchical clustering proble
 m\, where the goal is to produce a dendrogram that represents clusters at 
 varying scales of a data set. We propose the ParChain framework for design
 ing parallel hierarchical agglomerative clustering (HAC) algorithms\, and 
 using the framework we obtain novel parallel algorithms for the complete l
 inkage\, average linkage\, and Ward's linkage criteria. Compared to most p
 revious parallel HAC algorithms\, which require quadratic memory\, our new
  algorithms require only linear memory\, and are scalable to large data se
 ts. ParChain is based on our parallelization of the nearest-neighbor chain
  algorithm\, and enables multiple clusters to be merged on every round. We
  introduce two key optimizations that are critical for efficiency: a range
  query optimization that reduces the number of distance computations requi
 red when finding nearest neighbors of clusters\, and a caching optimizatio
 n that stores a subset of previously computed distances\, which are likely
  to be reused. Experimentally\, we show that our highly-optimized implemen
 tations using 48 cores with two-way hyper-threading achieve 5.8--110.1x sp
 eedup over state-of-the-art parallel HAC algorithms and achieve 13.75--54.
 23x self-relative speedup. Compared to state-of-the-art algorithms\, our a
 lgorithms require up to 237.3x less space. Our algorithms are able to scal
 e to data set sizes with tens of millions of points\, which previous algor
 ithms are not able to handle.\n\nhttps://datascience.utah.edu/talks/2023-1
 1-29-shangdi-yu/
URL:https://datascience.utah.edu/talks/2023-11-29-shangdi-yu/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:FASB 295
CATEGORIES:algorithms & theory,optimization
END:VEVENT
BEGIN:VEVENT
UID:2023-12-06-julian-shun@datascience.utah.edu
DTSTAMP:20231206T173000Z
DTSTART;TZID=America/Denver:20231206T103000
DTEND;TZID=America/Denver:20231206T114500
SUMMARY:Parallel Batch-Dynamic Graph Algorithms - Julian Shun
DESCRIPTION:Julian Shun (MIT)\n\nThere has been significant interest in gra
 ph analytics due to their applications in many domains\, including social 
 network and Web analytics\, machine learning\, biology\, and physical simu
 lations. Real-world graphs today are massive and also dynamic. As many rea
 l-world graphs change rapidly\, it is crucial to design dynamic algorithms
  that efficiently maintain graph statistics upon updates\, since the cost 
 of re-computation from scratch can be prohibitive. Furthermore\, due to th
 e high frequency of updates\, we can improve performance by using parallel
 ism to process batches of updates at a time. This talk presents new graph 
 algorithms in this parallel batch-dynamic setting.\n\nSpecifically\, we pr
 esent the first parallel batch-dynamic algorithm for approximate k-core de
 composition that is efficient in both theory and practice. Our algorithm i
 s based on our novel parallel level data structure\, inspired by the seque
 ntial level data structures of Bhattacharya et al. and Henzinger et al. Gi
 ven a graph with n vertices and a batch of B updates\, our algorithm maint
 ains a (2 + epsilon)-approximation of the coreness values of all vertices 
 (for any constant epsilon > 0) in O(B log^2(n)) amortized work and O(log^2
 (n) loglog(n)) span (parallel time) with high probability. We implement an
 d experimentally evaluate our algorithm\, and demonstrate significant spee
 dups over state-of-the-art serial and parallel implementations for dynamic
  k-core decomposition.\n\nWe have also designed new parallel batch-dynamic
  algorithms for low out-degree orientation\, maximal matching\, clique cou
 nting\, graph coloring\, minimum spanning forest\, single-linkage clusteri
 ng\, some of which use our parallel level data structure.\n\nhttps://datas
 cience.utah.edu/talks/2023-12-06-julian-shun/
URL:https://datascience.utah.edu/talks/2023-12-06-julian-shun/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
CATEGORIES:algorithms & theory,networks & graphs,statistics
END:VEVENT
BEGIN:VEVENT
UID:2023-12-13-felix-reidl@datascience.utah.edu
DTSTAMP:20231213T173000Z
DTSTART;TZID=America/Denver:20231213T103000
DTEND;TZID=America/Denver:20231213T114500
SUMMARY:Talk by Felix Reidl - Felix Reidl
DESCRIPTION:Felix Reidl (Birkbeck University)\n\nData Science and AI have a
 n every increasing presence in our social\, political and economic life. G
 iven the immense influence the technologies of these field have and will h
 ave\, I argue that researchers and academic institutions should reflect on
  the implications their work has in the world at large.\n\nIn this talk I 
 would like take stock of the larger context: a world lacking futures\, fra
 gmented academic disciplines\, and the dystopian use of technology. While 
 we cannot hope to solve any of these problems\, I argue that we can and sh
 ould resist the underlying trends. To that end\, I propose that our instit
 utions should be constructed first and foremost around creativity and part
 icipation and what that could mean in practice.\n\nhttps://datascience.uta
 h.edu/talks/2023-12-13-felix-reidl/
URL:https://datascience.utah.edu/talks/2023-12-13-felix-reidl/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:FASB 295
CATEGORIES:society & policy
END:VEVENT
BEGIN:VEVENT
UID:2024-01-24-anna-little@datascience.utah.edu
DTSTAMP:20240124T203000Z
DTSTART;TZID=America/Denver:20240124T133000
DTEND;TZID=America/Denver:20240124T143000
SUMMARY:Clustering and Visualization of High-dimensional Data using Path Me
 trics - Anna Little
DESCRIPTION:Anna Little (Utah)\n\nThis talk will explore the utility of dat
 a-driven path metrics for the clustering and visualization of high-dimensi
 onal data. These metrics are defined by solving an optimal path problem in
  a proximity graph\, and are characterized by a parameter harmonizing dens
 ity-based and geometric features. First\, we will discuss theoretical prop
 erties of these metrics and implications for clustering: in particular\, s
 pectral clustering with path metrics leads to strong theoretical guarantee
 s on estimating number of clusters and cluster accuracy. Furthermore\, as 
 the sample size converges to infinity\, the eigenvalues and eigenvectors o
 f the discrete path metric graph Laplacian converge to those of a continuu
 m operator\; this operator generates a diffusion which is accelerated in r
 egions of high data density\, allowing for the rapid exploration of elonga
 ted data structures. Secondly\, we will discuss dimension reduction with p
 ath metrics. Despite the allure of visually striking results from dimensio
 n reduction algorithms\, can we be sure the perceived patterns are genuine
 ly intrinsic to the data? We will see how designing algorithms which incor
 porate data-driven path metrics can lead to desirable and understandable p
 roperties for data visualization\, and specifically focus on their applica
 tion to the analysis of single cell RNA sequence data. Finally\, we will d
 iscuss generalizations of these metrics to simplex paths\, and their utili
 ty for multi-manifold clustering.\n\nhttps://datascience.utah.edu/talks/20
 24-01-24-anna-little/
URL:https://datascience.utah.edu/talks/2024-01-24-anna-little/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://utah.zoom.us/j/96005100565?pwd=WmFGN25RazZwV2NoMGE2dVFGMng
 yZz09
CATEGORIES:algorithms & theory,biology & genomics,networks & graphs,visuali
 zation
END:VEVENT
BEGIN:VEVENT
UID:2024-01-31-pratik-soni@datascience.utah.edu
DTSTAMP:20240131T203000Z
DTSTART;TZID=America/Denver:20240131T133000
DTEND;TZID=America/Denver:20240131T143000
SUMMARY:Cryptography for Fairness - Pratik Soni
DESCRIPTION:Pratik Soni (Utah)\n\nhttps://datascience.utah.edu/talks/2024-0
 1-31-pratik-soni/
URL:https://datascience.utah.edu/talks/2024-01-31-pratik-soni/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://utah.zoom.us/j/96005100565?pwd=WmFGN25RazZwV2NoMGE2dVFGMng
 yZz09
CATEGORIES:fairness & ethics
END:VEVENT
BEGIN:VEVENT
UID:2024-02-07-daniel-brown@datascience.utah.edu
DTSTAMP:20240207T203000Z
DTSTART;TZID=America/Denver:20240207T133000
DTEND;TZID=America/Denver:20240207T143000
SUMMARY:Challenges and Progress Towards AI Alignment via Reinforcement Lear
 ning from Human Feedback - Daniel Brown
DESCRIPTION:Daniel Brown (Utah)\n\nIn this talk I will discuss recent progr
 ess and challenges towards using human feedback to develop AI systems that
  behave in ways that are aligned with human preferences. One problem that 
 arises when robots and other AI systems learn from human input is that the
 re is often a large amount of uncertainty over the human’s true intent a
 nd the corresponding desired AI behavior. To address this problem\, I will
  discuss prior and ongoing research along three main topics: (1) Enabling 
 robots and other AI systems to learn models of human preferences in ways t
 hat are sample efficient and safe\, (2) Reward misidentification and causa
 l confusion when learning from human feedback\, and (3) Active learning me
 thods for learning better aligned features and reward functions.\n\nhttps:
 //datascience.utah.edu/talks/2024-02-07-daniel-brown/
URL:https://datascience.utah.edu/talks/2024-02-07-daniel-brown/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:GC 2560 (Gardner Commons) | https://utah.zoom.us/j/96005100565?pwd
 =WmFGN25RazZwV2NoMGE2dVFGMngyZz09
CATEGORIES:machine learning,robotics
END:VEVENT
BEGIN:VEVENT
UID:2024-02-14-shivam-garg@datascience.utah.edu
DTSTAMP:20240214T203000Z
DTSTART;TZID=America/Denver:20240214T133000
DTEND;TZID=America/Denver:20240214T143000
SUMMARY:In-Context Learning: A Case Study of Simple Function Classes - Shiv
 am Garg
DESCRIPTION:Shivam Garg (Harvard)\n\nIn-context learning refers to the abil
 ity of a model to learn new tasks from a sequence of input-output pairs gi
 ven in a prompt. Crucially\, this learning happens at inference time witho
 ut any parameter updates to the model. I will discuss our empirical effort
 s that shed light on some basic aspects of in-context learning: To what ex
 tent can Transformers\, or other models such as LSTMs be efficiently train
 ed to in-context learn fundamental function classes\, such as linear funct
 ions\, sparse linear functions\, and small decision trees? How can one eva
 luate in-context learning algorithms? And what are the qualitative differe
 nces between these architectures with respect to their ability to be train
 ed to perform in-context learning? This is based on joint work with Dimitr
 is Tsipras\, Percy Liang\, and Greg Valiant.\n\nhttps://datascience.utah.e
 du/talks/2024-02-14-shivam-garg/
URL:https://datascience.utah.edu/talks/2024-02-14-shivam-garg/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://utah.zoom.us/j/96005100565?pwd=WmFGN25RazZwV2NoMGE2dVFGMng
 yZz09
CATEGORIES:algorithms & theory,deep learning,large language models
END:VEVENT
BEGIN:VEVENT
UID:2024-02-21-tolga-tasdizen@datascience.utah.edu
DTSTAMP:20240221T203000Z
DTSTART;TZID=America/Denver:20240221T133000
DTEND;TZID=America/Denver:20240221T143000
SUMMARY:Deep Learning in Image Analysis: Applications and Challenges - Tolg
 a Tasdizen
DESCRIPTION:Tolga Tasdizen (Utah)\n\nMachine learning and more specifically
  deep learning has revolutionized computer vision and images analysis prob
 lems in a wide range of application areas. In this talk\, we will outline 
 four applications ranging from material science to health informatics and 
 medical image analysis. We will discuss challenges unique to each problem 
 and our approach using machine learning. A common theme underlying these p
 roblems is the difficulty of obtaining annotated data. Our contributions t
 owards solving this problem ranging from new semi-supervised learning form
 ulations to leveraging existing workflows of experts to collect annotation
 s will be introduced.\n\nhttps://datascience.utah.edu/talks/2024-02-21-tol
 ga-tasdizen/
URL:https://datascience.utah.edu/talks/2024-02-21-tolga-tasdizen/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://utah.zoom.us/j/96005100565?pwd=WmFGN25RazZwV2NoMGE2dVFGMng
 yZz09
CATEGORIES:computer vision,deep learning,health & medicine,machine learning
END:VEVENT
BEGIN:VEVENT
UID:2024-02-28-baskar-ganapathysubramanian@datascience.utah.edu
DTSTAMP:20240228T203000Z
DTSTART;TZID=America/Denver:20240228T133000
DTEND;TZID=America/Denver:20240228T143000
SUMMARY:Neural PDE solvers with applications in manufacturing and agricultu
 re - Baskar Ganapathysubramanian
DESCRIPTION:Baskar Ganapathysubramanian (Iowa)\n\nNumerical simulation is a
  critical tool in analysis\, design\, and control of complex systems\, whi
 ch are usually described by partial differential equations (PDE). The trad
 itional approach to solving these PDEs has been via their numerical approx
 imations (for example\, finite difference\, finite element). Recently\, ad
 vances in Scientific machine learning (SciML) has opened up the possibilit
 y of training deep networks to solve complex PDEs. Several promising SciML
  approaches –for example\, PINNs\, FNO – have recently been proposed t
 o solve PDEs under various amounts of data availability.\n\nIn this talk\,
  I will discuss some of my group’s contribution to this effort in traini
 ng neural PDE solvers. This discussion will cover (a) formulating and trai
 ning a mesh-based neural network approach that solves for a large parametr
 ic family of PDEs\, (b) how we accelerate training a large models (for meg
 a voxel PDE predictions) via a method analogous to the multigrid technique
  used in numerical linear algebra\, (c) approaches to solve PDEs over doma
 ins with irregularly shaped (non-rectilinear) geometric boundaries\, and (
 d) applying these approaches for solving design problems in agriculture an
 d energy technology. Whenever possible\, we perform analysis which reveals
  theoretical insights into several sources of error incurred in the model-
 building process.\n\nhttps://datascience.utah.edu/talks/2024-02-28-baskar-
 ganapathysubramanian/
URL:https://datascience.utah.edu/talks/2024-02-28-baskar-ganapathysubramani
 an/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://utah.zoom.us/j/96005100565?pwd=WmFGN25RazZwV2NoMGE2dVFGMng
 yZz09
CATEGORIES:machine learning,optimization,physics & astronomy
END:VEVENT
BEGIN:VEVENT
UID:2024-03-13-swabha-swayamdipta@datascience.utah.edu
DTSTAMP:20240313T193000Z
DTSTART;TZID=America/Denver:20240313T133000
DTEND;TZID=America/Denver:20240313T143000
SUMMARY:Understanding LLMs through their Generative Behavior\, Successes an
 d Shortcomings - Swabha Swayamdipta
DESCRIPTION:Swabha Swayamdipta (USC)\n\nGenerative capabilities of large la
 nguage models have grown beyond the wildest imagination of the broader AI 
 research community\, leading many to speculate whether these successes may
  be attributed to the training data or model design. I will present some w
 ork from my group which sheds light on understanding LLMs by studying thei
 r generative behavior\, successes and shortcomings. First\, I will show th
 at standard inference algorithms work well because of the particular desig
 n behind LLMs. Next\, I will discuss recently found successes and failures
  of LLMs on a combination of tasks\, requiring world and domain-specific k
 nowledge\, linguistic capabilities and awareness of human and social utili
 ty. Overall\, these findings paint a partial yet complex picture of our un
 derstanding of LLMs and provide a guide to the next steps forward.\n\nhttp
 s://datascience.utah.edu/talks/2024-03-13-swabha-swayamdipta/
URL:https://datascience.utah.edu/talks/2024-03-13-swabha-swayamdipta/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
CATEGORIES:large language models,machine learning,natural language processi
 ng
END:VEVENT
BEGIN:VEVENT
UID:2024-04-10-nitin-bakshi@datascience.utah.edu
DTSTAMP:20240410T193000Z
DTSTART;TZID=America/Denver:20240410T133000
DTEND;TZID=America/Denver:20240410T143000
SUMMARY:Service Operations for Justice-On-Time: A Data-Driven Queueing Appr
 oach - Nitin Bakshi
DESCRIPTION:Nitin Bakshi (Utah)\n\nLimited resources in the judicial system
  can lead to costly delays\, stunted economic development\, and even failu
 re to deliver justice. Using the Supreme Court of India as an exemplar for
  such resource-constrained settings\, we apply ideas from service operatio
 ns to study delay. Specifically\, court dynamics constitute a case-managem
 ent queue\, whereby each case may experience multiple service encounters s
 pread across time\, but all are necessarily with the same server. Our goal
  is to elucidate the drivers of congestion\, focusing on metrics such as t
 he expected case-disposition time (delay) and expected number of cases awa
 iting adjudication (pendency)\, and leverage this understanding to recomme
 nd operational interventions.\n\nWe employ data-driven calibrated simulati
 ons to model the analytically intractable case-management queue. The life 
 cycle of a case comprises two stages: pre-admission (before determining it
 s merit for detailed hearings) and post-admission. Our methodology allows 
 us to capture the queueing dynamics in which the judges are shared resourc
 es across the two stages. It also permits modeling of holiday capacity\, w
 hich is flexibly tailored to address any surplus work that spills over fro
 m the regular year. We find that the second stage of this judicial queue i
 s overloaded\, but holiday capacity creates a perception of stability by s
 teadying performance metrics.\n\nThe sources of inefficiency that drive co
 ngestion include a misalignment between scheduling guidelines and judicial
  capacity\, coupled with the requirement to schedule hearings in advance. 
 Together\, these factors inhibit utilization of shared capacity across the
  two-stage judicial queue. We demonstrate how interventions that account f
 or these inefficiencies can successfully tackle judicial delay. In particu
 lar\, scheduling to improve the allocation of time across pre- and post-ad
 mission cases can cut down the expected delay by as much as 65%.\n\nhttps:
 //datascience.utah.edu/talks/2024-04-10-nitin-bakshi/
URL:https://datascience.utah.edu/talks/2024-04-10-nitin-bakshi/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:https://utah.zoom.us/j/96005100565?pwd=WmFGN25RazZwV2NoMGE2dVFGMng
 yZz09
CATEGORIES:fairness & ethics,society & policy
END:VEVENT
BEGIN:VEVENT
UID:2024-08-27-jeff-phillips@datascience.utah.edu
DTSTAMP:20240827T183000Z
DTSTART;TZID=America/Denver:20240827T123000
DTEND;TZID=America/Denver:20240827T133000
SUMMARY:== Sketching and Classifying Spatial Trajectories == - Jeff Phillip
 s
DESCRIPTION:Jeff Phillips (Utah KSoC)\n\nSpatial trajectories\, often repre
 sented as a sequence of spatial positions\, are a standard way to represen
 t human mobility patterns. They also are used to represent motion patterns
  including for animals\, drones\, or last-mile rentals (e-scooters). Howev
 er\, these trajectories are notoriously difficult to work with as they ove
 rlap and can stretch long distances.\nIn this talk we discuss a sketch (th
 e minDist Sketch) that makes just about any data analysis on trajectories 
 tasks simple and efficient. This first considers spatial trajectories as a
 n abstract shape\, and then maps them to a high-dimensional Euclidean spac
 e as a vector. We can show recovery\, pseudo-metric\, and metric propertie
 s of this representation. Variants can include direction information\, or 
 traits like velocity and acceleration.\nMoreover\, once represented as thi
 s vector\, the trajectory data is extremely easy to work with. Allowing fo
 r out-of-the-box use of software for nearest-neighbor search\, clustering\
 , and classification.\nIn particular\, we conduct the first formal study o
 f classifying spatial trajectories: given trajectories from two different 
 distributions (e.g.\, generated by car or bus) given a new trajectory that
  is unlabeled\, how well can we predict which class it was from?\nOver sev
 eral data sets we have assembled that demand this task\, we conduct a larg
 e study\, and show that the minDist sketch and its variants are consistent
 ly the easiest and most accurate method (or at the least among the best in
  each instance).\n\nhttps://datascience.utah.edu/talks/2024-08-27-jeff-phi
 llips/
URL:https://datascience.utah.edu/talks/2024-08-27-jeff-phillips/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 1230 | https://utah.zoom.us/j/93909986581?pwd=d90LoHKoVAkagz1a
 CpJH1vTuaME9gG.1
CATEGORIES:algorithms & theory
END:VEVENT
BEGIN:VEVENT
UID:2024-09-03-esha-datta@datascience.utah.edu
DTSTAMP:20240903T183000Z
DTSTART;TZID=America/Denver:20240903T123000
DTEND;TZID=America/Denver:20240903T133000
SUMMARY:Topological Signatures of Out-of-Distribution Examples - Esha Datta
DESCRIPTION:Esha Datta (Sandia NL)\n\nMachine learning (ML) models employed
  for real-world tasks will invariably encounter inference data that is dis
 tributionally shifted from their training datasets. Such out-of-distributi
 on (OOD) examples can have adverse effects on model performance and can po
 se significant problems in high-consequence application areas like healthc
 are or autonomous vehicles. We develop a topological characterization of O
 OD examples and present a computationally feasible methodology for detecti
 ng such data in a deployed pipeline. The approach leverages the known prop
 erty that well-trained ML models induce a topological “simplification”
  on its training dataset. By computing the persistent homology of the hidd
 en layer embeddings of training and test data\, we demonstrate empirically
  our ability to identify the presence of OOD examples for a given model.\n
 \nhttps://datascience.utah.edu/talks/2024-09-03-esha-datta/
URL:https://datascience.utah.edu/talks/2024-09-03-esha-datta/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 1230 | https://utah.zoom.us/j/93909986581?pwd=d90LoHKoVAkagz1a
 CpJH1vTuaME9gG.1
CATEGORIES:machine learning
END:VEVENT
BEGIN:VEVENT
UID:2024-09-10-guanhong-tao@datascience.utah.edu
DTSTAMP:20240910T183000Z
DTSTART;TZID=America/Denver:20240910T123000
DTEND;TZID=America/Denver:20240910T133000
SUMMARY:Are AI-enabled Systems Safe and Secure? - Guanhong Tao
DESCRIPTION:Guanhong Tao\n\nAbstract\nArtificial Intelligence (AI) has been
  integrated into various sectors\, such as facial recognition and autonomo
 us driving. But are the security and safety of these AI-enabled systems fu
 lly ensured? In this talk\, I will present various vulnerabilities in thes
 e systems. My presentation will cover novel optimization techniques for id
 entifying and mitigating backdoor vulnerabilities in both white-box and bl
 ack-box settings\, achieving substantial improvements in performance. I wi
 ll share insights into the nature of backdoors and their presence in pre-t
 rained models. Finally\, I will conclude with an outlook on our recent exp
 loration of the security of emerging AI techniques\, such as generative AI
 .\n\nhttps://datascience.utah.edu/talks/2024-09-10-guanhong-tao/
URL:https://datascience.utah.edu/talks/2024-09-10-guanhong-tao/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 1230 | https://utah.zoom.us/j/93909986581?pwd=d90LoHKoVAkagz1a
 CpJH1vTuaME9gG.1
CATEGORIES:privacy & security
END:VEVENT
BEGIN:VEVENT
UID:2024-09-17-aurora-clark@datascience.utah.edu
DTSTAMP:20240917T183000Z
DTSTART;TZID=America/Denver:20240917T123000
DTEND;TZID=America/Denver:20240917T133000
SUMMARY:The Importance of Shape in Chemistry Data - Aurora Clark
DESCRIPTION:Aurora Clark\n\nData in the field of Chemistry has heavily leve
 raged graph theory representations within data science applications. Howev
 er\, there is a rich geometric and topological structure of many chemical 
 systems (and their data) that has been less employed for feature optimizat
 ion\, dimensionality reduction\, and predictive models. Within this discus
 sion I will highlight some recent work and interests that seek to employ c
 omputational topology and geometry within chemistry data sets from molecul
 ar dynamics simulations – both in the context of ensemble average and te
 mporally evolving data sets.\n\nhttps://datascience.utah.edu/talks/2024-09
 -17-aurora-clark/
URL:https://datascience.utah.edu/talks/2024-09-17-aurora-clark/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 1230 | https://utah.zoom.us/j/93909986581?pwd=d90LoHKoVAkagz1a
 CpJH1vTuaME9gG.1
CATEGORIES:networks & graphs,optimization
END:VEVENT
BEGIN:VEVENT
UID:2024-09-24-rebecca-barter@datascience.utah.edu
DTSTAMP:20240924T183000Z
DTSTART;TZID=America/Denver:20240924T123000
DTEND;TZID=America/Denver:20240924T133000
SUMMARY:Veridical Data Science: the Practice of Responsible Data Analysis a
 nd Decision Making - Rebecca Barter
DESCRIPTION:Rebecca Barter\n\nData science is often presented as a straight
 forward\, linear process involving statistical and computational technique
 s\, without addressing the complexities inherent in real-world application
 s. In contrast\, our new book\,\n"Veridical Data Science: The Practice of 
 Responsible Data Analysis and Decision Making"\, teaches data scientists t
 o navigate the reality that most projects involve answering ambiguous doma
 in questions with messy data\, all while managing a complex web of human j
 udgment calls. We emphasize that datasets are merely approximations of rea
 lity\, and analyses are shaped by human interpretation. Using the Predicta
 bility\, Computability\, and Stability (PCS) framework to assess the trust
 worthiness and relevance of data-driven results\, "Veridical Data Science"
  provides an actionable guide for conducting responsible and trustworthy d
 ata science.\n\nhttps://datascience.utah.edu/talks/2024-09-24-rebecca-bart
 er/
URL:https://datascience.utah.edu/talks/2024-09-24-rebecca-barter/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 1230 | https://utah.zoom.us/j/93909986581?pwd=d90LoHKoVAkagz1a
 CpJH1vTuaME9gG.1
CATEGORIES:health & medicine,statistics
END:VEVENT
BEGIN:VEVENT
UID:2024-10-01-simon-brewer@datascience.utah.edu
DTSTAMP:20241001T183000Z
DTSTART;TZID=America/Denver:20241001T123000
DTEND;TZID=America/Denver:20241001T133000
SUMMARY:Exploring long-term ecosystem change with self-organizing maps - Si
 mon Brewer
DESCRIPTION:Simon Brewer\n\nOngoing climate change has the potential to imp
 act a variety of physical\, biological and social systems\, and there is i
 ncreasing concern that these changes may be sufficient to result in these 
 systems crossing tipping points\, effectively undergoing irreversible chan
 ges in state. For slow turnover systems\, such as forest ecosystems\, unde
 rstanding the likelihood and ramifications of these state changes is chall
 enging due to the relative short observational record. Sedimentary records
  of ecosystem change offer an alternative data source with a wide temporal
  and spatial scope\, but are inherently noisy and high dimensional. Self-o
 rganizing maps provide a data-driven way to visualize nonlinear patterns i
 n these data\, and to identify past ecosystem states and state transitions
 . The results are used to build a simple Markov model illustrating the pro
 bability and directionality of these transitions.\n\nhttps://datascience.u
 tah.edu/talks/2024-10-01-simon-brewer/
URL:https://datascience.utah.edu/talks/2024-10-01-simon-brewer/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 1230 | https://utah.zoom.us/j/93909986581?pwd=d90LoHKoVAkagz1a
 CpJH1vTuaME9gG.1
CATEGORIES:biology & genomics,climate & environment,statistics
END:VEVENT
BEGIN:VEVENT
UID:2024-10-15-raghav-venkatraman@datascience.utah.edu
DTSTAMP:20241015T183000Z
DTSTART;TZID=America/Denver:20241015T123000
DTEND;TZID=America/Denver:20241015T133000
SUMMARY:Minmax estimation rates for manifold learning - Raghav Venkatraman
DESCRIPTION:Raghav Venkatraman\n\nThis talk is focused on obtaining minmax 
 estimation rates for the "manifold learning" problem. Given N data points 
 hypothesized to be i.i.d (independent and identically distributed) samples
  of a nice density (from within a reasonable class of densities) on a nice
  manifold (from within some class of nice manifolds of known intrinsic dim
 ension d)\, the manifold learning problem boils down to estimating certain
  statistics of this "ground truth" manifold\, such as the first few eigenm
 odes of the Laplace Beltrami operator on the manifold. The minmax estimati
 on problem further asks: given N such data points\, among all estimators o
 f the desired statistics (say\, a particular eigenvalue and associated eig
 enfunctions in a suitable norm)\, which one achieves the smallest maximum 
 expected risk\, and how does this minmax risk scale in N and the intrinsic
  dimension d of the manifold?\n\nAn intuitive but impractical estimator co
 nsists in estimating the density from the given samples through a ``kernel
  density estimation''\, and then solving the resulting continuum eigenprob
 lem using a numerical method such as finite elements: this estimator turns
  out to be minmax optimal in scaling-- namely\, the associated expected ri
 sk scales like N^{-2/d+4}\, and we can show a matching lower bound for the
  minmax risk\, demonstrating that no estimator can do better\, in scaling\
 , than this estimator.\n\nNext\, we ask: do there exist *practical* estima
 tors that are agnostic to knowledge of the manifold (so we don't have to d
 iscretize them in order to compute with finite elements!) that achieve\, a
 t least nearly\, this minmax scaling of the expected risk? We affirmativel
 y answer this question by showing that\, the spectrum of a carefully const
 ructed graph laplacian on a random geometric graph constructed from the gi
 ven N samples provides an estimator for the eigenvalue and eigenvectors of
  the Laplace Beltrami operator on the manifold that achieves this minmax r
 ate upto a log factor (to a small power). Both the lower and upper bound e
 stimates in the talk bring in new PDE tools to this statistical question\,
  and that we believe will be more broadly applicable in similar applicatio
 ns.\n\nThis talk is based on joint work with Nicolas Garcia Trillos and hi
 s PhD student Chenghui Li (U. W. Madison)\, and builds on prior joint work
  with Scott N. Armstrong (Courant Institute).\n\nhttps://datascience.utah.
 edu/talks/2024-10-15-raghav-venkatraman/
URL:https://datascience.utah.edu/talks/2024-10-15-raghav-venkatraman/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 1230 | https://utah.zoom.us/j/93909986581?pwd=d90LoHKoVAkagz1a
 CpJH1vTuaME9gG.1
CATEGORIES:algorithms & theory,networks & graphs,statistics
END:VEVENT
BEGIN:VEVENT
UID:2024-10-29-vivek-gupta@datascience.utah.edu
DTSTAMP:20241029T183000Z
DTSTART;TZID=America/Denver:20241029T123000
DTEND;TZID=America/Denver:20241029T133000
SUMMARY:Reasoning on Tabular and Multimodal Data - Vivek Gupta
DESCRIPTION:Vivek Gupta (ASU & UCDS alumni)\n\nIn this talk\, I’ll walk t
 hrough some of the latest AI advancements that address the challenges of w
 orking with complex data\, with a focus on improving reasoning for both ta
 bular and multimodal data.\n\nI’ll begin by introducing H-STAR\, a hybri
 d algorithm that combines symbolic and semantic reasoning to enhance quest
 ion answering for tabular data. By leveraging multi-view table extraction 
 and adaptive reasoning\, H-STAR has shown great potential in improving rea
 soning across tabular datasets.\n\nThen\, I’ll introduce MMTabQA\, a dat
 aset we developed to evaluate how AI systems manage multimodal tables that
  integrate structured text and images. Our research reveals where current 
 models struggle to process these diverse data types\, highlighting key are
 as for improvement.\n\nTo wrap up\, I’ll discuss open challenges and fut
 ure directions\, including expanding reasoning capabilities to other compl
 ex data types—such as charts\, maps\, and flowcharts—and enhancing AI 
 systems' robustness in handling numerical\, temporal\, and visual reasonin
 g\, particularly with large (vision) language models.\n\nhttps://datascien
 ce.utah.edu/talks/2024-10-29-vivek-gupta/
URL:https://datascience.utah.edu/talks/2024-10-29-vivek-gupta/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 1230 | https://utah.zoom.us/j/93909986581?pwd=d90LoHKoVAkagz1a
 CpJH1vTuaME9gG.1
CATEGORIES:computer vision,data management,natural language processing,visu
 alization
END:VEVENT
BEGIN:VEVENT
UID:2024-11-12-amir-abdullah@datascience.utah.edu
DTSTAMP:20241112T193000Z
DTSTART;TZID=America/Denver:20241112T123000
DTEND;TZID=America/Denver:20241112T133000
SUMMARY:Interpreting Learned Feedback Patterns in Large Language Models - A
 mir Abdullah
DESCRIPTION:Amir Abdullah\n\nAmir is an active researcher in mechanistic in
 terpretability\, opening the blackbox of large language models to reverse 
 engineer the inner workings and analyze internal representations.. In this
  talk\, he will discuss his paper in NeuIPS 2024 on interpreting reward mo
 dels in language models using sparse autoencoders on internal representati
 ons. Further\, Amir will introduce followup work studying whether internal
  representations can be transferred between large language models.\n\nhttp
 s://datascience.utah.edu/talks/2024-11-12-amir-abdullah/
URL:https://datascience.utah.edu/talks/2024-11-12-amir-abdullah/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 1230 | https://utah.zoom.us/j/93909986581?pwd=d90LoHKoVAkagz1a
 CpJH1vTuaME9gG.1
CATEGORIES:large language models,natural language processing
END:VEVENT
BEGIN:VEVENT
UID:2024-11-19-hoaning-xue@datascience.utah.edu
DTSTAMP:20241119T193000Z
DTSTART;TZID=America/Denver:20241119T123000
DTEND;TZID=America/Denver:20241119T133000
SUMMARY:Computational and experimental approaches to examining short videos
 ' persuasive effects - Hoaning Xue
DESCRIPTION:Hoaning Xue (Utah Communications)\n\nThere’s a gap in underst
 anding how people process multimodal information collectively\, despite ex
 tensive research on the effects of individual multimodal features and the 
 rapid advances in computer vision. This gap is increasingly relevant as sh
 ort video platforms emerge as major information sources and influence publ
 ic opinion. In this talk\, I present findings from a project that combines
  a data-driven approach with social scientific theories to investigate how
  multimodal features in short videos impact audience engagement and attitu
 de change. This research is grounded in the theoretical framework of Messa
 ge Sensation Value (MSV) to theorize and quantify how multimodal features 
 in short videos capture attention and affect information processing. This 
 project includes two studies: (1) a computational model of MSV that predic
 ts video engagement from 11 multimodal features across a dataset of 15\,00
 0 short videos from three popular short video platforms\; second\, an onli
 ne experiment examining the attentional mechanism underlying the persuasiv
 e effects of MSV in short videos on message credibility and attitude chang
 e. This project provides a useful framework and computational tool for sho
 rt video research.\n\nhttps://datascience.utah.edu/talks/2024-11-19-hoanin
 g-xue/
URL:https://datascience.utah.edu/talks/2024-11-19-hoaning-xue/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 1230 | https://utah.zoom.us/j/93909986581?pwd=d90LoHKoVAkagz1a
 CpJH1vTuaME9gG.1
CATEGORIES:computer vision
END:VEVENT
BEGIN:VEVENT
UID:2024-11-26-zhichao-xu@datascience.utah.edu
DTSTAMP:20241126T193000Z
DTSTART;TZID=America/Denver:20241126T123000
DTEND;TZID=America/Denver:20241126T133000
SUMMARY:Representation Learning for IR and Role of Retrieval in LLM Era - Z
 hichao Xu
DESCRIPTION:Zhichao Xu (Utah KSoC)\n\nRetrieval is the critical way of acce
 ssing information in people’s daily lives. Retrieval applications includ
 e search engines\, conversational shopping assistants\, or when asked abou
 t 2+3=?\, human brains do retrieval instead of reasoning. In this talk\, I
 ’ll briefly go through the history of representation learning in retriev
 al\, from Bag-of-Words representations to the latest dense and learned spa
 rse retrieval algorithms. Increasingly\, people go to ChatGPT or other lar
 ge language models for information seeking instead of search engines. With
  this existential crisis in mind\, I will talk about the role of retrieval
  in LLM era\, specifically\, retrieval-augmented generation\, strengths\, 
 weaknesses and open problems.\n\nhttps://datascience.utah.edu/talks/2024-1
 1-26-zhichao-xu/
URL:https://datascience.utah.edu/talks/2024-11-26-zhichao-xu/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 1230 | https://utah.zoom.us/j/93909986581?pwd=d90LoHKoVAkagz1a
 CpJH1vTuaME9gG.1
CATEGORIES:large language models,natural language processing
END:VEVENT
BEGIN:VEVENT
UID:2025-01-17-fengjiao-wang@datascience.utah.edu
DTSTAMP:20250117T203000Z
DTSTART;TZID=America/Denver:20250117T133000
DTEND;TZID=America/Denver:20250117T143000
SUMMARY:Supervised Learning on Tabular Data - Fengjiao Wang
DESCRIPTION:Fengjiao Wang (Utah SoC)\n\nSelf-supervised and Semi-supervised
  learning (SSL) on tabular data is an understudied topic. Despite some att
 empts\, there are two major challenges: 1. Imbalanced nature in the tabula
 r dataset\; 2. The one-hot encoding used in these methods becomes less eff
 icient for high-cardinality categorical features. To cope with the challen
 ges\, we propose SAWTab which uses a target encoding method\, Conditional 
 Probability Representation (CPR)\, for efficient representation in the inp
 ut space of categorical features. We improve this representation by incorp
 orating the unlabeled samples through pseudo-labels. Furthermore\, we prop
 ose a Smooth Adaptive Weighting mechanism in the target encoding to mitiga
 te the issue of noisy and biased pseudo-labels. Experimental results on va
 rious datasets and comparisons with existing frameworks show that SAWTab y
 ields best test accuracy on all datasets. We find that pseudo-labels can h
 elp improve the input space representation in the SSL setting\, which enha
 nces the generalization of the learning algorithm.\n\nhttps://datascience.
 utah.edu/talks/2025-01-17-fengjiao-wang/
URL:https://datascience.utah.edu/talks/2025-01-17-fengjiao-wang/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB L112 | https://utah.zoom.us/j/93909986581?pwd=d90LoHKoVAkagz1a
 CpJH1vTuaME9gG.1
CATEGORIES:machine learning
END:VEVENT
BEGIN:VEVENT
UID:2025-01-24-dieter-fox@datascience.utah.edu
DTSTAMP:20250124T210000Z
DTSTART;TZID=America/Denver:20250124T140000
DTEND;TZID=America/Denver:20250124T150000
SUMMARY:Data Science & AI Day - Dieter Fox
DESCRIPTION:Dieter Fox\n\nThe last years have seen astonishing progress in 
 the capabilities of generative AI techniques\, particularly in the areas o
 f language and visual understanding and generation. Key to the success of 
 these models are the use of image and text data sets of unprecedented scal
 e along with models that are able to digest such large datasets. We are no
 w seeing the first examples of leveraging such models to equip robots with
  open-world visual understanding and reasoning capabilities. Unfortunately
 \, however\, we have not achieved the RobotGPT moment\; these models still
  struggle with reasoning about geometry and physical interactions in the r
 eal world\, resulting in brittle performance on seemingly simple tasks suc
 h as manipulating objects in the open world. A crucial reason for this pro
 blem is the lack of data suitable to train powerful\, general models for r
 obot decision making and control. In this talk\, I will discuss approaches
  to generating large datasets for training robot manipulation capabilities
 \, with a focus on the role simulation can play in this context. I will sh
 ow some of our prior work\, where we demonstrated robust sim-to-real trans
 fer of manipulation skills trained in simulation\, and then discuss a prom
 ising direction toward training a model architecture that combines high-le
 vel\, semantic\, open-world reasoning\, with low-level 3D robot policies.\
 n\nhttps://datascience.utah.edu/talks/2025-01-24-dieter-fox/
URL:https://datascience.utah.edu/talks/2025-01-24-dieter-fox/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:Union Ballroom
CATEGORIES:robotics
END:VEVENT
BEGIN:VEVENT
UID:2025-02-07-omkar-bhalerao@datascience.utah.edu
DTSTAMP:20250207T203000Z
DTSTART;TZID=America/Denver:20250207T133000
DTEND;TZID=America/Denver:20250207T143000
SUMMARY:Triadic First-Order Logic Queries in Temporal Networks - Omkar Bhal
 erao
DESCRIPTION:Omkar Bhalerao\n\nMotif counting is a fundamental problem in ne
 twork analysis\, and there is a rich literature of theoretical and applied
  algorithms for this problem. Given a large input network G\, a motif H is
  a small “pattern" graph indicative of special local structure. Motif/pa
 ttern mining involves finding all matches of this pattern in the input G. 
 The simplest\, yet challenging\, case of motif counting is when H has thre
 e vertices\, often called a triadic query. Recent work has focused on temp
 oral graph mining\, where the network G has edges with timestamps (and dir
 ections) and H has time constraints. Such networks are common representati
 ons for communication networks\, citation networks\, financial transaction
 s\, etc.\nInspired by concepts in logic and database theory\, we introduce
  the study of Thresholded First Order Logic (FOL) Motif Analysis for massi
 ve temporal networks. A typical triadic motif query asks for the existence
  of three vertices that form a desired temporal pattern. An FOL motif quer
 y is obtained by having both existence and universal quantifiers with thre
 sholds. This allows for query semantics that can mine richer information f
 rom networks. A typical triadic query would be "find all triples of vertic
 es u\,v\,w such that they form a triangle within one hour". A thresholded 
 FOL query can express "find all pairs u\,v such that for half of w where (
 u\,w) formed an edge\, (v\,w) also formed an edge within an hour".\n\nWe d
 esign the first algorithm\, FOLTY\, for mining thresholded triadic FOL que
 ries\, whose theoretical running time matches the best known running time 
 for sparse graphs. Specifically\, our algorithms run in time 𝑂 (m $\\al
 pha \\log \\sigma_{\\max}$). Here\, $m$ is the number of temporal edges in
  the input graph\, $\\alpha$ is its degeneracy (maximum core number)\, and
  the $\\sigma_{\\max}$ is the maximum edge multiplicity. Our procedures ca
 n be efficiently implemented\, and FOLTY has good empirical behavior. For 
 example\, we can answer triadic FOL queries on graphs with nearly 70M edge
 s in less than an hour on commodity hardware. We believe that our work cou
 ld start a new research direction in the classic well-studied problem of m
 otif analysis.\n\nhttps://datascience.utah.edu/talks/2025-02-07-omkar-bhal
 erao/
URL:https://datascience.utah.edu/talks/2025-02-07-omkar-bhalerao/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB L112 | https://utah.zoom.us/j/93909986581?pwd=d90LoHKoVAkagz1a
 CpJH1vTuaME9gG.1
CATEGORIES:algorithms & theory,data management,networks & graphs
END:VEVENT
BEGIN:VEVENT
UID:2025-02-14-bao-wang@datascience.utah.edu
DTSTAMP:20250214T203000Z
DTSTART;TZID=America/Denver:20250214T133000
DTEND;TZID=America/Denver:20250214T143000
SUMMARY:Conditional Flow Divergence Matching - Bao Wang
DESCRIPTION:Bao Wang\n\nConditional flow matching (CFM) stands out as an ef
 ficient simulation-free approach for training flow-based generative models
 \, achieving remarkable performance for data generation. However\, CFM is 
 insufficient to ensure accuracy in learning probability paths\, and the le
 arned vector field significantly violates the continuity equation governin
 g probability flows. In response\, we establish a new total-variation boun
 d between the learned and ground-truth probability paths\, showing that th
 e gap between probability paths is bounded above by a combination of CFM l
 oss and an associated divergence loss. This theoretical bound informs us t
 o design a new objective to match both flow and divergence accompanied by 
 an efficient implementation. Our new training approach improves the perfor
 mance of the flow-based generative model by a noticeable margin without si
 gnificantly raising the computational cost.\n\nhttps://datascience.utah.ed
 u/talks/2025-02-14-bao-wang/
URL:https://datascience.utah.edu/talks/2025-02-14-bao-wang/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB L112 | https://utah.zoom.us/j/93909986581?pwd=d90LoHKoVAkagz1a
 CpJH1vTuaME9gG.1
CATEGORIES:algorithms & theory,statistics
END:VEVENT
BEGIN:VEVENT
UID:2025-02-21-chenglu-li@datascience.utah.edu
DTSTAMP:20250221T203000Z
DTSTART;TZID=America/Denver:20250221T133000
DTEND;TZID=America/Denver:20250221T143000
SUMMARY:Teaching and Learning at the Human-Technology Frontier: The Power o
 f Artificial Intelligence and Big Data - Chenglu Li
DESCRIPTION:Chenglu Li (Utah Edu Psych)\n\nIn this talk\, Chenglu will exam
 ine the rapidly evolving field of artificial intelligence in education (AI
 ED) by focusing on three key research gaps: agentic AI needs\, FAccT (fair
 ness\, accountability\, and transparency) challenges\, and issues related 
 to computing supremacy. He will illustrate these challenges through his cu
 rrent project\, ALTER-Math (AI-augmented Learning by Teaching to Enhance a
 nd Renovate Math Learning)\, a $10M initiative designed to accelerate midd
 le school math learning using generative AI-powered solutions. Additionall
 y\, Chenglu will discuss his contributions to advancing learning and teach
 ing through the development\, evaluation\, and dissemination of FAccT AI c
 yberinfrastructure. Finally\, he will outline promising future directions 
 for research and collaboration in this dynamic field.\n\nhttps://datascien
 ce.utah.edu/talks/2025-02-21-chenglu-li/
URL:https://datascience.utah.edu/talks/2025-02-21-chenglu-li/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB L112 | https://utah.zoom.us/j/93909986581?pwd=d90LoHKoVAkagz1a
 CpJH1vTuaME9gG.1
CATEGORIES:education,fairness & ethics,large language models
END:VEVENT
BEGIN:VEVENT
UID:2025-02-28-bernardo-modenesi@datascience.utah.edu
DTSTAMP:20250228T203000Z
DTSTART;TZID=America/Denver:20250228T133000
DTEND;TZID=America/Denver:20250228T143000
SUMMARY:Unveiling Hidden Patterns in Agent Behavior with Discrete-Choice an
 d Network Theory - Bernardo Modenesi
DESCRIPTION:Bernardo Modenesi (UU BioStats)\n\nMany datasets in data scienc
 e stem from agents repeatedly making choices over time\, with each choice 
 leading to an observable outcome. In this talk\, I introduce a novel appro
 ach to uncover latent agent heterogeneity\, enhancing both our understandi
 ng of agent behavior and causal inference estimation. By combining discret
 e choice models with network theory\, we develop a method to measure agent
  similarity based on their choice patterns. This results in a network-base
 d unsupervised clustering technique that groups agents with similar behavi
 ors—offering an interpretable alternative to black-box clustering models
  while maintaining explicit estimation assumptions. I will illustrate our 
 approach using labor market data\, where workers (agents) and jobs (choice
 s) form a bipartite network\, with worker-job matches represented as edges
 . By clustering workers based on their job choices\, we can infer unobserv
 ed worker skills—a crucial factor in economic analysis. Through Bayesian
  estimation\, we reveal latent worker groups\, improving predictions of la
 bor market outcomes and measuring labor market discrimination more effecti
 vely than models relying only on observable characteristics. This seminar 
 will detail our methodological framework\, estimation strategy\, and pract
 ical applications for understanding and predicting agent-choice dynamics.\
 n\nBonus project:\nIn the final portion of the talk\, I will pivot to a mo
 re informal discussion of a preliminary 'Model Ensemble Approach to Assess
 ing Discrimination in Machine Learning Models' in the space of algorithmic
  fairness.\n\nhttps://datascience.utah.edu/talks/2025-02-28-bernardo-moden
 esi/
URL:https://datascience.utah.edu/talks/2025-02-28-bernardo-modenesi/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB L112 | https://utah.zoom.us/j/93909986581?pwd=d90LoHKoVAkagz1a
 CpJH1vTuaME9gG.1
CATEGORIES:algorithms & theory,fairness & ethics,large language models,mach
 ine learning
END:VEVENT
BEGIN:VEVENT
UID:2025-03-21-jeff-phillips@datascience.utah.edu
DTSTAMP:20250321T193000Z
DTSTART;TZID=America/Denver:20250321T133000
DTEND;TZID=America/Denver:20250321T143000
SUMMARY:Robust High-Dimensional Mean Estimation With Low Data Size - Jeff P
 hillips
DESCRIPTION:Jeff Phillips (UU KSoC)\n\nRobust statistics aims to compute qu
 antities to represent data where a fraction of it may be arbitrarily corru
 pted. The most essential statistic is the mean\, and in recent years\,\nth
 ere has been a flurry of theoretical advancement for efficiently estimatin
 g the mean in high dimensions on corrupted data. While several algorithms 
 have been proposed that\nachieve near-optimal error\, they all rely on lar
 ge data size requirements as a function of dimension.\n\nIn this talk\, we
  perform an extensive experimentation over various mean estimation techniq
 ues where data size might not meet this requirement due to the high-dimens
 ional setting.\nFor data with inliers generated from a Gaussian with known
  covariance\, we find experimentally that several robust mean estimation t
 echniques can practically improve upon the sample mean\, with the quantum 
 entropy scaling approach from Dong et.al. (NeurIPS 2019) performing consis
 tently the best. However\, this consistent improvement is conditioned on a
  couple of simple modifications to how the steps to prune outliers work in
  the high-dimension\nlow-data setting\, and when the inliers deviate signi
 ficantly from Gaussianity. In fact\, with these modifications\, they are t
 ypically able to achieve roughly the same error as taking the sample mean 
 of the uncorrupted inlier data\, even with very low data size. In addition
  to\ncontrolled experiments on synthetic data\, we also explore these meth
 ods on large language models\, deep pretrained image models\, and non-cont
 extual word embedding models that do not necessarily have an inherent Gaus
 sian distribution. Yet\, in these settings\, a mean point of a set of embe
 dded objects is a desirable quantity to learn\, and the data exhibits the 
 high-dimension low-data setting studied in this paper. We show both the ch
 allenges of achieving this goal\, and that our updated robust mean estimat
 ion methods can provide\nsignificant improvement over using just the sampl
 e mean.\n\nhttps://datascience.utah.edu/talks/2025-03-21-jeff-phillips/
URL:https://datascience.utah.edu/talks/2025-03-21-jeff-phillips/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB L112 | https://utah.zoom.us/j/93909986581?pwd=d90LoHKoVAkagz1a
 CpJH1vTuaME9gG.1
CATEGORIES:natural language processing
END:VEVENT
BEGIN:VEVENT
UID:2025-03-28-seth-pettie@datascience.utah.edu
DTSTAMP:20250328T193000Z
DTSTART;TZID=America/Denver:20250328T133000
DTEND;TZID=America/Denver:20250328T143000
SUMMARY:Everything you always wanted to know about Cardinality - Seth Petti
 e
DESCRIPTION:Seth Pettie (U Michigan CS)\n\nThe Cardinality Estimation/Disti
 nct Elements problem is to approximate\nthe number of distinct elements in
  a data stream using a small\nprobabilistic data structure called a "sketc
 h". This problem has been\nstudied for 40 years\, has many industrial appl
 ications\, and is\nfeatured prominently in most courses on Big Data algori
 thmics. It is\ntherefore a real puzzle to explain why research on this pop
 ular and\nfundamental problem has been unusually slow.\n\nThis talk presen
 ts a complete history of the Cardinality Estimation\nproblem from Flajolet
  and Martin's seminal 1983 paper to the present\,\nand includes an account
  of how the research community became\nfractured\, delaying many natural d
 evelopments by decades. I will\npresent our recent efforts to achieve info
 rmation-theoretically\noptimal cardinality sketches\, which draws on two n
 otions of\n"information" developed in the 20th century: Fisher information
 \n(governing optimal point estimation) and Shannon entropy (governing\nopt
 imal space/communication).\n\nJoint work with Dingyu Wang.\n\n========\n\n
 https://datascience.utah.edu/talks/2025-03-28-seth-pettie/
URL:https://datascience.utah.edu/talks/2025-03-28-seth-pettie/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB L112 | https://utah.zoom.us/j/93909986581?pwd=d90LoHKoVAkagz1a
 CpJH1vTuaME9gG.1
CATEGORIES:algorithms & theory,education,society & policy
END:VEVENT
BEGIN:VEVENT
UID:2025-04-04-sabyasachi-basu@datascience.utah.edu
DTSTAMP:20250404T193000Z
DTSTART;TZID=America/Denver:20250404T133000
DTEND;TZID=America/Denver:20250404T143000
SUMMARY:"Triangles\, Communities\, and Dense Subgraphs" - Sabyasachi Basu
DESCRIPTION:Sabyasachi Basu (UCSC)\n\nIn this talk\, we will go over a few 
 recent results on dense subgraph discovery. We aim to discover 'many' dens
 e subgraphs of 'reasonable size' in real-world networks. We show that by l
 everaging triadic structure in graphs\, one can do this efficiently withou
 t complicated distributional assumptions on the input. Our techniques brid
 ge an important gap: most existing theory considers the setting where the 
 number of pieces is a constant\, whereas techniques that produce `satisfac
 tory' decompositions rarely have density guarantees (and indeed\, often gi
 ve sparse\, poorly connected subgraphs). We offer a community detection fl
 avor to our results: we provide a new metric for the 'goodness' of communi
 ties in terms of density and show that the spectrum of graph matrices impl
 ies the existence of communities.\n\nA key goal of this talk is to unpack 
 the several phrases in quotes in the preceding paragraph and offer some pe
 rspectives on why these are important (and sometimes difficult!). We also 
 provide an algorithm that (provably) decomposes large social networks into
  dense subgraphs in minutes on regular laptops\, and\, time permitting\, d
 iscuss the setting of overlapping subgraph detection using similar techniq
 ues.\n\nhttps://datascience.utah.edu/talks/2025-04-04-sabyasachi-basu/
URL:https://datascience.utah.edu/talks/2025-04-04-sabyasachi-basu/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB L112 | https://utah.zoom.us/j/93909986581?pwd=d90LoHKoVAkagz1a
 CpJH1vTuaME9gG.1
CATEGORIES:networks & graphs,society & policy
END:VEVENT
BEGIN:VEVENT
UID:2025-04-11-peter-jacobs@datascience.utah.edu
DTSTAMP:20250411T193000Z
DTSTART;TZID=America/Denver:20250411T133000
DTEND;TZID=America/Denver:20250411T143000
SUMMARY:Estimating High-Dimensional Zipfians & Language Distributions - Pet
 er Jacobs
DESCRIPTION:Peter Jacobs (UU KSoC\, Sandia)\n\nWe study estimation of large
  discrete distributions under the structural assumption that they follow a
  Zipfian distribution\, in which the ranking of alphabet items and/or leve
 l of decay need to be estimated from data. Empirical evidence for near Zip
 fian distributions has been found in diverse applications such as word and
  n-gram probability distributions in natural language text and chord proba
 bilities in musical pieces. We introduce the Sort and Snap estimator for w
 hen the level of decay is known but the ranking function needs to be estim
 ated\, and show it is minimax in several high dimensional regimes. When bo
 th the ranking and decay level are unknown\, we introduce an adaptive vari
 ant of Sort and Snap and show via Monte Carlo simulation that it outperfor
 ms state of the art discrete distribution estimators in these same high di
 mensional regimes. Our results motivate assessment of whether linguistical
 ly motivated marginal distributions for generating natural language that a
 re claimed to be Zipfian in quantitative linguistics communities are truly
  Zipfian. Through Monte Carlo experiments on one such well-regarded distri
 bution\, Sort and Snap procedures lag behind even the simplest non-paramet
 ric estimator (empirical proportions)\, which brings into focus that this 
 distribution thought to be Zipfian actually departs meaningfully from the 
 Zipfian pattern.\n\nhttps://datascience.utah.edu/talks/2025-04-11-peter-ja
 cobs/
URL:https://datascience.utah.edu/talks/2025-04-11-peter-jacobs/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB L112 | https://utah.zoom.us/j/93909986581?pwd=d90LoHKoVAkagz1a
 CpJH1vTuaME9gG.1
CATEGORIES:natural language processing,statistics
END:VEVENT
BEGIN:VEVENT
UID:2025-04-18-tucker-hermans@datascience.utah.edu
DTSTAMP:20250418T193000Z
DTSTART;TZID=America/Denver:20250418T133000
DTEND;TZID=America/Denver:20250418T143000
SUMMARY:Stein Variational Inference for Robotic Learning and Control - Tuck
 er Hermans
DESCRIPTION:Tucker Hermans (UU KSoC\, NVIDIA)\n\nProbabilistic inference\, 
 the problem of estimating a distribution given data\, has been a central p
 illar of robotic algorithms for more than two decades. Inference technique
 s have defined the de facto standard for robotic localization\, mapping\, 
 system calibration\, and online error estimation for mobile robots. Academ
 ic work has shown how these same inference problem formulations and algori
 thms can be used to solve problems of planning and control.\n\nIn this tal
 k I will discuss how probabilistic inference techniques can be extended fo
 r use in robotic manipulation where models may come either from engineerin
 g first-principles or in the form of large neural networks. I will then gi
 ve a brief dedication of Stein variational inference\, a recent nonparamet
 ric technique for probabilistic inference that is easily parallelized on m
 odern GPUs providing much faster inference times compared to more traditio
 nal Markov chain Monte Carlo methods. I will then show a few different app
 lications of using Stein variational inference from my lab\, including pla
 nning to goal distributions\, adaptive control of magnetic manipulation\, 
 and generating diverse data for training from real-world robot failures.\n
 \nhttps://datascience.utah.edu/talks/2025-04-18-tucker-hermans/
URL:https://datascience.utah.edu/talks/2025-04-18-tucker-hermans/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB L112 | https://utah.zoom.us/j/93909986581?pwd=d90LoHKoVAkagz1a
 CpJH1vTuaME9gG.1
CATEGORIES:algorithms & theory,robotics
END:VEVENT
BEGIN:VEVENT
UID:2025-08-20-varun-shankar@datascience.utah.edu
DTSTAMP:20250820T170000Z
DTSTART;TZID=America/Denver:20250820T110000
DTEND;TZID=America/Denver:20250820T120000
SUMMARY:Kernel Methods for Operator Learning - Varun Shankar
DESCRIPTION:Varun Shankar (Title: Kernel Methods for Operator Learning)\n\n
 https://datascience.utah.edu/talks/2025-08-20-varun-shankar/
URL:https://datascience.utah.edu/talks/2025-08-20-varun-shankar/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:LNCO 1100 | http://utah.zoom.us/my/vsutah
END:VEVENT
BEGIN:VEVENT
UID:2025-08-27-konstantin-genin@datascience.utah.edu
DTSTAMP:20250827T170000Z
DTSTART;TZID=America/Denver:20250827T110000
DTEND;TZID=America/Denver:20250827T120000
SUMMARY:Predictions as Public Reasons - Konstantin Genin
DESCRIPTION:Konstantin Genin\n\nhttps://datascience.utah.edu/talks/2025-08-
 27-konstantin-genin/
URL:https://datascience.utah.edu/talks/2025-08-27-konstantin-genin/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:LNCO 1100 | https://utah.zoom.us/j/7824755969?pwd=SUplL1EwUmU0TE5J
 RkhyQ2dtNmsvdz09
END:VEVENT
BEGIN:VEVENT
UID:2025-09-03-daniel-brown@datascience.utah.edu
DTSTAMP:20250903T170000Z
DTSTART;TZID=America/Denver:20250903T110000
DTEND;TZID=America/Denver:20250903T120000
SUMMARY:Swarms\, Emergent Behaviors\, and Multi-Agent Systems - Daniel Brow
 n
DESCRIPTION:Daniel Brown\n\nhttps://datascience.utah.edu/talks/2025-09-03-d
 aniel-brown/
URL:https://datascience.utah.edu/talks/2025-09-03-daniel-brown/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:LNCO 1100 | https://utah.zoom.us/j/81370106930?pwd=5rAKB2C2SrgkOpr
 GOGuAtzeF5xxbbT.1
CATEGORIES:large language models
END:VEVENT
BEGIN:VEVENT
UID:2025-09-10-bei-wang-phillips@datascience.utah.edu
DTSTAMP:20250910T170000Z
DTSTART;TZID=America/Denver:20250910T110000
DTEND;TZID=America/Denver:20250910T120000
SUMMARY:Talk by Bei Wang Phillips - Bei Wang Phillips
DESCRIPTION:Bei Wang Phillips\n\nhttps://datascience.utah.edu/talks/2025-09
 -10-bei-wang-phillips/
URL:https://datascience.utah.edu/talks/2025-09-10-bei-wang-phillips/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:LNCO 1100 | https://utah.zoom.us/j/81370106930?pwd=5rAKB2C2SrgkOpr
 GOGuAtzeF5xxbbT.1
END:VEVENT
BEGIN:VEVENT
UID:2025-09-17-aditya-bhaskara@datascience.utah.edu
DTSTAMP:20250917T170000Z
DTSTART;TZID=America/Denver:20250917T110000
DTEND;TZID=America/Denver:20250917T120000
SUMMARY:Descent with Misaligned Gradients and Applications to Hidden Convex
 ity - Aditya Bhaskara
DESCRIPTION:Aditya Bhaskara\n\nhttps://datascience.utah.edu/talks/2025-09-1
 7-aditya-bhaskara/
URL:https://datascience.utah.edu/talks/2025-09-17-aditya-bhaskara/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:LNCO 1100 | https://utah.zoom.us/j/81370106930?pwd=5rAKB2C2SrgkOpr
 GOGuAtzeF5xxbbT.1
CATEGORIES:optimization
END:VEVENT
BEGIN:VEVENT
UID:2025-09-24-kenneth-blake-vernon@datascience.utah.edu
DTSTAMP:20250924T170000Z
DTSTART;TZID=America/Denver:20250924T110000
DTEND;TZID=America/Denver:20250924T120000
SUMMARY:Indirect dating with mixture density networks - Kenneth Blake Verno
 n
DESCRIPTION:Kenneth Blake Vernon\n\nIt is an astonishing fact about the wor
 ld today that no one can say precisely how many people actually live on ou
 r planet\, even though it would presumably be useful to have such informat
 ion to plan for climate change and disaster risk management\, among other 
 things. Luckily for us\, demographers and spatial data scientists are keen
 ly aware of this problem and have devised sophisticated methods for interp
 olating population in these regions based on their built area\, typically 
 measured using remote sensing technology.\n\nAs it happens\, this is exact
 ly the reasoning applied by archaeologists seeking to reconstruct populati
 on sizes in the past. Unfortunately\, archaeology faces an additional chal
 lenge here\, since the archaeological record is a palimpsest of built area
 \, representing continuous human settlement over decades\, centuries\, and
  sometimes even millennia. So\, reconstructing population sizes across a r
 egion of interest requires that archaeologists also develop a chronology f
 or that region at the same time.\n\nA region’s chronology can be represe
 nted by a probability density function\, p(t)\, with well-dated archaeolog
 ical materials – like tree-rings and radiocarbon samples - assumed to be
  random draws from that distribution. Here\, we propose to estimate p usin
 g a deep-learning extension to the mixture model known as a Mixture Densit
 y Network. With this model\, we condition the chronology on diagnostic dat
 a X\, treating mixture parameters as unknown functions of X that can be es
 timated using a simple multilayer perceptron. Specifically\, we use the de
 nsity of X in the area around each sampled date to estimate the mixture pa
 rameters. This allows us to interpolate dates at under-sampled sites and t
 o build a better representation of the regional chronology\, one that is b
 ased\, in theory\, on the totality of the archaeological record.\n\nAs an 
 example\, we fit an MDN to the distribution of tree-rings in the Mesa Verd
 e region of southwestern Colorado using the spatial distribution of cerami
 cs to estimate mixture parameters. An important ancillary goal of this res
 earch is to develop software tools scientists can use to train MDNs on the
 ir own data without also having to learn the intricacies of AI development
  and testing.\n\nhttps://datascience.utah.edu/talks/2025-09-24-kenneth-bla
 ke-vernon/
URL:https://datascience.utah.edu/talks/2025-09-24-kenneth-blake-vernon/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:LNCO 1100 | https://utah.zoom.us/j/81370106930?pwd=5rAKB2C2SrgkOpr
 GOGuAtzeF5xxbbT.1
CATEGORIES:geospatial
END:VEVENT
BEGIN:VEVENT
UID:2025-10-01-luis-garcia@datascience.utah.edu
DTSTAMP:20251001T170000Z
DTSTART;TZID=America/Denver:20251001T110000
DTEND;TZID=America/Denver:20251001T120000
SUMMARY:A Trip to the Neural Frontier: Neurosymbolic Sensor Fusion for Trus
 tworthy AI-Enabled Neural Interventions - Luis Garcia
DESCRIPTION:Luis Garcia\n\nhttps://datascience.utah.edu/talks/2025-10-01-lu
 is-garcia/
URL:https://datascience.utah.edu/talks/2025-10-01-luis-garcia/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:LNCO 1100 | https://utah.zoom.us/j/81370106930?pwd=5rAKB2C2SrgkOpr
 GOGuAtzeF5xxbbT.1
END:VEVENT
BEGIN:VEVENT
UID:2025-10-15-vineet-pandey@datascience.utah.edu
DTSTAMP:20251015T170000Z
DTSTART;TZID=America/Denver:20251015T110000
DTEND;TZID=America/Denver:20251015T120000
SUMMARY:Designing human-centered systems that yield data that is minimal\, 
 relevant\, and actionable - Vineet Pandey
DESCRIPTION:Vineet Pandey\n\nhttps://datascience.utah.edu/talks/2025-10-15-
 vineet-pandey/
URL:https://datascience.utah.edu/talks/2025-10-15-vineet-pandey/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:LNCO 1100 | https://utah.zoom.us/j/81370106930?pwd=5rAKB2C2SrgkOpr
 GOGuAtzeF5xxbbT.1
CATEGORIES:human-centered computing
END:VEVENT
BEGIN:VEVENT
UID:2025-10-22-kenneth-marino@datascience.utah.edu
DTSTAMP:20251022T170000Z
DTSTART;TZID=America/Denver:20251022T110000
DTEND;TZID=America/Denver:20251022T120000
SUMMARY:VLM Agents - Kenneth Marino
DESCRIPTION:Kenneth Marino\n\nhttps://datascience.utah.edu/talks/2025-10-22
 -kenneth-marino/
URL:https://datascience.utah.edu/talks/2025-10-22-kenneth-marino/
SEQUENCE:0
STATUS:CANCELLED
TRANSP:OPAQUE
LOCATION:LNCO 1100 | https://utah.zoom.us/j/81370106930?pwd=5rAKB2C2SrgkOpr
 GOGuAtzeF5xxbbT.1
CATEGORIES:large language models
END:VEVENT
BEGIN:VEVENT
UID:2025-10-29-anna-fariha@datascience.utah.edu
DTSTAMP:20251029T170000Z
DTSTART;TZID=America/Denver:20251029T110000
DTEND;TZID=America/Denver:20251029T120000
SUMMARY:Understanding Data through Change Summarization and Causal Disparit
 y Explanations - Anna Fariha
DESCRIPTION:Anna Fariha\n\nhttps://datascience.utah.edu/talks/2025-10-29-an
 na-fariha/
URL:https://datascience.utah.edu/talks/2025-10-29-anna-fariha/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:LNCO 1100 | https://utah.zoom.us/j/81370106930?pwd=5rAKB2C2SrgkOpr
 GOGuAtzeF5xxbbT.1
CATEGORIES:causal inference,natural language processing
END:VEVENT
BEGIN:VEVENT
UID:2025-11-05-jenny-lin@datascience.utah.edu
DTSTAMP:20251105T180000Z
DTSTART;TZID=America/Denver:20251105T110000
DTEND;TZID=America/Denver:20251105T120000
SUMMARY:Talk by Jenny Lin - Jenny Lin
DESCRIPTION:Jenny Lin\n\nhttps://datascience.utah.edu/talks/2025-11-05-jenn
 y-lin/
URL:https://datascience.utah.edu/talks/2025-11-05-jenny-lin/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:LNCO 1100 | https://utah.zoom.us/j/81370106930?pwd=5rAKB2C2SrgkOpr
 GOGuAtzeF5xxbbT.1
END:VEVENT
BEGIN:VEVENT
UID:2025-11-12-kyle-dawson-tyler-hagen@datascience.utah.edu
DTSTAMP:20251112T170000Z
DTSTART;TZID=America/Denver:20251112T100000
DTEND;TZID=America/Denver:20251112T110000
SUMMARY:DESI: Disentangling Cosmology from Observational Artifacts - Kyle D
 awson\, Tyler Hagen
DESCRIPTION:Kyle Dawson\, Tyler Hagen\n\nThe Dark Energy Spectroscopic Inst
 rument (DESI) has concluded three years of observation\, leading to the la
 rgest spectroscopic galaxy sample ever produced. In combination with other
  cosmological probes\, these measurements reveal hints of new physics beyo
 nd the standard cosmological model. In this talk\, we will first present t
 he observations and key measurements that led to these new constraints. We
  will then describe the role that neural networks and random forests play 
 in this analysis and our tests of these machine learning algorithms agains
 t more physically-motivated\, linear models.\n\nWhere:\nThe location and t
 ime is moved for this talk so it can coincide with the CosmicAI seminar. I
 t will take place in WEB 3780 (the Evans Conference room in SCI) and at 10
 am. It will also be on Zoom at this *new* link:\n\nhttps://datascience.uta
 h.edu/talks/2025-11-12-kyle-dawson-tyler-hagen/
URL:https://datascience.utah.edu/talks/2025-11-12-kyle-dawson-tyler-hagen/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780 | https://utexas.zoom.us/j/87159746528?pwd=ouQu8lN9ARbb6a
 RFpvdf6Ddb1Oqa8B.1
CATEGORIES:machine learning,physics & astronomy
END:VEVENT
BEGIN:VEVENT
UID:2025-11-19-vivek-srikumar@datascience.utah.edu
DTSTAMP:20251119T180000Z
DTSTART;TZID=America/Denver:20251119T110000
DTEND;TZID=America/Denver:20251119T120000
SUMMARY:Talk by Vivek Srikumar - Vivek Srikumar
DESCRIPTION:Vivek Srikumar\n\nhttps://datascience.utah.edu/talks/2025-11-19
 -vivek-srikumar/
URL:https://datascience.utah.edu/talks/2025-11-19-vivek-srikumar/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:LNCO 1100 | https://utah.zoom.us/j/81370106930?pwd=5rAKB2C2SrgkOpr
 GOGuAtzeF5xxbbT.1
END:VEVENT
BEGIN:VEVENT
UID:2025-12-03-madison-golden-kaylee-alexander@datascience.utah.edu
DTSTAMP:20251203T180000Z
DTSTART;TZID=America/Denver:20251203T110000
DTEND;TZID=America/Denver:20251203T120000
SUMMARY:Data Visualization 101 - Madison Golden\, Kaylee Alexander
DESCRIPTION:Madison Golden\, Kaylee Alexander\n\nhttps://datascience.utah.e
 du/talks/2025-12-03-madison-golden-kaylee-alexander/
URL:https://datascience.utah.edu/talks/2025-12-03-madison-golden-kaylee-ale
 xander/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:LNCO 1100 | https://utah.zoom.us/j/81370106930?pwd=5rAKB2C2SrgkOpr
 GOGuAtzeF5xxbbT.1
CATEGORIES:visualization
END:VEVENT
BEGIN:VEVENT
UID:2026-01-09-marina-kogan@datascience.utah.edu
DTSTAMP:20260109T203000Z
DTSTART;TZID=America/Denver:20260109T133000
DTEND;TZID=America/Denver:20260109T143000
SUMMARY:Human-Centered Data Science for Crisis Informatics - Marina Kogan
DESCRIPTION:Marina Kogan (UU KSoC\, RAI Faculty Fellow)\n\nSocial media pla
 tforms have been increasingly used by the public in crisis situations\, pa
 rtly because they upend the traditional top-down broadcasting model of ris
 k communication. Instead\, social media platforms facilitate a two-way inf
 ormation exchange between the official response channels and the general p
 ublic\, enabling more participatory crisis communication\, as well as coor
 dination and self-organization among the public. In this more complex info
 rmation ecosystem\, understanding the flow of information is crucial to su
 pporting those affected and preventing malicious actors from capitalizing 
 on the uncertainty. However\, the study of such information flows is chall
 enging\, as the high-tempo\, high-volume convergent nature of crisis event
 s produces vast amounts of social media data\, necessitating the use of th
 e data science methods. On the other hand\, to glean meaningful insight fr
 om the crisis-related social media activity\, it is necessary to use metho
 ds that account for the complex social context of the user activity. In th
 is talk\, Kogan will show how the Human-Centered Data Science (HCDS) provi
 des methodological approaches that both harness the power of computational
  methods and account for the highly situated nature of social media activi
 ty in disruption. She will focus on sequence-based approaches as examples 
 of HCDS methods in two empirical studies: analysis of attention-garnering 
 information during a natural disaster and investigation of behavioral sign
 atures in coordinated information operations.\n\nhttps://datascience.utah.
 edu/talks/2026-01-09-marina-kogan/
URL:https://datascience.utah.edu/talks/2026-01-09-marina-kogan/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB L112 | https://utah.zoom.us/j/85983626630
CATEGORIES:human-centered computing,society & policy
END:VEVENT
BEGIN:VEVENT
UID:2026-01-23-andrew-mcnutt@datascience.utah.edu
DTSTAMP:20260123T203000Z
DTSTART;TZID=America/Denver:20260123T133000
DTEND;TZID=America/Denver:20260123T143000
SUMMARY:Linters as Socio Technical Systems - Andrew McNutt
DESCRIPTION:Andrew McNutt\n\nInterfaces—whether they are for data analysi
 s\, programming\, or any other activities—exist within specific communit
 ies of practice. The norms and standards of those groups inscribe themselv
 es in the form of those tools\; implicitly driving what is and is not poss
 ible. In this talk I will explore how linters (a spell checker-like tool u
 sed in programming) can be used to interrogate this arrangement. In doing 
 so I will describe recent and on-going efforts relating to application of 
 linters to a variety of domains\, including visualization\, color palettes
 \, and social media posts. Through this discussion\, I will argue that usi
 ng linters as a critical lens allows us to examine the technological world
  around us in a new light (offering new opportunities for research and des
 ign)\, and therein explore the values that we manifest in our tool design 
 and selection.\n\nhttps://datascience.utah.edu/talks/2026-01-23-andrew-mcn
 utt/
URL:https://datascience.utah.edu/talks/2026-01-23-andrew-mcnutt/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB L112 | https://utah.zoom.us/j/85983626630
CATEGORIES:human-centered computing,society & policy,visualization
END:VEVENT
BEGIN:VEVENT
UID:2026-02-06-makoto-kelp@datascience.utah.edu
DTSTAMP:20260206T203000Z
DTSTART;TZID=America/Denver:20260206T133000
DTEND;TZID=America/Denver:20260206T143000
SUMMARY:Navigating Advances and Inflections in Machine Learning for Atmosph
 eric Chemistry Modeling - Makoto Kelp
DESCRIPTION:Makoto Kelp\n\nGlobal climate and Earth system models rarely in
 clude comprehensive atmospheric chemistry because of its high computationa
 l cost. A bottleneck is the chemical solver that integrates the large-dime
 nsional coupled systems of kinetic equations describing the chemical mecha
 nism. In recent years\, machine learning (ML) methods have been proposed a
 s a potentially transformative approach to reducing this cost by replacing
  traditional solvers with fast emulators. However\, early efforts showed t
 hat ML-based chemical solvers often suffer from rapid error growth and ins
 tability. In this talk\, I will review the evolving landscape of ML for at
 mospheric chemistry modeling over the past decade and how its trajectory m
 irrors broader developments in climate and weather AI. I begin with detail
 ing how to achieve stable emulation in 0-D box models and then show how th
 ese principles translate to complex global atmospheric models. I conclude 
 by discussing the current state of ML for modeling atmospheric chemistry a
 nd outlining how mechanistic interpretability in geospatial foundation mod
 els may offer a promising future line of inquiry.\n\nhttps://datascience.u
 tah.edu/talks/2026-02-06-makoto-kelp/
URL:https://datascience.utah.edu/talks/2026-02-06-makoto-kelp/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB L112 | https://utah.zoom.us/j/85983626630
CATEGORIES:climate & environment,machine learning,optimization
END:VEVENT
BEGIN:VEVENT
UID:2026-02-17-juliana-freire@datascience.utah.edu
DTSTAMP:20260217T180000Z
DTSTART;TZID=America/Denver:20260217T110000
DTEND;TZID=America/Denver:20260217T120000
SUMMARY:Dataset Discovery and Integration in the Era of Large Language Mode
 ls - Juliana Freire
DESCRIPTION:Juliana Freire\n\n**Abstract:**\nThe proliferation of structure
 d data across open-data portals\, the web\, and enterprise data lakes pres
 ents unprecedented opportunities for scientific discovery and data-driven 
 decision-making. However\, realizing this potential requires solving funda
 mental challenges in data discovery and integration: How do we find releva
 nt datasets among millions of candidates? How do we understand their seman
 tics to integrate them? And how do we build systems that are scalable\, ac
 curate\, and cost-effective?\n\nIn this talk\, I will present our recent w
 ork addressing these challenges by combining techniques from data manageme
 nt\, visualization\, HCI\, and modern language models. Our work is motivat
 ed by real-world problems across domains: supporting biomedical researcher
 s in integrating heterogeneous datasets\, enabling dataset search over urb
 an data for policy analysis and planning\, and augmenting training data to
  improve machine learning model performance.\n\nFirst\, I will describe ho
 w we are reimagining dataset discovery by designing specialized search eng
 ines\, developing methods to automatically derive metadata\, and introduci
 ng novel data-driven queries that go beyond keywords to support complex in
 formation needs. Then\, I will turn to data integration\, showing how we c
 an leverage the semantic power of Large Language Models (LLMs) to match sc
 hemas with state-of-the-art accuracy. I will also argue for the critical i
 mportance of human-AI collaboration\, demonstrating how visual analytics c
 an effectively place the user in the loop to guide and verify these comple
 x processes.\n\nFinally\, I will reflect on how LLMs are fundamentally cha
 nging computer science research. They make previously intractable problems
  like semantic data discovery and integration tractable\, yet they behave 
 more like natural phenomena\, exhibiting variability and non-determinism. 
 This shift forces us to move beyond deterministic algorithmic thinking and
  to practice computer science as a science\, adopting empirical research m
 ethods to understand and leverage these powerful but unpredictable tools.\
 n\n**Bio:**\nJuliana Freire is an Institute Professor at the Tandon School
  of Engineering and Professor of Computer Science and Data Science at New 
 York University\, where she co-directs the Visualization Imaging and Data 
 Analysis (VIDA) Center. Her research develops methods and systems that ena
 ble a wide range of users to obtain trustworthy insights from data. It spa
 ns topics in large-scale data analysis and integration\, visualization\, m
 achine learning\, provenance management\, and web information discovery\, 
 addressing application areas including urban analytics\, predictive modeli
 ng\, computational reproducibility\, and biomedical data harmonization. Sh
 e has co-authored over 250 papers\, including 12 award winners and a test-
 of-time award. She served as elected chair of ACM SIGMOD and as a council 
 member of the Computing Community Consortium (CCC)\, and was the NYU lead 
 investigator for the Moore-Sloan Data Science Environment. She is a Fellow
  of the ACM and AAAS\, and a winner of the ACM SIGMOD Contributions Award.
  Her work has been supported by funding agencies and industry partners\, i
 ncluding the National Science Foundation\, DARPA\, ARPA-H\, the Department
  of Energy\, the National Institutes of Health\, and technology companies 
 such as Google\, Amazon\, Microsoft Research\, and IBM. Freire received he
 r Ph.D. and M.Sc. degrees in computer science from the State University of
  New York at Stony Brook and her B.S. degree in computer science from the 
 Federal University of Ceará in Brazil.\n\nhttps://datascience.utah.edu/ta
 lks/2026-02-17-juliana-freire/
URL:https://datascience.utah.edu/talks/2026-02-17-juliana-freire/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780 (Evans) | https://utah.zoom.us/j/85983626630
CATEGORIES:data management,large language models,natural language processin
 g,visualization
END:VEVENT
BEGIN:VEVENT
UID:2026-02-19-dinesh-manocha@datascience.utah.edu
DTSTAMP:20260219T223000Z
DTSTART;TZID=America/Denver:20260219T153000
DTEND;TZID=America/Denver:20260219T163000
SUMMARY:Robot Navigation in the Wild - Dinesh Manocha
DESCRIPTION:Dinesh Manocha\n\nIn the last few decades\, most robotics succe
 ss stories have been limited to structured or controlled environments. A m
 ajor challenge is to develop robot systems that can operate in complex or 
 unstructured environments corresponding to homes\, dense traffic\, outdoor
  terrains\, public places\, etc. In this talk\, we give an overview of our
  ongoing work on developing robust planning and navigation technologies th
 at use recent advances in computer vision\, sensor technologies\, machine 
 learning\, and motion planning algorithms. We present new methods that uti
 lize multi-modal observations from an RGB camera\, 3D LiDAR\, and robot od
 ometry for scene perception\, along with deep reinforcement learning for r
 eliable planning. The latter is also used to compute dynamically feasible 
 and spatial aware velocities for a robot navigating among mobile obstacles
  and uneven terrains. We have integrated these methods with wheeled robots
 \, home robots\, and legged platforms and highlight their performance in c
 rowded indoor scenes\, home environments\, and dense outdoor terrains.\n\n
 https://datascience.utah.edu/talks/2026-02-19-dinesh-manocha/
URL:https://datascience.utah.edu/talks/2026-02-19-dinesh-manocha/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780
CATEGORIES:machine learning,physics & astronomy,robotics
END:VEVENT
BEGIN:VEVENT
UID:2026-02-24-paul-parsons@datascience.utah.edu
DTSTAMP:20260224T173000Z
DTSTART;TZID=America/Denver:20260224T103000
DTEND;TZID=America/Denver:20260224T113000
SUMMARY:Visualization and Judgment: Human-Centered Computing in Data-Rich P
 ractice - Paul Parsons
DESCRIPTION:Paul Parsons\n\nVisualization and computational systems are inc
 reasingly powerful\, but their impact depends on a persistent\, often unde
 r-specified factor—human judgment. Designers frame problems and negotiat
 e constraints\; users interpret\, challenge\, and coordinate action around
  system outputs over time\, under uncertainty and constraint. In this talk
 \, I present a research agenda on visualization and judgment in data-rich 
 work\, grounded in empirical studies and design-oriented analyses across t
 hree contexts. First\, I study data visualization design practice\, showin
 g how professional designers frame problems and co-evolve problem and solu
 tion spaces—work that is often invisible in pipeline-oriented accounts o
 f visualization. Second\, I examine expert judgment in complex sociotechni
 cal settings\, where visualization\, procedures\, and automation reshape d
 ecision spaces and can either support adaptive performance or encourage br
 ittle reliance. Third\, I extend these insights to scientific cyberinfrast
 ructure\, where platforms exposing advanced computation and data services 
 succeed or fail based on whether diverse communities can understand\, adop
 t\, and sustain capabilities in practice. This perspective complements adv
 ances in modeling\, simulation\, and visual analytics by surfacing design 
 constraints\, evaluation targets\, and failure modes that matter as system
 s intersect with the constraints and contingencies of practice. I close wi
 th directions for judgment-aware visualization and human–AI support that
  preserve interpretability\, accountability\, and coordination in conseque
 ntial domains.\n\nhttps://datascience.utah.edu/talks/2026-02-24-paul-parso
 ns/
URL:https://datascience.utah.edu/talks/2026-02-24-paul-parsons/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780
CATEGORIES:human-centered computing,visualization
END:VEVENT
BEGIN:VEVENT
UID:2026-02-26-zezhong-wang@datascience.utah.edu
DTSTAMP:20260226T173000Z
DTSTART;TZID=America/Denver:20260226T103000
DTEND;TZID=America/Denver:20260226T113000
SUMMARY:Designing Narrative-Driven Data Experiences - Zezhong Wang
DESCRIPTION:Zezhong Wang\n\nIn today’s data-saturated world\, people are 
 facing an “infodemic” of information overload. As public demand to int
 erpret and use data grows\, so does the urgency to rethink how we help bro
 ader audiences engage with data in meaningful ways. This talk explores how
  we can reconnect data with its storytelling roots to support understandin
 g and agency. Dr. Wang will present research at the intersection of visual
  design\, data visualization\, and human-computer interaction\, showing ho
 w interdisciplinary collaboration can open up new modes of communication. 
 He will share empirical findings on how visual data narratives\, such as d
 ata comics\, can improve comprehension and engagement\, as well as methods
  for crafting visual data stories that connect data to context and lived e
 xperience. Examples will draw from cross-disciplinary projects in environm
 ental and healthcare data storytelling. The talk will conclude with future
  directions toward embedding narrative-driven data experiences into everyd
 ay tasks and decision-making.\n\nhttps://datascience.utah.edu/talks/2026-0
 2-26-zezhong-wang/
URL:https://datascience.utah.edu/talks/2026-02-26-zezhong-wang/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:Evans Conference Room (WEB 3780)
CATEGORIES:climate & environment,health & medicine,human-centered computing
 ,visualization
END:VEVENT
BEGIN:VEVENT
UID:2026-02-27-kenny-marino@datascience.utah.edu
DTSTAMP:20260227T204000Z
DTSTART;TZID=America/Denver:20260227T134000
DTEND;TZID=America/Denver:20260227T144000
SUMMARY:Agents: Hype or Opportunity - Kenny Marino
DESCRIPTION:Kenny Marino\n\nAre so-called "AI Agents" a fad or a potentiall
 y impactful research area enabled by the rapid progress in large language 
 models? In this talk I will try to strip away the marketing copy and look 
 at what an agent actually is\, returning to the classical understanding of
  the word\, and investigate how powerful new language models can present n
 ew opportunities for research in embodied decision making. We begin by rec
 ounting the places where language models have been useful in agent-like pr
 oblems\, then looking at the emerging environments for investigating VLM/L
 LM agents including computer use and robotics\, and finally discussing the
  frontier research challenges of LLM agents.\n\nhttps://datascience.utah.e
 du/talks/2026-02-27-kenny-marino/
URL:https://datascience.utah.edu/talks/2026-02-27-kenny-marino/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB L112 | https://utah.zoom.us/j/85983626630
CATEGORIES:large language models,natural language processing,robotics
END:VEVENT
BEGIN:VEVENT
UID:2026-03-02-bogdan-raita@datascience.utah.edu
DTSTAMP:20260302T230000Z
DTSTART;TZID=America/Denver:20260302T160000
DTEND;TZID=America/Denver:20260302T170000
SUMMARY:Solving Linear PDE by Machine Learning and Commutative Algebra - Bo
 gdan Raita
DESCRIPTION:Bogdan Raita\n\nWe use the theory of linear pde systems with co
 nstant coefficients (Malgrange\, Palamodov\, Pommaret\, Sturmfels) to impl
 ement a machine learning algorithm which generates solutions to arbitrary 
 linear pdes. Since we preprocess the equations with computer algebra\, our
  methods are applicable to arbitrary pde systems\, irrespective of type (e
 lliptic\, hyperbolic\, etc.) or order. We test our method for classical eq
 uations (wave\, heat\, Laplace) and discuss future applications to equatio
 ns describing wave-related phenomena\, for example direct and inverse prob
 lems involving Maxwellâ€™s system and the elasticity equations.\n\nht
 tps://datascience.utah.edu/talks/2026-03-02-bogdan-raita/
URL:https://datascience.utah.edu/talks/2026-03-02-bogdan-raita/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:LCB 222
CATEGORIES:machine learning,physics & astronomy
END:VEVENT
BEGIN:VEVENT
UID:2026-03-03-grace-guo@datascience.utah.edu
DTSTAMP:20260303T173000Z
DTSTART;TZID=America/Denver:20260303T103000
DTEND;TZID=America/Denver:20260303T113000
SUMMARY:Concepts and Counterfactuals: Human-Centered Interpretability in th
 e Age of Foundation Models - Grace Guo
DESCRIPTION:Grace Guo\n\nFoundation models are increasingly deployed in hig
 h-stakes domains\, yet their scale and opacity challenge traditional notio
 ns of AI interpretability. In this talk\, I present two complementary stra
 tegies for human-centered interpretability: reasoning through concepts and
  probing through counterfactuals. I first present MiMICRI\, a visualizatio
 n tool developed with doctors at Cleveland Clinic that enables them to int
 eractively create counterfactual medical images to examine how anatomical 
 changes influence model predictions. By grounding explanations in domain-r
 elevant visual features\, this tool helps experts reason about model behav
 ior using their established medical knowledge. Next\, I will introduce Con
 cept2Concept\, a framework for auditing text-to-image models by characteri
 zing their outputs as distributions over named\, interpretable concepts. B
 y analyzing the metrics of concept frequency\, stability\, and co-occurren
 ce\, we uncover hidden and sometimes harmful associations in image generat
 ion models and real-world training datasets. Finally\, I conclude with my 
 research agenda for developing new visualization tools and theoretical fou
 ndations that address the ongoing challenges of auditing and aligning the 
 foundation models of today.\n\nhttps://datascience.utah.edu/talks/2026-03-
 03-grace-guo/
URL:https://datascience.utah.edu/talks/2026-03-03-grace-guo/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:Evans Conference Room (WEB 3780)
CATEGORIES:causal inference,health & medicine,human-centered computing,larg
 e language models
END:VEVENT
BEGIN:VEVENT
UID:2026-03-05-josh-levine@datascience.utah.edu
DTSTAMP:20260305T173000Z
DTSTART;TZID=America/Denver:20260305T103000
DTEND;TZID=America/Denver:20260305T113000
SUMMARY:Extracting\, Visualizing\, and Analyzing Topological Features with 
 Discrete Representations - Josh Levine
DESCRIPTION:Josh Levine\n\nTopological features provide multi-scale summari
 es of the behavior of continuous data from diverse applications ranging fr
 om astrophysics to medicine. Nevertheless\, computing them robustly is cha
 llenging due to numerical precision issues. A promising strategy is to fir
 st convert the input to a discrete representation that satisfies criteria 
 introduced by Forman's discrete Morse theory. While numerous approaches ex
 ist to discretize the restricted case of gradient fields from scalar data\
 , state-of-the-art algorithms for the general case of vector fields requir
 e expensive optimization procedures. In this talk\, I will present recent 
 work that uses local evaluation to create discrete vector fields in linear
  time from two-dimensional\, triangulated vector fields. I will also frame
  this work within my contributions to the Topological ToolKit (TTK)\, an o
 pen source software platform for topological data analysis\, led by collab
 orators at UPMC Sorbonne.\n\nhttps://datascience.utah.edu/talks/2026-03-05
 -josh-levine/
URL:https://datascience.utah.edu/talks/2026-03-05-josh-levine/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:Evans Conference Room (WEB 3780)
CATEGORIES:visualization
END:VEVENT
BEGIN:VEVENT
UID:2026-03-16-tenghao-huang@datascience.utah.edu
DTSTAMP:20260316T160000Z
DTSTART;TZID=America/Denver:20260316T100000
DTEND;TZID=America/Denver:20260316T110000
SUMMARY:Generalizable\, Proactive\, and Agentic Learning for Open-Ended Tas
 ks - Tenghao Huang
DESCRIPTION:Tenghao Huang\n\nOpen-ended\, human-like intelligence requires 
 flexibility\, proactivity\, and social intelligence: we must learn subject
 ive goals and adapt to novel\, complex scenarios. . Current AI systems str
 uggle to learn these behaviors because reward signals are unclear\, task c
 ontext information is incomplete\, and the environment lacks observability
 . I outline a research agenda focused on building agentic learning systems
  that operate under such uncertainty. (1) Instead of enumerating task-spec
 ific heuristics\, I propose training AI systems through self-play in an ad
 versarial setting\, where models learn by interacting\, critiquing\, and i
 mproving against dynamically evolving counterparts. (2) I equip agents wit
 h the ability to proactively gather missing information when task context 
 is incomplete and 3) I reconstruct environment representations through mem
 ory to support long-horizon agentic reasoning. Together\, I will show\, th
 ese techniques enable AI systems to learn abstract objectives in tasks as 
 varied as creative writing\, multi-agent coordination and long-horizon pro
 blem solving\, enhancing the creativity\, usefulness\, and strategic helpf
 ulness of agents in open-ended environments.\n\nhttps://datascience.utah.e
 du/talks/2026-03-16-tenghao-huang/
URL:https://datascience.utah.edu/talks/2026-03-16-tenghao-huang/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780 (Evans)
CATEGORIES:large language models
END:VEVENT
BEGIN:VEVENT
UID:2026-03-20-xueguang-ma@datascience.utah.edu
DTSTAMP:20260320T160000Z
DTSTART;TZID=America/Denver:20260320T100000
DTEND;TZID=America/Denver:20260320T110000
SUMMARY:Breaking Information Silos: Advancing Search Systems for Unified In
 formation Seeking - Xueguang Ma
DESCRIPTION:Xueguang Ma\n\nInformation seeking has been fundamental to huma
 n advancement\, enabling knowledge acquisition\, decision-making\, and inn
 ovation across disciplines. However\, traditional information retrieval sy
 stems often rely on specialized pipelines optimized for specific retrieval
  tasks\, causing information silos that hinder unified information seeking
 . In this talk\, I will present our work in building unified document retr
 ieval systems that break these information silos across three dimensions: 
 (1) domain and language silos\, where I demonstrate how LLM-based dense re
 trievers achieve strong generalizability across retrieval tasks and presen
 t frameworks for training small\, generalizable retrievers through diverse
  LLM augmentation\; (2) modality silos\, where I introduce a paradigm shif
 t from text-based retrieval that relies on content extraction to directly 
 encoding document screenshots\, preserving all information including text\
 , images\, and layout in unified dense representations\; and (3) space sil
 os\, where we show the importance of LLM-powered search agents in seeking 
 and gathering information across disparate sources\, and present fair and 
 transparent evaluation benchmarks for assessing deep-search systems. I wil
 l conclude by discussing future directions that further pave the way towar
 d building truly unified retrieval systems for seamless information seekin
 g across world knowledge.\n\nhttps://datascience.utah.edu/talks/2026-03-20
 -xueguang-ma/
URL:https://datascience.utah.edu/talks/2026-03-20-xueguang-ma/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MEB 3147 (LCR)
CATEGORIES:large language models,natural language processing
END:VEVENT
BEGIN:VEVENT
UID:2026-03-20-kate-isaacs@datascience.utah.edu
DTSTAMP:20260320T193000Z
DTSTART;TZID=America/Denver:20260320T133000
DTEND;TZID=America/Denver:20260320T143000
SUMMARY:A Matter of Audiences: Capturing and Reporting Reasoning and Result
 s Around Data Visualizations - Kate Isaacs
DESCRIPTION:Kate Isaacs\n\nData visualizations are used throughout the data
  science process to facilitate the exploratory analysis and to report resu
 lts. Frequently\, these uses are neither separate nor solitary\, acting as
  a medium for data science teams to collaboratively reason about data. Thi
 s team-based data science work is often fast-paced and involves several fo
 rms of non-digital communication. I will discuss how we leverage gesture\,
  sketch\, and speech to create support for common meetings around data\, s
 pecifically collaborative remote meetings and informal presentations\, to 
 aid this aspect of data science work. Then\, focusing on the wider dissemi
 nation of results\, I will discuss findings regarding the public's views o
 n the use of data and AI in science videos.\n\nhttps://datascience.utah.ed
 u/talks/2026-03-20-kate-isaacs/
URL:https://datascience.utah.edu/talks/2026-03-20-kate-isaacs/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB L112 | https://utah.zoom.us/j/85983626630
CATEGORIES:visualization
END:VEVENT
BEGIN:VEVENT
UID:2026-03-23-benjie-wang@datascience.utah.edu
DTSTAMP:20260323T160000Z
DTSTART;TZID=America/Denver:20260323T100000
DTEND;TZID=America/Denver:20260323T110000
SUMMARY:Bridging the Formalization Gap for Generative AI - Benjie Wang
DESCRIPTION:Benjie Wang\n\nGenerative models\, such as large language model
 s and diffusion models\, have tremendously increased the scope of problems
  that AI can address. As such\, there is a significant trend toward incorp
 orating generative AI to automate tasks across computing and more broadly\
 , from controlling robotics systems\, to software generation and testing\,
  to searching over scientific knowledge. However\, there remains a signifi
 cant formalization gap between the domain knowledge\, theories\, and logic
 al and semantic constraints that are vital to applications\, and the stati
 stical patterns over natural data represented by large generative models. 
 In this talk\, I will demonstrate how we can systematically bridge this fo
 rmalization gap towards more trustworthy AI. First\, drawing from examples
  and applications in my research\, I will show how we can utilize suitable
  intermediate representations of probability distributions to bridge betwe
 en formal language and generative models at scale. Then\, I will discuss h
 ow these practical methods are underpinned by my work advancing the mathem
 atical and computational foundations underlying these tractable representa
 tions of probability distributions.\n\nhttps://datascience.utah.edu/talks/
 2026-03-23-benjie-wang/
URL:https://datascience.utah.edu/talks/2026-03-23-benjie-wang/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780
CATEGORIES:large language models,natural language processing,robotics,stati
 stics
END:VEVENT
BEGIN:VEVENT
UID:2026-03-25-dick-sadler@datascience.utah.edu
DTSTAMP:20260325T220000Z
DTSTART;TZID=America/Denver:20260325T160000
DTEND;TZID=America/Denver:20260325T170000
SUMMARY:What Happens in the 10 Years Following a State Government-Caused En
 vironmental Injustice? Flint's Progress Since Its Water Crisis - Dick Sadl
 er
DESCRIPTION:Dick Sadler\n\nThe Flint Water Crisis was the result of decades
  of deliberate disinvestment and state government ineptitude. In this talk
 \, Dr. Sadler will discuss how the crisis unfolded\, and how his research 
 - examining environmental exposures\, neighborhood conditions\, and blood 
 lead levels - revealed its scale. He will also address what has changed in
  Flint since then\, including the massive new investments in the city that
  have brought hope in the wake of catastrophe.\n\nhttps://datascience.utah
 .edu/talks/2026-03-25-dick-sadler/
URL:https://datascience.utah.edu/talks/2026-03-25-dick-sadler/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:MLIB 1110 | Zoom: 890 9876 9672\, Passcode: 156565
CATEGORIES:climate & environment,society & policy
END:VEVENT
BEGIN:VEVENT
UID:2026-03-26-bailing-lyu@datascience.utah.edu
DTSTAMP:20260326T163000Z
DTSTART;TZID=America/Denver:20260326T103000
DTEND;TZID=America/Denver:20260326T113000
SUMMARY:Navigating the Human-AI Nexus in Education: Bridging Cognitive Scie
 nce\, Learning Analytics\, and Intelligent Systems - Bailing Lyu
DESCRIPTION:Bailing Lyu\n\nArtificial intelligence is increasingly transfor
 ming education\, reshaping how students engage with learning and how instr
 uctors design and deliver instruction. However\, critical gaps remain in t
 he field of AI in Education (AIED): (1) a frequent lack of grounding in th
 e cognitive and learning sciences when developing AI pedagogical tools\, (
 2) a "black box" regarding how students process and interact with AI-suppo
 rted environments\, and (3) a limited emphasis on meaningful human involve
 ment in how AI is applied in practice. This talk presents a series of rese
 arch projects on AI-augmented learning and teaching that address these gap
 s by integrating cognitive science into AI-powered educational technologie
 s\, examining human-AI interaction\, and centering human agency in their a
 pplication. Specifically\, using teachable agents as an example\, it will 
 discuss (1) how pedagogical AI can be designed using theoretically grounde
 d learning principles\, (2) how students' learning processes unfold in the
 se environments\, and (3) how educators and learners can actively leverage
  AI to co-create and navigate engaging experiences that enhance learning o
 utcomes. Employing experimental research\, learning analytics\, and educat
 ional data mining\, these studies examine the intersection of human cognit
 ion and AI-driven learning. By aligning AI with evidence-based pedagogical
  strategies\, this work advances our understanding of how intelligent tech
 nologies can foster deeper learning\, positive learning experiences\, pers
 onalized instruction\, and adaptive support across diverse educational con
 texts.\n\nhttps://datascience.utah.edu/talks/2026-03-26-bailing-lyu/
URL:https://datascience.utah.edu/talks/2026-03-26-bailing-lyu/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:Evans Conference room (WEB 3780) | https://utah.zoom.us/j/87171666
 093
CATEGORIES:education,human-centered computing
END:VEVENT
BEGIN:VEVENT
UID:2026-03-27-erdogan-kaya@datascience.utah.edu
DTSTAMP:20260327T163000Z
DTSTART;TZID=America/Denver:20260327T103000
DTEND;TZID=America/Denver:20260327T113000
SUMMARY:From Self-Efficacy to Systemic Change: Building an Equitable Comput
 ing Education Research Program - Erdogan Kaya
DESCRIPTION:Erdogan Kaya\n\nThis talk examines a foundational question in c
 omputational STEM education: to what extent can targeted interventions imp
 rove pre-service elementary teachers’ computational thinking teaching ef
 ficacy beliefs? I present a study of pre-service elementary teachers who p
 articipated in a three-week CT intervention integrating EV3 robotics\, Cod
 e.org\, and Zoombinis. Using the CTTEBI in a pre-post design\, results sho
 w significant gains in personal CT teaching efficacy\, pointing to importa
 nt directions for future work. I then present my current AI education rese
 arch program\, including the “Educate AI” project developing AI litera
 cy curriculum through linguistically inclusive elementary robotics\, the R
 ural AI project integrating AI concepts for rural upper elementary student
 s\, and the Compose with AI platform guiding grades 4-8 students in critic
 ally evaluating AI-generated content. I will also discuss my emerging rese
 arch agenda\, including several NSF proposals under review focused on AI l
 iteracy across K-16 settings. I close with a vision for how University of 
 Utah’s collaborative structure across the Scientific Computing and Imagi
 ng (SCI) Institute\, Department of Educational Psychology\, College of Edu
 cation\, and partner departments represents an ideal ecosystem to advance 
 this agenda.\n\nhttps://datascience.utah.edu/talks/2026-03-27-erdogan-kaya
 /
URL:https://datascience.utah.edu/talks/2026-03-27-erdogan-kaya/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:Evans Conference room (WEB 3780)
CATEGORIES:education,fairness & ethics,robotics
END:VEVENT
BEGIN:VEVENT
UID:2026-03-30-xiaoling-hu@datascience.utah.edu
DTSTAMP:20260330T160000Z
DTSTART;TZID=America/Denver:20260330T100000
DTEND;TZID=America/Denver:20260330T110000
SUMMARY:Principled Learning for Medical AI: Structure\, Reliability\, and I
 nterpretability - Xiaoling Hu
DESCRIPTION:Xiaoling Hu\n\nThe widespread deployment of AI in medicine dema
 nds not only predictive accuracy but also structural awareness\, reliabili
 ty under uncertainty\, and interpretability for clinical trust. In this ta
 lk\, I will present a unified research agenda toward principled learning f
 or medical AI\, grounded in these core pillars.\n\nFirst\, I will discuss 
 how incorporating explicit structure\, such as topology and spatial priors
 \, into neural networks enhances the model's ability to reason about fine-
 grained anatomical and pathological features\, which are critical for task
 s like brain and tumor segmentation. Second\, I will focus on reliability\
 , exploring how we can quantify and mitigate uncertainty arising from impe
 rfect labels\, limited data\, and domain shifts\, using methods such as di
 stributional modeling\, hyperparameter learning\, and probabilistic infere
 nce. Third\, I will show how these approaches naturally support interpreta
 bility\, enabling AI systems to communicate meaningful representations tha
 t align with human clinical understanding.\n\nThrough applications in radi
 ology\, pathology\, neuroimaging\, and large-scale population datasets\, I
  will demonstrate how these principles facilitate scalable annotation\, ro
 bust generalization\, and scientific discovery. I will conclude with futur
 e directions aimed at generalizing these principles to multimodal learning
 \, real-world deployment\, and next-generation AI systems in medicine.\n\n
 https://datascience.utah.edu/talks/2026-03-30-xiaoling-hu/
URL:https://datascience.utah.edu/talks/2026-03-30-xiaoling-hu/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780
CATEGORIES:computer vision,health & medicine,machine learning
END:VEVENT
BEGIN:VEVENT
UID:2026-03-31-he-yin@datascience.utah.edu
DTSTAMP:20260331T164500Z
DTSTART;TZID=America/Denver:20260331T104500
DTEND;TZID=America/Denver:20260331T114500
SUMMARY:Advancing GeoAI and Earth observation for environmental monitoring 
 - He Yin
DESCRIPTION:He Yin\n\nLandscapes around the world are changing rapidly\, wi
 th important consequences for sustainability\, climate resilience\, and so
 ciety. Yet monitoring these changes across regions and scales remains diff
 icult. In this talk\, I present a research program that combines multi-sen
 sor Earth observation\, geospatial artificial intelligence (GeoAI)\, and l
 and system science to better understand how land systems are changing\, wh
 at drives those changes\, and why they matter.\n\nI begin by presenting my
  studies using satellite image time series to map land use change and\, in
  collaboration with environmental scientists and ecologists\, to examine i
 ts implications for carbon sequestration and biodiversity. These studies a
 lso reveal key limitations of conventional remote sensing approaches\, inc
 luding sensor constraints\, limited transferability\, and scarce training 
 data. I then show how these challenges motivate my more recent work in sen
 sor fusion\, physics-informed machine learning\, and deep learning with ve
 ry-high-resolution imagery — applied to problems ranging from irrigation
  water use and wildfire-invasive species interactions to conflict-induced 
 environmental damage. I conclude by discussing the broader goal of buildin
 g GeoAI models for environmental monitoring that are informed by physical 
 processes\, transferable across contexts\, and useful for real-world decis
 ion-making.\n\nhttps://datascience.utah.edu/talks/2026-03-31-he-yin/
URL:https://datascience.utah.edu/talks/2026-03-31-he-yin/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:Evans Conference room (WEB 3780)
CATEGORIES:climate & environment,geospatial,machine learning,society & poli
 cy
END:VEVENT
BEGIN:VEVENT
UID:2026-04-01-md-mostafijur-rahman@datascience.utah.edu
DTSTAMP:20260401T163000Z
DTSTART;TZID=America/Denver:20260401T103000
DTEND;TZID=America/Denver:20260401T113000
SUMMARY:Efficient and Reliable AI for Real-World Healthcare Deployment - Md
  Mostafijur Rahman
DESCRIPTION:Md Mostafijur Rahman\n\nHealthcare is one of the highest-impact
  domains for AI\, yet reliable deployment at scale remains difficult. To t
 ruly improve patient care and clinical workflows\, AI must operate under r
 eal clinical constraints\, not just in ideal lab settings. In practice\, d
 eployment is limited by high compute and memory costs\, scarce labeled dat
 a\, and distribution shifts across sites and time. Many clinically importa
 nt findings are also rare and long-tailed\, which makes generalization esp
 ecially challenging. My research makes deployability a design objective by
  developing methods that stay accurate under strict resource and data cons
 traints. In this talk\, I will first discuss high-performance lightweight 
 deep learning architectures built by redesigning core building blocks. I w
 ill then present training-time generative supervision strategies that impr
 ove data efficiency and generalization to rare and long-tailed cases with 
 no inference overhead. I will conclude with a forward-looking direction to
 ward real-time perception for surgical assistance\, where reliable perform
 ance under strict constraints is non-negotiable.\n\nhttps://datascience.ut
 ah.edu/talks/2026-04-01-md-mostafijur-rahman/
URL:https://datascience.utah.edu/talks/2026-04-01-md-mostafijur-rahman/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780
CATEGORIES:health & medicine,machine learning
END:VEVENT
BEGIN:VEVENT
UID:2026-04-06-qiang-ji@datascience.utah.edu
DTSTAMP:20260406T163000Z
DTSTART;TZID=America/Denver:20260406T103000
DTEND;TZID=America/Denver:20260406T113000
SUMMARY:Towards Data-Efficient\, Trustworthy\, and Generalizable AI for Vis
 ual Understanding - Qiang Ji
DESCRIPTION:Qiang Ji\n\nArtificial Intelligence (AI) has achieved remarkabl
 e progress and is increasingly integrated across a wide range of fields\, 
 fueling what many describe as the fourth industrial revolution. However\, 
 behind this widespread enthusiasm lie fundamental limitations. Today’s A
 I systems face three major challenges: (1) an insatiable demand for large-
 scale labeled data\, (2) limited trustworthiness due to inadequate uncerta
 inty quantification\, and (3) poor generalization across domains. These ch
 allenges cannot be addressed simply by scaling data and computation\; inst
 ead\, they require foundational advances in theory and methodology.\n\nIn 
 this talk\, I will present recent research from my lab that addresses thes
 e challenges in a variety of computer vision tasks. To improve data effici
 ency and generalization\, I will introduce our work on knowledge-augmented
  deep learning\, where prior knowledge from diverse sources is systematica
 lly identified\, encoded\, and integrated with data-driven neural networks
 . This approach leads to hybrid neural-symbolic models that are both more 
 data-efficient and more generalizable. To enhance model trustworthiness an
 d explainability\, I will discuss our advances in Bayesian deep learning. 
 First\, I will present our work on a Bayesian Transformer framework for ac
 curate and robust human activity recognition. I will then introduce our wo
 rk on uncertainty attribution\, which identifies the sources of uncertaint
 y in deep models and leverages this information for uncertainty mitigation
  and improved model performance. Finally\, I will highlight our recent wor
 k on causal deep learning for addressing domain generalization. I will int
 roduce a neural causal model that learns domain-invariant representations 
 by identifying and eliminating spurious correlations arising from data bia
 ses.\n\nTogether\, these efforts aim to advance a new generation of AI sys
 tems that are more data-efficient\, trustworthy\, and robust\, enabling re
 liable deployment across a range of domains including human behavior under
 standing\, medical imaging\, scientific discovery\, and human–robot inte
 raction.\n\nhttps://datascience.utah.edu/talks/2026-04-06-qiang-ji/
URL:https://datascience.utah.edu/talks/2026-04-06-qiang-ji/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780
CATEGORIES:causal inference,computer vision,deep learning,machine learning
END:VEVENT
BEGIN:VEVENT
UID:2026-04-07-si-chen@datascience.utah.edu
DTSTAMP:20260407T163000Z
DTSTART;TZID=America/Denver:20260407T103000
DTEND;TZID=America/Denver:20260407T113000
SUMMARY:Advancing AI Literacy and Human-Centered AI for Teaching and Learni
 ng - Si Chen
DESCRIPTION:Si Chen\n\nArtificial intelligence (AI) is rapidly transforming
  education and how people learn\, teach\, and prepare for the future. Yet 
 many systems are still built around technical capabilities rather than the
  real needs of students\, educators\, and families. In this talk\, I prese
 nt a human-centered design research agenda that advances both AI literacy 
 and AI for teaching and learning across students\, families\, and educator
 s. First\, I define and conceptualize generative AI literacy across childr
 en and parents by co-designing measurement frameworks and interactive tool
 s. This work enables families to build a shared understanding of AI while 
 supporting its critical and responsible use in everyday self-directed lear
 ning contexts. Second\, I present AI Academy\, an faculty-facing professio
 nal development program and badge system that expands institutional capaci
 ty for AI. Through curriculum design and lightweight tools\, this work sup
 ports faculty in integrating generative AI into teaching while strengtheni
 ng their AI literacy and maintaining pedagogical and disciplinary goals. L
 astly\, I present AI-powered tutoring systems and other AI interactions I 
 have developed with college learners with disabilities\, including LLM-bas
 ed chatbots\, particularly for Deaf and Hard of Hearing students.The talk 
 is intended for a broad\, interdisciplinary audience.\n\nhttps://datascien
 ce.utah.edu/talks/2026-04-07-si-chen/
URL:https://datascience.utah.edu/talks/2026-04-07-si-chen/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780
CATEGORIES:education,human-centered computing,large language models
END:VEVENT
BEGIN:VEVENT
UID:2026-04-08-fahim-faisal@datascience.utah.edu
DTSTAMP:20260408T160000Z
DTSTART;TZID=America/Denver:20260408T100000
DTEND;TZID=America/Denver:20260408T110000
SUMMARY:Multilingual Model Adaptation for Under-Served Languages - Fahim Fa
 isal
DESCRIPTION:Fahim Faisal\n\nLanguage models with multilingual capabilities 
 serve as crucial touchpoints for improving the inclusion of underrepresent
 ed languages in Natural Language Processing (NLP). This research investiga
 tes the structural sources of linguistic underrepresentation and explores 
 strategies for improving the adaptation of low-resource language varieties
  through the development of linguistically grounded resources. We first ex
 amine the extent of multilingual and geographic representation gaps across
  three key dimensions of language modeling: datasets\, model architecture\
 , and model-generated text. Next\, we introduce DialectBench\, an initiati
 ve designed to evaluate language variation in the form of dialects and lan
 guage varieties—an aspect often overlooked in NLP benchmarks\, which pri
 marily focus on standardized language forms. To address the challenges rev
 ealed by these analyses\, we propose two adaptation frameworks. The first 
 introduces phylogenetic adapter hierarchies that exploit language-family s
 tructure to enable zero-shot transfer across related languages. The second
  presents a pivot-based reinforcement learning approach that leverages hig
 h-resource expert models to transfer reasoning alignment without requiring
  target-language annotations. Together\, these linguistically motivated ad
 aptation strategies aim to improve the performance of language models on u
 nderrepresented languages and dialects\, ultimately contributing to more e
 quitable and accessible NLP systems for diverse linguistic communities.\n\
 nhttps://datascience.utah.edu/talks/2026-04-08-fahim-faisal/
URL:https://datascience.utah.edu/talks/2026-04-08-fahim-faisal/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780
CATEGORIES:fairness & ethics,large language models,natural language process
 ing,society & policy
END:VEVENT
BEGIN:VEVENT
UID:2026-04-10-chase-neumann@datascience.utah.edu
DTSTAMP:20260410T193000Z
DTSTART;TZID=America/Denver:20260410T133000
DTEND;TZID=America/Denver:20260410T143000
SUMMARY:Building and Utilizing Foundation Models for Drug Discovery and Cli
 nical Development - Chase Neumann
DESCRIPTION:Chase Neumann (PhD)\n\nThe journey toward decoding biology at s
 cale began with a focus on high-dimensional cellular morphology. At Recurs
 ion\, our differentiation centered on deep learning models designed to lea
 rn biological representations directly from imaging\, enabling predictive 
 inference at a massive scale. By leading multiple cross-functional teams f
 rom early-stage discovery through to early clinical development\, we demon
 strated the power of this "inference-first" philosophy. A primary highligh
 t of these efforts was the RBM39 program\, a novel molecular glue degrader
  discovered entirely through computational inference rather than tradition
 al screening. This success proved that models could identify complex biolo
 gical mechanisms\; however\, moving from cellular discovery to comprehensi
 ve patient care requires a leap into even higher-dimensional\, clinical da
 ta.\n\nValinor represents the next evolution of this mission. We build mul
 timodal clinical foundation models\, co-designing data collection and mode
 l architecture to maximize signal within and across complex modalities. We
  have proprietary access to patient cohorts and collaborate with biobanks\
 , clinical trial sites\, and academic partners\, giving us unique data adv
 antages at scale. Our approach is to first build the best unimodal patient
  representations across modalities—including DNA\, transcriptomics\, pro
 teomics\, cfDNA\, histopathology\, and patient reports—and then fuse the
 m.\n\nWe have shown that attention-based fusion consistently outperforms u
 nimodal approaches while making the contributions of different modalities 
 interpretable. This enables genuine clinical reasoning: the model can chai
 n evidence across modalities\, explain which features drive a prediction\,
  and engage with clinicians in natural language. Ultimately\, we envision 
 a virtual patient that reasons over the full spectrum of a patient's biolo
 gy the way an expert clinician would\, but at a scale and resolution no hu
 man can match.\n\nhttps://datascience.utah.edu/talks/2026-04-10-chase-neum
 ann/
URL:https://datascience.utah.edu/talks/2026-04-10-chase-neumann/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB L112 | https://utah.zoom.us/j/85983626630
CATEGORIES:biology & genomics,health & medicine,large language models,natur
 al language processing
END:VEVENT
BEGIN:VEVENT
UID:2026-04-13-jihyun-rho@datascience.utah.edu
DTSTAMP:20260413T190000Z
DTSTART;TZID=America/Denver:20260413T130000
DTEND;TZID=America/Denver:20260413T140000
SUMMARY:Supporting responsible use of AI in visual-based teaching and learn
 ing - Jihyun Rho
DESCRIPTION:Jihyun Rho\n\nArtificial intelligence (AI) is increasingly inte
 grated into educational contexts\, reshaping how teachers and students gen
 erate and interpret visual representations such as diagrams\, illustration
 s\, and data visualizations. While AI offers efficiency and flexibility\, 
 it also introduces inaccuracies and misleading interpretations that can un
 dermine learning. In this talk\, I present a research program centered on 
 supporting the responsible use of AI in visual-based teaching and learning
  by promoting visual literacy. Grounded in design-based research\, I desig
 n and evaluate AI-augmented learning environment where educators and stude
 nts engage with AI-generated outputs as critical inquirers. This work show
 s that these approaches improved interpretive accuracy\, deepened reasonin
 g about visual representations\, and reduce over-reliance on AI. Together\
 , this work contributes design principles for fostering responsible AI use
  in visual-based educational contexts.\n\nhttps://datascience.utah.edu/tal
 ks/2026-04-13-jihyun-rho/
URL:https://datascience.utah.edu/talks/2026-04-13-jihyun-rho/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780
CATEGORIES:education,fairness & ethics
END:VEVENT
BEGIN:VEVENT
UID:2026-04-14-sameer-honwad@datascience.utah.edu
DTSTAMP:20260414T163000Z
DTSTART;TZID=America/Denver:20260414T103000
DTEND;TZID=America/Denver:20260414T113000
SUMMARY:Move Slow and Build Community - Sameer Honwad
DESCRIPTION:Sameer Honwad\n\nSocio-emotional learning (SEL) and reflection 
 are foundational components of quality education\, yet they remain underre
 presented in many school curricula. Reflection supports scientific thinkin
 g and inquiry\, while SEL equips students with the emotional resilience ne
 eded to take risks\, embrace failure\, and navigate the uncertainties inhe
 rent in learning and innovation. Together\, these competencies are essenti
 al not only for academic and professional growth\, but for everyday wellbe
 ing.Implementing SEL and reflective practices in rural India presents dist
 inct challenges. Limited access to trained professionals and consistent re
 sources is compounded by cultural stigma around openly discussing emotions
 \, barriers that make it difficult for students to develop these critical 
 skills in traditional school settings.\n\nTo address this gap\, we designe
 d a conversational chatbot tailored for students in rural India\, providin
 g a low-barrier\, accessible space for daily reflection and emotional expr
 ession. Importantly\, the chatbot is also designed to serve as a springboa
 rd for teachers to introduce concepts of AI literacy\, enabling meaningful
  classroom conversations about how AI systems work\, their limitations\, a
 nd the ethical considerations surrounding their use. This presentation pre
 sents the co-design process behind the chatbot\, highlighting the collabor
 ative contributions of an interdisciplinary team comprising teachers\, the
 rapists\, computer scientists\, and learning scientists. The presentation 
 discusses how this cross-disciplinary approach shaped a tool intended to h
 elp students engage with their socio-emotional selves and build a habit of
  reflective thinking in their daily lives.\n\nhttps://datascience.utah.edu
 /talks/2026-04-14-sameer-honwad/
URL:https://datascience.utah.edu/talks/2026-04-14-sameer-honwad/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780
CATEGORIES:education,society & policy
END:VEVENT
BEGIN:VEVENT
UID:2026-04-15-shiqi-yu@datascience.utah.edu
DTSTAMP:20260415T160000Z
DTSTART;TZID=America/Denver:20260415T100000
DTEND;TZID=America/Denver:20260415T103000
SUMMARY:AI for Multi-Wavelength X-ray Analysis - Shiqi Yu
DESCRIPTION:Shiqi Yu\n\nAccurate parameter estimation in X-ray astronomy ty
 pically relies on traditional methods\, such as likelihood-based spectral 
 fitting\, which can be computationally prohibitive as model complexity and
  data dimensionality increase. In this talk\, I present a neural network-b
 ased framework designed to bypass iterative fitting by directly mapping sp
 ectral observations to physical parameters. Using the Circinus galaxy as a
  benchmark\, we demonstrate how architectures trained on synthetic data fr
 om theoretical models can recover intrinsic properties\, such as column de
 nsity and torus geometry\, with both high speed and high precision. This a
 pproach maintains the physical rigor required for broadband analysis with 
 multiple telescopes while significantly reducing inference time. I will di
 scuss the challenges and resolutions associated with training and predicti
 ng on multi-instrument data\, as well as the potential for these AI-driven
  methods to enable large-scale systematic studies across various astrophys
 ical sources and other scientific applications.\n\nhttps://datascience.uta
 h.edu/talks/2026-04-15-shiqi-yu/
URL:https://datascience.utah.edu/talks/2026-04-15-shiqi-yu/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:Dinosaur Room (CSC 206) | https://utexas.zoom.us/j/84742203545?pwd
 =lUosaf3T6bkIS1QAaIQYHYiiClE2ZP.1
CATEGORIES:deep learning,machine learning,physics & astronomy
END:VEVENT
BEGIN:VEVENT
UID:2026-04-15-amirali-abdullah@datascience.utah.edu
DTSTAMP:20260415T163000Z
DTSTART;TZID=America/Denver:20260415T103000
DTEND;TZID=America/Denver:20260415T110000
SUMMARY:Controlling LLM's via Activation Geometry - Amirali Abdullah
DESCRIPTION:Amirali Abdullah\n\nControlling the behavior of large language 
 models at inference time is an increasingly important problem. In this tal
 k\, I present a simple and unified approach to steering model behavior bas
 ed on activation geometry. By learning a single classifier over hidden rep
 resentations\, we can derive directions that control multiple attributes s
 uch as helpfulness\, style\, or safety\, and compose them dynamically with
 out retraining.\nThis framework enables flexible\, low cost control of mod
 el outputs and highlights a geometric view of representation space beyond 
 fixed linear directions. I will discuss empirical results showing how this
  approach supports multi attribute control in practice\, and briefly outli
 ne how such steering mechanisms can be useful in scientific settings where
  reliable and interpretable model behavior is critical. Our recent followu
 p work suggests that similar activation level interventions can extend acr
 oss modalities\, enabling systematic analysis and control in text to image
  models through composable operations.\n\nhttps://datascience.utah.edu/tal
 ks/2026-04-15-amirali-abdullah/
URL:https://datascience.utah.edu/talks/2026-04-15-amirali-abdullah/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:Dinosaur Room (CSC 206) | https://utexas.zoom.us/j/84742203545?pwd
 =lUosaf3T6bkIS1QAaIQYHYiiClE2ZP.1
CATEGORIES:large language models,natural language processing
END:VEVENT
BEGIN:VEVENT
UID:2026-04-15-chengbin-deng@datascience.utah.edu
DTSTAMP:20260415T193000Z
DTSTART;TZID=America/Denver:20260415T133000
DTEND;TZID=America/Denver:20260415T143000
SUMMARY:AI-Driven Environmental Intelligence: Scalable and System-Level App
 roaches for Environmental Decision Making - Chengbin Deng
DESCRIPTION:Chengbin Deng\n\nEnvironmental systems are becoming more comple
 x\, dynamic\, and tightly connected to human activities\, yet much of our 
 current work remains focused on isolated models or static mapping. In this
  talk\, a framework will be present and discussed that uses AI to move bey
 ond observation toward a more integrated understanding of environmental sy
 stems. The central idea is to link geospatial data\, models\, and real-wor
 ld decisions so that environmental information can better reflect changing
  conditions across space and time while remaining meaningful for researche
 rs and stakeholders. By drawing on some recent federally supported project
 s\, this talk will show how this framework can capture large scale environ
 mental dynamics\, account for social and local context\, and support more 
 informed decision processes in various settings\, ranging from urban syste
 ms to extreme events. Instead of using AI just as a tool\, this work views
  it as part of a broader system that shapes how environmental problems are
  understood and acted\, with the goal of advancing a more adaptive and pra
 ctical form of environmental intelligence.\n\nhttps://datascience.utah.edu
 /talks/2026-04-15-chengbin-deng/
URL:https://datascience.utah.edu/talks/2026-04-15-chengbin-deng/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780
CATEGORIES:climate & environment,geospatial
END:VEVENT
BEGIN:VEVENT
UID:2026-04-17-daniel-sieta@datascience.utah.edu
DTSTAMP:20260417T193000Z
DTSTART;TZID=America/Denver:20260417T133000
DTEND;TZID=America/Denver:20260417T143000
SUMMARY:Multimodal Data Augmentation for Data-Efficient Robot Manipulation 
 - Daniel Sieta
DESCRIPTION:Daniel Sieta (USC)\n\nDespite recent advances\, learning-based 
 robot manipulation systems often require large demonstration datasets and 
 degrade in cluttered or deformable environments. This talk presents diffus
 ion-based multimodal data augmentation methods that synthesize consistent 
 observations and action labels. By augmenting limited demonstrations\, the
 se approaches substantially reduce data requirements and enable robust man
 ipulation in complex\, real-world settings.\n\nhttps://datascience.utah.ed
 u/talks/2026-04-17-daniel-sieta/
URL:https://datascience.utah.edu/talks/2026-04-17-daniel-sieta/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB L112 | https://utah.zoom.us/j/85983626630
CATEGORIES:computer vision,robotics
END:VEVENT
BEGIN:VEVENT
UID:2026-08-28-warren-pettine@datascience.utah.edu
DTSTAMP:20260828T193000Z
DTSTART;TZID=America/Denver:20260828T133000
DTEND;TZID=America/Denver:20260828T143000
SUMMARY:About Start-UP MTN: which involves Comptutational Neuroscience and 
 AI - Warren Pettine
DESCRIPTION:Warren Pettine (UU Psychiatry & MTN)\n\nhttps://datascience.uta
 h.edu/talks/2026-08-28-warren-pettine/
URL:https://datascience.utah.edu/talks/2026-08-28-warren-pettine/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 2250
CATEGORIES:biology & genomics
END:VEVENT
BEGIN:VEVENT
UID:2026-09-04-george-vega-yon@datascience.utah.edu
DTSTAMP:20260904T193000Z
DTSTART;TZID=America/Denver:20260904T133000
DTEND;TZID=America/Denver:20260904T143000
SUMMARY:Data Science of Tracking Measles in Utah - George Vega Yon
DESCRIPTION:George Vega Yon (UU Epidemiology)\n\nFrom high-performance comp
 uting clusters to the constraints of a local health department\, this talk
  presents a research program in statistical computing and data science\, s
 panning network science and public health\, and goes deep on one example f
 rom each. On the public health side\, agent-based simulation developed wit
 h the Utah Department of Health and Human Services turns vaccination and d
 emographic data into school-level measles risk\, routed through the state'
 s health districts to the school districts that act on it\; related outbre
 ak-reconstruction work along the Utah-Nevada border estimated infections m
 issed by surveillance\, producing results consistent with independent geno
 mic analyses. On the network science side\, Exponential-family Random Grap
 h Models (ERGMs)—statistical models used to characterize the structure o
 f observed networks—I present a novel application of a pooled bipartite 
 model of patient-provider networks in healthcare\, highlighting methodolog
 ical challenges in data heterogeneity and model convergence. Underneath bo
 th is scientific software built for other researchers to use: epiworld\, e
 rgmito\, netdiffuseR\, rgexf\, slurmR\, and others\, under a single commit
 ment—the same model should run on a high-performance cluster and inside 
 a health department with far less to work with. Making powerful methods ex
 ecutable in constrained environments is treated here not as an implementat
 ion detail\, but as a methodological contribution in its own right.\n\nhtt
 ps://datascience.utah.edu/talks/2026-09-04-george-vega-yon/
URL:https://datascience.utah.edu/talks/2026-09-04-george-vega-yon/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 2250
CATEGORIES:health & medicine,networks & graphs,society & policy,statistics
END:VEVENT
BEGIN:VEVENT
UID:2026-09-11-benjie-wang@datascience.utah.edu
DTSTAMP:20260911T193000Z
DTSTART;TZID=America/Denver:20260911T133000
DTEND;TZID=America/Denver:20260911T143000
SUMMARY:The Formalization Gap in Generative AI - Benjie Wang
DESCRIPTION:Benjie Wang (UU Computing)\n\nGenerative models\, such as large
  language models and diffusion models\, have tremendously increased the sc
 ope of problems that AI can address. As such\, there is a significant tren
 d toward incorporating generative AI to automate tasks across computing an
 d more broadly\, from controlling robotics systems\, to software generatio
 n and testing\, to searching over scientific knowledge. However\, there re
 mains a significant **formalization gap** between the domain knowledge\, t
 heories\, and logical and semantic constraints that are vital to applicati
 ons\, and the statistical patterns over natural data represented by large 
 generative models. In this talk\, I will demonstrate how we can systematic
 ally bridge this formalization gap towards more reliable and efficient AI.
  First\, drawing from examples and applications in my research\, I will sh
 ow how to design practical inference-time strategies using probabilistic r
 easoning that bridge between formal language and generative models at scal
 e. Then\, I will discuss how these practical methods are underpinned by ad
 vances in the mathematical and computational foundations of efficiently re
 presenting joint probability distributions.\n\nhttps://datascience.utah.ed
 u/talks/2026-09-11-benjie-wang/
URL:https://datascience.utah.edu/talks/2026-09-11-benjie-wang/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 2250 | https://utah.zoom.us/j/85983626630
CATEGORIES:deep learning,large language models,natural language processing,
 robotics
END:VEVENT
BEGIN:VEVENT
UID:2026-09-18-polina-kukhareva@datascience.utah.edu
DTSTAMP:20260918T193000Z
DTSTART;TZID=America/Denver:20260918T133000
DTEND;TZID=America/Denver:20260918T143000
SUMMARY:Using AI and Real-World Data to Personalize Diabetes Treatment - Po
 lina Kukhareva
DESCRIPTION:Polina Kukhareva (UU BMI)\n\nPatients with type 2 diabetes have
  increasingly many treatment options\, but responses to these therapies va
 ry substantially across individuals and across outcomes such as glycemic c
 ontrol\, weight\, cardiovascular events\, kidney outcomes\, and hypoglycem
 ia. In this talk\, I will describe our work using large-scale electronic h
 ealth record data\, causal inference\, and machine learning to characteriz
 e treatment heterogeneity and develop approaches for individualized treatm
 ent selection. I will discuss challenges in reconstructing longitudinal tr
 eatment regimens from real-world prescribing data\, predicting multiple pa
 tient outcomes\, and identifying the level of patient-specific resolution 
 that can be reliably supported by available data. I will also discuss how 
 these methods can ultimately support patient-centered treatment comparison
 s and clinical decision making.\n\nhttps://datascience.utah.edu/talks/2026
 -09-18-polina-kukhareva/
URL:https://datascience.utah.edu/talks/2026-09-18-polina-kukhareva/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 2250 | https://utah.zoom.us/j/85983626630
CATEGORIES:causal inference,health & medicine,machine learning
END:VEVENT
BEGIN:VEVENT
UID:2026-09-25-matthew-s-sigman@datascience.utah.edu
DTSTAMP:20260925T193000Z
DTSTART;TZID=America/Denver:20260925T133000
DTEND;TZID=America/Denver:20260925T143000
SUMMARY:Data Science Meets Chemistry - Matthew S. Sigman
DESCRIPTION:Matthew S. Sigman (UU Chemistry)\n\nHow do chemists use data sc
 ience to discover and understand new reactions? This talk offers an access
 ible introduction to the distinctive challenges of chemical data: small\, 
 complex datasets\; molecular representations\; and the need to connect pre
 dictions with physical mechanisms. Through examples from reaction developm
 ent\, catalyst design\, and optimization\, it will show how molecular desc
 riptors\, statistical modeling\, and machine learning reveal relationships
  between structure\, reactivity\, and selectivity. Along the way\, the tal
 k will translate chemical questions for a data science audience\, highligh
 t opportunities for automation and collaboration\, and consider how new co
 mputational approaches might expand what chemists can discover.\n\nhttps:/
 /datascience.utah.edu/talks/2026-09-25-matthew-s-sigman/
URL:https://datascience.utah.edu/talks/2026-09-25-matthew-s-sigman/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 2250 | https://utah.zoom.us/j/85983626630
CATEGORIES:education,optimization
END:VEVENT
BEGIN:VEVENT
UID:2026-09-30-bei-wang-shandian-zhe@datascience.utah.edu
DTSTAMP:20260930T180000Z
DTSTART;TZID=America/Denver:20260930T120000
DTEND;TZID=America/Denver:20260930T130000
SUMMARY:Structure Without a Grid: Extracting and Tracking Topological Featu
 res - Bei Wang\, Shandian Zhe
DESCRIPTION:Bei Wang (&\, Shandian Zhe)\, Shandian Zhe (&\, Shandian Zhe)\n
 \nLarge-scale scientific simulations — in fluid dynamics\, climate\,\nco
 mbustion\, and beyond — are increasingly outpacing our ability to\nstore
 \, transfer\, and analyze them at full resolution. A growing\nalternative 
 to storing raw discretized grids is the continuous\nimplicit model\, such 
 as multivariate functional approximation (MFA)\nand implicit neural repres
 entations (INRs). These represent a\nsimulation's fields as continuous fun
 ctions that can be queried at\narbitrary points and differentiated exactly
 \, with no resampling\nrequired. This talk presents a line of work buildin
 g the topological\ndata analysis toolkit needed to make such continuous re
 presentations\npractically useful. First\, we show how to extract critical
  points —\nthe maxima\, minima\, and saddles that organize a scalar fiel
 d's\nstructure — directly from an MFA model\, without ever discretizing 
 back\nto a grid. Second\, we extend this to richer topological descriptors
 :\ncontours (isosurfaces)\, Jacobi sets (the shared critical structure\nbe
 tween two co-registered fields\, useful for comparing related\nquantities 
 like pressure and temperature)\, and ridge-valley graphs\,\nwhich trace fi
 lamentary or ridge-like structure in the data. Third\, we\nintroduce a fra
 mework for tracking these topological features\ncontinuously through time 
 or parameter space\, following critical\npoints as smooth trajectories usi
 ng the model's own derivatives rather\nthan resampling frame by frame — 
 avoiding the aliasing artifacts that\nplague discrete feature tracking. To
 gether\, these methods point toward\na workflow in which large-scale data 
 is stored once as a continuous\nfunctional surrogate\, and structural anal
 ysis — feature detection\,\ncomparison across fields\, and evolution ove
 r time — happens directly\non that surrogate\, at whatever resolution th
 e analysis demands. This\nwork is joint with Guanqun Ma\, David Lenz\, Tom
  Peterka\, Hanqi Guo\,\nKaiyuan Tang\, and Chaoli Wang.\n\nSpeaker 2: Shan
 dian Zhe (UU KSoC)\n\nhttps://datascience.utah.edu/talks/2026-09-30-bei-wa
 ng-shandian-zhe/
URL:https://datascience.utah.edu/talks/2026-09-30-bei-wang-shandian-zhe/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 3780 (Evans) | https://utexas.zoom.us/j/88630218251?pwd=OLbanw
 8MQXVnsiho5qaGjdPKxe98Ea.1&jst=1
CATEGORIES:visualization
END:VEVENT
BEGIN:VEVENT
UID:2026-10-02-berton-earnshaw@datascience.utah.edu
DTSTAMP:20261002T193000Z
DTSTART;TZID=America/Denver:20261002T133000
DTEND;TZID=America/Denver:20261002T143000
SUMMARY:Analyzing -omics Data with High-Dimensional Embeddings: Advantages 
 and Pitfalls - Berton Earnshaw
DESCRIPTION:Berton Earnshaw (UU Math+Bio)\n\nHigh-dimensional embeddings of
 fer a powerful way to summarize complex omics measurements\, such as gene-
 expression or protein-abundance profiles\, and to compare biological condi
 tions in a common geometric space. More dimensions can preserve weak or di
 stributed biological signals that would be missed by a small set of featur
 es. But additional dimensions also introduce noise\, inflate distances\, a
 nd can make apparently large differences difficult to interpret.\n\nIn thi
 s talk\, I will use a simple geometric test based on the squared length of
  an embedded data vector to illustrate both sides of this tradeoff. We wil
 l examine how statistical power depends on dimension\, sample size\, mean 
 shifts\, variance changes\, and correlations among features. The analysis 
 reveals several useful principles: weak signals can accumulate across many
  coordinates\, irrelevant features can reduce sensitivity\, and changes in
  covariance can alter power even when the average distance remains unchang
 ed. Along the way\, I will introduce accessible ideas from hypothesis test
 ing\, Gaussian models\, and high-dimensional geometry. The goal is to prov
 ide practical intuition for deciding when high-dimensional embeddings help
  reveal biological perturbations—and when they may instead obscure them.
 \n\nhttps://datascience.utah.edu/talks/2026-10-02-berton-earnshaw/
URL:https://datascience.utah.edu/talks/2026-10-02-berton-earnshaw/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 2250
CATEGORIES:biology & genomics,deep learning,statistics
END:VEVENT
BEGIN:VEVENT
UID:2026-10-09-david-gleich@datascience.utah.edu
DTSTAMP:20261009T180000Z
DTSTART;TZID=America/Denver:20261009T120000
DTEND;TZID=America/Denver:20261009T130000
SUMMARY:Powers of magnetic graph matrix: Fourier spectrum\, walk compressio
 n\, and applications - David Gleich
DESCRIPTION:David Gleich (Purdue University)\n\nMagnetic graph matrices are
  powerful tools for modeling quantum systems and directed networks\, but t
 heir application in network analysis has been limited by a lack of combina
 torial understanding. We present a combinatorial interpretation that funda
 mentally reveals how these matrices encode local network structure. We fur
 ther show that this structure information is highly compressible in real-w
 orld networks\, enabling accurate approximations from a small number of ma
 gnetic potentials. This fresh foundation also unlocks further applications
 \, like identifying crucial network motifs (e.g.\, frustrated directed cyc
 les such as feed-forward loops) and enhancing directed link prediction.\n\
 nLink to paper:\nhttps://www.pnas.org/doi/10.1073/pnas.2516664123\n\nhttps
 ://datascience.utah.edu/talks/2026-10-09-david-gleich/
URL:https://datascience.utah.edu/talks/2026-10-09-david-gleich/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 2760 | https://utah.zoom.us/j/85983626630
CATEGORIES:algorithms & theory,networks & graphs
END:VEVENT
BEGIN:VEVENT
UID:2026-10-23-kenneth-collins@datascience.utah.edu
DTSTAMP:20261023T193000Z
DTSTART;TZID=America/Denver:20261023T133000
DTEND;TZID=America/Denver:20261023T143000
SUMMARY:The talk will consider the impact of AI in moving image works\, inc
 luding AI filmmaking\, animation\, and video art - Kenneth Collins
DESCRIPTION:Kenneth Collins (UU Film)\n\nhttps://datascience.utah.edu/talks
 /2026-10-23-kenneth-collins/
URL:https://datascience.utah.edu/talks/2026-10-23-kenneth-collins/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 2250 | https://utah.zoom.us/j/85983626630
CATEGORIES:computer vision
END:VEVENT
BEGIN:VEVENT
UID:2026-10-30-guang-tian@datascience.utah.edu
DTSTAMP:20261030T193000Z
DTSTART;TZID=America/Denver:20261030T133000
DTEND;TZID=America/Denver:20261030T143000
SUMMARY:Land Use and Transportation Planning with AI - Guang Tian
DESCRIPTION:Guang Tian (UU City & Metro Planning + SCI)\n\nhttps://datascie
 nce.utah.edu/talks/2026-10-30-guang-tian/
URL:https://datascience.utah.edu/talks/2026-10-30-guang-tian/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 2250 | https://utah.zoom.us/j/85983626630
CATEGORIES:geospatial
END:VEVENT
BEGIN:VEVENT
UID:2026-11-13-aik-choon-tan@datascience.utah.edu
DTSTAMP:20261113T203000Z
DTSTART;TZID=America/Denver:20261113T133000
DTEND;TZID=America/Denver:20261113T143000
SUMMARY:AI for Cancer Treatment - Aik Choon Tan
DESCRIPTION:Aik Choon Tan (UU Oncological Sciences)\n\nhttps://datascience.
 utah.edu/talks/2026-11-13-aik-choon-tan/
URL:https://datascience.utah.edu/talks/2026-11-13-aik-choon-tan/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 2250 | https://utah.zoom.us/j/85983626630
CATEGORIES:health & medicine
END:VEVENT
BEGIN:VEVENT
UID:2026-12-04-tom-henderson@datascience.utah.edu
DTSTAMP:20261204T203000Z
DTSTART;TZID=America/Denver:20261204T133000
DTEND;TZID=America/Denver:20261204T143000
SUMMARY:40 Years of AI ! - Tom Henderson
DESCRIPTION:Tom Henderson (UU KSoC)\n\nhttps://datascience.utah.edu/talks/2
 026-12-04-tom-henderson/
URL:https://datascience.utah.edu/talks/2026-12-04-tom-henderson/
SEQUENCE:0
STATUS:CONFIRMED
TRANSP:OPAQUE
LOCATION:WEB 2250 | https://utah.zoom.us/j/85983626630
END:VEVENT
END:VCALENDAR
