My research lies at the intersection of artificial intelligence, data science, high-performance computing, and scientific discovery. A central focus of my current work is AI for Science and Engineering—developing machine learning methods that accelerate scientific discovery by integrating data-driven learning with scientific knowledge. These approaches are motivated by the unique challenges of scientific and engineering problems, where data are often sparse, predictions must generalize beyond the range of observed data, and models must remain consistent with established scientific principles.
Over the course of my career, my work has advanced scalable algorithms for large-scale computing, data mining methods for analyzing massive datasets, and AI approaches for understanding complex environmental systems. These efforts have helped build bridges between computer science and scientific disciplines such as climate science, hydrology, ecology, and agriculture.
Motivation for AI for Science
Most modern AI systems have been developed for applications such as computer vision and natural language processing, where large labeled datasets are available and predictive accuracy is the primary objective. Scientific and engineering problems present very different challenges. Data are often sparse and noisy, models must generalize beyond the range of observed data, and predictions must remain consistent with established scientific knowledge such as physical laws and conservation principles. These requirements motivate the development of new AI methods that combine data-driven learning with scientific reasoning.
AI for Earth System Science
Beginning in the early 2000s, my group applied large-scale data mining methods—precursors to what is now broadly recognized as machine learning and data science—to planetary-scale Earth system data. Using massive collections of satellite observations and environmental datasets, we developed methods to detect large-scale environmental change such as global land-cover and forest dynamics and to uncover previously unknown relationships within the Earth system.
This work produced some of the first automated systems for monitoring ecosystem disturbances such as forest fires, droughts, floods, and land-cover change. These capabilities formed the basis of the ALERTS ecosystem monitoring system, developed with NASA collaborators and demonstrated at COP16 as a scalable approach for global forest monitoring.
As Principal Investigator of the NSF Expeditions in Computing project Understanding Climate Change: A Data-Driven Approach, I led a multi-institution effort to integrate machine learning with traditional Earth system modeling. The project demonstrated that large-scale integration of satellite observations, climate model outputs, and environmental datasets can reveal climate relationships that are difficult to identify using traditional approaches. This work produced widely used global datasets, including long-term records of lake and reservoir dynamics and global catalogs of ocean eddies that support research on climate variability, extreme events, and water security.
These experiences highlighted both the promise and the limitations of conventional machine learning methods for modeling complex environmental systems and motivated the development of new approaches that integrate scientific knowledge with data-driven learning.
Knowledge-Guided Machine Learning
To address these challenges, my group developed Knowledge-Guided Machine Learning (KGML), a paradigm that integrates scientific knowledge directly into machine learning models. KGML incorporates physical laws, domain constraints, and scientific structure into machine learning architectures and training procedures, enabling models that are both data-driven and scientifically consistent.
By combining machine learning with domain knowledge, KGML improves generalization, robustness, and interpretability when applied to complex scientific systems. These approaches enable physically consistent and uncertainty-aware modeling in settings where purely data-driven models often fail and traditional process-based models are insufficient.
Applications in Environmental Science and Sustainability
Our current research applies KGML and related AI methods to several important environmental and sustainability challenges. In aquatic science, we collaborate with the U.S. Geological Survey to develop models that improve forecasting of freshwater quality and ecosystem dynamics across large networks of lakes and rivers. In hydrology, our work integrates machine learning with physical hydrological models to improve predictions of water availability, watershed dynamics, and hydrological extremes. In agriculture and climate-smart systems, our research supports emerging measurement, monitoring, reporting, and verification (MMRV) frameworks for agricultural greenhouse-gas emissions. These approaches enable credible carbon markets and support the development of climate-smart agricultural practices.
Leadership in AI for Science
Beyond individual research projects, I have worked to advance the broader field of AI for Science through community-building and national initiatives. This includes organizing NSF-supported workshops on AI-enabled scientific revolutions and leading community efforts around Knowledge-Guided Machine Learning and Generative AI for Science. These activities bring together researchers from computer science, environmental science, engineering, and federal agencies to identify new opportunities for accelerating scientific discovery using AI.
Foundations in Scalable Computing and Data Mining
My earlier research focused on the development of scalable algorithms for artificial intelligence and high-performance computing. This work introduced the isoefficiency framework, the first rigorous method for evaluating the scalability of parallel algorithms, which influenced the design of scalable computing systems across a wide range of scientific and engineering applications.
Building on this work, my group developed multilevel graph partitioning algorithms that led to the widely used METIS, ParMETIS, and hMETIS software packages. These tools are embedded in widely used engineering simulation platforms and enable large-scale multiphysics simulations in areas such as structural mechanics, aerospace engineering, and circuit simulation.
Our research also made important contributions to data mining, including graph-based clustering methods such as shared nearest-neighbor clustering and bisecting k-means, as well as widely cited work on anomaly detection. These methods helped establish the foundations for large-scale analysis of high-dimensional datasets and enabled the analysis of massive scientific data.