Research Areas
Our research spans three interconnected domains that define the future of intelligent data systems
AI4DB
AI for Database Systems
Leveraging artificial intelligence technologies to enhance database systems' query performance, autonomous capabilities, and other functionalities. Our research encompasses intelligent query optimization, automated database tuning, and self-managing database systems, dedicated to building next-generation intelligent database systems.
DB4AI
Database Technologies for AI Systems
Utilizing advanced data management technologies to support efficient large model training, low-latency inference, and high-throughput energy efficiency optimization. We focus on building scalable AI infrastructure and optimized data pipelines to provide robust data foundation for artificial intelligence systems.
AI4DS
Novel Data Science Systems with Intelligence-Data Fusion
Employing reasoning large models, multimodal semantic understanding, and intelligent agents to enhance the intelligence level and execution performance of data science systems, effectively unlocking data value. Building future-oriented intelligent data analysis and processing platforms that achieve deep integration of data and intelligence.
Current Projects
AutoDB
Autonomous Database Management
An intelligent database system that automatically optimizes queries, manages resources, and adapts to workload patterns without human intervention.
DataFlow
Efficient LLM Training Pipeline
A scalable data management system designed specifically for large language model training with optimized data loading and processing capabilities.
MultiModal Agent
Intelligent Data Analysis System
An AI agent system that understands and processes multimodal data sources to provide intelligent insights and automated analysis.
DataCentric Framework
Next-Gen AI Development
A comprehensive framework for data-centric AI development that prioritizes data quality and management in AI system design.
Recent Publications
Automatic Database Configuration Debugging using Retrieval-Augmented Language Models
Authors: Sibei Chen, Ju Fan, Bin Wu, Nan Tang, Chao Deng, Pengyi Wang, Ye Li, Jian Tan, Feifei Li, Jingren Zhou, Xiaoyong Du
Conference: Proc. ACM Manag. Data 3(1): 13:1-13:27 (2025) | Status: Published
A novel approach utilizing retrieval-augmented language models to automatically identify and debug database configuration issues, significantly improving system performance and reducing manual debugging efforts.
Weak-to-Strong Prompts with Lightweight-to-Powerful LLMs for High-Accuracy, Low-Cost, and Explainable Data Transformation
Authors: Changlun Li, Chenyu Yang, Yuyu Luo, Ju Fan, Nan Tang
Conference: Proc. VLDB Endow. 18(8): 2371-2384 (2025) | Status: Published
An innovative framework that leverages lightweight LLMs with strategic prompting to achieve high-accuracy data transformation while maintaining low computational costs and providing explainable results.
Andromeda: Debugging Database Performance Issues with Retrieval-Augmented Large Language Models
Authors: Pengyi Wang, Sibei Chen, Ju Fan, Bin Wu, Nan Tang, Jian Tan
Conference: SIGMOD Conference Companion 2025: 243-246 | Status: Published
Andromeda presents a comprehensive system for debugging database performance issues by combining retrieval-augmented techniques with large language models for enhanced diagnostic capabilities.
Research Impact
Publications
High-impact research papers published in top-tier conferences and journals.
Active Projects
Ongoing research projects addressing real-world challenges in Data+AI systems.
Citations
Our research has been recognized and cited by the global research community.