Graph Self-Supervised Learning: Taxonomy, Frontiers, and Applications

image

In recent years, deep learning on graph-structured data has drawn much attention in both academic and industrial communities. Following the prevailing (semi-) supervised learning paradigms, most deep graph learning methods suffer from several shortcomings, including heavy label reliance, poor generalization, and weak robustness. To circumvent these issues, graph self-supervised learning (GSSL), which extracts supervision signals for model training with well-designed pretext tasks instead of manual labels, has become a promising and trending learning paradigm for graph data. As the field rapidly grows, a global perspective of the development of GSSL is urgently needed in the research community. To fill the gap, we provide a comprehensive tutorial on this fast-growing yet challenging topic.

This tutorial starts with the foundational background of deep graph learning. Then, we conduct a systematic taxonomy to categorize the existing GSSL methods and introduce the most representative ones. Following the latest research trends, we discuss three frontier subtopics under the umbrella of GSSL, including trustworthy GSSL, efficient GSSL, and automatic GSSL. Afterward, we present the real-world applications of GSSL in various directions, including recommender systems, anomaly/out-of-distribution detection, chemistry, and graph structure learning. Lastly, we finalize the tutorial with conclusions and discuss potential future directions.

We believe this tutorial is beneficial to a broad audience from academia and industry, including general machine learning researchers who would like to know about self-supervised learning on graph-structured data, graph analytics researchers who want to keep track of the most recent advances in deep graph learning, and domain experts who would like to generalize GSSL to new applications or other fields.