Comparative Evaluation of Graph Neural Network Algorithms for Music Recommendation Systems
Farhan Rangkuti, Fitriyani
2024 International Conference on Intelligent Cybernetics Technology & Applications (ICICyTA), IEEE, pp. 297–301
Most music recommenders lean on collaborative filtering, which sees only that a user interacted with an item and nothing about how the catalogue connects. A graph keeps those connections. This paper asks whether that actually helps, and which graph architecture to reach for.
IThe problem
Collaborative filtering treats recommendation as a user–item matrix. It works, but it discards structure: the fact that two playlists overlap on four tracks, or that a track sits between two otherwise unconnected clusters, is information a matrix cannot hold.
Graph neural networks are built for exactly that shape, and had been applied to e-commerce and social networks — but very little work existed for music specifically. The gap was not "does GNN work", it was "which architecture, on this kind of data, with what trade-offs".
IIHow it was built
- Data
- The Spotify Million Playlist Dataset — one million playlists, nearly two million distinct tracks, three hundred thousand artists, built by users between January 2010 and October 2017. Experiments ran on a 10,000-playlist subset.
- Features
- Each track node was augmented with its Spotify API audio features, used as the initial node embedding rather than starting from random vectors.
- Graph
- A bipartite graph: playlist nodes and track nodes, edges only between the two types and never within one. Membership is the edge.
- Recommendation
- The trained model produces an embedding per node. A track is scored for a playlist by the dot product of their embeddings, and the highest scores that are not already in the playlist become the recommendations.
- Training
- Two convolutional layers, mean aggregation, BCEWithLogitsLoss, AdamW, 500 epochs. Three hyperparameter conditions — 36, 92 and 160 hidden channels — across two architectures, giving six trained models.
IIIResults
| Metric | GraphSAGE | GCN |
|---|---|---|
| AUC | 0.9116 | 0.8359 |
| Precision | 0.8939 | 0.7221 |
| Recall | 0.7216 | 0.8270 |
| F1-score | 0.7986 | 0.7710 |
| Training loss | 0.1755 | 0.3668 |
| Validation loss | 0.5409 | 0.6577 |
| Training time | 1,714.97 s | 1,691.54 s |
GraphSAGE wins on precision, GCN on recall
GraphSAGE reached an AUC of 0.91 against GCN's 0.84, and a precision of 0.89 against 0.72. GCN took recall, 0.83 to 0.72, and trained marginally faster. That split is the useful result: GraphSAGE is the better choice when a wrong recommendation is expensive, GCN when missing a relevant track is. The ordering held across all three conditions.
GCN overfits where GraphSAGE generalises
The loss curves separate the two more clearly than the headline metrics do. GraphSAGE's training loss falls to about 0.2 while its validation loss stays flat in the 0.4–0.5 band. GCN's training loss also falls, but its validation loss climbs consistently past a point — the standard signature of overfitting, and the reason its aggregate scores are worse despite a comparable training curve.
More parameters did not mean better results
Widening the hidden layers from 36 to 92 to 160 channels did not produce a significant improvement in AUC, precision, recall or F1. Some of the larger conditions returned higher loss and clearer overfitting, and all of them cost more training time. The first and smallest condition was the best of the three — a scalability finding as much as an accuracy one.
GCN drifts toward popular tracks
The qualitative output is where the two models differ most. Asked for ten tracks for the same playlist, GraphSAGE returned a spread of specific artists — Lecrae's religious hip-hop, Enrique Iglesias's Latin pop, Christina Aguilera's pop and R&B. GCN returned Kanye West, Justin Bieber, Fifth Harmony: broadly popular, broadly safe. The paper reads this as a preference for high-popularity items, which is a familiar failure mode in recommenders and one the aggregate metrics do not reveal.
IVWhat it doesn’t establish
Limitation
What the comparison establishes, and what it does not
Two architectures, one dataset, one graph construction, a 10,000-playlist subset. That is enough to say which performed better here; it is not enough to say which is better for music recommendation in general. Graph methods are sensitive to how the graph is built, and a different edge definition could reorder the result.
Limitation
The practical barriers are stated in the paper
GNNs need a large amount of data before recommendations are any good, they need meaningful compute, and integrating one into an existing recommendation pipeline is substantial work. Those three constraints are why the result is a finding rather than a deployment.
The paper's own next steps: a broader hyperparameter search, more of the dataset than the 10,000-playlist subset, and architectures beyond the two compared here.