Skip to content
Home

Publication

Research

One peer-reviewed publication: an undergraduate thesis, presented at an international conference, awarded, and indexed.

Peer-reviewedScopus-indexed · IEEE XploreBest Paper · 1stBest Presenter · 8th

Comparative Evaluation of Graph Neural Network Algorithms for Music Recommendation Systems

Farhan Rangkuti, Fitriyani

2024 International Conference on Intelligent Cybernetics Technology & Applications (ICICyTA), IEEE, pp. 297–301

Most music recommenders lean on collaborative filtering, which sees only that a user interacted with an item and nothing about how the catalogue connects. A graph keeps those connections. This paper asks whether that actually helps, and which graph architecture to reach for.

IThe problem

Collaborative filtering treats recommendation as a user–item matrix. It works, but it discards structure: the fact that two playlists overlap on four tracks, or that a track sits between two otherwise unconnected clusters, is information a matrix cannot hold.

Graph neural networks are built for exactly that shape, and had been applied to e-commerce and social networks — but very little work existed for music specifically. The gap was not "does GNN work", it was "which architecture, on this kind of data, with what trade-offs".

IIHow it was built

Data
The Spotify Million Playlist Dataset — one million playlists, nearly two million distinct tracks, three hundred thousand artists, built by users between January 2010 and October 2017. Experiments ran on a 10,000-playlist subset.
Features
Each track node was augmented with its Spotify API audio features, used as the initial node embedding rather than starting from random vectors.
Graph
A bipartite graph: playlist nodes and track nodes, edges only between the two types and never within one. Membership is the edge.
Recommendation
The trained model produces an embedding per node. A track is scored for a playlist by the dot product of their embeddings, and the highest scores that are not already in the playlist become the recommendations.
Training
Two convolutional layers, mean aggregation, BCEWithLogitsLoss, AdamW, 500 epochs. Three hyperparameter conditions — 36, 92 and 160 hidden channels — across two architectures, giving six trained models.

IIIResults

MetricGraphSAGEGCN
AUC0.91160.8359
Precision0.89390.7221
Recall0.72160.8270
F1-score0.79860.7710
Training loss0.17550.3668
Validation loss0.54090.6577
Training time1,714.97 s1,691.54 s
Table 1Experiment 1 — 10,000 playlists, 500 epochs, 36 hidden channels. The better value in each row is marked.

GraphSAGE wins on precision, GCN on recall

GraphSAGE reached an AUC of 0.91 against GCN's 0.84, and a precision of 0.89 against 0.72. GCN took recall, 0.83 to 0.72, and trained marginally faster. That split is the useful result: GraphSAGE is the better choice when a wrong recommendation is expensive, GCN when missing a relevant track is. The ordering held across all three conditions.

GCN overfits where GraphSAGE generalises

The loss curves separate the two more clearly than the headline metrics do. GraphSAGE's training loss falls to about 0.2 while its validation loss stays flat in the 0.4–0.5 band. GCN's training loss also falls, but its validation loss climbs consistently past a point — the standard signature of overfitting, and the reason its aggregate scores are worse despite a comparable training curve.

More parameters did not mean better results

Widening the hidden layers from 36 to 92 to 160 channels did not produce a significant improvement in AUC, precision, recall or F1. Some of the larger conditions returned higher loss and clearer overfitting, and all of them cost more training time. The first and smallest condition was the best of the three — a scalability finding as much as an accuracy one.

GCN drifts toward popular tracks

The qualitative output is where the two models differ most. Asked for ten tracks for the same playlist, GraphSAGE returned a spread of specific artists — Lecrae's religious hip-hop, Enrique Iglesias's Latin pop, Christina Aguilera's pop and R&B. GCN returned Kanye West, Justin Bieber, Fifth Harmony: broadly popular, broadly safe. The paper reads this as a preference for high-popularity items, which is a familiar failure mode in recommenders and one the aggregate metrics do not reveal.

IVWhat it doesn’t establish

Limitation

What the comparison establishes, and what it does not

Two architectures, one dataset, one graph construction, a 10,000-playlist subset. That is enough to say which performed better here; it is not enough to say which is better for music recommendation in general. Graph methods are sensitive to how the graph is built, and a different edge definition could reorder the result.

Limitation

The practical barriers are stated in the paper

GNNs need a large amount of data before recommendations are any good, they need meaningful compute, and integrating one into an existing recommendation pipeline is substantial work. Those three constraints are why the result is a finding rather than a deployment.

The paper's own next steps: a broader hyperparameter search, more of the dataset than the 10,000-playlist subset, and architectures beyond the two compared here.