Can Unsupervised Learning Methods Reveal Distinct Driving Styles Among Formula 1 Drivers?

Taylor, Antony and Secco, Emanuele Lindo (2026) Can Unsupervised Learning Methods Reveal Distinct Driving Styles Among Formula 1 Drivers? International Journal of Modelling and Optimization. ISSN 2010-3697 (Accepted for Publication)

[thumbnail of IJMO_Antony_v3.pdf] Text
IJMO_Antony_v3.pdf - Accepted Version
Restricted to Repository staff only until 30 June 2027.

Download (837kB)

Abstract

This paper investigates whether unsupervised machine learning techniques can identify meaningful and interpretable driving style clusters among Formula 1 (F1) drivers, using telemetry and lap-time data extracted via the FastF1 Python library applied to the data of the 2023 F1 Season. Four clustering algorithms are implemented and compared: K-Means, Hierarchical Agglomerative Clustering, Density-Based Spatial Clustering of Applications with Noise (DBSCAN), and Gaussian Mixture Models (GMM). A suite of performance features is engineered from raw race telemetry, including relative lap time, lap-time consistency (measured via the Median Absolute Deviation), tyre degradation slope, overtaking efficiency, and intra-team relative sector differentials. Car performance confounders are partially mitigated by computing intra-team relative metrics throughout the entire feature set. A Variance Inflation Factor (VIF) analysis is applied to identify and remove highly collinear features prior to clustering. Principal Component Analysis (PCA) is employed for dimensionality reduction and visualization.
Results suggest that between 3 and 5 statistically meaningful driver clusters emerge, broadly characterized as elite performers, aggressive racers, precision drivers, and solid midfield operators. K-Means with k = 4 and Ward-linkage Hierarchical Clustering yield the most internally consistent groupings, outperforming DBSCAN and GMM on silhouette and Davies-Bouldin scores. A supplementary supervised validation experiment confirms that the engineered feature set captures genuinely driver-discriminative signal, with a Random Forest classifier achieving 78.3% accuracy in predicting driver identity under 5-fold cross-validation. Unsupervised learning can surface interpretable driving style profiles from Formula 1 data, though the explanatory power of these clusters is constrained by data limitations including the small per season sample size, incomplete telemetry coverage and the dominance of car performance over driver skill in raw lap-time metrics.

Item Type: Article
Faculty / Department: Faculty of Human and Digital Sciences > School of Computer Science and the Environment
Depositing User: Emanuele Secco
Date Deposited: 20 Jul 2026 13:30
Last Modified: 20 Jul 2026 13:30
URI: https://hira.hope.ac.uk/id/eprint/4949

Actions (login required)

View Item View Item