Skip to main content
Have a personal or library account? Click to login
Predicting atmospheric particle formation days by Bayesian classification of the time series features Cover

Predicting atmospheric particle formation days by Bayesian classification of the time series features

Open Access
|Jan 2018

Figures & Tables

Fig. 1.

Examples of an event (a) and a non-event (b) day at Hyytiälä, Finland, in May 2005. The x-axis shows one 24-h time period whereas the y-axis shows the range of particle size diameters (from 3 to 1000 nm). The color scale indicates particle concentration (cm– 3). In (a) one can clearly see aerosol particles forming around noon and then growing into larger sizes. This data was accessed via Smart-SMEAR (Junninen et al., 2009).

Fig. 2.

Four different types of NPF days, classified based on the method proposed by Dal Maso et al. (2005). The x-axis shows the 24-h time period, whereas y-axis represents the range of particle diameters (from 3 to 1000 nm). The color indicates the particle concentration (cm-3).

Fig. 3.

Schematic diagram of the ML methodology for classifying aerosol particle formation days.

Fig. 4.

From the concentration data, two types of features are calculated to be used in the learning and validating of the neural network. At each instance of time, the ambient aerosol particle distribution can be presented as a multi-modal log normal distribution, characterized by three parameters per mode (see panel (a)). The set of these fitted parameters are used as the first type of features given to the neural network. The ambient particle distribution evolves throughout the day and this change is manifested in the parameters of the log normal distributions. A set of time-domain quantities calculated over the entire measurement day (excluding nighttime) are given to the neural network as second type of features (see panel (b) and Table 1).

Table 1.

Time-domain feature representations used in this study. The notation of x(i) and N denote the signal x(i) at time i and the number of data points, respectively. In this case, the signal is equivalent with the concentration level of every particle size distribution.

Time-domain featuresFormulaMean (x¯)1Ni=1N(x(i))Variance1Ni=1N(x(i)x¯)2Standard Deviation1Ni=1N(x(i)x¯)2RMS value (RMS)1Ni=1N|x(i)|2Peak value (PV)12(max(x(t))min(x(t)))Kurtosis1Ni=1N(x(i)x¯)4(1Ni=1N(x(i)x¯)2)2Crest factorPVRMSSkewness1Ni=1N(x(i)x¯)3(1Ni=1N(x(i)x¯)2)3Clearance factorPV1N(i=1N|x(i)|)2Impulse factorPV1Ni=1N|x(i)|Shape factorRMS1Ni=1N|x(i)|K-FactorPV·RMS
Fig. 5.

The cumulative sum of eigenvalues of the data in Principal Component (PC) space. In order to retain 99 per cent of original data information, we need to select only the first 69 PCs from almost 400 calculated features.

Fig. 6.

Schematic representation of a BNN with one hidden layer (a) and the used activation functions (b).

Fig. 7.

The bar chart of the successful and unsuccessful number of predicted days.

Table 2.

Training performance (1996–2010).

Visualization methodEvent-daysNon-event daysBNNEvent-days1223 (42.0 %)29 (1.0 %)97.7 %Non-event days32 (1.1 %)1630 (55.9 %)98.1 %97.5 %98.3 %97.9 %
Table 3.

Validation performance (2011–2014).

Visualization methodEvent-daysNon-event daysBNNEvent-days245 (30.3 %)63 (7.8 %)79.5 %Non-event days65 (8.0 %)435 (53.8 %)87.0 %79.0 %87.3 %84.2 %
Language: English
Page range: 1530031 - 1530031
Submitted on: Aug 30, 2017
Accepted on: Sep 11, 2018
Published on: Jan 1, 2018
Published by: Stockholm University Press
In partnership with: Paradigm Publishing Services

© 2018 M. A. Zaidan, V. Haapasilta, R. Relan, H. Junninen, P. P. Aalto, M. Kulmala, L. Laurson, A. S. Foster, published by Stockholm University Press
This work is licensed under the Creative Commons Attribution 4.0 License.