Nomenclature
ϕ(x)
father wavelet function
ψ(x)
mother wavelet function
φ(x)
Haar wavelet function
f
signal
am
average or trend
dm
difference or fluctuation
cj
cell
xi
features
yi
label
X
set of features
H
hyperplane
w
weight vector
b
bias
The WHO 2018 annual report revealed that 1.35 million people die every year due to road accidents. This death rate contributes 2.5% of total world deaths and ranks eighth, just below diabetes (World Health Organization, 2018). This shows that public awareness of driving safety is still low. Traffic signs are one of the essential factors of road safety, other than vehicle, road, and driver’s conditions, as well as the weather. Therefore, every driver should obey traffic signs to minimize the likelihood of accidents. The Government of Indonesia has regulated traffic through law No. 22 of 2009 on road traffic and transportation, in article 106 paragraph 4, which stipulates that every person who drives a motor vehicle must obey traffic signs, either prohibition or permission. Meanwhile, some European countries have standardize the colors and shapes of traffic signs in 1949 and in the USA followed suit in 1960 (Escalera et al., 2011).
Traffic signs on the road should be clearly distinguishable from other objects. However, in an environment with a complex background, traffic signs can be disguised or obstructed because they lie among trees, billboards, or other objects. Moreover, traffic signs might be physically faded and damaged due to vandalism, making it harder to detect their color and shape. Segmentation is used to separate the color of traffic signs from the background, which is then continued with shape search to find candidates based on feature extraction. Researches on shape features mostly use Histogram of Oriented Gradient (HOG) and Pyramid Histogram of Oriented Gradient (PHOG). HOG uses blocks and cells to determine shape features. In order to improve accuracy, blocks in HOG are made into intersections to allow duplication of processes in the cell. This increases computing speed. Meanwhile, PHOG feature extraction improves HOG in terms of cell size resolution based on level or depth. PHOG uses Canny edge detection for sharper object edges. However, edges of objects other than traffic signs also become sharper, resulting in a significant decrease in the accuracy. Therefore, the Haar–PHOG feature method is proposed to improve the accuracy of the PHOG feature as it conducts calculation on four different frequencies of Haar wavelet transform.
In this research, traffic signs detection is carried out by combining color segmentation in HSI color space and the extraction of Haar–PHOG features. HSI color space can separate traffic signs from complex backgrounds, while Haar–PHOG can emphasize the shape of candidate signs, whether they are circles, diamonds, and squares. In HSI segmentation process, the H and S threshold values are used to obtain red, yellow, and blue sign colors. This is followed by morphology processing to obtain binary images that are free of noise or blob. These two processes result in candidate signs in the form of Region of Interest (ROI).
ROI serves as input for the extraction of Haar–PHOG features. Haar–PHOG feature extraction is a combination of Haar and PHOG wavelet transforms. At this stage, ROI is transformed into Haar discrete wavelets to produce four regions of different frequencies of LL, HL, LH, and HH. Each area is extracted for its PHOG features and results in four PHOG feature vectors. This means that the number of features produced in Haar–PHOG feature extraction is four times those of PHOG features. Afterward, each ROI feature candidate is classified using binary SVM to determine whether the ROI is a traffic sign or not.
This research contributes to the extraction of Haar–PHOG features, which emphasize frequency and resolution. Haar–PHOG combines four regions of different frequencies from Haar discrete wavelet transform using PHOG resolution depth level to produce features that are four times those of PHOG.
There are several sections in this paper. The second section describes some previous studies concerning the detection of traffic signs. The third section focuses on describing the proposed method, while the fourth section contains experiments that have been carried out using training data and testing, as well as the application of the proposed feature extraction method. And the fifth section presents conclusions.
Related work
The color and shape of traffic signs are designed uniquely to highlight their presence. To detect traffic signs, some researchers first used color and followed by matching of shape features (Mogelmose et al., 2012). Traffic signs can be captured by a camera as images or video data for transportation monitoring in megacities (Kalistatov, 2019). RGB color segmentation can separate red, yellow, and blue sign colors with the detection accuracy of up to 92% (Ruta et al., 2010). RGB color space normalization was also used to detect red traffic signs by adding an average threshold value and a standard deviation (Zaklouta et al., 2011). In the meantime, Wang (2014) used RGB color segmentation to separate red, blue, and yellow signs from complex backgrounds using an achromatic model with an accuracy of up to 93.2%. In another study, normalized RGB was used to detect traffic signs made up of mostly red and blue using a threshold value based on experimental results. Results show that the use of normalized RGB is better than HSV color space for the detection of red signs. While for the blue sign, HSV is capable of higher detection accuracy compared to normalized RGB (Berkaya et al., 2016).
The use of RGB color space has a drawback against lighting changes that may result in low accuracy. Therefore, there is a need to use a more robust color space, such as HSV (Chen et al., 2013). H and S values are used as input in the Ada boost classification to produce a binary image with the desired color given a value of 1 and vice versa 0. Results of a study using data in bright, cloudy, foggy, and snowy lighting conditions obtained a detection accuracy of 95% (Fleyeh, 2013).
A study on the detection and recognition of speed limits also used the HSV color space. Speed limit sign was detected by training H and S values using the LVQ (Learning Vector Quantization) artificial neural network. This study obtained a speed limit sign detection accuracy of up to 97% (Biswas and Tora, 2014).
Other than HSV, some researchers used HSI to obtain color segmentation that is resistant to lighting changes. Traffic signs of red, yellow, and blue color were detected based on H, S, and I segmentation, while traffic signs of white color were detected based on achromatic color segmentation (Maldonado-Bascon et al., 2007). HSI color space was also used to detect the presence of traffic signs by separating red, yellow, and blue colors using threshold values (Shengchao et al., 2014). In another study, the H color component was used to localize three primary colors (red, blue, and yellow). Yet, another research used morphological and labeling techniques to obtain relevant ROI as candidates for traffic signs (Han et al., 2015).
Another research tried to increase the accuracy of value during the detection process by extracting features after color segmentation. The invariant moment feature is resistant to changes in rotation, scaling, and the translation used to detect fires in tunnels (Dai et al., 2019). Hough transformation was used to determine the features of a circular speed limit sign (Biswas and Tora, 2014). The texture aspect was applied for feature extraction processes such as LBP, which calculates the value pixel intensity at the center point to neighboring pixels that alters binary code obtained back to decimal (Ojala et al., 2002). Another research developed LBP into three DLBP or three-dimensional LBP on different gray-colored images and color images, including RGB, oRGB, YCbCr, YIQ, and HSV color spaces (Banerji et al., 2013a). Meanwhile, the use of LBP for feature extraction in traffic sign images with CSLBP as local features that are combined with global DWT features. Results show that combined features come with higher accuracy compared with separate use of either Discrete Cosine Transform (DCT) or DWT with significantly faster computing speed (He and Dai, 2016).
Research on feature extraction continues to develop, especially with descriptor-oriented features such as HOG. HOG uses bi-directional convolution operations with horizontal and vertical kernels that allow resistance against lighting changes. From the two convolution matrices, edge strength and angle tangent are calculated, and these result in orientation. Each block and cell is calculated for bin orientation of each descriptor, whether it is bin 180° or 360° (Dalal and Triggs, 2005). After successfully detecting pedestrians, HOG feature was used to identify triangular traffic signs (Fleyeh, 2015). HOG divides ROI into intersecting blocks, and each block is further divided into non-intersecting cells. This process was followed by the recognition of traffic signs using SVM multi-class classification, which was then compared with the use of Kd-tree and random forest (Zaklouta and Stanciulescu, 2012). Feature extraction of HOG was also used for traffic signs. Prior to the classification of traffic signs, an ROI of 100 × 100 pixels is obtained, and features are extracted using an eight bin HOG on each cell. The result was then used for the classification process. There are four classifications used: ANN, k-NN, SVM, and Random Forest. Using the GTSDB dataset, it was found that Random Forest classification has a higher level of accuracy compared to other classification methods (Wahyono and Jo, 2014).
HOG feature extraction was developed into HOG-ring(Soetedjo and Somawirata, 2017) and soft HOG or SHOG using symmetry patterns to determine the number of cells in a block. Thus, the number of cells in each block is not the same. GTSDB dataset was used to test SHOG performance compared to HOG, which implemented with genetic algorithms. Results show that SHOG is more promising compared to HOG (Kassani et al., 2016). The performance of HOG feature extraction on traffic signs was also tested with HSI-HOG, which involved HSI color space, and H, S, and I values are extracted using HOG. The three datasets used (GTSRB, GTSDB, and STS/Swedish Traffic Sign) show that HSI-HOG is better than HOG for all datasets (Ellahyani et al., 2016).
In another study, HOG feature extraction was developed into Haar–HOG. In this method, a discrete wavelet transformation with Haar was performed on an ROI image before HOG processing. The ROI was taken from the segmentation process using several different color spaces such as RGB, Grayscale, HSV, and YCbCr. HOG processing on four quadrants of LL, HL, LH, and HH frequencies was then performed. Both features were then tested using SVM on the Caltech, MIT, and UIUC datasets. Results show that Haar–HOG characteristics had better performance compared to those of HOG (Banerji et al., 2013b).
The characteristics of an object can also be seen as a pyramid consisting of several levels. Similarly, PHOG feature extraction views an image as an HOG pyramid (Adnan et al., 2015). PHOG uses Canny edge detection by calculating edge strength and direction gradients. PHOG feature vector is calculated based on the sum of feature vectors from each level (Bosch and Zisserman, 2007).
Research related to the detection of traffic signs using PHOG feature extraction performed segmentation stages by converting RGB to Gaussian color models and screening the area or extent of the candidate ROI. The results of this stage were then followed by feature extraction using PHOG. However, the use of Canny edge detection has a drawback in the form of noise coming from complex environments or backgrounds. Traffic signs used in this study had ether circle, triangle, inverted triangles, or diamond shapes. Results show that the use of PHOG* feature extraction followed by a binary SVM detection had a better performance compared to PHOG (Li et al., 2015). The detection process can be made even faster using Compute Unified Device Architecture (CUDA) (Razian and Mahvash Mohammadi, 2017) or with the help of tracing using Kalman filtering method (Espejel-García et al., 2017).
Proposed method
Traffic sign detection is important as it serves as input for the next stage of traffic sign recognition. This research used a combination of color segmentation and feature extraction of traffic signs to improve detection performance. The initial stage started with color segmentation using HSI and morphology to produce a binary image that contains ROI as a traffic sign candidate. In the next step, Haar–PHOG feature extraction was used to get the form of traffic signs. This feature extraction emphasizes highlighting contour edges with Haar wavelet transforms that produced four times the number of features compared to PHOG features. In the final stage, SVM classification was used to determine if the ROI feature was a traffic sign or not. The stages of the proposed method are depicted in Figure 1.

Figure 1:
Proposed research stages.
Color segmentation of traffic sign
Segmentation is aimed at separating images of traffic signs from complex backgrounds. Traffic signs generally come in unique colors that segmentation based on color is an option. RGB color space can be an option because it requires low computational level. However, RGB color space is very vulnerable to changes in light intensity that may result in lower accuracy. In this study, HSI color space was used because it is based on human color perception and is relatively more stable to changes in light. The range of H and S values used to obtain basic colors of traffic signs is shown in Table 1 (Shengchao et al., 2014).
Table 1.
Range of hue and saturation threshold values.
| Color | Hue | Saturation | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Red | H ≥ 290 or H ≤ 15 | S ≥ 10 | ||||||||
| Yellow | 20 ≤ H ≤ 65 | S ≥ 150 | ||||||||
| Green | 180 < H ≤ 280 | S ≥ 10 | ||||||||
| Training data (ROI) | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Road | Positive | Negative | Testing data (frame) | Number of traffic sign | ||||||
| Solo-Yogyakarta | 1,500 | 1,500 | 1,100 | 1,153 | ||||||
| Semarang-Yogyakarta | 1,500 | 1,500 | 950 | 929 | ||||||
| Semarang-Solo | 1,500 | 1,500 | 1,000 | 1,099 | ||||||
| Semarang – Salatiga Toll Roads | 1,500 | 1,500 | 950 | 982 | ||||||
| Haar–PHOG | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| HOG | PHOG | LL | LL HL LH | LL HL LH HH | ||||||
| Solo-Yogyakarta | Predict (No) | Predict (Yes) | Predict (No) | Predict (Yes) | Predict (No) | Predict (Yes) | Predict (No) | Predict (Yes) | Predict (No) | Predict (Yes) |
| Actual (No) | 2,328 | 109 | 2,213 | 112 | 2,317 | 56 | 2,308 | 38 | 2,304 | 41 |
| Actual (Yes) | 32 | 1,044 | 147 | 1,041 | 43 | 1,097 | 52 | 1,115 | 56 | 1,112 |
| Haar–PHOG | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| HOG | PHOG | LL | LL HL LH | LL HL LH HH | ||||||
| Semarang-Yogyakarta | Predict (No) | Predict (Yes) | Predict (No) | Predict (Yes) | Predict (No) | Predict (Yes) | Predict (No) | Predict (Yes) | Predict (No) | Predict (Yes) |
| Actual (No) | 2,492 | 132 | 2,356 | 144 | 2,471 | 84 | 2,451 | 56 | 2,451 | 52 |
| Actual (Yes) | 59 | 797 | 195 | 785 | 80 | 845 | 100 | 873 | 100 | 877 |
| Haar–PHOG | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| HOG | PHOG | LL | LL HL LH | LL HL LH HH | ||||||
| Semarang-Solo | Predict (No) | Predict (Yes) | Predict (No) | Predict (Yes) | Predict (No) | Predict (Yes) | Predict (No) | Predict (Yes) | Predict (No) | Predict (Yes) |
| Actual (No) | 1,896 | 232 | 1,807 | 172 | 1,881 | 161 | 1,865 | 145 | 1,868 | 145 |
| Actual (Yes) | 27 | 867 | 116 | 927 | 42 | 938 | 58 | 954 | 55 | 954 |
| Haar–PHOG | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| HOG | PHOG | LL | LL HL LH | LL HL LH HH | ||||||
| Semarang-Salatiga toll | Predict (No) | Predict (Yes) | Predict (No) | Predict (Yes) | Predict (No) | Predict (Yes) | Predict (No) | Predict (Yes) | Predict (No) | Predict (Yes) |
| Actual (No) | 1,504 | 123 | 1,421 | 162 | 1,427 | 103 | 1,433 | 79 | 1,431 | 73 |
| Actual (Yes) | 13 | 859 | 96 | 820 | 90 | 879 | 84 | 903 | 86 | 909 |
| Accuracy (%) | |||||
|---|---|---|---|---|---|
| HOG | PHOG | Haar–PHOG | |||
| Road | Fleyeh (2015) | LL | LL HL LH | LL HL LH HH | |
| Solo-Yogyakarta | 95.99 | 92.63 | 97.18 | 97.44 | 97.24 |
| SMG-Yogyakarta | 94.51 | 90.26 | 95.29 | 95.52 | 95.63 |
| SMG-Solo | 91.43 | 90.47 | 93.28 | 93.28 | 93.38 |
| SMG-SLTG-Toll | 94.56 | 89.68 | 92.28 | 93.48 | 93.64 |
| Precision (%) | |||||
|---|---|---|---|---|---|
| HOG | PHOG | Haar–PHOG | |||
| Road | Fleyeh (2015) | LL | LL HL LH | LL HL LH HH | |
| Solo-Yogyakarta | 90.55 | 90.29 | 95.14 | 96.70 | 96.44 |
| SMG-Yogyakarta | 85.79 | 84.50 | 90.96 | 93.97 | 94.40 |
| SMG-Solo | 78.89 | 84.35 | 85.35 | 86.81 | 86.81 |
| SMG-SLTG-Toll | 87.47 | 83.50 | 89.51 | 91.96 | 92.57 |
| Recall (%) | |||||
|---|---|---|---|---|---|
| HOG | PHOG | Haar–PHOG | |||
| Road | Fleyeh (2015) | LL | LL HL LH | LL HL LH HH | |
| Solo-Yogyakarta | 97.03 | 87.63 | 96.23 | 95.54 | 95.21 |
| SMG-Yogyakarta | 93.11 | 80.10 | 91.35 | 89.72 | 89.76 |
| SMG-Solo | 96.98 | 88.88 | 95.71 | 94.27 | 94.55 |
| SMG-SLTG-Toll | 98.51 | 89.52 | 90.71 | 91.49 | 91.36 |
| Training time (sec) | |||||
|---|---|---|---|---|---|
| Haar–PHOG | |||||
| Road | HOG | PHOG | LL | LL HL LH | LL HL LH HH |
| Solo-Yogyakarta | 148.38 | 512.66 | 15.73 | 27.31 | 32.78 |
| SMG-Yogyakarta | 148.02 | 555.50 | 16.25 | 27.05 | 32.59 |
| SMG-Solo | 168.98 | 537.25 | 15.73 | 28.73 | 33.66 |
| SMG-SLTG-Toll | 165.03 | 583.41 | 18.06 | 32.05 | 35.61 |
| Average | 157.60 | 547.20 | 16.45 | 28.79 | 33.66 |
| Testing time (milliseconds) | |||||
|---|---|---|---|---|---|
| Haar–PHOG | |||||
| Road | HOG | PHOG | LL | LL HL LH | LL HL LH HH |
| Solo-Yogyakarta | 18.40 | 3.90 | 3.00 | 4.50 | 5.00 |
| SMG-Yogyakarta | 19.00 | 4.40 | 2.60 | 4.40 | 5.00 |
| SMG-Solo | 19.50 | 4.00 | 2.60 | 4.40 | 5.10 |
| SMG-SLTG-TOLL | 19.80 | 4.00 | 2.80 | 4.60 | 5.30 |
| Average | 19.18 | 4.08 | 2.75 | 4.48 | 5.10 |








