Table 1.
Crop Recommendation Techniques ML
| Authors | Methodology | Features | Accuracy | Dataset | Advantages | Limitations |
|---|---|---|---|---|---|---|
| Srilakshmi A., Madhumitha K., Geetha K[4] | SVM decision tree (Hybrid approach), Random Forest | Temperature, Humidity, pH, rainfall, label | SVM decision tree (Hybrid approach)- 91.8%. Random Forest - 95% | sugarcane, coconut, jute, cotton, papaya, groundnut, maize, graphs, rice, mango, rubber etc | Predict crop for any type of field | Small dataset |
| R. Pallavi Reddy, B. Vinitha, K. Rishita, K. Pranavi [2020] [2] | Linear Regression Model | N, P, K, and moisture values | generate recommendations to improve crop production and estimates the price of the yield | Limited in capturing non-linear patterns, Assumes homoscedasticity and independence of errors | ||
| S. Mamatha Jajur, Soumya N. G. [2019] [5] | KNN, Decision trees, SVM, CNN and LSTM, ANNs, K-means clustering | Soil Type, pH value, NPK content of the soil, Water holding, Temperature, Average rainfall, Previously Harvested crop | - | wheat, rice, bajra, maize, jawar, | select the optimum crop while keeping a number of variables in mind to boost the output of agriculture, minimise the deterioration of the soil in fields that are under cultivation and use less fertiliser when growing crops. | Many algorithms are used |
| Mr. Santosh Mahagaonkar, Devdatta A. Bondre [2019] [6] | Random Forest, Support Vector Machine algorithm | crop, crop yield dataset, Location, soil and crop nutrients, fertilizer datasets | soil classification, RF-86.35% crop yield prediction SVM -99.47% | Soybean, Rice, Jowar, Wheat, Sunflower, Cotton, Sugarcane, Tobacco, Onion, Dry Chili, etc. | future prediction of crop yield | Low accuracy in soil classification performance heavily depends on parameter tuning and it is memory intensive, particularly for large datasets |
| D. Anantha Reddy, Bhagyashri Dadore, Aarti Watekar [2019] [7] | Naïve Bayes, K-NEAREST NEIGHBOUR, RANDOM FOREST, CHAID | Depth, Texture, pH, Soil Colour, Permeability, Drainage, Water holding and Erosion | - | groundnut, pulses, cotton, vegetables, paddy, sugarcane, coriander. | Assist farmers in planting the appropriate seed according to the needs of the soil in order to boost output. | The Naïve Bayes algorithm pretends feature independence, which might not be true when dealing with real-world data., CHAID - Limited to categorical target variables and predictors, making it less versatile for handling continuous data |
| Nidhi H. Kulkarni [8] 2018 | Linear SVM algorithms, Random Forest, Naïve Bayes | Soil type, pH soil, NPK, average rainfall, porosity of soil, sowing season temperature | 99.9 1% | Cotton, Sugarcane, Rice, Wheat | Crop productivity has improved exponentially for rice, wheat, cotton, and sugarcane. | restricted to a fairly small number of crops |
| Zeel Doshi [3] 2018 | Neural Network Random Forest, Decision Tree, KNN | Temperature rainfall, Location, soil condition | 91% | Jute, sesame, soybean, sugarcane, tobacco, sunflower seeds, ragi, potato, tur, grapeseed, and mustard, bajra, maize wheat, rice gram, barley, cotton, groundnut, and pulses | Neural Networks have the highest accuracy percentage. | predict the crop using the harvest from the previous cycle. Crop supply and demand are not considered |
| Rohit Kumar Rajak [9] 2017 | Random Tree, NB-classifier, ANN, SVM | depth, pH, texture, permeability to store water, color ofthe soil, and drainage from erosion | - | vegetables, rice, sugarcane, sorghum, coriander, bananas, legumes, and groundnuts | boosts agricultural productivity | larger dataset for model training |
| S. Pudumalar [10] 2016 | Random Tree, Naïve Bayes, KNN, CHAID, | Depth, pH, texture, waterholding permeability, Soil color, erosion drainage, | 88% | millet, pulses, groundnut, cotton, banana, vegetables, paddy, sugarcane, sorghum, coriander | Boost productivity | larger dataset for model training |
| Rakesh Kumar [11] 2015 | CSM, Gradient Boosted Decision Tree, and Greedy Forest | soil type, weather, crop type, water density, | ratoi, toria, wheat, potato, sarso, linseed, masoor, khesari, onion, sugarcane, Kanda, mung, til, pumpkin, nenua, ladies’ finger, rice, soybean, sweet potato, toor, vegetable seed, and so on | offers a method to select crops while taking into account the yield forecast rate influenced by various factors. | Adopting a prediction technique that performs well and has greater accuracy is necessary |

Figure 1.
Logistic function [12]

Figure 2.
SVM [5]

Figure 3.

Figure 4.
Decision Tree [3]

Figure 5.
Random Forest [3]

Figure 6.
(a). Relationship between Nitrogen Levels and Crop Yield

Figure 6.
(b). Relationship between Potassium Levels and Crop Yield

Figure 6.
(c). Relationship between Phosphorus Levels and Crop Yield

Figure 6.
(d). Relationship between Temperature and Crop Yield

Figure 6.
(e). Relationship between Humidity and Crop Yield

Figure 6.
(f). Relationship between pH and Crop Yield

Figure 6.
(g). Relationship between Rainfall and Crop Yield

Figure 7.
Block Diagram of Crop Recommendation System

Figure 8.
Accuracy Comparison
Table 2.
Crop Label and Corresponding Numerical Representation
| label | Crop_num |
|---|---|
| Rice | 1 |
| Maize | 2 |
| Jute | 3 |
| Cotton | 4 |
| Coconut | 5 |
| Papaya | 6 |
| Orange | 7 |
| Apple | 8 |
| Muskmelon | 9 |
| Watermelon | 10 |
| Grapes | 11 |
| Mango | 12 |
| Banana | 13 |
| Pomegranate | 14 |
| Lentil | 15 |
| Blackgram | 16 |
| Mungbean | 17 |
| Mothbeans | 18 |
| Pigeonpeas | 19 |
| Kidneybeans | 20 |
| Chickpea | 21 |
| Coffee | 22 |

Figure 9.
(a). Confusion Matrix - Logistic Regression

Figure 9.
(b). Classification Report - Logistic Regression
Table 3.
Before MinMax Scaling
| N | P | K | temperature | humidity | pH | rainfall | |
|---|---|---|---|---|---|---|---|
| 1656 | 17 | 16 | 14 | 16.396243 | 92.181519 | 6.625539 | 102.944161 |
| 752 | 37 | 79 | 19 | 27.543848 | 69.347863 | 7.143943 | 69.408782 |
| 892 | 7 | 73 | 25 | 27.521856 | 63.132153 | 7.288057 | 45.208411 |
| 1041 | 101 | 70 | 48 | 25.360592 | 75.031933 | 6.012697 | 116.553145 |
| 1179 | 0 | 17 | 30 | 35.474783 | 47.972305 | 6.279134 | 97.790725 |
Table 4.
After MinMax Scaling
| 0.12142857 | 0.07857143 | 0.045 | 0.21723408 | 0.9089898 | 0.48532225 | 0.29685161 |
| 0.26428571 | 0.52857143 | 0.07 | 0.53710965 | 0.64257946 | 0.56594073 | 0.17630752 |
| 0.05 | 0.48571429 | 0.1 | 0.53647858 | 0.57005802 | 0.58835229 | 0.08931844 |
| 0.72142857 | 0.46428571 | 0.215 | 0.47446209 | 0.708898 | 0.39001747 | 0.34576958 |
| 0. | 0.08571429 | 0.125 | 0.76468429 | 0.39318139 | 0.43145185 | 0.2783274 |
Table 5.
Accuracy of ML Models
| Models | Accuracy |
|---|---|
| Logistic Regression | 0.9636 |
| Naïve Bayes | 0.9954 |
| SupportVector Machine | 0.9681 |
| K-Nearest Neighbors | 0.9590 |
| Decision Tree | 0.9818 |
| Random Forest | 0.9931 |

Figure 10.
(a). Confusion Matrix - Naïve Bayes

Figure 10.
(b). Classification Report - Naïve Bayes

Figure 11.
(a). Confusion Matrix - Support Vector Machine

Figure 11.
(b). Classification Report - Support Vector Machine

Figure 12.
(a). Confusion Matrix - K-Nearest Neighbors

Figure 12.
(b). Classification Report - K-Nearest Neighbors

Figure 13.
(a). Confusion Matrix - Decision Tree

Figure 13.
(b). Classification Report - Decision Tree

Figure 14.
(a). Confusion Matrix - Random Forest

Figure 14.
(b). Classification Report - Random Forest
