Problem Description:
Many machine learning algorithms don't work well with continuous typed data. A common technique to work around this problem is to discretize the continuous values into a few categories.
The discretization algorithms are also called binning algorithms. There are at least three types of binning algorithms: equal width, equal frequncy (depth), and information based.
In equal-width binning algorithms, the data range gets divided to almost same width segments, each segment corresponding to one bin, thus the width of all the bins are roughly the same. The data items are dropped to their corresponding bins based on their values. The number of data items in each bin can vary greatly.
In equal-frequency (also called equal-depth) binning algorithms, the data range is still divided to several segments/bins, but the width of each bin is designed/calculated in such a way that each bin would contain roughly the same number of data items.
The equal width and equal depth algorithms only consider the value distribution in the attribute to be discretized, and don't take the target attribute value distribution into consideration. Information based discretization algorithm tries to minimize the entropy in the dataset when it discritizes the continuous values.
The easiest information based discretization algorithm tries to find a value to split all the data items in the dataset to two bins, such that the entropy of the dataset is minimum after the split compared with splitting using any other values.
Your tasks:
In the dataset Weather Dataset, the attributes temperature and humidity are both continuous (real) typed. Use these two attributes and their corresponding data as examples to practice splitting their values to two categories using equal width, equal depth and information based discretization methods.