1. Naive Bayes classifier
Create a Naive Bayes classifier for each handwritten digit that support discrete and continuous
features.
Input:
1. TrainingimagedatafromMNIST
You Must download the MNIST from this website and parse the data by yourself. (Please do not use the build in dataset or you’ll not get100.)
Please read the description in the link to understand the format.
Basically, each image is represented by bits (Whole binary file is in big endian format; you need to deal with it), you can use char arrary to store an
a image.
There are some headers you need to deal with as well, please read the link for more details.
2. TraininglabledatafromMNIST. 3. TestingimagefromMNIST
4. TestinglabelfromMNIST
5. Toggleoption
0: discrete mode
1: continuous mode
TRAINING SET IMAGE FILE(train-images-idx3-ubyte)
|
offset |
type |
value |
description |
0000 32 bit integer 0x00000803(2051) 0008 32 bit integer 28
0016 unsigned byte ??
… … …
magic number number of rows pixel
…
|
0004 |
32 bit integer |
60000 |
number of images |
|
0012 |
32 bit integer |
28 |
number of columns |
|
0017 |
unsigned byte |
?? |
pixel |
|
xxxx |
unsigned byte |
?? |
pixel |
TRAINING SET LABEL FILE(train-labels-idx1-ubyte)
|
offset |
type |
value |
description |
0000 32 bit integer 0x00000801(2049) 0008 unsigned byte ??
… … …
The labels values are from 0 to 9. Output:
magic number label
…
|
0004 |
32 bit integer |
60000 |
number of items |
|
0009 |
unsigned byte |
?? |
label |
|
xxxx |
unsigned byte |
?? |
label |
Print out the the posterior (in log scale to avoid underflow) of the ten categories (0-9) for each image in INPUT 3. Don’t forget to marginalize them so sum it up will equal to 1.
For each test image, print out your prediction which is the category having the highest posterior, and tally the prediction by comparing with INPUT 4.
Print out the imagination of numbers in your Bayes classifier
For each digit, print a binary image which 0 represents a white pixel, and 1 represents a black pixel.
The pixel is 0 when Bayes classifier expect the pixel in this position should less then 128 in original image, otherwise is 1.
Calculate and report the error rate in the end. Function:
1. InDiscretemode:
Tally the frequency of the values of each pixel into 32 bins. For example, The gray level 0 to 7 should be classified to bin 0, gray level 8 to 15 should be bin 1 … etc. Then perform Naive Bayes classifier. Note that to avoid empty bin, you can use a peudocount (such as the minimum value in other bins) for instead.
2. InContinuousmode:
Use MLE to fit a Gaussian distribution for the value of each pixel. Perform Naive
Bayes classifier.
Sample input & output (for reference only)
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 |
Postirior (in log scale): 0: 0.11127455255545808 1: 0.11792841531242379 2: 0.1052274113969039 3: 0.10015879429196257 4: 0.09380188902719812 5: 0.09744539128015761 6: 0.1145761939658308 7: 0.07418582789605557 8: 0.09949702276138589 9: 0.08590450151262384 Prediction: 7, Ans: 7 Postirior (in log scale): 0: 0.10019559729888124 1: 0.10716826094630129 2: 0.08318149248873129 3: 0.09027637439145528 4: 0.10883493744297462 5: 0.09239544343955365 6: 0.08956194806124541 7: 0.11912349865671235 8: 0.09629347315717969 9: 0.11296897411696516 Prediction: 2, Ans: 2 … all other predictions goes here … Imagination of numbers in Bayesian classifier: 0: |
32 0000000000000000000000000000 33 0000000000000000000000000000 34 0000000000000000000000000000 35 0000000000000000000000000000 36 0000000000000000000000000000 37 0000000000000011110000000000 38 0000000000000111111100000000 39 0000000000001111111110000000
|
0000000111111111100000 |
|
|
0000001111100011110000 |
|
|
0000011110000001110000 |
|
|
0000011100000000111000 |
|
|
0000111000000000111000 |
|
|
0001111000000000111000 |
|
|
0001110000000000111000 |
|
|
0001110000000000110000 |
|
|
0011100000000001110000 |
|
|
0011100000000011100000 |
|
|
0011100000000011100000 |
|
|
0011110000001111000000 |
|
|
0001111100111110000000 |
|
|
0001111111111100000000 |
|
|
0000111111111000000000 |
|
|
0000001111000000000000 |
|
|
0000000000000000000000 |
|
|
0000000000000000000000 |
|
|
0000000000000000000000 |
|
|
0000000000000000000000 |
|
|
other imagination of numbers goes here … |
|
|
0000000000000000000000 |
|
|
0000000000000000000000 |
|
|
0000000000000000000000 |
|
|
0000000000000000000000 |
|
|
000000000 |
0000000000000 |
|
0000000000000000000000 |
|
|
0000000000000000000000 |
|
|
0000000011111100000000 |
|
|
0000000111111110000000 |
|
|
0000001111000110000000 |
|
|
0000011100000110000000 |
|
|
0000011000000111000000 |
|
|
0000111000001111000000 |
|
|
0000110000011110000000 |
|
|
0000011001111110000000 |
|
|
0000011111111110000000 |
|
|
0000000000111100000000 |
|
|
0000000000011100000000 |
|
|
0000000000011000000000 |
|
|
0000000000011000000000 |
|
|
0000000000000000000000 |
|
|
0000000000000000000000 |
|
|
0000000000000000000000 |
|
|
0000000000000000000000 |
|
|
0000000000000000000000 |
|
0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
- 40 0000 0
- 41 0000 0
- 42 0000 0
- 43 0000 0
- 44 0000 0
- 45 0000 0
- 46 0000 0
- 47 0000 0
- 48 0000 0
- 49 0000 0
- 50 0000 0
- 51 0000 0
- 52 0000 0
- 53 0000 0
- 54 0000 0
- 55 0000 0
- 56 0000 0
- 57 0000 0
- 58 0000 0
- 59 0000 0
60
61 … all 62
- 63 9:
- 64 0000 0
- 65 0000 0
- 66 0000 0
- 67 0000 0
- 68 0000 0
- 69 0000 0
- 70 0000 0
- 71 0000 0
- 72 0000 0
- 73 0000 0
- 74 0000 0
- 75 0000 0
- 76 0000 0
- 77 0000 0
- 78 0000 0
- 79 0000 0
- 80 0000 0
- 81 0000 0
- 82 0000 0
- 83 0000 0
- 84 0000 0
- 85 0000 0
- 86 0000 0
- 87 0000 0
- 88 0000 0
|
89 90 91 92 93 |
0000000000 |
00000000000000000 00000000000000000 00000000000000000 |
0 0 0 |
|
0000000000 |
|||
|
0000000000 |
|||
|
Error rate: 0.1535 |
2. Online learning
Use online learning to learn the beta distribution of the parameter p (chance to see 1) of the coin tossing trails in batch.
Input:
1. Afilecontainsmanylinesofbinaryoutcomes:
2. parameterafortheinitialbetaprior
3. parameterbfortheinitialbetaprior
Output: Print out the Binomial likelihood (based on MLE, of course), Beta prior and posterior probability (parameters only) for each line.
Function: Use Beta-Binomial conjugation to perform online learning. Sample input & output (for reference only)
Input: A file (here shows the content of the file)
1 2 3
0101010111011011010101 0110101
010110101101
1 2 3 4 5 6 7 8 9
10 11 12
$ cat testfile.txt 0101010101001011010101 0110101
010110101101 0101101011101011010 111101100011110 101110111000110 1010010111 11101110110 01000111101
110100111 01101010111
Output
Case 1: a = 0, b = 0
case 1: 0101010101001011010101 Likelihood: 0.16818809509277344 Beta prior: a = 0 b = 0 Beta posterior: a = 11 b = 11
case 2: 0110101
Likelihood: 0.29375515303997485 Betaprior: a=11 b=11 Beta posterior: a = 15 b = 14
case 3: 010110101101 Likelihood: 0.2286054241794335 Betaprior: a=15 b=14 Beta posterior: a = 22 b = 19
case 4: 0101101011101011010 Likelihood: 0.18286870706509092 Betaprior: a=22 b=19 Beta posterior: a = 33 b = 27
case 5: 111101100011110 Likelihood: 0.2143070548857833 Betaprior: a=33 b=27 Beta posterior: a = 43 b = 32
case 6: 101110111000110 Likelihood: 0.20659760529408 Betaprior: a=43 b=32 Beta posterior: a = 52 b = 38
case 7: 1010010111
Likelihood: 0.25082265600000003 Betaprior: a=52 b=38 Beta posterior: a = 58 b = 42
case 8: 11101110110 Likelihood: 0.2619678932864457 Betaprior: a=58 b=42 Beta posterior: a = 66 b = 45
case 9: 01000111101 Likelihood: 0.23609128871506807 Betaprior: a=66 b=45 Beta posterior: a = 72 b = 50
case 10: 110100111
Likelihood: 0.27312909617436365 Betaprior: a=72 b=50 Beta posterior: a = 78 b = 53
case 11: 01101010111 Likelihood: 0.24384881449471862 Betaprior: a=78 b=53
5 6 7 8 9
10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53
|
54 |
Beta posterior: |
a = 85 b = |
57 |
Case 2: a = 10, b = 1
- 1 case1:0101010101001011010101
- 2 Likelihood:0.16818809509277344
- 3 Betaprior: a=10 b=1
- 4 Betaposterior:a=21 b=12
5
- 6 case2:0110101
- 7 Likelihood:0.29375515303997485
- 8 Betaprior: a=21 b=12
- 9 Betaposterior:a=25 b=15
10
- 11 case3:010110101101
- 12 Likelihood:0.2286054241794335
- 13 Betaprior: a=25 b=15
- 14 Beta posterior: a = 32 b = 20
15
- 16 case4:0101101011101011010
- 17 Likelihood: 0.18286870706509092
- 18 Betaprior: a=32 b=20
- 19 Beta posterior: a = 43 b = 28
20
- 21 case5:111101100011110
- 22 Likelihood:0.2143070548857833
- 23 Betaprior: a=43 b=28
- 24 Beta posterior: a = 53 b = 33
25
- 26 case6:101110111000110
- 27 Likelihood:0.20659760529408
- 28 Betaprior: a=53 b=33
- 29 Beta posterior: a = 62 b = 39
30
- 31 case7:1010010111
- 32 Likelihood: 0.25082265600000003
- 33 Betaprior: a=62 b=39
- 34 Beta posterior: a = 68 b = 43
35
- 36 case8:11101110110
- 37 Likelihood:0.2619678932864457
- 38 Betaprior: a=68 b=43
- 39 Beta posterior: a = 76 b = 46
40
- 41 case9:01000111101
- 42 Likelihood: 0.23609128871506807
- 43 Betaprior: a=76 b=46
- 44 Beta posterior: a = 82 b = 51
45
46 47 48 49 50 51 52 53 54 |
case 10: 110100111 case 11: 01101010111 Likelihood: 0.24384881449471862 Beta prior: a = 88 b = 54 Beta posterior: a = 95 b = 58 |
3. Prove Beta-Binomial conjugation
Try to proof Beta-Binomial conjugation and write the process on paper.
※ You should write down the proof process on paper and take a picture. When you hand in HW02, it must contain your code and picture.
l NOTE:
- ¡ Use whatever programming language you prefer.
- ¡ You can’t use numpy.random.beta in HW02. That would be great if you
implement all distribution by yourself.
- ¡ HW02 must contain your code and proof process (can be .pdf or any image
format).






