The Programming Assignment from Module 2, Week 4 (Create your own Deep NN) has been extremely valuable. I have a question on the Cost function in the assignment. Can anyone please suggest the math for modifying it to handle IMBALANCED datasets (say, class 1 only about 2% in the entire dataset of 1 million examples)? Thanks
Imbalanced datasets are a big problem.
How many classes are there in your dataset?
Here are some thoughts.
For example, if there are 10 classes, you only need about 10% of the examples in each class (since using one-hot labels, you’re doing "1-out-of-N) anyway.
If it’s a binary classifier, you might have to discard examples of the other labels in order to get better balance, and compensate for this by running the training many times with different randomly discarded examples.
If you don’t really need a complex model, then you could perhaps use an anomaly detection method instead of an NN.
In addition to what Tom is mentioning above, I would say for imbalanced datasets to also check and use additional metrics such as precision, recall, and F1 score, which are specifically used to calculate accuracy for such datasets.