To the best of my knowledge, my single error is in this piece of code:
# create a composition of two masks:
# the first mask to extract the diagonal elements,
# the second mask to extract elements in the negative_zero_on_duplicate matrix
# that are larger than the elements in the diagonal
mask1 = (fastnp.identity(batch_size)==1)
mask2 = (negative_zero_on_duplicate > positive.reshape(batch_size))
mask_exclude_positives = mask1 | mask2
where the elements in question are in scores. I keep trying to make small adjustments to see how that affects the output, but have yet to see anything that changes the output: 5 passed/1 failed
I’ve also been going over the derivation of TripletLoss, and feel like I get the basic process. There’s a “sweetspot” where the positive and negative (I think they should be called semantically similar and different) vectors work, but, of course, this is the whole point of gradient descent optimization.
I cannot proceed with the rest of the code since it needs a working cost function.
Hi, @arvyzukai, you did help me solve the problem.
I also noticed that I some times get grader errors with Cell 10 if I don’t import:
nltk.download(‘punkt’). Just in case anyone encounters this problem.
I need to drop this course and specialization for the time being. I don’t have enough background in tensorflow or trax (as well as the scope of experience in python, numpy and other tools) to be able to do the exercises without resorting to a web search and an excruciating amount of trial and error coding. I do want to gain this knowledge, so I would be quite open to suggestions regarding other courses that I could take to bridge the gap in my background and what is required for this specialization.
In 2003, I finished my PhD in Cognitive Science using recurrent networks to model human performance in natural language processing, followed by a couple of post-docs in Germany and California in psycholinguistics and spiking neural networks. Over this time, my toolkit was basically low-level C code and tools that interfaced with C such as Tcl/Tk. So, I do have relevant experience in the field, and I have been working with Python and some deep learning more in computer vision than NLP in the past couple of years.
Any advice would be deeply appreciated.
Regarding learning Python - recently I haven’t taken any courses or read books on Python so I feel that I should not be recommending any particular source because I learned Python ten or more years ago, so my knowledge on the “best” up to date sources are lacking to say the least. I would suggest to draw the line how deep you want to learn the subject (Python) and choose the source accordingly because it could easily become a rabbit hole and you would start learning things without your goal in mind.
On numpy - good understanding on manipulating arrays/tensors is crucial in data science and it comes very often in every day work. Again I learned numpy very long ago and I would refrain from recommending any particular source. But, I would recommend to take a “course” on this subject, because hands on problem solving is necessary to internalize the knowledge (just reading the documentation I think is not enough).
What I would really recommend to learn regarding deep learning is Andrew’s Deep Learning specialization. Even it is a couple of years old but in my opinion it covers deep learning foundations very well.
Also I also recently “read” (there are labs you can play with) a very excellent free online book - Dive into Deep Learning. It’s very good for everyone to refresh their knowledge on deep learning or even as a first starting point to cover most of the aspects of it. Highly recommended.
Cheers.
P.S. Do you have any good recommendations on Cognitive Science field books? I find it very interesting too because it’s always at the crossroads with AI, ML, Neuroscience and linguistics but my information about the subject comes mainly from internet, podcasts. By the way, I would also recommend an excellent channel ML street talk if you want high quality (and dense, expertise required) information on AI.
Thanks for your recommendations! I will take you up on them.
I understand your reluctance to recommend courses based on the time since you last had to seriously study them. As you say, the best teacher is using Python and numpy in the course of studying another subject, such as deep learning. I feel the same way with cognitive science. It’s been years since I last wrote a journal article in Cognitive Science, so my recommendations are not up to date. However, as I’ve tried to stay abreast of current theory, I think you might find Embodied cognition interesting. The book by L.A. Shapiro, now in its 2nd edition (2019) by the same title gives a decent account of why I think embodied cognition is the best approach yet to solving the symbol-grounding problem in cognition.
Here’s where theory and practice often diverge. As I was working on NMT using attention, I was very aware of the fact that this model was not cognitively plausible. The main issue was that, in reformulating the “information bottleneck,” this type of architecture allowed the decoder access to all of the input hidden states. In general, humans cannot exactly remember more than five to seven words as they are processing a sentence. A crucial point in a cognitive model is that it is the bottleneck itself that gives rise to behavior that is cognitively plausible, i.e., incremental sentence interpretation as each word is processed in the (compressed) context of the sentence processed at any given moment. This, in turn, leads to robust processing (again sentence repairs such as words like “um” or “er” and repeated words and phrases). There is more, but I just wanted to give you a sense of how task perspective (cognitive modeling vs standard non-cognitively plausible machine translation) affects the architectural choices when building the model.
Well, I’ll leave it at that. I do appreciate your help.
I think the problem with this question in general is that the ‘accepted’ answer is actually not correct, or at least it isn’t consistent with the lecture notes. I’m able to get the accepted answer of 0.7 but to get this answer, your scores, negative_without_positive and closest_negative values should look like:
According to the lecture notes, however, I believe the closest_negative values should be [0.9535077 -0.9535077] since they are the only off-diagonal values in the scores matrix.
Lab ID: vztezjrq
I am having issue with the triplet loss and I could not pass all the test cases. There is an issue with mask_exclude_positives varible. Plz help to resolve this issue
@arvyzukai , I have the same issue with triplet loss function, my code passed 3 test cases, and failed 3. FYI, I tried to use reshape on the positive argument in the masking step.
This is my notebook. Plz help to resolve the issue.
I could easily post here the fresh copy of tripletloss function but future readers might find it outdated so generally it’s better to refresh the workspace.
Hello @arvyzukai. I have the same problem and i did not find the solution after searching for some hours. Can you please guide me through it?
I will message you the notebook.