Hello,
As I was going through the ungraded assignment - chart generation - in Module 2: Reflection of the Agentic AI course, I was curious as to how we could tackle hallucinations of the LLMs. For example, one of the modifications I made in the reflection prompt was to have the chart be color-blind friendly. However, the output from the reflection did not satisfy this condition. In fact, the graph generated from the generation_model was more appropriate for accessibility.
How do we identify when LLMs have hallucinated? How does this effect our Agentic workflow
Should it be part of an eval metric?
Thanks in advance!
if you notice the type of color blindness hue, agentic ai did actually a good job on your prompt
it more changes the hue which effects more such as red color, so in this instance agentic ai changed the color with red hue, which seems actually fine.
I wasn’t aware of this. Thank you!
Although, the question remains: how do we know if an LLM call hallucinated while producing the result, either at the generation stage or the reflection stage?
And is there a way we factor in hallucinations for evals?
hi @Devika_Vishnu
to identify hallucinations during generation or reflection using few strategies
-
token probability (generation stage) - Check the model’s token-level logits or perplexity. Low confidence scores or high perplexity during decoding indicate the model is guessing.
-
self-consistency (generation or reflection stage) prompt the model to produce multiple responses at a higher temperature, or run multi-sample generations. If the independently sampled answers contradict one another, the model is likely hallucinating.
-
LLM-as-a-Judge (at reflection stage) use a separate, powerful model to review the generated output against the provided reference context. Prompt the judge to break the text into individual claims, trace them back to the source, and check for logical failue, unsupported claims of contraindications.
-
guardrails (reflection stage) use strict guardrails like deterministic code executors for calculations or entity extraction tools to verify that dates and proper nouns in the output match a structured ground-truth database. one needs to know their metadata used throughly.
to factory hallucination into evals use semantic natural processing interference. you can test whether the generated text logically matches or contradicts your source material. track the percentage of tokens or answers flagged for contradiction to create a continuous metric.
you can use factual precision or factual score using external knowledge graphs or zero-shot entailment approach where you create classifier system that checks relationship between two sentences into one or three categories.
entailment (true) generated text is fully supported by the source document.
contradiction (false/hallucination) generated text directly conflicts with the source document.
neutral (unknown) generated text will provide information that isn’t mentioned in the source, in which case is known as external hallucination.
famously used one is DeBERTa(Decoding-enhanced BERT with Disentangled Attention) from Hugging face that hold different vector presentation for content and it’s relative position.
Hope this helps.
Regards
Dr. Deepti