16/06/2026
🔬 Advancing AI Safety Through Cutting-Edge Research
We are proud to celebrate an outstanding achievement by researchers from the Andrew and Erna Viterbi Faculty of Electrical and Computer Engineering.
A series of studies led by Prof. Haggai Maron’s group, in collaboration with researchers from other universities and NVIDIA, has resulted in three papers being accepted to top-tier machine learning conferences: NeurIPS 2025, ICLR 2026, and AAAI 2026.
The research was led by PhD student Guy Bar-Shalom (co-advised by Prof. Ran El-Yaniv) and postdoctoral researcher Fabrizio Frasca, together with Dr. Yftah Ziser (University of Groningen and NVIDIA).
Their work addresses one of the most critical challenges facing artificial intelligence today: detecting hallucinations and other failure modes in large language models (LLMs).
Rather than focusing only on a model’s final response, the researchers developed innovative deep-learning methods that examine the model’s internal computational processes, including activations, attention patterns, and output probability distributions, enabling more accurate and systematic detection of unreliable behavior.
These advances offer a powerful new framework for monitoring AI systems and could play an important role in the safe deployment of AI in healthcare, education, scientific research, and other high-stakes domains.
Congratulations to the researchers on this remarkable accomplishment!
הפקולטה להנדסת חשמל ומחשבים ע"ש ויטרבי טכניון
Read more: https://ffabffrasca.substack.com/p/detecting-llm-misbehaviors-from-the