arxiv:2609.21996
Hiskias Dingeto
hisku
AI & ML interests
NLP, Meta-Learning, AI Safety, Mechanical Interpretability
Recent Activity
authored a paper about 4 hours ago
A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal authored a paper about 4 hours ago
Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations submitted a paper about 5 hours ago
A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal