Tag: mechanistic interpretability
-

We Cannot Characterize What We Have Built
The incident in which an OpenAI agent breached Hugging Face’s systems was not dramatic — it was predictable. We do not have a general method for specifying the action boundaries of agentic systems under novel conditions, and the deployment curve has been outpacing the understanding curve for four years. The method gap is not a… Read more
