OpenAI’s models showed real autonomous cyber capability, but the Hugging Face incident also exposes evaluation design, containment, monitoring, and ordinary security failures.