OpenAI’s models showed real autonomous cyber capability, but the Hugging Face incident also exposes evaluation design, containment, monitoring, and ordinary security failures.
The Foundation Labs Case File: OpenAI, Hugging Face, and the Getaway Car


OpenAI’s models showed real autonomous cyber capability, but the Hugging Face incident also exposes evaluation design, containment, monitoring, and ordinary security failures.

agent-skills v0.7.0 lands 31 new skills, upgrades 20 more, and flattens the repo: the deepest open-source AI agent skill library keeps compounding.

A preliminary look at DeepSeek’s new coding-agent harness, what an early source audit found, where its authority reaches, and when it deserves a closer look.

Why voice transcripts make better agent context than typed notes: capture-before-curation, a record that shapes the room, and patterns an agent can surface.

A solo founder is ten people on Monday and zero reviewers on Saturday. Here’s what changes when the agent carries the discipline.

OpenAI GPT-5.6 Luna, Terra, and Sol look like a sensible tiered lineup. Benchmark evidence suggests choosing by the cost of finished work instead.

The life-coach Agent Skill helps you examine hard decisions, map ambivalence, and test goals, and refuses to pretend it’s something it isn’t.

Ponytail made my AI agent terser. Neckbeard makes it prove its work. Why I ditched the viral coding skill for one designed to catch mistakes, verify at the real boundary, and say what it couldn’t check. The architecture, authority gates, and evaluation suite inside.

A practical method for earning AI autonomy by finding workflow bottlenecks, strengthening adjacent capabilities, and expanding only with evidence.

Postmortems matter when their lessons survive the meeting and constrain the next release, without turning blameless learning into automated blame.

Andrej Karpathy predicted a small AI model that would trade encyclopedic knowledge for raw reasoning. Just over a year later, it is a product category, and it changes everything about how we should think about AI infrastructure.

Human organizations have a maximum perceivable rate of change. Above that threshold, they don’t accelerate their response. They slow down. They treat a category shift as a tool upgrade. The shift from deterministic to probabilistic computing has exceeded the threshold, and the three gaps we see everywhere are not failures of leadership. They’re symptoms of an organizational immune system working exactly as designed for a world that no longer exists.

I had an AI agent extract and analyze all 464 YC Library articles. The word ‘flywheel’ appears exactly once. But the mechanics are everywhere, hiding in plain sight. Y Combinator has described a new economic primitive that replaces capital as the lubricant for startup growth. They just refuse to name it.