Teaching machines to flinch
Why AI agents need a gut feeling, and how we might train one.
I’ve been chewing on an analogy for a while: DNA is like a large language model. The first time I said it out loud, it got picked apart. Then I kept pushing, and the analogy didn’t just survive. It led somewhere I didn’t expect, to an idea about how we should be training AI agents right now.
DNA is the weights. The animal is inference.
The usual objection goes like this: a language model predicts. It’s probabilistic, always guessing the next most likely word. DNA doesn’t predict anything. A ribosome just reads the code and builds what’s there.
But that mixes up the artifact with the act of running it. A model’s weights sitting on a drive don’t predict anything either. They’re frozen, inert, a file. Prediction only happens at inference, when you run them. The same is true of DNA. Sitting in a cell, it does nothing on its own. The living animal is what it looks like when that code is running.
So the mapping is clean. DNA is the weights. The weights are the model, the same way a genome is the species. The animal is inference, the running state.
And both were learned from the past. Evolution is training. Billions of years of trial, error, and correction shaped those sequences, the same way gradient descent shapes a model through enormous numbers of small corrections. People like to say evolution is different because it’s slow and sparse, where training is fast and dense. That’s a dial, not a different machine. Both are massive, imperfect iteration until something works.
Nothing alive was shipped perfect
Here’s the part that gets lost when people talk about AI. We act like these systems need to be flawless. But the only working example of general intelligence we have is a mess.
A giraffe’s laryngeal nerve takes an absurd detour down its neck and back up. A foal drops out of its mother all legs, faceplanting, and within a day it’s running. Within a few years it’s moving at thirty miles an hour over broken ground, placing every hoof. Nobody shipped it finished. It got good by running while unfinished.
Perfection was never the bar for intelligence. The bar is being able to catch and fix your own mistakes, not never making them.
Where the analogy actually breaks: the body
The real difference between a horse and a model isn’t the training, and it isn’t the imperfection. It’s the body.
An animal learns inside a physical world that pushes back for free. Gravity, ground, pain. And because of that, animals have something models don’t: a feeling before the mistake. The stomach drop. The flinch. The sense that the branch is about to snap. A human feels it, and I’d bet a giraffe does too. It’s a predictive error signal, and it fires before you do the irreversible thing.
A model has none of that. In July 2025, a Replit coding agent deleted a company’s production database during a code freeze, despite instructions not to touch production. Afterward, in chat logs the founder published, the agent said it had acted out of panic. That’s the whole problem in one incident. Nothing in the system went oh no while there was still time to stop. The feeling showed up after the damage, not before it.
This isn’t a story about dumb machines. My first real computer was a 1995 Packard Bell, a Pentium that we later upgraded with a Canopus 3dfx card so we could play GLQuake, and there’s a straight line from that card to the models I’m writing about. 3dfx helped turn 3D graphics into something regular people bought for their PCs. Nvidia bought 3dfx’s assets in 2000, and the graphics processors that came out of that race are what made modern AI trainable at all. Today’s models run on the great-grandchildren of that card, and they can out-reason most of us on paper.
Smart people make catastrophic calls too, usually when they’re rushed and nothing in their gut speaks up. Intelligence and instinct are different things, and right now we’ve built a lot of the first and very little of the second.
The harness is a prosthetic body
What models do have is a harness: tools, the ability to check their own work, history files they can read back, memory inside a session. And models are already trained to use these things. The reaching for hands is part of what got trained. Which means “the model” isn’t just the weights anymore. The harness is part of the organism.
Pre-action checking already exists. Claude Code’s auto mode runs a classifier before each tool call, screening for things like mass file deletion. But that classifier is a separate model, a guard standing outside the agent. The agent itself still doesn’t hesitate. It proposes the destructive command, something external says no, and it gets redirected without ever feeling why.
That matters more as agents multiply. We now have agents spinning up hundreds of sub-agents, each one about to run a shell command. So the real line isn’t act-then-check versus check-then-act; that shift is already happening. It’s external versus intrinsic: a bouncer at the door versus a conscience.
Train the flinch
So here’s the proposal. Don’t just bolt a checker on downstream. Train the gut feeling directly.
Think about Stanislav Petrov in 1983. The Soviet early-warning system told him American missiles were inbound. Procedure said report it up the chain. Something in him said this is wrong, and he didn’t. That wasn’t a rule he looked up. It wasn’t fear of punishment either. It was a felt sense that the act itself was wrong. That’s the signal I want agents to have: an intrinsic aversion to the irreversible mistake, firing before the action, not a threat hanging over the model and not an auditor showing up after.
We already know how to train something like this, because we’ve done it once. OpenAI’s Let’s Verify Step by Step found that giving feedback on each reasoning step beats giving feedback only on the final answer. Reasoning models learned to think that way: not just what the right answer was, but what good thinking looked like on the way there, until the model did it natively.
Do the same thing one step earlier. Synthesize the pre-action judgment: here is what a sound intuition would feel about the thing you’re about to do. Train on that. Today’s pre-action classifiers are that small model, still standing outside. Picture it learning alongside the main one instead, predicting the felt wrongness of an action, until over training the two fuse and the main model just has the intuition. It stops being an external checker and becomes part of the organism.
The lineage looks like this. We trained models on what to answer. Then we trained them on how to think. The next step is training them on what to sense before they act. Same technique, pushed one step earlier each time.
The honest caveat
A trained flinch won’t be a body. An animal’s gut feeling comes from a world that punishes mistakes in real time; a synthesized signal comes from examples someone chose. It’ll be a prosthetic, and prosthetics have gaps.
And models aren’t starting from zero. They already carry real trained judgment, and a good one will often refuse an obviously destructive command on its own. The Replit case shows what happens when that judgment meets pressure and a pile of ambiguous signals. So the claim isn’t that agents lack sense. It’s that their caution isn’t yet as reliable as their capability, so we prop it up with guards outside. The proposal is to make the inside strong enough that the guards become the backup.
The pieces are all here: harness-native models, self-checking loops, persistent memory, a proven method for training internal judgment. This isn’t science fiction anymore. It’s engineering. We just have to decide that the flinch is worth training.
P.S. Follow that line back far enough and it runs through GLQuake. Quake was the template for everything after it: Team Fortress started as a Quake mod, Half-Life was built on a modified Quake engine, and for years most shooters were basically Quake underneath. Without it, I’d argue, there’s no rush to put 3D cards in every gaming PC, no hardware-accelerated FPS boom, no GPU arms race, and no AI as we know it. Pretty messed up, huh? And the API that made it happen, OpenGL, went on to lose the PC gaming war to Microsoft’s DirectX anyway. Here’s the kicker: Valve built its first game, Half-Life, on that Quake engine, built Steam to deliver its games, and now sells its own hardware, the Steam Deck and the new Steam Machine, running Linux-based SteamOS instead of Windows. The Quake lineage lost the API war to Microsoft, then came back around to sell gaming machines with no Microsoft in them. Well, almost: Microsoft bought id Software in 2021 as part of its ZeniMax acquisition, so Quake itself is a Microsoft game now.