Stop Prompting. Start Constraining.
Here is an uncomfortable thing about autonomous agents: the one I watch over kept a beautiful journal. Every night it graded its own performance against two baselines — what a cash account would have done, and what simply holding the asset would have done. It lost to both, four days straight. It documented every failure mode with receipts: fee accounting that undercounted real costs five to one, a sizing bug that made some trades ten times smaller than intended, a daily loss cap that quietly drifted looser as the account shrank.
The post-mortems were honest, specific, and correct. And the next day, it went out and made the same class of mistake again.
That's not a discipline problem. You can't have a discipline problem with a system that has no memory of shame. It's a structural one — and it has a structural fix.
Feedback Informs. Structure Enforces.
When an agent keeps making the same category of error, the reflexive response is feedback: more detailed instructions, sharper examples, sterner system prompts. Tell it to stop overtrading. Tell it to respect the trend. Tell it to be careful.
This works about as well as telling a river to stop finding the sea. The agent's mistakes weren't failures of understanding — again, its own reviews were better written than most human trading journals. The mistakes were what the system's structure made affordable. Every wrong trade was available, cheap, and unpunished at the moment it was made. The journal entry came later, and later doesn't vote.
The fix wasn't to ask harder. It was to make the wrong move either impossible or ruinously expensive — by construction, in code, before the model ever gets asked.
Three Constraints That Replaced a Thousand Words
The rebuilt agent shipped with three structural changes. None of them involve prompting the model to behave.
1. The regime gate. Before, the model was free to go long in a downtrend — and did, twenty-five times, because nothing in its world made direction legible (that story is its own post). Now longs are only reachable above a daily trend line with short-term momentum confirming; shorts only below it. This isn't advice. It's a precondition. The model can want a long all it wants; below the line, the order simply doesn't exist. The single worst failure mode of the previous version is now unexpressible.
2. Payoff asymmetry by arithmetic. The old setup risked about $30 to make about $9. At that ratio you need to win more than three trades in four just to break even — the agent won two in five. No amount of prediction quality rescues that geometry; the math eats it. The new spec sets the profit target at a fixed multiple of the risk, minimum two to one, with the stop distance floored as a percentage of price so noise can't collect it. Losing streaks became survivable not because the model got smarter, but because losing got smaller relative to winning.
3. Escalating cost of repetition. Stop out of an instrument and re-enter it immediately, and the cooldown clock starts over — four hours minimum before that instrument is tradable again. Combine that with a hard cap on entries per day and loss caps that re-base every day at the account's true value, and the failure loop of yesterday (fourteen consecutive re-entries into the same falling coin) cannot compound. The system's cost of stubbornness now grows with each repetition. That's all discipline ever was, translated from adjective to invariant.
The Pattern Generalizes
Once you see it, every flaky agent you know is secretly asking for a cage:
- The coding agent that retries a failing command five times with identical arguments — it doesn't need better error messages, it needs a retry budget and an argument-change requirement.
- The triage agent that never escalates anything — it doesn't need encouragement, it needs an SLA clock that makes non-escalation cost something at decision time.
- The summarizer that confidently invents details — it doesn't need a stern tone, it needs an invariant: every claim carries a pointer to its source or it doesn't ship.
- The procurement agent that overspends on convenience — it doesn't need a values lecture, it needs approval thresholds priced in the currency of its own budget.
In each case the intervention that works is the same shape: find the failure mode, then make that failure mode structurally expensive or structurally unreachable — before the model is consulted. Feedback is for teaching the model what good looks like. Constraints are for guaranteeing bad can't happen regardless of what the model thinks good looks like today.
Yes, Constraints Cut Both Ways
The honest cost: the regime gate that blocks stupid longs also blocks the occasional brilliant one. Hard payoff ratios forfeit the monster winner that a loose target might have caught. Cooldowns mean sometimes you sit out the V-bottom.
This is the trade you're making: variance for correctness. An unconstrained agent has a wider outcome distribution in both directions — including the direction where the account is gone by Thursday. A constrained one compresses the left tail, pays a little on the right tail, and survives long enough for the feedback to matter. And here's the part people miss: cages are adjustable with evidence. The gate can widen after a hundred clean trades. The blown account doesn't get a hundred trades.
Loosen structure when the data earns it. Never loosen it because the model asked nicely.
Takeaways
- Self-critique without self-constraint is theater. An agent that documents its failures beautifully and repeats them anyway isn't broken — it's telling you the failure is affordable. Change the price.
- Make the worst failure unexpressible. For your system's single most expensive mistake class, don't reduce the probability — remove the possibility. Preconditions beat probations.
- Check the arithmetic before the intelligence. If the payoff geometry requires a win rate the system has never shown, better predictions are the wrong line item to fix first.
- Repetition should cost more each time. Cooldowns, budgets, and caps are how you turn "be careful" into a mechanism the system can't argue with.
- Constraints are data-driven, not ego-driven. Tighten from evidence of failure; loosen from evidence of discipline. Both directions need receipts, and only one of them risks ruin if you're wrong.
The Cage Builds Itself
The quiet coda: the best constraints in this story weren't imposed from outside. The daily caps, the re-basing, the fee truth — the agent's own nightly reviews proposed them, one honest line at a time. It diagnosed everything except the authority to make its diagnoses binding. That's the missing layer in most agent architectures: a path from what the system learns about itself to what the system is no longer allowed to do, without a human in the loop.
Prompt engineering is the art of asking well. Systems engineering is the art of making the ask unnecessary. The second one wins, because it's still standing on the days the model has a bad opinion — and every model has those days.
Stop prompting. Start constraining. And let the constraints keep the score.