AI agents that act, not just simulate
The companion bet to AI simulations: agents that execute real transactions, not rehearse them. The product is the trust and liability structure around the action, not the model.
Aug 5, 2026
The last thesis was about practicing before the real thing. This one is about the real thing. An agent that books the flight, moves the money, files the return, is a fundamentally different bet than one that role-plays the conversation first. Most of what gets called an "agent" right now is a chatbot with function-calling bolted on. That's not the same claim.
The tell is reversibility. A simulation's whole value is that the mistake is free. An acting agent's mistake is not. That single fact changes everything about what has to be built underneath the model: not better prompts, but a trust and liability structure for what happens when it's wrong, because at scale, it will be.
This is the professional-services liability argument again, generalized past law and audit. Someone has to be accountable when an autonomous action goes wrong, whether that action is a refund, a trade, a contract signature, or a support ticket resolved the wrong way. Founders who treat that accountability as a legal afterthought instead of the actual product are building on top of a hole they haven't noticed yet.
The domains worth building in aren't the highest-stakes ones or the lowest, they're the ones with bounded, recoverable consequences: real enough that a human doesn't want to do it by hand, contained enough that a mistake doesn't end a company or a life. Moving a person's life savings unsupervised is the wrong place to start. Rebooking a flight when one leg gets cancelled is closer to right.
What actually makes this buildable is staged autonomy, not full autonomy from day one. The agent proposes, a human approves, and only after a track record earns it does the loop tighten. Audit trails and insurance get designed in from the first version, not retrofitted after the first expensive mistake. I'm looking for founders building that trust ladder as the core product, in a domain they picked because the stakes are real but bounded, not founders who wrapped function-calling around a chat model and are calling it agentic.
More theses