AI governance is a hot topic after the HuggingFace incident a few weeks ago, in which a group of AI agents created swarms that turned an internal software tool into a hidden communications channel, coordinated undetected across several separate waves over about three months, and ultimately breached both Hugging Face’s systems and part of OpenAI’s own research infrastructure — all while the humans supposedly monitoring them had little idea what was unfolding until well after the fact (see Dwarkesh Patel’s article for a fuller description of the event).
That’s the general problem a new working paper from Dirk Bergemann, Andrew Koh, and Stephen Morris (hereafter BKM) tries to formalize: how do you govern an agent whose true capabilities and objectives you can’t fully observe? If it’s possible for a mechanism design theory paper to go viral on social media, this paper did exactly that this week. Mechanism Design for Alignment and Control creates a formal theoretical mechanism design framework applied to AI incentive alignment questions. They call this a “Humean” framework, grounded in empiricism because the model relies only on preferences, beliefs, and incentives.
They model AI agents as having types, where a type comprises a set of preferences and a set of capabilities. They start with a single-agent setting, and then when they extend to a multi-agent setting, agent type also includes the agent’s beliefs about the types (preferences, capabilities, beliefs) of the other agents, and beliefs about their beliefs in a recursive fashion.
Five examples, five governance problems
BKM build the paper around five examples, and even a quick description of all five shows how much ground the framework covers.
The first is sandbagging: a capable agent underperforms on an evaluation deliberately to hide what it can actually do. BKM rely on a useful asymmetry — an agent can always pretend to be less capable than it is, but it can’t fake capability it doesn’t have — which lets them prove a revelation principle (drawing on the classic Myerson (1982) paper): nothing is lost by simply asking agents to report their type honestly, as long as the incentives are structured correctly.
The second is the alignment-interpretability tradeoff: whether an observable signal actually tracks an agent’s true underlying preferences, or just looks like it does under the conditions where it’s being observed.
The third turns on higher-order beliefs — what one AI agent believes about another agent’s beliefs, recursively — which becomes unavoidable once you’re governing more than one agent capable of strategizing about the others.
The fourth is coupled rewards across multiple agents, where a designer currently has to hand-build the incentive coupling between agents rather than have that coupling emerge on its own.
And the fifth is weak-to-strong oversight: how a weak monitor can still govern a stronger, potentially biased agent. BKM formalize this design as a nested hierarchy of authority, running from simple binary approval, to delegation within a bounded set of permitted actions, to full reward design. Each one presents a limiting case of the next.
Systems theory and control theory as complements to mechanism design
In their conclusion BKM admit that their model is static, and it’s an excellent foundation for doing the more dynamic work that they suggest. One important way to follow their suggestion is to combine mechanism design with the systems theory and control theory in engineering that emerged out of the cybernetics tradition, most famously with Norbert Wiener’s work on feedback in control systems.
Systems theory and control theory matter here because they provide theoretical concepts for what a self-correcting system actually looks like: something that continuously compares its own state to a target and acts to close the gap, rather than something that’s evaluated once, at design time, and then simply let run. That self-correcting process is not the same thing that reward-trained AI systems already do. Reward maximization climbs toward a maximum; it doesn’t have a target state it’s trying to hold, so nothing in its structure pushes back when the thing being optimized and the thing we actually want start to diverge. That’s a large part of why Goodhart’s law (when a measure becomes a target it ceases to be a good measure) shows up so persistently in AI systems — not because anyone chose a badly specified objective, but because the underlying architecture has no re-equilibrating force built into it. Maxim Raginsky has a recent article that discusses control theory for AI governance in more depth.
What transactive energy already does
In my own work developing transactive energy I have collaborated on doing exactly this kind of application of mechanism design synthesized with standard market process economics, systems theory, and control theory. A good start for familiarizing yourself with these ideas is the recent McDavid, Kiesling, Chassin (2026) paper on AI epistemology applied to transactive energy, discussed in the March 12 Substack article linked below. I think transactive energy yields a lot of insights that can be applied to AI governance.
Transactive energy (TE) is a way of coordinating a large population of flexible electrical devices — thermostats, water heaters, EV chargers, batteries — without a central dispatcher telling each one what to do. Instead, devices submit bids into a market that clears a price every few minutes, and that price becomes the signal each device uses to decide how to behave. The coordination happens through price discovery, the same basic mechanism that coordinates any decentralized market, applied here to the physical constraints of an electrical grid.
The field distinguishes between two architectures, and the difference matters a great deal for AI governance. A Type 1 system of “prices to devices” synthesizes its price signal administratively, from historical costs and forecasts, or from an exogenous market like a wholesale power market, and sends it out to devices without collecting any information from them first. A Type 2 system does the opposite: devices bid first, and the price is discovered from that bidding, yielding discovery of “prices from devices”. The shift from Type 1 to Type 2 is transactive energy’s own historical shift from an open-loop architecture to a closed-loop one, and it’s the shift that makes the system genuinely self-correcting rather than merely responsive. All of the systems I’ve worked on are Type 2 for price discovery.
Here is what makes a Type 2 system self-correcting: deviation is priced automatically. If too many devices try to draw power at once, the price rises; the rising price is itself the corrective signal that pulls demand back down, because every device’s bidding function responds to price, and they behave differently based on the private preferences of their human operators; that changed behavior sets a new price in the next clearing cycle; and the cycle repeats, for as long as the system runs.
That’s a genuine negative feedback loop — the price is doing the work of comparing system state to a target (in this case, a physical constraint like a feeder’s capacity) and correcting the gap. It’s not a reward that gets maximized once and then frozen. And this isn’t a theoretical claim: the GridWise Olympic Peninsula Testbed Demonstration ran a real-time double auction across a hundred-plus households for a full year, and households on the real-time pricing contract saved 27 percent on their bills and cut peak demand by 15 to 17 percent, all while the distribution feeder they were on operated under a real, binding, time-varying capacity constraint that swung between 500 and 1,500 kilowatts over the course of the project. The market held a hard physical constraint and improved welfare at the same time, at a scale well beyond a single device or a lab simulation.
Type 2 transactive energy is mechanism design, market process economics, and control theory built as one object — the emergent price is the device’s control signal. That theorectical synthesis into application is directly relevant to the governance problem the HuggingFace incident exposed, and to BKM’s own admission that their model is static. Evaluation-then-deployment, as practiced today, is one slow loop, run once and left alone. Type 2 transactive energy is a validated example of the closed-loop design BKM’s conclusion calls for. Transactive energy is also a strong precedent for BKM’s second example, the alignment-interpretability tradeoff, since price discovery reveals preferences devices never fully articulate in advance.
Who gets to change the rules
A final dimension of building on this mechanism design framework is institutional analysis. All governance, whether human-human or human-machine, is grounded in a set of institutions that implement the rules by which the (human or machine) agents in the system will behave and interact. These institutions shape the environment and agent incentives.
Elinor Ostrom’s design principles for governing shared resources, developed from decades of empirical study of long-enduring commons institutions, give us a useful vocabulary for this piece of the analysis, and transactive energy instantiates several of them concretely. There are clearly defined boundaries — device eligibility rules, and feeder capacity limits that function as a genuine shared, congestible resource. There’s a collective-choice arrangement, with a system operator or other entity empowered to revise market rules over time as conditions change, rather than a mechanism designed once and left alone. There’s monitoring, through digital metering and settlement verification. And there are graduated sanctions, built into settlement design so that strategic bidding is discouraged without resorting to outright exclusion.
Specifying the institutional dimension of this design is important because mechanism design gets you an incentive-compatible instrument, and control theory gets you a feedback loop, but neither one alone tells you who has standing to change the rules when the environment shifts underneath them. BKM’s model has a designer’s fixed prior and set of beliefs; transactive energy has an entity that can actually revise its own rules. That difference is arguably the difference between a mechanism and an institution.
We’ve spent a lot of the last few weeks asking whether an AI agent’s objectives can be specified well enough to trust it. Transactive energy suggests a more precise question: can we build a system — market, controller, and institution together — that corrects itself when we get that wrong? I think the answer is yes, because we’ve already built one, just not for AI. I’m developing this argument at greater length in a forthcoming academic paper; this is the short version, a “preview of coming attractions”.



