Spectrum SharingGame TheoryLaw and Economics

Spectrum Jails: Engineering Trust After the Fact

Cognitive radios promised flexibility that regulators could not certify line by line. Drawing on law and economics, this work asked whether identity, liability, and future sanctions could replace trust in a radio's internal design, then followed the technical tradeoffs that appear when deterrence, rights, commitment, and reputation must operate through real observations.

In this story

A regulatory problem, not a clean-sheet puzzle

What was already understood

A cognitive radio was supposed to adapt its waveform and spectrum use in software. But every new freedom was also another behavior that conventional certification would have to anticipate. The practical question was how to permit innovation without asking a regulator to bless every future line of radio code.

Traditional radio regulation relied heavily on ex-ante trust. A device was tested against a prescribed behavior before sale, then presumed to keep behaving that way. That bargain becomes brittle when the point of the device is to learn, negotiate, and reconfigure itself after deployment. Human legal enforcement sits at the other extreme: investigate harm, identify the responsible party, and impose a sanction afterward. It is flexible, but far too slow and expensive for machines making spectrum decisions in milliseconds.

There was also a structural worry, and it is the one that pulled us in hardest. When the only way to permit a new use is to write that use into the rules beforehand, the rulemaking becomes the scarce resource, and being well represented in that proceeding starts to matter as much as being right about the radios. Rules that turn on measurable behavior rather than on approved designs leave less room for lobbying to stand in for a technical tradeoff.

We began by engaging the law-and-economics view of deterrence, liability, and property rights, then asking what would have to become technical for those ideas to govern radios directly. The translation exposed gaps that verbal policy could glide past. Who can identify an offending transmitter from the harm it causes? What can a regulator credibly take away? How should false accusations count? What if the protected user is strategic too?

Those were not decorative policy questions around a finished wireless system. Their answers changed the signals radios had to emit, the resources they had to hold in reserve, and the performance that sharing could deliver. Regulation itself became an engineering object.

This story sits differently from the others here. Interpolation Can Be Harmless and the Witsenhausen story both begin in curiosity — a puzzle about overfitting, a counterexample that would not behave. This one begins at the other end, with a policy failure mode we wanted to avoid, and the curiosity arrives later. Donald Stokes would have called it use-inspired basic research: the use came first and set the questions, and the questions turned out to be fundamental anyway. What follows runs from a spectrum-rulemaking worry to a theorem about no-regret learning, and the second half would not have happened without the first.

When peers cannot police one another

A genuine inspiration for us was the work of Etkin, Parekh, and Tse on unlicensed spectrum sharing. In a repeated game among peers, the prospect of future retaliation can support efficient sharing without a central authority: cooperate today because a deviation can be punished tomorrow.

But opportunistic sharing is often deliberately unequal. A primary user may be badly harmed by a secondary transmitter while being physically unable to retaliate in kind. The weaker user has, as the light-handed-regulation paper put it, neither defense nor ammunition. Repetition does not create trust when only one side has a credible punishment. Some external vulnerability has to be engineered.

That led us to the spectrum jail. A radio remains free to choose its substantive behavior. The certified core is much smaller: the radio must be identifiable, and it must obey an authenticated command that temporarily removes its spectrum privileges. Compliance is encouraged not by proving in advance that every action is legal, but by making misconduct detectable and the loss of future access credible.

The regulatory turn

Earlier networking work had used repeated-game punishment among peers, and trusted-radio architectures had included technical sanctions as a backstop. What we made the organizing principle was automated ex-post deterrence: certify catchability and obedience to punishment, not every permissible behavior. It was not certification-free; it changed what little had to be certified.

Identity is a communications problem

Deterrence begins with attribution. In radio, identity cannot be only a label in a database: the signal carrying the label may be exactly what an agile transmitter is free to alter. A separate identity beacon also proves presence, not causation. It can implicate every nearby transmitter while leaving the victim with the burden of decoding another waveform.

The obvious information-theoretic route was to embed the identity in the waveform itself. Dirty-paper coding — how much message can be carried without disturbing too much of what is already there — was very much in the air, and the cognitive-radio-channel work of Devroye, Mitran, and Tarokh had built capacity regions on it. Two things made it the wrong tool for attribution. Prescribing a waveform and a message to carry constrains the radio designer exactly where the point of the exercise was to leave them free; and even a perfectly embedded identifier still only proves that a device is present, not that this transmitter caused that harm.

Our first instinct was gentler than the one we ended up with, and it came from a MAC protocol. Rozovsky and Kumar's SEEDEX gave nodes pseudorandomly generated slot schedules so they could avoid colliding with one another — randomness that costs essentially nothing, since a radio has to pick some schedule anyway. If identity could ride on the same trick, attribution would be free. It cannot, for a reason worth naming precisely: whoever observes the harmful interference cannot see the context in which the choice was made, and so cannot recompute what an innocent radio would have done. Modulated randomness is only evidence to someone who can check it.

So we gave up on harmlessness. Our alternative was to stop modulating choices and start forbidding them, encoding identity in silence. Each radio receives a sparse pattern of taboo time-frequency slots in which its transmitter is physically disabled. If harmful interference repeatedly disappears in one radio's taboo slots and returns when that radio is free to transmit, those slots act like experimental controls. Hierarchical codes can combine network, user, and device identities, and different bands can use different patterns. Because different radios fall silent at different times, an observer learns which suspects were provably absent while the harm persisted and which were free to transmit whenever it appeared — group testing, run on the spectrum. We cite SEEDEX in the enforcement paper for exactly this reason: codes over who may transmit when are useful for more than politeness.

Once identity costs an honest radio real opportunity, the question stops being whether attribution is possible and becomes how much it has to hurt. That reframing made accountability a hypothesis-testing problem with a price attached. A wider taboo pattern makes guilt easier to establish but consumes more opportunities even for an honest radio. Catching weaker interferers or conspiracies takes longer or costs more overhead. Robust identity is possible, but it is not free.

A hierarchy of network, user, and device identifiers produces different sparse patterns of forbidden transmission slots in three spectrum bands.

A radio identifies itself through the opportunities it is forbidden to use. The taboo slots require only a certifiable transmitter veto, yet give observers controlled periods with which to test whether that radio caused the harm.

Key Insight

The identity signal is the pattern of missing transmissions. Silence becomes evidence because the device can be cheaply certified not to transmit in its own taboo slots.

A brief note for readers arriving from machine learning: watermarking a language model's output follows the same instinct, and it gets to keep the nearly harmless version we could not have. In Aaronson's scheme, a secret key gently biases the randomness behind low-stakes token choices; Anthropic's watermark for supported Claude models pursues the same broad goal, with little apparent effect on meaning or quality. This works because the detector can see the context in which those choices were made and test for the corresponding statistical signature. An observer of radio interference cannot. That single difference in what the detector knows is why our version of identity has to be paid for in forgone opportunity, and why the interesting question became how little we could pay.

Future access becomes collateral

Identification matters only if something valuable can be taken away. The jail mechanism turns future spectrum access into collateral. A secondary may exploit openings in a primary band, but if it is caught transmitting illegally, an authenticated command excludes it for a finite sentence. A purchased or otherwise dependable home band gives honest operation real value; losing access to that resource makes the threat consequential.

The Markov model we built connects primary activity, the secondary's legal and illegal choices, imperfect detection, and time in jail. It makes the deterrence inequality explicit: the immediate gain from cheating must be outweighed by the chance of conviction times the future utility lost during punishment.

Making that inequality technical produced a less obvious result of our own. Improving the probability of catching a guilty radio helps, but the rate of wrongful conviction can matter even more. Innocent radios also lose valuable service while jailed. If that overhead is too large, the enforcement system can destroy the benefit of opening the spectrum in the first place. The wireless version of due process is therefore a capacity issue, not just a moral analogy.

Two ways into jail

Deterrence depends on the whole cycle, including the path an honest radio can be pushed down.

OUT OF JAIL
a transmission decision
the secondary chooses
OBEY
Transmit only when allowed
wrongly convicted: PwrongP_{wrong}
pure overhead on an honest radio
CHEAT
Transmit while primary active
caught: PcatchP_{catch}
deterrence working as intended
JAIL
Access revoked, including
the home band worth β\beta
leaves w.p. 1 − PpenP_{pen}
sentence served, back to the cycle

Cheating is deterred only when the stake outweighs the temptation

β>B1Ppen+PwrongPcatchPwrong\beta > B\,\dfrac{1 - P_{pen} + P_{wrong}}{P_{catch} - P_{wrong}}
β\beta
value of the home band staked
BB
number of expansion bands
When PwrongP_{wrong} approaches PcatchP_{catch}, deterrence collapses.
If jail comes at nearly the same rate either way, the radio may as well cheat and collect the utility.

Both paths matter. Cheating while the primary is active leads to jail through PcatchP_{catch}, which is deterrence working as designed; but obeying the sharing rule can lead there too, through the wrongful-conviction rate PwrongP_{wrong}, which is pure overhead on an honest radio. Deterrence needs the staked home band β\beta to outweigh the temptation across all BB expansion bands — and it collapses entirely once PwrongP_{wrong} approaches PcatchP_{catch}.

Drawn for this page from the model and results of Crime and Punishment for Cognitive Radios (2008).

Key Insight

Trust does not require confidence in all of a radio's code. It can come from a narrower, verifiable vulnerability: the radio has something to lose, can be caught with useful probability, and cannot ignore the sanction.

Punishment must reach the utility that matters

A jail sentence punishes a throughput-hungry radio because silence delays its work. But a battery-limited device may care more about energy than elapsed time. We made that distinction operational in the caged-radio paper by adding a singing sanction: while jailed, the radio is forced to burn energy. Combining silence and singing can deter devices with different mixtures of time and energy costs.

The same principle appears in our sensing work. A rule cannot simply demand a detector with a nominal specification and assume the radio will expend the effort to use it. The fine, the probability of being caught, and the value of access together determine the sensing behavior the device finds worthwhile. Regulation has to act on the device's actual objective.

These variations also marked limits. When the primary is active too rarely, cheating may be almost impossible to deter because there are too few occasions on which misconduct can be exposed. No choice of sentence repairs a missing evidentiary channel. Light-handed enforcement enlarges the design space, but it does not repeal observability.

Sanctions are system components

We moved punishment from a policy word into the utility model. Silence prices delay; singing prices energy; a home band creates lawful outside value; sensing effort responds to the expected sanction. Each mechanism works only for the incentives and observations it actually touches.

The victim can cry wolf

Once the primary participates in enforcement, it stops being merely a passive victim. A primary can report interference when none occurred, or transmit useless gibberish to make the band appear unavailable. If a complaint is free and reliably removes the secondary, even a well-protected primary may rationally cry wolf.

Our first response was to make reporting costly enough that the primary would complain about real harm but not casually expel a compatible secondary. The later ex-post-enforcement paper pushed us into a question about spectrum rights. The regulator can choose parameters that protect a primary from excessive harmful interference across secondary types. It cannot simultaneously promise every possible secondary that a primary will never accuse it falsely. A secondary can trust the primary only under compatibility conditions, including that the sharing opportunity is valuable enough to discipline both sides.

This is a different notion of a right from a line in a license. It is an equilibrium guarantee produced by the enforcement mechanism, the available evidence, and the strategic alternatives of both users.

Scope and caveats

Ex-post enforcement can robustly limit the harm suffered by a primary. It cannot give a universal no-false-accusation guarantee to every secondary. The asymmetry is a proved limitation of the mechanism, not a parameter-tuning failure.

Commitment has bits

The regulatory game also depends on who moves first and what the other side can believe. A primary that can commit to a strategy before the secondary responds can obtain a Stackelberg first-mover advantage. But real commitment is rarely a perfectly observed probability distribution. It is a contract, policy, public record, or behavioral history with finite resolution.

This is where the curiosity took over. We introduced partial commitment to make that gap explicit. A finite number of commitment bits restricts the primary to a range of strategies rather than a single exact mixture. More bits produce a finer promise and greater advantage; ideal Stackelberg commitment appears only as the limiting case.

This converts institutional strength into an information quantity. The question is no longer just whether an actor can commit, but how precisely the commitment can be expressed and verified. That naturally leads from regulatory games to reputation learned from finite behavior.

Key Insight

Commitment is not a yes-or-no power. It is an information resource, and finite commitment can be measured in bits.

The ideal promise can be the fragile one

Suppose a follower sees only a finite sample of the leader's past actions. An ideal mixed Stackelberg commitment often lies on a boundary where the follower is exactly indifferent between responses. With infinite precision, a favorable tie-breaking rule gives the leader the desired response. With finite observations, ordinary sampling fluctuations put the empirical mixture on either side of that boundary with nonvanishing probability. The ideal commitment can therefore be bad precisely because it is idealized.

We imported a geometric idea from interior-point methods. Move the commitment slightly into the interior of the desired best-response region. The leader sacrifices a little nominal payoff but creates a buffer large enough that empirical fluctuations usually leave the follower's response unchanged. The best offset depends on the number of observations and on the local geometry of the response regions.

By this point the spectrum story had become a general game-theoretic one, and we were no longer asking a question about radios at all. Partial reputation is not a noisy footnote to commitment. It changes which commitment should be made.

A triangular mixed-strategy simplex is partitioned into colored follower best-response regions, with the ideal Stackelberg commitment located at a fragile vertex of the desired red region.

The ideal Stackelberg commitment sits at an extreme point of the desired best-response region. Finite-sample fluctuations can easily cross into a region with a different response; a robust commitment moves inward to purchase a margin against observation error.

Tight crop of Figure 7 from Robust Commitments and Partial Reputation (2019).

No regret is not the same as settling down

Repeated games often use no-regret learning as a route toward equilibrium. That statement usually concerns averages: over time, the player does almost as well as the best fixed action in hindsight, and empirical frequencies may approach equilibrium behavior. It does not say that the strategy used today converges.

The distinction becomes sharp when learners update from the opponent's realized actions rather than from an exactly revealed mixed strategy. For competitive two-action games we proved that broad classes of algorithms achieving the optimal no-regret rate cannot have their last-iterate mixed strategies converge almost surely to the mixed Nash equilibrium. The randomness needed to realize a mixed strategy keeps entering the learning dynamics as fresh evidence.

This matters whenever a mechanism is judged by its current behavior rather than by a long-run average. Low regret can certify excellent historical performance while the actual strategy continues to wander. The theorem is deliberately scoped: it covers broad mean-based, monotone families and extensions, not every conceivable learning rule. But it identifies realized-action stochasticity as a structural obstacle that deterministic mixture-level analyses can miss.

Two plots compare learning in matching pennies. With exact opponent mixtures, an optimistic method settles at one half; with sampled actions, all shown mixed strategies continue to fluctuate broadly around one half.

The opponent's mixture and its random realization are not interchangeable. In panel (a) the players see each other's exact mixtures and optimistic multiplicative weights sits at equilibrium; in panel (b) they see only realized actions, and the fluctuations are still going after one hundred million rounds.

Tight crop of Figure 3 from On the Impossibility of Convergence of Mixed Strategies with Optimal No-Regret Learning (2024 journal version; arXiv preprint).

Key Insight

A mixed strategy is a distribution; play reveals only samples from it. When those samples drive the next update, convergence of averages can coexist with perpetual motion in the strategy actually being used.

A broader engineering view of institutions

Two smaller branches show that the habit had become broader than spectrum jails for us. Spectrum zoning treated a band plan as a slow institutional choice made before future preferences are known. Robust optimization showed why the best long-lived plan need not sit on today's Pareto frontier: some immediate efficiency may be worth trading for useful choices later, but only when flexibility is not consumed by its own technical overhead.

We applied the same style to an academic institution in the peer-review paper. A public score can reward researchers who review promptly by prioritizing their own papers, but the score also leaks clues about anonymous referees. Deliberately distorting the public score creates a tunable tradeoff between timely participation and anonymity. The setting is lighter; the method is the same: write down the incentives, information, and delay, then see which institutional promises are technically compatible.

What this line changed

Spectrum jails began from a use-inspired constraint: software-defined radios needed freedom, but a regulator could not certify an open-ended future. Our answer was not to abandon trust. It was to relocate trust into a few enforceable primitives — identity, evidence, authenticated sanctions, and something valuable to lose.

Making those primitives technical changed the questions the field could ask. Identity acquired a measurable spectrum cost. Wrongful punishment became a system overhead. A primary's protection and a secondary's security became distinct equilibrium guarantees. Commitment acquired finite resolution, reputation acquired sampling error, and no regret separated from last-iterate stability.

The larger contribution is a way of doing regulatory engineering. Law and economics supplied the motivating language, but the radio model forced every promise to pass through observations, utilities, and feasible mechanisms. The conceptual gaps that appeared in that translation were not defects of the exercise. They were where the research was — which is why a worry about spectrum rulemaking ends, several papers later, in a theorem about what no-regret learning cannot promise.

The papers

Eight of the fourteen papers collected under this story are shown below; the "Explore this story's papers" link opens the full set. The collection follows the main spectrum-jail spine from identity and deterrence through sensing, energy-aware sanctions, strategic primaries, spectrum rights, and commitment. It then includes the later work on robust reputation and no-regret learning, together with the short coda on spectrum zoning and peer-review incentives.

Related Papers
2008IEEE Symposium on Dynamic Spectrum Access Networks
2008Allerton Conference on Communication, Control, and Computing
2009IEEE Communications Magazine
2010IEEE Symposium on Dynamic Spectrum Access Networks
2011IEEE Symposium on Dynamic Spectrum Access Networks
2012IEEE Symposium on Dynamic Spectrum Access Networks
2014IEEE International Symposium on Dynamic Spectrum Access Networks
2017IEEE International Symposium on Information Theory