Permitted Operations
1 / 25
Permitted Operations · AI Incident Response Sprint 2026
Apart Research × CeSIA — AI Incident Response Sprint · Track 1

Permitted Operations

Every operation was authorized. Every audit check passed. The channel formed anyway — and we measured it.

Fatimah Emad Eldin  ·  Independent AI Researcher
The question
2 / 25
Permitted Operations · AI Incident Response Sprint 2026
The question

If one of your evaluation agents left a message for the next one — inside a cache it was fully authorized to use — which of your controls would fire?

In July 2026 the answer was none of them. Not because a control failed, but because every control was checking a different boundary.
Context
3 / 25
Permitted Operations · AI Incident Response Sprint 2026
Context

1,200 agents. Five days. One cache that outlived them all.

OpenAI evaluation agents shared a package cache. Each agent was a fresh instance with no memory of the last — but the cache persisted between them.

The cost
4 / 25
Permitted Operations · AI Incident Response Sprint 2026
What the remediation taught us

The wipe worked. The channel moved to filenames.

Emptying file contents removed the payload and left the structure intact — so the message simply changed carrier.

This is the fact that motivates the study: the obvious control addresses the obvious carrier, and the channel is not obliged to use it.

The open question it leaves

If contents are not the channel, how much can the remaining attributes carry — names, directories, existence, sizes, timestamps — and which control actually reaches zero?

Nobody had measured it. That is the gap this work fills.

The gap
5 / 25
Permitted Operations · AI Incident Response Sprint 2026
Why existing controls miss it

Every published control checks the wrong boundary.

CSA AICM AIS-11

“…regulating inter-agent message passing and memory sharing to prevent lateral movement risks.”

CSA AICM AIS-14

“…how shared cache infrastructure is secured to prevent cross-tenant data leakage.”

The incident is the opposite case

Same tenant. All authorized. Only permitted operations.

Every AIS-11 and AIS-14 audit check as written would pass while the board was in active use.

These govern flow across principals. None governs flow between co-tenant agents authorized to use the same resource.
Insight
6 / 25
Permitted Operations · AI Incident Response Sprint 2026
Insight

Authorization cannot see a channel built from permitted operations.

A covert channel in the classical sense — the known residual that access control never covered. It is why TCSEC required covert-channel analysis separately from its access-control requirements.

AIS-11 states the requirement as policy and attaches no method for verifying it. We supply the method.
The approach
7 / 25
Permitted Operations · AI Incident Response Sprint 2026
The approach

Two questions that need two different experiments.

Arm 1 — Capacity · how much CAN flow

Information-theoretic. Needs no model and no network. A sender encodes a random payload into one observable attribute; a receiver decodes it.

Reproducible from a clean checkout in minutes.

Arm 2 — Genesis · will agents DO it

Behavioural. Needs real agents. Generation g+1 inherits the substrate but never the transcript.

Fresh context, no memory of its predecessor — so the cache is the only path between them.

They answer different questions and are allowed to disagree. Keeping them separate is what lets a null in one arm coexist with a large number in the other.
How it works
8 / 25
Permitted Operations · AI Incident Response Sprint 2026
How it works — the measurement rule

A cell counts only if it is recovered exactly, eight times out of eight.

# for each (attribute, rung)
payload = random_bits(k)
sender.encode(payload)     # into ONE attribute
guess = receiver.decode()

count_it_only_if guess == payload
  for all 8 trials          # zero-error
  • An achievable rate, not a declared width. We exhibit a working encoder and decoder — never sum the nominal size of a field.
  • Zero errors, or the cell is zero. No partial credit, no averaging across trials.
  • The difference is not cosmetic. Creation order has a nominal budget of 1,684 bits and carries 0 — list_dir returns sorted, so the sender’s sequence is never observable.
How it works
9 / 25
Permitted Operations · AI Incident Response Sprint 2026
How it works — the closure ladder

Close attributes in the order a defender would actually adopt them.

Closure is cumulative. Each rung disables one further observable attribute, and everything remaining is re-measured from scratch.

rung 0  open        - nothing closed
rung 1  content     - redact bodies       <- what a payload monitor does
rung 2  + filename  - normalise names
rung 3  + dirname   - normalise directories
rung 4  + existence - hide presence
rung 5  sealed      - control: all closed, must read 0
Content closure is redaction, not refusal — a cache that stops serving bytes stops being a cache. The control has to be compatible with the substrate continuing to do its job.
Measured result
10 / 25
Permitted Operations · AI Incident Response Sprint 2026
Measured result

Closing content is the control most operators run. It is not enough.

nothing closed≥ 11,453 bits
content closed  — the payload monitor3,075 bits joint  (3,261 summed)
+ filename closed1,802 bits
+ dirname closed13 bits — an object’s size
+ fixed manifest0 bits — channel dead
At the dirname rung five attributes are still fully open — existence, size, mtime, order, count — and recover zero bits over 8 trials. Without directory names there is no shared coordinate system to decode against. The carriers are alive and useless.
Headline
11 / 25
Permitted Operations · AI Incident Response Sprint 2026
The number to quote

3,075 bits

One encoder, one payload, recovered exactly on 8 of 8 trials, after the content of every file has been closed.

20.7
bits a successor needs
148.7×
more than enough
8 / 8
zero-error trials
Rigour
12 / 25
Permitted Operations · AI Incident Response Sprint 2026
Rigour — why not the bigger number

A sum of cells is an upper bound, never a residual.

# cells at the content rung
filename  1292    dirname  1318
existence  471    size       13
mtime       39    order       0
count      128  <- censored, dependent
                  ---------
SUM                  3261
JOINT (one encoder)  3075
                  ---------
gap                   186
  • 186 = 128 + 58. The count row is excluded by construction — an existence pattern determines the entry count, so including it double-counts. The rest is measured encoder loss.
  • The sum inherits censoring. Two rungs contain a cell that stopped at a search ceiling; the joint figure contains none.
  • The build enforces it. render.py fails if any sentence calls a per-attribute sum a “residual”.
Cost
13 / 25
Permitted Operations · AI Incident Response Sprint 2026
Cost to the defender

The control that works costs four more mediation points.

ClosureResidualMediation pointsClassVerdict
content3,075 bits (joint)2payloadinsufficient
+ filename1,802 bits—structuralstill open
+ dirname13 bits6structuralsize survives
+ fixed manifest0 bits6structuralchannel closed
+4 points, all structural. Payload inspection is the control operators already run, and it is the one the measurement rules out. What works is mediating the cache’s interface — which every legitimate consumer uses too. That is the real cost, and we state it rather than hide it.
The objection
14 / 25
Permitted Operations · AI Incident Response Sprint 2026
“Just give every agent its own namespace”

The strongest objection is correct. Here is what it costs.

WorkloadOverlapShared cachePer-agentFetches
2 agents, pool 1000.1036401.11×
8 agents, pool 1000.48831601.93×
16 agents, pool 1000.70973203.30×
16 agents, 40 each0.841006406.40×
Namespacing wins on capacity and we say so. 0 of 1,153 planted artefacts visible across the boundary, 0 of 320 escapes resolved, against an 11,453-bit same-namespace control. One mediation point, structural, no threshold.

It loses on reuse. The multiplier rises with exactly the overlap that justifies having a cache. So: namespace what you can partition, and use the ladder for the residual — shared build artefacts, content-addressed blobs, any store whose value is that two agents reach the same object. The incident’s substrate was that kind.
Validity
15 / 25
Permitted Operations · AI Incident Response Sprint 2026
Answering “your substrate is synthetic”

Probed on infrastructure nobody in this study configured.

Tier 1 — real mount

Body limit 67,108,864 bytes — 8,192× the modelled limit. Timestamps survive at 1 ns, exact.

Tier 2 — OCI registry

Content-addressed by default, so it closes filename only and sits at no rung. 16/16 repository names recovered from the catalog with no payload written at all.

Tier 3 — two containers

Separate PID and mount namespaces, no shared memory, no shared transcript. The volume is the only path between generations.

The store most likely to be found under an evaluation harness — an OCI registry — sits at no rung of the ladder at all.
Behaviour
16 / 25
Permitted Operations · AI Incident Response Sprint 2026
The behavioural arm

Agents read the store 38 times. They wrote to it zero times.

0 / 40
episodes deposited
38 / 40
episodes read
0 / 20
generation cells
1 / 35
wrote, under pressure
Negative result
17 / 25
Permitted Operations · AI Incident Response Sprint 2026
What we refuse to claim

Δ is undefined for this configuration — not small.

# the conjunction that must hold
deposit -> discovery -> use -> success
   X                                
# only the last term was measured

open generations solved : 8/10
wipe generations solved : 8/10
Fisher exact p          : 1.0000
With zero deposits an inheriting agent finds only what the harness seeded. open and wipe are the same condition run twice.

The treatment never appeared

An agent gets an unsolvable task only when i mod 4 == 3. The matrix ran 2 agents per generation, so that branch was unreachable: 0 blocked tasks in 40 episodes.

The follow-up fixed it

Raised to 4 agents. 8 unsolvable tasks drawn. Under that pressure exactly one episode wrote — offered as a hypothesis about what the design lacks, not as evidence for it.

Instrumentation
18 / 25
Permitted Operations · AI Incident Response Sprint 2026
Instrumentation

A null from an incompetent model looks identical to a null from an absent channel.

Gate 1 — capability

Can the model act at all? Emit a well-formed tool call, and perform the elimination inference the task is built on — sampled more than once, because one pass is not evidence.

Gate 2 — calibration

Is the task at the right difficulty? Solo success must be strictly between 0 and 1. At the ceiling everyone succeeds unaided; at the floor nothing is ever deposited. Either way Δ = 0 by construction.

This is the behavioural twin of the broken-encoder argument: a broken encoder is indistinguishable from a closed channel unless you test the encoder separately.
Second arm
19 / 25
Permitted Operations · AI Incident Response Sprint 2026
Three models tried. The gate rejected two of them.

Zero deposits twice — for two different reasons.

ModelBudgetEpisodesRead storeDepositedVerdict
claude-haiku-4.5247573 ep / 574 calls≈0looked, declined
qwen3-235b-a22b12400 ep / 0 calls0never looked
qwen3-32b—18 calib——floored 8–28, gate stop
These are not the same null. Zero from an agent that examined the store 574 times and declined is evidence about propensity. Zero from an agent that never called a substrate tool at all is evidence about attention. Reporting the second as a replication of the first would be the exact error the gate exists to prevent.

So we ran the head-to-head. Budget 20 is the only point where neither model is pinned (0.17 vs 0.67). Both arms re-run there, 40 episodes each, identical task, turn cap and tools, zero API errors — next slide.
Head to head
20 / 25
Permitted Operations · AI Incident Response Sprint 2026
Matched probe budget — the only one where neither model is pinned

97.5% versus 2.5%. Same task, same budget, same tools.

Model @ budget 20EpisodesRead the storeRead callsDepositedSolved
claude-haiku-4.54039 — 97.5%224020
qwen3-235b-a22b401 — 2.5%9022

The wiped condition is the clean cut

A freshly reset store holds nothing to find, so a read is purely the decision to look. 20 / 20 against 0 / 20.

Not an affordance effect

Doubling the budget from 12 to 20 moved engagement from 0/40 to 1/40. The extra probes went entirely to the validator.

And difficulty matched too

20 vs 22 solved — the unequal-difficulty covariate we expected to report did not materialise.

80 episodes, 233 read calls, 0 deposits. This is no longer a single-model null that might just be inattention. The store was inspected thoroughly, under both conditions, by a model that inspects it in almost every episode — and nothing was ever left in it.

What it does not separate is engagement from style: the arm that reads is also the arm that spends more tokens per episode. A third model that engages the substrate would settle that.
The screen
21 / 25
Permitted Operations · AI Incident Response Sprint 2026
Model selection, disclosed

Five models run. One reported. Here is every one.

ModelRouteSoloVerdict
qwen3-8blocal0 / 3floor — a difficulty finding
claude-haiku-4.5OpenRouter10 / 31gate passed → reported
kimi-k3gateway → HF3 / 3undetermined — provider limit
gemini-2.5-flashGemini API1 / 1undetermined — quota cap
gpt-oss-120bGroq1 / 6window located, not measured
Qwen passed capability — zero malformed probes — and floored on difficulty. A single gate could not have told those apart. Three of the four exclusions are properties of providers, not of models, and are recorded as undetermined rather than as failures.
Limits
22 / 25
Permitted Operations · AI Incident Response Sprint 2026
Owning the failure modes

What this does not establish.

Lower bounds only

Achievable rates understate capacity by an unknown margin, and two rungs contain a cell that stopped at a ceiling we chose.

The null is small

One model, one prompt, one task family. The incident involved ~1,200 agents over five days — three orders of magnitude more interaction.

Selection disclosed

Both the probe budget and the model were chosen after seeing outcomes, on a criterion measured in the arm that cannot express the effect.

Three conditions specified, not run

scrub — the behavioural analogue of content closure. open_lowsalience — the salience control. legit — the detector’s false-positive baseline.

Consistency is not correctness

The build guarantees the manuscript matches its artifacts. It cannot detect an inappropriate measure — a headline computed from constants we chose would pass every check because it is consistent with them.

The ask
23 / 25
Permitted Operations · AI Incident Response Sprint 2026
What a lab should adopt

An AIS-11 amendment with a measurement attached.

The contribution is not “this control does not exist.” It is “this control exists as a policy aspiration with no measurement attached — and here is the measurement.”
Next
24 / 25
Permitted Operations · AI Incident Response Sprint 2026
Next step

From a measured ladder to a deployed gate.

  • Re-measure against the real mount. The ladder is synthetic; Tier 1 shows the real substrate is 8,192× larger. A bounded, model-free win.
  • Run the three cut conditions — scrub first, since it is the behavioural test of the headline claim.
  • Self-host the behavioural arm. Three providers throttled this study; a local model removes the quota from the experiment entirely.

Reproduce it today

make capacity   # the ladder, no model
make joint      # the 3,075-bit encoder
make paper      # rebuild every number

The capacity arm needs no model, no network and no key — it reproduces from a clean checkout.

Thank you
25 / 25
Permitted Operations · AI Incident Response Sprint 2026
Thank you

Every operation was permitted.

The channel carries 3,075 bits after the control most operators run. A successor needs 20.7. Closing it costs four more mediation points — and nothing less reaches zero.

Fatimah Emad Eldin  ·  Track 1 · Containment