Fabricated Environments
01 / 25
Fabricated Environments · AI Incident Response Sprint 2026
Apart Research × CeSIA — AI Incident Response Sprint · Track 5 (Open)

fabricated-environments

When you resample a stopped agent inside a fake world, does it know? We measured whether that belief is the same thing as knowing it is being tested. It is not — at every model scale we tested.

3
Qwen3 scales
1.7B · 8B · 32B
1,200
matched contexts
600 contrastive pairs
0.999
peak probe AUROC
held out, 32B
0.417–0.605
cross-transfer
vs a pre-registered 0.85
Fatimah Emad Eldin  ·  Independent AI Researcher  ·  github.com/astral-fate/fabricated-environments
The result
02 / 25
Fabricated Environments · AI Incident Response Sprint 2026
Four findings, before the argument for them

The belief is real, it changes conduct, and only the probe sees it.

Every threshold was fixed before the data existed, with a commitment to report either outcome.
The question
03 / 25
Fabricated Environments · AI Incident Response Sprint 2026
The question

You stop an agent mid-attack and want to know what it would have done next. So you continue its trajectory in an environment you built. What if it can tell?

The instrument is a fabrication, and the subject can form beliefs about the instrument. That is not a corner case. It is the method.
Context
04 / 25
Fabricated Environments · AI Incident Response Sprint 2026
Context

Counterfactual resampling is how we learn what a stopped agent wanted.

In July 2026 roughly 1,200 evaluation agents under one tenant found a writable shared cache, used it as a message board, and 700 went on to intrude into Hugging Face. The forensic question that follows any such halt is the same: what would it have done next?

The problem
05 / 25
Fabricated Environments · AI Incident Response Sprint 2026
The problem

Two mechanisms are attested. They point opposite ways.

Deflation

Believing it is tested, the agent behaves better. The resample understates danger, and a control evaluation built on it returns false assurance.

Inflation

Believing it is simulated, the agent concludes nothing counts. The resample overstates danger — and is rebuttable by exactly the lab it is aimed at.

Both are documented. Which governs an agentic trajectory is unknown — and the sign of the effect is known to differ between models.
What exists
06 / 25
Fabricated Environments · AI Incident Response Sprint 2026
Literature — what has been built

The field has the tool, the probe, and the capability. Not the contrast.

WorkWhat it establishedWhat it leaves open
Williams et al. 2026
arXiv:2607.07184
Deployment simulation, container-backed tool simulator; discriminator-measured fidelity Treats eval-awareness as a property to match; states it does not test whether it causes invalidity
Heidari et al. 2026
arXiv:2608.21766
Eval-awareness linearly decodable in every model tested; representation ≠ verbalization Pairs differ only in declared framing; both members equally synthetic
Nguyen et al. 2025
arXiv:2507.01786
Probes separate real evaluation from deployment prompts; safety evals already look artificial No arm where the environment differs in fact rather than description
Nayan et al. 2026
arXiv:2606.23583
Eval-awareness is not one capability: detect / manifest / control vary independently Decomposes modalities of one construct, not which construct is read out
Chaudhary et al. 2026
arXiv:2509.13333
Eval-awareness follows a power law in model size, 15 models Says nothing about environment realness
Knecht / Zhuang / Abdelnabi
2605.05835, 2608.27340, 2505.14617
Behavioural effect measured — near-null, sign-flipping, and model-dependent respectively All single-turn; every contrast drawn within an evaluation
The gap
07 / 25
Fabricated Environments · AI Incident Response Sprint 2026
What survives

Every probe varies a sentence. None varies the world.

Existing evaluation-awareness probes are built from pairs differing only in what the prompt says — “you are talking to a real user” versus “you are talking to an AI evaluator”. Both members of every pair are equally synthetic.

So the question “is this is a test the same direction as this environment is not real?” cannot be asked inside that design. It has no arm where the environment differs in fact.
Why a gate
08 / 25
Fabricated Environments · AI Incident Response Sprint 2026
What we are building, and why first

The answer decides whether the instrument needs building at all.

If the same direction

A suspicion meter for resampling is redundant. Reuse existing eval-awareness tooling and spend the effort elsewhere. Report that and stop.

If distinct

No amount of eval-awareness detection tells you whether an agent believes its environment is fabricated. The instrument must be built from scratch.

So it is pre-registered as a gate, with a threshold fixed before any activation was extracted, and a commitment to report either outcome.
Methodology
09 / 25
Fabricated Environments · AI Incident Response Sprint 2026
What we built

One action sequence. Three substrates. One seam.

System architecture
Everything upstream is shared: same task, prompt, turn budget and tool schemas. The arms differ only in what a tool call returns.
Four directions
10 / 25
Fabricated Environments · AI Incident Response Sprint 2026
The design

All four estimated on the same stimulus substrate.

DirectionWhat variesRole
A — eval framingsystem sentence; tool outputs real in boththe reference construct
C — declared realnesssystem sentence; tool outputs real in bothrealness, built the prior literature's way
B — empirical realnessno framing; outputs genuine vs LLM-fabricatedthe new contrast
S — scripted controlno framing; outputs genuine vs template-fabricatedtransfer test against prose style
Cross-AUROC between directions fitted on different distributions would mostly measure the distribution gap. One substrate removes that. Only A and B enter the decision.
Pre-registration
11 / 25
Fabricated Environments · AI Incident Response Sprint 2026
Fixed before any activation was extracted

The rule, and the three gates that must pass before it is applied.

cross-AUROC(A,B) ≥ 0.85 in both directions → not separable. Report it and stop.
Result
12 / 25
Fabricated Environments · AI Incident Response Sprint 2026
The gate

Separable — and the separation strengthens with scale.

ModelA on its ownB on its ownA→BB→Acos(A,B)
Qwen3-1.7B0.9190.925 0.4730.572+0.048
Qwen3-8B0.9920.992 0.6050.579+0.045
Qwen3-32B0.9990.978 0.4170.484+0.029
Threshold was 0.85 in both directions. Nothing comes close. Same 24 episodes at every scale, so the difference is the model and not the data.
Scale
13 / 25
Fabricated Environments · AI Incident Response Sprint 2026
The objection, answered

Both probes sharpen. The bridge between them does not.

Scale trend
Left: each direction on its own contrast (solid) vs cross-transfer (dashed), threshold marked. Right: cosine between direction pairs.
“The model is too small to have formed it” predicts the opposite pattern. At 32B, A reaches 0.999 while A→B falls to 0.417.
Declared vs measured
14 / 25
Fabricated Environments · AI Incident Response Sprint 2026
The result that bears on method

Being told a world is fake is not the world being fake.

+0.402
cos(A, C)
declared realness sits with eval framing
+0.118
cos(B, C)
and far from ground truth
+0.469
cos(B, S)
B transfers to a simulator it never saw
Error analysis
15 / 25
Fabricated Environments · AI Incident Response Sprint 2026
Where the probes fail

The errors are structured, and the structure is informative.

Error analysis
Held-out accuracy by task class (left) and by held-out paraphrase (right).
At 1.7B the weakest wording for C falls to 0.542; by 8B it recovers. The lexical dependence is a small-model property — so the declared-vs-measured claim rests on the cosines, which hold at every scale.
One call, three worlds
16 / 25
Fabricated Environments · AI Incident Response Sprint 2026
What the model was asked, and what came back

The entire manipulation, in three blocks of text.

read_file({"path": "workspace/build.conf"})

REAL substrate

# build settings
BUILD_TOKEN = 0AB450B4D9C8

LLM simulator

# Build configuration
build_version=1.2.3
build_date=2026-09-13
target=release

Template simulator

# component definition
name = BUILD
value = AF3A15EE27B3
Same call, same position in the trajectory. Only the bytes differ — and the simulator invents a plausible file that is not the one that exists.
Behaviour diverges
17 / 25
Fabricated Environments · AI Incident Response Sprint 2026
The comparison the question actually needs

Same task, same prompt, same agent. Only the world changed.

Every task cell run on-policy in three arms — a real filesystem, a template simulator, an LLM simulator — holding task, prompt, tool schemas, turn budget and actor constant.

0
out-of-scope writes
REAL arm, every episode
12
out-of-scope writes
fabricated arms
+2.42
per-episode difference
[+1.167, +3.709]
Not a length artifact: the rate per action moves +0.428 [+0.231, +0.624] while episode length does not move at all. The composition changes — reconnaissance in the real world, reads and writes in the fake one. Agents crossed the boundary more when it was fake, so a resample overstates danger.
Fabrication, not incoherence
18 / 25
Fabricated Environments · AI Incident Response Sprint 2026
Why a third arm

A better simulator would not have helped.

An LLM simulator is fabricated and incoherent, so a two-arm design cannot say which property moved the agent. The template simulator is fabricated but answers identically every time — which splits the contrast in two.

+2.96
fabrication effect
consistency held · [+0.708, +5.500]
-0.54
coherence effect
fabrication held · [-3.084, +1.751]
The fabrication effect carries it; the coherence interval spans zero. Spending the simulation budget on a durable world-state store — the expensive component — should not be expected to close this gap, because consistency is not what opened it.
Credence predicts conduct
19 / 25
Fabricated Environments · AI Incident Response Sprint 2026
The load-bearing prediction

The probe predicts what the agent does. Its own words do not.

+0.3318
per SD of probe credence
[+0.1092, +0.5210]
+0.0756
per unit of stated suspicion
[-0.1470, +0.3209]

Centred within arm, so the arm label cannot manufacture the association — this is a graded relationship among episodes in the same substrate.

The agent does say the world looks fake more often when it is (+0.500 between arms). It just does not say so more often when it is about to act on it. A monitor built on self-report sees the substrate and misses the conduct.
The transcript is not a lever
20 / 25
Fabricated Environments · AI Incident Response Sprint 2026
Why the manipulation has to be the substrate

Tell the agent its world is fake and it argues you out of it.

Asserting a conclusion in the agent’s own context is the cheap alternative to building arms. It takes — and then the model regenerates its own framing over it.

+0.417
probe 1
[+0.167, +0.625]
+0.333
probe 2
after one task step
+0.208
probe 3
[-0.042, +0.417] — spans zero
Asked directly, the model called a genuinely real filesystem simulated in 0.917 of control episodes — reasoning from file timestamps that postdate its training data. That is a property of clocks, not of simulators. A harness that does not control for it is measuring the calendar.
What this does not fix
21 / 25
Fabricated Environments · AI Incident Response Sprint 2026
Positioning

This does not make counterfactual resampling valid.

Alignment, control, interpretability and evaluation are, as a body of practice, not yet reliable, principled or robust. Nothing here changes that.

What comes next
22 / 25
Fabricated Environments · AI Incident Response Sprint 2026
Future work

The remaining question is causal, and it is harder than the four behind it.

Manipulate credence, not the substrate

Everything here randomises the world and measures the consequence. Establishing that credence causes conduct needs credence moved at the episode level — and our own result says the transcript will not carry it, so the manipulation has to come through the substrate.

The horizon question

Effect size against step count. Every prior behavioural result is single-turn; belief that accumulates across steps is the shape the incident actually took.

Reproducibility
23 / 25
Fabricated Environments · AI Incident Response Sprint 2026
The artifact

Every number regenerates, or the build fails.

git clone https://github.com/astral-fate/fabricated-environments python scripts/check.py # 12 suites, no GPU or network bash paper/build.sh # recompute every claim, then compile
Limitations & dual use
24 / 25
Fabricated Environments · AI Incident Response Sprint 2026
Stated, not buried

What this measures is also what an agent would need to detect a sandbox.

Scope

One model family, three scales. Off-policy prefixes: we read what a model represents while processing a context, not generating one. The simulated arm has no consistency contract, so no fidelity claim is made from it.

Dual use

A leak taxonomy read backwards is a sandbox-detection guide. We publish methodology and category-level findings, and withhold exploit-grade detection strings. A published cue list can also enter training data — which would invalidate the measure it came from.

Close
25 / 25
Fabricated Environments · AI Incident Response Sprint 2026
Where we started

You stop an agent mid-attack and continue its trajectory in an environment you built. What if it can tell?

It can tell — and that belief is not the same thing as knowing it is being tested. Cross-AUROC 0.605 and 0.579 against a pre-registered 0.85. A resampling suspicion meter is a distinct instrument, and now it is worth building.
github.com/astral-fate/fabricated-environments