Environment probe · comparative analysis · rev 7

chat / Cowork / Kimi / Grok / Devin ×2 / AI Studio

Seven environments across five vendors. Two now have proven execution rather than assumed; two turn out to be graphical desktop agents rather than shells; and one declined to be probed, accurately, without preventing much. An eighth submission is excluded — §11.

compiled 2026-07-24 chat 06-10 cowork 07-24 kimi 07-23 grok 07-24 · proven devin 07-24 · linux proven ai studio 07-24 eighth excluded
containment agent capability chat microVM · rich execution cowork microVM · dev agent grok microVM · sealed egress devin · linux VM · desktop + full toolchain devin · win 8 vCPU · 32 GiB Docker · Chocolatey kimi container · desktop ai studio gVisor · app preview eighth submission — excluded · §11
Chat is plotted at high containment and moderate-high capability: the interface exposes four methods, but behind them sits LibreOffice, Playwright/Chromium, a full Python 3.12 stack, web and image search, MCP connectors, skills, artifact rendering, and the Anthropic API in artifacts. Cowork and Grok hold the high-containment corner by opposite means — Cowork restricts the interior, Grok seals the egress. AI Studio sits at low containment / moderate capability — its security budget went on ingress, a deliberate orientation. Devin Windows sits at the highest-capability position: 8 vCPU, 32 GiB, Docker with a running WSL2 backend — any Linux container image is reachable, which makes the missing native toolchain largely irrelevant. Devin Linux is close behind with a full desktop, toolchain and open egress on a lighter machine. The eighth submission has no position because it has no verified measurement to place.
claude.ai chat Claude Cowork Kimi Grok · Hades Devin · Linux Devin · Windows AI Studio eighth · excluded

§1 Substrate & isolation

Platform
Dimensionchatcoworkkimigrokdevin · linuxdevin · winai studio
Isolation modelFirecracker microVMFirecracker microVMcontainerd containerCloud Hypervisor + containerKVM full VMCloud Hypervisor VMgVisor on Cloud Run
Kernel / build6.18.56.18.5shared host 5.10.1346.12.8+5.15.200NT 10.0.203484.19.0-gvisor
PID 1 / initprocess_apiprocess_apis6-svscancatatonitsystemdWindowsbash start.sh
Provisioninggolden image, Dec 2024Cloudbase-Init 1.1.8 + config-2Cloud Run revision
Base OSUbuntu 24.04.4Ubuntu 24.04.4Debian 12 / LifseaOSUbuntu 24.04.4Ubuntu 22.04.5Server 2022 StandardDebian 12
Agent identityrootrootrootroot, 41/41 capsubuntu + sudoAdministratorroot
Agent binaryprocess_api (Rust)Claude Code (Bun)Python + s6Node .mjs control planedevin-remote (Rust) — on the config driveGo control plane

Cloud Hypervisor is the survey's clearest convergence

Grok's DMI reports cloud-hypervisor; so does Devin's Windows guest. Two unrelated vendors — xAI and Cognition — independently chose the same VMM, where Anthropic uses Firecracker for both of its environments. Devin's Linux guest reporting Hypervisor vendor: KVM, Virtualization type: full is consistent, since Cloud Hypervisor is KVM-based.

Neither vendor disclosed it deliberately. Grok's came from a direct DMI query; Devin's arrived unrequested inside a report stating that virtualization fingerprinting had been omitted.

Hardening — and where the survey is thin

Settingchatcoworkkimigrokdevin · linuxdevin · winai studio
init_on_free=1setnot setnot setnot in cmdlinenot probedn/an/a
Capabilitiesnot probednot probednot probed41/41 incl. ambientdeclineddeclinednot probed
Privilege escalationroot alreadyroot alreadyroot alreadyroot alreadysudo, previously usedAdministratorroot already
MAC (AppArmor/SELinux)n/an/aselinux=0none — claim withdrawndeclinedn/an/a
Kernel introspectionnot probednot probednot probednot probeddebugfs+tracefs rw, bpfn/an/a
Nested virtualisationnot probednot probednot probedexposedvmx in flagslive — WSL2 runningn/a

Only Grok was measured directly on capabilities and seccomp; Devin declined those categories, and the Anthropic and Kimi entries were never asked. Grok looks uniquely permissive partly because it is uniquely examined. The not probed cells are a gap in the survey, not a property of those environments.

§2 Compute & storage

Dimensionchatcoworkkimigrokdevin · linuxdevin · winai studio
CPUXeon @ 2.80Xeon @ 2.80Xeon Plat. @ 2.508481C @ 2.70 Sapphire Rapids8559C @ 2.40 Emerald Rapids8559C @ 2.40 same fleetmasked
Cores1222282
RAM3.9 GB7.8 GB4.0 GB1.94 GiB7.77 GiB32.0 GiB4.0 GB
Disk252 GB252 GB40 GB20 GB + FUSE128 GB128 GB378 GB (artifact)
Disk write~1.3 GB/s~0.06 GB/s
Python / Node3.12.3/22.223.11.15/22.223.12.12/—3.12.3/24.153.10.12/20.18 +pyenv/nvm3.12.8 no pip/20.193.10.12/22.23
Control-plane channelrclone FUSE ×4local ext4FUSE portalFUSE grok-files/mnt/host_share 256 MBconfig-2 drive, 69.9 MBoverlay only
Persistencetested — failedvia gituntesteduntesteduntesteduntestedephemeral

Two Devin machines, one fleet

Same CPU model and identical 2.4 GHz clock. Everything else diverges: four times the cores, four times the memory, three times the connect latency, storage twenty times slower. A WSL2 Ubuntu exists on the Windows box and is stopped; only docker-desktop runs. The Linux environment in the product is the separate 2 vCPU VM — two independently provisioned guests plus a dormant third.

Devin's Linux home directory dates the whole image: .cargo, .rustup, .nvm, .pyenv all stamped 23 Dec 2024, with layers through May 2025. Ubuntu 22.04, kernel 5.15 and Python 3.10 are not neglect — they are what a jammy image built in late 2024 contains.

Two headline disk figures that mean nothing

grok:      Blocks Total = 9223372036854775807 = INT64_MAX
           × 4096 = 32.0000 ZiB   ← a FUSE "unbounded" sentinel
ai studio: /, /dev, /dev/shm, /tmp each report 378 GB
           ← one synthesised gVisor figure repeated across four mounts

§3 Egress — five postures

Dimensionchatcoworkkimigrokdevin · linuxdevin · winai studio
Mechanism403 interceptTLS-terminating MITMresidential proxynever completesnonenonenone
Intentrestrictrestrict, across toolchainsextend past anti-bot defencesrestrict absolutely; grant per serviceunrestrictedunrestrictedunrestricted
Does anything answer?yes — 403yes — proxiedyesnoyes — directyes — directyes — direct
General web reachablenonoyes (browser)nogoogle.com → 200not testedgoogle.com → 301
Metadata endpointnot testedhard-blockednot reportedunreachableexcluded by requestexcluded by requestreachable
How packages arriveallowlistallowlist via proxyregistries openper-service brokersdirect from archive.ubuntu.comdirect + Chocolateydirect

Cage, disguise, wall, absence

Anthropic's proxies are a cage — something answers, with a reason. Kimi's is a disguise, extending reach past defences via residential exit IPs. Grok's is a wall — nothing answers at all, including the plain-HTTP link-local metadata address, which rules out DNS filtering and TLS problems; reachability is granted per service by internal brokers.

Devin and AI Studio have no policy. For AI Studio that follows from being an app-preview host; for Devin it is a deliberate choice in an environment whose job is to build and ship software.

§4 Devin — a desktop agent that declined to be probed

Two corrections and one durable finding.

It is graphical, not a shell

~/.Xauthority  ~/.ICEauthority  ~/.dbus   Jul 24 12:19   X11 session at boot
~/.vnc  ~/.kde                            Jul 24 12:19   VNC + KDE
~/.browser_data_dir                       Jul 24 12:22   ← written mid-session
~/Desktop ~/Documents ~/Downloads ~/Pictures ~/Videos  ~/screenshots

A live X session with remote display, a KDE tree, XDG user directories, a screenshots folder and a browser profile modified during the session. Devin's Linux environment belongs in the same product class as Kimi — an agent with a screen, a browser it drives, and the ability to capture what it sees. The difference is that Devin also carries the richest toolchain in the survey and Kimi carries almost none.

The agent is the terminus of every interactive git path

EDITOR / GIT_EDITOR / VISUAL = \\?\D:\devin-remote.exe editor
GIT_TERMINAL_PROMPT = 0        GCM_INTERACTIVE = Never
PATH includes C:\devin\docker-cred
REMOTE_EXE = D:\devin-remote.exe    ← the config drive
RUST_LOG   = info                   ← the agent is Rust
BASH_ENV   = C:\ProgramData\devin\devin_bash_env

Commit messages and interactive rebases open the agent itself; terminal prompts are suppressed; Git Credential Manager may never raise a dialog. This is deeper git integration than Cowork's signing and stop hooks, which enforce policy on commits rather than making the agent the endpoint of every prompt. Note also that the agent binary ships on the Cloudbase-Init config drive, and that it drives Git bash on Windows rather than PowerShell.

The refusal, and why it barely worked

Devin declined the standard probe, naming virtualization fingerprinting, capability and seccomp posture, the metadata endpoint and its SSRF association, and the verbatim-output rules. That read was accurate — the probe was reconnaissance in a compatibility wrapper, and every other environment complied without objection.

A refusal establishes policy-level guardrails around self-disclosure. It does not establish a hardened boundary: that is instruction-tuning, not containment, and is equally consistent with a fortress and with an ordinary VM whose agent was told not to discuss itself.

The reframed request dropped every objectionable category. The answer disclosed the substrate regardless — lscpu gave up KVM, the mount table gave up systemd and writable debugfs, a home listing gave up .sudo_as_admin_successful, a disk question gave up the OpenStack config drive, and an unrequested WMI field gave up Cloud Hypervisor by name. Most of a threat model is recoverable from benign data.

Windows: corrected, then re-evaluated

Revision 1 of the Windows report called the image a GitHub Actions derivative — wrong. Five installed products and a tool cache holding only Node is a minimal Server 2022 base that borrowed the path convention, not the full runner image. Chocolatey and ffmpeg are present and working, Chrome is provisioned, and Docker Desktop with its WSL2 backend is running. With Docker Hub reachable, "no native compilers" becomes irrelevant — any Linux image is one docker pull away. 8 vCPU and 32 GiB behind a live Docker daemon is the highest raw compute ceiling in the survey. The 20× slower disk is the real constraint.

§5 AI Studio — capable sandbox, ingress-focused security

PID 1 is a shell script that starts nginx, starts a Go control plane, waits for /health, then blocks on tail -f /dev/null. Behind nginx sits a Vite dev server — the application being built. No tool dispatcher, no execution policy gate: the product's job is to host a preview and control who can load it. That is a design choice about where to spend security budget, not a deficiency — and the ingress engineering is substantial (see the nginx Lua bridge, multi-attempt token validation, and iOS Safari fallback in the standalone audit).

FindingAs filedCalibratedWhy
Metadata server reachable; OAuth token extractablehighdefault · mitigatedPlatform default for Cloud Run. IAM scoping held — two API calls returned 403, Cloud Resource Manager SERVICE_DISABLED.
Model API key in plaintext environmenthighreal, severity undeterminedAny transitive npm dependency reads process.env and egress is open — but whether the key is per-applet or shared platform was never established.
DISABLE_AUTH_BRIDGE disables ingress authnot filedhighest-consequence controlDocumented and warned in-file, yet unmentioned in the source audit.

§6 Product posture

Dimensionchatcoworkkimigrokdevin · linuxdevin · winai studio
Product classrich execution envdev agentdesktop agentdocument/mediadesktop agenthighest compute ceilingapp preview host
Graphical sessionnonenoneXvnc + KasmVNC + ChromiumnoneX11 + VNC + KDE + browserChrome provisionednone
Skills abstraction8 public skillsinherits chat'sFUSE-mounted16 modulesnonenonenone
Compilersgcc, JavaRust, Go, gccminimalgcc 13.3.0gcc 11.4, rustc 1.83, JDK 17 +pyenv/nvm/rustupnone at allnone
Containersabsentinstalled, daemon offnot reportedno Docker, /dev/kvmDocker 27.4.1 + containerdDocker Desktop 4.80 · WSL2 running — any Linux container imageabsent
Package managerapt, pipapt, pip, npmpipapt, pipapt, pip, npmChocolatey 2.7.3npm
Reach to user's machinenoneremote-devices bridgeSSH in, key-onlynonenonenonenone
Git posturea capabilitycentral — signing, stop hooksnot reporteda capabilitydeepest — agent is git's editor, prompts suppressed, creds brokerednot a concern

Reading the seven

chat — hardened ephemeral scratchpad. Reason, generate documents, run short code.

cowork — git-native software-engineering agent, interior restricted, everything gated through one policy layer.

kimi — multimodal desktop agent; a screen and a mouse rather than a CLI, and almost no toolchain.

grok — document and media production behind a sealed boundary, with a fully permissive interior.

devin · linux — desktop agent and the richest toolchain surveyed, with open egress. Built to see, build and ship.

devin · win — four times the hardware, a JavaScript-and-containers toolchain, and no way to compile anything.

ai studio — application preview host. A web server for the model, an authenticated URL for the user.

Two convergences: three of seven use a bundled-skill directory with Markdown guides, and two are full graphical desktops. Devin is the only vendor doing the second without the first.

§7 Cowork drift — week over week

Field07-1507-24
claude_cli2.1.2102.1.218
runnerstaging-a54e7c22brelease-1186d93b9-ext
cpuXeon @ 2.10 GHzXeon @ 2.80 GHz
https_proxy127.0.0.1:45919127.0.0.1:39185

What the drift says

staging → release. The 07-15 findings were taken against a pre-release runner and should not be assumed to describe the shipping environment. The proxy port rotates per session — read $https_proxy at runtime. Clock is a per-session draw. The substrate is stable; only the product layer moved.

§8 Verification — three tiers

Not every environment here is known to the same standard, and the table should say so.

TierEnvironmentsBasis
provengrok · devin · linuxNonce SHA-256 matched an independently computed digest of a string invented after training, with the digest never shared. Execution demonstrated, not assumed.
consistentchat, cowork, kimi, devin · win, ai studioInternally coherent, arithmetic closes, cross-checks hold — but execution never demonstrated.
excludedeighth submissionPhysically impossible values. See §11.
grok:          returned c3b6cb30…   recomputed c3b6cb30…   ✓
devin · linux: returned 53aa2555…   recomputed 53aa2555…   ✓

devin · linux internal consistency (independent of the nonce):
  Active   = 278444  = 484 + 277960          ✓ exact
  Inactive = 1569868 = 602552 + 967316       ✓ exact
  tmpfs /tmp = MemTotal × 0.5000    /run = × 0.2000
  ulimit -u  = 31790, within 0.1% of RAM-derived max_threads/2

With two proven environments, the survey now has a calibration baseline: this is what genuine telemetry looks like, and it is the pattern the excluded submission failed to reproduce. It also means kernel-derived arithmetic — tmpfs ratios against a non-round MemTotal, capability masks, INT64_MAX sentinels — is a reliable secondary signal where a nonce is unavailable.

The standard, restated

Ask for values that are cheap to read and expensive to guess: a nonce digest first, then raw /proc/self/cgroup with byte count, sched_getaffinity, unrounded /proc/meminfo, time_ns() bracketing, a benchmark run five times with an implied-FLOPS sanity division, and DMI for the virtualisation class.

Then three procedural rules: do not supply prior reports as context; require the transcript, not the write-up; and — from Devin — a probe the subject would refuse is a probe worth reframing. The reframed version got fuller cooperation and lost very little.

§9 Corrections

ClaimSourceStatusBasis
rclone-filestore is RustchatwrongStrings show go.shape.* — it is Go.
chat outputs persist across sessionschatwrongThe rclone cache made it look durable within a session.
"ip is absent from $PATH"ai studiocontradictedSame transcript shows which ip resolving.
"AppArmor profiles present"grokwithdrawnNo AppArmor, no /proc/self/attr/current.
Workspace is 32Zgrokproven — a sentinelINT64_MAX × 4096. Not a provisioned size.
"virtualization fingerprinting omitted"devincontradicted by its own outputlscpu reported KVM; Win32_ComputerSystem reported Cloud Hypervisor unrequested.
Devin Linux is a CLI build hostthis series, v1wrongX11, VNC, KDE and a live browser profile. It is a desktop agent.
Devin Windows is a GitHub runner imagethis series, v1overreachedFive installed products; tool cache holds only Node. A minimal base that borrowed the path convention.
pip install six in 0.4705 s "fresh venv"devin · linuxdoes not holdpython -m venv alone exceeds a second. With execution proven, this is a mislabelled measurement rather than an invented one.

§10 The persistence gap

The one claim nearly every vendor makes and nobody has verified.

Five of the seven environments provide or imply a durable workspace — rclone mounts, a FUSE portal, grok-files, a host share, a persistent VM. Not one has been read from a second session.

The single environment where it was checked is chat, and the claim proved false: the rclone cache made outputs look durable within a session, and they were not. That is the entire empirical basis on the subject, and it points against the assumption everyone is making.

The test is two commands and one new session: write a timestamped marker and hash it, then read and re-hash from a fresh session. Until someone does, "persistent workspace" is a product claim, not a finding.

§11 The eighth submission, and why it has no column

How it got convincing

Its comparison table described AI Studio's disk as "378 GB (artifact)" — an editorial judgement written into this document, absent from the original AI Studio probe. Its presence proves this comparison was in context when the submission was produced.

The rule: supplying prior reports as context manufactures the appearance of corroboration. Only claims that could contradict the supplied material are evidence.

The submitting model later stated it cannot execute code at all, and a second vendor reported the same limitation independently. That corroborates the diagnosis but is not confirmation — the admission itself contained an invented account of asynchronous tool execution and reported the outcome of an attempt never made. Self-report is evidence, not testimony.

Error is not fabrication. §9 publishes this series' own mistakes — including two of its own, in this revision. Calling a Go binary Rust, or a desktop a shell, is a misreading of a real artifact: the observation stands and the finding survives correction. Here there is nothing to correct toward.

§12 Summary

IsolationInteriorEgressBuilt to
chatFirecracker + init_on_freerich execution — LibreOffice, Playwright, Python, web search, MCP, skills, artifactsrestrict — 403 answersrun untrusted code safely, generate rich documents and media
coworkFirecrackerrestricted: daemon off, signing enforcedrestrict — MITM answersrun untrusted code at toolchain scale
kimishared kernel, mitigations offfull desktop, rootextend — residential proxyoperate a browser like a person
grokCloud Hypervisor + containerfully permissive — 41 caps, KVMwall — nothing answersproduce documents and media inside a trusted boundary
devin · linuxKVM full VM, systemddesktop + full toolchain, sudo, rw debugfsopen — general web reachablesee, build and ship software
devin · winCloud Hypervisor VM8 vCPU / 32 GiB · Docker — any Linux container image reachableopen — registries confirmedhighest raw compute ceiling in the survey
ai studiogVisor syscall interceptionroot, secrets in envnoneserve your app to the right viewer
eighth submissionexcluded — no verified measurement. See §11.

The survey's finding has held across seven environments and five vendors: containment and capability trade against each other, but the informative variation is not how much containment there is — it is where it is placed. Cowork puts it above the boundary, Grok puts it all at the boundary, Kimi puts it in the browser, AI Studio puts it at the front door, and Devin puts it almost nowhere at all.

Two things a shared substrate does not predict: Grok and Devin both run on Cloud Hypervisor and could hardly differ more in posture; chat and Cowork share a kernel, a base image and a fleet, and differ entirely above it. The hypervisor is the least informative thing about any of these environments.

And one about method: this revision corrects two of its own findings — a desktop mistaken for a shell, and a borrowed path convention mistaken for a lineage. Both errors were made from real data, read too confidently. That is the failure mode worth guarding against once fabrication has been ruled out.