This Week AI Moved to the Desktop: Qwen Book, Assistant 'o', DeepSeek Harness, Copilot Super-App — How the Resident Workbench Redefines Human-AI Interaction

This Week AI Moved to the Desktop: Qwen Book, Assistant “o”, DeepSeek Harness, Copilot Super-App — How the Resident Workbench Redefines Human-AI Interaction

1. Introduction: Four Forces Racing Toward the “Desktop” in One Week

If you had to sum up frontier-AI progress this week in one word, it would be “desktop.”

In nearly the same time window, superficially unrelated things happened:

  • Alibaba unveiled Qwen Book, its first AI-agent laptop, at Yunqi Conference — with a Skill keyboard array, a global AI key, and a voice/pen — claiming it “understands intent and evolves on its own”;
  • Developers dug an always-on assistant “o” (internal codename Aeon) out of the ChatGPT frontend code — positioned as a long-online, always-running agent backed by a cloud-hosted sandbox;
  • DeepSeek Harness desktop leaked via the nightly channel, shipping with four modes (Standard / PTC / Minimalist / Creative) and an “Everything is a Plugin” Cordis plugin system;
  • Xiaomi open-sourced MiMo-V2.6 while also shipping MiMo Desktop as an official client;
  • Microsoft upgraded the new Copilot into a “super-app” with Home, Code, and Autopilot — turning AI from a chat tool into an agent that keeps working.

These moves come from giants of different camps and technical routes in both China and the US, yet they all point in the same direction at nearly the same time — AI is migrating from “chatbots in a cloud dialog box” to “OS-grade partners resident on your desktop workbench.”

This is no coincidence, but the inevitable result of three currents converging: model capability, tool ecosystem, and interaction paradigm. This article unpacks the “AI desktopization” wave from three dimensions: product phenomena, technical architecture, and paradigm meaning.

2. Phenomena: The “Desktop Missions” of Four Forces

First let’s see what each force is doing, then why they converge.

┌─────────────────────────────────────────────────────────────────┐
│            Last week of Sep 2026: Agent desktopization map       │
├─────────────────────────────────────────────────────────────────┤
│                                                                 │
│  Alibaba Qwen Book          OpenAI Assistant "o" (codename Aeon) │
│  · first AI-agent laptop    · always-on personal agent           │
│  · Skill keyboard array     · cloud sandbox (hours~weeks)        │
│  · global AI key            · Pro Lite $100/Pro $200/Pro Max     │
│  · voice + pen              · vs Grok Bot & Meta Muse            │
│  · partners with Omarchy    · expected at DevDay (9/29)          │
│        │                          │                             │
│        └───────────┬──────────────┘                             │
│                    ▼                                            │
│            <!-- AI resident desktop workbench -->                │
│                    ▲                                            │
│        ┌───────────┴──────────────┐                             │
│        │                          │                             │
│  DeepSeek Harness desktop       Microsoft Copilot super-app      │
│  · Standard/PTC/Minimal/Creative │ · Home desktop entry         │
│  · Everything is a Plugin(Cordis)│ · Code app building          │
│  · model/tool/session/sandbox/storage all plugins               │
│                                   · Autopilot autonomous exec   │
│  (Xiaomi MiMo Desktop shipped)   · from chat → agent           │
└─────────────────────────────────────────────────────────────────┘

1. Alibaba Qwen Book: Building the “workbench” into a laptop

Qwen Book, shown at Yunqi, is, in product form, a computer designed for an agent: the Skill keyboard array provides a whole row of definable agent-capability hotkeys; the global AI key makes summoning an agent as easy as pressing Esc; the voice + pen covers two natural input modes. It claims to “understand user intent and evolve,” and partners with the open-source Omarchy project to explore an agent-native desktop OS.

The signal is blunt: traditional operating systems (Windows/macOS) were designed for “humans operating files with mouse and keyboard”; an agent needs a runtime designed for “agents resident on the device, autonomously orchestrating tools.” When AI becomes a first-class citizen, the desktop OS should be redesigned for it.

2. OpenAI Assistant “o” (Aeon): Making the agent a “butler”

Developers found lowercase “o” in the ChatGPT frontend — positioned as an always-on personal agent: built-in Fast Mode, a multi-agent shared message board, and an -o email suffix. Reports say the internal codename is Aeon, attached to a cloud-hosted sandbox, able to run for hours to weeks. Tiers are Pro Lite ($100), Pro ($200), and not-yet-live Pro Max ($500).

The keyword for “o” is “long-online, sustainable operation.” It is no longer a bot that answers when asked; it is a resident proxy that watches, works, and reports back when done — the mark of agents moving from “session” to “service.”

3. DeepSeek Harness desktop: putting “Everything is a Plugin” on the desktop

DeepSeek Harness desktop leaked via the nightly channel (the binary is signed and notarized with DeepSeek’s Apple developer certificate), letting you log in or fill an API key, pick a local workspace, and deploy agent tasks directly — no more npx Web UI startup.

It defines four operating modes:

  • Standard: handles most tasks (code, files, materials); the agent calls retrieve/edit/terminal tools on demand;
  • PTC: leans further toward batch tool invocation, filtering, organizing, deduplicating, counting, or summarizing results;
  • Minimalist: uses only terminal tools, for testing and benchmarking the “bare agent”;
  • Creative: lets users customize DSH via conversation — the agent can write plugins to add features/UI, and combine tools + prompts into custom modes.

Its core philosophy, “Everything is a Plugin,” deserves special attention: model, tool, skill, session, sandbox, storage, agent loop, scheduling, even UI — all agent capabilities are provided by the underlying Cordis plugin system. This is an architectural stance to standardize and pluggify “agent capability” itself.

4. Microsoft Copilot super-app: upgrading the “chat tool” into an “agent”

The new Copilot adds Home, Code, and Autopilot:

  • Home: a resident desktop entry/workspace for Copilot, not a floating dialog;
  • Code: lets Copilot build applications and produce runnable projects;
  • Autopilot: lets Copilot go from “you tell it what to do” to “it completes tasks for you on its own.”

Microsoft itself says the goal is “extending the AI assistant from a conversation tool to an intelligent agent that executes tasks, builds applications, and keeps working.”

3. Why the desktop? Three currents converging

The four forces look different but lead to the same place. Behind them, three underlying variables matured at once:

1. Model capability reached the threshold of “stable long-term work.” As discussed in the previous article, GPT-6 Astra can already drive software through scripting APIs and operate a real computer on OSWorld. When a model can reason, write code, call tools, and “run—feedback—self-correct,” “leaving it running to work for you” finally becomes meaningful.

2. The tool/sandbox ecosystem matured. Long-running execution requires an isolated, resumable, resource-bounded environment — which is exactly what cloud-hosted sandboxes (Aeon), local workspaces (Harness), and agent runtimes (Qwen Book’s OS-native design) all address.

3. The interaction paradigm is shifting. Old human-AI interaction was a request-response pull model. Resident agents mean a push model: the agent actively watches and works, then reports when done. This requires products to upgrade from “dialog box” to “workbench/workspace.”

┌─────────────────────────────────────────────────────────────────┐
│  Why AI moves to the desktop: interaction paradigm shift         │
├─────────────────────────────────────────────────────────────────┤
│                                                                 │
│  Old paradigm: Pull (session / Q&A)                              │
│  ┌─────┐ ask  ┌───────┐  answer  ┌─────┐                        │
│  │ user │ ───► │ dialog │ ◄────── │ user │  ask one, get one     │
│  └─────┘      └───────┘         └─────┘                        │
│                                                                 │
│  New paradigm: Push (resident)                                   │
│  ┌─────────┐  delegate  ┌────────────────┐                       │
│  │  user   │ ──────────►│ resident Agent │                       │
│  │(set goal)│◄──────────│ (workbench run) │                      │
│  └─────────┘   report   │ · auto-watch    │                      │
│                     ▲   │ · self-schedule │                      │
│                     │   │ · cloud sandbox │                      │
│                     │   └───────┬────────┘                      │
│                     │           │ call                          │
│                     │           ▼                               │
│                     │      ┌───────────┐                        │
│                     └──────│ file/API/Term│                      │
│                            └───────────┘                        │
└─────────────────────────────────────────────────────────────────┘

4. Technical breakdown: under the hood of an “agent desktop”

Whatever the shell — laptop, client, or super-app — nearly every agent-desktop product shares a runtime architecture. Understanding it unlocks the whole wave.

Here is a generalized (aggregating all approaches) agent-desktop runtime layering:

┌─────────────────────────────────────────────────────────────────┐
│                 Agent Desktop Runtime (layered)                  │
├─────────────────────────────────────────────────────────────────┤
│  ┌────────────────────────────────────────────────────────────┐│
│  │ ① Interaction/UI layer (Human-in-the-loop)                  ││
│  │   resident panel · global hotkeys · voice/pen · approve     ││
│  └───────────────────────────┬────────────────────────────────┘│
│                              ▼                                 │
│  ┌────────────────────────────────────────────────────────────┐│
│  │ ② Agent Loop layer (intent·plan·exec·reflect)               ││
│  │   state memory · task queue · multi-agent · self-correct    ││
│  └───────────────────────────┬────────────────────────────────┘│
│                              ▼                                 │
│  ┌────────────────────────────────────────────────────────────┐│
│  │ ③ Capability orchestration (Skill / plugin system) ★core    ││
│  │  model·tool·skill·schedule·storage·session·UI all pluggable ││
│  │  (DeepSeek Harness = Everything is a Plugin / Cordis)        ││
│  └───────────────────────────┬────────────────────────────────┘│
│                              ▼                                 │
│  ┌────────────────────────────────────────────────────────────┐│
│  │ ④ Execution sandbox (Ephemeral workspace)                   ││
│  │   FS · terminal · process · egress(limited) · cred-isolate  ││
│  └───────────────────────────┬────────────────────────────────┘│
│                              ▼                                 │
│  ┌────────────────────────────────────────────────────────────┐│
│  │ ⑤ Host/infrastructure                                       ││
│  │   cloud (AgenticCloud/sandbox) or local desktop (workspace) ││
│  └────────────────────────────────────────────────────────────┘│
└─────────────────────────────────────────────────────────────────┘

Layer ①: Interaction/UI — from “dialog” to “workbench”

Field-level differences aside, what matters is that the interface shape decides the mental model: a dialog implies “one question at a time”; a workbench implies “long-term collaboration.” A global AI key and a Skill keyboard array turn “calling an agent” from “opening a webpage” into a system-level operation.

Layer ②: Agent Loop — the engine of “always-on”

An agent can “run for hours to weeks” because it is not a function but a continuously running loop: maintain state → decompose intent → plan → execute tool calls → observe results → reflect and fix → enter the next round. Here is a highly simplified resident-loop skeleton:

# ============================================================
# Concept: the "resident loop" at the heart of an agent desktop
# ============================================================
class ResidentAgentLoop:
    def __init__(self, brain, skill_registry, sandbox):
        self.brain = brain            # reasoning/decision model
        self.skills = skill_registry  # capability orchestration
        self.sandbox = sandbox        # execution environment
        self.memory = []              # session-level state

    def step(self, observation) -> str:
        # 1. feed current state to the brain
        decision = self.brain.plan(memory=self.memory, new_obs=observation)
        # 2. tool call -> dispatch via skill registry
        if decision.kind == "tool":
            result = self.skills.invoke(decision.tool, decision.args,
                                        in_sandbox=self.sandbox)
            self.memory.append(("tool_result", result))
            return "keep_running"     # loop continues
        # 3. task judged done -> finalize and report
        if decision.kind == "final":
            return decision.answer
        # 4. otherwise (think/wait) stay in the loop
        return "polling"

agent = ResidentAgentLoop(brain=model, skill_registry=plugins,
                          sandbox=ephemeral_ws)
while not agent.should_exit():
    agent.step(await_next_signal_or_schedule())

The key is keep_running: after finishing a task the agent does not exit; it returns to “waiting for a new goal / being scheduled.” That is the fundamental difference between a “resident workbench” and a chatbot.

Layer ③: Capability orchestration — why “Everything is a Plugin” is a watershed

Traditional agent capabilities are “hardcoded”: adding one requires retraining or hardcoding. “Everything is a Plugin” abstracts model, tool, skill, session, sandbox, storage, agent loop, scheduling, UI into pluggable units. The benefits are striking:

  • plug-and-play capability: attach new tools, data sources, even new modes of thinking without touching the core;
  • user customizability: let the agent “write its own plugin” to add features through conversation (Harness’s Creative mode);
  • composability: tools and prompts combine into “custom modes,” assembling agent capability like building blocks.

Here is a “Skill plugin” interface illustration, emphasizing extension via a unified contract:

# ============================================================
# Concept: unified interface contract for pluggable Skill plugins
# enables dynamic loading of "capability" into the agent runtime
# ============================================================
from abc import ABC, abstractmethod
from dataclasses import dataclass, field
from typing import Any

@dataclass
class SkillManifest:
    name: str
    version: str
    tools: list[str] = field(default_factory=list)   # exposed tools
    modes:  list[str] = field(default_factory=list)  # custom modes

class SkillPlugin(ABC):
    """Unified lifecycle every plugin must implement"""

    @abstractmethod
    def on_load(self, runtime) -> None:
        """register tools, init resources, declare capabilities"""
        ...

    @abstractmethod
    def on_unload(self) -> None:
        """release resources, exit safely"""
        ...

    @abstractmethod
    async def handle(self, ctx, action: str, args: dict) -> Any:
        """execute one tool call, return structured result"""
        ...

# e.g. a "terminal" plugin
class TerminalFeature(SkillPlugin):
    def __init__(self):
        self.manifest = SkillManifest(name="terminal", version="1.0",
                                      tools=["exec_shell","read_file"])
    def on_load(self, runtime):
        runtime.register_tool("exec_shell", self.handle, sandboxed=True)
    async def handle(self, ctx, action, args):
        if action == "exec_shell":
            return await ctx.sandbox.exec(cmd=args["cmd"], timeout=args.get("t"))
        ...

# runtime discovers and loads dynamically; core loop never changes
runtime.load(SkillPlugin() for SkillPlugin in discover_plugins())

The value: the core loop never cares what a plugin does; it only honors the on_load / handle / on_unload contract. That raises extension efficiency and makes third-party ecosystems (user plugins, shared team skills) possible.

Layer ④: Execution sandbox — the safety foundation of the resident era

Once an agent runs long-term and can touch terminals and files, security risks intensify sharply (exactly why “the more capable, the more permissions must collapse,” as the previous article discussed). So agent desktops generally isolate execution in a sandbox:

  • FS isolation: the agent can only read/write the workspace assigned to it;
  • egress restriction: outbound connections blocked by default or through allowlist proxies;
  • credential stripping: the agent only gets temporary/minimal credentials;
  • resource quotas: CPU/memory/time limits to prevent runaway spinning or exhaustion.

Cloud hosting (AgenticCloud, Aeon’s sandbox) and local desktops (Harness’s local workspace) are just two “host” forms; the safety layering is identical.

5. Comparison: “workbenches” converging from different angles

ProductFormCore selling pointExecution envPluggable capabilityStatus
Alibaba Qwen Bookagent laptopOS-native for agentslocal+cloudSkill keyboard/pluginsYunqi debut
OpenAI “o”resident assistantalways-on, shared boardcloud sandboxplugin planneddug/waiting
DeepSeek Harnessdesktop clientEverything is a Plugin (Cordis)local workspacestrong (model/tool/UI)nightly
Xiaomi MiMo Desktopdesktop clientlocal run + RSIlocalmediumofficial
Microsoft Copilotsuper-appHome/Code/Autopilotcloudplugin ecosystemshipped

Despite different forms, they give one consistent answer to “where AI’s future runs”: beside the user, resident, able to orchestrate tools autonomously, and able to let capability grow like plugins. Qwen Book wants to redefine the OS, Aeon the “butler,” Harness the “agent developer tool,” and Copilot the “everyday super-app.” They rush into the same room through four doorways.

6. Challenges and concerns

We must stay sober: this wave has three hard bones to chew.

1. Homogenization risk. If a “desktop agent” is just “an always-on dialog + a row of feature buttons,” it barely differs from the old chatbot. Real value lies in deep tool orchestration and stable autonomous execution, which take long-term polish — not something a shipped client delivers overnight.

2. Safety and trust. Letting an agent be “always-on, touch files, call terminals” means handing over growing authority. “Running for hours to weeks” means enough time to make mistakes, overreach, or get attacked. Least privilege, human-in-the-loop, and sandbox isolation are not optional — they are the line separating shippable product from liability (see the previous article on Astra’s critical-level cyber ability).

3. Ecosystem and standards. Every vendor builds its own Skill/plugin system; no unified contract exists yet. An agent desktop’s value rides heavily on “how many third-party capabilities can plug in immediately,” which hinges on ecosystem and standards battles. Whoever defines the “common language” for plugins takes the platform gateway of the agent era.

7. Conclusion

This week, Alibaba, OpenAI, DeepSeek, Xiaomi, and Microsoft almost simultaneously made the same judgment: AI’s next home is the resident desktop workbench. That is not a press-release coincidence, but the footnote of three underlying variables (the model can work, the sandbox works, the paradigm should change) maturing at once.

Core takeaways:

  1. The paradigm is shifting: human-AI interaction is moving from “pull — ask one, get one” to “push — set a goal, the agent runs continuously for you”;
  2. Architecture consensus: agent desktops generally consist of five layers — interaction, agent loop, capability orchestration (plugin system), execution sandbox, host — and “Everything is a Plugin” is the most imaginative orchestration philosophy;
  3. Entry differs, destination aligns: redefining the OS (Qwen Book), the butler (Aeon), the developer tool (Harness), the super-app (Copilot) — four roads, one terminus;
  4. Safety is the lifeline: the safety design of an always-on agent that can touch files/terminals decides whether it can be truly trusted and used.

If the Astra article in the previous piece answered “whether AI can do great work,” this week’s desktopization wave answers “where and in what form AI should work for the long run.” When agents move from cloud dialogs into your desktop to run resident, our relationship with machines is quietly shifting from “using a tool” to “working alongside a colleague.”

That is not just a product change; it is the beginning of a new digital way of life.