Plugin-first agent runtime guide

DeepSeek Harness: run agents the plugin-first way, with full traceability

Describe the agent task, mount models and tools, then inspect prompt -> action -> observation -> replay as one continuous harness workflow.

Compare modes

Runnable start

Use the runtime flow on its own page

The embedded console now lives on a dedicated runtime page so the guide stays readable and the tool surface has enough room. Open it when you want the first-run command, mode choice and traceability checklist in one focused view.

First-run task
Command
npx @deepseek-ai/dsh web
Default check
Mode, mounted capabilities, first trace and recovery path.
Context
DeepSeek Harness guide, plugin-first agent runtime, full traceability.
Open runtime page

Paid console plans

Keep the guide free; pay only when you need hosted console capacity.

The guide, commands and source notes stay free. If your first run needs more hosted model capacity, open the paid Coachix Console plans and complete payment through Polar there. This site records the checkout intent, but never handles card details.

Compare payment options

Payment opens in a secure Coachix Console window.

What DeepSeek Harness is

DeepSeek Harness is the open-source agent harness around DeepSeek's agent product direction. The public preview frames the architecture around one rule: every capability is a plugin. That means the model adapter, tools, skills, session store, sandbox, scheduler, loop and user interface can be selected, swapped or recomposed through configuration instead of being hard-wired into one agent shell.

For builders, the important idea is not only extensibility. A harness is the operational layer that lets a model act in a real environment: it decides what tools exist, what context enters the run, how actions are logged, which safety boundaries apply, and how a user can resume, fork, search or replay a session. The model sets the ceiling; the harness determines whether that capability survives contact with files, terminals, long tasks, errors and product feedback.

This page is an independent guide. It summarizes the public preview, points to source links at the end, and keeps recommendations grounded in what is currently visible from DeepSeek's own page and repository. Treat the maintainer documentation and repository as the source of truth; use this guide as a practical reading path before you spend time installing, testing modes, or comparing the runtime with another agent framework.

The useful question is not whether the preview can answer one prompt. A serious evaluation asks whether the runtime makes a run inspectable after mistakes happen. Can you see which capability was mounted? Can you tell whether a tool call changed local state? Can you replay the session without guessing what context was present? Those are the details that separate a real harness from a demo assistant shell.

Quick-start lab

Run the preview, then inspect the runtime

The guide keeps the first-run workflow one click away on the runtime page. Start there when you want a runnable command, a small test task, and a checklist for deciding whether the preview is worth deeper evaluation.

Fast path

Launch the Web UI from npm

Use the package command when you only need a quick look at the developer preview surface. Run it in a disposable workspace first, then watch what the UI exposes about mode, mounted capabilities and session history.

npx @deepseek-ai/dsh web

If the command fails, check Node.js, npm access, local port conflicts and proxy settings before judging the runtime design.

Source path

Clone the project repository

Use the source checkout when you want to inspect package structure, plugin boundaries, examples, issue history or docs together with the running interface.

git clone https://github.com/deepseek-ai/deepseek-harness

Keep source changes separate from your real projects until you know what files, shells and network surfaces the selected mode may touch.

Use this first test task

Pick a harmless repository or scratch directory and ask for a bounded job such as: "Read the README, identify the main runtime pieces, and produce a short setup risk list without editing files." A task like this is intentionally small, but it still exercises context loading, tool availability, observations and the session record.

  1. Confirm launch state.The UI should make the active mode and available capability surface understandable before you run a real task.
  2. Inspect the first trace.Look for prompts, tool calls, observations, errors, retries and context injections in an order you can explain later.
  3. Change one variable.Switch mode or remove a capability, then rerun the same small task to see whether behavior changes for a clear reason.
  4. Check recovery.Interrupt or resume a run if the preview exposes that flow. Long agent work is only usable when recovery is boring and visible.

Architecture

A harness is the layer around the model

Read the system as a set of replaceable capability slots. Each slot changes what the agent can perceive, do, remember, replay and recover from.

Kernel

Cordis-style plugin mounting

The preview points to Cordis as the plugin system that manages mounting, unmounting and dependencies.

Tools

Typed action surfaces

Tools become reviewable interfaces with permission boundaries, arguments, observations and traceable outcomes.

Sessions

Append-only run records

Prompts, tool calls, observations, context injections and subagent scheduling can be inspected from one event stream.

Storage

State that can be resumed

Harness work is long-running. Durable session state lets a user pause, fork, replay and compare runs instead of losing context.

Sandbox

Environment boundaries

The harness defines what the agent may touch, how risky actions are gated, and how outputs become evidence rather than guesses.

UI

Runtime you can inspect

Developer preview UI is not decoration. It is where configuration, trajectory inspection and plugin composition become usable.

Traceability

Every run should leave a trail

When an agent touches real work, a result is not enough. Builders need to see how the run moved from prompt to tool call to observation to next action.

Try the preview

Install Node.js, then use the package preview command.

npx @deepseek-ai/dsh web
git clone https://github.com/deepseek-ai/deepseek-harness
Open the runtime page
prompt
Intent enters the run

The harness decides what context and capability surface should be available.

mount
Plugins become active

Models, tools, skills, sessions, storage and UI can be composed for the job.

observe
Tool output becomes evidence

The next model step can use structured observations rather than memory of what might have happened.

replay
The session can be inspected

Resume, fork, search and replay work because the event stream is the shared source of truth.

Runtime modes

Pick the smallest surface that still fits the work

The preview describes Standard, Code, Minimal and Creator modes. The right choice depends on whether you need breadth, orchestration, benchmarking or preset authoring.

Builder path

Evaluate it like infrastructure, not a prompt pack

A harness changes failure modes. Test the runtime around the model: trace quality, permission gates, recovery, composition and user feedback loops.

1

Run the Web UI

Confirm the preview launches locally and note which capabilities are enabled by default.

2

Inspect a trace

Check whether prompts, tool calls, observations and context injections can be understood after the fact.

3

Swap a capability

Try a small plugin or mode change before judging the architecture by the default surface.

4

Replay a real task

Use a work sample with files, errors and recovery. That is where harness design becomes visible.

Fit and boundaries

When this guide helps, and when it should send you elsewhere

The upstream project owns the roadmap, implementation details and release status. This page is useful when you need a neutral evaluation path before reading everything end to end.

Good fit

  • You are comparing agent runtimes and care about trace quality, not only model output.
  • You want to understand plugin mounting, modes, sessions, storage and UI before installing.
  • You need a short checklist for testing a preview in a low-risk local workspace.
  • You want source links collected in one place, with independent notes about what to inspect.

Use upstream sources first

  • You need current release guarantees, security policy, supported environments or API compatibility.
  • You are making production architecture decisions and need maintainer-confirmed behavior.
  • You are looking for the official DeepSeek product page, brand statement or support channel.
  • You need to verify licensing, package metadata or repository activity at the exact time of adoption.

FAQ

Short answers for developers

Use these as quick orientation before reading the project repository and docs.

Is this the official DeepSeek site?

No. This is an independent guide. Official source links are listed below so readers can verify the project directly.

What does “everything is a plugin” mean?

It means agent capabilities are mounted and composed through a plugin system: models, tools, skills, sessions, sandboxes, storage, loops, scheduling and UI.

Why does traceability matter?

Agent runs are multi-step. Traceability lets builders inspect what the model saw, which tools ran, what returned, and why a later step happened.

Which mode should I start with?

Start with Standard mode for exploration, Minimal mode for cleaner benchmarks, Code mode for orchestrated tool rounds, and Creator mode when you are shaping presets or plugins.

Can I evaluate the preview from this site?

Yes for first-pass orientation. Use the dedicated runtime page for the embedded console, then use this guide for architecture notes, source paths and the safety checklist.

What should I avoid in a first run?

Do not point a preview agent at private production code, credentials or important local files until you understand its mode, tool surface and permission boundaries.

Source notes

Official and community sources used

This guide keeps external links in source notes rather than using them as conversion paths.