ISMS Copilot Docs

What delegation saves (and what it does not)

How handing GRC work to ISMS Copilot over MCP keeps context out of your agent's transcript, what your agent still pays for, what the usage receipt shows, and where ISMS Copilot usage is billed.

When your coding agent hands a compliance question to ISMS Copilot, part of the work happens on ISMS Copilot's side instead of in your agent's context. This page explains the mechanism, what it does not change, and how the two sides are billed. It makes no claim about how much you save: that depends on your harness, your model and your work.

Why your agent's context matters

A chat-based agent does not remember earlier turns on its own. On each turn, a typical harness sends the model its transcript again: your instructions, the tool definitions, earlier messages, files it read and tool results. Anything that enters that transcript is sent or processed again on later turns and may count toward usage, subject to caching, compaction, and provider accounting. ISMS Copilot's own chat works the same way, which is why long threads use more of your usage window per message.

So the question is not only what a single answer costs, but what stays in your agent's transcript afterwards.

What stays on ISMS Copilot's side

When the work is delegated, the items below stay on ISMS Copilot's side unless the returned result reproduces them. The MCP request your agent sends and the result it gets back (the reply plus fields such as IDs, status, usage and errors) do enter its transcript, alongside the tool schemas covered further down:

  • Framework source material. The standard and regulation references ISMS Copilot grounds its answers in. Your agent sends the question, not the framework text. The reply may quote or cite the parts that matter.
  • Retrieval. Finding the relevant clauses, controls and articles for the question, and the context assembled for the model.
  • Workspace memories and files. On a turn scoped with workspace_id, they are read on ISMS Copilot's side. Your agent does not load them, though the reply may use facts from them.
  • Intermediate reasoning. The analysis behind the answer, including multi-step Think and Beyond work.
  • Drafting passes. Long text is generated on ISMS Copilot's side. Your agent receives the result, not the drafting process.
  • The specialist's copy of the thread. Follow-ups with send_message continue a thread ISMS Copilot keeps, so your agent does not resend earlier turns of that thread in its request. The earlier tool calls and replies still sit in your agent's own transcript until it trims them.

What your agent still pays for

Delegation does not make the compliance step free for your harness:

  • Tool schemas. The ISMS Copilot MCP tool definitions sit in your agent's context once the server is connected, like any other MCP server's tools.
  • The question it sends. Whatever your agent writes into create_conversation or send_message is part of its transcript.
  • The reply it reads. The returned answer is input to your agent's model, and it stays in the transcript for the rest of the session.

The reply is the part you control most. Pass answer_format: "brief" (about 150 words) or "decision" (recommendation first, about 300 words) when your agent only needs the conclusion. When the deliverable is a long document, have your agent write it to a file rather than repeat it in chat. See Delegate GRC work from your agent.

The usage receipt

A completed Fast or Think reply can carry a usage object:

{
  "usage": {
    "copilot_input_tokens": 0,
    "copilot_output_tokens": 0
  }
}

The values above are placeholders. The fields are:

  • copilot_input_tokens and copilot_output_tokens: tokens ISMS Copilot's models processed and produced for that turn.
  • copilot_cache_read_input_tokens and copilot_cache_creation_input_tokens: included only when the model provider reports them.

This receipt counts ISMS Copilot-side tokens only. It is not your agent's token count and does not include what your harness spends on the question, the tool schemas or reading the reply; your harness reports those itself. The usage object is omitted when the numbers are not available yet (get_reply includes it once they are recorded) and Beyond replies do not carry it.

Where ISMS Copilot usage is billed

There is no separate MCP product or meter. Delegated turns count against your ISMS Copilot chat plan:

  • The 4-hour usage window. MCP turns draw from the same 4-hour UTC session window as the chat app, on your own budget or your organization's pool. See Understanding usage limits.
  • Overflow. On eligible paid solo accounts (not Essential, Advanced Data Protection off), when the window is used up, your agent can continue only on your explicit go-ahead by resending with overflow_consent: true. The turn then runs on the disclosed fallback models, up to an additional 2x the plan's token limit. Overflow is never available on a team pool.
  • Beyond runs. Capped at 10 per UTC day on paid plans, 50 on Unlimited.

Your harness's own model usage is billed by your harness provider, as usual.

Measured results

A reproducible benchmark of delegation is in progress. We will publish the numbers here with the method and the date they were measured, and not before.

Until then, this page describes the mechanism only. It does not claim a percentage saving, a lower cost than any other model or subscription, or better answer quality than your harness's own model.

On this page