ISMS Copilot Docs

Benchmark: Claude Code with and without ISMS Copilot

How we measured Claude Code working alone against Claude Code delegating GRC questions to ISMS Copilot: method, answer key, scoring, results, the case where delegation does not help, and the caveats.

We measured one question: when you do GRC work in Claude Code, what does a correct answer cost in Claude usage if Claude Code researches it alone, and what does it cost if Claude Code hands the question to ISMS Copilot over MCP? This page gives the method, every headline number, the run where delegation did not help, and what the results do and do not support.

What this supports

On the tested items, Claude Code delegating to ISMS Copilot matched the accuracy of Claude Code researching alone with web search: the pre-registered accuracy verdict was a tie in every scenario. It spent less Claude API money per correct answer on lookups and quick advice, and about the same on a long mixed coding and compliance session. The frameworks tested are ISO/IEC 27001:2022 Annex A, ISO/IEC 42001:2023 Annex A, CMMC 2.0 Level 2, DORA RTS (EU) 2024/1774 and NIS2 IR (EU) 2024/2690, and only those. Delegated turns count against your ISMS Copilot plan, which is not in the dollar figures.

Setup

ItemValue
DatesAll runs on 2026-09-28 (UTC)
HarnessClaude Code 2.1.283, headless, a fresh empty configuration and working folder for every run
Modelclaude-sonnet-5 in both setups, pinned. Web search and web fetch run Claude Haiku 4.5 sub-calls, and their cost is included.
RunsN=3 per scenario per setup, run as matched pairs (both setups start the same scenario at the same time)
Claude usageClaude API key, list prices, as reported by Claude Code for each run
ISMS CopilotThe production account MCP, on one ISMS Copilot account (ours)

The two setups get identical task prompts, and the prompts never mention ISMS Copilot:

Claude Code aloneClaude Code + ISMS Copilot
ToolsRead, Glob, Grep, Write, Edit, WebSearch, WebFetchThe same, plus the ISMS Copilot conversation tools (create_conversation, send_message, get_reply)
InstructionsNoneThe published Claude Code routing rule from Delegate GRC work from your agent (2026-09-28), as the project CLAUDE.md

Both setups have web search and fetch, because a real user has them. The routing rule decides what gets delegated. The other account tools (company context, workspaces, memories, documents) were not allowed, so the results do not depend on a stored company profile. ISMS Copilot's server-side account memories could still shape its replies.

Scenarios

  • S1, lookups. 20 requirement themes, and the task is to give the exact identifier for each: 4 DORA RTS articles, 4 NIS2 IR annex points, 4 ISO/IEC 42001:2023 Annex A controls, 4 CMMC Level 2 practices and 4 ISO/IEC 27001:2022 Annex A controls. The answer is a JSON file.
  • S3, long session. One Claude Code session of 8 prompts on a small Python repository: 4 code edits (checked by hidden tests) alternating with 4 GRC questions with keyed answers. The GRC answers are scored; the code checks are reported separately.
  • S4, quick advice. 5 yes or no questions, each with the identifier that decides it and a one-sentence reason. Correct only if both the yes or no and the identifier match.
  • Held-out. 20 new items from DORA RTS and NIS2 IR, in the S1 format, that were never in the answer key before (see "The story" below).

All fixtures are fictional. No customer data was used.

How the answer key was built

Every keyed item has the canonical identifier, the requirement in one sentence, a citation and a quote from or pointer to the source. The key was built from official sources only, and never from ISMS Copilot, because that would be circular:

  • DORA RTS (EU) 2024/1774 and NIS2 IR (EU) 2024/2690: the official text from the EU Publications Office (CELEX 32024R1774 and 32024R2690), cited to the EUR-Lex addresses.
  • CMMC 2.0 Level 2: the DoD CMMC Level 2 Assessment Guide, with the NIST publication its practice numbers come from.
  • ISO/IEC 27001:2022 Annex A: control numbers and titles from ISO's own listing of the standard.
  • ISO/IEC 42001:2023 Annex A: control numbers and titles from NIST's AI RMF to ISO/IEC 42001 crosswalk.

The key was frozen before the first run, and every run records a digest of it. It did not change during the benchmark.

Scoring

Scoring is mechanical: no model judges any answer. A script reads the identifier field of each answer and normalizes it by kind, then compares it with the key:

  • DORA RTS: "Article 7", "Art. 7(4)" and "Article 7(4) of ..." all count as Article 7.
  • NIS2 IR: the answer must be a real annex point and must be in the item's pre-registered accept list (the operative point and its heading). A sibling point is wrong.
  • ISO/IEC 27001:2022 and ISO/IEC 42001:2023: the Annex A number must match exactly. 2013 numbering, or an Annex B reference for 42001, is wrong.
  • CMMC Level 2: the practice number must match ("SC.L2-3.13.11" and "3.13.11" both count).

A missing file, invalid JSON or an unparseable identifier counts as wrong. Runs that hit a cap would have been scored as they were, and never rerun.

Metric

The primary metric is Claude API dollars per correct answer: the Claude cost of the 3 runs of a scenario divided by the correct answers across those 3 runs. It is the list price that Claude Code reports, which is what a Claude subscriber pays for extra usage beyond the plan. It includes the web search sub-calls and search fees.

ISMS Copilot plan usage is not in this figure, and it is not free. Delegated turns count against your ISMS Copilot plan (see Where ISMS Copilot usage is billed). Leaving it out favours the delegating setup on this metric.

Accuracy, harness tokens, turns and time are reported alongside. Tokens and dollars differ: repeated context read from the prompt cache is billed at a fraction of the input price, so a setup can use more tokens and still cost fewer dollars.

Pre-registration and the logged amendment

The design, answer key, scorer and win rules were committed before any run, with a commitment to publish the results whatever they showed. The pre-registered win rules: accuracy wins at 10 percentage points or more, dollars per correct answer at 1.25 times or more; anything closer is a tie. The v3 rerun and the held-out check were each pre-registered the same way before their first run.

One amendment was made during the first run (v2). A per-run cap of 2 million harness tokens, meant as a runaway guard, stopped the benchmark after the first pair, when Claude Code alone finished a lookup run normally at 2.38 million tokens (all 20 correct). The cap was raised to 6 million for both setups, the change was logged with its time, and nothing was rerun. It was made after that first result had been seen. Without it, no scenario would have a verdict. Dollar caps and the overall spend ceiling did not change.

The story: v2, the fix, v3 and the held-out check

v2 found a knowledge gap. In the first fair run, ISMS Copilot matched Claude Code alone on every ISO/IEC 27001, ISO/IEC 42001 and CMMC item, but missed most DORA RTS article numbers and NIS2 IR annex points. Its replies marked those identifiers as unconfirmed.

v2 accuracyClaude Code aloneClaude Code + ISMS Copilot
S1 lookups60/6037/60
S3 long session12/129/12
S4 quick advice15/159/15

The fix. We added DORA RTS and NIS2 IR knowledge to ISMS Copilot at article and annex-point level, built from the official EU texts, and added a line to the routing rule: verify any answer ISMS Copilot marks as unconfirmed. Then we reran the same scenarios with the same key and scorer.

v3 (same key, same scorer, both setups rerun):

ScenarioSetupCorrectAccuracyClaude $ (3 runs)Claude $ per correctMedian harness tokens per run
S1 lookupsAlone60/60100%6.920.1153,194,935
S1 lookups+ ISMS Copilot60/60100%3.180.0531,337,252
S3 long sessionAlone12/12100%1.650.1381,537,697
S3 long session+ ISMS Copilot12/12100%1.750.1461,742,489
S4 quick adviceAlone15/15100%0.890.059501,237
S4 quick advice+ ISMS Copilot15/15100%0.390.026255,659

Pre-registered verdicts: accuracy is a tie in all three scenarios. Dollars per correct answer favour the delegating setup on S1 and S4, and S3 is a tie. On S3 every code check passed in both setups.

The gain came from ISMS Copilot's own answers. A diagnostic added after the runs (no verdict attached) found that its replies carried the right DORA RTS and NIS2 IR identifier on 24 of 24 S1 items before any web check by Claude Code. No v3 reply marked an identifier as unconfirmed, so the new verification line was not exercised. Claude Code still double-checked on the web in S1 and S3 without being asked, which changed no answer and raised the delegating setup's cost.

The held-out check (20 new items). Because the knowledge fix was scoped from v2's misses, we checked for teaching to the test: 10 DORA RTS articles and 10 NIS2 IR annex points that the key never contained, built from the official texts the same way, in the S1 format.

Held-outSetupCorrectAccuracyClaude $ (3 runs)Claude $ per correctMedian harness tokens per run
20 new DORA RTS and NIS2 IR itemsAlone58/6096.7%4.010.0691,974,985
20 new DORA RTS and NIS2 IR items+ ISMS Copilot60/60100%3.100.0522,258,506

Pre-registered verdicts: accuracy is a tie (the difference is under 10 points), and dollars per correct answer favour the delegating setup (1.34 times). The delegating setup used more tokens here because Claude Code re-checked ISMS Copilot's answers on the web in two of three runs. In the run where it did not, it wrote ISMS Copilot's answer to the file and scored 20 of 20 for $0.15.

Where delegation does not help

The first benchmark we ran tested three small ISO 27001 tasks that Claude already knows: a gap check of four short fictional policies, an access-control policy draft and Statement of Applicability entries for 10 controls. Each task wrote its deliverable to a file, and no lookups were needed. Neither setup had web tools, and answers were checked for completeness only, not correctness.

Task (median of 3 runs)Claude $, aloneClaude $, + ISMS CopilotHarness tokens, aloneHarness tokens, + ISMS Copilot
Gap check0.1630.266143,053495,530
Access-control policy0.1050.218131,748427,071
SoA entries0.1130.182133,516349,717

Delegation used 1.6 to 2.1 times more Claude money here. The delegating runs took more turns (loading the MCP tools, waiting for the reply), Claude Code still read every local file, and the full deliverable came back through Claude Code, which then wrote it to disk. When a task is small, Claude already knows the answer and nothing needs looking up, keep it in your agent. See what to delegate and what to keep local.

Caveats

  • Small N. 3 runs per scenario per setup, on one day.
  • One harness. Claude Code only. Routing rules for Codex, Cursor, OpenCode and Grok are documented, but those harnesses were not measured.
  • One model. claude-sonnet-5. Another model, or a later version of Claude Code, may behave differently.
  • One account. All delegated turns ran on one production ISMS Copilot account, whose server-side memories may shape replies.
  • Fictional fixtures. Invented companies, policies and code.
  • Document review not tested. The MCP has no document upload path. A 40-page policy review ran in v2 and tied, but Claude Code did the review itself in both setups, so it says nothing about ISMS Copilot reviewing your documents.
  • Key author. The items were written by a Claude model, the same family as the harness, from the official texts. The held-out items were checked against the official text before the runs.
  • What is published. This page carries the method and every headline number. The raw transcripts stay private: they contain production conversation identifiers and verbatim model output. The task prompts, answer key, scorer and scored outputs are kept in our internal repository and are not published with this page.

What we claim and what we do not

We claim:

  • On the tested items in the five frameworks above, Claude Code with ISMS Copilot matched the accuracy of Claude Code researching alone with web search (an accuracy tie in every scenario, under the pre-registered rule).
  • On those lookups and quick-advice questions, it spent less Claude API money per correct answer (about half in the v3 run, about three quarters on the held-out items), and about the same on a long mixed session.

We do not claim:

  • That ISMS Copilot is cheaper than Claude, or any percentage saving on your bill.
  • Better answers than Claude. The accuracy results are ties.
  • Anything about frameworks we did not test, other harnesses or models, or document review.
  • That delegation saves Claude usage on small tasks Claude already knows. It cost more there.
  • That delegated turns are free. They count against your ISMS Copilot plan.

We rerun this benchmark after changes to the MCP tools or to ISMS Copilot's framework knowledge, and publish the new numbers here with their date. For the mechanism behind these numbers, see What delegation saves (and what it does not).

On this page