Skip to content
Proposals/Run `claude plugin eval` to score plugins with J
proposalteamP5Worth a lookClaude Code

Run `claude plugin eval` to score plugins with JSON and HTML reports

Claude Code v2.1.269 adds claude plugin eval for scored, reproducible plugin results (JSON + HTML). Run it on every plugin you ship, then restart sessions so managed-settings org plugins load in headless/Desktop and plugin extracts are no longer world-readable.

Why this loop

v2.1.269 adds claude plugin eval, which runs a plugin’s eval suite against Claude Code and returns scored, reproducible JSON and HTML. After upgrade, run it on every plugin this team maintains and keep those reports as the v2.1.269 baseline. The same release also changes plugin runtime behavior: organization plugins enabled through managed settings now load in headless sessions and on Claude Desktop from the next session (once Desktop bundles this CLI); synced plugin MCP servers connect when a remote session resumes; plugin archives extracted for a session are no longer readable by other local users, extracted files no longer keep world-writable bits, and stale files no longer survive re-extraction. Quit and start fresh sessions so those loads and re-extractions happen. This is also a breaking permissions release: a deny or ask rule starting with ! now applies only within the settings source that wrote it, and a bare ! negation is ignored—put each such rule in every source that must enforce it. Do not infer anything from the 51 omitted changelog items.

Proposed actions

  1. Run: claude plugin eval --help
  2. From each plugin repo this team maintains, run claude plugin eval against Claude Code v2.1.269 and store the JSON and HTML report next to the plugin as the scored baseline.
  3. Upgrade to Claude Code v2.1.269, then fully quit and start a new headless session and a new Claude Desktop session so organization plugins enabled through managed settings load.
  4. On every shared or multi-user machine, upgrade to Claude Code v2.1.269 and start a fresh session so plugin archives re-extract without other-user readability, world-writable bits, or leftover stale files.
  5. Resume a remote Claude Code session after upgrade and confirm each synced plugin MCP server connects; v2.1.269 fixed them not connecting on resume.

Agent prompt

Paste into your agent or query via MCP (get_agent_prompt) — free, no extra AI cost

Paste into Claude Code / CLAUDE.md task

DevAgentRadar → Claude Code

You are helping me adopt a real coding-assistant change. Work only from the facts below. Do not invent features.

Context

Assistant: Claude Code Proposal: Run claude plugin eval to score plugins with JSON and HTML reports Summary: Claude Code v2.1.269 adds claude plugin eval for scored, reproducible plugin results (JSON + HTML). Run it on every plugin you ship, then restart sessions so managed-settings org plugins load in headless/Desktop and plugin extracts are no longer world-readable. Primary source: https://github.com/anthropics/claude-code/releases/tag/v2.1.269

Why it matters

v2.1.269 adds claude plugin eval, which runs a plugin’s eval suite against Claude Code and returns scored, reproducible JSON and HTML. After upgrade, run it on every plugin this team maintains and keep those reports as the v2.1.269 baseline. The same release also changes plugin runtime behavior: organization plugins enabled through managed settings now load in headless sessions and on Claude Desktop from the next session (once Desktop bundles this CLI); synced plugin MCP servers connect when a remote session resumes; plugin archives extracted for a session are no longer readable by other local users, extracted files no longer keep world-writable bits, and stale files no longer survive re-extraction. Quit and start fresh sessions so those loads and re-extractions happen. This is also a breaking permissions release: a deny or ask rule starting with ! now applies only within the settings source that wrote it, and a bare ! negation is ignored—put each such rule in every source that must enforce it. Do not infer anything from the 51 omitted changelog items.

Suggested actions

  1. Run: claude plugin eval --help
  2. From each plugin repo this team maintains, run claude plugin eval against Claude Code v2.1.269 and store the JSON and HTML report next to the plugin as the scored baseline.
  3. Upgrade to Claude Code v2.1.269, then fully quit and start a new headless session and a new Claude Desktop session so organization plugins enabled through managed settings load.
  4. On every shared or multi-user machine, upgrade to Claude Code v2.1.269 and start a fresh session so plugin archives re-extract without other-user readability, world-writable bits, or leftover stale files.
  5. Resume a remote Claude Code session after upgrade and confirm each synced plugin MCP server connects; v2.1.269 fixed them not connecting on resume.

Config surfaces this release may change

  • skills (high confidence) — check your repo before applying
  • permission rules (high confidence) — check your repo before applying
  • background and headless runs — check your repo before applying
  • agent context files — check your repo before applying
  • MCP servers — check your repo before applying

After you finish

Do not report this as applied to DevAgentRadar. You cannot write the visitor's loop.

Tell the human: open https://devagentradar.com/proposals/claude-code-v2-1-269-run-claude-plugin-eval-to-score-plugins-with-json-a and mark Applied, Skipped, or Failed. Proposal id: 9bf33f8b-cf39-4868-92d8-9838132aeb7b

Your job

  1. Restate the change in one sentence.
  2. Propose a minimal plan for my repo (or a throwaway pilot).
  3. Implement only what I approve; prefer small diffs and tests.
  4. Call out risks (permissions, breaking APIs, cost).

Start by confirming you understood the proposal.

agentmcpmodelsecurityideRelease source ↗

Your loop

This browser · no sign-in · not shared as “you”

After you run the prompt

Only you can mark this. Agents cannot write your loop.

Your decision stays on this device. A public tally appears after a few votes.

Originating release signal

Claude Codev2.1.269Sep 11, 2026

v2.1.269

Added claude plugin eval: run a plugin's eval suite against Claude Code and get scored, reproducible results (JSON + HTML report); see claude plugin eval --help · Added /output-style [name] to list and switch output styles, including over Remote Control and in cloud and other headless sessions · Added a diff of the files a Bash command changed to the Bash tool result when the Bash tool handles file edits (setting bashEditDiffEnabled) · +103 more changes
Verified excerpt — the source's own words

What's changed

  • Added claude plugin eval: run a plugin's eval suite against Claude Code and get scored, reproducible results (JSON + HTML report); see claude plugin eval --help
  • Added /output-style [name] to list and switch output styles, including over Remote Control and in cloud and other headless sessions
  • Added a diff of the files a Bash command changed to the Bash tool result when the Bash tool handles file edits (setting bashEditDiffEnabled)
  • Added OTEL_METRICS_INCLUDE_REPOSITORY to tag OpenTelemetry metrics and events with vcs.* repository attributes; commit events get vcs.ref.head.* with OTEL_LOG_TOOL_DETAILS
  • Added CLAUDE_CODE_GATEWAY_MODEL_DISCOVERY_TIMEOUT_MS to extend the LLM gateway /v1/models discovery timeout (default 3s)
  • Added a spinner tip suggesting /focus for a view with just your prompt, a one-line work summary, and the response
  • Added CLAUDE_CODE_WORKFLOW_MAX_CONCURRENT_AGENTS (1–256) to raise the Workflow tool's per-run concurrent agent limit for inference-bound fan-outs
  • Fixed the prompt cache being partially invalidated on the turn after a response was cut off at the output-token limit and automatically resumed

Excerpt ends here — this release continues at the source ↗.

Primary source ↗