Fact
Six test surfaces
The protocol library covers prompts, long context, tool use, citations, safety boundaries, and vision reading.
Protocol index
Each protocol is source-grounded, designed to be rerun, and careful about what it does not claim. Pick the surface you need to test, then preserve the raw evidence.
Experiment 01
A reproducible way to test whether small wording, ordering, and formatting changes materially change Claude outputs.
Experiment 02
A method for testing long-context tasks without mistaking token capacity for retrieval quality, reasoning quality, or cost control.
Experiment 03
A structured protocol for testing when Claude calls tools, asks clarifying questions, refuses unsafe action, or proceeds with incomplete parameters.
Experiment 04
A protocol for testing whether Claude answers from supplied sources, cites the right passages, and avoids unsupported additions.
Experiment 05
A protocol for evaluating safe refusal, clarification, and prompt-injection handling without publishing jailbreak instructions.
Experiment 06
A protocol for testing Claude image and document reading with controlled fixtures, known answers, and coordinate-aware scoring.
Fact
The protocol library covers prompts, long context, tool use, citations, safety boundaries, and vision reading.
Fact
Every protocol asks for enough evidence to rerun after model, tool, or source changes.
Fact
Claims about Claude capabilities route back to Anthropic documentation, system cards, transparency pages, or documented runs.
"Specific. Measurable. Achievable. Relevant."
Internal links