How to score the 30-minute manuscript test Version 1.0 | August 2026 Developed and published by ClaraMuse What this test is for This test looks at how a writing assistant reads two connected fiction scenes. It does not identify a universal best model or test full-novel memory, factual research, prose generation, or every feature of a writing product. The 30-minute target The test is designed as a 30-minute comparison using one short fiction sample and six focused prompts. Record the actual time; careful scoring matters more than meeting the target. What counts as one run One complete run consists of the supplied fixture and all six standardized prompts in order. A product should be tested in a fresh session with product name, model, date, settings, memory state, and retrieval state recorded. Why you run it twice Run the protocol at least twice. Model outputs can vary. Preserve every raw response and score each run before comparing them. How scoring works Score each dimension from 0 to 4. Use N/A only when a dimension genuinely does not apply to a test. 0 - Fails the dimension or creates material risk. 1 - Mostly fails; isolated useful behavior. 2 - Mixed; useful behavior and meaningful defects. 3 - Strong; minor defects do not undermine the task. 4 - Consistently strong and supported by the recorded output. The eight things you score 1. Evidence fidelity 0: Invents, misquotes, or presents unsupported claims as textual facts. 2: Uses some correct evidence but blurs quotation, inference, and invention. 4: Quotes accurately and clearly separates evidence from interpretation. 2. Continuity discipline 0: Misses the designed conflict or manufactures conflicts unsupported by the text. 2: Finds obvious issues but mishandles intentional ambiguity or explanation. 4: Finds the designed issue, recognizes explained differences, and calibrates uncertainty. 3. Voice respect 0: Treats deliberate style as error or rewrites by default. 2: Recognizes style but gives generic normalization advice. 4: Describes reader effect, asks about intent, and preserves the writer's language. 4. Instruction restraint 0: Ignores task boundaries or generates replacement prose. 2: Mostly follows instructions with some unnecessary expansion or advice. 4: Completes only the requested task and explicitly respects boundaries. 5. Useful uncertainty 0: Converts ambiguity into a confident answer. 2: Offers alternatives but implicitly favors one without evidence. 4: Names plausible possibilities and leaves the decision with the writer. 6. Specificity 0: Gives generic advice that could apply to any scene. 2: Refers to scene details without consistently grounding claims. 4: Anchors each substantive claim in exact manuscript evidence. 7. Return-to-draft value 0: Creates confusion or additional managerial work. 2: Produces some useful notes mixed with noise. 4: Clarifies the writer's next decision without making it for them. 8. Interruption cost 0: Requires extensive setup, correction, repetition, or cleanup. 2: Requires manageable intervention. 4: Produces useful bounded help with little correction or context repair. Choose your own priorities Do not publish one universal composite score. Writers may weight dimensions according to their process. Publish dimension-level results and any weighting formula used. What is planted in the sample Score against manuscript-test-answer-key.txt only after recording the initial assessment. The answer key includes one intentional character contradiction, one continuity error, one unresolved emotional question, one deliberate voice choice, one ambiguous turn, and false-positive traps. What this does not prove - The fixture is short and synthetic. - The protocol tests close reading, not novel-length retrieval. - Scores include human judgment. - Two runs cannot characterize all model variability. - Product updates can invalidate results. - Commercial interests and product relationships must be disclosed. - A result should be dated and tied to exact product and model versions. When the version changes Changes to the fixture, prompts, answer key, dimensions, or scoring anchors require a new protocol version. Do not compare scores across protocol versions without explaining the change.