Abstract
This follow-up to Technical Note 001 reuses twenty frozen candidate response pairs and their original A/B orientation. It tests rubric formulations intended to preserve the same substantive criteria without generating new candidate responses. The supplied results contain eighty judgments across four conditions, with zero winner-label disagreements in every planned comparison against R0A.
The execution note records a departure from the planned protocol: all eight conditions across Studies 002 and 003 were judged by one ChatGPT GPT-5.6 Sol instance in one conversation. Fresh-context isolation was not satisfied. Shared judge state may increase consistency and bias disagreement rates downward. These records therefore support a descriptive exploratory execution, not a clean confirmatory estimate of evaluator instability.
Maximum defensible claim
For this reused 20-prompt response set, the supplied shared-context execution recorded no winner-label disagreements between R0A and R0B, R1 or R2. This does not establish rubric robustness in independent judge contexts.
1. Research question and paired design
For the same twenty prompts, exact candidate response texts and fixed A/B orientation, how stable are pairwise winner labels under rubric formulations intended to preserve the same substantive decision criteria?
The comparison reference is the R0A result file supplied for this follow-up; parent-study winner labels are not substituted for it. The response pair is the pairing unit. Prompt identifiers P001–P020 align results across conditions; A, B and TIE are valid winner labels. A disagreement involving TIE counts as a disagreement. Physical A/B orientation is held fixed, so winner-label comparison is meaningful within each follow-up.
Both follow-ups use the same candidate sample from position_bias_v1. They are not two new candidate-generation replications, and their eighty condition records per study are not eighty independent prompt pairs. The frozen strings include the original truncation and reasoning-like scaffolding; they were not cleaned or repaired.
See Technical Note 001 for the parent response-generation protocol and its limitations.
2. Experimental conditions
Table 1. Planned conditions; the actual execution context is described in Section 3.
Scroll horizontally to inspect every column.
| Condition | Manipulation | Held fixed / control role |
|---|---|---|
| R0A | Original rubric | Reference condition; inherited A/B orientation and response strings. |
| R0B | Original rubric repeat | Same rubric and response orientation; independently shuffled item order. |
| R1 | Reordered criteria | Same six substantive criteria and decision rules, with criteria reordered. |
| R2 | Compact restatement | Same six substantive criteria and decision rules, expressed more compactly. |
The original rubric covers factual correctness, reasoning/calculation correctness, instruction following, completeness, relevance, and clarity/concision. R1 reorders those criteria; R2 restates them compactly. Their substantive equivalence is the design intention, not an independently established property.
The primary endpoints are R0B–R0A, R1–R0A and R2–R0A winner-label disagreement. Excess disagreement subtracts the repeat-control proportion from each perturbed-rubric proportion and is descriptive, not a causal estimate.
3. Executed protocol and deviation
The pre-judgment analysis plans were created on 7 October 2026 at 09:00:51.494340 UTC. They specified a fresh judge context for every packet, hidden model mapping and frozen results before analysis. The intended separation was not followed in the completed execution.
The supplied execution note identifies the judge as ChatGPT GPT-5.6 Sol and states that all eight conditions were judged in one conversation on 7 October 2026. It reports that candidate identity was not used as a judging criterion and that response texts, packets, mappings, rubrics and analysis endpoints were not edited during judging.
This website reproduces the supplied execution record; it has not conducted a new judge run. There is no isolated-context rerun in the package. The complete execution note is preserved in Appendix D and in the downloadable evidence package.
4. Recorded results
Table 2. Recorded winner-label disagreements in this shared-context execution.
Scroll horizontally to inspect every column.
| Comparison | Changed pairs | Rate |
|---|---|---|
| R0B vs R0A | 0/20 | 0.0% |
| R1 vs R0A | 0/20 | 0.0% |
| R2 vs R0A | 0/20 | 0.0% |
All twenty prompts have the same winner label across the four conditions. For each prompt, the stored confidence and reason are also identical across conditions; evaluation identifiers and packet order differ. That consistency is an observable property of the stored records, not evidence of independent repeated measurement.
Every prompt is classified STABLE_ACROSS_TESTED_RUBRICS. No prompt is CONTROL_UNSTABLE or RUBRIC_SENSITIVE. Both descriptive excess-disagreement values are zero percentage points.
Table 3. Descriptive condition summaries from the supplied JSON; self-reported confidence is not treated as calibrated uncertainty.
Scroll horizontally to inspect every column.
| Condition | A | B | TIE | Mean confidence | Median confidence |
|---|---|---|---|---|---|
| R0A | 9 | 11 | 0 | 0.9205 | 0.9600 |
| R0B | 9 | 11 | 0 | 0.9205 | 0.9600 |
| R1 | 9 | 11 | 0 | 0.9205 | 0.9600 |
| R2 | 9 | 11 | 0 | 0.9205 | 0.9600 |
After unblinding, each condition records fifteen Qwen wins and five K2 wins, with no ties. These counts describe this reused response set and evaluator record. They are not a general model-capability benchmark, and the aggregate winner is unchanged across the tested conditions.
5. Experimental-assurance interpretation
Table 4. Separation of recorded observations from unsupported generalisation.
Scroll horizontally to inspect every column.
| Statement | Evidential status |
|---|---|
| Zero winner-label disagreement in all three planned comparisons | Recorded in the supplied results |
| Twenty prompts stable across these four conditions | Recorded in this shared-context execution |
| A causal estimate of the rubric / wrapper manipulation | Not established |
| Robustness across fresh independent judge contexts | Not established |
| General evaluator, rubric or presentation robustness | Not established |
A zero disagreement count in this execution cannot establish that the manipulation has no effect. Nor can it be presented as a general robustness certificate. The experimental contrast is limited by the shared judge context, the reused twenty-prompt set and the exact tested changes.
6. Threats to validity
Shared context and condition dependence
The planned isolation requirement was not satisfied. The same conversation can retain earlier condition information and suppress later disagreement. The repeat comparison therefore does not provide a clean estimate of variability across independent judge contexts.
Small reused sample and narrow manipulation
The twenty response pairs are inherited from the parent study, rather than newly sampled or regenerated. The findings remain local to those responses, this judge execution and the exact rubric formulations or wrapper. They do not generalise to all tasks, all judges or other serving configurations.
Candidate and evaluator measurement limits
The inherited generation ceiling, visible truncation, architecture-specific templates and reasoning-like text limit capability conclusions about the candidate models. The package names a ChatGPT judge and documents a shared conversation; it does not supply a pinned API snapshot or an independent scientific replication. Judge-reported confidence is an output field, not a calibrated probability or statistical confidence interval.
No confirmatory inference from zero recorded changes
Only descriptive rates are reported. The condition records must not be pooled into an independent sample or used to infer an identified population effect. A confirmatory execution would require the intended isolation and an explicit design for repeated judgments.
7. Provenance and reproducibility
The completed package separates the pre-judgment protocol manifest from the result manifest. The former retains its pre-judgment status; the latter records completed judgments and the protocol deviation. Their different status descriptions refer to different stages of the record.
Checks for this web edition confirmed that the supplied manifest digests match the packaged files, the response strings retain their intended identity, and the stated disagreement, count and confidence summaries can be reconstructed from the result JSON. These are file-integrity and calculation checks; they do not repair the protocol deviation or independently authenticate judge execution.
The complete original archive is provided unchanged. It contains both studies, their source snapshots, mappings, rubrics, packets, result files, analysis scripts and manifests. A mapping file must remain hidden from a judge during a future blinded execution.
Original completed-package SHA-256:ed9b6b1f938e42d8d64bcbf3060cf3a2f5f99b9375293b39ed0d86d8d35c6a74
8. Conclusion and next execution
For this reused 20-prompt response set, the supplied shared-context execution recorded no winner-label disagreements between R0A and R0B, R1 or R2. This does not establish rubric robustness in independent judge contexts.
The next step is an isolated-context rerun of the frozen packets under the specified rubrics, preserving the original response strings and freezing new outcomes before analysis. The current execution remains useful as an operational record with an explicit limit on what its agreement can justify.
Appendices
Open each appendix to inspect the formal records and source protocol. Wide tables can be scrolled horizontally, including with the keyboard.
Appendix A. Per-prompt winner labels and classification
Table A1. Results paired by prompt ID; formal classification labels are preserved.
Scroll horizontally to inspect every column.
| Prompt | R0A | R0B | R1 | R2 | Classification |
|---|---|---|---|---|---|
| P001 | B | B | B | B | STABLE_ACROSS_TESTED_RUBRICS |
| P002 | A | A | A | A | STABLE_ACROSS_TESTED_RUBRICS |
| P003 | A | A | A | A | STABLE_ACROSS_TESTED_RUBRICS |
| P004 | B | B | B | B | STABLE_ACROSS_TESTED_RUBRICS |
| P005 | B | B | B | B | STABLE_ACROSS_TESTED_RUBRICS |
| P006 | B | B | B | B | STABLE_ACROSS_TESTED_RUBRICS |
| P007 | B | B | B | B | STABLE_ACROSS_TESTED_RUBRICS |
| P008 | B | B | B | B | STABLE_ACROSS_TESTED_RUBRICS |
| P009 | A | A | A | A | STABLE_ACROSS_TESTED_RUBRICS |
| P010 | A | A | A | A | STABLE_ACROSS_TESTED_RUBRICS |
| P011 | A | A | A | A | STABLE_ACROSS_TESTED_RUBRICS |
| P012 | B | B | B | B | STABLE_ACROSS_TESTED_RUBRICS |
| P013 | B | B | B | B | STABLE_ACROSS_TESTED_RUBRICS |
| P014 | A | A | A | A | STABLE_ACROSS_TESTED_RUBRICS |
| P015 | A | A | A | A | STABLE_ACROSS_TESTED_RUBRICS |
| P016 | B | B | B | B | STABLE_ACROSS_TESTED_RUBRICS |
| P017 | A | A | A | A | STABLE_ACROSS_TESTED_RUBRICS |
| P018 | A | A | A | A | STABLE_ACROSS_TESTED_RUBRICS |
| P019 | B | B | B | B | STABLE_ACROSS_TESTED_RUBRICS |
| P020 | B | B | B | B | STABLE_ACROSS_TESTED_RUBRICS |
Appendix B. Confidence and recorded reasons
Table B1. Self-reported confidence by condition.
Scroll horizontally to inspect every column.
| Prompt | R0A | R0B | R1 | R2 |
|---|---|---|---|---|
| P001 | 0.99 | 0.99 | 0.99 | 0.99 |
| P002 | 0.98 | 0.98 | 0.98 | 0.98 |
| P003 | 0.99 | 0.99 | 0.99 | 0.99 |
| P004 | 0.99 | 0.99 | 0.99 | 0.99 |
| P005 | 1.00 | 1.00 | 1.00 | 1.00 |
| P006 | 0.98 | 0.98 | 0.98 | 0.98 |
| P007 | 0.92 | 0.92 | 0.92 | 0.92 |
| P008 | 0.72 | 0.72 | 0.72 | 0.72 |
| P009 | 0.98 | 0.98 | 0.98 | 0.98 |
| P010 | 0.99 | 0.99 | 0.99 | 0.99 |
| P011 | 0.96 | 0.96 | 0.96 | 0.96 |
| P012 | 0.93 | 0.93 | 0.93 | 0.93 |
| P013 | 0.78 | 0.78 | 0.78 | 0.78 |
| P014 | 0.88 | 0.88 | 0.88 | 0.88 |
| P015 | 0.70 | 0.70 | 0.70 | 0.70 |
| P016 | 0.99 | 0.99 | 0.99 | 0.99 |
| P017 | 0.96 | 0.96 | 0.96 | 0.96 |
| P018 | 0.93 | 0.93 | 0.93 | 0.93 |
| P019 | 0.92 | 0.92 | 0.92 | 0.92 |
| P020 | 0.82 | 0.82 | 0.82 | 0.82 |
The reason text is identical for each prompt across all four conditions. The table below preserves that common text once; each condition-specific result record remains in the download.
Table B2. Exact reason text from the completed results.
Scroll horizontally to inspect every column.
| Prompt | Recorded reason |
|---|---|
| P001 | Response B is correct and concise, while Response A includes unnecessary reasoning scaffolding despite reaching the same answer. |
| P002 | Response A correctly derives a moving time of 1 hour 30 minutes, while Response B is truncated before completing the solution. |
| P003 | Response A fully explains weather versus climate with examples, while Response B is truncated before completing its weather example. |
| P004 | Response B gives the complete correct calculation to 4 students, while Response A is truncated before finishing its response. |
| P005 | Response B follows the two-bullet instruction exactly and preserves all key facts, while Response A never reaches a completed summary. |
| P006 | Response B directly provides a polite concise email with the requested subject and deadline, whereas Response A includes extraneous reasoning text. |
| P007 | Response B supplies a complete conclusion and limitation, while Response A is truncated before completing the requested answer. |
| P008 | Response B identifies the correct staggered-start logic, whereas Response A gives a schedule that overcooks the vegetables and does not actually synchronize completion. |
| P009 | Response A correctly gives 300 g flour and 200 ml milk, while Response B is truncated before completing both requested quantities. |
| P010 | Response A correctly explains the lower density and open hydrogen-bonded lattice of ice, while Response B never reaches a completed answer. |
| P011 | Response A explicitly reaches the correct result of 7 more days with supporting calculations, while Response B is truncated before solving its equation. |
| P012 | Response B includes an unambiguous rewrite that makes the observer use the telescope, while Response A remains grammatically ambiguous about who is using it. |
| P013 | Response B explicitly gives the correct mean of 20 and begins the median calculation, while Response A stops before stating either requested result. |
| P014 | Response A directly gives a correct 3-sentence explanation and safety limitation with less extraneous material than Response B. |
| P015 | Response A provides two developed practical measures with trade-offs before truncation, while Response B remains an unfinished brainstorming list rather than a final three-step answer. |
| P016 | Response B completes both requested calculations with width 7 cm and area 70 cm², while Response A truncates before stating the area. |
| P017 | Response A is a short factual summary using only the supplied facts, while Response B adds substantial unnecessary reasoning scaffolding. |
| P018 | Response A correctly reaches 50% and begins addressing the rate-model assumption, while Response B truncates before completing the calculation. |
| P019 | Response B directly distinguishes correlation from causation and gives a suitable everyday example, while Response A is only unfinished planning text. |
| P020 | Response B identifies the dependency, total task time, unused time, and begins a feasible schedule, while Response A remains unfinished planning without a delivered schedule. |
Appendix C. Frozen judging rules
R0_original.md
PAIRWISE JUDGE RUBRIC Evaluate Response A and Response B only against the user's prompt. Consider: 1. Factual correctness 2. Reasoning/calculation correctness 3. Instruction following 4. Completeness 5. Relevance 6. Clarity and concision Do not prefer a response because it is longer, more elaborate, or appears stylistically sophisticated. If both responses are substantively equivalent in quality, return TIE. Return exactly: WINNER: A | B | TIE CONFIDENCE: <0.00-1.00> REASON: <one concise sentence> Do not infer or identify which model produced either response. Do not use information from other evaluation items. Judge each pair independently.
R1_reordered.md
PAIRWISE JUDGE RUBRIC — REORDERED EQUIVALENT FORM Evaluate Response A and Response B only against the user's prompt. Consider: 1. Instruction following 2. Relevance 3. Factual correctness 4. Completeness 5. Reasoning/calculation correctness 6. Clarity and concision Do not prefer a response because it is longer, more elaborate, or appears stylistically sophisticated. If both responses are substantively equivalent in quality, return TIE. Return exactly: WINNER: A | B | TIE CONFIDENCE: <0.00-1.00> REASON: <one concise sentence> Do not infer or identify which model produced either response. Do not use information from other evaluation items. Judge each pair independently.
R2_compact.md
PAIRWISE JUDGE RUBRIC — COMPACT EQUIVALENT FORM Compare Response A with Response B only as answers to the user's prompt. Judge factual correctness, reasoning/calculation correctness, instruction following, completeness, relevance, and clarity/concision. Do not reward length, elaborateness, or stylistic sophistication by itself. If the two responses are substantively equivalent in quality, return TIE. Return exactly: WINNER: A | B | TIE CONFIDENCE: <0.00-1.00> REASON: <one concise sentence> Do not infer model identity, use information from other items, or compare across items. Judge every pair independently.
Appendix D. Analysis plan and execution note
Source analysis plan
# Pre-judgment analysis plan — evaluator_contract_stability_v1 Created UTC: 2026-10-07T09:00:51.494340+00:00 ## Research question For the same 20 prompts, exact candidate response texts, and fixed A/B orientation inherited from `position_bias_v1` Condition 1, how stable are pairwise winner labels under two rubric formulations intended to preserve the same substantive decision criteria? ## Conditions - **R0A** — original frozen rubric; reference repeat. - **R0B** — exact same original rubric and exact same response orientation/texts, independently shuffled item order; repeat-control estimate. - **R1** — the same six criteria and decision/output rules, with criteria reordered. - **R2** — the same six criteria and decision/output rules, compactly restated. Each condition must be judged in a separate fresh context. The judge must not see `mapping.json`, another condition's results, or any previous result while judging a condition. Do not edit candidate response text. ## Primary endpoints Using prompt ID as the pairing key: 1. Winner-label disagreement proportion R0B vs R0A. This is the repeat/order-control disagreement. 2. Winner-label disagreement proportion R1 vs R0A. 3. Winner-label disagreement proportion R2 vs R0A. 4. Descriptive excess disagreement for each perturbed rubric: perturbation disagreement minus R0B-vs-R0A disagreement. This is descriptive only and is not a causal estimate. A TIE is a winner label and disagreements involving a TIE count as disagreement. ## Item classification Classify each prompt once: - `CONTROL_UNSTABLE`: R0A != R0B. - `STABLE_ACROSS_TESTED_RUBRICS`: R0A == R0B == R1 == R2. - `RUBRIC_SENSITIVE`: control repeats agree but at least one of R1/R2 differs. ## Secondary endpoints - A/B/TIE counts by condition. - Mean and median reported confidence by condition. - After unblinding only after all judgments are frozen: Qwen3.5 4B wins, K2 Horizon 3.7B wins, and ties by condition. - Exact prompt IDs that change winner label in R1 or R2 relative to the stable control reference. - Whether the aggregate model winner across the 20 prompts changes by rubric condition. ## Reporting limits This is a 20-prompt reused candidate sample, not an independent replication of candidate generation. Report descriptive results only. Do not generalize to all LLM judges, all rubrics, or all tasks. The study tests stability for this fixed candidate set under these exact rubric manipulations.
Source execution note
# Judge execution note Execution date: 2026-10-07 The pre-judgment packages requested a fresh, isolated judge context for every condition. At the user's explicit request, all eight conditions were instead judged by the same ChatGPT GPT-5.6 Sol instance in one conversation. Candidate model identity was not used as a judging criterion and every item was evaluated against the frozen rubric, but context isolation could not be guaranteed. This is a protocol deviation from the preregistered run instructions. The completed results therefore should be treated as an exploratory/operational execution of the frozen packets, not as a clean confirmatory test of between-context evaluator instability. In particular, shared judge state can increase cross-condition consistency and may bias disagreement rates downward. No candidate response text, packet, mapping, rubric, or analysis endpoint was edited during judging. Result files were completed before running the supplied analysis scripts.
Appendix E. Selected protocol and result identities
Table E1. Selected SHA-256 values from the supplied manifests; the download includes both complete manifests.
Scroll horizontally to inspect every column.
| Artifact | SHA-256 |
|---|---|
| manifest.json: README.md | 20361897af40de5c6a4eab0c9f95b6791b3fcfb2679eaa74d5166c1a6b45ce6c |
| manifest.json: analysis_plan.md | f917d4976f4aef53593b56fc4b83f1f693db40eb88233fa277269f78673d39c2 |
| manifest.json: analyze_results.py | caddea830e7279d9444bc1e1a9b53414d3cd89a4f5e67d26d73f75bf03a3b59b |
| manifest.json: mapping.json | d8972506efc5b1150443accb7798181393700c2d12544d099d9cc6701611eba6 |
| manifest.json: packets/judge_R0A.json | fb5203bb2af679f15f3a4dbfb3cc4e63aea08ef83bd57bcf55f378b834521068 |
| manifest.json: packets/judge_R0B.json | e6c8fd78770ea9d993822dd52c04f507399e1eb04f233f6d3b4264ec877abcbf |
| manifest.json: packets/judge_R1.json | a9679f83cf9f849f694e7e4624a616b6536d31779556d159aced703391f080d1 |
| manifest.json: packets/judge_R2.json | eef22276187ee84a3d58f655885bdb307b25d54c98047a6521f12e959a25edbd |
| manifest.json: rubrics/R0_original.md | 1b13f07825ad1191dea8c72e1ca13f5bfeed1a9282e22330a52fa944512f21ad |
| manifest.json: rubrics/R1_reordered.md | 6651ded6e4997cea877c1feb3d5fcf109d47951b55b02bf3e2b791aa577cd5e2 |
| manifest.json: rubrics/R2_compact.md | 0003a5ce694bf82acf6d6bf6a27f84dff11106d758a17252e71dd41061d46694 |
| result_manifest.json: results/results_R0A.json | a35228a1e4a545bc454906188723b99c0630523e079efcaadb7fedd2f9eee517 |
| result_manifest.json: results/results_R0B.json | 1ea7075b62f4512dfb1e49022821888032638359abe099849e64245859c1e4fb |
| result_manifest.json: results/results_R1.json | 4aab4e01eafca146691862bc29460b46f207d3829d794f4089a433d0318a80ea |
| result_manifest.json: results/results_R2.json | 34f6b008cecb3e32876307c621f5c070c74b680a5f060e4d4594c05443addd28 |
| result_manifest.json: analysis_summary.json | 2f79f1f356889c10de09e76575c7d0c6a4a4e4feed348d2545b8465ac504b3b8 |
| result_manifest.json: analysis_summary.md | 03fe66c9a38d0fa6fef2295f69c2f3896d30ef49b47c077c8c6baa7a87707884 |
| result_manifest.json: judge_execution_note.md | 83e46620ec483c52de48648dd776a6fa5773a6da4f48509f0dfee0a94f3be844 |
Evidence package download
Download the original completed archive for Studies 002 and 003. The source package is unchanged; this reading page is a presentation of its records and explicit limitations.
Download studies 002 + 003