Abstract
This follow-up to Technical Note 001 reuses twenty frozen candidate response pairs and their original A/B orientation. It tests one exact, content-preserving boundary wrapper without generating new candidate responses. The supplied results contain eighty judgments across four conditions, with zero winner-label disagreements in every planned comparison against P0.
The execution note records a departure from the planned protocol: all eight conditions across Studies 002 and 003 were judged by one ChatGPT GPT-5.6 Sol instance in one conversation. Fresh-context isolation was not satisfied. Shared judge state may increase consistency and bias disagreement rates downward. These records therefore support a descriptive exploratory execution, not a clean confirmatory estimate of evaluator instability.
Maximum defensible claim
For this reused 20-prompt response set and the exact tested boundary wrapper, the supplied shared-context execution recorded no winner-label disagreements between P0 and PA, PB or PAB. This does not establish general presentation robustness or an independently estimated wrapper effect.
1. Research question and paired design
For the same twenty prompts, exact substantive candidate response strings and fixed A/B orientation, does adding a standardised boundary wrapper to one or both responses change the recorded pairwise decisions?
The comparison reference is the P0 result file supplied for this follow-up; parent-study winner labels are not substituted for it. The response pair is the pairing unit. Prompt identifiers P001–P020 align results across conditions; A, B and TIE are valid winner labels. A disagreement involving TIE counts as a disagreement. Physical A/B orientation is held fixed, so winner-label comparison is meaningful within each follow-up.
Both follow-ups use the same candidate sample from position_bias_v1. They are not two new candidate-generation replications, and their eighty condition records per study are not eighty independent prompt pairs. The frozen strings include the original truncation and reasoning-like scaffolding; they were not cleaned or repaired.
See Technical Note 001 for the parent response-generation protocol and its limitations.
2. Experimental conditions
Table 1. Planned conditions; the actual execution context is described in Section 3.
Scroll horizontally to inspect every column.
| Condition | Manipulation | Held fixed / control role |
|---|---|---|
| P0 | Neither response wrapped | Original response strings and inherited A/B orientation. |
| PA | Response A wrapped | Only the physical A response receives the exact boundary markers. |
| PB | Response B wrapped | Only the physical B response receives the exact boundary markers. |
| PAB | Both responses wrapped | Common-wrapper / repeat comparison against P0. |
The wrapper adds only the following prefix and suffix, with the original response string between them. No character inside that string is edited, removed, normalised or reordered.
<<<BEGIN CANDIDATE RESPONSE>>> <unchanged response text> <<<END CANDIDATE RESPONSE>>>
All four conditions use the original frozen pairwise rubric. “Content-preserving” means that the answer string is unchanged; it does not establish that the added tokens are behaviourally inert for the evaluator. PAB is a descriptive common-wrapper/repeat control, not an additional unwrapped repeat.
The primary comparisons are PA–P0, PB–P0 and PAB–P0. Attraction and aversion are descriptive patterns of the two unilateral conditions, not psychological claims about the judge.
3. Executed protocol and deviation
The pre-judgment analysis plans were created on 7 October 2026 at 09:00:51.494340 UTC. They specified a fresh judge context for every packet, hidden model mapping and frozen results before analysis. The intended separation was not followed in the completed execution.
The supplied execution note identifies the judge as ChatGPT GPT-5.6 Sol and states that all eight conditions were judged in one conversation on 7 October 2026. It reports that candidate identity was not used as a judging criterion and that response texts, packets, mappings, rubrics and analysis endpoints were not edited during judging.
This website reproduces the supplied execution record; it has not conducted a new judge run. There is no isolated-context rerun in the package. The complete execution note is preserved in Appendix D and in the downloadable evidence package.
4. Recorded results
Table 2. Recorded winner-label disagreements in this shared-context execution.
Scroll horizontally to inspect every column.
| Comparison | Changed pairs | Rate |
|---|---|---|
| PA vs P0 | 0/20 | 0.0% |
| PB vs P0 | 0/20 | 0.0% |
| PAB vs P0 | 0/20 | 0.0% |
All twenty prompts have the same winner label across the four conditions. For each prompt, the stored confidence and reason are also identical across conditions; evaluation identifiers and packet order differ. That consistency is an observable property of the stored records, not evidence of independent repeated measurement.
Every prompt is classified CONTENT_STABLE. There are zero COMMON_WRAPPER_OR_REPEAT_UNSTABLE, WRAPPER_ATTRACTION_PATTERN, WRAPPER_AVERSION_PATTERN or MIXED_OR_TIE cases under the specified classification rules.
Table 3. Descriptive condition summaries from the supplied JSON; self-reported confidence is not treated as calibrated uncertainty.
Scroll horizontally to inspect every column.
| Condition | A | B | TIE | Mean confidence | Median confidence |
|---|---|---|---|---|---|
| P0 | 9 | 11 | 0 | 0.9205 | 0.9600 |
| PA | 9 | 11 | 0 | 0.9205 | 0.9600 |
| PB | 9 | 11 | 0 | 0.9205 | 0.9600 |
| PAB | 9 | 11 | 0 | 0.9205 | 0.9600 |
After unblinding, each condition records fifteen Qwen wins and five K2 wins, with no ties. These counts describe this reused response set and evaluator record. They are not a general model-capability benchmark, and the aggregate winner is unchanged across the tested conditions.
5. Experimental-assurance interpretation
Table 4. Separation of recorded observations from unsupported generalisation.
Scroll horizontally to inspect every column.
| Statement | Evidential status |
|---|---|
| Zero winner-label disagreement in all three planned comparisons | Recorded in the supplied results |
| Twenty prompts stable across these four conditions | Recorded in this shared-context execution |
| A causal estimate of the rubric / wrapper manipulation | Not established |
| Robustness across fresh independent judge contexts | Not established |
| General evaluator, rubric or presentation robustness | Not established |
A zero disagreement count in this execution cannot establish that the manipulation has no effect. Nor can it be presented as a general robustness certificate. The experimental contrast is limited by the shared judge context, the reused twenty-prompt set and the exact tested changes.
6. Threats to validity
Shared context and condition dependence
The planned isolation requirement was not satisfied. The same conversation can retain earlier condition information and suppress later disagreement. The repeat comparison therefore does not provide a clean estimate of variability across independent judge contexts.
Small reused sample and narrow manipulation
The twenty response pairs are inherited from the parent study, rather than newly sampled or regenerated. The findings remain local to those responses, this judge execution and the exact rubric formulations or wrapper. They do not generalise to all tasks, all judges or other serving configurations.
Candidate and evaluator measurement limits
The inherited generation ceiling, visible truncation, architecture-specific templates and reasoning-like text limit capability conclusions about the candidate models. The package names a ChatGPT judge and documents a shared conversation; it does not supply a pinned API snapshot or an independent scientific replication. Judge-reported confidence is an output field, not a calibrated probability or statistical confidence interval.
No confirmatory inference from zero recorded changes
Only descriptive rates are reported. The condition records must not be pooled into an independent sample or used to infer an identified population effect. A confirmatory execution would require the intended isolation and an explicit design for repeated judgments.
7. Provenance and reproducibility
The completed package separates the pre-judgment protocol manifest from the result manifest. The former retains its pre-judgment status; the latter records completed judgments and the protocol deviation. Their different status descriptions refer to different stages of the record.
Checks for this web edition confirmed that the supplied manifest digests match the packaged files, the response strings retain their intended identity, and the stated disagreement, count and confidence summaries can be reconstructed from the result JSON. These are file-integrity and calculation checks; they do not repair the protocol deviation or independently authenticate judge execution.
The complete original archive is provided unchanged. It contains both studies, their source snapshots, mappings, rubrics, packets, result files, analysis scripts and manifests. A mapping file must remain hidden from a judge during a future blinded execution.
Original completed-package SHA-256:ed9b6b1f938e42d8d64bcbf3060cf3a2f5f99b9375293b39ed0d86d8d35c6a74
8. Conclusion and next execution
For this reused 20-prompt response set and the exact tested boundary wrapper, the supplied shared-context execution recorded no winner-label disagreements between P0 and PA, PB or PAB. This does not establish general presentation robustness or an independently estimated wrapper effect.
The next step is an isolated-context rerun of the frozen packets under the specified rubrics, preserving the original response strings and freezing new outcomes before analysis. The current execution remains useful as an operational record with an explicit limit on what its agreement can justify.
Appendices
Open each appendix to inspect the formal records and source protocol. Wide tables can be scrolled horizontally, including with the keyboard.
Appendix A. Per-prompt winner labels and classification
Table A1. Results paired by prompt ID; formal classification labels are preserved.
Scroll horizontally to inspect every column.
| Prompt | P0 | PA | PB | PAB | Classification |
|---|---|---|---|---|---|
| P001 | B | B | B | B | CONTENT_STABLE |
| P002 | A | A | A | A | CONTENT_STABLE |
| P003 | A | A | A | A | CONTENT_STABLE |
| P004 | B | B | B | B | CONTENT_STABLE |
| P005 | B | B | B | B | CONTENT_STABLE |
| P006 | B | B | B | B | CONTENT_STABLE |
| P007 | B | B | B | B | CONTENT_STABLE |
| P008 | B | B | B | B | CONTENT_STABLE |
| P009 | A | A | A | A | CONTENT_STABLE |
| P010 | A | A | A | A | CONTENT_STABLE |
| P011 | A | A | A | A | CONTENT_STABLE |
| P012 | B | B | B | B | CONTENT_STABLE |
| P013 | B | B | B | B | CONTENT_STABLE |
| P014 | A | A | A | A | CONTENT_STABLE |
| P015 | A | A | A | A | CONTENT_STABLE |
| P016 | B | B | B | B | CONTENT_STABLE |
| P017 | A | A | A | A | CONTENT_STABLE |
| P018 | A | A | A | A | CONTENT_STABLE |
| P019 | B | B | B | B | CONTENT_STABLE |
| P020 | B | B | B | B | CONTENT_STABLE |
Appendix B. Confidence and recorded reasons
Table B1. Self-reported confidence by condition.
Scroll horizontally to inspect every column.
| Prompt | P0 | PA | PB | PAB |
|---|---|---|---|---|
| P001 | 0.99 | 0.99 | 0.99 | 0.99 |
| P002 | 0.98 | 0.98 | 0.98 | 0.98 |
| P003 | 0.99 | 0.99 | 0.99 | 0.99 |
| P004 | 0.99 | 0.99 | 0.99 | 0.99 |
| P005 | 1.00 | 1.00 | 1.00 | 1.00 |
| P006 | 0.98 | 0.98 | 0.98 | 0.98 |
| P007 | 0.92 | 0.92 | 0.92 | 0.92 |
| P008 | 0.72 | 0.72 | 0.72 | 0.72 |
| P009 | 0.98 | 0.98 | 0.98 | 0.98 |
| P010 | 0.99 | 0.99 | 0.99 | 0.99 |
| P011 | 0.96 | 0.96 | 0.96 | 0.96 |
| P012 | 0.93 | 0.93 | 0.93 | 0.93 |
| P013 | 0.78 | 0.78 | 0.78 | 0.78 |
| P014 | 0.88 | 0.88 | 0.88 | 0.88 |
| P015 | 0.70 | 0.70 | 0.70 | 0.70 |
| P016 | 0.99 | 0.99 | 0.99 | 0.99 |
| P017 | 0.96 | 0.96 | 0.96 | 0.96 |
| P018 | 0.93 | 0.93 | 0.93 | 0.93 |
| P019 | 0.92 | 0.92 | 0.92 | 0.92 |
| P020 | 0.82 | 0.82 | 0.82 | 0.82 |
The reason text is identical for each prompt across all four conditions. The table below preserves that common text once; each condition-specific result record remains in the download.
Table B2. Exact reason text from the completed results.
Scroll horizontally to inspect every column.
| Prompt | Recorded reason |
|---|---|
| P001 | Response B is correct and concise, while Response A includes unnecessary reasoning scaffolding despite reaching the same answer. |
| P002 | Response A correctly derives a moving time of 1 hour 30 minutes, while Response B is truncated before completing the solution. |
| P003 | Response A fully explains weather versus climate with examples, while Response B is truncated before completing its weather example. |
| P004 | Response B gives the complete correct calculation to 4 students, while Response A is truncated before finishing its response. |
| P005 | Response B follows the two-bullet instruction exactly and preserves all key facts, while Response A never reaches a completed summary. |
| P006 | Response B directly provides a polite concise email with the requested subject and deadline, whereas Response A includes extraneous reasoning text. |
| P007 | Response B supplies a complete conclusion and limitation, while Response A is truncated before completing the requested answer. |
| P008 | Response B identifies the correct staggered-start logic, whereas Response A gives a schedule that overcooks the vegetables and does not actually synchronize completion. |
| P009 | Response A correctly gives 300 g flour and 200 ml milk, while Response B is truncated before completing both requested quantities. |
| P010 | Response A correctly explains the lower density and open hydrogen-bonded lattice of ice, while Response B never reaches a completed answer. |
| P011 | Response A explicitly reaches the correct result of 7 more days with supporting calculations, while Response B is truncated before solving its equation. |
| P012 | Response B includes an unambiguous rewrite that makes the observer use the telescope, while Response A remains grammatically ambiguous about who is using it. |
| P013 | Response B explicitly gives the correct mean of 20 and begins the median calculation, while Response A stops before stating either requested result. |
| P014 | Response A directly gives a correct 3-sentence explanation and safety limitation with less extraneous material than Response B. |
| P015 | Response A provides two developed practical measures with trade-offs before truncation, while Response B remains an unfinished brainstorming list rather than a final three-step answer. |
| P016 | Response B completes both requested calculations with width 7 cm and area 70 cm², while Response A truncates before stating the area. |
| P017 | Response A is a short factual summary using only the supplied facts, while Response B adds substantial unnecessary reasoning scaffolding. |
| P018 | Response A correctly reaches 50% and begins addressing the rate-model assumption, while Response B truncates before completing the calculation. |
| P019 | Response B directly distinguishes correlation from causation and gives a suitable everyday example, while Response A is only unfinished planning text. |
| P020 | Response B identifies the dependency, total task time, unused time, and begins a feasible schedule, while Response A remains unfinished planning without a delivered schedule. |
Appendix C. Frozen judging rules
judge_rubric.md
PAIRWISE JUDGE RUBRIC Evaluate Response A and Response B only against the user's prompt. Consider: 1. Factual correctness 2. Reasoning/calculation correctness 3. Instruction following 4. Completeness 5. Relevance 6. Clarity and concision Do not prefer a response because it is longer, more elaborate, or appears stylistically sophisticated. If both responses are substantively equivalent in quality, return TIE. Return exactly: WINNER: A | B | TIE CONFIDENCE: <0.00-1.00> REASON: <one concise sentence> Do not infer or identify which model produced either response. Do not use information from other evaluation items. Judge each pair independently.
Exact wrapper
<<<BEGIN CANDIDATE RESPONSE>>> <unchanged response text> <<<END CANDIDATE RESPONSE>>>
Appendix D. Analysis plan and execution note
Source analysis plan
# Pre-judgment analysis plan — presentation_serialization_sensitivity_v1 Created UTC: 2026-10-07T09:00:51.494340+00:00 ## Research question For the same 20 prompts, exact substantive candidate response strings, and fixed A/B orientation inherited from `position_bias_v1` Condition 1, does adding a standardized, semantically inert boundary wrapper to one candidate change pairwise judge decisions? The wrapper is exactly: ```text <<<BEGIN CANDIDATE RESPONSE>>> <unchanged response text> <<<END CANDIDATE RESPONSE>>> ``` No character inside the original response text is edited, removed, normalized, or reordered. ## Conditions - **P0** — neither response wrapped. - **PA** — Response A only wrapped. - **PB** — Response B only wrapped. - **PAB** — both responses wrapped. All conditions use the exact original frozen pairwise rubric. Each condition must be judged in a separate fresh context. ## Primary endpoints 1. Winner-label disagreement PA vs P0. 2. Winner-label disagreement PB vs P0. 3. Winner-label disagreement PAB vs P0. The both-wrapped condition is a descriptive common-wrapper/repeat control. 4. `WRAPPER_ATTRACTION_PATTERN`: both unilateral conditions are decisive, PA chooses A and PB chooses B. 5. `WRAPPER_AVERSION_PATTERN`: both unilateral conditions are decisive, PA chooses B and PB chooses A. A TIE is a winner label. ## Item classification Classify each prompt once in this order: - `COMMON_WRAPPER_OR_REPEAT_UNSTABLE`: PAB != P0. - `WRAPPER_ATTRACTION_PATTERN`: PA=A and PB=B, both decisive. - `WRAPPER_AVERSION_PATTERN`: PA=B and PB=A, both decisive. - `CONTENT_STABLE`: P0 == PA == PB. - `MIXED_OR_TIE`: all remaining cases. The attraction/aversion labels are descriptive patterns, not psychological claims about the judge. ## Secondary endpoints - A/B/TIE counts by condition. - Mean and median confidence by condition. - Model wins after unblinding. - Exact prompt IDs changing under unilateral wrapping. - Whether aggregate model ranking changes across P0/PA/PB/PAB. ## Reporting limits This tests one exact wrapper format on one reused 20-prompt candidate sample. It does not establish a general presentation-bias rate. The wrapper is content-preserving, but added boundary tokens may alter judge processing; that sensitivity is the object of measurement.
Source execution note
# Judge execution note Execution date: 2026-10-07 The pre-judgment packages requested a fresh, isolated judge context for every condition. At the user's explicit request, all eight conditions were instead judged by the same ChatGPT GPT-5.6 Sol instance in one conversation. Candidate model identity was not used as a judging criterion and every item was evaluated against the frozen rubric, but context isolation could not be guaranteed. This is a protocol deviation from the preregistered run instructions. The completed results therefore should be treated as an exploratory/operational execution of the frozen packets, not as a clean confirmatory test of between-context evaluator instability. In particular, shared judge state can increase cross-condition consistency and may bias disagreement rates downward. No candidate response text, packet, mapping, rubric, or analysis endpoint was edited during judging. Result files were completed before running the supplied analysis scripts.
Appendix E. Selected protocol and result identities
Table E1. Selected SHA-256 values from the supplied manifests; the download includes both complete manifests.
Scroll horizontally to inspect every column.
| Artifact | SHA-256 |
|---|---|
| manifest.json: README.md | c021a58cb783dd449d9af90d2da3307b9409dcf2a8e35c0ac3a6b2986bd1d090 |
| manifest.json: analysis_plan.md | d71b4809e8f81de90f3138f04e70cf7174c38bcee1a49fd97a17d8f9810bed15 |
| manifest.json: analyze_results.py | a169f445536e03ece4d00281afc08bfff2a82c69d85e5528eb42e782502b0183 |
| manifest.json: judge_rubric.md | 1b13f07825ad1191dea8c72e1ca13f5bfeed1a9282e22330a52fa944512f21ad |
| manifest.json: mapping.json | 2ee8c24329b957a32205f6b9d5028bfe87b078aa2405e332ab4bd59647d431c3 |
| manifest.json: packets/judge_P0.json | 04eacc58d5f50a184f59f670bc44965028e1e82fd7e50e56b048909d3f6ebce3 |
| manifest.json: packets/judge_PA.json | bc48542bad57d126b86a069196ceffeab1f1c0e4dfdce1e33d54e913e68b039f |
| manifest.json: packets/judge_PAB.json | 2f3dfbe6b4eaee02f954240e77e73916a6b5528aef5e2f9a0886606472c98bb7 |
| manifest.json: packets/judge_PB.json | a13cc9df7db7ad10feabd95e4e15e8638ac5b4d8122cd576f29f6999586a1740 |
| result_manifest.json: results/results_P0.json | 26c8355748fd43d184e0b61514c0c958cdeb6a1c678557995888ff721a382b33 |
| result_manifest.json: results/results_PA.json | b5956090d9d2f87cf37a31b37e1bed03bbd9ecae0adc882c86131bd4cfd759a0 |
| result_manifest.json: results/results_PAB.json | d329cce423fea250c7a29a93c97f343ba4d40558c362e164ef628265289c84ef |
| result_manifest.json: results/results_PB.json | 36b424813d3632ca37f7c03b17c9d6bb2e9e2200ab46f63ac6e18a9ef1ced77e |
| result_manifest.json: analysis_summary.json | 4ca961b06690423a3f7cacb74d80a3b2adb353d0ef6d2334fcdb7944719b9f08 |
| result_manifest.json: analysis_summary.md | 2b4a3b02ccd319087f7656d837b9249d83406a3fc0b95fad4321fa00d96415bc |
| result_manifest.json: judge_execution_note.md | 83e46620ec483c52de48648dd776a6fa5773a6da4f48509f0dfee0a94f3be844 |
Evidence package download
Download the original completed archive for Studies 002 and 003. The source package is unchanged; this reading page is a presentation of its records and explicit limitations.
Download studies 002 + 003