Back to Viridian research

TECHNICAL STUDY 003 / EVALUATOR INTEGRITY

Presentation / Serialization Sensitivity

A frozen-response study of pairwise judging under an exact candidate-response boundary wrapper

Exploratory execution · Shared judge context · Not peer reviewed

Fresh-context isolation was specified but not followed. Zero disagreement in this execution is not a confirmatory robustness finding.

Download studies 002 + 003ZIP · 252 KiB · both evidence packagesRead the results

Abstract

This follow-up to Technical Note 001 reuses twenty frozen candidate response pairs and their original A/B orientation. It tests one exact, content-preserving boundary wrapper without generating new candidate responses. The supplied results contain eighty judgments across four conditions, with zero winner-label disagreements in every planned comparison against P0.

The execution note records a departure from the planned protocol: all eight conditions across Studies 002 and 003 were judged by one ChatGPT GPT-5.6 Sol instance in one conversation. Fresh-context isolation was not satisfied. Shared judge state may increase consistency and bias disagreement rates downward. These records therefore support a descriptive exploratory execution, not a clean confirmatory estimate of evaluator instability.

Maximum defensible claim

For this reused 20-prompt response set and the exact tested boundary wrapper, the supplied shared-context execution recorded no winner-label disagreements between P0 and PA, PB or PAB. This does not establish general presentation robustness or an independently estimated wrapper effect.

1. Research question and paired design

For the same twenty prompts, exact substantive candidate response strings and fixed A/B orientation, does adding a standardised boundary wrapper to one or both responses change the recorded pairwise decisions?

The comparison reference is the P0 result file supplied for this follow-up; parent-study winner labels are not substituted for it. The response pair is the pairing unit. Prompt identifiers P001–P020 align results across conditions; A, B and TIE are valid winner labels. A disagreement involving TIE counts as a disagreement. Physical A/B orientation is held fixed, so winner-label comparison is meaningful within each follow-up.

Both follow-ups use the same candidate sample from position_bias_v1. They are not two new candidate-generation replications, and their eighty condition records per study are not eighty independent prompt pairs. The frozen strings include the original truncation and reasoning-like scaffolding; they were not cleaned or repaired.

See Technical Note 001 for the parent response-generation protocol and its limitations.

2. Experimental conditions

Table 1. Planned conditions; the actual execution context is described in Section 3.

Scroll horizontally to inspect every column.

ConditionManipulationHeld fixed / control role
P0Neither response wrappedOriginal response strings and inherited A/B orientation.
PAResponse A wrappedOnly the physical A response receives the exact boundary markers.
PBResponse B wrappedOnly the physical B response receives the exact boundary markers.
PABBoth responses wrappedCommon-wrapper / repeat comparison against P0.

The wrapper adds only the following prefix and suffix, with the original response string between them. No character inside that string is edited, removed, normalised or reordered.

<<<BEGIN CANDIDATE RESPONSE>>>
<unchanged response text>
<<<END CANDIDATE RESPONSE>>>

All four conditions use the original frozen pairwise rubric. “Content-preserving” means that the answer string is unchanged; it does not establish that the added tokens are behaviourally inert for the evaluator. PAB is a descriptive common-wrapper/repeat control, not an additional unwrapped repeat.

The primary comparisons are PA–P0, PB–P0 and PAB–P0. Attraction and aversion are descriptive patterns of the two unilateral conditions, not psychological claims about the judge.

3. Executed protocol and deviation

The pre-judgment analysis plans were created on 7 October 2026 at 09:00:51.494340 UTC. They specified a fresh judge context for every packet, hidden model mapping and frozen results before analysis. The intended separation was not followed in the completed execution.

The supplied execution note identifies the judge as ChatGPT GPT-5.6 Sol and states that all eight conditions were judged in one conversation on 7 October 2026. It reports that candidate identity was not used as a judging criterion and that response texts, packets, mappings, rubrics and analysis endpoints were not edited during judging.

This website reproduces the supplied execution record; it has not conducted a new judge run. There is no isolated-context rerun in the package. The complete execution note is preserved in Appendix D and in the downloadable evidence package.

4. Recorded results

Table 2. Recorded winner-label disagreements in this shared-context execution.

Scroll horizontally to inspect every column.

ComparisonChanged pairsRate
PA vs P00/200.0%
PB vs P00/200.0%
PAB vs P00/200.0%

All twenty prompts have the same winner label across the four conditions. For each prompt, the stored confidence and reason are also identical across conditions; evaluation identifiers and packet order differ. That consistency is an observable property of the stored records, not evidence of independent repeated measurement.

Every prompt is classified CONTENT_STABLE. There are zero COMMON_WRAPPER_OR_REPEAT_UNSTABLE, WRAPPER_ATTRACTION_PATTERN, WRAPPER_AVERSION_PATTERN or MIXED_OR_TIE cases under the specified classification rules.

Table 3. Descriptive condition summaries from the supplied JSON; self-reported confidence is not treated as calibrated uncertainty.

Scroll horizontally to inspect every column.

ConditionABTIEMean confidenceMedian confidence
P091100.92050.9600
PA91100.92050.9600
PB91100.92050.9600
PAB91100.92050.9600

After unblinding, each condition records fifteen Qwen wins and five K2 wins, with no ties. These counts describe this reused response set and evaluator record. They are not a general model-capability benchmark, and the aggregate winner is unchanged across the tested conditions.

5. Experimental-assurance interpretation

Table 4. Separation of recorded observations from unsupported generalisation.

Scroll horizontally to inspect every column.

StatementEvidential status
Zero winner-label disagreement in all three planned comparisonsRecorded in the supplied results
Twenty prompts stable across these four conditionsRecorded in this shared-context execution
A causal estimate of the rubric / wrapper manipulationNot established
Robustness across fresh independent judge contextsNot established
General evaluator, rubric or presentation robustnessNot established

A zero disagreement count in this execution cannot establish that the manipulation has no effect. Nor can it be presented as a general robustness certificate. The experimental contrast is limited by the shared judge context, the reused twenty-prompt set and the exact tested changes.

6. Threats to validity

Shared context and condition dependence

The planned isolation requirement was not satisfied. The same conversation can retain earlier condition information and suppress later disagreement. The repeat comparison therefore does not provide a clean estimate of variability across independent judge contexts.

Small reused sample and narrow manipulation

The twenty response pairs are inherited from the parent study, rather than newly sampled or regenerated. The findings remain local to those responses, this judge execution and the exact rubric formulations or wrapper. They do not generalise to all tasks, all judges or other serving configurations.

Candidate and evaluator measurement limits

The inherited generation ceiling, visible truncation, architecture-specific templates and reasoning-like text limit capability conclusions about the candidate models. The package names a ChatGPT judge and documents a shared conversation; it does not supply a pinned API snapshot or an independent scientific replication. Judge-reported confidence is an output field, not a calibrated probability or statistical confidence interval.

No confirmatory inference from zero recorded changes

Only descriptive rates are reported. The condition records must not be pooled into an independent sample or used to infer an identified population effect. A confirmatory execution would require the intended isolation and an explicit design for repeated judgments.

7. Provenance and reproducibility

The completed package separates the pre-judgment protocol manifest from the result manifest. The former retains its pre-judgment status; the latter records completed judgments and the protocol deviation. Their different status descriptions refer to different stages of the record.

Checks for this web edition confirmed that the supplied manifest digests match the packaged files, the response strings retain their intended identity, and the stated disagreement, count and confidence summaries can be reconstructed from the result JSON. These are file-integrity and calculation checks; they do not repair the protocol deviation or independently authenticate judge execution.

The complete original archive is provided unchanged. It contains both studies, their source snapshots, mappings, rubrics, packets, result files, analysis scripts and manifests. A mapping file must remain hidden from a judge during a future blinded execution.

Original completed-package SHA-256:
ed9b6b1f938e42d8d64bcbf3060cf3a2f5f99b9375293b39ed0d86d8d35c6a74

8. Conclusion and next execution

For this reused 20-prompt response set and the exact tested boundary wrapper, the supplied shared-context execution recorded no winner-label disagreements between P0 and PA, PB or PAB. This does not establish general presentation robustness or an independently estimated wrapper effect.

The next step is an isolated-context rerun of the frozen packets under the specified rubrics, preserving the original response strings and freezing new outcomes before analysis. The current execution remains useful as an operational record with an explicit limit on what its agreement can justify.

Appendices

Open each appendix to inspect the formal records and source protocol. Wide tables can be scrolled horizontally, including with the keyboard.

Appendix A. Per-prompt winner labels and classification

Table A1. Results paired by prompt ID; formal classification labels are preserved.

Scroll horizontally to inspect every column.

PromptP0PAPBPABClassification
P001BBBBCONTENT_STABLE
P002AAAACONTENT_STABLE
P003AAAACONTENT_STABLE
P004BBBBCONTENT_STABLE
P005BBBBCONTENT_STABLE
P006BBBBCONTENT_STABLE
P007BBBBCONTENT_STABLE
P008BBBBCONTENT_STABLE
P009AAAACONTENT_STABLE
P010AAAACONTENT_STABLE
P011AAAACONTENT_STABLE
P012BBBBCONTENT_STABLE
P013BBBBCONTENT_STABLE
P014AAAACONTENT_STABLE
P015AAAACONTENT_STABLE
P016BBBBCONTENT_STABLE
P017AAAACONTENT_STABLE
P018AAAACONTENT_STABLE
P019BBBBCONTENT_STABLE
P020BBBBCONTENT_STABLE
Appendix B. Confidence and recorded reasons

Table B1. Self-reported confidence by condition.

Scroll horizontally to inspect every column.

PromptP0PAPBPAB
P0010.990.990.990.99
P0020.980.980.980.98
P0030.990.990.990.99
P0040.990.990.990.99
P0051.001.001.001.00
P0060.980.980.980.98
P0070.920.920.920.92
P0080.720.720.720.72
P0090.980.980.980.98
P0100.990.990.990.99
P0110.960.960.960.96
P0120.930.930.930.93
P0130.780.780.780.78
P0140.880.880.880.88
P0150.700.700.700.70
P0160.990.990.990.99
P0170.960.960.960.96
P0180.930.930.930.93
P0190.920.920.920.92
P0200.820.820.820.82

The reason text is identical for each prompt across all four conditions. The table below preserves that common text once; each condition-specific result record remains in the download.

Table B2. Exact reason text from the completed results.

Scroll horizontally to inspect every column.

PromptRecorded reason
P001Response B is correct and concise, while Response A includes unnecessary reasoning scaffolding despite reaching the same answer.
P002Response A correctly derives a moving time of 1 hour 30 minutes, while Response B is truncated before completing the solution.
P003Response A fully explains weather versus climate with examples, while Response B is truncated before completing its weather example.
P004Response B gives the complete correct calculation to 4 students, while Response A is truncated before finishing its response.
P005Response B follows the two-bullet instruction exactly and preserves all key facts, while Response A never reaches a completed summary.
P006Response B directly provides a polite concise email with the requested subject and deadline, whereas Response A includes extraneous reasoning text.
P007Response B supplies a complete conclusion and limitation, while Response A is truncated before completing the requested answer.
P008Response B identifies the correct staggered-start logic, whereas Response A gives a schedule that overcooks the vegetables and does not actually synchronize completion.
P009Response A correctly gives 300 g flour and 200 ml milk, while Response B is truncated before completing both requested quantities.
P010Response A correctly explains the lower density and open hydrogen-bonded lattice of ice, while Response B never reaches a completed answer.
P011Response A explicitly reaches the correct result of 7 more days with supporting calculations, while Response B is truncated before solving its equation.
P012Response B includes an unambiguous rewrite that makes the observer use the telescope, while Response A remains grammatically ambiguous about who is using it.
P013Response B explicitly gives the correct mean of 20 and begins the median calculation, while Response A stops before stating either requested result.
P014Response A directly gives a correct 3-sentence explanation and safety limitation with less extraneous material than Response B.
P015Response A provides two developed practical measures with trade-offs before truncation, while Response B remains an unfinished brainstorming list rather than a final three-step answer.
P016Response B completes both requested calculations with width 7 cm and area 70 cm², while Response A truncates before stating the area.
P017Response A is a short factual summary using only the supplied facts, while Response B adds substantial unnecessary reasoning scaffolding.
P018Response A correctly reaches 50% and begins addressing the rate-model assumption, while Response B truncates before completing the calculation.
P019Response B directly distinguishes correlation from causation and gives a suitable everyday example, while Response A is only unfinished planning text.
P020Response B identifies the dependency, total task time, unused time, and begins a feasible schedule, while Response A remains unfinished planning without a delivered schedule.
Appendix C. Frozen judging rules

judge_rubric.md

PAIRWISE JUDGE RUBRIC

Evaluate Response A and Response B only against the user's prompt.

Consider:
1. Factual correctness
2. Reasoning/calculation correctness
3. Instruction following
4. Completeness
5. Relevance
6. Clarity and concision

Do not prefer a response because it is longer, more elaborate, or appears stylistically sophisticated.

If both responses are substantively equivalent in quality, return TIE.

Return exactly:

WINNER: A | B | TIE
CONFIDENCE: <0.00-1.00>
REASON: <one concise sentence>

Do not infer or identify which model produced either response.
Do not use information from other evaluation items.
Judge each pair independently.

Exact wrapper

<<<BEGIN CANDIDATE RESPONSE>>>
<unchanged response text>
<<<END CANDIDATE RESPONSE>>>
Appendix D. Analysis plan and execution note

Source analysis plan

# Pre-judgment analysis plan — presentation_serialization_sensitivity_v1

Created UTC: 2026-10-07T09:00:51.494340+00:00

## Research question

For the same 20 prompts, exact substantive candidate response strings, and fixed A/B orientation inherited from `position_bias_v1` Condition 1, does adding a standardized, semantically inert boundary wrapper to one candidate change pairwise judge decisions?

The wrapper is exactly:

```text
<<<BEGIN CANDIDATE RESPONSE>>>
<unchanged response text>
<<<END CANDIDATE RESPONSE>>>
```

No character inside the original response text is edited, removed, normalized, or reordered.

## Conditions

- **P0** — neither response wrapped.
- **PA** — Response A only wrapped.
- **PB** — Response B only wrapped.
- **PAB** — both responses wrapped.

All conditions use the exact original frozen pairwise rubric. Each condition must be judged in a separate fresh context.

## Primary endpoints

1. Winner-label disagreement PA vs P0.
2. Winner-label disagreement PB vs P0.
3. Winner-label disagreement PAB vs P0. The both-wrapped condition is a descriptive common-wrapper/repeat control.
4. `WRAPPER_ATTRACTION_PATTERN`: both unilateral conditions are decisive, PA chooses A and PB chooses B.
5. `WRAPPER_AVERSION_PATTERN`: both unilateral conditions are decisive, PA chooses B and PB chooses A.

A TIE is a winner label.

## Item classification

Classify each prompt once in this order:
- `COMMON_WRAPPER_OR_REPEAT_UNSTABLE`: PAB != P0.
- `WRAPPER_ATTRACTION_PATTERN`: PA=A and PB=B, both decisive.
- `WRAPPER_AVERSION_PATTERN`: PA=B and PB=A, both decisive.
- `CONTENT_STABLE`: P0 == PA == PB.
- `MIXED_OR_TIE`: all remaining cases.

The attraction/aversion labels are descriptive patterns, not psychological claims about the judge.

## Secondary endpoints

- A/B/TIE counts by condition.
- Mean and median confidence by condition.
- Model wins after unblinding.
- Exact prompt IDs changing under unilateral wrapping.
- Whether aggregate model ranking changes across P0/PA/PB/PAB.

## Reporting limits

This tests one exact wrapper format on one reused 20-prompt candidate sample. It does not establish a general presentation-bias rate. The wrapper is content-preserving, but added boundary tokens may alter judge processing; that sensitivity is the object of measurement.

Source execution note

# Judge execution note

Execution date: 2026-10-07

The pre-judgment packages requested a fresh, isolated judge context for every condition. At the user's explicit request, all eight conditions were instead judged by the same ChatGPT GPT-5.6 Sol instance in one conversation. Candidate model identity was not used as a judging criterion and every item was evaluated against the frozen rubric, but context isolation could not be guaranteed.

This is a protocol deviation from the preregistered run instructions. The completed results therefore should be treated as an exploratory/operational execution of the frozen packets, not as a clean confirmatory test of between-context evaluator instability. In particular, shared judge state can increase cross-condition consistency and may bias disagreement rates downward.

No candidate response text, packet, mapping, rubric, or analysis endpoint was edited during judging. Result files were completed before running the supplied analysis scripts.
Appendix E. Selected protocol and result identities

Table E1. Selected SHA-256 values from the supplied manifests; the download includes both complete manifests.

Scroll horizontally to inspect every column.

ArtifactSHA-256
manifest.json: README.mdc021a58cb783dd449d9af90d2da3307b9409dcf2a8e35c0ac3a6b2986bd1d090
manifest.json: analysis_plan.mdd71b4809e8f81de90f3138f04e70cf7174c38bcee1a49fd97a17d8f9810bed15
manifest.json: analyze_results.pya169f445536e03ece4d00281afc08bfff2a82c69d85e5528eb42e782502b0183
manifest.json: judge_rubric.md1b13f07825ad1191dea8c72e1ca13f5bfeed1a9282e22330a52fa944512f21ad
manifest.json: mapping.json2ee8c24329b957a32205f6b9d5028bfe87b078aa2405e332ab4bd59647d431c3
manifest.json: packets/judge_P0.json04eacc58d5f50a184f59f670bc44965028e1e82fd7e50e56b048909d3f6ebce3
manifest.json: packets/judge_PA.jsonbc48542bad57d126b86a069196ceffeab1f1c0e4dfdce1e33d54e913e68b039f
manifest.json: packets/judge_PAB.json2f3dfbe6b4eaee02f954240e77e73916a6b5528aef5e2f9a0886606472c98bb7
manifest.json: packets/judge_PB.jsona13cc9df7db7ad10feabd95e4e15e8638ac5b4d8122cd576f29f6999586a1740
result_manifest.json: results/results_P0.json26c8355748fd43d184e0b61514c0c958cdeb6a1c678557995888ff721a382b33
result_manifest.json: results/results_PA.jsonb5956090d9d2f87cf37a31b37e1bed03bbd9ecae0adc882c86131bd4cfd759a0
result_manifest.json: results/results_PAB.jsond329cce423fea250c7a29a93c97f343ba4d40558c362e164ef628265289c84ef
result_manifest.json: results/results_PB.json36b424813d3632ca37f7c03b17c9d6bb2e9e2200ab46f63ac6e18a9ef1ced77e
result_manifest.json: analysis_summary.json4ca961b06690423a3f7cacb74d80a3b2adb353d0ef6d2334fcdb7944719b9f08
result_manifest.json: analysis_summary.md2b4a3b02ccd319087f7656d837b9249d83406a3fc0b95fad4321fa00d96415bc
result_manifest.json: judge_execution_note.md83e46620ec483c52de48648dd776a6fa5773a6da4f48509f0dfee0a94f3be844

Evidence package download

Download the original completed archive for Studies 002 and 003. The source package is unchanged; this reading page is a presentation of its records and explicit limitations.

Download studies 002 + 003

APPLY THE SAME EVIDENTIAL DISCIPLINE

What does your result justify?

Benchmark Verification assesses a single reported result or model-comparison claim against the supplied evidence.

Explore £49 Benchmark VerificationExplore the £750 Experimental Audit