Back to Viridian research

TECHNICAL STUDY 002 / EVALUATOR INTEGRITY

Evaluator Contract Stability

A frozen-response study of pairwise judging under original, reordered and compact rubric formulations

Exploratory execution · Shared judge context · Not peer reviewed

Fresh-context isolation was specified but not followed. Zero disagreement in this execution is not a confirmatory robustness finding.

Download studies 002 + 003ZIP · 252 KiB · both evidence packagesRead the results

Abstract

This follow-up to Technical Note 001 reuses twenty frozen candidate response pairs and their original A/B orientation. It tests rubric formulations intended to preserve the same substantive criteria without generating new candidate responses. The supplied results contain eighty judgments across four conditions, with zero winner-label disagreements in every planned comparison against R0A.

The execution note records a departure from the planned protocol: all eight conditions across Studies 002 and 003 were judged by one ChatGPT GPT-5.6 Sol instance in one conversation. Fresh-context isolation was not satisfied. Shared judge state may increase consistency and bias disagreement rates downward. These records therefore support a descriptive exploratory execution, not a clean confirmatory estimate of evaluator instability.

Maximum defensible claim

For this reused 20-prompt response set, the supplied shared-context execution recorded no winner-label disagreements between R0A and R0B, R1 or R2. This does not establish rubric robustness in independent judge contexts.

1. Research question and paired design

For the same twenty prompts, exact candidate response texts and fixed A/B orientation, how stable are pairwise winner labels under rubric formulations intended to preserve the same substantive decision criteria?

The comparison reference is the R0A result file supplied for this follow-up; parent-study winner labels are not substituted for it. The response pair is the pairing unit. Prompt identifiers P001–P020 align results across conditions; A, B and TIE are valid winner labels. A disagreement involving TIE counts as a disagreement. Physical A/B orientation is held fixed, so winner-label comparison is meaningful within each follow-up.

Both follow-ups use the same candidate sample from position_bias_v1. They are not two new candidate-generation replications, and their eighty condition records per study are not eighty independent prompt pairs. The frozen strings include the original truncation and reasoning-like scaffolding; they were not cleaned or repaired.

See Technical Note 001 for the parent response-generation protocol and its limitations.

2. Experimental conditions

Table 1. Planned conditions; the actual execution context is described in Section 3.

Scroll horizontally to inspect every column.

ConditionManipulationHeld fixed / control role
R0AOriginal rubricReference condition; inherited A/B orientation and response strings.
R0BOriginal rubric repeatSame rubric and response orientation; independently shuffled item order.
R1Reordered criteriaSame six substantive criteria and decision rules, with criteria reordered.
R2Compact restatementSame six substantive criteria and decision rules, expressed more compactly.

The original rubric covers factual correctness, reasoning/calculation correctness, instruction following, completeness, relevance, and clarity/concision. R1 reorders those criteria; R2 restates them compactly. Their substantive equivalence is the design intention, not an independently established property.

The primary endpoints are R0B–R0A, R1–R0A and R2–R0A winner-label disagreement. Excess disagreement subtracts the repeat-control proportion from each perturbed-rubric proportion and is descriptive, not a causal estimate.

3. Executed protocol and deviation

The pre-judgment analysis plans were created on 7 October 2026 at 09:00:51.494340 UTC. They specified a fresh judge context for every packet, hidden model mapping and frozen results before analysis. The intended separation was not followed in the completed execution.

The supplied execution note identifies the judge as ChatGPT GPT-5.6 Sol and states that all eight conditions were judged in one conversation on 7 October 2026. It reports that candidate identity was not used as a judging criterion and that response texts, packets, mappings, rubrics and analysis endpoints were not edited during judging.

This website reproduces the supplied execution record; it has not conducted a new judge run. There is no isolated-context rerun in the package. The complete execution note is preserved in Appendix D and in the downloadable evidence package.

4. Recorded results

Table 2. Recorded winner-label disagreements in this shared-context execution.

Scroll horizontally to inspect every column.

ComparisonChanged pairsRate
R0B vs R0A0/200.0%
R1 vs R0A0/200.0%
R2 vs R0A0/200.0%

All twenty prompts have the same winner label across the four conditions. For each prompt, the stored confidence and reason are also identical across conditions; evaluation identifiers and packet order differ. That consistency is an observable property of the stored records, not evidence of independent repeated measurement.

Every prompt is classified STABLE_ACROSS_TESTED_RUBRICS. No prompt is CONTROL_UNSTABLE or RUBRIC_SENSITIVE. Both descriptive excess-disagreement values are zero percentage points.

Table 3. Descriptive condition summaries from the supplied JSON; self-reported confidence is not treated as calibrated uncertainty.

Scroll horizontally to inspect every column.

ConditionABTIEMean confidenceMedian confidence
R0A91100.92050.9600
R0B91100.92050.9600
R191100.92050.9600
R291100.92050.9600

After unblinding, each condition records fifteen Qwen wins and five K2 wins, with no ties. These counts describe this reused response set and evaluator record. They are not a general model-capability benchmark, and the aggregate winner is unchanged across the tested conditions.

5. Experimental-assurance interpretation

Table 4. Separation of recorded observations from unsupported generalisation.

Scroll horizontally to inspect every column.

StatementEvidential status
Zero winner-label disagreement in all three planned comparisonsRecorded in the supplied results
Twenty prompts stable across these four conditionsRecorded in this shared-context execution
A causal estimate of the rubric / wrapper manipulationNot established
Robustness across fresh independent judge contextsNot established
General evaluator, rubric or presentation robustnessNot established

A zero disagreement count in this execution cannot establish that the manipulation has no effect. Nor can it be presented as a general robustness certificate. The experimental contrast is limited by the shared judge context, the reused twenty-prompt set and the exact tested changes.

6. Threats to validity

Shared context and condition dependence

The planned isolation requirement was not satisfied. The same conversation can retain earlier condition information and suppress later disagreement. The repeat comparison therefore does not provide a clean estimate of variability across independent judge contexts.

Small reused sample and narrow manipulation

The twenty response pairs are inherited from the parent study, rather than newly sampled or regenerated. The findings remain local to those responses, this judge execution and the exact rubric formulations or wrapper. They do not generalise to all tasks, all judges or other serving configurations.

Candidate and evaluator measurement limits

The inherited generation ceiling, visible truncation, architecture-specific templates and reasoning-like text limit capability conclusions about the candidate models. The package names a ChatGPT judge and documents a shared conversation; it does not supply a pinned API snapshot or an independent scientific replication. Judge-reported confidence is an output field, not a calibrated probability or statistical confidence interval.

No confirmatory inference from zero recorded changes

Only descriptive rates are reported. The condition records must not be pooled into an independent sample or used to infer an identified population effect. A confirmatory execution would require the intended isolation and an explicit design for repeated judgments.

7. Provenance and reproducibility

The completed package separates the pre-judgment protocol manifest from the result manifest. The former retains its pre-judgment status; the latter records completed judgments and the protocol deviation. Their different status descriptions refer to different stages of the record.

Checks for this web edition confirmed that the supplied manifest digests match the packaged files, the response strings retain their intended identity, and the stated disagreement, count and confidence summaries can be reconstructed from the result JSON. These are file-integrity and calculation checks; they do not repair the protocol deviation or independently authenticate judge execution.

The complete original archive is provided unchanged. It contains both studies, their source snapshots, mappings, rubrics, packets, result files, analysis scripts and manifests. A mapping file must remain hidden from a judge during a future blinded execution.

Original completed-package SHA-256:
ed9b6b1f938e42d8d64bcbf3060cf3a2f5f99b9375293b39ed0d86d8d35c6a74

8. Conclusion and next execution

For this reused 20-prompt response set, the supplied shared-context execution recorded no winner-label disagreements between R0A and R0B, R1 or R2. This does not establish rubric robustness in independent judge contexts.

The next step is an isolated-context rerun of the frozen packets under the specified rubrics, preserving the original response strings and freezing new outcomes before analysis. The current execution remains useful as an operational record with an explicit limit on what its agreement can justify.

Appendices

Open each appendix to inspect the formal records and source protocol. Wide tables can be scrolled horizontally, including with the keyboard.

Appendix A. Per-prompt winner labels and classification

Table A1. Results paired by prompt ID; formal classification labels are preserved.

Scroll horizontally to inspect every column.

PromptR0AR0BR1R2Classification
P001BBBBSTABLE_ACROSS_TESTED_RUBRICS
P002AAAASTABLE_ACROSS_TESTED_RUBRICS
P003AAAASTABLE_ACROSS_TESTED_RUBRICS
P004BBBBSTABLE_ACROSS_TESTED_RUBRICS
P005BBBBSTABLE_ACROSS_TESTED_RUBRICS
P006BBBBSTABLE_ACROSS_TESTED_RUBRICS
P007BBBBSTABLE_ACROSS_TESTED_RUBRICS
P008BBBBSTABLE_ACROSS_TESTED_RUBRICS
P009AAAASTABLE_ACROSS_TESTED_RUBRICS
P010AAAASTABLE_ACROSS_TESTED_RUBRICS
P011AAAASTABLE_ACROSS_TESTED_RUBRICS
P012BBBBSTABLE_ACROSS_TESTED_RUBRICS
P013BBBBSTABLE_ACROSS_TESTED_RUBRICS
P014AAAASTABLE_ACROSS_TESTED_RUBRICS
P015AAAASTABLE_ACROSS_TESTED_RUBRICS
P016BBBBSTABLE_ACROSS_TESTED_RUBRICS
P017AAAASTABLE_ACROSS_TESTED_RUBRICS
P018AAAASTABLE_ACROSS_TESTED_RUBRICS
P019BBBBSTABLE_ACROSS_TESTED_RUBRICS
P020BBBBSTABLE_ACROSS_TESTED_RUBRICS
Appendix B. Confidence and recorded reasons

Table B1. Self-reported confidence by condition.

Scroll horizontally to inspect every column.

PromptR0AR0BR1R2
P0010.990.990.990.99
P0020.980.980.980.98
P0030.990.990.990.99
P0040.990.990.990.99
P0051.001.001.001.00
P0060.980.980.980.98
P0070.920.920.920.92
P0080.720.720.720.72
P0090.980.980.980.98
P0100.990.990.990.99
P0110.960.960.960.96
P0120.930.930.930.93
P0130.780.780.780.78
P0140.880.880.880.88
P0150.700.700.700.70
P0160.990.990.990.99
P0170.960.960.960.96
P0180.930.930.930.93
P0190.920.920.920.92
P0200.820.820.820.82

The reason text is identical for each prompt across all four conditions. The table below preserves that common text once; each condition-specific result record remains in the download.

Table B2. Exact reason text from the completed results.

Scroll horizontally to inspect every column.

PromptRecorded reason
P001Response B is correct and concise, while Response A includes unnecessary reasoning scaffolding despite reaching the same answer.
P002Response A correctly derives a moving time of 1 hour 30 minutes, while Response B is truncated before completing the solution.
P003Response A fully explains weather versus climate with examples, while Response B is truncated before completing its weather example.
P004Response B gives the complete correct calculation to 4 students, while Response A is truncated before finishing its response.
P005Response B follows the two-bullet instruction exactly and preserves all key facts, while Response A never reaches a completed summary.
P006Response B directly provides a polite concise email with the requested subject and deadline, whereas Response A includes extraneous reasoning text.
P007Response B supplies a complete conclusion and limitation, while Response A is truncated before completing the requested answer.
P008Response B identifies the correct staggered-start logic, whereas Response A gives a schedule that overcooks the vegetables and does not actually synchronize completion.
P009Response A correctly gives 300 g flour and 200 ml milk, while Response B is truncated before completing both requested quantities.
P010Response A correctly explains the lower density and open hydrogen-bonded lattice of ice, while Response B never reaches a completed answer.
P011Response A explicitly reaches the correct result of 7 more days with supporting calculations, while Response B is truncated before solving its equation.
P012Response B includes an unambiguous rewrite that makes the observer use the telescope, while Response A remains grammatically ambiguous about who is using it.
P013Response B explicitly gives the correct mean of 20 and begins the median calculation, while Response A stops before stating either requested result.
P014Response A directly gives a correct 3-sentence explanation and safety limitation with less extraneous material than Response B.
P015Response A provides two developed practical measures with trade-offs before truncation, while Response B remains an unfinished brainstorming list rather than a final three-step answer.
P016Response B completes both requested calculations with width 7 cm and area 70 cm², while Response A truncates before stating the area.
P017Response A is a short factual summary using only the supplied facts, while Response B adds substantial unnecessary reasoning scaffolding.
P018Response A correctly reaches 50% and begins addressing the rate-model assumption, while Response B truncates before completing the calculation.
P019Response B directly distinguishes correlation from causation and gives a suitable everyday example, while Response A is only unfinished planning text.
P020Response B identifies the dependency, total task time, unused time, and begins a feasible schedule, while Response A remains unfinished planning without a delivered schedule.
Appendix C. Frozen judging rules

R0_original.md

PAIRWISE JUDGE RUBRIC

Evaluate Response A and Response B only against the user's prompt.

Consider:
1. Factual correctness
2. Reasoning/calculation correctness
3. Instruction following
4. Completeness
5. Relevance
6. Clarity and concision

Do not prefer a response because it is longer, more elaborate, or appears stylistically sophisticated.

If both responses are substantively equivalent in quality, return TIE.

Return exactly:

WINNER: A | B | TIE
CONFIDENCE: <0.00-1.00>
REASON: <one concise sentence>

Do not infer or identify which model produced either response.
Do not use information from other evaluation items.
Judge each pair independently.

R1_reordered.md

PAIRWISE JUDGE RUBRIC — REORDERED EQUIVALENT FORM

Evaluate Response A and Response B only against the user's prompt.

Consider:
1. Instruction following
2. Relevance
3. Factual correctness
4. Completeness
5. Reasoning/calculation correctness
6. Clarity and concision

Do not prefer a response because it is longer, more elaborate, or appears stylistically sophisticated.

If both responses are substantively equivalent in quality, return TIE.

Return exactly:

WINNER: A | B | TIE
CONFIDENCE: <0.00-1.00>
REASON: <one concise sentence>

Do not infer or identify which model produced either response.
Do not use information from other evaluation items.
Judge each pair independently.

R2_compact.md

PAIRWISE JUDGE RUBRIC — COMPACT EQUIVALENT FORM

Compare Response A with Response B only as answers to the user's prompt. Judge factual correctness, reasoning/calculation correctness, instruction following, completeness, relevance, and clarity/concision. Do not reward length, elaborateness, or stylistic sophistication by itself.

If the two responses are substantively equivalent in quality, return TIE.

Return exactly:

WINNER: A | B | TIE
CONFIDENCE: <0.00-1.00>
REASON: <one concise sentence>

Do not infer model identity, use information from other items, or compare across items. Judge every pair independently.
Appendix D. Analysis plan and execution note

Source analysis plan

# Pre-judgment analysis plan — evaluator_contract_stability_v1

Created UTC: 2026-10-07T09:00:51.494340+00:00

## Research question

For the same 20 prompts, exact candidate response texts, and fixed A/B orientation inherited from `position_bias_v1` Condition 1, how stable are pairwise winner labels under two rubric formulations intended to preserve the same substantive decision criteria?

## Conditions

- **R0A** — original frozen rubric; reference repeat.
- **R0B** — exact same original rubric and exact same response orientation/texts, independently shuffled item order; repeat-control estimate.
- **R1** — the same six criteria and decision/output rules, with criteria reordered.
- **R2** — the same six criteria and decision/output rules, compactly restated.

Each condition must be judged in a separate fresh context. The judge must not see `mapping.json`, another condition's results, or any previous result while judging a condition. Do not edit candidate response text.

## Primary endpoints

Using prompt ID as the pairing key:

1. Winner-label disagreement proportion R0B vs R0A. This is the repeat/order-control disagreement.
2. Winner-label disagreement proportion R1 vs R0A.
3. Winner-label disagreement proportion R2 vs R0A.
4. Descriptive excess disagreement for each perturbed rubric: perturbation disagreement minus R0B-vs-R0A disagreement. This is descriptive only and is not a causal estimate.

A TIE is a winner label and disagreements involving a TIE count as disagreement.

## Item classification

Classify each prompt once:
- `CONTROL_UNSTABLE`: R0A != R0B.
- `STABLE_ACROSS_TESTED_RUBRICS`: R0A == R0B == R1 == R2.
- `RUBRIC_SENSITIVE`: control repeats agree but at least one of R1/R2 differs.

## Secondary endpoints

- A/B/TIE counts by condition.
- Mean and median reported confidence by condition.
- After unblinding only after all judgments are frozen: Qwen3.5 4B wins, K2 Horizon 3.7B wins, and ties by condition.
- Exact prompt IDs that change winner label in R1 or R2 relative to the stable control reference.
- Whether the aggregate model winner across the 20 prompts changes by rubric condition.

## Reporting limits

This is a 20-prompt reused candidate sample, not an independent replication of candidate generation. Report descriptive results only. Do not generalize to all LLM judges, all rubrics, or all tasks. The study tests stability for this fixed candidate set under these exact rubric manipulations.

Source execution note

# Judge execution note

Execution date: 2026-10-07

The pre-judgment packages requested a fresh, isolated judge context for every condition. At the user's explicit request, all eight conditions were instead judged by the same ChatGPT GPT-5.6 Sol instance in one conversation. Candidate model identity was not used as a judging criterion and every item was evaluated against the frozen rubric, but context isolation could not be guaranteed.

This is a protocol deviation from the preregistered run instructions. The completed results therefore should be treated as an exploratory/operational execution of the frozen packets, not as a clean confirmatory test of between-context evaluator instability. In particular, shared judge state can increase cross-condition consistency and may bias disagreement rates downward.

No candidate response text, packet, mapping, rubric, or analysis endpoint was edited during judging. Result files were completed before running the supplied analysis scripts.
Appendix E. Selected protocol and result identities

Table E1. Selected SHA-256 values from the supplied manifests; the download includes both complete manifests.

Scroll horizontally to inspect every column.

ArtifactSHA-256
manifest.json: README.md20361897af40de5c6a4eab0c9f95b6791b3fcfb2679eaa74d5166c1a6b45ce6c
manifest.json: analysis_plan.mdf917d4976f4aef53593b56fc4b83f1f693db40eb88233fa277269f78673d39c2
manifest.json: analyze_results.pycaddea830e7279d9444bc1e1a9b53414d3cd89a4f5e67d26d73f75bf03a3b59b
manifest.json: mapping.jsond8972506efc5b1150443accb7798181393700c2d12544d099d9cc6701611eba6
manifest.json: packets/judge_R0A.jsonfb5203bb2af679f15f3a4dbfb3cc4e63aea08ef83bd57bcf55f378b834521068
manifest.json: packets/judge_R0B.jsone6c8fd78770ea9d993822dd52c04f507399e1eb04f233f6d3b4264ec877abcbf
manifest.json: packets/judge_R1.jsona9679f83cf9f849f694e7e4624a616b6536d31779556d159aced703391f080d1
manifest.json: packets/judge_R2.jsoneef22276187ee84a3d58f655885bdb307b25d54c98047a6521f12e959a25edbd
manifest.json: rubrics/R0_original.md1b13f07825ad1191dea8c72e1ca13f5bfeed1a9282e22330a52fa944512f21ad
manifest.json: rubrics/R1_reordered.md6651ded6e4997cea877c1feb3d5fcf109d47951b55b02bf3e2b791aa577cd5e2
manifest.json: rubrics/R2_compact.md0003a5ce694bf82acf6d6bf6a27f84dff11106d758a17252e71dd41061d46694
result_manifest.json: results/results_R0A.jsona35228a1e4a545bc454906188723b99c0630523e079efcaadb7fedd2f9eee517
result_manifest.json: results/results_R0B.json1ea7075b62f4512dfb1e49022821888032638359abe099849e64245859c1e4fb
result_manifest.json: results/results_R1.json4aab4e01eafca146691862bc29460b46f207d3829d794f4089a433d0318a80ea
result_manifest.json: results/results_R2.json34f6b008cecb3e32876307c621f5c070c74b680a5f060e4d4594c05443addd28
result_manifest.json: analysis_summary.json2f79f1f356889c10de09e76575c7d0c6a4a4e4feed348d2545b8465ac504b3b8
result_manifest.json: analysis_summary.md03fe66c9a38d0fa6fef2295f69c2f3896d30ef49b47c077c8c6baa7a87707884
result_manifest.json: judge_execution_note.md83e46620ec483c52de48648dd776a6fa5773a6da4f48509f0dfee0a94f3be844

Evidence package download

Download the original completed archive for Studies 002 and 003. The source package is unchanged; this reading page is a presentation of its records and explicit limitations.

Download studies 002 + 003

APPLY THE SAME EVIDENTIAL DISCIPLINE

What does your result justify?

Benchmark Verification assesses a single reported result or model-comparison claim against the supplied evidence.

Explore £49 Benchmark VerificationExplore the £750 Experimental Audit