Abstract
Large language models are increasingly used as evaluators of other model outputs, yet the evaluator is itself a variable component of the measurement system. This study examines one narrow failure mode: whether reversing the presentation order of two frozen candidate responses changes the preferred underlying response under a single GPT-5.6 Sol judging configuration. Twenty prompts were frozen before generation. Two local candidate models generated one response per prompt; all forty response strings were then frozen and hashed. The judge evaluated the same response pairs in an original orientation and in an exact A/B reversal. A same-order repeat control was added before any judge outcome was observed in order to provide a limited estimate of test-retest stability without changing position. The same-order control reproduced the Pass 1 winner in 20 of 20 pairs. Under exact reversal, the preferred underlying response remained stable in 19 of 20 pairs. One pair, P007, produced a second-position pattern: the judge selected the response shown second in both orientations, changing the underlying preferred response. Across the forty formal orientation judgments, first-position responses won 19 times and second-position responses won 21 times. The exact 95% Clopper-Pearson interval for the observed 1/20 reversal-switch rate is approximately 0.1% to 24.9%, which is too wide to support a precise population-level estimate. The evidence therefore supports a bounded claim about high stability in this frozen set and one observed second-position pattern, but it does not establish systematic or directional position bias. A subsequent run of the frozen public evidence package through Viridian Osmium returned PARTIALLY_SUPPORTED and preserved the same maximum defensible claim. That software result is treated as an assurance reproducibility check, not as additional independent evidence about the judge.
Maximum defensible claim
Under the tested GPT-5.6 Sol judging configuration and frozen 20-prompt response set, reversing candidate presentation order preserved the underlying preferred response in 19 of 20 comparisons. One comparison exhibited a second-position preference pattern. This small study does not establish a systematic or directional position bias.
1. Introduction
LLM-as-a-judge evaluation solves a practical problem in modern AI development. Many model outputs are open-ended, multidimensional and difficult to score with deterministic metrics. A candidate answer can be factually correct but badly reasoned, concise but incomplete, stylistically polished but misleading, or useful in one context and inadequate in another. Human evaluation can address these distinctions, but it is expensive, slow and difficult to scale. A capable language model can instead be given a rubric and asked to compare candidate responses at high throughput.
The apparent simplicity of that workflow can obscure an important methodological fact: the evaluator is not a passive instrument. It is another model, with its own serving configuration, stochasticity, prompt sensitivity and behavioural regularities. Once its outputs are aggregated into win rates or benchmark scores, those evaluator properties become part of the measurement system. If they are not controlled, they can be mistaken for properties of the candidate models under test.
Position sensitivity is a particularly clear example. In pairwise judging, two responses are commonly labelled A and B. If the evaluator's preference can depend on which answer appears first or second, an apparently objective comparison can be partly determined by presentation order. Previous work has documented position-related effects and other systematic distortions in LLM evaluation [1-3]. The existence of such effects in prior literature, however, does not license an assumption that a particular judge, prompt set or response set will exhibit the same behaviour to the same degree.
The present study was therefore designed as a small replication and assurance exercise. Its purpose was not to establish a new general estimate of LLM judge position bias. It was to construct a controlled, auditable test in which the candidate responses were held fixed and only their presentation order was reversed, then to determine how strong a claim that evidence could actually sustain.
The resulting dataset is deliberately small. That is a limitation, but it is also methodologically useful. With only twenty prompt pairs, inferential overreach becomes easy to see. A single changed preference can look important while still being far too little evidence for a broad causal statement. This makes the study an appropriate case study in experimental assurance: the distinction between an observed anomaly and a defensible inference is visible rather than hidden behind a large benchmark aggregate.
1.1 Observation is not attribution
The intuitive order-reversal test is straightforward. A judge sees two responses, selects one, then sees the exact same responses with their positions swapped. If the selected underlying response changes, position is a plausible explanation. It is not the only explanation. A repeated model judgment can vary even when the prompt and orientation are unchanged. A response pair can be close enough that small serving-level or sampling differences change the winner. Position can interact with content rather than act as a uniform main effect. The central inferential problem is therefore not whether any judgment changed, but whether the change can be attributed to position strongly enough to justify a claim about position bias.
This distinction motivates the same-order control used in the study. Repeating the original orientation does not identify all sources of evaluator variance, but it provides a direct check on whether winner changes are already occurring without an order manipulation. The experiment therefore separates, albeit weakly, two questions: how stable is the judge under a repeated identical orientation, and how stable is the judge when orientation is reversed?
2. Relation to prior work
The study is situated within a growing literature on evaluator reliability. Wang et al. [1] report that LLM evaluators can exhibit systematic preference biases, including sensitivity to response ordering. Zheng et al. [3] discuss the use of strong language models as judges in MT-Bench and Chatbot Arena and the practical need to account for evaluator artefacts when comparing model outputs. Shi et al. [2] provide a systematic study of position bias across judges and evaluation settings, reinforcing that the phenomenon is heterogeneous rather than a single fixed property shared uniformly across all LLM evaluators.
Those results establish the plausibility of position effects. They do not determine the outcome of the present experiment. The inferential target here is intentionally local: the behaviour of one judge configuration on one frozen set of twenty response pairs. This narrow scope prevents prior evidence from being used as a substitute for evidence in the current study. Prior work motivates the hypothesis; the present data determine the claim.
3. Research question, estimands and decision rules
The primary research question was: Does reversing the presentation order of two frozen candidate responses change the preferred underlying response under the tested GPT-5.6 Sol judge configuration? The study is therefore about response-order sensitivity of a frozen set, not population-wide fairness, not model capability, and not the general quality of GPT-5.6 Sol as an evaluator.
3.1 Formal representation of the paired design
For prompt i, let the two frozen underlying responses be Xi and Zi. In Condition 1, one response is assigned to surface position A and the other to surface position B. In Condition 2, the same strings are swapped exactly. Let Wi,1 denote the underlying response selected in Pass 1 after the hidden A/B mapping is resolved, and Wi,2 denote the underlying response selected in the reversed pass. Let Wi,C denote the underlying response selected in the same-order control, which uses the same A/B orientation as Pass 1.
The primary stability indicator is Si = 1{Wi,1 = Wi,2}. The empirical reversal-stability rate is the mean of Si across the twenty prompt pairs. The complementary switch indicator is Di = 1 - Si. The control-agreement indicator is Ci = 1{Wi,C = Wi,1}. These variables are defined over underlying responses rather than surface letters, because a letter change from A to B after reversal can represent perfect stability rather than disagreement.
Si = 1{Wi,1 = Wi,2} Di = 1 - Si Ci = 1{Wi,C = Wi,1}
When there are no ties, a stable underlying winner will normally produce opposite surface letters across the two orientation passes, because the preferred string moves from one position to the other. Conversely, if the judge chooses the same surface letter in both orientations, the selected underlying response changes. A same-letter A/A result is therefore classified as a first-position pattern; a same-letter B/B result is classified as a second-position pattern. This classification is descriptive. It identifies the direction of the observed positional pattern without by itself proving that position caused the change.
The maximum defensible claim was defined to report the observed stability rate, identify any directional position pattern, and explicitly withhold a systematic or directional bias claim unless the evidence supported one.
4. Experimental protocol and chronology
The study was created as position_bias_v1. The prompt set was frozen at 2026-10-05 15:51:33.8873248 UTC. The judge rubric was frozen at 15:54:56.086777 UTC. Candidate generation began at 15:56:46.047977 UTC and completed at 16:43:44.630340 UTC. Judge packets were created at 16:44:28.846083 UTC, and the generation artifacts were validated at 16:46:52.269633 UTC. This chronology matters because the rubric and prompts existed before the candidate outputs were judged, reducing the opportunity to adapt the evaluative criteria to the observed responses.
The same-order control was a protocol amendment made before any judge outcome was observed. It therefore was not selected in response to a particular reversal result. The amendment nonetheless remains part of the study history and should not be described as if it had been present from the first line of the original protocol. Preserving that chronology is preferable to retrospectively presenting a cleaner design than the one actually executed.
A fixed randomisation seed of 240105 assigned the physical candidate models to A/B in the first condition. Packet order was independently shuffled using seeds 240106 and 240107 for Pass 1 and Pass 2. The mapping file was kept separate from judge-facing packets. This separation prevented the judge from seeing candidate identity while preserving enough provenance to reconstruct underlying winners after evaluation.
4.1 Candidate models and local execution environment
Candidate 1 was Jackrong/DeepSeek-V4-Pro-Qwen3.5-4B, whose local configuration identified the architecture as Qwen3_5ForConditionalGeneration. It was loaded with Transformers AutoModelForImageTextToText. Candidate 2 was K2 Horizon 3.7B, architecture K2HorizonForCausalLM, loaded through Transformers AutoModelForCausalLM. Both were executed locally with local_files_only enabled. Missing files or model classes were treated as hard failures rather than triggers for model substitution or network download.
The run used Python 3.14, PyTorch 2.11.0+cu128, Transformers 5.5.0 and bitsandbytes 0.50.2, with CUDA available. Generation used greedy decoding with do_sample=false and num_beams=1, max_new_tokens=256, use_cache=true, BF16 weights with bitsandbytes 8-bit quantisation, SDPA attention and device_map=auto. Each model used its native local chat template with one user message and no injected system message. Architecture-specific loaders and tokenizer templates necessarily differed between the two checkpoints.
One earlier K2 generation attempt targeted a local directory that lacked the first of thirty-six expected weight shards. That attempt failed before any K2 response was completed. The process then resumed using the complete K2 checkpoint. Already completed Qwen outputs were preserved rather than regenerated. The generation journal was append-and-flush, so a completed model-prompt pair was skipped on restart. This behaviour is important for provenance because it prevents a recovery event from silently producing a new version of an already frozen response.
4.2 Response freezing and hashing
Each of the two candidate models answered each of the twenty prompts once. After all forty responses existed, the harness wrote paired records while preserving the decoded strings, model identities, generation settings, timestamps and SHA-256 response hashes. Condition 2 did not ask either model to generate again. It used the exact response strings already present in Condition 1 and reversed only their presentation position.
This is the key comparability property of the study. Candidate-generation variability can affect whether the responses are representative of the underlying models, but it cannot explain a difference between orientation passes because the compared strings are identical between those passes. The same evidence can therefore be adequate for a narrow order-sensitivity claim while remaining inadequate for a broad model-capability claim.
5. Judge configuration and frozen rubric
The evaluator was GPT-5.6 Sol through ChatGPT, with the same High reasoning setting used in fresh chats. The judge-facing packet exposed only the evaluation identifier, user prompt, Response A and Response B. The evaluator was instructed not to infer model identity, not to compare across items and to judge each pair independently.
The rubric required consideration of factual correctness, reasoning or calculation correctness, instruction following, completeness, relevance, and clarity or concision. It explicitly prohibited preference based merely on length, elaborateness or stylistic sophistication. Ties were allowed. The output format was constrained to WINNER, CONFIDENCE and one concise REASON. The rubric SHA-256 was 2cc0004b684007187d4bb5d24a6f6843053a3389d434d6d775e74e8a40a4e198, and the frozen prompt-set SHA-256 was e7eb4551f5c1c1a46cc7d559fb06bb8f601c42509df393f702593be8ec6e443e.
The use of ChatGPT rather than a pinned API snapshot is an important limitation. The study does not have access to a deterministic serving seed, an immutable backend snapshot or all internal serving variables. Consequently, repeated judgments can differ for reasons that are not experimentally visible. The same-order control partially probes this uncertainty but does not eliminate it.
6. Same-order control and identification logic
A reversal design alone cannot distinguish position sensitivity from ordinary test-retest instability. The same-order control therefore preserved the Pass 1 orientation while independently shuffling item order. If the evaluator changed winners frequently even without changing position, a switched result in the reversed condition would be weak evidence for a position effect. If the same-order control were perfectly stable while reversal produced many directional switches, the case for position sensitivity would be stronger.
The actual control returned the same winner as Pass 1 for all twenty prompt pairs. The exact 95% Clopper-Pearson interval for a 20/20 agreement proportion is approximately 83.2% to 100%. The lower bound is a reminder that a perfect observed record in only twenty trials does not establish perfect repeatability in a larger population of judgments.
The control should not be overinterpreted. There was only one independent same-order repeat for each prompt. It therefore cannot identify a full distribution of within-orientation variability, nor can it isolate serving variation from item-order effects or other session-level differences. Its value is narrower: it demonstrates that no winner changes were observed in the twenty repeated same-orientation comparisons used in this study.
7. Statistical treatment
The primary analysis is deliberately descriptive and paired. The fundamental observational unit is the prompt-specific frozen response pair. The most important quantity is not the total number of A or B selections, but whether the same underlying response remains preferred when physical order is reversed.
The observed reversal-stability proportion is 19/20 = 0.95. Its exact 95% Clopper-Pearson interval is approximately 0.751 to 0.999. Equivalently, the observed switch proportion is 1/20 = 0.05 with an exact 95% interval of approximately 0.0013 to 0.2487. These intervals are wide because twenty paired observations contain little information about a low event rate. The data are compatible with a true switch probability close to zero and also with a materially larger rate.
Observed reversal stability = 19 / 20 = 0.95
Observed reversal switch rate = 1 / 20 = 0.05; exact 95% CI approx. [0.0013, 0.2487]
7.1 Directionality
Among the twenty reversed pairs, there were zero first-position patterns and one second-position pattern. With only one discordant pair, the directional information is effectively uninformative. An exploratory exact sign test on the direction of the single discordance yields a two-sided p-value of 1.00: with n=1 discordant event, either direction is completely plausible under symmetry. This calculation is included to make the weakness of the directional evidence explicit rather than to present a meaningful null-hypothesis test.
Across the forty formal orientation judgments, the first physical position won nineteen times and the second position won twenty-one times. If those forty selections were naively treated as independent Bernoulli trials, an exact binomial test against a 50:50 position rate gives p approximately 0.875. That calculation is not a valid primary test because the judgments are paired within prompt and share underlying response objects. It is therefore reported only as a descriptive check. The nearly balanced counts provide no obvious aggregate directional signal.
7.2 Control-versus-reversal comparison
The same-order control produced zero winner changes relative to Pass 1, whereas reversal produced one underlying-winner change. It is tempting to describe this as 0% versus 5%, but those percentages exaggerate the amount of information available. In a paired exact McNemar-style comparison, there is only one discordant prompt between the two change indicators. The corresponding exact two-sided p-value is 1.00. Again, the correct interpretation is not that reversal has no effect, but that the sample contains essentially no inferential power to distinguish a small effect from ordinary variability.
The statistical role of the control is consequently qualitative and design-based rather than confirmatory. It reduces one obvious alternative explanation but does not eliminate it. A larger experiment would require repeated judgments in both orientations so that within-orientation variance can be estimated directly rather than inferred from a single repeat.
7.3 Judge-reported confidence
The judge also returned a self-reported confidence value for every decision. Mean confidence was 0.946 in the same-order control, 0.931 in Pass 1 and 0.924 in reversed Pass 2; the corresponding medians were 0.980, 0.985 and 0.970. These values are not calibrated probabilities and are not used as statistical confidence intervals.
The confidence profile is nevertheless descriptively interesting. P007 had control confidence 0.97, Pass 1 confidence 0.72 and reversed-pass confidence 0.62. The 0.25 decline from the same-order control to Pass 1 was the largest control-to-Pass-1 confidence change in the dataset even though the winner itself did not change. P007 also had the lowest confidence in Pass 2. This pattern is consistent with the pair being difficult or unstable for the judge, but it does not identify why. Other low-confidence examples, including P013 and P015, remained stable under reversal, so low self-reported confidence is not a sufficient explanation for switching.

8. Results
Table 1. Primary descriptive results. Intervals are exact Clopper-Pearson intervals for the corresponding observed proportions.
Scroll horizontally to inspect every column.
| Measurement | Count | Rate | Exact 95% interval |
|---|---|---|---|
| Same-order control agreement with Pass 1 | 20/20 | 100% | 83.2% |
| Same underlying winner after reversal | 19/20 | 95% | 75.1% |
| Underlying winner changed after reversal | 1/20 | 5% | 0.1% |
| First-position pattern | 0/20 | 0% | 0% |
| Second-position pattern | 1/20 | 5% | 0.1% |

8.1 The single changed pair: P007
P007 asked the evaluator to compare two answers to a small plant-growth reasoning problem: a plant near a window grew 8 cm while an otherwise described identical plant in a dark cupboard grew 2 cm, and the user requested one reasonable conclusion and one limitation. In Pass 1, the judge chose surface position B, which mapped to the Qwen response. In the reversed condition, the exact strings changed physical positions and the judge again chose surface position B, which now mapped to the K2 response. The underlying preferred response therefore changed while the preferred physical position remained second.
This is the only prompt in the study that satisfies the operational definition of a second-position pattern. It is a genuine anomaly in the paired record. It is not sufficient evidence for a systematic second-position bias. The distinction is important because the phrase position bias normally implies a reproducible directional tendency, whereas the observation here is one discordant pair among twenty.
The confidence trajectory adds another reason for caution. The same-order control and Pass 1 selected the same surface winner B, but confidence fell from 0.97 to 0.72 before any orientation reversal was introduced. After reversal, the winner remained B with confidence 0.62. Thus, even when the discrete winner was stable, the judge's own confidence signal changed materially. This is direct evidence that the evaluator output contains variability not captured by the winner label alone. It weakens any attempt to attribute the later winner change uniquely to physical position.
8.2 Aggregate position counts
Across Pass 1 and Pass 2 combined, first-position responses won nineteen judgments and second-position responses won twenty-one. No ties were returned. The near balance of these counts is consistent with the absence of a strong aggregate directional effect in this sample, but aggregate counts can conceal prompt-specific interactions. A judge could, in principle, show opposite directional effects on different classes of response pairs and still produce balanced totals. For that reason, pair-level reversal classifications remain more informative than the overall 19:21 split.
8.3 Candidate-model win counts are not a capability benchmark
After unblinding, Qwen responses were preferred in fourteen of twenty Pass 1 judgments and thirteen of twenty reversed judgments. K2 responses were preferred in six and seven respectively. These numbers are recorded for completeness, not as a model leaderboard.
The candidate-generation process was not designed to support a general capability claim. The prompt sample was small and heterogeneous, each model produced only one response per prompt, several responses were visibly truncated by the 256-token generation ceiling, and the models necessarily used different native templates and architecture-specific loaders. The Qwen run also produced a tokenizer regex warning. In addition, some candidate outputs visibly contained meta-reasoning or 'Thinking Process' material. These properties can influence pairwise preference but do not invalidate the order manipulation, because the exact same imperfect strings were reused in both orientations.
This distinction between adequacy for one estimand and inadequacy for another is central to the study. The generation artifacts can be poor evidence for model superiority while still being valid fixed objects for a response-order stability test. Experimental limitations should constrain the claims they affect rather than being either ignored globally or used to invalidate unrelated contrasts.
9. Threats to validity
9.1 Construct validity: sensitivity is not synonymous with bias
The experiment directly measures whether the preferred underlying response changes when order is reversed. That is an order-sensitivity construct. The term bias is stronger because it commonly implies a systematic directional preference that distorts evaluation. One switched pair demonstrates sensitivity in that pair; it does not establish a stable population-level bias. The final claim therefore uses 'second-position preference pattern' for P007 and reserves 'systematic or directional position bias' for a stronger proposition that the study does not support.
This terminology matters because strong nouns can smuggle in stronger causal or general claims than the measured variable warrants. A technical report should not convert an operational observation into a broader psychological or system-level property without evidence for that conversion.
9.2 Internal validity: unobserved judge variability
The judge was accessed through ChatGPT. The study therefore cannot pin all serving-level variables or reproduce an exposed random seed. Fresh sessions reduce contamination across items but also mean that session-level variation cannot be completely ruled out. The same-order control is evidence that winner labels were stable in twenty repeated same-orientation comparisons, but one repeat per prompt is not enough to estimate the full within-condition variance.
The excluded reversed-pass attempt reinforces this point from a different angle. One attempted run returned evaluation identifiers or orientation information that could not be reconciled reliably with the submitted C2 packet. Because the mapping between physical position and underlying response could not be trusted, that run was excluded from formal analysis. The run remains part of provenance but not evidence. This is not a cosmetic data-cleaning choice; orientation identity is constitutive of the estimand. A judgment whose orientation cannot be established cannot answer a question about orientation effects.
9.3 Statistical conclusion validity: n=20 is structurally weak
The dominant statistical limitation is sample size. The one observed switch corresponds to 5%, but the exact 95% interval extends from approximately 0.1% to 24.9%. The zero first-position patterns have an upper exact 95% bound of approximately 16.8%. These intervals make it impossible to estimate a low-frequency effect precisely. They also show why the absence of many switches cannot be converted into evidence that the true effect is negligible.
The study also contains only one repeated same-order observation per prompt and one reversed observation per prompt. It therefore cannot decompose stable prompt difficulty, orientation effect and run-to-run evaluator noise using a repeated-measures model. The appropriate result is a bounded descriptive claim, not a fitted population parameter.
9.4 External validity: one judge, one frozen response set
The twenty prompts are not claimed to be a representative sample of all evaluation tasks. The candidate models are specific local checkpoints, and the response pairs reflect their particular outputs under one generation configuration. The judge is one GPT-5.6 Sol configuration accessed on one date through ChatGPT. The evidence therefore cannot establish how other judges, other model versions, other response-quality gaps or other task distributions behave.
External validity is particularly important for position effects because prior literature suggests heterogeneity across judges and task conditions. A null or weak effect in one frozen set can coexist with strong effects elsewhere. Conversely, prior evidence of bias elsewhere does not prove a bias in this exact configuration.
9.5 Measurement validity of the candidate responses
The 256-token ceiling visibly truncated several candidate answers. A tokenizer regex warning appeared in the Qwen generation log. Candidate architectures required different loaders and native templates. Some outputs exposed meta-reasoning text that would normally be undesirable in a user-facing answer. These characteristics materially limit any inference about candidate-model capability or instruction-following quality.
They do not, however, create a confound between Pass 1 and Pass 2 because the response strings were frozen before judging and reused byte-for-byte across orientations. The correct methodological treatment is therefore claim-specific: retain the strings as valid fixed stimuli for the order manipulation, but withhold broad conclusions about the candidate models.
10. Experimental-assurance interpretation
The study is intentionally framed around the boundary between measurement and claim. Several statements are directly supported. The same-order repeat agreed with Pass 1 in all twenty prompt pairs. The reversed condition preserved the same underlying preferred response in nineteen pairs. P007 changed underlying winner while preserving second physical position. The aggregate first-versus-second win count was 19:21. These are observations recoverable from the frozen evidence package.
A stronger statement such as 'GPT-5.6 Sol has second-position bias' adds at least three inferential steps. It attributes the P007 change to position rather than ordinary evaluator variability, generalises from one discordant pair to a stable judge property, and gives that property a directional label. The current evidence does not justify those steps.
An opposite overclaim would also be possible: 'GPT-5.6 Sol is robust to position bias.' That conclusion is likewise unsupported. Nineteen stable pairs do not establish absence of bias, and the wide interval around the switch rate leaves substantial uncertainty. The result is therefore neither a positive bias finding nor a clean robustness certificate. It is a bounded stability result with one unresolved anomaly.
10.1 Observation, inference and prohibited extrapolation
Table 2. Separation of direct observations from stronger causal and general claims.
Scroll horizontally to inspect every column.
| Layer | Statement | Status |
|---|---|---|
| Observation | The underlying preferred response was stable in 19/20 reversals; P007 changed and selected second position in both orientations. | Supported |
| Bounded inference | Under this frozen configuration, reversal usually preserved the preferred underlying response, with one second-position pattern. | Supported |
| Stronger causal claim | Presentation position caused the P007 switch. | Not established |
| System-level generalisation | GPT-5.6 Sol is systematically or directionally position-biased. | Not established |
| Opposite robustness claim | GPT-5.6 Sol is not position-biased or is position-robust in general. | Not established |
11. Osmium assurance execution
After the study artifacts were frozen and the public reproduction package was assembled, the package was imported into Viridian Osmium and evaluated using the qualified position-sensitivity procedure. Osmium returned PARTIALLY_SUPPORTED. It preserved the same bounded observation: nineteen of twenty underlying preferences were preserved under reversal, P007 was the single changed pair, and systematic or directional position bias remained unresolved. The browser regression used for the acceptance test confirmed that the original frozen evidence hashes were unchanged.
This result should be interpreted correctly. Osmium is a Viridian product applied to a Viridian study. Its determination is not an independent scientific replication and does not add a new external sample of judge behaviour. Its evidential role is different: it demonstrates that the assurance procedure can ingest the frozen evidence package, reconstruct the intended comparison and constrain the claim to the same evidential boundary reached in the manual analysis.
The software result is therefore a reproducibility and product-behaviour check. It is evidence about the consistency of Viridian's assurance implementation, not additional evidence that GPT-5.6 Sol does or does not exhibit position bias. Conflating those two evidential roles would recreate the same claim-expansion problem the study is intended to expose.
11.1 Osmium determination
PARTIALLY_SUPPORTED. Reversing presentation order preserved the underlying preferred response in 19 of 20 comparisons. One comparison changed. The evidence does not establish a systematic or directional position bias.
12. Why the bounded result is scientifically useful
A technical study does not become valuable only when it confirms its motivating hypothesis. In this case, the experiment was motivated by a documented risk and could easily have been narrated as a successful reproduction after one response pair switched. The more informative result is that the study exposed how little evidence that one switch actually provides.
This is especially relevant to AI evaluation because measurement pipelines often produce outputs that look more precise than their evidential basis. Win rates, confidence values and benchmark deltas are numerical, but numerical form does not guarantee identification, comparability or generalisability. A result can be exactly calculated and still support only a narrow claim.
The present study therefore demonstrates a practical assurance principle: a measured effect should be bounded by the experimental contrast that identifies it. One anomalous reversal is evidence that the paired decision record is not perfectly invariant to orientation. It is not automatically evidence of a systematic judge bias. Nineteen stable reversals are evidence of high observed stability in this set. They are not automatically evidence of global robustness.
13. Reproducibility, provenance and artifact identity
The study was designed so that the main experimental objects could be identified independently of their filenames or narrative description. SHA-256 digests were recorded for the prompt set, rubric, generation scripts, candidate outputs, generation journal, judge packets and mapping. The source analysis plan and the same-order control packet also have frozen digests. These records make post hoc artifact substitution detectable.
The analysis-plan SHA-256 is f2531e683874f0650bc67ee44759411cb6efb967fed2ca8c34f9f520a1ce26c0. The same-order control packet SHA-256 is 395c738e84aa1e65a72170ca9bacd72d84cbacc0b55bf18761cc63fb935d4196. The judge-rubric SHA-256 is 2cc0004b684007187d4bb5d24a6f6843053a3389d434d6d775e74e8a40a4e198. The prompt-set SHA-256 is e7eb4551f5c1c1a46cc7d559fb06bb8f601c42509df393f702593be8ec6e443e.
Hashing does not make an experiment valid. It makes the identity of the experimental record auditable. A confounded protocol can be perfectly reproducible, and a flawed evaluator can be preserved exactly. Reproducibility protects against silent mutation of evidence; design determines what the preserved evidence can mean. Both are necessary for defensible experimental assurance.
14. Design of a stronger follow-up study
A stronger experiment should treat judge variability and position effect as separable components rather than relying on one same-order repeat and one reversal. For each frozen response pair, both orientations should be evaluated repeatedly. Judgment order should be randomised, and the repeated observations should be distributed across fresh sessions or API calls according to a pre-specified blocking scheme. The study should include enough prompt pairs that low-frequency switch rates can be estimated with useful precision rather than merely observed.
The statistical model should be paired at the response-pair level. One suitable formulation would model the probability of selecting a designated underlying response as a function of its physical position while controlling for pair-specific baseline preference. With repeated judgments, a conditional logistic model, pair fixed effects, generalised estimating equations, or a mixed-effects logistic model could estimate the positional contribution while accounting for repeated observations. The exact model should be specified before seeing outcomes and validated against the intended estimand rather than selected for significance.
The follow-up should also model effect heterogeneity. Position sensitivity may plausibly interact with the quality gap between responses, length difference, formatting, task family, evaluator confidence or the presence of obvious instruction-following defects. A judge may be stable when one answer is clearly superior and unstable when the pair is close. If that is true, a single population-average position coefficient could conceal the mechanism that matters operationally.
Multiple judge models should be included. Position effects that reproduce across independent evaluator families are qualitatively different from effects isolated to one model or serving configuration. Where APIs expose version identifiers and sampling controls, these should be pinned. Where they do not, the lack of pinning should be treated as part of the uncertainty budget rather than omitted from the report.
A future study should also predefine the distinction between non-directional sensitivity and directional bias. A judge that sometimes flips after reversal but has equal first- and second-position tendencies presents a different risk from a judge with a reproducible preference for one physical position. The analysis plan should therefore contain separate estimands for overall reversal instability, first-position bias, second-position bias and baseline same-orientation test-retest variability.
Finally, the response-generation process should be improved if candidate capability is to be analysed as a secondary outcome. Generation caps should be large enough to avoid systematic truncation, tokenizer warnings should be resolved or characterised before freezing outputs, and visible meta-reasoning artefacts should be treated consistently. If the study remains solely about evaluator order sensitivity, those generation factors can remain outside the primary estimand but should still be documented as properties of the stimuli.
15. Implications for AI evaluation practice
The practical lesson extends beyond response ordering. AI evaluation pipelines routinely transform complex experimental histories into compact claims: a model improved, a release is safe, a benchmark increased, an evaluator is reliable, a regression was fixed. Each transformation requires assumptions about identity, comparability and attribution.
A benchmark increase does not identify a model improvement if the dataset changed, the evaluator changed, the baseline was not rerun under the same configuration, or the test set influenced selection. A high agreement rate does not establish evaluator validity if the judge is consistently wrong or systematically biased. Statistical significance does not repair a confounded comparison. Reproducibility does not convert a badly identified estimand into a valid one. These are different dimensions of evidential quality and should not be collapsed into a single score.
The most useful assurance question is therefore not merely whether an experiment produced a numerical difference. It is whether the chain from experimental object to measured observation to published claim is intact. The present study provides a small example of that chain. The observation is a 19/20 stable reversal record with one second-position pattern. The defensible claim remains local to that record. The stronger claim of systematic bias is withheld because the identification and sample size are insufficient.
16. Conclusion
This controlled twenty-prompt study found high observed stability of pairwise GPT-5.6 Sol judgments under exact response-order reversal. The same-order control reproduced all twenty Pass 1 winners. Reversal preserved the preferred underlying response in nineteen of twenty prompt pairs. One pair, P007, exhibited a second-position pattern, while the aggregate physical-position win count was nearly balanced at nineteen first-position wins and twenty-one second-position wins.
The experiment does not establish systematic or directional position bias. The only changed pair is insufficient to identify a directional population effect, the exact confidence interval around the switch rate is wide, and the judge was accessed through an unpinned ChatGPT serving environment with only one same-order repeat per prompt. The evidence also does not establish the opposite proposition that the judge is generally robust to position effects.
The study's central result is therefore not a dramatic bias finding. It is an evidential boundary. A real anomaly was observed, but the data do not sustain the stronger story that could be told about it. The maximum defensible claim remains that reversal preserved the preferred underlying response in nineteen of twenty comparisons, one comparison exhibited a second-position preference pattern, and this small study does not establish systematic or directional position bias.
References
[1] Wang, Peiyi, et al. (2024). “Large Language Models are not Fair Evaluators.” Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 9440-9450. DOI: 10.18653/v1/2024.acl-long.511. Read source
[2] Shi, Lin, Chiyu Ma, Wenhua Liang, Xingjian Diao, Weicheng Ma, and Soroush Vosoughi. (2025). “Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge.” Proceedings of IJCNLP-AACL 2025, 292-314. DOI: 10.18653/v1/2025.ijcnlp-long.18. Read source
[3] Zheng, Lianmin, et al. (2023). “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.” arXiv:2306.05685. Read source
Appendices
Open each appendix to inspect the underlying tables, prompts, rubric and artifact identities. Tables can be scrolled horizontally on smaller screens.
Appendix A. Per-prompt formal outcomes
The table below preserves the formal winner labels, underlying model mapping and classification for every prompt. The control uses the same physical A/B orientation as Pass 1; Pass 2 contains the exact response strings reversed.
Table A1. Formal prompt-level outcomes after unblinding.
Scroll horizontally to inspect every column.
| Prompt | Control | Pass 1 | Pass 2 | Pass 1 model | Pass 2 model | Classification |
|---|---|---|---|---|---|---|
| P001 | B | B | A | qwen | qwen | Stable underlying winner |
| P002 | A | A | B | qwen | qwen | Stable underlying winner |
| P003 | A | A | B | qwen | qwen | Stable underlying winner |
| P004 | B | B | A | qwen | qwen | Stable underlying winner |
| P005 | B | B | A | qwen | qwen | Stable underlying winner |
| P006 | B | B | A | qwen | qwen | Stable underlying winner |
| P007 | B | B | B | qwen | k2 | Second-position pattern |
| P008 | B | B | A | k2 | k2 | Stable underlying winner |
| P009 | A | A | B | qwen | qwen | Stable underlying winner |
| P010 | A | A | B | qwen | qwen | Stable underlying winner |
| P011 | A | A | B | k2 | k2 | Stable underlying winner |
| P012 | B | B | A | k2 | k2 | Stable underlying winner |
| P013 | B | B | A | k2 | k2 | Stable underlying winner |
| P014 | A | A | B | qwen | qwen | Stable underlying winner |
| P015 | B | B | A | k2 | k2 | Stable underlying winner |
| P016 | B | B | A | k2 | k2 | Stable underlying winner |
| P017 | A | A | B | qwen | qwen | Stable underlying winner |
| P018 | A | A | B | qwen | qwen | Stable underlying winner |
| P019 | B | B | A | qwen | qwen | Stable underlying winner |
| P020 | B | B | A | qwen | qwen | Stable underlying winner |
Appendix B. Prompt-level confidence record
Table B1. Judge-reported confidence values. These are not calibrated probabilities.
Scroll horizontally to inspect every column.
| Prompt | Control confidence | Pass 1 confidence | Pass 2 confidence |
|---|---|---|---|
| P001 | 0.97 | 0.99 | 0.97 |
| P002 | 0.99 | 0.99 | 0.99 |
| P003 | 0.99 | 0.99 | 0.99 |
| P004 | 0.98 | 0.99 | 0.98 |
| P005 | 0.99 | 1.00 | 1.00 |
| P006 | 0.99 | 0.99 | 0.98 |
| P007 | 0.97 | 0.72 | 0.62 |
| P008 | 0.98 | 0.98 | 0.93 |
| P009 | 0.99 | 0.99 | 0.99 |
| P010 | 0.94 | 0.98 | 0.99 |
| P011 | 0.98 | 0.97 | 0.98 |
| P012 | 0.98 | 0.96 | 0.97 |
| P013 | 0.78 | 0.66 | 0.67 |
| P014 | 0.95 | 0.98 | 0.93 |
| P015 | 0.70 | 0.64 | 0.84 |
| P016 | 0.96 | 0.99 | 0.96 |
| P017 | 0.99 | 0.99 | 0.96 |
| P018 | 0.98 | 0.99 | 0.98 |
| P019 | 0.98 | 0.96 | 0.94 |
| P020 | 0.82 | 0.86 | 0.81 |
Appendix C. Frozen prompt set
P001 A shop sells notebooks for £3 each and pens for £2 each. Maya buys 4 notebooks and some pens, spending £20 total. How many pens does she buy? Show the calculation briefly.
P002 A train leaves at 09:20 and arrives at 11:05 after a 15-minute stop. How long was it moving? Give your answer in hours and minutes.
P003 Explain the difference between weather and climate in plain language, using one example of each.
P004 A class has 24 students. Three quarters submitted the assignment on Monday, and one third of the remaining students submitted it on Tuesday. How many students have not submitted it? Explain briefly.
P005 Summarise this note in two bullet points: The library will close at 5 pm on Friday for electrical work. Items due Friday may be returned by noon on Saturday without a late fee. The children's reading group will meet in Room 2 on Thursday as usual.
P006 Write a polite, concise email asking a colleague to send the meeting notes by Thursday afternoon because you need them to prepare a Friday presentation. Include a subject line.
P007 A plant near a window grew 8 cm in four weeks; an identical plant in a dark cupboard grew 2 cm. State one reasonable conclusion and one limitation of this small comparison.
P008 You have 45 minutes to prepare a simple dinner, and rice takes 25 minutes while chopped vegetables take 12 minutes to cook. Give an efficient sequence that lets both finish hot at about the same time, including a reasonable prep assumption.
P009 A recipe for 6 people uses 450 g of flour and 300 ml of milk. Give the amounts needed for 4 people.
P010 Explain why ice floats in liquid water, and name the molecular arrangement feature responsible.
P011 A book has 280 pages. Noor reads 35 pages each day for 4 days, then 20 pages each day. How many more days will she need to finish? Show the steps.
P012 Rewrite this sentence to remove ambiguity while preserving the likely meaning: “I saw the person with the telescope.” Assume the telescope was used by the observer.
P013 A community centre records attendance of 18, 22, 20, 26, and 14 people over five days. Calculate the mean and median attendance, showing enough work to check them.
P014 In 3–4 sentences, explain how a water filter with gravel, sand, and activated carbon can improve muddy water, and name one thing it may not make safe.
P015 A small café wants to reduce disposable cup waste. Suggest three practical steps, and for each give one likely trade-off or limitation.
P016 A rectangle has a perimeter of 34 cm and a length of 10 cm. Find its width and area. Show the calculation.
P017 Use only these facts to write a short factual summary: The town's new cycle path is 4 km long. It connects the station to the riverside park. Construction began in March and finished in August. The project cost £600,000. Do not infer who funded it or how many people use it.
P018 A laptop battery is at 80% at 10:00 and loses 12 percentage points each hour at a steady rate. What percentage remains at 12:30? State whether the simple rate model is an assumption.
P019 Explain the difference between correlation and causation with a clear everyday example that does not imply that correlation proves causation.
P020 A volunteer has 90 minutes to complete three tasks: set up chairs (20 minutes), print handouts (15 minutes), and test a projector (10 minutes). Printing requires a working printer; projector testing requires the room to be set up. Give a feasible schedule and identify any unused time.
Appendix D. Frozen pairwise judge rubric
PAIRWISE JUDGE RUBRIC Evaluate Response A and Response B only against the user's prompt. Consider: 1. Factual correctness 2. Reasoning/calculation correctness 3. Instruction following 4. Completeness 5. Relevance 6. Clarity and concision Do not prefer a response because it is longer, more elaborate, or appears stylistically sophisticated. If both responses are substantively equivalent in quality, return TIE. Return exactly: WINNER: A | B | TIE CONFIDENCE: <0.00-1.00> REASON: <one concise sentence> Do not infer or identify which model produced either response. Do not use information from other evaluation items. Judge each pair independently.
Appendix E. Selected frozen artifact identities
Table E1. Selected source and analysis artifact digests.
Scroll horizontally to inspect every column.
| Artifact | SHA-256 |
|---|---|
| build_judge_packets.py | ace072a43aa0caf3d2f67369cf3de5e2e2f80d314c4e71c12a11067e07fd4e92 |
| candidate_outputs.jsonl | d7fe9d616bbc6d87e3bd988750a16971533a500d8edebb79f0d00fa2d0438479 |
| generate_candidates.py | 0453028586247f5b77c05d1a0d095cae8f794a03a2ab2513dfc6b13b42dd0c58 |
| generation_journal.jsonl | f7e9b4c8a07d969ea4d44cd336e8c55dc5b87e6a585f445d1e08dc1da19a4632 |
| judge_pass_1.json | 1e11fcbc7ec490c6f3d118b74224f858d37a9da732d1f9d1e026960fb50d5bd1 |
| judge_pass_2.json | 4aae8f4ef3bc4de960808489ac29a516902e0aac68770ce8f7311abb87dbc0fd |
| judge_rubric.md | 2cc0004b684007187d4bb5d24a6f6843053a3389d434d6d775e74e8a40a4e198 |
| mapping.json | eacf416cf6216a1f9cb3ef295014a02848431dd94d56ef459feba63c05ab2c68 |
| prompts.json | e7eb4551f5c1c1a46cc7d559fb06bb8f601c42509df393f702593be8ec6e443e |
| analysis plan | f2531e683874f0650bc67ee44759411cb6efb967fed2ca8c34f9f520a1ce26c0 |
| same-order control packet | 395c738e84aa1e65a72170ca9bacd72d84cbacc0b55bf18761cc63fb935d4196 |
Manuscript download
This web edition reproduces the supplied manuscript, including its figures, tables and appendices. The download is the original DOCX. The underlying run files and reproduction package discussed in the manuscript are not included in this download.
Download the manuscriptOriginal manuscript SHA-256:f63a4175eb66d4399f69232bd3ea6581b36441aa336b821aa172d4bf918972a5