INDEPENDENT EXPERIMENTAL ASSURANCE

Evidence behind
AI decisions.

Viridian examines whether experimental evidence supports claims of model improvement, regression safety, evaluator consistency and release readiness.

Reproducibility. Comparability. Provenance.
A technical boundary between a reported result and a justified claim.

One claim. Defined evidence. Written assurance.

RESULT → EVIDENCE → CLAIM → DECISION

The connection is the subject of assurance.

01 / SERVICES

Start with a claim.
Or with a failure.

Two focused engagements for different questions: does the evidence support a decision, or what is undermining the experiment?

CLAIM ASSURANCE£750

Experimental Audit

“Does this evidence support our claim?”

A bounded review of a model comparison, suspected regression, evaluator change or release claim. The written finding identifies supported conclusions, material limitations and the evidence needed to resolve them.

Claim
One principal claim
Input
One evidence package
Output
One written assurance finding
Request an Experimental Audit

Scope, evidence handling and delivery timing are agreed before work begins. Additional execution or remediation requires a separate scope.

TECHNICAL DIAGNOSIS£150

Diagnostic Review

“What is compromising this experiment?”

A focused review of an evaluation or experiment problem: evaluator drift, inconsistent runs, missing provenance, or a comparison that no longer holds. We examine the failure boundary and identify the next checks needed to investigate it.

Question
One defined technical issue
Scope
Agreed around the available artifacts
Purpose
Diagnosis and targeted next checks
Request a Diagnostic Review

Scope and deliverable agreed before work begins. Useful when the experimental problem needs to be understood before a claim can be assessed.

  1. 01 / DESCRIBE

    State the decision or issue

    Tell us what depends on the result and which artifacts you have. All intake and discussions are asynchronous.

  2. 02 / BOUND

    Agree the evidence and scope

    Define the review boundary, handling requirements, fee and delivery timing by email.

  3. 03 / REVIEW

    Receive the written assessment

    Supported conclusions, unresolved questions and targeted next checks, tied to the evidence reviewed.

What evidence should I prepare?

Scores and per-example results, run logs, model and dataset identities, experiment configuration, evaluator or judge configuration, and the rubric where relevant. Missing artifacts may limit the conclusion; identify them at intake.

How should I handle private evidence?

Describe the case in the intake form and indicate that the artifacts are private. Keep confidential artifacts and credentials out of the form. Agree evidence handling and any local-first requirements by email before sharing them.

02 / WHY EVALUATION NEEDS ASSURANCE

The score can improve.
The comparison can fail.

An evaluation measures performance under a particular set of conditions. A claim of improvement also depends on stable measurement, traceable execution and an explicit scope of inference.

EXPERIMENTAL FAILURECONSEQUENCE FOR THE CLAIM

Evaluator drift

A different judge, rubric, prompt or scoring rule can change what is measured. A score difference may reflect the evaluator rather than the model.

Broken comparability

Dataset versions, sampling, configuration and execution conditions can differ. Baseline and candidate results may no longer answer the same question.

Incomplete provenance

A result without identifiable artifacts and run conditions cannot be reliably reconstructed or connected to the reported experiment.

Overextended inference

An aggregate benchmark does not establish every release claim. Uncertainty, selection and intended use determine how far a result can be taken.

We focus on whether evidence supports claims—not just whether benchmarks look good.

03 / PRODUCTS

Assurance as a
technical discipline.

Osmium and Adamantine address the integrity of experimental evidence and the conditions under which it can support a decision. Their implementation scope and qualification status matter as much as their architecture.

OSMIUMEnterprise enquiries

Osmium

Experimental assurance across repeated AI development cycles.

Viridian’s enterprise assurance product for reproducibility, provenance, regression control and evaluator integrity. Its focus is the evidential basis of model and evaluation decisions across successive experiments.

Measurement identity
Model, dataset, evaluator and configuration identities define the conditions of a result.
Evidence continuity
Provenance and reproducibility connect reported results to the experiments that produced them.
Decision boundary
Regression and release conclusions remain limited to the scope justified by the evidence.

Deployment requirements and the applicable assurance scope are established for the engagement.

Discuss Osmium
ADAMANTINE

Adamantine

Qualification scope, artifact identity and explicit gate policy.

Adamantine extends assurance into the definition of the candidate itself: which runtime surfaces exist, what qualification binds, and which authority supplies the gate policy.

Scope inventory
Source-defined callable surfaces and activation inputs are assessed against a declared capability registry.
Qualification identity
Candidate files, test artifacts, gate mapping, runtime, platform and runner options are bound into the assessment identity.
Gate authority
An externally supplied gate-mapping digest is required; a missing or mismatching pin is refused.

Qualification remains specific to the assessed candidate, runtime and gate policy.

Discuss Adamantine

We do not sell certainty.

We show you exactly what
the evidence can justify.

04 / EVIDENTIAL FOUNDATIONS

A claim has a scope.
Evidence has an identity.

Formal assurance begins with explicit obligations: what must be established before a conclusion is admissible? Empirical conclusions remain bounded by the data, assumptions and implementation behind them.

01 / EVIDENTIAL REASONING

Observation, inference, conclusion

Separate what was measured from what is inferred. Identify the assumptions connecting the result to the claim, and make the limits of that inference explicit.

02 / PROOF-ORIENTED VALIDATION

Explicit validation obligations

Reproducibility, comparability and evaluator identity are conditions to examine. Naming an obligation does not establish that it has been satisfied.

03 / DECISION TRACEABILITY

A reasoned record of the decision

Keep the evidence, limitations and resulting recommendation connected. A finding should be understandable in the context of the experiment it reviewed.

HOW WE WORK

Disciplined boundaries.
Explicit limitations.

Local-first
Discuss local execution and evidence handling during scoping. Private artifacts stay out of the intake form.
Evidence-bound
Conclusions stay within the material reviewed and the conditions under which it was produced.
Reproducibility-oriented
Artifact and environment identities matter alongside reported scores.
Structured refusal
Unsupported claims are refused or qualified. Missing evidence remains a limitation, rather than becoming confidence.

05 / CONSULTANCY PARTNERS

Add a specialist
assurance boundary.

For AI engineering, governance, MLOps and risk consultancies whose client decisions depend on experimental evidence. Bring Viridian into a defined evidence review through a specialist subcontract or client referral.

Keep the client relationship. Agree responsibilities, access and scope before starting. No upfront partnership fee; commercial terms apply when paid client work is generated.

Discuss a consultancy partnership

START WITH THE EVIDENCE

What claim does your
decision depend on?

Describe the claim or technical issue. We’ll establish the right review boundary.